Skip to content
bazzi.ai
How do you set a disruption tolerance that is not just a description of current recovery capability?
All answers

How do you set a disruption tolerance that is not just a description of current recovery capability?

Set it to business harm, without reference to what recovery currently achieves. A disruption tolerance states how long the organisation can be without a function before harm becomes unacceptable, which is a board judgement. In practice it is frequently set by asking the technology owner how long recovery takes and writing that number down, which is accurate and entirely circular.

Published

Maximilian Bazzi, Founder and CEO

Why the circular version is so common

A disruption tolerance is meant to answer a business question: how long can the organisation be without a given function before the harm to customers, the market or the firm's safety and soundness becomes unacceptable. That is a board judgement about harm, not a technical statement about recovery.

In practice, tolerances are frequently set the other way round. Someone asks the technology owner how long recovery of the underlying system actually takes, and that number becomes the tolerance. The resulting figure is accurate, in the narrow sense that it describes something real, and it is entirely circular: the organisation has stated that it can tolerate exactly as much disruption as it is already capable of recovering from, which is not a tolerance, it is a description of the status quo wearing a governance label.

The carried fact this section exists for

If a disruption tolerance has never been breached in a severe but plausible test, it was almost certainly written to describe the current state rather than business harm. A tolerance is supposed to produce a gap between what the business can actually survive and what the organisation can currently deliver. That gap is the investment case for closing it. A tolerance that produces no gap, that recovery always comfortably beats, is comfortable to report and useless as a governance tool, because it has never told the board anything it did not already know.

The tolerance gapA horizontal comparison of what the business can survive against what current recovery capability delivers, with the difference between the two labelled as the investment case.What the business can surviveDisruption toleranceWhat recovery currently deliversRecovery capabilityGap= the investment case
The gap between what the business can survive and what recovery currently delivers is the case for investment, stated as a distance rather than a percentage.

Setting it correctly means starting from harm, not recovery

Setting a tolerance to business harm means asking what happens to customers, to regulatory standing, to the market or to safety at increasing durations of disruption to a given service, independently of how long recovery is believed to take today. That conversation belongs with the people who own the consequences, commercial, regulatory, operational, not with the team that will eventually be asked to deliver against whatever number is agreed. Only once the harm-based tolerance is set does it become useful to compare it against actual recovery capability, and the resulting gap, where one exists, is the thing worth prioritising.

The four things clients conflate

Four related but distinct concepts get treated as interchangeable, and the confusion undermines the governance conversation. The maximum tolerable period of disruption is the outer limit past which the disruption becomes unsurvivable for the organisation. The recovery time objective sits inside that limit as the target time to restore full service. The minimum business continuity objective is the reduced level of service the organisation commits to sustaining during the disruption, short of full recovery. Impact tolerance, in the UK and Swiss regulatory sense, is a board-set outer limit specifically for an important business service, and it is a distinct regulatory construct rather than simply another name for the recovery time objective. Treating any two of these as synonyms is where most tolerance-setting conversations lose their way.

Where teams commonly run into difficulty

The most common failure is sequencing: recovery capability gets measured first, and the tolerance is then set to match it, because that produces a number quickly and nobody has to have the harder conversation about what harm the business is actually prepared to accept. The correct sequence is slower and more uncomfortable, and it is the only version that gives the board a genuine decision to make.

How Bazzi Consulting helps

We set disruption tolerances against business harm first, independently of current recovery capability, so the resulting gap is a real investment case rather than a number recovery was always going to beat. See risk and resilience advisory.

Essential cookies keep the site working and cannot be switched off. Analytics is optional.

Always on. Required for the site to function.

Cookieless usage analytics (Vercel Analytics). No cross-site tracking.