RESOLV
← Insights

Financial services

Operational resilience: setting impact tolerances for critical services

Operational resilience asks a different question from business continuity: not whether systems recover, but how much disruption customers and the financial system can bear. This is a practical method for identifying important services, setting impact tolerances and testing against them.

· 5 min read

Banks and payment institutions have long maintained business continuity and disaster recovery plans. These remain necessary, but they tend to be organised around systems and sites: how quickly the core banking platform can fail over, how long the data centre can run on generators. Operational resilience shifts the frame to the services customers depend on and asks how far disruption to those services can go before it causes intolerable harm. The Basel Committee's Principles for Operational Resilience set out this approach for banks, and supervisors in many jurisdictions are adopting its logic. The underlying assumption is uncomfortable but realistic: disruption will happen. The work is to decide in advance how much is tolerable, organise the institution to stay within that limit, and prove through testing that it can. An important business service is one delivered to an external user, a customer or a market participant, whose disruption could cause intolerable harm to those users, threaten the institution's viability or affect the stability of the wider financial system. The emphasis on external users matters. Payroll processing for staff is important, but it is an internal function; making salary payments on behalf of a corporate client is a service.

  • Start from the customer's perspective: cash withdrawal, domestic transfers, mobile money cash-in and cash-out, cross-border remittance receipt, loan disbursement.
  • Describe each service in terms of the outcome the user needs, not the system that delivers it.
  • Keep the list short enough that the board can hold it in mind; a long list dilutes attention.
  • Consider which customers are most vulnerable to each disruption, such as those who depend on remittances for daily needs.

Map the dependencies end to end

For each important service, map everything required to deliver it: people, processes, technology, facilities, data and third parties. This mapping usually reveals dependencies nobody had written down, such as a single staff member who runs a manual reconciliation, a shared network link between two supposedly independent sites, or an outsourced provider whose own subcontractor hosts a critical component.

  • Systems and the infrastructure beneath them, including identity services, network links and databases.
  • Key people and the knowledge they hold that is not documented.
  • Third parties and their own dependencies, including cloud providers, card schemes, mobile network operators and correspondent banks.
  • Data feeds and reference data whose absence would halt processing.
  • Physical sites, power and connectivity, including branches and agent networks.

Set impact tolerances

An impact tolerance is the maximum tolerable level of disruption to an important business service. It is usually expressed as a duration, sometimes combined with volume or value measures, such as the longest period customers can be unable to withdraw cash before the harm becomes intolerable. The tolerance is set from the perspective of harm to users and the market, not from what the current technology can achieve. That distinction is the point: where the institution cannot currently stay within tolerance, the gap becomes a remediation priority. Recovery time objectives for individual systems should then be derived from the service tolerance, not the other way round. If a service must be restored within a given window, every dependency on its critical path must be recoverable well within it, with time left for decision-making and communication.

An impact tolerance describes what customers can bear, not what your systems can currently deliver; the gap between the two is your investment plan.

Test severe scenarios, including third-party failure

Testing should be designed to find weaknesses, not to produce a pass. Choose scenarios that are severe but plausible for the institution's context and run them against the dependency maps.

  • Prolonged loss of a primary data centre or cloud region.
  • A ransomware attack that encrypts production systems and backups reachable from the same network.
  • Failure or withdrawal of a critical third party, including a mobile network operator or payment switch.
  • Extended power or connectivity outages affecting branches and agents across a region.
  • Unavailability of key staff, including those holding critical administrative credentials.

Combine tabletop exercises, which test decisions and communication, with technical tests such as failover and restoration from backups. Record where the service would breach its tolerance and why, and track remediation to closure. Outsourcing a service does not outsource the responsibility for it. Where several important services depend on the same cloud provider, data centre operator or payment processor, a single failure can breach multiple tolerances at once. Identify these concentrations explicitly. For material providers, require evidence of their own resilience testing, contractual rights to information and audit, notification obligations for incidents, and a documented, tested exit or substitution strategy. Consider whether a degraded alternative, such as offline processing with later reconciliation, could keep a service within tolerance while a provider recovers. Include the people who would actually make decisions during an incident, not their deputies, and test the communication plan for customers, agents and the supervisor alongside the technical recovery. Some concentrations are national rather than institutional: where every bank in a market depends on the same switch or network operator, coordinated testing with peers and the supervisor may be the only realistic mitigation.

Board reporting and governance

The board should approve the list of important business services and their impact tolerances, and receive regular reporting on whether the institution can remain within them. Useful reporting is brief and decision-oriented.

  • For each important service: current ability to stay within tolerance, with the evidence from the latest test.
  • Open vulnerabilities, their owners, remediation dates and any slippage.
  • Material third-party and concentration risks, and changes since the last report.
  • Incidents that affected important services, and lessons applied.

Operational resilience is not a separate programme sitting beside risk, technology and continuity management; it is the thread that connects them around the services that matter most. Institutions that adopt it gain a clearer view of where to invest and a more honest conversation with their boards and supervisors about what disruption they can withstand. A sensible first year is modest: agree the list of services, set provisional tolerances, map the dependencies of the most critical service in depth, and run one well-designed scenario test. Each subsequent year refines the tolerances and extends the depth of mapping and testing.

Tell us what can't fail.

A senior engineer reviews every enquiry and replies within one business day.