A practical way to translate regulatory expectations into an operational resilience program is to move from regulation → business outcomes → measurable tolerances → evidence-based testing. The biggest mistake organizations make is treating resilience as an extension of business continuity or technology disaster recovery. Regulators generally expect firms to demonstrate that they can continue delivering critical outcomes for customers and markets, even during severe disruptions.
A structured approach looks like this:
| Step | Regulatory expectation | Practical implementation |
|---|
| 1. Identify Important Business Services (IBS) | Define services whose disruption would cause intolerable harm | Create an executive-approved list of customer-facing and regulatory-critical services |
| 2. Map dependencies | Understand end-to-end delivery | Map people, processes, technology, data, facilities, third parties, and upstream/downstream dependencies |
| 3. Set impact tolerances | Define maximum acceptable disruption | Establish measurable tolerances tied to customer and regulatory outcomes rather than system recovery objectives |
| 4. Scenario testing | Demonstrate ability to remain within tolerances | Execute progressively severe scenarios involving multiple simultaneous failures |
| 5. Remediate vulnerabilities | Address weaknesses | Prioritize investment based on tolerance breaches |
| 6. Governance | Senior management accountability | Board reporting, annual review, continuous improvement |
1. Define Important Business Services
Rather than listing business processes, identify the outcomes delivered to customers.
For a large insurer these may include:
- FNOL (First Notice of Loss)
- Claims payment
- Policy issuance
- Policy renewal
- Premium collection
- Broker trading capability
- Customer servicing
- Regulatory reporting
- Investment operations
- Reinsurance settlement
Notice these are services, not departments.
Instead of:
Claims Operations
Use:
Customers receive approved claim payments.
That distinction becomes important when setting tolerances.
2. Map End-to-End Delivery
Good resilience mapping usually captures six dependency layers:
- Business processes
- Applications
- Infrastructure
- Data
- People
- Third parties
For example:
Claims Payment
│
├── Claims Platform
├── Payment Gateway
├── Banking Interface
├── IAM
├── Data Warehouse
├── Fraud Engine
├── Call Centre
├── Claims Handlers
├── Cloud Provider
├── Payment Provider
└── Banking Partner
Many firms stop at application mapping.
Leading programs continue until every dependency supporting delivery is understood.
3. Translate Impact Tolerances
This is where many organizations struggle.
Impact tolerances should answer:
"How much disruption can customers experience before harm becomes unacceptable?"
Not
"How quickly can IT recover?"
Recovery Time Objectives (RTOs) remain technology metrics.
Impact tolerances are business outcome metrics.
Example:
| Service | Poor tolerance | Better tolerance |
|---|
| Claims payment | RTO = 4 hours | No approved claim remains unpaid beyond 24 hours |
| FNOL | System restored in 2 hours | 95% of claims accepted within 30 minutes |
| Policy issuance | System uptime 99.9% | No more than 5% of policies delayed beyond one business day |
| Contact centre | Phone system restored | Customers can report claims through at least one channel within 15 minutes |
Good tolerances are measurable from a customer perspective.
4. Multiple Dimensions of Tolerance
Large insurers increasingly use several dimensions simultaneously.
Time
Maximum duration of disruption.
Example:
Claims service unavailable for no longer than four hours.
Volume
Maximum backlog.
Example:
No more than 2,000 unprocessed claims.
Customer Harm
Example:
- vulnerable customers
- catastrophe victims
- medical claims
Example metric:
No vulnerable customer waits longer than six hours.
Financial Harm
Examples:
- delayed settlements
- liquidity exposure
- compensation risk
Regulatory Impact
Examples:
- missed statutory reporting
- sanctions screening failures
- policy cancellation errors
Reputation
Sometimes measured through:
- complaint volumes
- media thresholds
- social media spikes
5. Design Realistic Testing
Regulators increasingly expect severe-but-plausible scenarios rather than tabletop discussions alone.
A maturity model might look like:
Level 1
Tabletop exercises
"What would happen?"
Level 2
Walkthroughs
Teams execute documented response plans.
Level 3
Simulation
Selected systems fail in controlled conditions.
Level 4
Live technical testing
Examples:
- cloud outage
- payment gateway outage
- network partition
- cyber incident
- ransomware
- identity service failure
Level 5
Enterprise scenario
Several failures occur simultaneously.
Example:
Flooding
+
Cloud region unavailable
+
Payment provider outage
+
Call centre staff unavailable
+
Media attention
+
High catastrophe claims volume
These scenarios expose dependency risks that single-failure tests often miss.
6. Use Scenario Libraries
For insurers, useful scenarios include:
Operational
- major office loss
- industrial action
- pandemic resurgence
- loss of key staff
Technology
- ransomware
- Active Directory compromise
- cloud region failure
- database corruption
- network outage
- telecom provider failure
Third Party
- payment processor outage
- outsourced claims administrator failure
- catastrophe modelling vendor outage
- cyber attack on supplier
External
- flood
- wildfire
- earthquake
- severe weather
- market volatility
7. Define Success Criteria
A test should assess whether the organization stays within its impact tolerances.
For example:
| Measure | Target |
|---|
| Claims accepted | >95% within tolerance |
| Payments processed | >98% within 24 hours |
| Customer calls answered | <10 minute wait |
| Regulatory reports | No missed submissions |
| Data loss | Zero |
| Manual workarounds | Sustainable for 5 days |
This shifts testing from "did DR work?" to "did customers receive the service within acceptable limits?"
8. Drive Investment Decisions
Each test should identify:
- Dependency failures
- Single points of failure
- Third-party concentration risks
- Manual process limitations
- Data recovery weaknesses
- Decision bottlenecks
- Governance issues
These findings should feed into a prioritized remediation backlog based on the potential to reduce the risk of exceeding impact tolerances.
9. Establish Governance
An effective operating model typically includes:
- Board: approves important business services, impact tolerances, and resilience strategy.
- Executive resilience committee: oversees testing, remediation, and investment priorities.
- Business service owners: accountable for end-to-end resilience of each important business service.
- Technology and operations teams: manage technical resilience, recovery capabilities, and supporting controls.
- Risk and compliance: provide independent challenge, monitor adherence to the framework, and report on regulatory expectations.
A concise board dashboard might cover the resilience status of each important business service, including:
- Current residual risk (e.g., Red/Amber/Green).
- Latest testing outcome against impact tolerances.
- Number and severity of unresolved vulnerabilities.
- Progress on remediation actions.
- Key third-party dependencies and concentration risks.
- Trends in resilience capability over time.
A practical principle
A useful way to align business and technology is to define impact tolerances using a hierarchy of metrics:
- Business outcome: What customers or counterparties experience (for example, "95% of approved claims paid within 24 hours").
- Operational measure: The internal process performance needed to achieve that outcome (for example, maximum claims backlog or manual processing capacity).
- Technology objective: The supporting recovery targets for critical systems (such as RTOs, RPOs, and service availability).
This creates a clear line of sight from regulatory expectations through customer outcomes to operational capabilities and technical resilience, ensuring that recovery objectives are justified by the level of disruption the business has determined is tolerable rather than being set independently by IT.