Vendor Business Continuity & Resilience
Uptime SLA: 99.9%. BCP Tested: Never Asked. Backup: Same Region. Recovery Time: 11 Days. SLA: 4 Hours.
4 min read · 4 May 2026 · Third-party oversight
Vendor business continuity and resilience assessment is one of the most consequential and most superficially performed dimensions of TPRM. Uptime metrics and SLA commitments describe historical performance under normal operating conditions. They do not describe what happens when abnormal events , data centre fires, regional power outages, network backbone failures, ransomware incidents, or natural disasters , disrupt the vendor's primary infrastructure. Business continuity assessment must go beyond SLA review to evaluate the vendor's actual resilience architecture: where backup infrastructure is located, whether it is genuinely independent of primary infrastructure failure modes, whether the BCP has been tested under realistic failure scenarios, and what the tested recovery time looks like rather than the designed recovery objective.
The geographic colocation problem is the specific resilience failure in the hook scenario. Backup infrastructure that is geographically close to primary infrastructure is subject to the same regional failure modes , power grid disruptions, network exchange failures, natural disasters, and major infrastructure events that affect a geographic region rather than a specific site. True geographic resilience requires backup infrastructure in a different region served by different power infrastructure, different network connectivity, and different physical risk exposure. Twenty miles is not geographic separation in the context of regional resilience.
The tested versus designed recovery distinction is the second critical dimension. Recovery Time Objectives and Recovery Point Objectives describe what a vendor's BCP is designed to achieve. Tested recovery times describe what it actually achieved when tested against realistic failure scenarios. A BCP that has never been tested is a plan that has not been validated , the designed 4-hour RTO may be achievable in a controlled failover drill but may not be achievable in a real failure scenario with unplanned dependencies, cascading failures, and the operational complexity of actual disaster conditions.
Why this matters
Vendor business continuity and resilience matter because operational disruptions at critical vendors directly affect the enterprise's own operations , and the enterprise's ability to maintain service continuity for its own customers during vendor disruptions depends on having accurate information about the vendor's realistic recovery capabilities under failure conditions. SLA commitments describe the vendor's normal performance. BCP testing results describe their failure recovery performance , which is what matters when the failure occurs.
- Uptime SLA accepted as resilience assessment
- BCP existence confirmed without testing validation
- Geographic backup location not assessed for regional independence
- Tested RTO not requested , designed RTO accepted without validation
- Failure scenario scope not assessed , what failure modes the BCP covers
What good looks like
Mature vendor resilience assessments request BCP documentation, testing results from the most recent test, RTO and RPO for specific failure scenarios, geographic location of backup infrastructure, and confirmation that backup infrastructure is served by genuinely independent power, network, and physical infrastructure.
- BCP documentation including tested RTO/RPO
- Geographic independence assessment , backup region served by independent infrastructure
- Testing scope and frequency , which failure scenarios have been tested
- Most recent test results including actual recovery time achieved
- Cascade failure planning , dependencies that might delay recovery beyond designed RTO
Tooling
Business Continuity , Fusion Risk Management, Riskonnect for vendor BCP assessment; AWS/Azure region architecture review
For cloud-hosted vendors, reviewing their cloud architecture for multi-region deployment versus single-region with regional backup provides resilience insight that uptime metrics do not. AWS and Azure publish their regional availability zone architecture , understanding whether a vendor's infrastructure uses multiple availability zones within a region versus genuinely separate regions with independent network connections is a specific resilience assessment question.
Governance challenges
The governance challenge with vendor BCP assessment is the testing evidence availability problem. Vendors who have conducted BCP testing may be reluctant to share test results that reveal gaps , recovery times that exceeded targets, dependencies that caused delays, or infrastructure failures during the test. The governance resolution is framing BCP evidence requests as operational due diligence rather than security audit , what the enterprise needs to plan its own contingencies, not evidence of vendor failure.
- Request BCP with tested RTO/RPO , not just designed objectives
- Assess geographic independence of backup infrastructure
- Ask about last test date and scope , what failure scenarios were tested
- Include BCP test results in critical-tier vendor assessment requirements
- Plan enterprise contingencies based on realistic vendor recovery timelines
If you are a small team
For your three most operationally critical vendors, ask four questions that uptime metrics don't answer: where is your backup infrastructure geographically, and is it served by independent power and network infrastructure? When was your BCP last tested, and what was the actual recovery time achieved , not the designed objective? Have you tested against a scenario where your primary data centre is completely unavailable, not just degraded? Those four questions reveal the resilience architecture behind the SLA commitment.
- Ask for geographic location and infrastructure independence of backup
- Ask for last BCP test date, scope, and actual recovery time achieved
- Ask whether complete primary site failure has been tested
- Plan contingency operations based on realistic recovery timelines
What to require
Ask directly:
"For your most recent BCP test , what failure scenario was tested, what was the actual recovery time achieved, and is your backup infrastructure in a genuinely different geographic region served by independent power and network infrastructure from your primary site?"
Expect as evidence
- Most recent BCP test results with actual RTO achieved
- Geographic architecture documentation for primary and backup infrastructure
- Infrastructure independence confirmation , power and network
- Tested failure scenario scope
A vendor who confirms strong uptime SLA should be asked for tested RTO results. SLA describes normal performance. BCP test results describe failure recovery performance. Both are needed for resilience assessment.
How to evidence it
- BCP documentation with tested RTO
- Geographic independence assessment records
- Testing scope and frequency records
- Contingency planning based on realistic recovery timelines
Key Takeaway
99.9% uptime SLA. Backup: 20 miles away, same region. BCP: never tested, never asked about. Recovery time designed: 4 hours. Recovery time actual: 11 days. The SLA described performance under normal conditions. The resilience architecture revealed by the fire described performance under failure conditions. Geographic independence requires more than physical distance , independent power infrastructure, independent network connectivity, and genuinely separate failure domains. Tested RTO is what the BCP achieves in realistic failure conditions. The SLA and the tested RTO are both needed. Only one was ever asked for.
Speak to It™
The term you nodded along to, explained in ninety seconds, so you can speak to it professionally. It is how most readers find these articles.
Join the Association