Workbook: Designing for disruption
- Traditional disaster recovery assumes regional, isolated events – modern cyber incidents hit primary, secondary and backup environments simultaneously
- A minimum viable business framework forces executive alignment on which systems must come back first, in what order and within what timeframe
- Hidden concentration risks in cloud regions, vendor dependencies and legacy integrations are often only discovered during an incident
Most organisations have a disaster recovery plan. Fewer have tested whether it would hold up against the kind of disruption that is now commonplace – a cyber event that takes out production, backup and disaster recover (DR) environments at the same time, across every location, with no clean system to fail over to.
Traditional DR thinking was built for regional events. A site goes down, your secondary picks up. But that model assumes a simpler architecture than the hybrid, multi-cloud environments most organisations now operate, and it assumes the backup is still intact when you need it. In practice, attackers routinely target backup and recovery systems before triggering their primary payload.
The Datacom team has put together a workbook that helps teams rethink their resilience posture through a minimum viable business lens – including deep-dive exercises, supporting prompts and a Q&A with Datacom's Associate Director of Strategy, Daniel Bowbyes.
The workbook, ‘Designing for disruption,’ is designed to be your guide to reducing dependency risk and building digital resilience. Below is a snapshot of what it covers.
Download the workbook: 'Designing for disruption'
Starting from minimum viable business, not just disaster recovery
When an entire environment is compromised, the question shifts from "how do I fail over" to "how do I rebuild from zero." A minimum viable business framework forces that conversation before an incident happens – which systems keep the organisation alive, in what order do they come back and how long will it take.
Every organisation's answer is different. For one airline, the priority was payroll – if they couldn't pay pilots and crew, nothing else mattered. For a manufacturer, it was the production line. The value of the exercise is in reaching executive agreement upfront so recovery priorities aren't being debated under pressure.
Where hidden dependencies sit
Many organisations assume they're resilient because they run across multiple cloud providers or have dual data centres. In reality, there are often single points of failure that only surface during an incident.
Cloud regions that are marketed as geographically diverse but physically sit across the road from each other. Low-level services like time synchronisation, DNS and certificate infrastructure that nobody thinks about until authentication stops working. Legacy integrations written years ago by someone who has since left, with no documentation and no clear understanding of how data flows between systems.
How environments are interconnected matters as much as where they sit. Infrastructure that has grown organically across on-premises, private cloud and public cloud without a cohesive platform connecting them means recovery involves reassembling the fabric between systems, not just restoring applications.
Designing cyber recovery that works under pressure
Separating cyber recovery capability from standard backup and DR is becoming essential. Recovery vaults need to be immutable – once data is written, it cannot be altered. Recovery environments need to function as clean rooms, logically and physically separate from production.
Automation matters here too. Recovery processes that rely on individual knowledge and manual steps will buckle in a major incident. Scripted, repeatable runbooks for minimum viable business systems remove the risk of human error when the pressure is highest.
There is also a practical hardware question that many organisations overlook. If compromised infrastructure is quarantined as evidence, the compute and storage needed to rebuild has to come from somewhere – and lead times for some vendors are currently running at six months.
Assess your organisation's resilience posture
The workbook is designed to help your team evaluate where your resilience assumptions may not match your actual recovery capability.
It walks through three core areas – identifying and prioritising your minimum viable business systems, mapping where hidden dependencies and concentration risks sit across your infrastructure, and stress-testing whether you could access the capacity needed to rebuild if your existing environment became unavailable.