Digital Resilience: Protecting Business Operations
An information system outage does not necessarily constitute a major crisis. It becomes one when the organization no longer knows which activities to prioritize, who decides on fallback procedures, how to communicate, or how quickly to restore an acceptable level of service. Digital resilience addresses precisely this issue: maintaining and restoring the digital capabilities essential to business operations, even when planned scenarios do not unfold as expected.
For an organization subject to regulatory requirements, with extended value chains, or providing critical services, the issue goes far beyond data backup and disaster recovery planning. It involves governance, business units, service providers, cybersecurity, crisis management, and business continuity. An effective approach integrates these components rather than treating them as independent measures.
Digital resilience is not limited to the disaster recovery plan
The IT disaster recovery plan, or DRP, remains a key component. It outlines the process for restoring infrastructure, applications, data, and technical services following an outage. However, a technically sound DRP alone does not guarantee the continuity of critical operations.
Consider the case of an application that has been restored within the announced timeframe, even though its identity provider remains unavailable, the business teams do not have workaround procedures, or the restored data does not match the latest validated state. In this scenario, the service is not truly available to users. Technical recovery must therefore be evaluated in terms of operational recovery.
Digital resilience encompasses the ability to absorb a disruption, continue operating in an acceptable manner during the incident, and then restore operations in a controlled manner. It covers cyberattacks, major outages, change management errors, cloud provider downtime, data corruption, and telecommunications failures. It also accounts for combined incidents, such as a ransomware attack followed by a communications crisis and a service provider outage.
The point is not to promise that there will be no disruption. Such a promise would rarely be credible. Rather, the goal is to define what must be preserved at all costs, within what time frames and tolerance levels for degradation, and with what responsibilities and decision-making mechanisms.
Start with critical activities, not just technical assets
The most common challenge is developing a strategy based on an inventory of technologies. While this approach is useful for understanding a company’s digital assets, it can lead to components with very different business impacts being protected with the same level of effort.
A Business Impact Analysis(BIA) provides the necessary starting point. It helps identify critical processes, their dependencies, the consequences of an outage, and the maximum acceptable downtime. Recovery time objectives (RTOs) and recovery point objectives (RPOs) should not be determined solely by IT. They reflect a risk decision driven by the business units and approved by governance.
This analysis highlights dependencies that are often underestimated: directories and identity management, collaboration tools, messaging, interfaces with partners, security equipment, reference data, authorized personnel, and site access. For example, a customer service department may have a backup solution but still be unable to handle requests if the knowledge base or the phone system is unavailable.
The level of detail must remain appropriate. Mapping the entire information system down to the smallest detail can become costly and quickly become obsolete. Conversely, a map that is too general does not allow for the development of realistic strategies. The right level is one that allows for a decision that can be acted upon in a crisis: which functions are priorities, what dependencies affect them, and what level of service must be maintained.
Define an Acceptable Level of Service During a Crisis
Continuity does not always mean operating at 100% of nominal capacity. For certain activities, temporary manual processing, limiting access channels, or prioritizing certain categories of customers are appropriate responses. For others—particularly when regulatory or safety requirements are at stake—no reduction in service will be acceptable beyond a very short period of time.
The fallback plan must be designed, documented, and tested in collaboration with the teams that will implement it. A procedure that relies on an inaccessible spreadsheet, a compromised directory, or approval from someone who is unavailable does not constitute a reliable business continuity solution. The question to ask is not only “How do we resume operations?” but also “How do we operate until full resumption?”
Building a Framework for Digital Resilience
Digital resilience is cross-functional by nature. The IT department, IT security, business unit leaders, risk management, compliance, procurement, legal, and crisis communications each have a role to play. Without clear governance, the organization typically ends up with numerous plans that are poorly coordinated.
Effective governance defines explicit responsibilities before an incident occurs. It specifies who assesses the event, who can trigger a crisis response plan, who oversees service restoration, who approves the return to normal operations, and who handles internal or external communications. These decisions must be prepared in advance, as they cannot be improvised under pressure.
TheISO 22301 framework helps structure this process by linking policy, impact analysis, risk analysis, strategies, plans, exercises, and continuous improvement. It does not replace technical expertise or industry-specific requirements, but it provides a common language and a management framework that make it easier to demonstrate compliance to senior management, auditors, and regulatory authorities.
Service providers must be included within the scope. Outsourcing an application, a cloud platform, or a service center does not transfer responsibility for business continuity to an external party. Contracts, service level agreements, rollback capabilities, alert procedures, and test documentation must be reviewed in light of the business’s actual requirements. A high availability commitment does not necessarily meet a requirement for data recovery or crisis management.
Test what will actually happen
A plan that has not been tested remains a mere hypothesis. Technical recovery tests are essential, but they are not sufficient to verify decision-making processes, cross-team coordination, and the ability of business units to operate under degraded conditions.
The most useful exercises are progressive. A restore test can confirm the integrity of a backup and compliance with a recovery time objective. A tabletop exercise then allows decision-makers to work through trade-offs, communication, and escalation procedures. Finally, a broader crisis simulation reveals the challenges of coordinating with service providers, support functions, and business teams.
The goal is not to catch participants making mistakes. It is to identify gaps between the theoretical framework and actual capabilities: missing information, overlooked dependencies, non-operational emergency access points, unclear roles, or poorly defined decision thresholds. Each exercise must result in a prioritized action plan, with a designated person in charge, a deadline, and a governance follow-up process.
The frequency of testing depends on the level of criticality, the stability of the environment, and the pace of change. An organization that regularly deploys new cloud architectures, interconnections, or SaaS solutions must review its assumptions more often than an organization with a stable environment. Major changes should trigger a reassessment, rather than waiting for the annual review cycle.
Measuring Digital Resilience
Metrics should not be limited to the number of available plans or the backup success rate. While these measures are useful, they do not necessarily reflect the ability to maintain a critical operation.
A relevant metric combines several dimensions: coverage of critical processes by validated strategies, compliance with RTOs and RPOs during tests, the rate of qualified critical dependencies, the response time of the crisis management team, the implementation of corrective actions, and the effective participation of business units in drills. The challenge is to highlight the gaps that truly expose the organization to risk.
Maturity is not measured by the amount of documentation produced. It is recognized by the ability of those in charge to explain priorities, provide evidence of testing, make risk-based decisions, and implement consistent responses when an incident occurs. This ability requires shared skills that are regularly maintained.
Building a common foundation requires professionalizing those involved in business continuity, crisis management, IT, and cybersecurity. Structured training programs and recognized standards provide teams with the methods needed to transform sometimes abstract requirements into operational, auditable, and improvable systems.
The next disruption will likely blur the lines between cyber, IT, and business. The key question for every organization is therefore a practical one: if a critical digital service becomes unavailable tomorrow morning, will teams still be able to serve their customers, meet their obligations, and collaboratively set priorities?
This post is also available in:



