Why resilience, not availability, is the true measure of a modern cloud strategy
Cloud computing has transformed enterprise technology over the past decade. It has accelerated product development, expanded the geographic reach of financial services, reduced the friction of experimentation, and enabled organisations to consume advanced capabilities—analytics, machine learning, secure storage, high-performance networking—without the delays and capital outlay of building them internally. For most large enterprises, the direction of travel is now settled: the majority of new workloads are designed for the cloud by default, and legacy estates are being progressively migrated, modernised or retired.
Yet the same shift that has delivered these benefits has quietly changed the nature of operational risk. Where once an organisation’s technology risk was distributed across dozens of internal systems, discrete data centres, and independently managed applications, it is now increasingly concentrated in the platforms, services and control planes of a very small number of hyperscale providers. This concentration is not a theoretical concern. It is a structural feature of the market that regulators, boards, auditors and rating agencies are now examining with a level of seriousness that would have seemed disproportionate only three or four years ago.
From July 2026, the major cloud providers that support the United Kingdom’s financial sector are subject to direct regulatory oversight as critical third parties. This is a material development. It moves the accountability conversation beyond the individual financial institution and its contractual arrangements with a provider, and places the providers themselves within a formal supervisory perimeter. The change reflects the settled view of the Bank of England, the Prudential Regulation Authority and the Financial Conduct Authority: that the failure, prolonged outage or compromise of a single dominant cloud platform could now produce systemic consequences comparable to those arising from disruption at a major clearing house or systemically important payments infrastructure.
Similar considerations are emerging across other jurisdictions. The European Union’s Digital Operational Resilience Act is fully in force, with critical ICT third-party providers designated for oversight. The United States has continued to sharpen expectations around third-party risk management, particularly in banking, insurance and market infrastructure. The United Arab Emirates and other Gulf jurisdictions have adopted increasingly prescriptive requirements on outsourcing, data residency and operational continuity for regulated entities. The direction is unambiguous. Cloud concentration risk is being reframed, globally, as a matter of national and sectoral resilience—not simply a procurement or architecture question to be resolved within the technology function.
The changing shape of operational risk
The traditional model of operational risk in large financial institutions was built around the assumption that the organisation itself owned, operated and controlled the majority of its critical technology. Disaster recovery frameworks, business continuity plans and testing regimes reflected that assumption. Recovery time objectives were negotiated between the business and the internal IT function. Failover was demonstrated between primary and secondary data centres. Regulatory examinations focused principally on the institution’s own governance, controls and testing evidence.
The cloud era has inverted much of this. A modern financial institution may run its retail banking platform, its market data feeds, its collaboration tools, its identity services, its analytics estate and its customer contact infrastructure across services provided by two or three hyperscale platforms. Each of those platforms is, in turn, dependent on its own control planes, its own regional infrastructure, its own supply chain of hardware, connectivity and specialist software, and its own operational staff. When a control plane suffers a degradation, or a regional service loses availability, the effect can propagate rapidly across many customers simultaneously and in ways that individual institutions cannot influence.
The result is that resilience has become a shared and layered concept. It depends not only on the choices made by the institution itself, but also on the design, operating discipline and transparency of a small number of providers whose internal workings are largely opaque to their customers. The question boards must now ask is not whether their organisations use the cloud responsibly—almost all do—but whether they truly understand, at a business-service level, the dependencies that their cloud strategy has created and the failure modes they have implicitly accepted.
Availability is not the same as recoverability
A recurring feature of cloud resilience conversations in board and executive committee settings is the confusion between availability and recoverability. These are not the same. Availability is an operational measure: the proportion of time a service is functioning within defined parameters. It is typically expressed as a percentage in a service-level agreement, and reported in monthly or quarterly service reviews. Recoverability, by contrast, is the ability to restore a service to a defined state following a disruption, within an agreed timeframe, under adverse conditions, and with data integrity intact.
It is entirely possible for an organisation to enjoy excellent availability metrics year after year and yet to have limited genuine recoverability in the event of a serious incident. High availability may be delivered through multiple availability zones within a single cloud region. That configuration protects against localised hardware or facility failure, but it does not protect against a regional service outage, a control-plane failure affecting the entire region, an authentication or identity service disruption that renders workloads inaccessible even when they are technically running, or a cyber incident that spreads across the institution’s footprint within a single provider.
Stating that an application is “hosted across multiple availability zones” is therefore not, in isolation, a satisfactory resilience narrative. It is a starting point. The relevant board-level questions are more searching. What happens when the region itself is unavailable? What happens when the control plane through which workloads are provisioned and managed is degraded? What happens when the identity provider on which every human and machine login depends is unreachable? What happens when the provider’s customer support and incident response processes are overwhelmed by the very same event affecting the institution?
Mapping dependencies at the level of important business services
One of the most consequential shifts in the regulatory conversation over recent years has been the move from a technology-centric view of resilience to a business-service view. Under the operational resilience frameworks now embedded in UK, European and other regimes, regulated firms are required to identify their important business services, define impact tolerances for the maximum tolerable disruption to those services, and demonstrate that they can remain within those tolerances under a range of severe but plausible scenarios.
This reframing has significant implications for how cloud concentration risk is understood. A cloud dependency is only meaningful when traced through to the business services that rely on it. An identity service outage might sound modest in technical terms, yet if that identity service underpins customer authentication for online banking, colleague access to the payments platform, and administrative access to the fraud detection system, the aggregate business impact could be severe and rapid. A shared analytics platform, apparently peripheral, may in fact carry the models on which credit decisioning, transaction monitoring and regulatory reporting all depend.
Dependency mapping is therefore not an optional exercise. It is the foundation on which every subsequent resilience decision rests. The mapping must extend from each important business service down through its applications, data flows, integration points, identity dependencies, security controls and third-party services—and, crucially, down to the specific cloud services, regions and control planes on which each of those depends. Only when that view is established can leaders form a defensible opinion on where concentration is acceptable, where it is not, and where mitigation is required.
The hidden concentration points
Discussion of cloud concentration risk often focuses on the compute and storage layer. This is understandable, because it is the most visible and the most straightforward to reason about. However, some of the most serious concentration points in a modern estate sit above or beneath the compute layer, and they are frequently under-analysed.
Identity and access management is a prime example. Whether provided by the hyperscaler itself, by a specialist identity vendor built on that hyperscaler’s infrastructure, or by a federated arrangement across multiple providers, identity is the single control plane through which people and workloads authenticate and authorise. A meaningful disruption to identity services can bring an entire estate to a halt even when every other component is functioning. Yet identity is often treated as a utility rather than as a critical service in its own right.
Domain name resolution, certificate management, secrets management, container registries, continuous integration and deployment pipelines, observability platforms and messaging backbones are similar in character. Each is comparatively invisible in day-to-day operations. Each is capable of producing a wide-blast-radius outage if it fails. And each is often supplied by a small number of vendors—sometimes a single vendor—embedded so deeply in the operating model that migrating away would be a multi-year programme.
Data itself represents another concentration point that deserves specific attention. Regulatory obligations around retention, residency, sovereignty and privacy vary by jurisdiction, and the practical portability of very large data estates between providers is significantly harder than the portability of the applications that consume them. An organisation that has consolidated its analytics, its data lake and its regulatory reporting warehouse onto a single provider may find that its ability to switch providers is constrained not by contract but by physics and by the very real cost and time required to move petabyte-scale data sets while maintaining integrity, lineage and compliance.
Multi-cloud is not automatically the answer
Faced with these concerns, some boards and executive teams reach quickly for multi-cloud as a solution. The intuition is understandable: if concentration is the problem, distribution across multiple providers must be the answer. In practice, multi-cloud is a nuanced strategy that must be adopted deliberately, for specific reasons, and with a full appreciation of its costs and second-order consequences.
Undisciplined multi-cloud can increase, rather than reduce, operational risk. Each additional provider brings its own control plane, its own security model, its own identity system, its own operational tooling, its own commercial construct, and its own supply chain. The organisation must invest in the skills, the automation and the governance required to operate all of them to a consistently high standard. Where those investments are inadequate, the organisation ends up with the aggregate weaknesses of every provider it uses, without the resilience benefits it hoped to achieve. Costs rise. Complexity multiplies. Incident response becomes more difficult, not less, because engineers must reason simultaneously about several different environments during a crisis.
A disciplined multi-cloud strategy is different. It begins from the resilience requirements of specific important business services. It identifies where the failure modes of a single provider are unacceptable and where, therefore, active-active or active-passive distribution across providers is warranted. It accepts that other services can safely remain concentrated with a primary provider, provided that a credible exit route and a tested recovery plan exist. It invests in abstractions—portable data formats, standardised interfaces, provider-agnostic tooling where practicable—so that the strategic optionality is real rather than notional. And it recognises that some services, particularly those that depend on differentiated capabilities of a single provider, cannot be made portable without significant redesign, and accepts that constraint openly.
Exit planning as a strategic capability
Regulatory expectations around exit planning have hardened considerably. It is no longer sufficient for a regulated firm to hold a contractual right of termination and a general statement that services could, in principle, be migrated elsewhere. Supervisors expect a documented, tested and credible exit strategy for material outsourcing arrangements, including cloud arrangements. That strategy must address not only technical migration but also data extraction, contractual continuity, customer communication, regulatory notification, staff redeployment and the operational management of the transition period itself.
Exit planning is often uncomfortable, because it forces the organisation to confront the true cost of its dependency. In many cases, the honest answer is that a full exit from a strategic cloud provider would take multiple years, would carry material transformation costs, and would introduce transitional risks of its own. Regulators understand this. What they seek is not the fiction of an easy exit, but a realistic plan, tested at meaningful scale, that demonstrates the organisation has thought through the practicalities and could execute the plan under pressure if required.
The most mature organisations treat exit planning as a design constraint rather than a compliance artefact. They ask, at the point of architecture decisions, how a given service could be migrated away from its provider if circumstances required. They invest selectively in portability where it is proportionate. They document their exit assumptions honestly, review them regularly, and exercise the most critical elements as part of their broader resilience testing programme. In doing so, they convert what might otherwise be a bureaucratic obligation into a genuine strategic capability.
Realistic testing under severe but plausible scenarios
A resilience strategy that has not been tested under realistic conditions is a hypothesis, not a capability. This is one of the reasons that regulators place such emphasis on scenario testing, and why they increasingly expect that testing to reflect severe but plausible conditions—including, in some cases, the assumption that a critical third-party provider is unavailable for a sustained period.
Testing cloud resilience is significantly more difficult than testing traditional data-centre disaster recovery. In a traditional environment, an institution could take a defined failover action on a scheduled date and observe the outcome. In the cloud, the equivalent tests may require coordination with the provider, may involve services whose failure modes cannot easily be reproduced, and may be difficult to conduct without risk to live customer services. Some organisations therefore fall back on tabletop exercises. Tabletop exercises are useful and should be part of the programme, but they are not a substitute for genuine technical testing.
Leading practice combines several forms of testing. It includes controlled fault injection, in which specific components are deliberately degraded or removed within a safe environment to validate that dependent services behave as designed. It includes region-level exercises in which failover to an alternative region is performed under load. It includes third-party dependency simulations, in which the effect of a provider outage is modelled at a business-service level, including the human and process implications. And it includes joint exercises with critical providers and, where relevant, with regulators and industry peers, to validate that the collective response to a systemic event would be adequate.
Contractual commitments and technical reality
One of the recurring findings of resilience reviews across the financial sector is a gap between what contracts commit and what technical arrangements can actually deliver. A contract may guarantee a recovery time objective, a data residency assurance, an audit right or a security standard. Whether the organisation could, in practice, invoke that contractual right and receive the promised outcome under real incident conditions is a separate question—and one that is often not adequately tested.
Boards should expect their executive teams to be able to demonstrate that contractual commitments in relation to material cloud arrangements are supported by tested technical capabilities. Recovery time objectives should be validated through actual recovery exercises, not assumed on the basis of provider assertions. Data residency claims should be verified against actual data flows and storage locations, not solely against contractual clauses. Audit rights should be exercised meaningfully, not merely held in reserve. Security assurances should be substantiated through independent evidence, including relevant certifications, penetration testing outcomes and, where appropriate, direct examination.
Governance: from procurement to boardroom
Historically, cloud arrangements were governed principally through procurement and vendor management functions, with technology architecture teams providing input on design. This model is no longer sufficient. The scale, criticality and systemic implications of modern cloud arrangements demand governance that extends into the boardroom itself. Boards do not need to become technical experts. They do need to be able to interrogate the resilience of their organisations’ cloud strategies with the same rigour they apply to capital adequacy, liquidity, credit exposure and conduct risk.
Effective governance in this area typically requires several elements. A clear articulation of the important business services and their impact tolerances, expressed in language that non-technical directors can engage with. A transparent map of critical cloud dependencies at the level of those business services. A defined risk appetite for concentration, with thresholds that trigger explicit board or committee decisions. Regular, meaningful reporting on resilience testing outcomes, including honest disclosure of failed tests and remediation status. A defined escalation path for material incidents affecting critical providers. And a schedule for periodic review of exit strategies, provider performance and evolving regulatory expectations.
These elements should be reinforced by a senior executive owner accountable for cloud resilience across the enterprise, with sufficient authority to convene the relevant technology, risk, compliance, legal, procurement and business leaders. In many organisations, the natural home for this accountability is the Chief Information Officer or Chief Technology Officer, working closely with the Chief Risk Officer and, where the organisation has one, the Chief Operational Resilience Officer. What matters is not the title but the clarity of the accountability, the seniority of the individual holding it, and the visibility of the role to the board.
Implications for the wider ecosystem
The regulatory reframing of cloud providers as critical third parties has implications that extend beyond the largest financial institutions. Managed service providers, systems integrators, financial technology firms and regulated infrastructures all sit within the same broader ecosystem, and they will all feel the effects of the new supervisory regime.
Managed service providers who deliver services on behalf of regulated firms should expect increasing scrutiny of their own cloud dependencies, their operational controls, and their ability to support their clients’ resilience obligations. Financial technology firms that partner with regulated institutions should expect due diligence conversations to become more searching, particularly where their services touch important business services. Cloud providers themselves will need to invest further in transparency, in customer-facing resilience tooling, in incident communications and in the operational disciplines required to operate under formal supervision.
For technology leaders in the wider corporate and public sectors, the direction of travel in financial services is instructive rather than confined. Concentration risk is not a phenomenon unique to banks and insurers. Utilities, telecommunications operators, healthcare providers, sovereign entities and large family businesses all face structurally similar exposures, and are likely to see analogous regulatory and stakeholder expectations emerge over the coming years. Boards outside the financial sector would be well advised to anticipate those expectations rather than await them.
A closing perspective for boards and executive teams
Cloud resilience is no longer solely an infrastructure discussion. It has become a business continuity discussion, a vendor governance discussion, a regulatory confidence discussion and, ultimately, a strategic discussion about the durability of the enterprise itself. The organisations that will navigate the coming period most effectively are those that treat cloud concentration risk as a first-order board-level issue and invest accordingly in the governance, architecture, testing and cultural disciplines required to manage it.
The essential question every executive team should be prepared to answer is straightforward, even if the underlying work is not. Does your cloud strategy demonstrate recoverability, or does it merely describe availability? Can you evidence, at the level of your important business services, that you understand your dependencies, that your contractual protections are matched by tested technical capabilities, that your exit plans are credible, and that your governance can respond decisively when a critical provider is disrupted?
For those able to answer those questions with confidence, the cloud remains a source of enduring strategic advantage. For those who cannot, the coming period of regulatory intensification, systemic scrutiny and stakeholder attention will be uncomfortable at best. The good news is that the path from one condition to the other is well understood. It requires clarity of intent, executive sponsorship, disciplined execution and a willingness to confront uncomfortable truths about dependencies that have accumulated, often quietly, over many years of migration and modernisation. That work is neither glamorous nor short. It is, however, one of the most consequential investments an enterprise can make in its own long-term resilience.
#CloudStrategy #OperationalResilience #CyberResilience #BankingTechnology #RiskManagement #CIO #VendorManagement





