When does hybrid cloud beat going all-in on cloud?
In most cases it loses, and the reason is not technical. A hybrid estate means two operating models, two sets of skills, two on-call rotations, two security boundaries and a link between them that somebody has to own. That overhead is permanent and it is paid whether or not you use the flexibility it buys. Hybrid wins when a specific, nameable constraint makes the alternative impossible, and it is worth being ruthless about whether such a constraint exists, because the failure mode is paying the overhead for a reason nobody can articulate.
What does hybrid cost that a single environment does not?
Two of nearly everything, plus the boundary. That is the honest accounting, and it is usually missing from the comparison because the costs are operational rather than a line on an invoice.
The duplication is specific: two monitoring stacks or one stack instrumented twice, two backup and recovery approaches, two patching regimes, two capacity planning exercises, two sets of runbooks, and engineers who must be competent in both. Hiring is harder because the useful candidate understands cloud services and physical networking, and that person is scarcer than either specialist.
The boundary is its own cost centre. Somebody owns the link, the identity federation, the certificates that span both sides and the drift between them, and in many organisations that somebody is nobody, which is how these estates degrade. Incidents are also slower to resolve, because the first question in a hybrid outage is which side is at fault and the answer requires two teams to be awake.
Which reasons for hybrid actually hold up?
Four, and they share a property: each names something that physically or legally cannot be relocated, with a person who can point at it. A reason that cannot survive the question "what specifically cannot move, and who says so" is a preference rather than a constraint.
Regulated or contractually bound data is the most common and the most durable. Latency to a physical process is the most absolute, because no commercial arrangement changes the speed of light or the reliability of a link to a factory. Software licensed to physical hardware, or licensed in a way that makes cloud deployment punitively expensive, is a genuine constraint even though it is an accident of a vendor's commercial model. Capital equipment with remaining book value is real but temporary, and it should carry an expiry date rather than a permanent architecture.
Notice what is absent from that list. Performance in general, security in general, and control in general do not appear, because each dissolves into specifics under examination, and the specifics either name one of the four above or reveal that the concern is about capability rather than location.
| Stated reason | Holds up? | What to check |
|---|---|---|
| Regulator or contract requires the data stays put | Yes, when the clause is produced | The actual text, and which data categories it names |
| Latency to machinery or a physical process | Yes | The loop period required, against measured round trip to the nearest region |
| Licence tied to physical hardware | Yes, until renegotiated | Whether the vendor offers cloud terms and at what price |
| Hardware bought recently with book value left | Temporarily | The depreciation end date, which is the architecture's expiry date |
| The cloud would cost more | Sometimes | Load steadiness, and whether an existing team already runs hardware |
| We must avoid vendor lock-in | Rarely as stated | What an exit would cost today, versus the price of designing for portability |
| Security is better on our own hardware | No, as stated | Which specific control is unavailable in the cloud |
Do the cost and lock-in arguments survive scrutiny?
Partly, and they are worth separating because they fail for different reasons. The cost argument is genuine in a narrow band: steady, predictable load, on hardware run efficiently, by a team that already exists and has the skills. Under those conditions owned infrastructure can be cheaper, and organisations that have moved workloads back have generally been in exactly that position. What makes the argument fail more often is that the comparison omits the staff, the facility, the refresh cycle and the idle capacity bought for peaks that occur rarely.
The lock-in argument fails in a more specific way. Avoiding lock-in usually means designing to the lowest common denominator of all platforms, which means forgoing managed services and building the equivalents, which costs engineering time every year in exchange for a portability that is exercised almost never. That is a real price paid continuously for an option rarely used.
The version of the lock-in concern that does hold up is about exit cost rather than architecture. Knowing what a full extraction would cost today, and re-checking it annually, is a cheap and genuine mitigation. It preserves the ability to make a decision without paying for portability you are not using.
When does latency genuinely force it?
When the process has a control loop shorter than the round trip to the nearest cloud region, or when it cannot pause if the link drops. Both conditions matter, and the second is the one people forget: a process that could tolerate the latency but not the outage still has to run locally.
The clear cases are industrial. A machine reacting to a sensor on a production line, vision-based inspection or safety systems, and robotics all operate on cycles far shorter than a network round trip across a country, and a link failure is not a degraded experience but a stopped line. Similar reasoning applies to any site that must keep operating when its connectivity fails, such as a remote facility, a vessel or a venue.
The test is arithmetic rather than judgement. Measure the actual round trip from the site to the nearest region, add the queueing and processing time your system needs, and compare that to the loop period the process requires. If it fits with margin, latency is not your constraint and the reason for hybrid must come from elsewhere. If it does not fit, no architecture argument changes it.
What are the options between a data centre and a full migration?
Three, and they are underused because they sit outside the binary the discussion usually assumes. Each removes some of the duplication that makes hybrid expensive while keeping whatever cannot move.
Cloud services running on equipment at your site, sold as Outposts, Azure Local and Google Distributed Cloud, give you cloud tooling on hardware in your building. Who owns that hardware differs by product, and it matters. An Outposts rack is AWS-owned, AWS-installed and AWS-maintained, so the commitment is to a rack and to that provider. Azure Local runs on Microsoft-validated OEM hardware you buy yourself from Dell, HPE, Lenovo and others, and Microsoft charges a per-core Azure subscription on top, so the commitment is capital you have already spent plus a recurring fee. The appeal in both cases is one operating model instead of two, which addresses the largest hidden cost of hybrid. The cost is a service catalogue narrower than the full region, so it needs checking against what you actually run.
Colocation adjacent to a cloud region is the second: your own hardware in a facility with direct, very short connectivity to the provider, which reduces latency and simplifies the link without changing who owns the machines. The third is simply keeping a small, well-defined on-premise footprint for the constrained workload and moving everything else, which is often the right shape and is rarely proposed because it is less tidy than a plan that moves all or nothing.
How do you decide?
Write one sentence for each workload staying on-premise: "this cannot move because X". Then check that X is one of the four constraints that hold up, that a named person will stand behind it, and whether it has an end date. Anything that produces a vague sentence or no name is a candidate for moving, and the list of such workloads is usually longer than expected.
Then price the boundary honestly. The link, the identity work, the drift detection, the second on-call rotation and the engineer who understands both sides are the cost of being hybrid, and they belong in the comparison alongside compute and storage. A decision made on infrastructure prices alone will favour hybrid more often than a decision made on total operating cost.
The outcome that surprises organisations most often is a very small permanent hybrid footprint: one or two systems that genuinely cannot move, everything else in the cloud, and a deliberately simple connection between them. That is usually cheaper than either a full migration attempted against a real constraint or a large hybrid estate maintained out of habit.
Common questions
- Is hybrid cloud better than public cloud?
- Usually not, because hybrid means two operating models to run: two monitoring approaches, two backup regimes, two on-call rotations, engineers competent in both, and a boundary somebody must own. That overhead is permanent and is paid whether or not the flexibility is used. Hybrid wins when a specific constraint makes a full migration impossible, most commonly regulated data, latency to a physical process, a licence tied to hardware, or capital equipment with remaining book value.
- What are the real reasons to keep infrastructure on-premise?
- Four survive scrutiny. Data a regulator or contract requires to stay in a place, verified by producing the actual clause. A process with a control loop shorter than the round trip to the nearest cloud region, or that must keep running when connectivity fails. Software licensed to physical hardware or priced punitively for cloud deployment. And capital equipment with remaining book value, which is a temporary reason and should carry an expiry date rather than a permanent architecture.
- Is on-premise cheaper than cloud?
- In a narrow band: steady predictable load, hardware run efficiently, and a team that already exists with the skills to run it. Organisations that have moved workloads back generally match that description. The comparison fails when it omits staff, facility costs, the refresh cycle and idle capacity purchased for peaks that occur rarely, which is why infrastructure-only price comparisons favour owned hardware more than total operating cost does.
- Is avoiding vendor lock-in a good reason for hybrid cloud?
- Rarely as usually stated. Designing for portability means working to the lowest common denominator across platforms and rebuilding equivalents of managed services, which costs engineering time every year in exchange for an option almost never exercised. The version of the concern that does hold up is exit cost: knowing what a full extraction would cost today and re-checking it annually is a cheap mitigation that preserves the decision without paying continuously for unused portability.
- What is the middle ground between on-premise and full cloud migration?
- Three options. Cloud services on equipment at your site, such as Outposts, Azure Local or Google Distributed Cloud, give you one operating model on hardware in your building, at the cost of a narrower service catalogue. Check who owns the hardware before you compare the commitments: an Outposts rack is AWS-owned, AWS-installed and AWS-maintained, while Azure Local runs on Microsoft-validated OEM hardware you buy yourself and pay a per-core Azure subscription to run. Colocation adjacent to a cloud region keeps your own hardware but shortens the link dramatically. And keeping a small, well-defined on-premise footprint for the constrained workload while moving everything else is often the right answer and rarely proposed.
- How do you decide whether to go hybrid?
- Write the sentence "this cannot move because X" for every workload staying on-premise, then test whether X names a real constraint with a person behind it and an end date. Then price the boundary: the network link, identity federation, drift detection, the second on-call rotation and an engineer who understands both sides. The common outcome is a very small permanent hybrid footprint rather than a large one, because most workloads produce a vague sentence and no name.