Is your hybrid cloud a stalled migration?
On an architecture diagram there is no difference. Both show workloads in two places connected by a link, and both are described in the same language internally. The difference is entirely organisational: a deliberate hybrid has a named constraint and someone who owns both sides, while a stalled migration has a set of workloads nobody scheduled and a plan that quietly stopped being referenced. The second costs more than either finishing or committing, and it is the more common of the two.
How do you tell the difference?
Ask why each remaining on-premise workload is still there and see whether the answer names something. A permanent hybrid produces specific sentences: this database is subject to a clause in a customer contract, this controller runs the production line, this application is licensed to the hardware. A stalled migration produces sentences about circumstances: the team that owned it moved on, it was too risky to do in the last window, we were going to get to it after the reorganisation.
The second test is money. Permanent estates receive investment, so the on-premise side has a hardware refresh plan, a support contract in date, and a documented recovery procedure that someone has actually exercised. Stalled estates receive maintenance only, and the tell is usually that nobody has agreed to spend anything on the on-premise side in a while because it is understood to be temporary.
The third is ownership of the boundary. In a permanent hybrid the connection, the identity federation and the drift between the environments are somebody's job. In a stalled migration they belong to whoever built them, who has often moved to another team, and the first person to discover this is the one handling an incident at an unhelpful hour.
| Signal | Permanent hybrid | Stalled migration |
|---|---|---|
| Reason for on-premise workloads | A named constraint with a person behind it | Circumstances: timing, risk, a departed team |
| Hardware refresh | Planned and budgeted | Deferred each year on the basis that it is temporary |
| Ownership of the link and identity | An accountable team | Whoever built it, who has usually moved on |
| Recovery testing on the on-premise side | Exercised and documented | Assumed to work, last tested some time ago |
| Rate of change on-premise | Normal, with a change process people follow | Frozen, because nobody wants to touch it |
| Documentation | Current, describing the intended architecture | A migration plan with dates in the past |
| Hiring and skills | Deliberately maintained for both environments | Nobody left who knows the on-premise side well |
Why do migrations stall near the end?
Because the workloads that remain are the ones selected against by every prior round. Anything easy went first, so the residue shares a set of properties that made it skippable, and those properties do not improve with time.
The recurring set is small and recognisable. No clear owner, so nobody can approve a change window. No test coverage, so nobody can prove the move worked. A dependency on something unsupported, where the vendor has gone or the version is out of maintenance. A licence the vendor will not extend to cloud deployment on acceptable terms. A dependency on hardware, such as a device, a dongle or a card. And undocumented integrations that only reveal themselves when the system is switched off, which is why nobody wants to switch it off.
There is also a straightforward incentive problem. Migration programmes are usually resourced against a target that has already been substantially met, so the remaining work competes for attention with new initiatives while showing very little visible progress. The last workloads are the slowest and the least rewarding, and no organisation naturally prioritises them without someone deciding to.
What does the stall actually cost?
More than the sum of the two environments, and most of the excess is invisible on any single budget line. The obvious costs are duplicated licences, the network link, and support contracts on hardware serving a shrinking number of workloads. Those are visible and usually understood.
The larger costs are elsewhere. Every engineer must understand both environments, so the effective team is smaller than the headcount. Security surface is doubled while attention is not. Incidents take longer because responsibility is ambiguous. And the estate carries the drift and identity problems described elsewhere in this cluster indefinitely, with nobody funded to address them because the situation is officially temporary.
The most expensive item is usually the one that never appears: the data centre or floor space that cannot be released until the last workload leaves. The saving from decommissioning is almost entirely back-loaded, so a stall at the end forfeits nearly all of the financial case while having paid nearly all of the migration cost. That asymmetry is the strongest argument for either finishing or formally stopping.
How do you restart one?
Not by rerunning the original programme, which will stall in the same place. Start by inventorying what remains by dependency rather than by owning team, because the residue is usually a small number of clusters that must move together and a long tail of things that could have moved individually at any point.
Then make an explicit decision per workload, choosing between four outcomes: move it, replace it with something else, retire it, or declare it permanently on-premise. Retire is chosen far less often than it should be, because nobody wants to be the person who switched off a system that turned out to matter. A period of monitoring access to confirm nothing has used it usually settles that question with evidence.
Batch the tail rather than treating each item as a project. The long tail of small applications is often movable in a single co-ordinated effort with one change window and one rollback plan, and treating each as its own project is precisely how it never happens. Reserve the individual treatment for the genuine clusters with real dependencies.
Is declaring it permanent a legitimate outcome?
Yes, and it is frequently the cheapest available decision. If a workload is stable, meets its requirements and has no compelling reason to move beyond programme tidiness, deciding to keep it converts an open item into a supported system, which is a change in status that has real consequences.
The consequences are what make the decision meaningful rather than an excuse. A permanent workload gets a budget, a refresh plan, a support contract, a documented recovery procedure and a named owner. It gets included in the security programme rather than waived. Its part of the boundary becomes somebody's responsibility. None of that happens for a workload classified as pending migration, which is why an honest reclassification often improves the system immediately.
The honest way to distinguish this from surrender is to require the same sentence as anything else: name what makes moving it not worth doing, and set a date to revisit the decision. A permanent classification with an annual review is a decision. One without a review is the stall wearing different words.
What can you do this quarter?
List every workload still running on-premise and write one sentence next to each: "this is here because X". Do it with the people who operate them rather than from an asset register, and accept the sentences they actually give you rather than improving them. The exercise takes an afternoon for most estates.
Sort the results into three piles. Sentences naming a constraint with a person behind it are your permanent hybrid, and they should be given proper ownership and budget. Sentences naming circumstances are your stall, and each needs a date and an owner. Anything where nobody could produce a sentence at all is a candidate for retirement and should be checked for use before anyone plans a migration for it.
Then publish the three lists. Most of the cost of a stalled migration comes from ambiguity rather than from the technology, and an estate where everyone agrees which workloads are permanent, which are scheduled and which are being switched off is already cheaper to run than one where the same diagram means different things to different teams.
Common questions
- How do you know if your hybrid cloud is really an unfinished migration?
- Ask why each remaining on-premise workload is still there. A deliberate hybrid produces sentences naming a constraint: a contract clause, a production line controller, a licence tied to hardware. A stalled migration produces sentences about circumstances: the team moved on, the window was too risky, it was going to happen after the reorganisation. Two supporting signals are whether the on-premise side has a funded hardware refresh, and whether anyone currently owns the link and identity federation between the environments.
- Why do cloud migrations stall before finishing?
- Because the remaining workloads were selected against in every earlier round, so they share the properties that made them skippable: no clear owner to approve a change, no test coverage to prove the move worked, a dependency on unsupported software or physical hardware, a licence the vendor will not extend to cloud on acceptable terms, and undocumented integrations that only surface when the system is switched off. Programme incentives compound it, since the target has usually already been substantially met.
- What does an unfinished cloud migration cost?
- More than the two environments separately. The visible costs are duplicated licences, the network link and hardware support contracts serving a shrinking workload count. The larger ones are that every engineer must understand both environments, security surface is doubled while attention is not, and incidents take longer because responsibility is ambiguous. The biggest is usually that floor space cannot be released until the last workload leaves, so the financial case is almost entirely back-loaded.
- How do you restart a stalled cloud migration?
- Inventory what remains by dependency rather than by owning team, because the residue is usually a few clusters that must move together plus a long tail that could have moved at any time. Then decide explicitly per workload between four outcomes: move, replace, retire or declare permanently on-premise. Batch the tail into one co-ordinated effort with a single change window rather than treating each small application as its own project, which is how it never happens.
- Is it acceptable to keep some workloads on-premise permanently?
- Yes, and it is often the cheapest decision available. What makes it a decision rather than an excuse is what follows: a permanent workload gets a budget, a refresh plan, an in-date support contract, a documented and exercised recovery procedure, a named owner and inclusion in the security programme. Require the same sentence as anything else, naming why moving is not worth doing, and set an annual date to revisit it.