The short version
Removing capacity before proving reliability inverts the correct order. Deployments that hold up run alongside the existing process until the numbers justify the change — which costs more up front and far less in total. If you are already in the recovery, rehire for exception handling rather than for the old process, and fix the measurement before anything else.
Why does this keep happening?
Because a demonstration is a very poor predictor of production accuracy, and almost nobody treats it as one. The demo was built on a sample someone assembled, which means the awkward records were quietly excluded as noise. It measured the clean path, and the clean path was never the expensive part.
Add a board expecting savings, a vendor whose incentive is to close, and a project sponsor who has publicly committed to a number, and the decision to cut gets made on the most optimistic figure anyone has produced. The first honest measurement arrives after the redundancies.
What actually breaks first?
Not accuracy on the common cases — that usually holds. What breaks is everything around it:
- The exceptions have nowhere to go. Ten percent of cases need a human and the humans are gone, so they queue.
- Nobody can tell which outputs are wrong. Without confidence signals or citations, checking one means checking all.
- The edge cases were undocumented. The rules for handling them lived with the people who were let go.
- Nobody owns it. When accuracy drifts, there is no one whose job is to notice.
The result is a system that is right most of the time and trusted none of the time, which is worse than the manual process it replaced because the review burden landed on people who were not resourced for it.
Was the technology oversold, or the deployment underdone?
Mostly the second, and it matters because the two lead to opposite responses. If you conclude the technology cannot do it, you stop — and a competitor who deployed it properly takes the advantage. If you conclude the deployment was underdone, you fix the deployment.
The diagnostic is straightforward. Ask whether an accuracy threshold was ever agreed, whether it was measured on real uncurated data before the headcount decision, and whether there was a written exception path. If the answer to any of those is no, the technology was never actually evaluated.
What should have happened instead?
A parallel run. The system produces its answer, the existing team produces theirs, and you compare for a full business cycle without changing anyone’s workload.
It feels like paying twice and it is the cheapest insurance available. It gives you an accuracy figure on real data with no risk attached, a labelled record of every disagreement, and a team that has watched it work before being asked to rely on it. Adoption fails on trust far more often than on capability, and a parallel run is how trust gets built.
Only after that do you have the information a headcount decision requires: not “does it work” but “what proportion of cases does it handle to what standard, and what does the remainder cost to service”.
We are already in the recovery. What now?
Four steps, in order:
- Measure honestly, on real data. You probably still do not have a true accuracy figure. Get one before making any further decisions.
- Separate the cases it handles from the ones it does not. Almost always the failures concentrate in identifiable categories rather than spreading evenly, and those categories can be routed away rather than fixed.
- Rehire for the exceptions, not the old process. Different job, smaller team, higher skill — and say so, because rehiring people into the role you just made redundant without acknowledging the change corrodes what trust remains.
- Document what the departed staff knew. If any are contactable, this is worth paying for. The exception rules are the specification and they were never written down.
Does this mean nobody should reduce headcount?
No. It means the reduction follows the evidence rather than preceding it, and that the honest saving is usually smaller and later than the business case assumed — because a deployment that handles eighty percent of volume does not remove eighty percent of the people. Someone still services the remainder, owns the system, and handles the cases that reach a customer.
Organisations that do this well tend to reduce through attrition and redeployment over a year rather than through a cut in month three. Slower, considerably less disruptive, and it does not destroy the knowledge you turn out to need.
What does the rehiring cost, really?
More than the redundancy saved, in most cases, and the direct costs are the smaller part. There is recruitment and notice. There is the productivity gap while replacements learn a process nobody documented. And there is the reputational cost in a labour market where the people you want have watched what happened.
The largest cost is usually institutional. The undocumented knowledge left with the people who held it, and rebuilding it means rediscovering by trial what used to be known. That is the expense nobody forecasts and everybody pays.
How do we avoid repeating it on the next process?
Write the rule down and apply it to every deployment: no capacity is removed until the system has run in parallel for a full cycle and met an accuracy threshold agreed in advance.
It is one sentence, it is unpopular with anyone who has promised a saving by a date, and it would have prevented most of what has been written about this over the past year. Attach it to the business case rather than the project plan, so it survives contact with the pressure to deliver early.
How do we tell our team what is happening?
Directly, and before the rumour does it for you. A workforce that has watched one round of AI-driven redundancies followed by rehiring will assume the next deployment means the same thing, and that assumption is expensive — because the people whose knowledge you need to encode are the ones who now have every reason not to share it.
What changes the dynamic is being specific about the intent for this deployment: which work it takes, what happens to the time saved, and whether any role is at risk. If the honest answer is that roles will change, say so. Vagueness reads as bad news being withheld, and it is the fastest way to lose the co-operation the project depends on.
What does the second attempt look like?
Narrower, measured, and with the capacity question deferred. Pick one process rather than a programme. Agree the threshold before you start. Run in parallel for a full cycle. Publish the numbers internally, including the bad ones.
The organisations that recover well tend to make the second deployment deliberately unambitious and very well evidenced, because credibility is the scarce resource after a failure rather than budget. One process working demonstrably beats a broader plan nobody believes.