Improvement that is not landing is three problems
Three mandates, three failures, and one sentence that fits all of them
The coda to a three part series on operational improvement. The parts are one case each and this piece is the argument they add up to. It is the only one of the four that is not a case study, and it is the only one that could not have been written from a single engagement.
Over three sprints we were asked, in three different companies, to do something about improvement that was not working. In every case the request arrived in roughly the same words. Improvement is running. It is not landing.
That sentence is true in all three, and it is useless in all three, because it describes a symptom that has at least three distinct causes and gives no way to tell them apart. The three engagements went in three completely different directions, and the direction was decided in the first fortnight, before anyone had done any improving at all.
The expensive mistake is not choosing the wrong fix. It is buying a fix for a problem you do not have.
What follows is the shape we now use to tell which one is in front of us. It is not a framework and we are not selling it. It is three questions, and the useful part is that they have different answers in different buildings.
Variation, measurement, diagnosis
They are genuinely different problems. Different evidence, different work, different people to talk to, different length of engagement.
Variation
Everyone is doing the job, and each of them is doing a different job
Sixteen factories had each arrived at a way of keeping their equipment running. Every one of them worked, in the sense that the plants ran. What did not exist was a version that could be compared, taught, resourced or improved. A €400 million global chemical company.
Measurement
Improvement is happening and nobody can say whether it is working
Continuous improvement ran at every site with no standardised way to track it. Execution was inconsistent and the impact on outcomes was unclear. Not disputed. Unclear, which is worse. A €14 billion fertilizer and crop nutrition company.
Diagnosis
Everyone can name the fault and nobody has traced it
Deliveries were going out wrong and every function had a theory about why. What was missing was any fact-based view of the types, drivers and relative impact. A €600 million measurement and industrial technology company.
Notice that none of the three is a shortage of effort, and none of them is solved by more improvement activity. Two of them are made worse by it.
They present identically, and the presentation is what gets acted on
In a management meeting all three sound the same. Someone says the programme is not delivering what was expected. Somebody else says the teams are engaged but the numbers have not moved. Everyone agrees something should be done.
The reason this is hard is not that the three are subtle. It is that the person describing the problem is describing the part of it they can see, and each function can see a different part. The site manager sees variation. The programme owner sees a measurement gap. The operations lead sees an unexplained fault. All three are looking at the same organisation and all three are right about what is in front of them.
So the diagnosis gets made by whoever speaks first, or by whoever has budget. And the fix that follows is bought against that person’s view of it.
This is a general shape, not a claim about any of the three clients. In all three engagements the company knew its own operation far better than we did. What none of them had was somebody whose only job that quarter was to work out which of the three they were in.
Three questions, and the one you cannot answer is the one you are in
Could two sites describe the job the same way?
Not whether they get the same result. Whether they would use the same words, the same steps and the same thresholds. If the answer is no, you have variation, and nothing you measure across those sites will mean anything until it is closed, because the measure will not be measuring the same thing twice.
Could you tell whether last year’s improvement worked?
Not whether it felt better. Whether there is a reading from before and a comparable reading from after. If the honest answer is that you would be reasoning from impressions, you have a measurement problem, and the programme has no way of learning from itself.
Could you say which type of your recurring fault costs you most?
Type, frequency and relative impact, and the third one is where it usually fails. If you can name the fault but not rank its causes, you have a diagnosis problem, and the improvement effort will land wherever it is easiest to count rather than where the money is.
Most organisations can answer one of the three comfortably, hesitate on the second and cannot answer the third at all. The one you cannot answer is the one to buy work against.
Standardise, then measure, then diagnose
The three are not independent, and that is the part that took three engagements to see.
Measurement depends on standardisation. An indicator means different things at sites that work differently. If two plants define a maintenance intervention differently, then a count of interventions is two different quantities added together, and the total is not wrong so much as meaningless. The measurement engagement worked because its indicators were grouped into six stages that every site recognised, which is a standardisation act wearing a measurement label.
Diagnosis depends on both. Ranking causes by relative impact requires a classification everyone applies the same way, which is standardisation again, and a frequency count over enough history to be more than an impression, which is measurement. That is why the diagnosis engagement spent its first weeks fixing the classification rather than looking for causes.
You can do them out of order. It costs about a quarter each time, and the work has to be redone.
None of which means a company must complete all three in sequence before anything improves. It means that when the answer to an earlier question is no, work bought against a later one will underdeliver for a reason nobody in the room will be able to name.
What each mistake actually looks like from the inside
| What gets bought | When the real problem is | How it fails |
|---|---|---|
| A measurement system | Variation | Dashboards fill up and nobody trusts them, because each site is reporting a differently defined number |
| A diagnosis | Measurement | The causes are plausible and unrankable, so the recommendation is a list rather than a sequence |
| More improvement activity | Any of the three | Effort rises, the sentence in the meeting does not change, and the programme loses credibility it will need later |
The third row is the common one and it is the most expensive, because it burns the thing you cannot buy back. A company that has run two improvement pushes that went nowhere will not fund the third one properly, even when the third one is finally aimed at the right problem.
Three things it would be easy to read into this, and should not be
It is not a study. Three engagements, three different companies, three different industries, no control and no comparison. It is a pattern noticed across three mandates that happened to land close together. A fourth engagement could break it and we would say so.
It is not a claim that we solved all three. Each sprint produced a diagnosis, an instrument or a standard, and each was handed over. What happened afterwards depended on the client, and in two of the three we have no visibility into it at all. The figures on those case pages are what the pages state, and the next section is honest about how they are worded.
It is not a sequence anyone has to complete. Plenty of operations run well with unstandardised practice, because the variation does not cost them anything they care about. The argument is only that buying work against a later question while an earlier one is unanswered is where the money goes missing.
Where every figure comes from, and one uncomfortable thing about our own pages
From the three published case study pages. The sixteen factories each with their own maintenance practice, the 40+ systems screened and the five shortlisted. The more than thirty cultural indicators, the six KPI areas, the four pilot sites and the two measurement tools. The eight prioritised improvement solutions and the 30% reduction potential identified.
From our internal engagement record. The three revenue scales, and that each was an eight to ten week sprint run by one senior analyst and two juniors.
The uncomfortable part, and it is about us rather than the clients. Reading the three pages side by side is the only way to notice this. Two of the three state their outcome figures flatly, as things that happened. The maintenance page reports downtime reduced 25% and coverage increased 30%. The culture page reports conversion 20% higher, bottleneck identification 30% faster and visibility 3x more. The third page carries its qualifier. It reports a 30% reduction potential identified, which is a different kind of statement and an honest one.
An eight to ten week sprint that builds a standard, or an instrument, hands over a standard or an instrument. It does not hand over a year of operating results. Where a page states a delivered figure that the sprint’s own duration could not have observed, the page is reporting something it cannot have measured, and that is worth fixing on our side rather than defending.
Not claimed anywhere in this piece. That any of the three companies has since improved. We know what was handed over. We do not know what was done with it, and a page that told you otherwise would be selling you something.
Have a similar requirement?
Contact us today to learn more about on-demand workforce and accelerate development on your most pivotal projects!
Featured Case Studies
Supply Chain & SustainabilityEight fixes for a fault everyone already knew
3 September 2026Read ›
Supply Chain & SustainabilityThirty indicators for something everyone called a culture
1 September 2026Read ›
Supply Chain & SustainabilitySixteen factories, sixteen ways to fix a machine
29 August 2026Read ›Accelerating Success for Enterprises in 20+ Geographies

Launch Your Sprint with
Download the full report
Enter your email to access this exclusive case study.