Eight fixes for a fault everyone already knew
Everyone in the building could describe the fault. Nobody had diagnosed it.
An 8 to 10 week sprint. One senior analyst and two juniors. A €600 million measurement and industrial technology company. Part three of three on operational improvement. Part one was about standardising the work. Part two was about whether you can tell if the improvement on top of it is working. This one is about a fault everybody could name and nobody had traced.
Deliveries were going out wrong. Not late, not damaged in transit. Wrong: the box arrived and what was inside it was not what the order said, or not all of it.
The industry term is an out-of-box failure, and the phrase is precise about the worst part. The failure is not caught in the warehouse. It is found by the customer, at the moment the carton is opened, which is the one point in the whole chain where nobody from the company is standing there.
So this was not a discovery project. Nobody needed convincing the problem existed. What did not exist was any single, fact-based view of the types, drivers and relative impact of those failures. That gap is the page’s own wording, and the consequence it names is worth reading twice: it increased the risk of assumption-driven fixes and local optimisations.
A fault everybody can describe is not the same as a fault anybody has diagnosed. The gap between those two states is where this engagement sat.
A theory costs nothing. Testing one costs weeks.
The intuitive answer is that nobody cared enough. That is almost never the answer and it was not the answer here.
A recurring operational fault survives because of an asymmetry. Holding a theory about the cause is free. Testing one is expensive. So theories accumulate, roughly one per function, each of them plausible, and none of them ever settled.
Ask the packing line and the cause is upstream, in how orders arrive. Ask the order desk and it is downstream, in how things are picked. Ask planning and it is the product mix. Every one of those answers is defensible from where it is standing. That is exactly what makes the situation stable: the disagreement is not between people who are wrong and people who are right, it is between people who each hold one true fragment and no way to assemble them.
The second half of the asymmetry is who would have to resolve it. The page states that the logistics and operations teams were fully focused on day-to-day execution. The people with the knowledge to diagnose the fault were the people whose full-time job was keeping flow moving through the operation producing it. Taking them off that to spend a month counting is not a scheduling problem. It is a trade against this quarter’s delivery performance, made by the person who is measured on delivery performance.
That trade has an obvious answer every single week, and the fault survives another quarter.
Classify, count, then rank by what it costs rather than by how often it happens
Fix the classification before counting anything
An operation this size does not lack data. What it lacks is a failure type that everybody applies the same way. Without one, two people counting the same month produce different totals and both are right. The classification is not admin ahead of the analysis, it is the first analytical decision and everything downstream inherits it.
Interview along the chain, not inside one function
More than twenty stakeholders, taken across organising and packing rather than from the function that owns the problem on the org chart. Each function holds a fragment that is true locally. The point of going wide is not balance, it is that the cause of a handover failure is never visible from one side of the handover.
Put the fragments in one room, four times
Four cross-functional workshops. A workshop is the cheapest instrument that exists for finding out which fragment belongs where, because the packing view and the order desk view have to be stated in front of each other before either can be tested. Interviews collect the fragments. Workshops are where they get arbitrated.
Rank on relative impact, which is the hard one
Type and frequency are countable. Relative impact is a judgement about consequence, and it is the one that decides where the effort goes. A rare failure with an expensive consequence outranks a common cheap one, and frequency on its own will point a competent team at the wrong fix.
Two outcomes, and this page words its own numbers better than most
| What the page states | Figure | How it is worded |
|---|---|---|
| Reduction in out-of-box failures | 30% | Reduction potential identified |
| Improvement solutions | 8 | Prioritized, defined |
Read the first line again, because the qualifier is doing real work. Thirty per cent is a reduction potential that was identified. It is not a reduction that was banked.
That distinction matters more than it looks. An eight to ten week sprint that classifies failures, counts them and ranks the causes has produced a diagnosis and a size for the prize. It has not produced a delivered improvement, because delivering one requires the implementation the sprint deliberately did not run. Whether any of the thirty per cent arrives depends entirely on what happened after the handover.
Part two of this series criticised its own case page for stating three figures flatly, as things that had happened, when the sprint had produced an instrument rather than a year of readings. This page does not make that mistake. It carries the qualifier itself, in its own words, before anybody asked it to. That is the standard the rest of the estate should be held to, and it is worth naming when a page gets it right rather than only when one gets it wrong.
The senior decided what counted as impact. The juniors did the counting.
One senior analyst and two juniors, which is the whole model and is the reason the arithmetic works.
The senior’s work was the two judgements that cannot be delegated: what the classification distinguishes, and what relative impact means for this business. Both are decisions about the shape of the answer, and both are made before the first count. Get either wrong and a large amount of careful work produces a confident ranking of the wrong things.
The juniors’ work was volume. More than twenty interviews is roughly three working weeks of somebody’s life once scheduling, travel and writing up are included. Four workshops is a further set of days that have to be found in the calendars of people who are all busy at the same times. Then the classification has to be applied consistently across enough history for a count to be more than an impression.
That split is the answer to the obvious question about this engagement, and it is worth being exact about it. The company knows more about its products, its customers and its packing operation than any outside team will learn in ten weeks. What it did not have was three uninterrupted weeks of attention, four rooms of the right people at the same time, and the classification work that has to happen before a single count means anything.
Eight is the finding, not the by-product
It is tempting to read eight prioritised improvement solutions as the packaging and the thirty per cent as the result. It is the other way round.
The number that changes behaviour is eight, because eight is small enough to start on Monday. A diagnosis that hands back forty opportunities has not finished its job. It has moved the prioritisation problem from the operation to the reader, and the reader is the person who was too busy to do the diagnosis in the first place.
Prioritised also means something specific here. It means ordered against relative impact, which was the hardest of the three missing things. Without that ordering, eight solutions is a menu. With it, it is a sequence.
The third objective set for the sprint is the one that quietly eliminates a whole class of answer: the roadmap had to be implementable without slowing daily logistics operations. A recommendation that requires the line to stop, or that requires the same fully committed people to run a parallel programme, is not an answer to this problem. It is the problem again with a project code on it.
Three things, and the first one is the number everybody will quote
Whether any of the thirty per cent arrives. It is a potential identified against a ranked set of causes. It becomes a result only through implementation that this engagement did not run and cannot vouch for. Anyone quoting it as a delivered reduction is quoting it wrong, including us.
The ranking rests on a judgement. Relative impact is not measured, it is decided, using a definition of consequence agreed at the start. A different and equally defensible definition would produce a different order in the middle of the list. The top of a well-built ranking is usually robust. The middle is where reasonable people differ, and the eight solutions were not all equally certain.
The window bounds the finding. Types, frequency and root cause were established across organising and packing for the failures inside the period examined. A failure mode driven by seasonality, or by a product introduced after the window closed, would not appear in the count at all. A diagnosis is a photograph of an operation, and operations move.
Where every figure above comes from
From the published case study page. The 30% reduction potential in out-of-box failures, carried here with the page’s own qualifier rather than restated as an achieved result, and the 8 prioritized improvement solutions defined. The three objectives set for the sprint. The statement that there was no single, fact-based view of the types, drivers and relative impact of out-of-box failures, and that this increased the risk of assumption-driven fixes and local optimizations. The statement that logistics and operations teams were fully focused on day-to-day execution. The client’s Vice President of Operations described the outcome as clarity on a recurring issue they did not have the bandwidth to diagnose internally.
From our internal engagement record. The €600 million revenue scale, the eight to ten week sprint, the more than twenty stakeholder interviews and the four cross-functional workshops.
Deliberately not used. The failure types themselves and the order they came out in. That ranking is the client’s operating detail, it is the most commercially useful thing in the whole engagement, and it is not ours to publish. Every argument in this piece is about the method that produced the ranking, which is the part that transfers. The ranking itself stays with the company that paid for it.
Not claimed anywhere. That the failures went down. We do not know, and a page that told you otherwise would be selling you something.
Have a similar requirement?
Contact us today to learn more about on-demand workforce and accelerate development on your most pivotal projects!
Featured Case Studies
Accelerating Success for Enterprises in 20+ Geographies

Launch Your Sprint with
Download the full report
Enter your email to access this exclusive case study.


