Services Method Training Firm Work with us FAQ Insights Contact
ESENPT

Home  ›  Insights  ›  12

Article 12 · Problems That Don't Come Back

From symptom to cause

Second of three parts. The 5 Whys hang laminated in every plant and almost no one uses them well. The difference between desk analysis and the kind that works isn't the tool: it's verifying every link with data and demanding that the cause explain all the facts. We introduce The Closed Loop, the method that does move the needle.

By 32sur · July 2026 · Reading time: 13 minutes

The customer complaint came in on a Tuesday: a product batch out of spec. The report was closed in four days with a tidy root cause —"operator error"— and a corrective action: retrain the staff. File closed, indicator green.

Three months later the same defect reappears. Another shift, another operator, the same training that was never scheduled. No one had looked at the data in plain sight from day one: in the Pareto of complaints, the defect was concentrated, almost entirely, on the night shift. It wasn't the person. It was the line's start-up setup, badly adjusted since the shift change. The five whys never went down to the plant floor; they were filled in a room, from the memory of whoever wrote the report.

The scene is composite —it repeats, with variations, in any operation in the region—, but the logic is exact. And it probably sounds familiar.

In the first installment of this series we saw why the system rewards fighting fires and not preventing them. This is the sequel: what happens when someone does try to genuinely solve —and still fails.

The complaint was closed. The question is whether the problem was solved.

To understand it, we have to start with the most famous and worst-used of the quality tools.

The poster no one fills in

In almost any quality room there's an immaculate poster: the five whys, in five boxes, framed. Everyone glances at it on the way to the coffee machine. Almost no one remembers the last time it was filled in with real data. It's used, when it's used, to close fast: you drill down to "lack of attention" or "operator error", you sign, and the closed-complaints indicator ticks up by one.

The myth behind that habit is simple and recited from memory in any quality committee: asking "why" five times leads, almost by design, to the root cause. The tool comes from the production line, from Taiichi Ohno's insistence on not settling for the first explanation of a failure. That origin is real and worth reclaiming: it isn't a consultant's cliché, it's a shop-floor tool.

The problem isn't what Ohno invented. It's what was done with it afterward.

In 2017, a review published in BMJ Quality & Safety was scathing about the way the five whys are used in practice. Different people, facing the same incident, arrive at different root causes. The depth at which the fifth why stops is arbitrary —why five and not seven?—. And the method pushes, almost by structure, toward settling on a single cause when most complex problems have several. Its popularity, the review concludes, doesn't come from working better than the alternatives: it comes from being simple, fast and fitting inside a laminated box.

None of those failings is in the tool. They're in where and how it's used: in a meeting room and from memory, not at the place of the facts and with data.

Before we go on, a management question: the last "human error" that closed a complaint in your company —did anyone verify it was human, or was it simply the fastest way to close the file?

Desk 5 Whys vs. gemba 5 Whys

Here is the heart of the difference between amateur and professional analysis, and there's nothing esoteric about it: it's a matter of verifying every link of the causal chain with data, at the place where the work happens —the gemba, in Toyota's vocabulary—, instead of reconstructing it from memory in an air-conditioned room.

The rule Toyota demands of this kind of analysis is simple to state and hard to meet. The root cause has to explain all the data of the current state, not just the specific case that prompted the meeting. If the defect appears only on the night shift, "operator error" can't be the root cause: it doesn't explain why the morning shift, with different operators and the same training, doesn't have the same problem. A cause that doesn't explain the full pattern isn't a cause: it's a conjecture in the shape of a conclusion.

Go back to the complaint at the start. The Pareto was built, someone had printed and filed it next to the report. It showed, unambiguously, that 80% of the out-of-spec units came from a single shift. No one cross-checked that data against the cause about to be signed. Had they done so, the next question wouldn't have been "who made the mistake?" but "what's different about that shift?". The start-up setup, the absence of a second check, the rotation of a less senior supervisor. That's the fifth why that actually matters, and the one almost never asked because it means leaving the room.

The same complaint, two paths Complaint: product batchout of spec The desk path — the symptom The gemba path — the cause 5 Whys in the room, from memory Cause signed:«operator error» Closed in 4 days ·indicator green 3 months later: the samedefect reappears (anothershift, another operator) Go to the floor andcross-check the Pareto Data: 80% comes froma single shift The right question:«what's different about that shift?» Verified cause: start-upsetup badly adjusted Countermeasure withowner and date The defect doesn't come back comes back
Figure 1 — The same complaint, two paths: the desk path recurs; the gemba path doesn't. Illustrative case; Card, "The problem with 5 whys", BMJ Quality & Safety, 2017.

This connects with something this firm has written about metrics: sometimes the mis-diagnosed problem is nothing more than a badly chosen indicator. We saw it when we discussed why metrics get corrupted: if the dashboard measures "closed complaints" instead of "complaints that don't recur", the whole system pushes toward the fast close, not the right cause.

A routine to install tomorrow: before signing any root cause, demand a single test —that it explain all the available data: shifts, lines, batches, suppliers, time slots. If there's a single data point the proposed cause doesn't explain, the cause is incomplete. Go back to the gemba.

But even the best cause analysis is wasted if it starts from a badly framed problem. And there almost everyone fails, much further upstream in the process than they imagine.

The problem statement almost no one writes

Before asking why, you have to be able to write precisely what.

As Dwayne Spradlin wrote in Harvard Business Review in 2012, the rigor with which an organization defines the problem is what weighs most when it comes to finding a good solution. Most companies aren't rigorous enough at that step and, because of it, end up investing resources in solving something that wasn't the real problem. It's the same principle this firm has worked on the capital-projects side: the front-end loading we covered in investment governance. Cheap definition prevents the expensive correction, whether in an industrial plant or an US$18 million project.

What does a well-written problem statement look like in practice? Nelson Repenning and colleagues, in a 2017 MIT Sloan Management Review article, propose five conditions that are rarely all met together:

Five conditions of a good problem statement

  1. It connects with something the organization genuinely cares about —cash, margin, customer, safety— and not with an abstract concern.
  2. It quantifies the gap between the current state and the target, with numbers, not adjectives.
  3. It uses measurable variables, not impressions ("the defect rose from 0.9% to 1.3%", not "quality got worse").
  4. It's neutral about the solution or the diagnosis. This is the most uncomfortable condition: if the statement already contains the solution ("we need to train the staff"), it isn't framing a problem —it's hiding a conclusion that hasn't been proven yet.
  5. It has a small scope, narrow enough to be tackled fast: weeks, at most a quarter. Big problems are framed better when broken into pieces someone can solve within that short window, not as a year-long crusade.
The good-statement filter: five conditions before moving to cause analysis «the line staff lacks training» a solution disguised as a diagnosis «the night shift's reject rate is at 1.8% against a target of 1%; the cause is unknown» 1. Connects with something the company cares about (cash, margin, customer, safety). 2. Quantifies the gap between the current state and the target, with numbers. 3. Uses measurable variables, not impressions. 4. Neutral about the solution (hides no unproven conclusion). 5. Small scope: weeks, a quarter at most. → ready for cause analysis ? ✗ bounces — doesn't move to cause analysis If it fails two of these conditions, it isn't ready to move to cause analysis.
Figure 2 — The good-statement filter: five conditions before moving to cause analysis. Repenning, Kieffer & Astor, "The Most Underrated Skill in Management", MIT Sloan Management Review, 2017; Spradlin, "Are You Solving the Right Problem?", HBR, 2012.

A one-line template for the next committee: does the statement connect with something the company measures and cares about? does it quantify the gap? is it neutral about the solution? does it have a scope of weeks, not years? If it fails two of those four questions, it isn't ready to move to cause analysis.

With the problem well framed and the cause verified at the gemba, the skeleton of a complete routine begins to show. At 32sur we gave it a name.

The Closed Loop: the four stations

We call this method The Closed Loop, and it has four stations, not three loose steps or one more poster for the wall.

Frame. The problem is written with the five conditions from the previous section. Define, don't diagnose.

Verify. The cause is tested at the gemba, with data, and must explain one hundred percent of the available facts —not just the case that prompted the meeting.

Solve. The countermeasure is assigned to an owner with a full name, and a date. Not a generic training "for the team", not a memo that circulates and gets filed.

Close. You measure whether the problem happened again —the recurrence rate— and run a short debrief that records what was learned. This last station is the one almost no one installs as a routine, and it's so decisive that we'll devote the whole next installment to it.

The method's signature, what sets it apart from any training workshop: the loop only closes when you measure that the problem didn't come back. Not when the report is signed.

The Closed Loop 32sur framework — the four stations of structured problem-solving if it recurs, start over RECURRENCE the metric that closes the loop 12 o'clock · FRAME Define the gap. Neutral to the solution. Scope of weeks, not years. INTEGRATE 3 o'clock · VERIFY Cause tested at the gemba. Explains 100% of the data. IMPLEMENT 6 o'clock · SOLVE Countermeasure with owner and date. IMPLEMENT 9 o'clock · CLOSE Did the problem come back? Debrief + record (A3/COE). SUSTAIN CADENCE · OWNER · DASHBOARD
Figure 3 — The Closed Loop: the four stations. 32sur framework. Based on Shook, "Managing to Learn" (A3), 2008; Amazon "Correction of Errors" (mandatory recurrence field); Repenning & Sterman, CMR, 2001.

Each station calls for a different tool depending on the type of problem, and that's where teams go wrong most: they use the same tool for everything because it's the one they know.

If you lead and don't execute, you don't need to memorize the table that follows —your quality people do—. It's enough for you to demand that this criterion exist.

ToolWorks when…The trap
5 Whys (verified at the gemba)single cause, short chain, data on handdone from a desk: biases toward one cause
Ishikawa / fishbonethere are multiple possible cause families (6M)filled out whole and not prioritized with data
Issue tree / MECEthe problem is big and must be decomposedbuilt without hypotheses and ends up decorative
A3you need a common language and a shared recordbecomes paperwork if it doesn't go to the gemba
DMAIC / Six Sigmarepetitive process with measurable variationheavy for a small, one-off problem
8Dcustomer complaint that demands urgent containmentclosed at D3 (containment) and never reaches D4-D5 (root cause)

Charles Conn and Robert McLean sum it up bluntly in Bulletproof Problem Solving (Wiley, 2019): systematic problem-solving "isn't taught in most universities or business schools". It's learned as a craft —with logic trees, explicit hypotheses and a tool chosen on purpose—, not by inspiration or from the poster stuck on the wall for five years.

For the committee board: The Closed Loop on a single A3 sheet —Frame / Verify / Solve / Close— with the "which tool for which problem" table beside it, visible in the room where complaints are discussed.

All of this can sound like added cost: more steps, more discipline, more meetings. It's worth looking at the other side of the counter.

What not having a root cause costs

There's a factory that appears on no org chart and that almost no company measures: the one devoted, every day, to reworking what came out wrong the first time. Rework, scrap, re-inspection, returns, and in the worst case, a recall. Quality specialists call it, rightly, the hidden factory.

The cost of poor quality usually sits, per industry estimates compiled by the American Society for Quality, around 15-20% of revenue in plants that have never measured it. In those operating with world-class discipline, it drops below 5%. It's an indicative figure, not an academic study with a controlled sample; what matters is the gap: between not measuring it and measuring it there are, at minimum, ten points of revenue in difference.

That cost isn't abstract. In a machining plant in the region, the customer reject rate jumped from a target of 0.9% to 1.57%: 4,933 defective parts and a loss of nearly 34,500 dollars in a single reporting period. The initial meeting, as in the night-shift complaint, looked for a culprit. But this time the team ran a full DMAIC project —define, measure, analyze, improve, control— instead of closing with a generic training. It verified the cause on the floor, validated it against production data and only then defined the countermeasure. The process returned to the 0.9% target.

The difference between that case and the opening complaint isn't the industry or the size of the company. It's that someone demanded the cause explain the data before signing the close.

1.57%the customer reject rate that triggered 4,933 defects and nearly US$34,500 in losses in the auto-parts case
15-20%of revenue goes to the cost of poor quality in plants that never measured it (indicative industry range, ASQ)
90%+on-time supplier payments, after replacing quick fixes with structured problem-solving (RefineCo case, MIT Sloan Management Review)

An afternoon's exercise: add up a month of rework, scrap, extra freight for replenishment and re-inspection hours. The total usually surprises —and it's almost always the budget that already justifies doing things right the first time.

And does it really work, beyond one specific plant? There's a documented case that shows it with concrete numbers.

The proof: the refinery that stopped patching

In a case documented by MIT Sloan Management Review in 2017, a refinery had been operating in permanent firefighting mode. Procurement was piling up supplier-payment delays that put it on the verge of losing credit lines, and each week was a succession of improvised fixes to keep the plant from stopping. The team replaced that way of working with structured problem-solving —the very elements of this method: a problem framed with data, a verified cause, a countermeasure with an owner. The result was measurable: on-time supplier payments went from crisis levels to more than 90%. On top of that, the area freed up two roles previously devoted full-time to putting out that particular fire.

The proper caveat applies: it's a documented case, not an industry average. But it's exactly the kind of evidence that separates "sounds reasonable" from "proven to work".

The phrases that close the file

"We already do the five whys." Good —in the room or on the floor? Does the cause you wrote down explain all the data —the shifts, the lines, the batches— or just the one case in front of you that day?

"We hired a consultancy that left us the methodology." The methodology is almost never the problem. It's used well by whoever goes to the gemba and verifies every link with data; the rest hang it on the wall, next to the five-whys poster, and keep closing complaints with "operator error".

Back to the night shift

The opening complaint was closed in four days. The indicator turned green. And the defect came back, three months later, on the same shift, with another operator carrying blame that wasn't his. The method to prevent it already existed inside the company —the Pareto was printed, filed, a click away— but it never came down from the meeting room to the production line.

Now suppose your company does all this well. It frames the problem rigorously, verifies the cause at the gemba, solves with an owner and a date. There's still one last trap, and it's the most expensive of the three you'll meet in this series. The method applied well just once is not an installed capability. A systematic review of 21 root-cause-analysis studies found that only two showed a real improvement, and that in half the cases the recommendations weren't enough to keep the problem from happening again.

In the final installment of this series: how the loop is installed so it truly doesn't come back —who has to be the owner, what cadence sustains it and which single metric almost no one in the region measures yet.

Are your complaints closed with "human error" and back in three months?

The method to separate symptom from cause exists, is proven and isn't expensive. What's expensive is not bringing it down to the place of the facts. 32sur works on structured problem-solving with gemba data —frame it well, verify the cause, solve with an owner— and doesn't leave you a filed report: it leaves the method installed in your team.

Let's talk

References

  1. Dwayne Spradlin, "Are You Solving the Right Problem?", Harvard Business Review, September 2012.
  2. Alan J. Card, "The problem with '5 whys'", BMJ Quality & Safety, 26(8), 2017.
  3. Nelson P. Repenning, Don Kieffer & Todd Astor, "The Most Underrated Skill in Management", MIT Sloan Management Review, 2017.
  4. Nelson P. Repenning, Don Kieffer & James Repenning, "A New Approach to Designing Work", MIT Sloan Management Review, December 2017.
  5. Charles Conn & Robert McLean, Bulletproof Problem Solving: The One Skill That Changes Everything, Wiley, 2019.
  6. American Society for Quality (ASQ), industry estimates on the cost of poor quality (COPQ).
  7. Case "Drilling-process optimization" (machining plant, automotive sector, LatAm), documented DMAIC, 2022.