The photo came out well: eighteen people with their freshly awarded green belts, the certifier's logo in the background, smiles. It was a Thursday. The company had paid for the full program —four modules, two facilitators, coffee break included— and for three weeks everyone talked about DMAIC and root cause in the hallways.
Within a month, the calendar swallowed everything again. Within three months, the defect that had driven the course's flagship project reappeared, almost identical. No one linked it to the workshop. And above all: no one asked the one thing that mattered. Hadn't we already solved this problem?
That question has backing in the evidence. And what it says isn't reassuring.
The corridor of frozen A3s
We saw in the first installment of this series why the system rewards fighting fires and not preventing them. In the second, we saw the method that truly separates symptom from cause —the statement that doesn't marry a solution, verification in the field, the countermeasure with an owner. What's left is the hardest part: keeping all of that from dying at three months, when the workshop's enthusiasm has evaporated and the operation has gone back to its usual pace.
Because there, in the corridor of so many companies, hang twelve framed A3s, immaculate, with their charts and their countermeasures. They were made last year, in the continuous-improvement push. They're perfect. And they're frozen: not one was looked at again to confirm whether the countermeasure held. They're posters, not learning.
The question no one asked at the belt workshop —hadn't we already solved this?— is exactly the one that separates an organization that learns from one that only trains people.
The method applied well once isn't a capability
There's a comfortable temptation: if the root-cause analysis identified the cause well, the problem is solved. The evidence says otherwise.
A systematic review published in 2020 in Medical Principles and Practice put a number on the problem. Of 21 studies on root-cause analysis, only two showed a real improvement. In half, the recommendations weren't enough to keep the event from happening again. Put another way: in the vast majority of cases, finding the cause wasn't enough to make the event stop repeating.
2 of 21 — root-cause-analysis studies that showed a real improvement; in half of the 21, the recommendations didn't prevent recurrence. (Medical Principles and Practice, 2020)
The reason shows up elsewhere: when the analysis is done as a ritual —incident by incident, without spreading what was learned—, similar events happen again. As Dixon-Woods and her team warned in BMJ Quality & Safety, the organization, literally, "forgets". Each area solves its own case, files its own A3, and no one connects the dots with the area next door, or even with its own past.
A necessary caveat: both findings come from patient safety, a domain where root-cause analysis is particularly well measured. They aren't a general-industry figure, but the principle transfers without trouble to a plant, a utility or a bank: identifying the cause correctly isn't the same as preventing recurrence. A step is missing.
The management question that follows is simple and annoying: of the improvement projects your company celebrated last year, how many were reviewed afterward to confirm the problem didn't come back? If the answer is zero, you don't have an installed capability. You have a photo archive.
If analysis alone isn't enough, then what is? What the loop was missing: the station that closes it.
Closing the loop: the station no one installs
In the previous installment we introduced The Closed Loop, the framework we use to order problem-solving: Frame (define the gap, without naming the solution), Verify (test the cause with data in the field, until it explains everything observed) and Solve (a countermeasure with an owner and a date, not a generic training). Anyone arriving through this third installment already has the essentials in that sentence.
What's left is the fourth station, and it's the one almost no organization truly installs: Close. To close means measuring whether the problem came back, running the debrief on what was learned and leaving a record someone else can find. If the problem recurs, the loop doesn't end there: it returns to Frame, with the new information on why the previous countermeasure fell short.
The reference model belongs, without much searching, to Amazon. The company doesn't do "postmortems": it does Correction of Errors (COE), with a mandatory recurrence field. If the problem had happened before, you have to explain why the earlier actions didn't prevent it. And a root cause that simply says "the engineer made a mistake" gets the document sent back unapproved. That's not a cause: it's a label that explains nothing and teaches nothing.
That mandatory field is, in essence, Station 4 turned into a writing routine. It forces the question the belt workshop never asked.
The other pillar is a common language for documenting learning: the A3, Toyota's sheet where the problem, its analysis, the countermeasures and the follow-up plan all fit. John Shook, in Managing to Learn, describes it in no uncertain terms. It isn't paperwork: it's the backbone of the system Toyota uses to develop its people and learn from its own work. The discipline of writing it on a single sheet —with no way to hide vagueness behind thirty slides— is what forces you to genuinely define, observe, analyze and countermeasure.
Neither the COE nor the A3 works alone. They need three conditions around them, and those are what truly distinguish the organization that installs from the workshop that only trains:
Cadence. A periodic review of open problems, not an annual event with an English name.
Owner. A name and a date on each problem, not "the department" or "we're looking into it".
Dashboard. A visible record of open, closed and recurred problems —not a folder of filed A3s no one opens again.
The concrete routine to install this tomorrow: add a mandatory box to your problem-closing form —"has this happened before? If so, why didn't the earlier action prevent it?"— and ban "human error" as a final root cause. It's, literally, the field Amazon uses, adapted to any Excel dashboard.
All of this rests, in the end, on a single metric.
The metric almost no one in the region measures
Frame, Verify and Solve integrate and implement; Close is what sustains: the discipline that turns "solve" into "don't let it come back".
It's measured with three simple numbers. Open versus closed problems in a time window. Average time to reach the root cause. And, above all, the recurrence rate: how many of the problems declared closed reappeared in a given period. It's the number the belt workshop never calculated, because no one stayed long enough to calculate it.
The minimum viable dashboard has three columns: open, closed, recurred. You don't need six-figure quality-management software; it's enough that it exists, gets updated and someone looks at it every two weeks.
Recurrence is a metric, and metrics get corrupted when they become a target. If leadership blindly rewards "zero recurrence", what it will achieve is people no longer reporting the problems that come back —exactly the Goodhart dysfunction we already walked through on this blog. The metric has to push toward improving the system, not toward hiding what recurs. A dashboard where nothing ever recurs isn't a sign of success: it's the first suspicion that someone stopped writing things down.
Does this really exist installed in the region, or is it Japanese-manual theory translated into Spanish? It exists, and 90 kilometers from Buenos Aires.
Routine, not event: what happens 90 kilometers from Buenos Aires
At Toyota Argentina's plant in Zárate, solving problems isn't a workshop with a date on the calendar. It's what gets done every day, at every station. Last year, 94% of the plant's staff contributed at least one improvement suggestion within their work circles —61,059 suggestions in total over twelve months—, with a clear logic of roles: operators solve and engineers provide tools.
94% / 61,059 — staff participation and improvement suggestions in a year, Toyota Argentina (Zárate plant).
No belt photo sustains that volume. What sustains it is a daily routine of listening and adjustment, with station owners who are, literally, the people doing the work. It's the positive counter-case to the workshop-with-a-photo: problem-solving turned into a shop-floor habit, not a corporate event with catering.
And there's a second piece of evidence, this one properly regional, that somewhat contradicts management common sense. The evidence on Lean and Six Sigma in Latin American SMEs agrees on one point: what makes or breaks the method isn't the tool —it isn't whether you use DMAIC or A3 or five S's. It's the commitment of senior leadership and whether the method is tied to the business strategy. Company size and scarcity of resources are the main reason these programs are abandoned, more than any technical resistance to the method itself.
Put another way: the tool is the easy part. The hard part —and what really decides whether the loop holds or falls— is whether leadership sustains it with presence, not just with an initial budget.
The leadership routine that follows: a fixed, biweekly, thirty-minute cadence, reviewing the open/closed/recurred problem dashboard, with leadership present in the room. Not delegated. The regional evidence is clear on one point: without that presence, the program is abandoned before it pays off.
Now, the part no workshop photo shows: installing this hurts before it pays off.
The cost of doing it right
Here is the point where most serious attempts to install this capability die. Repenning and Sterman called it, back in 2001, "worse-before-better": investing in improvement worsens the numbers in the short term before improving them. Pulling two people out of the fire to build the problem-solving routine means, for a stretch, fewer hands putting out fire —and the fire, meanwhile, keeps starting.
That's why most improvement programs don't die of resistance. They die of haste: someone looks at the dashboard after a week, sees that nothing improved —or that it got a little worse— and aborts, at exactly the worst possible moment, just before the investment begins to return something.
Sustaining it takes leadership courage —the same that the regional evidence in the previous section identifies as the true critical factor. And there's a reward, on the economics side, that justifies it with a number that isn't small. A 2019 study of 35,000 plants from the United States census found that structured management practices explained more than 20% of the productivity differences between plants: as much as or more than investment in R&D or technology. Those practices include goal-setting, data monitoring and problem-solving routines.
More than 20% — of the productivity variance between plants is explained by structured management, as much as or more than investment in R&D or technology. (American Economic Review, 2019)
A caveat is in order: that study measures "structured management" in a broad sense, not "problem-solving" as an exact label. But structured problem-solving —frame, verify, solve, close, with cadence and owner— is precisely the kind of practice that umbrella describes. It's no coincidence that the cheapest discipline to install is also one of the most ignored: we already saw it with the debrief in the close of this blog's series on teams —the cheapest learning routine is usually also the first to be canceled when the calendar tightens.
The leadership decision to make before starting: budget the worse-before-better in advance. Accept, in writing if need be, that for two or three months the numbers won't improve —and commit to not aborting in that stretch. Pull two people out of the fire and protect them from being thrown back in at the first hard week.
One last note, so as not to confuse speed with capability, about the fashion of the moment.
The AI mirage
In 2026 there are already AI-assisted root-cause-analysis systems that drastically cut the time to locate the cause in complex technical environments —microservice architectures, for example, where finding the origin of a failure can take hours of manual tracing through logs.
It's a real advance, but it shouldn't be confused with what this series solves. Those systems automate the analysis —Station 2, Verify—, not the definition of the problem —Station 1, Frame—. And far less Station 4, Close, which demands human judgment on whether the countermeasure held. A system that finds the root cause of a badly framed problem faster simply helps solve the wrong problem faster. Caution is warranted: it's 2026 research still without peer review, in the specific domain of software, not a consolidated result for operations in general. As a flavor of what's coming, it's useful. As a substitute for installed capability, it isn't.
With this on the table, two objections remain to dismantle.
The objections that stall the board
"We already did the continuous-improvement training, we have the belts." How many of those projects were reviewed afterward to see whether the problem came back? If the answer is none, the company has belts. It doesn't have a capability.
"I don't have people to pull from day-to-day and put on this." It's exactly the trap we described in the first installment of this series: there's never anyone available because prevention never happens, and prevention never happens because no one is available. The loop is broken by pulling two people out of the fire today, not by waiting for time to spare —time that, in a system that never prevents, will never be spare.
The committee that stopped meeting
Return, for a moment, to the Monday stockout committee this series opened with. Same Monday, same shortage spreadsheet, repeating for two years now. With the closed loop installed —a statement free of premature solutions, a cause verified in the warehouse, a countermeasure with an owner, and now, at last, someone who measures whether the problem comes back—, someone finally asks the question never asked before: why do those SKUs always run out, the same ones, month after month?
The cause turns out to be a reorder parameter wrongly entered in the system, inherited from a migration three years earlier. Not a demand problem, not a logistics problem, not "the supplier who doesn't deliver". It's fixed once, with an owner and a date. And the Monday committee, in time, stops having any reason to exist.
That committee that no longer meets is, exactly, the problem that doesn't come back. And no one will get praise for it, no promotion, no mention in the board meeting —because a problem that stopped happening leaves no trace to applaud. That, precisely, is the sign that the system worked.
Nine questions for Monday
A self-check that closes the three installments of this series. Before the next committee meeting, answer them honestly:
Nine questions for Monday
- Who did you promote last year: the one who put out the fire or the one who kept it from starting?
- Can you write the problem you have today in one line, without naming the solution?
- Did the last root cause you approved explain all the data, or only the most convenient?
- Does each open problem have an owner with a name and a closing date?
- Do you measure the recurrence rate, or only count closed problems?
- Do you have a dashboard of open, closed and recurred problems?
- Does continuous improvement in your company have a fixed cadence, or is it an annual event with catering?
- Does your leadership hold through the worse-before-better without aborting at the first weak week?
- Does anyone return, at three months, to check whether the countermeasure is still alive?
If you answered "no" to more than three, you don't have an installed capability. You have, at most, a photo archive with belts.
Did you pay for the workshop, take the photo, and the problem recurred?
We don't run workshops with photos or hand out belts: we install the loop —a statement, a verified cause, a countermeasure with an owner and the recurrence metric— with its cadence and its dashboard, and we stay until the problem stops coming back. Integrate, implement and sustain: installed capability, not an event with an expiry date.
References
- Jimmy Martin-Delgado et al., "How Much of Root Cause Analysis Translates into Improved Patient Safety: A Systematic Review", Medical Principles and Practice, 29(6), 2020.
- Mary Dixon-Woods et al. (Peerally, Carr, Waring), "The problem with root cause analysis", BMJ Quality & Safety, 26, 2017.
- Amazon, "Correction of Errors" — blameless postmortem culture; described in Colin Bryar & Bill Carr, Working Backwards, 2021.
- John Shook, Managing to Learn: Using the A3 Management Process, Lean Enterprise Institute, 2008.
- Toyota Argentina / SAMECO, "La mejora continua en el corazón de la mediana empresa" (kaizen at the Zárate plant); La Nación, 2024.
- "Lean Six Sigma en pymes latinoamericanas: factores críticos de éxito", Ciencia Latina, 2024; Scielo Chile, 2014.
- Nelson P. Repenning & John D. Sterman, "Nobody Ever Gets Credit for Fixing Problems That Never Happened", California Management Review, 43(4), 2001.
- Nicholas Bloom, Erik Brynjolfsson, John Van Reenen et al., "What Drives Differences in Management Practices?", American Economic Review, 109(5), 2019.
- (2026 frontier flavor) "KRCA: An Efficient Root Cause Analysis System in Hyper-Scale Microservice Systems via Agentic AI", arXiv:2607.01788, 2026 — non-peer-reviewed preprint.