Eight hundred quotes
Seven forty on a Tuesday in August. An electrical supplies distributor in Rosario: one hundred and forty people, three warehouses, forty years in the trade. The head of administration turns her screen so the owner can see a finished quote.
"This used to take us two hours. Now it takes twenty minutes."
It is true and it checks out. For a year now the sales team has been putting quotes together with artificial intelligence: it loads the line items, pulls the prices from the list, drafts the terms note. It works. Nobody is exaggerating and nobody is lying.
"And how many quotes went out last month?" he asks.
"Eight hundred and something. Same as always."
"And a year ago, how many went out?"
A short silence.
"Eight hundred and something too, I suppose."
They are both right. The tool works, the task takes a sixth of the time it used to, and a year later there is not one line of that company's P&L anyone can attribute to it. There is no way to argue about it either: nobody wrote down how things were before.
(The scene is a composite reconstruction, and so are the figures inside it. From here on, everything carrying a number is measured and has a source.)
The sources were read in the version available in August 2026.
The myth in this last part is the most expensive of the four the series has covered, because it is the only one that arrives with an invoice: adopting artificial intelligence means buying the tool and giving people access to it. The licences, the training, the vendor who configures it: that is the expensive part and it is the part that gets approved, because it is the part with a price and a date.
And it is the one part of all this that, on its own, changes nothing in the result.
Part 3 of this series answered one half of the question: the time saved at one step is spent, two desks further down, by whoever has to verify what came out. This is the other half, and it is cheaper to fix. Even when the work is not transferred to anyone, the time saved shows up nowhere if the flow around the task stayed the same: the quote still waits for the same signature, still goes out in the same Friday batch, and the bottleneck —which was never the writing— is still exactly where it was. Speeding up a step that was not the one holding things back speeds up nothing.
Which leaves the question: what did the companies where artificial intelligence does show up in the result do differently? There is an answer, and it comes with a label that has to be read first.
What they did differently
In March 2025, McKinsey published the round of its global AI survey that put twenty-five different things a company can do or fail to do when it brings AI inside to the test: who governs it, how much is invested, how people are trained, which data gets tidied up first. Of those twenty-five, the one most associated with AI showing up in operating profit is redesigning workflows.
There is one expression worth unpacking: the report talks about EBIT attributable to generative AI, which means how much of its operating profit the company attributes to the tool. The word doing all the work is attributes: it is declared by the same person answering the survey, and nobody audits it.
And redesigning workflows is also the thing fewest companies do. Among those reporting gen AI use by their organisations, 21% say they have fundamentally redesigned at least some workflows. One in five.
The next round of that survey, the one from November 2025, explains that Rosario morning better than any diagnosis. Almost nine in ten say they use artificial intelligence regularly in at least one function. Fewer than four in ten report any effect on operating profit, and most of those put it below 5% of their total operating profit. The tool reached everywhere. The result, almost nowhere.
Why this article quotes a consultancy after Part 2. Part 2 of this series left three questions to put to any figure before it sets an amount: what the primary source is, whether it is the published version or a draft, and who measured it. The figure in this section passes them only halfway, and that is worth saying before using it. The primary source exists and is public: it is the firm's own report, with its method stated. It is not a draft, but it has not been peer reviewed either: it is a survey of subscribers who choose to answer it, and the profit attributed to AI is declared by the same person answering. And whoever measured it sells AI implementations. With that, the rule from Part 2 applies on its own: this figure cannot set an amount. What it can do is point to where to look, which is what it is used for here — and the rest of this section rests on evidence of a different kind, which arrives shortly. A data point that cannot hold up a business case can still hold up a question.
Of everything a company can do, which one came first and how many did it:
What “redesigning the flow” actually means
Put like that it means nothing: it is one of those phrases you have heard a hundred times without it ever telling you what to do on Monday morning. It does, however, have a short definition that can be verified in an afternoon.
A redesigned flow is one where who does what, in what order, what stopped being done and who signs at the end have all changed. All four changes, not one.
That is where the test comes from, and it is the only thing worth remembering from this section:
If nobody stopped doing something, there was no redesign: there was one more tool.
That is the test for two reasons. The first is that it cannot be faked: a company can say it redesigned the work and show a new diagram, but it cannot point to the step that disappeared if the step is still there. The second is that a step that disappears is the only thing that genuinely frees up capacity. As long as the quote still waits for the same signature before it goes out, it makes no difference how long it took to write: the flow delivers exactly what it delivered before, with less busy people inside it.
And now the warning that avoids the disaster, because removing a control step without having tested the task is worse than doing nothing. The order of what follows is the order in which it is done: first you find out which side the task falls on —which side of the limit of what the tool does well—, then you measure how things were going, and only then do you take the step out. The redesign arrives at that point, and not before.
Three common tasks at a mid-sized company, with and without redesign:
| The task | If only the tool was bought | If the flow was redesigned | How you know which of the two happened |
|---|---|---|---|
| Putting a quote together | The salesperson writes it in twenty minutes instead of two hours and sends it to admin, which reviews the whole thing exactly as before and signs it on Thursday | The quote goes out signed by the salesperson; admin reviews one in five as a sample, against a written criterion, and is no longer in the way of the other four | A wait disappeared. The full-review step is no longer in the procedure |
| The monthly management report | The same fourteen pages, now in half the time, for the same four people who skim them | The report became one page with the decision that has to be made, and the fourteen pages became an annex opened only if somebody asks — or stopped being issued at all | Somebody stopped writing something, and the meeting that used it changed length |
| Setting up new products in the system | Somebody asks the tool to build the product record and then re-types it into the system to be sure | The record goes straight in and review is by exception: only the cases the tool itself flagged as doubtful, plus a fixed sample | Re-typing no longer exists in the work instruction, and there is a written rule for what always gets reviewed |
In all three rows the proof is the same: somebody stopped doing a step and there is a new signature. If you cannot point to both, you bought a tool. 32sur schematic: none of the three rows is a measured case —they are three flows common at a mid-sized company, written to show what changes and what does not— and the sampling percentages are illustrative: each company sets its own from the result of its own test (see Part 1 of this series).
Before redesigning, you need to know which task holds up
The first part of this series showed that the limit of what the tool does well cannot be deduced: two tasks that look equally demanding to a professional can fall on opposite sides, and the one that fails does not come marked. The authors gave that shape a name: the jagged frontier. The consequence for this method is a single one: the unit of work is the task, not the department and not the company. Anyone who redesigns the flow of an entire department because the tool worked well on one task is betting that the tasks next to it are the same, and that is the error Part 1 measured.
Locating a task costs an afternoon and nothing more. You test it on already-solved cases, with the right answer written down before anything is opened, and you count two things: how often it got it right and how often it got it wrong while sounding sure. The second one decides, because it is the one your people will not catch.
In Competing in turbulence we argued that a bet on new ground is tested small and with the exit conditions written before starting. It is the same here, with one difference that matters: the criterion is not written to find out whether the bet paid off, it is written so you can see the error that sounds sure of itself.
The number your company does not have
Suppose the task held up under the test. What is still missing is the number that lets you say, a year later, whether it was worth it: how things were going, written down before anything was touched. It is called a baseline.
And it is not written in hours saved, but in the output of the flow: how many quotes went out per month, in how many days from the request, how many had to be corrected afterwards, how many were won. Hours are no use, and Part 2 of this series showed why: they are precisely what people estimate badly, even about their own work.
If you rolled the tool out eight months ago and never wrote any of that down, the good news is concrete: the baseline still exists, dated, inside your own systems. The previous twelve months are in the management system, in the mailbox, in the quotes folder. Rebuilding it costs an afternoon.
There is a second thing that also costs no money. The best field study there is on this —the one on the support agents at a software company, the same one from Part 1— was able to measure what it measured for an operational reason: the tool was opened site by site and one agent at a time, and that difference in calendar made it possible to compare each person with themselves before and after. There is a method being given away there: one slice of the company that starts six weeks later. That is a comparison group, and no consultancy sells you those six weeks.
What is still missing is the underlying reason to do it in-house. There is a meta-analysis that puts the field in order: it pools the results of another twenty-three studies and finds a real, modest effect on productivity, 0.33 on the standardised scale used to compare studies that measured different things —it is not a percentage: there, 0.2 counts as a small effect and 0.5 as a medium one—. Two caveats before using it: the twenty-three studies are about programming and not about quotes, and the meta-analysis has not been peer reviewed yet. But the number that matters is a different one: the range where that average may sit, from 0.09 to 0.58 —from almost nothing to quite a lot—. The meta-analysis itself explains the width: the more the measurement resembles a real company, the smaller the number.
The average of other companies is not a forecast for yours. It is barely proof that your number exists and that somebody has to go and find it.
Where that effect may sit, and where yours sits:
Adoption is not requested: it is designed
This is still the same station of the method: the rule is the other half of the redesign, and without it the redesign cannot even be written down. Everything above assumes something no company has guaranteed: that its people use the tool and say when they used it.
Declaring that you used artificial intelligence is not free for whoever declares it, and Part 3 of this series shows the experiment where you can see it: the same piece of work got a lower competence rating when its author said they had used the tool. Whoever paid that cost once does not say so the next time. Few people connect the consequence: use that is not declared leaves no trace, and a flow nobody knows the workings of cannot be redesigned.
And asking for it in a meeting is not enough. A Danish study —a draft, still without peer review, that crossed two survey rounds covering some 25,000 workers with administrative data on earnings and hours— found something that puts any adoption plan in order: encouraging use, rolling out the corporate version of the tool and running training are associated with better outcomes when they go together; done separately, the corporate version and the training are associated with lower reported benefits. Half an initiative is associated with less reported benefit, not more.
The study does not say why; it offers more than one explanation and does not choose between them, and this article is not going to invent the missing one. What does fit with everything above is that training and rolling out without touching the flow adds steps —learning the tool, reviewing what comes out— and removes none.
Not even the companies that push manage all three: among those that encourage use, only 29 out of 100 rolled out the corporate version and ran training. The other end is the finding that convinces: where all three coexist, 93 out of 100 workers say they have used the tool, and more than one in four use it every day. No amount of exhortation produces that level of use.
What replaces exhortation is a one-page usage rule. It is not an information security policy and it is not a code of ethics: it is four lines, none longer than a single line.
The one page: four lines, none longer than a single line
- When use is declared —and, above all, what does not happen when it is, which is the half that makes it work.
- What always gets reviewed and what gets reviewed by sampling.
- Who signs what goes out.
- What does not go into the tool: customer data, price lists, personnel files.
If it takes you more than one page, it is not a rule: it is a regulation, and nobody is going to read it.
One more rule is missing, and it comes from Part 1: asking the tool to review its own work scales the persuasion, not the truth. Verification happens against the criterion written beforehand, and never inside the conversation with the tool. A review that starts after you have read the answer is not a verification: it is an argument, and on the other side there is something better equipped to win it.
The full loop
Put in order, the five things this article talks about —locating the task, testing it small, writing the baseline, redesigning and setting the rule, scaling or stopping— are a loop: you come in through a task and come out with a written decision, which is how the next one comes in. A task is not adopted until it has gone round the full loop, and almost every company does the first quarter and calls that adoption.
The loop ends in scaling or in stopping, and both are legitimate outcomes. Stopping a task after testing it is the cheap way of finding out it was on the wrong side of the limit: it costs an afternoon and not an entire rollout. What almost no company has is the record of the tasks it tested and discarded, and on what criterion it discarded them; that is the asset that makes the second loop cost half.
And then the scale, which is where you decide whether this is executable at all. A full loop over one task is weeks and not quarters: an afternoon for the criterion and the test, an afternoon for the baseline, six weeks of rollout in batches, one page of rule. The expensive part was never the time. The expensive part is having four pilots open at once and being unable to decide on any of them.
The five stations, with the question each one answers and the paper it produces:
The four questions of the system
The four parts of this series were not four subjects. They were four questions put to the same task, in order, and failing just one is enough for the result never to appear.
And they are put to a task. None of the four questions of the series can be answered about “artificial intelligence in the company”: all four are answered about “the task of putting quotes together”, which is the unit this series has worked with since its first line. That, incidentally, is why an AI committee never answers them: a committee deliberates at company level, and all four only have an answer at task level.
The order is not decorative either. If the task falls on the wrong side of the limit, measuring it well is of no use. If the number that excites everyone was measured by somebody else, at another company, on other work, your own baseline never gets written. If nobody declares when they use the tool, there is nothing to review. And if the flow stayed the same, the three previous answers stay on the desk of whoever answered them.
All four, with what it takes to answer each one:
The method is the barrier, and the method can be learned
The two numbers that opened Parts 1 and 2 come from the same survey and have never been read together. The first: 41.6% of the Argentine SMEs surveyed by CEPE —the public policy centre at Universidad Torcuato Di Tella— and Fundar already use at least one artificial intelligence technology. The second: the barrier they report most, above cost, is the lack of experience and knowledge about the subject, 46.2%.
Read them one after the other and the conclusion assembles itself. What those companies say they are missing is on no price list. It is the method, and the method is the only thing in everything this series covered that can be learned.
The other half is a matter of calendar. A full loop over a single task is weeks, and nobody finishes it without having started it. Whoever sits down this month to reconstruct how many went out per month before the rollout is in time. Whoever leaves it for next year will have the same Rosario conversation, one year more expensive.
The tool reached everywhere; the design of the work, almost nowhere:
The three figures come from three different kinds of source and describe the same thing from three places: the tool came in, and the work stayed where it was.
Questions for Monday
The four questions on the card are put to a task; these seven you put to your company. The first three can be answered right now, without opening anything and without asking anyone.
Three you can answer now, four that take longer
- The task where artificial intelligence is used most in your company: can anyone say how many went out per month before the tool arrived? If the answer is “roughly the same”, that is the answer, and you already know what is missing.
- Since it has been in use, has anyone stopped doing a step? Name it out loud. If you cannot name it, the flow is exactly where it was a year ago.
- Who signs today what comes out of that task? Is it the same person who signed before?
The next four take longer. None of them requires buying anything:
- Ask three people on your team whether they say when they used the tool, and what happens to them when they do. Ask them separately and not in a meeting. The answer tells you whether you have a data problem or not.
- Go into your systems for the twelve months before the rollout and write down, on one line, the number for how things were going. The baseline was not lost: it is there, dated.
- If you still have somewhere left to roll out to —a branch, a shift, a warehouse— do not do it all at once: start with one and leave the rest for six weeks later. That difference in calendar is your comparison group, and it is free.
- Write the one page: when use is declared, what always gets reviewed and what gets sampled, who signs, and what does not go into the tool. One line each. If it takes you more than one page, it is not a rule: it is a regulation, and nobody is going to read it.
Go back to the office in Rosario, seven forty. Nothing that company did was wrong: it picked a reasonable task, bought a tool that works, and its people use it every day without anyone having to insist. What is missing fits in two lines. Nobody wrote down, a year ago, how many quotes went out per month and in how many days. And nobody took a step out: the quote gets written in twenty minutes and still waits for the same signature until Thursday.
With those two lines written down, this morning's conversation would have lasted forty seconds and would have ended in a decision. Without them it can last another year.
Four parts, four questions, one task. The series was not about artificial intelligence: it was about how decisions get made when the number that reaches the table was measured by somebody else, at another company, on other work. The tool will keep changing every six months, and the four questions will keep being the same.
Can you say today how many went out per month before the tool arrived?
Ninety days in, this is what can be written down about one task in your company: which side of the frontier it fell on, how much went out before —dated, taken from your own records—, which step stopped being done, who signs now and what always gets reviewed. That is the unit: one task. Once the first one is done, the second costs half, and your company's AI discussion stops being about tools and becomes about which task is next. The first step is a two-hour session to choose it and write the baseline from the records you already have. That is what 32sur works with. Let's talk.
References
- McKinsey & Company, "The state of AI: How organizations are rewiring to capture value", Global Survey on AI, March 2025; and the following round of the same survey, November 2025 (around 1,993 respondents across some 105 countries) — the source of the finding that opens this article. Of 25 organisational attributes tested, the one showing the largest effect on EBIT attributable to generative AI is the redesign of workflows, and 21% of those reporting gen AI use at their organisation say they have fundamentally redesigned at least some workflows (verbatim: "Twenty-one percent of respondents reporting gen AI use by their organizations say their organizations have fundamentally redesigned at least some workflows"). In the November 2025 round: 88% use AI regularly in at least one function, around 39% report some EBIT impact and most of those put it below 5%. Standard warning, which this article makes in the body and not in this line: it is a consultancy survey with a self-selected sample of subscribers, the profit impact is declared by the respondent, it has not been peer reviewed and it is published by a firm that sells AI implementations. It is not used here as a basis for calculation. The report does not publish a full ranking of the 25 attributes.
- Humlum, A. and Vestergaard, E., "Still Waters, Rapid Currents: Early Labor Market Transformation under Generative AI", University of Chicago Booth / NBER Working Paper 33777, May 2025, revised version of March 2026 — the evidence on adoption design, read in the version in force: two survey rounds covering some 25,000 workers at 7,000 workplaces in 11 exposed occupations, linked to Danish administrative data on earnings, hours and occupation. Employer policies —encouraging use, rolling out the corporate version and running training— are associated with better outcomes when they go together; verbatim: "enterprise chatbots and training—when implemented in isolation—are associated with lower reported benefits". The paper itself offers more than one explanation for that pattern and does not choose between them; the reading this article adds —that half an initiative adds steps without removing any— belongs to this article and not to the study, and it is stated as such in the body. Among the companies that encourage use, 61% have a corporate version, 39% run training and 29% do both; where all three initiatives coexist, 93% of workers say they have used the tool and 28% use it daily. It is a draft, still without peer review, and what it reports are associations within a survey, not an experiment: this article uses the verb "is associated with" and no other. Version warning: the same work circulated in 2025 under the title "Large Language Models, Small Labor Market Effects" —the NBER page says so: "This paper was previously circulated under the title 'Large Language Models, Small Labor Market Effects'"— and the March 2026 revision changed figures and analysis: it no longer reports the 38% of employers with a corporate version or the 30% of employees trained out of the total, which are now conditional on the companies that encourage use, and it does not carry the gender gap in adoption that the July 2025 version did. The version in force is the one cited here and every figure was taken from it. The time-saving findings and the effects on earnings and hours from this same study are the subject of Part 3 of this series and are not used here.
- Dell'Acqua, F., McFowland III, E., Mollick, E., Lifshitz, H., Kellogg, K., Rajendran, S., Krayer, L., Candelon, F. and Lakhani, K., "Navigating the Jagged Technological Frontier", Organization Science, 37(2), 2026, pp. 403-423, DOI 10.1287/orsc.2025.21838 — the pre-registered experiment with 758 Boston Consulting Group consultants that holds up the first station of the method: the condition separating the tasks where the tool improved the result from the one where it made it worse is that a verifiable answer existed. The figures from that experiment are in Part 1 of this series and are not repeated here.
- Brynjolfsson, E., Li, D. and Raymond, L., "Generative AI at Work", The Quarterly Journal of Economics, 140(2), 2025, pp. 889-942 — cited in this article for its method and not for its results: access to the assistant was opened site by site and one agent at a time between 2020 and 2021, and that difference in calendar is what made it possible to compare each agent with themselves before and after, and against the agents not yet treated. The study does not propose staggered rollout as a management method; this article recommends it as one because it describes something a company can do at no cost. Its figures are in Parts 1 and 2 of this series.
- Maier, S., Gunzenhäuser, M., Schweisthal, J., Schneider, M. and Feuerriegel, S., "A meta-analysis of the effect of generative AI on productivity and learning in programming", LMU Munich / Munich Center for Machine Learning, arXiv:2605.04779, 2026 — 23 studies and 27 effect sizes, systematic search and formal risk-of-bias assessment. Effect on productivity: Hedges' g = 0.33, with a 95% confidence interval from 0.09 to 0.58 and substantial heterogeneity; the gains are larger in controlled experiments and smaller in open-source and corporate contexts. It is a preprint: it has not been peer reviewed, and it is about programming.
- Randazzo, S., Joshi, A., Kellogg, K., Lifshitz, H., Dell'Acqua, F. and Lakhani, K., "GenAI as a Power Persuader: How Professionals Get Persuasion Bombed When They Attempt to Validate LLMs", Harvard Business School Working Paper 26-021, October 2025 — the basis for the rule of verifying outside the tool: when professionals tried to validate the output, the model escalated the intensity of its persuasion instead of revealing its limits. It is a working paper, not peer reviewed. The available documentation does not report how often professionals ended up accepting incorrect answers, and this article does not claim it. The full development is in Part 1 of this series.
- CEPE-Universidad Torcuato Di Tella (the public policy centre at its School of Government) and Fundar, nadIA initiative, with fieldwork by Fundación Observatorio PyME and support from the IDB, "Encuesta Nacional sobre Adopción de IA en Pequeñas y Medianas Empresas de la Argentina" (national survey on AI adoption in Argentine SMEs), April 2026 — 402 companies with 10 to 249 employees in manufacturing and in software and IT services, fieldwork between November and December 2025, stratified sampling across 14 sector strata. 41.6% use at least one AI technology; the most frequently reported barrier is the lack of experience and knowledge about the subject (46.2%). The two figures this article uses are the same ones used in Parts 1 and 2 of the series, cross-checked against three independent pieces of coverage of the same report. The rest of the survey data is not quoted here, and neither is the figure on budget devoted to AI: the available coverage does not agree with itself.
All sources were read in their primary version available as of August 2026.