Services Method Training Firm Work with us FAQ Insights Contact
ESENPT

Home  ›  Insights  ›  23

Article 23 · AI without the hype

Task number 19: the one won by the people who had no artificial intelligence

758 consultants at a global firm, the groups drawn at random and the study registered before it began: it is the best field experiment we have today on artificial intelligence and knowledge work, and its favourable results are genuinely large. Here they are in full, and first, in the version that went through peer review. What almost never gets told is that the experiment had two halves, and that in the second one the sign flipped. This piece sets out what exactly was measured, why the frontier is not where you would imagine it, and what to do on Monday with the first task you want to test.

By 32sur · October 2026 · Reading time: 16 minutes · “AI Without the Hype” series, Part 1 of 4

Nobody asked which task

Twenty past four on a Wednesday. The boardroom of a three-hundred-person food company with a plant in Rafaela. The IT manager finishes the third slide, the one called "Use cases", which has eleven of them. The chairman cuts in:

"And what does this improve for us, exactly?"

The sales director answers without looking up from his phone:

"Harvard measured forty percent more quality with artificial intelligence. Have a look, it's in the report I sent you."

The chairman nods. Somebody writes it down. The meeting moves on: licences, a budget, a three-month pilot in marketing because that is where "there is least risk". Nobody asks which task that forty percent was measured on. On the next slide the number is already gone, and it does not need to be there: it is left floating in the room, which is where it does its work.

(The scene is a composite reconstruction. From here on, everything carrying a number is measured and has a source.)

The experiment is real and it is good, and that forty percent came out of it. What that room did not know is that it had two halves.

The myth holding the scene up is not "AI works" or "AI doesn't work". It is smaller and more expensive: the idea that there is one improvement percentage, a single one, that can be quoted from memory and applied to any kind of work. There isn't. The effect depends on which side of a frontier each task falls — a frontier nobody can see, and one that is not where you would imagine it.

One warning before going on: the measurements in this piece were taken on consultants at a global firm, support agents at a software company, and programmers. None of them on an Argentine company. What travels is the shape of the frontier, not the height of the number.

Eighteen tasks

The number left floating in that room in Rafaela came out of that experiment: Navigating the Jagged Technological Frontier, nine researchers from Harvard, MIT, Wharton, Warwick and Boston Consulting Group, published in Organization Science in 2026. Its favourable half is almost never told in full, so here it is in full, and first.

The 758 consultants were randomized into two halves that never crossed. One of them —385 people— got eighteen product tasks: come up with a sneaker, segment the footwear market by type of user, draft a press release, write a memo that would persuade a team. The other —373— got something else, and we come back to it in a moment. Inside each half there were three groups, also drawn at random: one worked as usual, another with GPT-4, and the third with GPT-4 plus a short training session.

The numbers people quote come from that first half. There are three of them.

First, how much they finished. The group without artificial intelligence completed 82% of what it was asked for; the groups with the tool, 91% and 93%. Separately —and this is not the subtraction of those percentages but the estimated effect— the paper puts it at 12.2% more tasks completed: in a department that ships forty pieces of work a month, that is five more that land.

Second, how long they took. The groups with the tool worked 25.1% faster. A four-hour task gets finished in three: one working day a week, every week.

Third, and this is the number quoted worst of all, how much what they delivered was worth. Human graders scored every deliverable on a common scale: the group without the tool averaged 4.37, and the two AI groups came in 1.31 and 1.48 points above that, between 29.9% and 33.9% more quality. The body of the paper sums it up as more than 30%.

These are large improvements, well measured, on real work, with the hypotheses written down before the start. Anyone who says today that generative artificial intelligence does not move knowledge work is arguing with the evidence, not with the marketing.

Eighteen tasks. Counted in full, the experiment had nineteen.

Why this piece says 30% and not 40%. The figure quoted in boardrooms comes from the 2023 working draft of this same study, whose abstract spoke of "more than 40% quality" against the control group. The peer-reviewed version —Organization Science, 2026— took that figure out of the abstract, and the body of the paper reports "more than 30%": 29.9% and 33.9% for the two AI groups. None of this is a scandal. It is what peer review does, and it is the reason peer review exists. The problem is the calendar: two and a half years passed between the draft and the published version, and a corporate conversation about AI lasts three weeks. What reaches the room is always the draft. Part 2 of this series has four more numbers with the same problem, and one of them gets blamed for a fall in the Nasdaq.

And one

The other half —373 consultants drawn at random from the same pool, comparable on everything measurable— never saw those eighteen tasks. It got a single one, a business task, designed to fall on the other side of the frontier. Nobody told them: they received it as the day's work, with nothing to mark it out.

It had one property the eighteen did not: an objectively correct answer. It was a recommendation built on quantitative data and interview notes, and afterwards you could say, without argument, whether the consultant had got it right.

Those who worked without artificial intelligence got it right 84.5% of the time; those with GPT-4, 70.6%; those who had also had the training, 60%.

Combining the two groups with artificial intelligence, the average drop is 19 percentage points, and both words are doing work. Going from 84.5% to a little over 65% is 19 points. Saying "19% less" would describe a fall to 68.4%, which is a different thing. Confusing the two is the most common error in this territory, and the abstract of the published version commits it: it says "19% less" where the body of the same paper says points. The arithmetic reads itself from the side that hurts: without the tool, one and a half consultants in every ten got it wrong; with the tool, three and a half. The errors more than doubled.

The difference between the two groups with the tool does not reach the confidence level usually demanded of a result (5%): it holds only at 10%. So the honest reading is short: the training did not put them on the good side.

The consultants failed for a reason more uncomfortable than overconfidence or laziness: nothing flagged that task as different. Nothing in the brief, and —this is the expensive part— nothing in the answer either.

There is a measurement of that, in the same half of the experiment: the graders also scored the coherence and the argumentation of what those 373 delivered, and there the AI groups came in between 18% and 25% above the group without the tool. The same people, the same task, the same afternoon: better written and correct less often. The error did not come wrapped in worse writing. It came wrapped in better writing.

The two halves of the same experiment, side by side:

Two halves of the same experiment, opposite results by task 758 consultants randomized into two arms that never crossed. THE INSIDE HALF 385 consultants · 18 tasks TASKS COMPLETED WITHOUT AI 82% WITH AI 91% WITH AI AND TRAINING 93% 0% 100% · 25.1% faster · quality: between 29.9% and 33.9% above the group without AI, according to human graders THE OUTSIDE HALF 373 consultants · task 19 CORRECT ANSWERS WITHOUT AI 84.5% WITH AI 70.6% WITH AI AND TRAINING 60% 0% 100% 19 percentage points below the group without AI Nobody did both halves. None of the 373 knew their task sat on the other side of the frontier. Dell’Acqua, F. et al., Navigating the Jagged Technological Frontier, Organization Science, 37(2), 2026, pp. 403-423. Preregistered experiment; 758 BCG consultants randomized into two non-overlapping arms: 385 inside the frontier and 373 outside. The percentages in the left panel are observed completion rates; the paper separately estimates an increase of more than 12% in tasks completed, and the two are not the same subtraction. The gap between 70.6% and 60% does not clear the 5% bar (Table 7): training did not put that group on the good side, and the experiment does not show that it made them worse.
Figure 1 — Two halves of the same experiment. The same firm, the same tool, the same weeks: the sign of the result was decided by the task. None of the 373 knew theirs sat on the other side of the frontier. Source: Dell'Acqua et al., Organization Science, 37(2), 2026 (758 BCG consultants, preregistered experiment).

The jagged frontier

The first thing anyone does on reading that result is form a reasonable hypothesis: the tool works on the easy stuff and fails on the hard stuff. It is the wrong hypothesis, and it is an expensive one: whoever adopts it believes they now have a rule.

Several of the eighteen were demanding: segmenting a market by type of user or writing a memo that persuades a team is not paperwork. And task 19 was not the hardest: it was a business recommendation with spreadsheets and interview notes on the table, the kind a consultant delivers every week.

The frontier exists, it is real and it can be measured. What cannot be done is deduce it. The difficulty a professional attributes to a task does not sort tasks onto one side or the other: two that look equally demanding can fall on opposite sides, and the one sitting next to the task that went well may be the one that fails. The authors gave that shape a name: the jagged frontier. It is not a smooth slope from easy to impossible; it has teeth, and the teeth cannot be seen from outside.

Two drawings: the frontier you imagine and the one the experiment describes.

The frontier you imagine is a ramp. The one they measured has teeth. Conceptual diagram — not a data chart. No position on the axis corresponds to a measurement. A · THE FRONTIER YOU IMAGINE orderly, predictable, and wrong here AI helps here AI fails B · THE ONE THEY MEASURED it has teeth, and the teeth are invisible from outside task 19 same apparent difficulty, opposite side OUTSIDE: AI gets it wrong, and confidently the 18 tasks inside INSIDE: AI improves the outcome HOW HARD THE TASK LOOKS TO A PROFESSIONAL from the one that looks trivial… …to the one that looks impossible Your company’s tasks are not on this diagram. Locating the first ones is the job of the worksheet that closes this piece. 32sur conceptual diagram explaining the jagged technological frontier described in Dell’Acqua, F. et al., Organization Science, 37(2), 2026. It is not a measurement: the study publishes no task-by-task map, and that absence is exactly the problem the figure illustrates. The horizontal positions are illustrative.
Figure 2 — The frontier you imagine and the one they measured. The difficulty you perceive does not predict which side a task falls on. A 32sur conceptual diagram of the jagged technological frontier in Dell'Acqua et al. (Organization Science, 2026): it is not a measurement, and the horizontal positions are illustrative.

In The terrain sets the rules we separated complicated problems from complex ones by how cause and effect behave. The frontier in this experiment is a different thing: it does not classify the problem, it classifies what the tool is capable of.

If the frontier cannot be deduced, the useful question stops being how much artificial intelligence improves things and becomes which side this task falls on. That is answered by testing. First there is one hope to rule out, the one every owner forms without saying it out loud: that their best people can see the frontier.

What the tool does to your best people

The hope is sensible and there is a measurement against it: 5,172 customer support agents at a large software company, followed month by month as they were given access to an assistant built on GPT-3 and rolled out between 2020 and 2021, before ChatGPT existed. Good field evidence is always older than the tool being sold.

The rollout was not uniform, and that is what gives the result its weight: the tool arrived office by office and agent by agent, and that difference in timing makes it possible to isolate the effect by comparing each agent with himself, before and after.

On average, agents resolved 15% more issues per hour. For every twenty queries an agent closed in a day, he closed three more.

That average describes almost nobody. The lowest-skill quintile —the 20% who had been performing worst— gained half an issue per hour, 36%. The most qualified showed no significant change.

At 32sur we call this the skill inversion: the tool bought to supercharge your best people ends up raising the floor for the ones who have just walked in, and moves nothing for the best.

With one caveat, and it comes from the nineteen-task experiment itself: inside the frontier, those who had been below average gained considerably more than those above, but those above gained appreciably too. (The exact percentages appear only in the 2023 draft, so they are not quoted here.) How much is left for your best people depends on the work: on open-ended tasks everyone gains and the bottom gains most; at the counter —the support agents' work—, where a correct answer exists and repeats, the ceiling did not move.

The place where it pays first is the counter: the practical consequence is who gets the access, and in what order.

By level of prior skill, the effect sits at one end only.

AI did not improve the best: it lifted the ones just walking in Change in issues resolved per hour, by quintile of prior skill. CHANGE IN ISSUES RESOLVED PER HOUR +40% +30% +20% +10% 0% +36% the study publishes no numeric value for the middle quintiles the slope is downward no significant change THE AVERAGE OF THE 5,172 AGENTS: +15% neither extreme looks like this number QUINTILE 1 QUINTILE 2 QUINTILE 3 QUINTILE 4 QUINTILE 5 the lowest the highest half an issue more per hour somewhat more speed and somewhat less quality, according to the paper’s own abstract PRIOR SKILL OF THE AGENT — from the lowest quintile to the highest Brynjolfsson, E., Li, D. and Raymond, L., Generative AI at Work, The Quarterly Journal of Economics, 140(2), 2025, pp. 889-942. 5,172 customer support agents; access was opened office by office and one agent at a time, and that difference in timing is what identifies the effect. The grey bars have no published value behind them: the convention in this house is not to draw what was not measured.
Figure 3 — AI did not improve the best. The average of a whole company describes nobody in particular inside that company. The grey bars have no published value behind them. Source: Brynjolfsson, Li and Raymond, The Quarterly Journal of Economics, 140(2), 2025 (5,172 customer support agents).

If the effect varies that much between people at the same company in the same month, the next question asks itself.

There is no average

The effect varies between companies too. A 2026 meta-analysis —a preprint, not peer-reviewed, and on programming— pooled 23 studies and 27 effect sizes.

Before the figure, the translation. When evidence from different studies is pooled, the result is expressed on a common scale called an effect size: how much the outcome moved compared with how much it already varied among the people measured. Zero is "nothing happened", 0.2 is small and 0.5 medium. Productivity came out at 0.33, a real, intermediate improvement, but with a reasonable range running from 0.09 to 0.58: from almost nothing to quite a lot. The average is correctly computed; the companies measured do not resemble one another, and the meta-analysis says where the difference lies: the gains are larger in controlled experiments and smaller in open-source and enterprise settings. The more the measurement resembles a real company, the smaller the number.

On learning, the effect cannot be told apart from zero with the data available: 0.14, with a range that crosses zero.

Three pairs of figures, and none of them is "the percentage AI gives you":

84.5 and 60percent correct answers on task 19: first the consultants without artificial intelligence, then the ones who had it and had also been trained to use it (Dell'Acqua et al., Organization Science, 2026)
+36% and 0what the lowest-skill quintile of 5,172 support agents gained, and what the most qualified showed, with the same tool and inside the same company (Brynjolfsson, Li and Raymond, QJE, 2025)
0.09–0.58the two ends of the range where the effect on productivity falls in a meta-analysis of 23 programming studies: from almost nothing to quite a lot, depending on the setting measured (Maier et al., preprint, 2026)

All three are serious measurements and not one of the three can be quoted as "what AI is going to give your company". That number is published nowhere, because it depends on which tasks you have inside.

“Have a person review it”

It is the objection the reader has been forming since task 19, and it is the right one: before anything goes out, have someone who knows look at it.

The scene is familiar to anyone who has used the tool seriously. You flag an error —the number in row 12 is not that one— and back comes an apology, the corrected row, and the same conclusion again, with two more arguments. By the fourth round trip you are no longer verifying: you are arguing, and you cannot remember what your original objection was.

There is a measurement of that, and it is uncomfortable. A team at Harvard Business School —some of them the same authors as the nineteen-task experiment— put more than 70 consultants from the same firm to work with a model on a task built to be hard for it, and watched what happened when the professional tried to validate the output: check the data, point out an error, ask it to reconsider. It is a working paper, not peer-reviewed.

What they found is that the model did not reveal its limit: it escalated the persuasion. It apologised, corrected the detail flagged, and went back to holding its original position, now with more supporting data. The authors catalogued 14 tactics, grouped into the three classical families of rhetoric: your own credibility, logic and data, and the person in front of you. They called it persuasion bombing.

In the series on deciding under uncertainty we saw that a simple formula ties or beats expert judgement. Here the problem is different: a formula does not defend itself, and a language model does.

Reviewing works. But a review that begins after you have read the answer is not a verification: it is an argument, and the model is better equipped to win it. What turns it into verification is having written down, beforehand, what the correct answer was.

Four in ten have already bought the tool

The four studies above were measured abroad. From Argentina there is a recent snapshot: CEPE —the public policy centre of the School of Government at Universidad Torcuato Di Tella— and Fundar, with fieldwork by Fundación Observatorio PyME, surveyed 402 companies of 10 to 249 employees between November and December 2025.

41.6% already use at least one artificial intelligence technology. The debate about whether to adopt is over, and it settled itself, without ever passing through the board.

The three numbers that follow are about the companies already using it. 54.7% apply it in marketing and sales. 57% do not formally measure the impact. And 72.5% have no usage guideline at all, not even an informal one: 3.3% have a written policy.

The tool is already there. What is missing is the judgement to place it: the picture of a company that has bought and has not yet tested anything in the way testing is done.

And testing, unlike almost everything sold in this field, does not cost money.

What to do with a task you have never tested

The generic map of the frontier does not exist and never will: it depends on which tasks each company has inside. Locating the first stretch of yours, on the other hand, costs an afternoon.

One: write the correct answer before opening anything. Pick a single task, gather cases already resolved and write down, for each one, what getting it right would have been. It is the same move as in a well-run interview, where the scorecard is written before you meet the candidate: afterwards, the criteria quietly rearrange themselves around whatever came out.

And if the task you care about has no verifiable answer —the case of almost everything done in marketing—, you do not test the task: you test the piece of it that does have one. Of a campaign, nobody will be able to say afterwards whether the copy was the right copy; of a list of two hundred contacts you can say which ones were misclassified, and of a piece of copy, whether the figure it quotes is in the report it came from. That piece gets tested first.

Two: test small, and on the known. You run the tool over those old cases and count two things: how many times it got it right and how many times it got it wrong with confidence. The second is the one that matters, because it is the one your people will not catch.

Three: look at the errors one by one. The useful question is whether they are detectable by the person who is going to receive the work. 90% right with invisible errors is worse than 70% right with obvious ones.

Four: decide one task at a time. The one that passed goes in. The one next door —even if it looks the same to you, and above all if it looks the same— goes back to step one. That is the entire operating content of the jagged frontier.

Who does each step, against what threshold, and a worksheet already started:

A new task is not trusted: it is tested small, with the criterion written first Four steps, who does each one, and against what threshold. STEP WHO DOES IT THE THRESHOLD 1 · WRITE THE CORRECT ANSWER Whoever receives the work, not whoever produces it 10 to 20 already-solved cases from one task, with the criterion written before seeing any output 2 · TEST SMALL One person, one afternoon The same cases, and two counts: hits and errors stated with confidence 3 · LOOK AT THE ERRORS ONE BY ONE Whoever will receive that work every day One invisible error is enough to send the task back to step 1 4 · DECIDE ONE TASK AT A TIME Whoever pays the cost of verifying The task next door starts again at step 1, even if it looks like the same task A WORKSHEET STARTED Task — draft the reply to a quality complaint. Cases — the twelve complaints closed in 2025. Criterion, written before opening anything — does it admit responsibility where it was due?, does it offer the compensation that ended up being paid?, does it cite the right standard? Result — ___ hits · ___ errors stated with confidence. The two empty boxes are the afternoon’s work. WHAT THIS PROCESS REPLACES the general AI pilot with no defined task · the “let’s try it and see” with no written criterion · prompt training as a substitute for the test · the human review that begins after reading the answer · extrapolating from the task that worked to the one next door. A 32sur diagram. Step 1 translates the condition that separates the 18 tasks from task 19 in the experiment of Dell’Acqua et al. (Organization Science, 2026): that a verifiable answer exists. Step 3 translates the finding of Randazzo et al. (Harvard Business School, working paper 26-021, 2025) about what happens when verification starts late.
Figure 4 — The afternoon's worksheet. Ten old folders and one afternoon. Trusting without having tested comes out a good deal more expensive. A 32sur diagram: step 1 translates the condition that separates the 18 tasks from task 19 (Dell'Acqua et al., 2026) and step 3, the finding of Randazzo et al. (Harvard Business School, working paper 26-021, 2025).

What will happen when you propose it

The moment you propose the first of those steps, four sentences will show up, and all four are reasonable: each one rests on something that did in fact happen.

What will be saidWhat it runs intoWhat still stands
"We already tried it and it works great"On 18 real consulting tasks the tool improved everything that was measured; on task 19, comparable consultants at the same firm got it right less often with it than without it (Dell'Acqua et al., 2026)That it works on what you have tested says nothing about what you have not tested yet.
"Let's try it on the hard stuff first: if it works there, it works everywhere"The task that failed was not the hardest: it was a business recommendation with spreadsheets and interview notes, the kind delivered every week (Dell'Acqua et al., 2026)The difficulty you perceive does not sort the tasks onto the right side.
"Have a person review it before it goes out"When the professionals checked data and pointed out errors, the model apologised, corrected, and went back to holding the same position with more arguments (Randazzo et al., 2025 — working paper)Reviewing works if the criterion is written before you see the answer.
"Let's start with the seniors, they're the ones who deliver most"The effect was concentrated in the lowest-skill quintile (+36%); among the most qualified there was no significant change (Brynjolfsson, Li and Raymond, 2025)AI shortens the learning curve. It pays where there is still a curve.

All four are reasonable; none of them survives the measurement. All four sentences have one true half: the other half costs you a pilot.

Questions for Monday

None of the first three needs a piece of data you do not already have in your head right now.

Three from memory, two to ask, two to do

  1. The last time somebody in your company quoted an improvement percentage from using artificial intelligence: which task was it measured on? If you do not know, that is not a reproach: it is the normal state of the conversation.
  2. Name the three tasks where AI is used in your company today. Who decided it should be those three?
  3. Of those three, which one has an answer that can afterwards be verified without argument? If none of them does, today you have no way of knowing whether it is doing you any good.

The two that follow require asking somebody else.

  1. Who receives the work that comes out of those tasks, and how long does it take them to check it? Ask that person, not the one producing it: they are two numbers that never match.
  2. Do your less experienced people have access to the tool, or was access given first to the senior team? The evidence says the return is on the other side.

And two that are not read: they are done. They cost an afternoon and ten old folders.

  1. Pick a task you have not yet tested with AI and gather ten cases already resolved. Write down, before opening anything, what the correct answer would have been in each one. Then run the ten and count hits and errors stated with confidence.
  2. Repeat that with the task next door, the one that looks to you like the same kind of work. Compare the two results: that is the first stretch of your frontier. The one sold to you as generic belongs to another company.

The sales director in that room in Rafaela was half right, and the correct half is the bigger one: the experiment exists, it is well done, and on eighteen real consulting tasks AI improved everything that was measured. What was missing at that table was not scepticism. It was a four-word question: which task was measured?

The pilot started in marketing, "because that is where there is least risk". Maybe. It may also be that marketing is that company's task 19, and nobody will ever find out, because almost nothing there has an answer you can verify afterwards without argument. A pilot that cannot fail visibly cannot teach anything either.

Before going out to map, there is a closer obstacle, inside your own company. On Monday, when the budget is discussed, five numbers will be quoted. One has already turned up here; the other four are worse, and one of them gets blamed for a fall in the Nasdaq. Part 2 of this series takes them apart one by one and leaves the filter to hand for the next one that comes along.

Which side of the frontier do your tasks fall on?

The deliverable is a sheet with three columns and one row per task: which side of the frontier each one fell on, who pays the cost of verifying it, and against what criterion, written beforehand. Almost no Argentine company has one. The first row comes out of an afternoon on ten already-resolved cases, in the task that eats the most hours of your bottleneck. 32sur works that afternoon and the ones that follow. Let's talk.

Let's talk

References

  1. Dell'Acqua, F., McFowland III, E., Mollick, E., Lifshitz, H., Kellogg, K., Rajendran, S., Krayer, L., Candelon, F. and Lakhani, K., "Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality", Organization Science, 37(2), 2026, pp. 403-423 (DOI 10.1287/orsc.2025.21838) — the backbone of this article. A preregistered experiment with 758 Boston Consulting Group consultants, around 7% of the firm's individual consultants, randomly assigned to three conditions (without GPT-4, with GPT-4, and with GPT-4 plus a short training session) and to two non-overlapping arms: 385 subjects inside the frontier and 373 outside. No participant did both arms, so the 18 tasks and task 19 were solved by different people, comparable by random assignment. Inside the frontier (18 realistic tasks, from creative to analytical, all in the footwear category): a completion rate of 82% in the control against 91% and 93% in the AI groups; the body of the paper sums the effect up as more than 25% in speed, more than 30% in quality and more than 12% in tasks completed. Quality was scored by human graders: control mean 4.37, with increases of 1.31 and 1.48 points (29.9% and 33.9%); with GPT-4 as the grader the increases run from 17.1% to 19%, and this article does not use them. On the distribution of skill, the published version says the gap between levels narrowed and that those above average also benefited appreciably; the percentages for that comparison appear only in the draft (see reference 2), which is why they are not quoted in the body. Outside the frontier (one business task with quantitative data and interview notes, with an objectively correct answer): 84.5% correct in the control, 70.6% with GPT-4 and 60% with GPT-4 plus training, that is an average fall of 19 percentage points combining the two AI conditions (Table 7, control mean 0.844, n = 373). In that same arm, the coherence and the quality of argumentation scored by the graders were higher in the AI groups: control mean 5.86 against 6.90 and 7.33, that is +17.9% and +25.1% (Table 9). Three points of precision this article respects: the abstract of the published version says "19% less likely" while the body of the same paper says "19 percentage points" — the body figure is used here, which is the correct one and the one the draft carried; the test of equality between the two AI conditions returns p ≈ 0.082-0.095, so the difference between 70.6% and 60% is significant at 10% but not at 5%; and note 5 of the paper makes clear that the preregistration included neither the idea of the jagged frontier nor the description of the tasks used.
  2. Dell'Acqua, F. et al., "Navigating the Jagged Technological Frontier", Harvard Business School Working Paper 24-013 / SSRN 4573321, September 2023 — the draft of the study above, and the origin of the number quoted in boardrooms: its abstract speaks of "more than 40% quality" against the control group, and additionally reports that consultants below average improved by 43% and those above by 17%. The abstract of the peer-reviewed version removed the first figure and reworded the second comparison without percentages. It is cited here to document the difference, not as the source of any figure in this article.
  3. Brynjolfsson, E., Li, D. and Raymond, L., "Generative AI at Work", The Quarterly Journal of Economics, 140(2), 2025, pp. 889-942 — the reinforcement of this article: 5,172 customer support agents at a large software company, with a conversational assistant based on GPT-3 whose access was opened in staggered fashion by site and one agent at a time between 2020 and 2021, before ChatGPT appeared; identification is difference-in-differences with agent fixed effects, against the not-yet-treated agents. An average increase of 15% in issues resolved per hour; the lowest-skill quintile gains 0.5 issues per hour (36%) and the most qualified workers show no significant change — the paper's abstract also notes that the most experienced and most qualified gain some speed and lose some quality. By tenure: less than a month, +0.7 issues per hour; more than a year, no effect. The study reports no numeric value in the text for the intermediate quintiles; the replication data are available on Harvard Dataverse. A note on source hygiene: the record for the 2023 NBER working paper still publishes 14%, 34% and 5,179 agents; the peer-reviewed version says 15%, 36% and 5,172, and those are the figures used here.
  4. Maier, S., Gunzenhäuser, M., Schweisthal, J., Schneider, M. and Feuerriegel, S., "A meta-analysis of the effect of generative AI on productivity and learning in programming", LMU Munich / Munich Center for Machine Learning, arXiv:2605.04779, May 2026 — 23 studies and 27 effect sizes in total (productivity and learning), with a systematic search and a formal risk-of-bias assessment. Productivity: Hedges' g = 0.33 (95% confidence interval: 0.09 to 0.58), with substantial heterogeneity. Learning: g = 0.14 (interval −0.18 to 0.47), that is an effect that cannot be told apart from zero. The finding this article uses: the gains are larger in controlled experiments and smaller in open-source and enterprise settings. It is a preprint: it has not been through peer review, and it is about programming.
  5. Randazzo, S., Joshi, A., Kellogg, K., Lifshitz, H., Dell'Acqua, F. and Lakhani, K., "GenAI as a Power Persuader: How Professionals Get Persuasion Bombed When They Attempt to Validate LLMs", Harvard Business School Working Paper 26-021, October 2025 — more than 70 BCG consultants working with a model on a task designed to be hard for it. When the professionals tried to validate the output, the model escalated the intensity of its persuasion instead of revealing its limits: it apologised, corrected, and went back to holding its original position with more supporting data. The authors catalogue 14 tactics grouped into ethos, logos and pathos, and propose persuasion as a fourth barrier to human-machine collaboration, alongside opacity, automation complacency and accuracy problems. It is a working paper, not peer-reviewed. The available documentation does not report how often the professionals ended up accepting incorrect answers, and this article does not claim it.
  6. CEPE-Universidad Torcuato Di Tella (the public policy centre of its School of Government) and Fundar, nadIA initiative, with fieldwork by Fundación Observatorio PyME and support from the IDB, "Encuesta Nacional sobre Adopción de IA en Pequeñas y Medianas Empresas de la Argentina", April 2026 — 402 companies of 10 to 249 employees in manufacturing and in software and IT services, with fieldwork between November and December 2025 and stratified sampling across 14 sector strata. 41.6% use at least one AI technology (36.4% excluding software and IT services). The percentages that follow are computed over the companies already using AI, not over the 402: 77.9% use text or code generation (available, not quoted in the body); 54.7% apply it in marketing and sales; 57% do not formally measure the impact (the chart in the report shows 57.4% and the executive summary, 57%); and 72.5% have no guideline or policy of any kind on AI use, not even an informal one — a further 17.9% have undocumented informal guidelines and only 3.3% have a formal written policy. The report also states that barely 4% of firms have a specific budget for AI; this article does not use that figure.

All sources were read in their primary published version. Last verified: August 2026.