Services Method Training Firm Work with us FAQ Insights Contact
ESENPT

Home  ›  Insights  ›  24

Article 24 · AI without the hype

The five numbers your board will quote on Monday (the slide nobody argues with: it sets the amount and gets eleven seconds)

Every AI presentation has one slide nobody argues with: the one with the big numbers and university logos along the bottom. We read the primary sources for the five that circulate most today in boardrooms across the region. Two came from drafts whose numbers changed once they were reviewed; one was never a study at all and helped move the Nasdaq; another measures something different from what is attributed to it; the fifth was measured by the vendor that sells the product, in a draft that was never published; and the first says more than the slide says. At the end, the three customs questions for the next figure someone puts in front of you.

By 32sur · October 2026 · Reading time: 15 minutes · “AI Without the Hype” series, Part 2 of 4

Slide 4

Twenty past eleven on a Thursday. The board of an auto-parts maker meets where it always meets: at the offices in Belgrano, not at the plant in Garín where the seven hundred people who will be discussed all morning actually work. Six on this side of the table, two on the other.

The deck is called "Artificial Intelligence Programme — Phase 1" and asks for six hundred thousand dollars over eighteen months. Slides 5 to 22 get an hour and forty minutes of debate, and two are sent back to be redone. Slide 4 is not. It is called "The evidence": five big numbers with four logos along the bottom —Harvard, MIT, Stanford, GitHub— and one with no logo, carrying a four-letter acronym nobody googled.

The consultant gives it eleven seconds. "You all know this already," he says, and adds without pausing: "I can send you the study afterwards if you like."

Nobody asked him for it.

It is not that the board is distracted. They argue with that slide less because it is the only one of the twenty-two that appears to have no opinion inside it. The others are his. That one is Harvard's.

By two o'clock the programme is approved. The minutes record the amount and none of the five figures that set its size.

(The scene is a composite reconstruction. From here on, everything carrying a number is measured and has a source.)
The five primary sources were read in the original, in the version available in August 2026.

Here that slide lands on soft ground: in the survey run by CEPE —the public policy centre at Universidad Torcuato Di Tella— and Fundar across 402 Argentine SMEs, the most frequently reported barrier to adopting AI is not cost, it is the lack of experience and knowledge about the subject: 46.2%, followed closely by the lack of qualified staff.

One caveat, and only one: the five measurements were taken on support agents at a Fortune 500 company, consultants at a global firm and open-source programmers. None of them on a company that looks like yours. What this article gives you is how each number got here.

Why the correction never reaches the room

If the studies are one click away, how do they arrive wrong? We have taken apart before the 70% of change efforts that fail, which has no such measurement behind it: one figure on its own, and the mechanism was word of mouth. Here there are five, and the mechanism belongs to a subject that ages in months.

A study is not born published. First it circulates as a draft —a working paper or a preprint: the three words name the same thing—, a text nobody on the outside has reviewed yet. Then comes peer review, which is when other researchers in the field read it and object: two years in one of the cases in this article and two and a half in the other, and by the time it arrived the number had changed. The conversation about AI does not take two years: it takes two weeks. By then the draft's figure is already in press articles, in posts, in vendor decks and on slide 4. And there is nothing to carry the correction back: nobody sends an erratum to a slide.

The figure that survives that circuit is what we call the zombie number: one its own authors have already corrected and that keeps walking, because the draft travels and the correction does not.

And it is not assembled in bad faith. Somebody has to deliver the deck on Monday: they search "AI productivity study", find an article quoting Harvard, copy the figure and the logo —which comes from the article, not from the study— and build the row. Four minutes. The original is one click away, and the click is not part of the assignment.

None of this has to be taken on trust; it checks out in thirty seconds. The study that opens the list has been published, peer reviewed, since February 2025, and the landing page for its draft at the NBER —where economics working papers circulate— still shows the 2023 abstract. Anyone who "verifies" there confirms the wrong number. Which is why verifying is not opening the first page that comes up: it is searching the title in quotation marks with the most recent year. It is question two of the customs check at the end.

The circuit is two lanes that never touch.

What reaches the slide is always the draft One case: the study that opens this article, from 2023 to 2026. 2023 2024 2025 2026 THE LIFE OF THE STUDY THE DRAFT COMES OUT 2023 THE PEER-REVIEWED VERSION COMES OUT February 2025: the number changed two years the correction has no way of getting there WHAT CIRCULATES post press article vendor deck slide 4 The case drawn is Brynjolfsson, Li and Raymond: NBER draft, 2023 → The Quarterly Journal of Economics, online 4 February 2025. The second case in this article takes two and a half years: draft of September 2023 → Organization Science, online in March 2026. There is no measurement of the average lag in the field: these are those two cases.
Figure 1 — What reaches the slide is always the draft. The draft takes weeks to reach a boardroom. The correction takes years, and never arrives. The case drawn here is Brynjolfsson, Li and Raymond (NBER draft, 2023 → The Quarterly Journal of Economics, online 4 February 2025, print issue in May). The second case in this article takes two and a half years: draft of September 2023 → Organization Science, online in March 2026. There is no measurement of the average lag in the field: these are those two cases.

Number one · As quoted: "AI lifts productivity by 14%, and by 34% among the rookies"

The first one corrects upwards: whoever quotes the version in circulation is understating the tool.

The source is Brynjolfsson, Li and Raymond, in The Quarterly Journal of Economics, 2025: 5,172 customer support agents at a Fortune 500 company, with a conversational assistant rolled out in waves. The abstract says "by 15% on average", and the group with the lowest measured skill appears as "an increase of 0.5 in RPH, or 36%" —RPH being resolutions per hour—. There is one sentence that never makes it onto the slides: the most experienced agents "see small gains in speed and small declines in quality".

The unit needs translating, and it is not the one you would imagine: the raw means in Table I go from 2.0 to 2.5 resolutions per hour, and the 15% in the abstract is what is left once the effect of the staggered rollout is isolated. Two an hour, not twenty: a chat averages 41 minutes and each agent handles several at once.

And there is a second correction, larger than the decimal places: the 36% does not belong to the rookies. It belongs to the fifth of the team that had been measuring worst, who may have been in the job for ten years. Tenure is a different variable, reported separately: less than a month on the job, +0.7 resolutions per hour; more than a year, no change. Whoever translated "lowest-skill quintile" into "rookies" did not change a number: they changed who gets the tool first. And multiplied across the whole headcount, the average projects a saving the study finds neither in the best-measuring quintile nor among those with more than a year in the role.

Number two · As quoted: "BCG consultants produced 40% higher quality"

This number was never argued down. It simply left.

In September 2023 the draft of a pre-registered experiment with 758 BCG consultants appeared, and its abstract said "more than 40% higher quality compared to a control group". That sentence went round the world. In March 2026 the same work came out published in Organization Science: the abstract does not carry the figure —it says "solutions of significantly improved quality"— and the body reports "performance by more than 30%".

It is worth saying what "40% higher quality" is: every deliverable was scored by human graders, against the average of the group that worked without AI. With them it runs from 29.9% to 33.9%, and the two ends are the two groups with AI: the one that simply used it and the one that also got training in how to ask it for things. With GPT-4 as the grader it falls to a little over half.

None of this says the study is bad: it is a pre-registered field experiment on real work, and it also measured a task where the effect flips over — the subject of Part 1 of this series. What was corrected was a figure in the draft's abstract, and the author withdrew it: the figure carried on alone. The zombie number in its sharpest form.

And you have to be careful with the subtraction, because "more than 30%" is a floor, not a point. What can be asserted is that anyone who sized the case at 40% projected too much: they have a quarter too much —40 where the published source says more than 30—, not a quarter left over. In Thursday's programme, that extra quarter is what separates a business case from a wish.

Number three · As quoted: "95% of AI pilots fail"

Start with what it is not. It is not a study.

The GenAI Divide: State of AI in Business 2025, by Challapally, Pease and Raskar, came out in mid-2025 from MIT Project NANDA —an MIT initiative devoted to networks of AI agents— and it did not go through peer review, nor was it ever meant to. Its entire empirical base is 52 interviews, 153 survey responses gathered at four conferences and a review of some 300 initiatives.

The arithmetic, slowly, over a hundred companies. Out of every hundred, about eighty never ran a pilot of a tool of their own. That leaves twenty, and of those twenty, five reached a result. Five companies out of a hundred: that is the 5% in the headline. But those same five, out of the twenty that tried, are one in four.

The 95% was never a pilot failure rate: it is the share of companies that did not reach a result, counting the ones that never tried anything. Read that way it is no use either for stopping a project or for selling one. Two more facts: the bar was a marked and sustained impact six months after the pilot, and the authors were building the AI agent tools the report recommends.

There is also a figure almost nobody read: the report puts success at around 67% for initiatives with outside partners against 33% for in-house builds — self-reported by respondents, with no published base for either calculation. Here they are no use for calculating anything. They are useful for one thing: the document quoted to prove that AI does not work carries inside it a success rate of two out of three.

None of that stopped it from moving money: it is credited with contributing to a fall in the Nasdaq in August 2025, and it was amplified by Forbes, Axios, The Hill and Harvard Business Review. The last one deserves a paragraph of its own.

The number Harvard Business Review did not filter out. In September 2025, Harvard Business Review published one of its most widely shared articles of the year on the hidden cost of AI inside organisations. It opens, in the first paragraph, leaning on the 95% from the MIT report. The flagship magazine of management amplifying the very figure it should have filtered out. This is not one magazine's slip: it is the best available proof that there is no source hygiene here, not even at the top. And the phenomenon that article did measure —they call it workslop, work that looks like work— is real, it is expensive, and not one of the five figures on slide 4 measures it.

Number four · As quoted: "expert programmers are 19% slower with AI"

This one is fine. The number is correct, the study is honest and its authors are the first to qualify it. What is not in the source is the conclusion hung on it.

It was published by METR in 2025 —Becker, Rush, Barnes and Rein—, and METR is the logo-less acronym on slide 4: not a university, an independent non-profit institute that evaluates AI systems. Small and careful: 16 developers, 246 tasks, five years on average in those repositories. The abstract says three things in a row: "developers forecast that allowing AI will reduce completion time by 24%… developers estimate that allowing AI reduced completion time by 20%… allowing AI actually increases completion time by 19%".

Read them in that order, because that is where the finding is. They believed AI would save them almost a quarter of their time. Then they believed it had. The stopwatch says they took longer: what ended up being measured is not AI, it is how far off the perception of the person using it can be.

There is a later fact almost nobody quotes, from METR itself: in February 2026 it reported that its follow-up experiment —57 developers, more than 800 tasks— had been rendered unusable by selection bias. It is getting harder and harder to find developers willing to work without AI, and between 30% and 50% stopped submitting tasks they did not want to do without it. METR calls its new results very weak evidence and is redesigning the study.

The line exists and you have heard it: "it saves me about half a day a week". Serious people say it, about their own work, with no interest in exaggerating. It is exactly the number this study measured and found to be wrong.

What to take away, then: the 19% is well measured in 2025 and it is not a verdict on 2026. What survives intact is not the figure, it is what the figure exposed. If you ask your team how much time AI saves them, you will get a number measured with the one instrument this study showed to be faulty. And for the sceptic: sixteen developers are not "programmers" as a class.

The distance between what they believed and what happened fits on a single axis.

They believed they had saved a fifth of the time. They had spent a fifth more. Change in task completion time, measured over 246 real tasks. 24% less time WHAT THEY FORECAST BEFORE STARTING 19% MORE time WHAT THE STOPWATCH MEASURED 0 ← LESS TIME MORE TIME → 20% less time WHAT THEY BELIEVED AFTERWARDS nearly forty points between what they believed and what happened points, not per cent: it is the distance between two percentages Becker, J., Rush, N., Barnes, E. and Rein, D., Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity, METR, arXiv:2507.09089, 2025. 16 developers with five years on average in those repositories, 246 tasks; the interval on the 19% runs from +2% to +39%. In February 2026 METR reported that its follow-up experiment had been rendered unusable by selection bias.
Figure 2 — They believed they had saved a fifth of the time. They had spent a fifth more. The finding is not about AI: it is about the one instrument almost every company uses to measure its effect — the impression of the person using it. Source: Becker, J., Rush, N., Barnes, E. and Rein, D., Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity, METR, arXiv:2507.09089, 2025. 16 developers with five years on average in those repositories, 246 tasks; the interval on the 19% runs from +2% to +39%. In February 2026 METR reported that its follow-up experiment had been rendered unusable by selection bias.

Number five · As quoted: "Copilot makes them 55.8% faster"

That decimal place is all the warning you need: it promises a precision the measurement does not have.

The 55.8% comes from Peng, Kalliamvakou, Cihon and Demirer, 2023: a draft never published with peer review, and three of its four authors are from Microsoft and GitHub — the vendor measuring its own product. The task is a single one and it is synthetic: writing an HTTP server in JavaScript. Underneath that decimal place there are 35 freelance developers who finished it, and an interval running from 21% to 89%.

The serious counterweight came from the same team, enlarged, in Management Science in 2026: three randomised experiments at Microsoft, Accenture and a Fortune 100 company, with 4,867 developers. It gives around 26% more tasks completed, with a margin its own authors call noisy: enough to assert that there is an effect, not how big it is. Two of the six are from Microsoft, and that has to be said here too.

The benchmark that orders the whole field comes from a 2026 LMU Munich meta-analysis of 23 studies: g = 0.33. That "g" is not a percentage: it measures how far the average of the group with AI moves relative to the group without it, in units of how much people already vary among themselves. By convention 0.2 is small and 0.5 medium: 0.33 sits in between, real and modest. That unit is what makes it possible to average studies that measured different things. And the gains, it adds, are larger in controlled experimental settings and smaller in open source and inside companies.

The number on the slide, by contrast, sets the board's expectation, and it comes out of a task nobody performs at your company, measured by the party selling it to you. The one that should set it is half that: the 26% from the three experiments. A business case built on that decimal place is born oversized, and no amount of management fixes that afterwards.

Not one survives intact

The distinction to take away is not the obvious one. The rule is not "published good, draft bad": drafts are often the best there is on a new subject. It is knowing which of the two you are reading, because how much weight a figure can carry in a capital decision depends on that and not on the logo at the bottom.

Not one of the five says, at source, what the slide says The five figures on slide 4, read against their primary source. AS QUOTED WHAT THE PRIMARY SOURCE SAYS WHY THEY DO NOT MATCH “+14%, +34% for rookies” 15% and 36%, over 5,172 agents CHANGED ON PUBLICATION and “rookies” is not what was measured “+40% quality (BCG)” The published abstract drops it; the body says more than 30% CHANGED ON PUBLICATION “95% of AI pilots fail (MIT)” 52 interviews and 153 surveys. The 5% was computed over the total, not over those that piloted NEVER WAS A STUDY “experts are 19% slower” The number is right: 16 developers in their own repos. It measures the perception gap MEASURES SOMETHING ELSE “+55.8% with Copilot” A never-published draft, from the vendor, on a synthetic task finished by 35 people MEASURED BY THE SELLER Two changed when they went through peer review. The other three never went through it. References 1 to 5, read in August 2026. In the fourth row the number is not in dispute: the source does not claim it is a verdict. 32sur schematic.
Figure 3 — Not one of the five says, at source, what the slide says. Not one of the five survives the trip from the source to the slide intact. References 1 to 5, read in August 2026. In the fourth row the number is not in dispute: the source does not claim it is a verdict.

The real size of what holds the slide up:

52 + 153interviews and survey responses: the entire empirical base of the report that installed "95% of AI pilots fail" (MIT Project NANDA, 2025)
16developers — the complete sample of the study quoted as a verdict on programmer productivity with AI (Becker et al., METR, 2025)
35people and a single task: writing an HTTP server. The entire body of work behind Copilot's 55.8%; the interval on that improvement runs from 21% to 89% (Peng et al., 2023 draft)

The two measurements in this article that did go through peer review were made on 5,172 support agents and 758 consultants. Three of the five figures on slide 4 rest on nothing of that size.

What is well measured (and enough to decide with)

On the other side of the counter there is evidence you can download for nothing and almost nobody takes to the meeting: a real effect, moderate, very unevenly distributed across people, that shrinks outside the laboratory.

One figure orders the discussion better than the previous five: how many people actually use AI. Between 17% and 20% of companies in the United States, 37% among those with 250 employees or more. There are fewer people inside than the conversation suggests.

And a methodological detail worth more than the figure itself: the number changes depending on who is asked. A Federal Reserve note puts it at 18% if the unit is the firm, 41% if it is the working person, and 78% if you count people instead of firms —each one weighing what its payroll weighs, and the large firms are the ones that use AI most—. Before you compare yourself with the market, ask which one you are being shown.

The figureWhere it comes fromWhat kind of evidence it isWhat it lets you decide
+15% resolutions per hour on average, and +36% in the fifth of the team that had been measuring worstBrynjolfsson, Li and Raymond, The Quarterly Journal of Economics, 2025 — 5,172 support agentsPeer reviewed. Field data from a real company, with a staggered rolloutWho gets the tool first: the worst-measuring fifth and the person who has just joined, not the whole payroll
g = 0.33 effect on productivity in programming — the average shifts by a third of how much people already vary among themselves: real and modest — and smaller in companies than in the labMaier et al., LMU Munich, 2026 — meta-analysis of 23 studies, with a formal risk-of-bias assessmentDraft: not yet peer reviewed. Used here with that label attachedHow much to discount from any laboratory figure before taking it into the business case
Between 17% and 20% of US companies use AI; 37% among those with 250 employees or moreU.S. Census Bureau, Business Trends and Outlook Survey, May 2026Official statistics, with a probability sample and published methodologyHow many you are really competing against, and how much advantage is still there to take
Adoption is 18%, 41% or 78% depending on the survey: the unit can be the firm, the person, or people counted with the weight of each payrollAllen, FEDS Notes, Board of Governors of the Federal Reserve System, April 2026A central bank research note comparing three official surveysBefore comparing your company with "the market": ask what is being counted

Four figures that can carry a capital decision, each with a label saying how far it goes.

The customs check: three questions before the figure sets the amount

They run from the cheapest to the most uncomfortable.

One: what is the primary source? Not the article, not the post, not the vendor's slide: the study, with author, year and where it came out. If nobody on the team has opened it, the figure is not verified. It is repeated.

Two: published or a draft? The official page for a published study may still be showing the draft's abstract: the institution's name does not guarantee the version. The way out takes a minute: search the title in quotation marks with the most recent year; if it comes up with a journal, a volume and page numbers, that is the one that rules.

Three: who did the measuring? The vendor, the consultancy that is going to implement it, or a team that sells what its own report recommends. Knowing this does not invalidate the figure: it places it.

And one rule that keeps this from turning into paperwork: the figure that fails all three is not banned, it can stay as a market signal. The only thing it cannot do is set the amount.

The customs check is not run by whoever brings the slide: it is run by whoever is going to sign the amount, and it is run before the meeting, with the invitation. Five minutes per figure. What does not exist today is the obligation to answer them, and that cannot be bought: it is written into the agenda.

No figure is banned: the ones that fail the check stop setting amounts Five minutes per figure, before the meeting and not inside it. THE QUESTION WHAT IS ASKED FOR WHO ANSWERS IF THERE IS NO ANSWER 1 WHAT IS THE PRIMARY SOURCE? The study: author, year, where published Whoever brings it, in writing and before the meeting The figure is not verified: it is repeated 2 PUBLISHED OR A DRAFT? Version and date. Title in quotation marks + the most recent year Anyone on the team: one minute The draft may have changed: use it as an order of magnitude 3 WHO MEASURED IT? Who funded it and who sells what the figure measures Whoever signs the amount, and nobody else The figure stays as a market signal, not as a basis for sizing The customs check bans no figures, demands no study for every claim and does not replace the judgement of whoever decides. It separates the figures that can set an amount from the ones that cannot. 32sur schematic. The only third-party fact is in question 2: the NBER page for the Brynjolfsson, Li and Raymond draft still shows the 2023 abstract even though the study has been published since February 2025.
Figure 4 — No figure is banned: the ones that fail the customs check stop setting amounts. Five minutes per figure, and not one of those five minutes happens inside the meeting. 32sur schematic. The only third-party fact is in question 2: the NBER page for the Brynjolfsson, Li and Raymond draft still shows the 2023 abstract even though the study has been published since February 2025.

Questions for Monday

The first three need you to open nothing and ask nobody anything.

Three from memory, four that take longer

  1. The last AI presentation you were given: how many of the figures on the evidence slide can you put an author and a year to today? Do not verify them: count them.
  2. Those figures — were they brought by somebody from your company or by the party that was going to implement the project?
  3. Of the five numbers in this article, how many have you heard in the past six months around your own table?

The next four take longer. None of them requires buying anything:

  1. Ask for the original studies behind the last presentation you approved. Not to audit anyone: to see how many arrive and how long they take. It is one email, and the answer is already the diagnosis.
  2. Open the business case currently in force and find the number that sets the saving. Write next to it where it came from. If it does not fit on one line, it cannot hold up that amount.
  3. Ask three people on your team how much time AI saves them per week. Write it down and do not use it for the business case: use it to measure how far the saving is overestimated inside your own walls.
  4. For the next decision, ask that every figure arrive with author and year, draft or published, and who funded it. Then compare how many slides are left standing.

Go back for a moment to the room in Belgrano. The programme was approved, and it probably should be: nothing here says it is not worth doing. What there is, is a difference of size. With the five numbers on the slide, six hundred thousand dollars was cheap. With the five as their sources state them there is still a case — smaller, more concentrated, with the bulk of the benefit sitting with the people who appear least in the deck.

The five will keep circulating, and two of them will keep being figures their own authors have already corrected. That is the problem with zombies: you do not kill them by arguing with them. You stop them at the door.

And there is something none of the five measures: they argue about how much faster things get produced, and not one of them asks what is being produced or who reviews it afterwards. That is the report that reaches you on Monday and that, once you have finished reading it, said nothing: work that looks like work. That is the subject of Part 3 of this series.

In the meantime, Thursday's problem has a fix and it costs no money. Slide 4 can be argued with: it is the only one of the twenty-two where a five-minute question changes the amount of the decision.

Can you put an author and a year, today, on every figure that set your last AI budget?

The deliverable fits on a single page: one row per figure, with the author, the year, whether it is a draft or a published version, and who funded it. With that page on the table no amount gets sized with a number whose origin is not written down, and vendor decks arrive with the source column filled in. The first one comes out of a session on the last AI deck you were shown: the sources that set the amount are opened, and the sizing is redone with the ones that hold. That is the session 32sur works with. Let's talk.

Let's talk

References

  1. Brynjolfsson, E., Li, D. and Raymond, L., "Generative AI at Work", The Quarterly Journal of Economics, 140(2), 2025, pp. 889-942 (advance online publication on 4 February 2025; print issue in May) — number one in this article, in its peer-reviewed version: 5,172 customer support agents at a Fortune 500 company, staggered rollout of a conversational assistant, +15% resolutions per hour on average and +36% in the lowest measured-skill quintile —which is a different variable from tenure: agents with less than a month in the job gain 0.7 resolutions per hour and those with more than a year show no change—. The real scale of the unit is in Table I: 2.0 resolutions per hour before the rollout and 2.5 after, with an average handle time of 41 minutes per query. The abstract also records that the most experienced and highly skilled workers see small gains in speed and small declines in quality. The draft version (NBER Working Paper 31161, 2023) reports 14%, 34% and 5,179 agents, and the NBER page still shows that abstract with no mention of the publication: checking against that page confirms the wrong number.
  2. Dell'Acqua, F., McFowland III, E., Mollick, E., Lifshitz, H., Kellogg, K., Rajendran, S., Krayer, L., Candelon, F. and Lakhani, K., "Navigating the Jagged Technological Frontier", Organization Science, 37(2), 2026, pp. 403-423, DOI 10.1287/orsc.2025.21838 — number two: pre-registered experiment with 758 BCG consultants. The draft abstract (HBS Working Paper 24-013 / SSRN 4573321, September 2023) claimed "more than 40% higher quality"; the abstract of the published version does not carry the figure and the body reports "performance by more than 30%" — between 29.9% and 33.9% with human graders, which are the two groups with AI (with and without training in how to ask it for things), and between 17.1% and 19% when the grader is GPT-4.
  3. Challapally, A., Pease, C. and Raskar, R., The GenAI Divide: State of AI in Business 2025, MIT Project NANDA, 2025 (Pradyumna Chari appears on the cover alongside the three authors, but page 2 lists him as a reviewer) — number three. It is not a paper and it did not go through peer review: 52 structured interviews, 153 survey responses gathered at four conferences and a review of some 300 public initiatives. The report's own funnel for custom-built tools —around 50% investigated, 20% piloted, 5% implemented successfully— is what holds up the arithmetic in the body; the April 2026 episode of 80,000 Hours devoted to the report also documents the Nasdaq impact and the undisclosed conflicts of interest. Arnon Shimoni's complementary methodological critique notes that the report counts as a failure any pilot that does not reach full production and that on page 19 it reports ~67% success in external partnerships against ~33% in in-house builds — self-reported figures, with no published base for either, which is why this article uses them only as an internal contradiction in the document and not as an estimate. Kevin Werbach (Wharton) publicly asked MIT to release the data or withdraw the report; as of the closing date of this article the report has been neither withdrawn nor retracted.
  4. Becker, J., Rush, N., Barnes, E. and Rein, D., "Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity", METR, arXiv:2507.09089, 2025 — number four, and the only one of the five whose number is not in dispute: 16 developers with five years of average experience in the repositories where they worked, 246 tasks; they forecast a 24% time saving, afterwards estimated they had saved 20%, and the measurement gave 19% more time, with an interval of +2% to +39%. METR update of 24 February 2026: the follow-up experiment (57 developers, more than 800 tasks since August 2025) was rendered unusable by selection bias —more and more developers refuse to take part without AI, and between 30% and 50% of participants said they had stopped submitting tasks they did not want to do without it—; the new results (−18% for the original developers, interval −38% to +9%; −4% for the new ones, interval −15% to +9%) are described by METR itself as very weak evidence, METR itself thinks it likely that developers are faster today, and the study is being redesigned.
  5. Peng, S., Kalliamvakou, E., Cihon, P. and Demirer, M., "The Impact of AI on Developer Productivity: Evidence from GitHub Copilot", arXiv:2302.06590, 2023 — number five: a preprint never published in a peer-reviewed journal; three of its four authors are from Microsoft Research and GitHub (the fourth is from MIT Sloan). A single synthetic task —writing an HTTP server in JavaScript as fast as possible—, on developers recruited on Upwork: 35 completed the task and the survey, and the paper itself reports a 95% confidence interval for the improvement of between 21% and 89%.
  6. Cui, Z., Demirer, M., Jaffe, S., Musolff, L., Peng, S. and Salz, T., "The Effects of Generative AI on High-Skilled Work: Evidence from Three Field Experiments with Software Developers", Management Science, 2026, DOI 10.1287/mnsc.2025.00535 — the serious counterweight to number five: three randomised experiments at Microsoft, Accenture and a Fortune 100 company, 4,867 developers, 26.08% more tasks completed with a standard error of 10.3%; the paper itself warns that each experiment is noisy and that the estimate is an instrumental-variables treatment-on-the-treated effect, not a simple average. Two of the six authors are from Microsoft.
  7. Maier, S., Gunzenhäuser, M., Schweisthal, J., Schneider, M. and Feuerriegel, S., "A meta-analysis of the effect of generative AI on productivity and learning in programming", LMU Munich / Munich Center for Machine Learning, arXiv:2605.04779, 2026 — the benchmark for the field: 23 studies, 27 effect sizes, systematic search with a formal risk-of-bias assessment; effect on productivity g = 0.33 (interval 0.09 to 0.58) and on learning not significant (g = 0.14). Hedges' g is a standardised measure: it expresses the difference between the group with AI and the group without it in units of the dispersion of the data, which is what makes it possible to average studies that measured different things. It is a preprint and it is cited in this article with that label attached: the gains are larger in controlled experimental settings and smaller in open-source and corporate contexts.
  8. U.S. Census Bureau, Business Trends and Outlook Survey, May 2026 report on AI use in businesses — between 17% and 20% of companies in the United States report using AI between December 2025 and May 2026; 37% among those with 250 or more employees, 32% among those with 100 to 249 and less than 20% among those with fewer than 20; national average 19.8%.
  9. Allen, J. S., "Monitoring AI Adoption in the U.S. Economy", FEDS Notes, Board of Governors of the Federal Reserve System, 3 April 2026 — why the adoption figure changes with the survey: around 18% of firms according to the Census, 41% of the workforce according to the Real-Time Population Survey and 78% weighted by employment according to the Survey of Business Uncertainty. The author explicitly refrains from drawing conclusions about productivity.
  10. Niederhoffer, K., Rosen Kellerman, G., Lee, A., Liebscher, A., Rapuano, K. and Hancock, J. T., "AI-Generated 'Workslop' Is Destroying Productivity", Harvard Business Review, 22 September 2025 — cited in this article for one reason only, the one in the box: one of the most widely shared articles of the year on the hidden cost of AI in organisations opens its first paragraph leaning on the 95% from the MIT report. What that article did measure is the subject of Part 3 of this series and is not used here.
  11. CEPE-Universidad Torcuato Di Tella (the public policy centre at its School of Government) and Fundar, nadIA initiative, with fieldwork by Fundación Observatorio PyME and support from the IDB, "Encuesta Nacional sobre Adopción de IA en Pequeñas y Medianas Empresas de la Argentina" (national survey on AI adoption in Argentine SMEs), April 2026 — the only local data point in this article: 402 companies with 10 to 249 employees in manufacturing and in software and IT services, fieldwork between November and December 2025, stratified sampling across 14 sector strata. The most frequently reported barrier to adopting AI is the lack of experience and knowledge about the subject (46.2%), followed by the lack of qualified staff (42.2%) and high cost (29.2%). The figure on budget devoted to AI is not quoted here: the available coverage does not agree with itself.

All sources verified against their primary version as of August 2026.