Services Method Training Firm Work with us FAQ Insights Contact
ESENPT

Home  ›  Insights  ›  25

Article 25 · AI without the hype

Work that looks like work (AI did not eliminate the thinking: it moved it to the desk of whoever receives it)

The report arrived on time, it runs to fourteen pages and it is better written than usual. Six minutes in you discover that it says nothing, and now you have two problems: what to do with the report, and what to do with the person who sent it. The phenomenon has had a name since 2025. What holds this piece up is not that name: it is a Danish study that matched what twenty-five thousand workers say artificial intelligence saves them against what later shows up in the records of their hours and their earnings. At the end, the one-page sheet a team can use to agree, in half an hour, on what counts as finished work.

By 32sur · October 2026 · Reading time: 15 minutes · “AI Without the Hype” series, Part 3 of 4

Six minutes

Ten past eight on a Tuesday, before anyone else arrives. The operations manager of a medical supplies distributor —two hundred and forty people, warehouse and back office in Córdoba— opens the report she asked for on Friday: why the margin on the pharmacy channel fell four points over the half-year.

Fourteen pages. Half a page of executive summary, subheadings in bold, three tables, one chart. It is well written. It is better written than what used to reach her from that same person.

Six minutes in she stops reading. There is no recommendation: there are four "possible lines of action", and they are the four anyone would name without opening the system. There is not a single number she did not already have on Friday. And there are two claims about a competitor that she does not know where came from, that she cannot verify before ten, and that, if true, change the decision.

It is sixteen minutes past eight. The meeting where she has to bring that decision is at ten.

And now she has two problems, not one. The first is the report. The second is Julián, who delivered on time, who did nothing she can point a finger at, and who until Friday was the person she trusted most for this.

(The scene is a composite reconstruction. From here on, everything carrying a number is measured and has a source.)

That has had a name since September 2025, and the name is not ours. A team from BetterUp Labs and the Stanford Social Media Lab published it in Harvard Business Review and called it workslop: in their words, "AI generated work content that masquerades as good work, but lacks the substance to meaningfully advance a given task". In plain language: work that looks like work.

40% of those surveyed said they had received at least one in the past month.

Where that 40% comes from, because it matters: an online survey of 1,150 full-time US employees, each one describing what happens to them. It is not an experiment, and it is published by a company that sells leadership development. That does not invalidate it: it places it. It is good for knowing that the phenomenon exists and has a size; not for calculating what it costs you.

The work does not end when it is delivered

There is a calculation everybody makes and that seems impossible to argue with: if each person saves two hours a week, the company saves two hours per person per week. Multiply by headcount and there is your business case.

The calculation fails in exactly one place: it assumes the work ends when it is delivered. It does not. It carries on, at another desk, and there it is more expensive, because verifying something that is well written takes longer than discarding something badly written used to take. No study measures that last point, and we say it for what it is: what we see inside companies. We call it, at 32sur, the downstream transfer: the thinking is not eliminated, it changes owner.

Whoever receives it has no way of anticipating it. In the first part of this series we described an experiment with hundreds of consultants at a global firm who solved real tasks, some with artificial intelligence and some without it. On one of those tasks, chosen because it fell outside what the tool does well, the ones who used it got it right less often. And their deliverables were scored as more coherent and better argued than those of the group that worked without it. The error did not come wrapped in worse writing. There, that was a problem for whoever decides. Here it is the problem of whoever receives: the signal your head uses to estimate whether something deserves your time —it is well written, it is orderly, it has structure— stopped predicting what you want to know.

That is why the cost is not paid on detecting it. It is paid earlier, looking for it. Those who have already received one estimate that, of all the written work reaching them, a little more than one in seven is of that kind. And they report having spent one hour and fifty-six minutes resolving the last episode: reading it, working out that it was no use, separating what could be verified from what could not, and redoing the part that was needed.

A survey in which people report how much time they were made to lose proves nothing at company scale. That is why what follows does not rest on what people say. It rests on what was recorded.

The same deliverable, the same journey, before and now:

Time does not disappear: it changes desk The same deliverable travelling the same path, before and now. Conceptual diagram — not a data chart, and the lengths are illustrative. A · BEFORE whoever produces pays the bulk of the work REQUEST PRODUCE handover USE B · NOW the thinking is not eliminated: it changes owner REQUEST shorter than before PRODUCE VERIFY handover USE this station did not exist before, it has no assigned owner and it is on no dashboard THE JOURNEY OF THE SAME DELIVERABLE — NO NUMERICAL SCALE it is requested… …it is used to decide Whoever staffs that station is usually the person who was already the bottleneck for everything else: the one who signs, the one who approves, the one who decides. 32sur conceptual diagram. It is not a measurement: the lengths are illustrative, and no study measures this transfer task by task. What is measured, and cited in the body, is the time reported by whoever receives (BetterUp Labs and Stanford Social Media Lab, 2025) and the appearance of new integration, review and compliance tasks, which reach even workers who do not use the tool (Humlum and Vestergaard, 2026 revision).
Figure 1 — Time does not disappear: it changes desk. What got shorter and what appeared do not happen to the same person. 32sur conceptual diagram: it is not a measurement and the lengths are illustrative. What is measured is the time reported by whoever receives (BetterUp Labs and Stanford Social Media Lab, 2025) and the appearance of new integration, review and compliance tasks (Humlum and Vestergaard, 2026 revision).

The country where you could look instead of asking

Before going to Denmark, the local picture. In the survey the Di Tella public policy centre ran with Fundar on Argentine SMEs, 91.2% of the companies that adopted artificial intelligence say it did not cut headcount. Nobody left, and those who use it say it saves them time. One question remains: where did that time go?

Denmark is one of the few places in the world where you can answer that without asking. Every person's earnings and hours worked are on record. Two researchers surveyed around twenty-five thousand workers at seven thousand workplaces, in eleven occupations chosen for being among the most exposed to the tool —accountants, lawyers, journalists, technical support staff, teachers—, in two rounds: one at the end of 2023 and another at the end of 2024. Then they matched the answers against those records. They did not ask them how much they had earned. They looked.

The two results are best read one at a time.

The first is what people report: around 3% of time saved. It is a real saving and it is small: a little over an hour in a forty-hour week.

The second is what the records show: nothing. Neither in earnings nor in hours, neither at the level of the worker nor at the level of the workplace.

Here is the sentence that makes all the difference: it is not that it could not be measured. It was measured, and the measurement is precise enough to rule out any effect larger than 2% two years after the launch of ChatGPT. Said without the statistics: if those saved hours had turned into more earnings or fewer hours, this measurement would have found it. It looked, and it was not there. A zero measured with precision is not the same as a zero for lack of data: the first is a finding, the second is a limitation.

Three things about this measurement. The first: it is a working paper, a draft that has not yet been through peer review, and in this series that label always goes on. What gives it weight is where the data come from. The second is uncomfortable for us: the same work circulated in 2025 under a different title and with different figures, and what is cited here is the March 2026 version, which is the current one. And the third: Denmark is not Argentina, and the difference cuts both ways —there the enterprise version of the tool arrived earlier and hours are on record; here, neither of the two.

What this changes for you, in one sentence: if the saving your people feel does not show up in the hours or in the revenue, it is not that your people are lying. It is that the time turned into something else before it reached the result.

What people report and what was recorded, side by side:

What is felt does not show up in the records Denmark: 25,000 workers surveyed, and their earnings and their hours actually on record. The two panels measure different things and do not share a scale. WHAT USERS REPORT time saving reported by the people who use it 3% 0% 5% 10% 3% of working hours a little over an hour in a forty-hour week SCALE FROM 0% TO 10% WHAT THE RECORDS SHOW estimated effect on earnings and hours the margin of the measurement does not reach this far EARNINGS HOURS WORKED −3% 0 +2% +3% Neither in earnings nor in hours. Neither at worker nor at workplace level. SCALE FROM −3% TO +3% This is not a zero for lack of data. It is a zero measured with precision: the measurement rules out any effect larger than 2%. Humlum, A. and Vestergaard, E., Still Waters, Rapid Currents: Early Labor Market Transformation under Generative AI, University of Chicago Booth / NBER Working Paper 33777, May 2025, revised version of March 2026 — the same work circulated earlier under the title Large Language Models, Small Labor Market Effects and with different figures; the current one is cited here. Around 25,000 workers at 7,000 workplaces in eleven exposed occupations, in two survey rounds linked to the Danish administrative records of earnings, hours and occupation. It is a draft: it has not been through peer review, and the two panels do not share a scale. Ruling out effects larger than 2% refers to the two-year window after the launch of ChatGPT.
Figure 2 — What is felt does not show up in the records. Saving time and freeing up hours or money are two different things, and only the first one is felt. The two panels measure different things and do not share a scale. Source: Humlum and Vestergaard, Still Waters, Rapid Currents, NBER Working Paper 33777, revised version of March 2026 (draft, not peer-reviewed).

The tasks that were not there before

The same study asks what users did with the time they did save, and the answer closes the circle: more than eight in ten reassigned it to other tasks at work. Fewer than one in ten used it to take a break. The time neither evaporated nor turned into margin. It stayed inside the working day, doing something else.

What that something else is, is the interesting part. The researchers asked them to describe in their own words the new tasks that had appeared, and then classified them. The single most frequent category is not producing with the tool: it is fitting it into the work —deciding what it is good for, writing down how it gets used, splicing it into what was already there—, and it accounts for one in four new tasks. Another good third is controlling it: correcting what the machine wrote so that it is accurate and clear, and answering for it not running over a rule or a commitment to a client. Between the two, six in ten new tasks are not about producing: they are about enabling and reviewing.

And there is one figure that orders everything above. Where the employer did nothing, close to 8% of users report entirely new tasks caused by the tool; where the employer pushes it, that more than doubles. The more seriously it is adopted, the more new work appears. And it does not stay among users: 4% of those who never used it also report new load, and among the teachers who never opened it, it is one in ten. The new work turns up on the desk of somebody who never opened the program.

The study does not say on whom those tasks land. Your company does: in a two-hundred-person firm, reviewing, splicing and answering do not get spread across two hundred people. They land on the same three or four who already signed everything.

And Argentine ground has both conditions for that to concentrate: among the SMEs that have already adopted artificial intelligence, 95.5% brought in no specialist profile, and in 44.4% each area and each person adopts on their own, with no central coordination. Tools are not what is missing. What is missing is somebody having designed the next step.

Four everyday cases, with the name of who pays each half:

What was done with the toolWhich task got shorter, and for whomWhich task appeared, and for whomWhat happens if nobody counts it
Reply to a customer complaintDrafting it — the support agentChecking that the rule cited and the compensation offered are the right ones — whoever signsThe signer's time is on no dashboard, and the signer was already the bottleneck
Analysis report for a decisionPutting it together and writing it — the analystChecking the claims the recommendation rests on — whoever decidesThe decision gets made anyway, without knowing which of those claims were verified
Summary of a meetingTaking the notes — whoever wrote itComparing the summary against what was actually agreed — the people who were thereThe agreements end up written as a machine understood them, and nobody reads them again until there is a dispute
A spreadsheet or a data queryWriting it — the analystChecking that the formula does what it says it does — whoever uses it laterThe error travels forward and surfaces once a decision has already been made on it

In all four rows the time was saved at the top and the new work appeared at the bottom, on the one person who had no slack left: it is the downstream transfer, row by row. 32sur diagram; the four rows are examples, not measurements. The categories of new work they illustrate are the ones Humlum and Vestergaard (2026 revision) obtained when classifying the new tasks described by the workers themselves: integrating the tool into the work is the single most frequent category, with 26% of all new tasks, and quality review plus ethics and compliance add another 35%.

The cost that is on no dashboard

Go back to that office in Córdoba. The manager has already dealt with the report. What she has not dealt with is who she asks for something next month.

That, too, is measured, and it is the part of the survey that speaks loudest to a director. Of those who received one of these deliverables, 42% came to see the sender as less reliable. 32% would rather not work with that person again.

The direction matters as much as the size. Most of it circulates between peers, but almost one in five of these episodes goes from somebody up to their boss. It travels upward: it reaches the place where promotions are decided, and it gets there without anybody naming it.

We have seen this before, in a different outfit. When we wrote about what happens when a measure becomes the target, a trap with a technical name turned up —surrogation—: producing the appearance of the deliverable instead of the deliverable. A dashboard all in green while the business sinks is exactly that. The fourteen-page report is the same thing, with one large difference: until three years ago, producing the appearance cost almost as much as producing the thing. Now it costs nothing.

40%the share of respondents who say they received, in the past month, at least one AI-generated deliverable that looked like good work and was not (BetterUp Labs and Stanford Social Media Lab, published in Harvard Business Review, 2025; 1,150 full-time US employees)
1 h 56 minwhat the people who received one report it took them to resolve the last episode (same source)
3% and zerothe time saving Danish users report, and the effect that shows up in the administrative records of their earnings and their hours (Humlum and Vestergaard, draft, revised version of March 2026)

Three figures measuring three different things, and the distance between them is the point. The first two come from what people say about themselves; the third matches what people say against what was recorded. That distance is not a flaw in the measurements: it is the subject of this piece.

The way out looks obvious: if nobody knows what was done with a machine and what was not, let everyone say so. That way out has a measured cost, and the company does not pay it: it is paid, unevenly, by the people who use the tool.

The tax no adoption plan has budgeted for

The experiment reads like a card trick. 1,026 engineers were asked to evaluate code. The code was the same. The only thing that changed was a label: whether or not the author had declared using artificial intelligence to write it.

On the quality of the code there was no difference: the reviewers scored it the same. What changed was the mark they gave the person. When the author was a man and declared the use, his competence rating fell 6%. When the author was a woman, it fell 13%.

One more figure, and it is the most uncomfortable. The harshest reviewers were the engineers who do not use the tool, and their penalty on the women engineers who did was 26% larger than the one they applied to their male colleagues. That is the pattern the measurement produced; nobody is attributing an intention to anybody.

It is another draft without peer review, and what makes it hard to argue with is the design: identical code, groups drawn at random, hypotheses written down before the start.

The field counterpart is what matters to whoever signs a budget. The same authors followed for a year the digital trail of 28,698 software engineers at a large technology company, from the internal launch of an AI coding tool. At twelve months, with company incentives to use it, 41% did. Among the women engineers, 31%. Licences and training were not missing: both were there.

From there comes the management consequence, and there is only one. If declaring the use has a price, people do not declare it. If they do not declare it, the company does not know where the tool came in or on which tasks. And a workflow nobody knows the shape of cannot be redesigned.

The same code, evaluated three times:

Same code, different mark: the difference is the label The author’s competence rating, compared with the same code carrying no label. THE SAME CODE, IN ALL THREE CASES COMPARED WITH THE SAME CODE −5% −10% −15% 0 baseline 0 −6% −13% among the harshest reviewers —engineers who do not use the tool— the penalty on women engineers was 26% larger than on men engineers NO AI USE DECLARED DECLARES AI USE MALE AUTHOR DECLARES AI USE FEMALE AUTHOR The code was the same in all three cases, and it was scored at the same quality. The only thing that changed was what the label said. Gai, P. J., Hou, J. and Tu, Y., Competence Penalty Is a Barrier to the Adoption of New Technology, SSRN 5255039; coverage in Harvard Business Review, August 2025. Preregistered experiment with 1,026 engineers evaluating identical code: the competence rating falls 9% on average, 6% for male authors and 13% for female authors, with no difference in the perceived quality of the code. It is a draft: it has not been through peer review. The field counterpart —28,698 software engineers at a technology company, 41% adoption at twelve months and 31% among the women engineers— is set out in reference 3.
Figure 3 — Same code, different mark. Declaring that the tool was used has a price, and it is not the same price for everyone. Source: Gai, Hou and Tu, Competence Penalty Is a Barrier to the Adoption of New Technology, SSRN 5255039 (preregistered experiment with 1,026 engineers; draft, not peer-reviewed).

And it is not one company

A company paying for the tool and adoption stalling anyway is not that company's quirk, and there is no average to copy either. The US Census Bureau's business trends survey asks whether the firm used artificial intelligence in any function in the previous two weeks: in the Information sector, 39.7% of firms say yes; in Retail trade, 14%. Same country, same month, almost three times the difference.

If inside a single country the gap is that size, the average for "the market" describes nobody's company, and that includes yours. There is no outside figure that tells you where you stand: the only diagnosis that works is the one on your own workflow.

What counts as finished work

None of the above gets fixed with an artificial intelligence use policy. The sheet that follows bans nothing and asks nobody to declare how they did their work: the previous section explains why that carries a measured and uneven cost.

What it does is move the definition of "finished" from format to content. Today a deliverable counts as finished when it looks finished: it has the sections, the length and the tone that were expected. That is the criterion a machine satisfies for free.

The same finding, written twice. Neither of the two texts below came out of a measurement: we wrote them ourselves, about the company in the scene, so that the difference in size can be seen.

Version A. "Analysis of the pharmacy channel reveals a combination of factors impacting profitability. Of note are the evolution of the product mix, growing competitive pressure and certain aspects linked to the current commercial policy. It is recommended that each of these axes be examined in greater depth, involving the relevant areas, with a view to defining an integrated action plan."

Version B. "The channel's margin fell four points. Three of those four come from one thing: the product that grew most is the one that leaves least, and it grew because in March we cut its price to hold on to two customers. One of those two now buys half of what it used to. If the price goes back to where it was in February, we lose that customer and we recover three points. That is the decision."

Version A is not badly written: it is well written. What it does not have is a single claim you can verify or a single decision you can take after reading it. Version B has four claims, and all four can be checked in the system in ten minutes. The difference between the two is not writing quality. It is whether anybody looked.

Four moves are enough, and each one has a named owner.

The request says which decision it unlocks. Whoever asks writes it, before asking, and it takes half a line. With no decision named, what was commissioned is pages.

The deliverable carries the claims that hold it up, with where to check them. Three, not thirty. The two sentences about the competitor that the manager could not verify before ten should have arrived with the source next to them.

The deliverable says what was left out. Two lines. The sentence "all factors were analysed" is the warning that none of them were.

Whoever receives it notes how many minutes verifying it took. It is the only number on the list, and today it does not exist in any company. Without it, the station that appeared between PRODUCE and USE is on no dashboard, and what is on no dashboard does not get redesigned.

Installing this costs half an hour with the three or four people who receive the most written work. Nothing needs to be bought and no policy needs to be written.

The four moves, with owners and a sheet already started:

A deliverable is not finished because it looks finished Four moves, what gets written in each one and who does it. THE MOVE WHAT GETS WRITTEN WHO DOES IT IF IT IS MISSING 1 · THE REQUEST SAYS WHICH DECISION IT UNLOCKS One line: «when this arrives, we will be able to decide ___» Whoever asks, before asking It is not a request: it is an order for pages, and it will come back full 2 · THE CLAIMS, WITH WHERE TO CHECK THEM Three claims that hold up the recommendation, with the source next to each one Whoever produces What arrives cannot be verified without redoing it 3 · WHAT WAS LEFT OUT Two lines. «All factors were analysed» is not an answer Whoever produces Nobody knows what was not looked at, and it surfaces after deciding 4 · HOW MANY MINUTES VERIFYING IT COST A number, noted down on finishing the read Whoever receives it, and nobody else The cost stays invisible, and the invisible does not get redesigned A SHEET ALREADY STARTED Request — why the margin on the pharmacy channel fell. Decision it unlocks — whether to reverse the March price or not. Claims to verify — the three that hold up the recommendation, with the query or the report next to each one. Left out — ___ · Minutes spent verifying — ___ Those two empty boxes are the whole measurement system you need to start. WHAT THIS SHEET IS NOT it is not an AI use policy · it does not ban the tool · it does not ask anyone to declare how they did their work · it does not replace the judgement of whoever receives · it is not announced at an all-hands. 32sur diagram. Row 4 translates the Humlum and Vestergaard finding (2026 revision) on the new integration and review tasks: if nobody counts them, they show up in no record. The bottom band —that the sheet does not ask anyone to declare tool use— translates the Gai, Hou and Tu finding (reference 3) on the competence penalty paid by whoever declares.
Figure 4 — A deliverable is not finished because it looks finished. Half an hour with the three people who receive the most written work. Nothing needs to be bought. 32sur diagram: row 4 translates the Humlum and Vestergaard finding (2026 revision) on the new tasks nobody counts, and the bottom band, the Gai, Hou and Tu finding on the penalty paid by whoever declares.

Questions for Monday

The first three can be answered right now, without opening anything or asking anybody.

Three from memory, four that take longer

  1. The last piece of written work you received that was no use to you: how long did it take you to work out that it was no use? Not the total time: the time until you worked it out. That is the cost nobody talks about.
  2. Did you tell the person who sent it, or did you redo it yourself and say nothing? Both answers are common; only one leaves a trace.
  3. Name the people in your company who review what goes out before it goes out. If they are the same three as always, you already know where the new work landed.

The four that follow take longer.

  1. Ask those three people how many minutes of last week went on verifying things that arrived ready-made. It is a number nobody has today, not even them.
  2. Open the last piece of work you commissioned in writing. Did it say which decision it was going to unlock when it arrived? If it did not, half the problem belongs to the request.
  3. Get the three or four people who receive the most written work into half an hour together and pick two of the four moves on the sheet to start with. Two, not four. Write them down in an email.
  4. A month from now, count how many deliverables came back and why. If the number does not fall, look at the request before you look at the tool.

Go back for a moment to that office in Córdoba. The ten o'clock meeting went ahead all the same, with the recommendation the manager put together on her own in ninety minutes, which is what it would have taken her if the report had never existed. Nothing fell over. That is the problem: nothing falls over, there is no incident to report, and the cost is paid somewhere nobody measures.

And Julián is still there, still not knowing that anything happened. The only thing he got that Tuesday was a message at sixteen minutes past eight:

"Thanks, I'll take a look and get back to you."

Nobody is going to tell him anything more: it is awkward and, strictly speaking, he delivered on time what he was asked for. Next time the request will be smaller, or it will go to somebody else. That is how a company loses somebody without firing them.

None of this gets fixed by buying another tool or banning the one you have. The two previous parts showed where the tool delivers and where it fails, and how to read the numbers behind the investment. This one showed that the saving you feel reappears, more expensive, two desks downstream, at a workstation nobody opened and nobody owns. All three point to the same place: the workflow the tool was dropped into without anything else changing. That is what Part 4, the last one, takes on, with the one lever the evidence links to the tool reaching the result and not just the feeling.

What does it cost you to verify what arrives ready-made?

Start without us: half an hour, three people, two empty boxes. This piece is enough to run that meeting on Monday. If a month later the deliverables are still coming back, the problem is in the workflow around the sheet, and redesigning workflows with the people who use them every day is what we do at 32sur. Write to us telling us which of your deliverables hurts most, and in the first reply we will tell you whether we need to come in or whether the half hour is all it takes.

Let's talk

References

  1. Niederhoffer, K., Rosen Kellerman, G., Lee, A., Liebscher, A., Rapuano, K. and Hancock, J. T., "AI-Generated 'Workslop' Is Destroying Productivity", Harvard Business Review, 22 September 2025 — the source that names the phenomenon. Research by BetterUp Labs with the Stanford Social Media Lab: an online survey of 1,150 full-time US employees, September 2025. Definition, verbatim: "AI generated work content that masquerades as good work, but lacks the substance to meaningfully advance a given task". 40% report having received at least one episode in the past month; those who received one estimate that 15.4% of everything reaching them qualifies as such; the reported time to resolve the last episode is 1 hour and 56 minutes; 42% came to see the sender as less reliable and 32% would rather not work with that person again; 18% of the episodes go from a report to their boss. Three points of precision this article carries in the body rather than in a footnote: it is self-reported evidence from an ongoing survey, not an experiment, and it is published by a company that sells leadership development; the article itself says "40%" in one paragraph and uses "41%" to compute the cost in another, and 40% is what is used here; and the translation of that time into a cost per person per month and into an annual cost for a ten-thousand-employee organization is not used in this article, because it is an estimate derived from the time and the salary reported by the same respondent. The article also opens by leaning on the 95% figure from the MIT Project NANDA report, which Part 2 of this series takes apart. Nor is the share of episodes circulating between peers used here, so as not to confuse it with the 40% prevalence figure: in the body that direction is stated in words.
  2. Humlum, A. and Vestergaard, E., "Still Waters, Rapid Currents: Early Labor Market Transformation under Generative AI", University of Chicago Booth / NBER Working Paper 33777, May 2025, revised version of March 2026 — the source this article rests on, read in the original. A version warning, which this series is obliged to give and which this article also carries in the body: the same work circulated in 2025 under the title "Large Language Models, Small Labor Market Effects" and with figures the revision changed (it ruled out effects larger than 1%, not 2%; it said 80% where it now says 85%; it grouped the new tasks differently). The NBER page itself makes it clear: "This paper was previously circulated under the title 'Large Language Models, Small Labor Market Effects'". The current version is the one cited here, and every figure was taken from it. Two survey rounds (end of 2023 and end of 2024); the latter gathers 25,000 workers at 7,000 workplaces in eleven exposed occupations —accountants, customer support specialists, financial advisers, HR professionals, IT support, journalists, lawyers, marketing, office clerks, software developers and teachers—, linked to the Danish administrative records of earnings, hours and occupation; identification is difference-in-differences, with employer policies as quasi-experimental variation. Verbatim from the abstract: "using difference-in-differences, we estimate precise null effects on earnings and recorded hours at both the worker and workplace levels, ruling out effects larger than 2% two years after the launch of ChatGPT". The reported saving is verbatim from the body: "adopters in our sample report savings of about 3% of their work hours" — the earlier version also carried a 2.8% figure in the body that the current one no longer reports. The other data this article uses are also in the body of the paper: 85% of users reassign the time saved to other work tasks and fewer than 10% use it for pauses or rest; among the new tasks, integration of the tool into the work is the single most frequent category at 26% ("the single most common new AI task category is AI Integration"), while the 35% for quality review plus ethics and compliance groups two categories and the 42% for content generation groups three (ideation, drafting and reading data) — the 42% exceeds the 26% because it adds categories, not because it competes with it, which is why the body says "single category"; close to 8% of users report entirely new tasks where the employer does not push the tool, and around 17% where it does; 4% of non-users report new load, with 10% among the teachers who never used it. It is a draft: it has not been through peer review, and this article says so in the body. Data from the study this article does not use in the body: 43% of companies explicitly encourage use, 21% permit it and 6% prohibit it; among those that encourage it, 61% have an enterprise version and 39% provide training; where there is no initiative at all, around 40% of workers have used a chatbot at work anyway; and where encouragement, an enterprise version and training coexist, 93% report having used it and 28% use it daily. The paper also reports, verbatim, that the enterprise version and training implemented in isolation are associated with lower reported benefits ("enterprise chatbots and training—when implemented in isolation—are associated with lower reported benefits"); this article does not use that finding because the work itself offers two alternative explanations for it and does not choose between them.
  3. Gai, P. J., Hou, J. and Tu, Y., "Competence Penalty Is a Barrier to the Adoption of New Technology", SSRN 5255039; coverage in Harvard Business Review, August 2025 ("Research: The Hidden Penalty of Using AI at Work") — the competence penalty. A preregistered experiment with 1,026 engineers evaluating the same code: those who declared having used artificial intelligence received competence ratings 9% lower on average —6% for male authors and 13% for female authors—, with no change in the perceived quality of the code. The harshest reviewers were the engineers who had not adopted the tool, and their penalty on the women engineers who used it was 26% larger than on the men who used it. Field counterpart, from the same work: the digital trail of 28,698 software engineers at a technology company after the internal launch of a generative AI coding tool; at twelve months 41% were using it, with 31% among the women engineers and 39% among those over forty. A complementary survey (n = 919) links the anticipated penalty to slower adoption. It is a draft hosted on SSRN: as of this article's closing date it does not appear published in a peer-reviewed journal. Two corrections to secondary coverage in circulation: the 41% is the adoption rate across all engineers, not the rate among men (which is 43%); and the company is described in the source as a leading technology company, with no position in any ranking — this article attributes none to it.
  4. U.S. Census Bureau, Business Trends and Outlook Survey, May 2026 report on artificial intelligence use in business — the question asks whether the firm used AI in any function in the previous two weeks. Sector breakdowns: Information 39.7%, Finance and insurance 33.9%, Retail trade 14%; national average 19.8%. The breakdowns by company size and the general adoption range were used in Part 2 of this series and are not repeated here.
  5. CEPE-Universidad Torcuato Di Tella (the public policy centre of its School of Government) and Fundar, nadIA initiative, with fieldwork by Fundación Observatorio PyME and support from the IDB, "Encuesta Nacional sobre Adopción de IA en Pequeñas y Medianas Empresas de la Argentina", April 2026 — 402 companies of 10 to 249 employees in manufacturing and in software and IT services, with fieldwork between November and December 2025 and stratified sampling across 14 sector strata. The three figures this article uses, computed over the companies already using AI: 91.2% did not cut headcount, 95.5% brought in no specialist profile and 44.4% report that different areas and individuals adopt on their own, with no central coordination.
  6. Dell'Acqua, F. et al., "Navigating the Jagged Technological Frontier", Organization Science, 37(2), 2026, pp. 403-423 — cited here for one reason only: on the task where the groups with artificial intelligence got it right less often, the graders scored their deliverables as more coherent and better argued than those of the group without the tool. The figures for that finding are in Part 1 of this series and are not repeated in this article.
  7. Choi, J., Hecht, G. and Tayler, W. B., "Lost in Translation: The Effects of Incentive Compensation on Strategy Surrogation", The Accounting Review, 87(4), 2012 — the origin of the concept of surrogation, which this blog has already used when writing about dashboards and indicators: substituting the measure that represents something for the thing that matters. Here it appears in its new form, in which the appearance of the deliverable can be produced by a machine.

All sources were read in their primary available version. Last verified: August 2026.