The sign in week five
Haifa, 1998. Ten private day-care centres, twenty weeks and a problem anybody will recognise: some parents arrive late to collect their children, and when that happens somebody has to stay behind. That somebody is always the same person: a teacher whose working day is already over, waiting at the door with a child who also wants to go home.
For the first four weeks nothing out of the ordinary happens. Two researchers count, week by week and centre by centre, how many parents arrive late. They do nothing but count.
In week five, six of those ten centres —drawn at random, not chosen— wake up with a sign on the door: anyone ten minutes late or more pays a fine, per child. A family with two children pays twice. The other four have no sign and carry on as before: they are the comparison group.
The fine is small, and it was clearly designed to be small: it comes to less than one per cent of what that family pays in monthly fees, and less than an hour of babysitting in that city. Next to other Israeli fines in force in 1998 it looks smaller still: running a red light cost a hundred times as much.
And there is a detail the paper reports in a single line, without underlining it. The fine is collected by the owner of the centre. The teacher who stays behind waiting receives nothing.
Before we get to what happened, an inventory. In your company's rulebook, written at different moments, by different people, and never once read side by side, there are sentences from this family: attendance bonus, deduction for arriving late, reward for referring a candidate, allowance for wearing protective equipment, penalty for not entering data into the system, liquidated damages for the supplier who delivers late, surcharge for the customer who pays late. None of them was designed alongside the others. All of them do the same thing: they put a price on a behaviour.
This is not a quirk of your company. In a survey of 561 privately held companies in the United States, 93% reported having a short-term incentive programme. It is, quite literally, the payroll operating system of almost all of them.
In What gets measured gets managed we looked at the dashboard. Here we look at the wallet, which falls apart rather worse because it has been signed.
And behind all those sentences there is a single idea, so reasonable that almost nobody bothers to state it: if something costs, people do less of it. It is the intuition of any owner and of any legislator. It is what was put to the test in Haifa. It is the first thing that broke.
What happened once the sign went up
With the sign up, late pick-ups did not fall. In the six centres drawn at random the weekly average per centre went from 8.0 to 16.4: more than double. In the four without a sign, over those same weeks, it barely moved: from 10.0 to 9.2.
One fact about the starting point that almost never gets told: the six centres that received the fine had been arriving late less often than the control centres, and ended up arriving late far more often than them.
In week seventeen the sign came down. Late pick-ups did not come back: they climbed to 18.7, the highest reading of the twenty weeks, with no fine in force. The centres where there had never been a sign closed at 8.3.
What changed was not the people. It was what kind of thing those ten minutes of the teacher's time were. The fine did not add a cost to the guilt: it put a price on it, which is a different matter. As long as they were a favour, those ten minutes lived at the counter: you ask for them, somebody gives them to you or does not, and you are left owing. With a price they move to the shelf: you pick them up and you pay. They are no longer something that had to be repaid, but a service with a tariff, and a cheap one.
And taking a product off the shelf does not put it back on the counter: the information that this thing had a price was already in every parent's head, and nobody had any way to erase it.
That line from the first page still stands: the one paying the cost was the teacher and the one collecting the price was the owner. From there comes a rule for any line of your rulebook: if whoever pays the cost does not receive the payment, you are not correcting a behaviour. You are collecting a toll.
The twenty weeks, in two lines:
"Yes, but that's a day-care centre"
They are parents, not employees. It is a service relationship, not an employment contract. And above all: the fine was set by a third party. It is not a company.
The objection is a good one, and it is not sour grapes: it was the state of the art for twenty years. Canice Prendergast wrote it in 1999, in the field's journal of record: the idea has intuitive appeal, but conclusive empirical evidence is missing, particularly in workplaces.
In 2024 it stopped being missing.
A German retail chain: 346 employees on training contracts, one or two per store, spread across 232 stores. The German apprentice is not an intern: it is an employee with a two- or three-year contract and a collectively bargained wage. They were drawn into three groups: one carried on as before, a second was offered a cash bonus for every month with perfect attendance, and the third the same value in time off.
The hypotheses and the analysis plan were deposited in a public registry before the start, so that nobody could later pick which result to tell. That is called preregistration: the difference between a measurement and an anecdote with a spreadsheet.
Each month with perfect attendance earned a point and every three points were worth a reward, but nothing was paid out until the experiment ended. The annual maximum came to more than a quarter of a monthly wage: it was not a branded pen, and it was not an extra salary either.
The cash bonus raised absenteeism by around 50%. In days, more than five additional absences per person per year. The time-off bonus, of the same value, produced no effect that can be asserted, which is not the same as saying it worked: it is that no effect was found.
The effect was driven by the ones who had joined most recently, the ones still working out what is done and not done there. The bonus ran through the twelve months of 2018 and ended there; in the first half of 2019, with no bonus in place, absenteeism did not return to its starting point. What had changed was how acceptable missing work seemed to them, something the authors call a lasting erosion of the norm. It is Haifa again, with a payroll, in this century, even if the people involved are young trainees in German shops.
A fine and a bonus look like opposites and are not: both put a price on something that did not have one. The sign of it makes no difference; what matters is that there is now a price in money, and that is exactly what the time-off arm isolates.
The Argentine attendance bonus is, instrument for instrument, the same one measured in those stores: cash in exchange for turning up to work. Here it has been institutionalised for decades and nobody has ever measured it.
Paying a little does worse than paying nothing
The next objection is the most reasonable of all: so the bonus was too big. The other way round.
The same two researchers, another experiment in Haifa. A hundred and sixty students, fifty questions from an admissions test and the same payment for taking part; the only thing that differed between groups was how much was paid per correct answer. Those who were paid nothing got 28 questions right on average; those paid a few cents, 23. Paying a little left them below paying nothing, and only with a payment ten times larger did they beat the group that was paid nothing at all.
These are students sitting a test, not a sales force: the measurement with a payroll is the one in the previous section. What this one adds is the edge, and the edge is where almost everything your company has in place lives: a low price is still a price, and it does all the damage of a price with very little of its benefit. There is no such thing as a mild version.
Three pay schemes that were actually measured, and in none of the three did the money buy what the document said it was buying.
None of the three measured people with no appetite for work. In all three, people did exactly what the scheme paid for.
The 44% everybody quotes, and the half nobody quotes
Money does buy behaviour, and there is a field experiment that measures it with a cleanness almost no management study has.
A US windscreen fitting chain moved to piece rate, region by region, over nineteen months. They measured 2,755 fitters. Output per worker rose 44%.
The small print is more interesting than the headline. Roughly half of that 44% is not people trying harder: looking at each worker against himself, before and after the change, the same fitter does 22% more. The other half is the composition of the payroll, who works there. And the hires that came after the change weighed more than the departures: a pay-for-output scheme publishes, without meaning to, a permanent advert saying who it suits to work here.
The worker took his share too, around 10% more pay. It was not an elegant way of squeezing.
What never gets reproduced is missing. That scheme carried three things the paper describes as parts of the programme that was measured, not as recommendations from the author.
A floor. Anyone moving to piece rate was guaranteed, as a minimum, roughly what they earned before. Nobody bet their salary.
A quality clause with teeth. Anyone who fitted badly refitted on their own time and paid the company for the replacement glass before receiving a new paid job. The defect cost money to whoever had caused it.
A measure nobody disputed. A computer system counted the units each person fitted per week. The figure was not negotiated.
Almost nobody who quotes the 44% knows those three things existed. Without all three, what you have is not the scheme that produced that number: it is a different scheme, and the published result does not apply to you. They are windscreen fitters and what is counted is units installed, which is exactly the point: simple, repeated, countable work.
The same person, the same session, two tasks
If money buys output, why does it not buy judgement? There is an experiment that answers this, and what makes it hard to dodge is the design rather than the result. Twenty-four people, each of whom did both tasks and did both with a small prize and with a prize ten times larger: there are no two groups that could differ in anything else, it is the same person, the same session.
The first task was pure effort: pressing two keys alternately for four minutes, deliberately mind-numbing. With the small prize they delivered 39.9% of the maximum possible; with the prize ten times larger, 77.9%. Money worked exactly as the textbook says.
The second required thinking: adding up matrices. With the small prize they delivered 62.9%. With the prize ten times larger, 42.9%. It flipped.
Seventeen of the twenty-four people did worse with the large prize on the task that required thinking.
It is twenty-four people in a laboratory: what travels to your company is the shape of the effect, not its size. But the shape reappears at scale. A review of thirty-nine studies found that financial incentives are associated with how much gets produced and not with how well: when you put a commission in place you are buying volume, even if you wrote "quality" in the title of the document. A later meta-analysis —the study that pools the studies— covering 183 papers and forty years of data put it in one sentence: the incentive predicts how much; what the person brought predicts how well.
The year both answers were published. In 2000, Edward Lazear published the windscreen results in the American Economic Review and read his own data as closing the debate. That same year, Uri Gneezy published the two studies that hold up the first half of this piece. For two decades the standard reply to Gneezy was Prendergast's; in 2024, the German chain supplied what was missing. And both measured well, on different behaviours: Lazear measured windscreen fitters, where before piece rate nobody fitted glass out of moral duty; Gneezy measured parents, teachers and students sitting a test, where the behaviour was already held up by something that was not money.
The Sunday spreadsheet
One Sunday in 2018, close to eleven at night, the kitchen table. Three cells: 3% on cash collected, rising to 4% on everything above the quarterly target, with no cap. On Monday at 8:14 an email goes out with the subject line "new scheme". And it was a good decision: the team understood it in two minutes and revenue moved.
(That Sunday did not happen. The Sundays that inspired it did, and all of them ended in an email on Monday morning.)
Eight years later, that file decides three things that were not in the three cells.
Who works there. A commission scheme is not primarily a motivator: it is a selection filter. It defines who applies, who accepts and who survives the first weak quarter. In the windscreen case, half the jump came through that channel. Before asking how much it pays, ask who it attracts.
What gets sold. What is commissioned gets sold. What is profitable gets sold when it is also commissioned. Nobody needs to act in bad faith; what is needed is a percentage applied to the wrong line.
When deals close. A study of a large enterprise software vendor with accelerated commissions found that salespeople manage the timing of the close so that it falls in the quarter that suits them, and accept lower prices when they have a financial incentive to close. Those badly timed prices cost the vendor between 6% and 8% of revenue. Nobody steals and nobody slacks: they do what the scheme pays for.
The article on pricing policy explained where price leaks away. This piece explains who gets paid for leaking it.
And there is a local layer none of these studies had to solve. A percentage on nominal sales pays the salesperson for the price index as well as for their work, and a target set in pesos in January is a different target in December. Both distortions run the same way and do not cancel out: the commission overpays for a sale that did not grow, and the target loosens by itself with every month that passes. Nobody decided that; the scheme was written in a currency that will not stay still. Nobody has redesigned it because nobody looks at it: the Sunday spreadsheet has no owner.
Before you touch a peso
The useful question is not whether incentives work. It is what was governing this behaviour before you put a price on it. From there comes a conclusion that is smaller and more uncomfortable than removing incentives: stop putting them in and taking them out as if they were reversible. There are four questions, and all four are answered inside your own company.
What does it buy today? Not the behaviour you meant to buy: the one people perform in order to collect it. They almost never coincide, and the difference is found out by asking.
What does it displace? If the behaviour was already happening without a price, the price does not add to it: it takes its place. It is the only one of the four that has to be answered before signing, because afterwards it can no longer be done.
Does it have the three conditions? A floor, a quality clause that costs money to whoever caused the defect, and a measure nobody disputes. But the order matters: if the answer to 2 is that nobody did it without a price, only then are the three conditions any use, and each "no" is a piece of the result you do not control. If they were already doing it, the three conditions fix nothing: the problem is the price.
What happens the day you remove it? In Haifa they removed it and nothing went back to where it had been. If you cannot answer this one, you are still missing half the design.
None of this requires a new compensation policy.
The half of every objection that is true
When these measurements reach a meeting room, four sentences show up, almost always in this order. None of them is stupid.
All four have their half that is true. None of them survives whole.
| What will be said | What it runs into | What is left standing |
|---|---|---|
| "If it hits their pocket, they stop doing it" | In the day-care centres with the fine, late pick-ups doubled; in the ones that never had a fine, they did not move (Gneezy and Rustichini, 2000) | That the price is felt. Not that it corrects. |
| "It is a small bonus, if it does not work nothing happens" | A bonus for perfect attendance raised absenteeism by around half, and paying a few cents per answer did worse than paying nothing at all (Alfitian and others, 2024; Gneezy and Rustichini, 2000) | That it is small. Small is not neutral: a low price is still a price. |
| "If it goes wrong, I take it out and that's that" | They withdrew the fine in week seventeen and late pick-ups climbed to the highest level of the twenty weeks; the German bonus also left a trace after being withdrawn (Gneezy and Rustichini, 2000; Alfitian and others, 2024) | That it can be taken out. Not that you go back to the starting point. |
| "The commission worked for me: output went up" | Output per worker rose 44% with piece rate, and roughly half of that difference came from who came to work, not from who tried harder (Lazear, 2000) | That quantity went up. What is missing is thanks to which three conditions, and at the cost of what. |
All four sentences have an owner, and the owner almost always signed the scheme.
Questions for Monday
Three from memory, two you have to ask, two that cost a meeting
The first three are answered right now, from memory and out loud, without opening a file.
- Name every price your company currently has in place on the behaviour of its own people. If the list runs past five, it is certain that none of them was designed alongside the others.
- Of that list, which ones put a price on something people were already doing without anybody paying? Those are the ones to look at first, and the only ones with no free way back.
- For each one: who collects and who pays the cost of the behaviour? If they are not the same person, you are not correcting anything. You are collecting a toll.
The next two require asking somebody else something.
- Ask somebody who is paid on commission what they do, concretely, to earn more. Not what they should do: what they do. That answer is the real design of your scheme, and it does not match the spreadsheet.
- Ask the finance office what surcharge it charges a customer who pays late today, and compare it with what it costs you to get that same money a month earlier. If the surcharge is lower, you do not have a penalty: you have a line of credit you opened without noticing.
And two that are not read: they are done. They cost a meeting.
- Take the target and the commission currently in force and restate them in real terms: how much that target bought the day it was set and how much it buys today, and how much of what you paid in commissions last quarter was put there by the salesperson and how much by the price index. If they were never repriced, your scheme has an indexation clause nobody wrote.
- Before adding the next incentive —the one you already have in mind—, write on one line what behaviour you want to buy, and on another what will happen the day you want to remove it. If the second line does not come out, you have not designed it yet.
In week seventeen, the six Haifa day-care centres took the sign down. It was a sensible decision: the fine had not worked, so they removed it. Parents carried on arriving late, later than ever, now without paying anything at all. The sign could be taken down; the price could not. Once it was known what ten minutes of the teacher's time were worth, there was no way back to not knowing.
And in the four centres where nobody hung anything up, late pick-ups did not move. Nothing special happened there. It simply carried on having no price.
The Sunday spreadsheet was also a sensible decision, taken by somebody who knew their business with the information they had to hand. The problem is not how it was written: it is that it is still in force eight years later, deciding who comes to work and what gets sold, without anybody having opened it again.
Was yours written one Sunday and never looked at again?
In ninety days there are three documents that do not exist today. One: the complete inventory of the prices your company has in place on its people's behaviour, in two columns —the ones that buy quantity and the ones that put a tariff on an obligation—. Two: the commission scheme currently in force, run through the three conditions, with the gaps noted. Three: a new rule, that no incentive is announced without the four questions having been answered. None of the three comes out of a compensation policy. They come out of a two-hour meeting, with the scheme currently in force open on screen and the person who gets paid by it sitting at the table. That meeting is where 32sur comes in. Let's talk.
References
- Gneezy, U. and Rustichini, A., "A Fine is a Price", The Journal of Legal Studies, 29(1), 2000, pp. 1-17 — the scene that opens and closes this article. Ten private day-care centres in Haifa between January and June 1998, twenty weeks: no intervention for the first four; in week 5 a fine per child is introduced for delays of ten minutes or more in six centres drawn at random, with four as controls; in week 17 it is withdrawn. Weekly averages of late pick-ups computed from Table 2 of the paper: fine group 8.0 (weeks 1-4), 16.4 (5-16) and 18.7 (17-20); control group 10.0, 9.2 and 8.3. The study's text sums up the effect by saying late pick-ups came to be almost double the initial number. The fine came to around 0.7% of the monthly fee per child; the paper itself compares it with other fines in force in Israel in 1998 and with the cost of an hour of babysitting. The study also records that the money was collected by the owner of the centre, not by the teacher who stayed behind, who received nothing extra.
- Alfitian, J., Sliwka, D. and Vogelsang, T., "When Bonuses Backfire: Evidence from the Workplace", Management Science, 70(9), 2024, pp. 6395-6414 (DOI 10.1287/mnsc.2022.00484) — the measurement that closes a twenty-year discussion. Preregistered field experiment (AEA RCT Registry AEARCTR-0002863) in a German retail chain: 346 apprentices from 232 stores randomly assigned to control, a monetary attendance bonus or a time-off bonus. The experimental period was the twelve months of 2018; the first half of 2019, with no bonus in force, is the follow-up window in which the authors measure persistence. The monetary bonus added one point for every month of perfect attendance —up to twelve in the year—, every three points were worth one reward unit and the annual maximum came to more than a quarter of an apprentice's typical monthly wage; points were converted into rewards only once the experiment ended, with quarterly information on the running total. It raised absenteeism by around 50% on average, that is more than five additional days of absence per employee per year; the effect was driven by the most recent joiners and persisted after the bonus was withdrawn. The time-off bonus, calibrated to an equivalent value, produced no conclusive effect. The post-experiment survey shows a shift in how acceptable missing work is perceived to be, which the authors describe as a lasting erosion of the social norm.
- Gneezy, U. and Rustichini, A., "Pay Enough or Don't Pay At All", The Quarterly Journal of Economics, 115(3), 2000, pp. 791-810 — the edge of the mechanism. Two experiments. In the first, 160 University of Haifa students answered fifty questions from an admissions test with a fixed payment for taking part and a marginal payment per correct answer that varied by group: with no marginal payment the average was 28.4 correct answers; with the smallest marginal payment, 23.07; with a payment ten times larger, 34.7; and at triple that amount, 34.1. In the second, 180 secondary school students went out on the annual door-to-door collection: with no commission they raised an average of 238.67; with a 1% commission, 153.67; with a 10% one, 219.33. Participants were told that the commission was paid by the researchers and not by the beneficiary organisations. This second experiment is not used in the body: the article rests on the first one only.
- Lazear, E. P., "Performance Pay and Productivity", American Economic Review, 90(5), 2000, pp. 1346-1361 — the Safelite Glass study, and the most abused figure in this territory. 2,755 fitters, 29,837 person-month observations, nineteen months between 1994 and 1995, with the move to piece rate implemented region by region. The effect on output per worker is 44%, of which the author himself attributes roughly half to an incentive effect —the same worker fits 22% more after the change— and the rest to the payroll shifting towards more capable workers, mainly through new hires. Worker pay rose by around 10%. The programme included three elements this article highlights and which are almost never quoted: a floor guaranteed at roughly the previous wage, a quality clause under which anyone who fitted badly refitted on their own time and paid for the replacement glass before receiving new paid work, and a computer system that recorded units fitted per person per week.
- Ariely, D., Gneezy, U., Loewenstein, G. and Mazar, N., "Large Stakes and Big Mistakes", The Review of Economic Studies, 76(2), 2009, pp. 451-469 — experiment 2, at MIT: 24 undergraduates, each of whom did both tasks with a low prize and with a prize ten times larger. On the pure physical effort task (alternating two keys for four minutes) they delivered 39.9% of the maximum with the low prize and 77.9% with the high one; on the cognitive task (adding up matrices), 62.9% with the low prize and 42.9% with the high one. Seventeen of the 24 participants did worse with the high prize on the cognitive task, three did not change and four improved. Experiment 1 in the same paper, with 87 participants from a rural area near Madurai, India, and prizes of up to a hundred times the smallest, shows the same pattern with very large prizes; this article does not use it in the body.
- Jenkins, G. D. Jr., Mitra, A., Gupta, N. and Shaw, J. D., "Are Financial Incentives Related to Performance? A Meta-Analytic Review of Empirical Research", Journal of Applied Psychology, 83(5), 1998, pp. 777-787 (DOI 10.1037/0021-9010.83.5.777) — 39 studies and 47 relationships: financial incentives are associated with the quantity of performance and do not appear associated with its quality. The article uses the finding in qualitative terms and does not quote coefficients, which is how it is written in the body.
- Cerasoli, C. P., Nicklin, J. M. and Ford, M. T., "Intrinsic Motivation and Extrinsic Incentives Jointly Predict Performance: A 40-Year Meta-Analysis", Psychological Bulletin, 140(4), 2014, pp. 980-1008 — 183 studies and more than 212,000 people. The two conclusions this article uses: incentives better predict the quantity of performance and intrinsic motivation better predicts its quality; and when incentives are directly tied to performance —the paper explicitly mentions sales commissions and year-end bonuses— intrinsic motivation loses predictive power, which is the meta-analytic version of what the Haifa day-care centre shows in a single scene.
- Larkin, I., "The Cost of High-Powered Incentives: Employee Gaming in Enterprise Software Sales", Journal of Labor Economics, 32(2), 2014, pp. 199-227 — under an accelerated commission scheme, salespeople manage the timing of the close so that it falls in the quarter that suits them and accept lower prices in the quarters in which they have a financial incentive to close; the cost to the vendor is estimated at between 6% and 8% of revenue.
- Prendergast, C., "The Provision of Incentives in Firms", Journal of Economic Literature, 37(1), 1999, p. 18 — the standard objection for twenty years: the idea that paying for an activity reduces intrinsic interest has intuitive appeal, but there is little conclusive empirical evidence, particularly in workplaces. The quotation is literal and has been checked against the original in the JEL; the paper in entry 2 reproduces it in the same terms.
- WorldatWork and Compensation Advisory Partners, Incentive Pay Practices: Privately Held Companies, 2021 — 561 complete responses from privately held companies, of which 93% report having a short-term incentive programme (99% in 2019; 94% in 2015). The paper in entry 2 quotes that figure as 94%, which is the value from the 2015 edition of the same series. The original 2021 report says 93%, and that is the number this article uses.
All ten entries were checked against the published primary source. Entry 10 also corrects the figure as reproduced by the paper in entry 2.