The judge's dice
A group of young judges —law graduates who have already handed down their first rulings— are given a real case file: a woman arrested for repeated shoplifting. They read the whole thing. Before setting the sentence, they roll a pair of dice. The dice are loaded: half the group always rolls a 3, the other half always rolls a 9. The number, they are told, represents the prosecutor's sentencing request, in months. Everyone in the room knows a die is not a legal argument. Then they sentence.
Those who rolled a 3 handed down, on average, 5.28 months in prison. Those who rolled a 9, 7.81 months. The same woman, the same file, the same trainee judges: almost 50% more prison time because of a number that came out of a die —a die they had watched being rolled themselves. The study is by Birte Englich, Thomas Mussweiler and Fritz Strack (2006), and in its other variants the effect shows up just the same among prosecutors and judges with more than ten years on the job.
In case the reader suspects this is a courtroom curiosity, an earlier classic took it to the property market. Gregory Northcraft and Margaret Neale (1987) invited professional appraisers to walk through a real house with a full ten-page dossier —comparables, measurements, everything. The only thing that changed between groups was the listing price. With a low listing price, the professionals valued the property at US$67,811 on average; with a high one, US$75,190. Afterwards, only 24% mentioned the listing price among the factors they had weighed —and the experts denied its influence more emphatically than the amateurs did.
This is what makes the central finding of decision psychology so uncomfortable: biases are not a beginner's flaw that experience cures, nor an intelligence deficit that talent compensates for. They are properties of the standard equipment. And if a die can move a sentence and a listing price can move a professional appraisal, it is worth asking what the first number spoken in the room does to your budget, your valuation of an acquisition, or your five-year plan.
That is what this series of three articles is about: how to decide when you cannot be sure. This first part looks at the apparatus —the systematic defects of judgment and their antidotes. The second looks at the terrain: why not every decision should be decided the same way, and how to read which territory you are standing on. The third is the final arbitration: when to trust expert intuition, when to trust a rule, and what to do about noise.
Two speeds, one driver
The most useful map of the apparatus was popularized by Daniel Kahneman —psychologist, Nobel laureate in Economics in 2002— in Thinking, Fast and Slow (2011). The mind runs in two modes. System 1 is fast, automatic, associative and tireless: it recognizes faces, completes patterns, produces impressions, and cannot be switched off. System 2 is slow, deliberate and effortful: it calculates, compares, verifies —and it is lazy: it steps in as little as it can, and tends to rubber-stamp whatever System 1 has already decided. These are not two brains, or two kinds of people: they are two speeds of the same driver.
The point management usually skips is this: System 1 does not retire when you get promoted. It runs the same in the analyst and in the CEO, with one dangerous difference: the CEO has fewer people willing to contradict him.
Two properties of System 1 explain a good share of strategic disasters. The first is what Kahneman christened WYSIATI (what you see is all there is): the mind builds the most coherent story it can with the information available, and never registers —or discounts for— the information that is missing. The second follows directly: the confidence we feel in a judgment measures the coherence of the story we assembled, not its truth. A tidy business case, with a round narrative and no loose ends, feels solid —even when it rests on three unverified assumptions. The sense of "I've got this clear" is an emotion, not a diagnosis.
What held up from the book (and what didn't). Thinking, Fast and Slow is fifteen years old, and psychology's replication crisis happened in the meantime. It is worth saying plainly: the chapter on social priming —the studies where reading words about old age made people walk more slowly— aged badly; many of those experiments failed to replicate, and Kahneman himself admitted it publicly in 2017: "I placed too much faith in underpowered studies". But the core this article draws on —anchoring, overconfidence, the planning fallacy, escalation of commitment, loss aversion— is among the most replicated material in the discipline: anchoring, for instance, came out as one of the most robust effects in the Many Labs project (2014), which redid classic studies across dozens of laboratories. The lesson about method holds for companies too: evidence is not collected, it is curated.
Optimism as company policy
Start with the most expensive bias. In 1994, Roger Buehler, Dale Griffin and Michael Ross asked students to estimate when they would hand in their thesis. The average estimate was 33.9 days; the reality, 55.5 days —64% longer— and only 29.7% finished on the date they had predicted. The striking part is not the optimism: it is that not even the worst case was enough. Asked to estimate "assuming everything went as badly as it possibly could", fewer than half (48.7%) were done even by that date. Kahneman and Amos Tversky called this the planning fallacy: we plan from the inside view —the mental movie of how this project is going to unfold— and ignore the statistics of comparable projects.
Kahneman told the story of his own stumble with it. A team he belonged to estimated it would finish an educational textbook in about two years. He asked the group's curriculum expert how long comparable teams had taken: 40% never finished, and those that did took between seven and ten years. The team heard the number, did not use it, and took eight years. The expert held the outside view —the base rate for that class of project— and still planned with the inside view, like everyone else.
At corporate scale, the planning fallacy keeps books. Bent Flyvbjerg (Oxford) has spent decades assembling the largest database of major projects in existence —more than 16,000 projects. The results: 47.9% come in on budget; 8.5% hit budget and schedule; and a bare 0.5% deliver budget, schedule and the benefits promised. One in two hundred. Average cost overruns vary by type —solar +1%, wind +13%, rail +39%, IT +73%, nuclear +120%, Olympic Games +157%— and they conceal fat tails: in IT, the 18% of projects that overrun by more than 50% overrun by an average of +447%. Projects do not fail "a bit more than expected": either they land within reason, or they blow up.
And the people who know most? Itzhak Ben-David, John Graham and Campbell Harvey collected 13,300 forecasts of S&P 500 returns made by CFOs of large companies between 2001 and 2011, each with its 80% confidence interval. If those CFOs were calibrated, the actual outcome would land inside the range 8 times out of 10. It landed there 36.3% of the time. Not even in the calmest quarter of the decade (59%) did they come close to 80%. It is not that CFOs do not understand markets: it is that the range the mind produces spontaneously is far too narrow —the coherent story leaves out the scenarios it never imagined (WYSIATI again).
That same overconfidence buys companies. Ulrike Malmendier and Geoffrey Tate measured CEO overconfidence with an elegant behavioural indicator —executives who hold their own stock options to expiry, over-betting on their own management— and found that their odds of making an acquisition are 65% higher than their peers', above all in diversifying deals financed with cash. The market, which has already learned, reacts to the announcement: −90 basis points on average for the overconfident acquirer, against −12 for everyone else. Back in 1986 Richard Roll had proposed the "hubris hypothesis": much of merger activity is explained not by synergies but by managers convinced that they, at least, will pull it off. The evidence since has kept proving him right.
Anchors, sunk costs and double standards
The anchor. We have already seen what a die does to a sentence and a listing price to an appraisal. Inside a company, the anchor goes by proper names: last year's budget (which frames the discussion of next year's), the first figure spoken in a negotiation, the valuation that appears in the banker's first draft. The practical rule is unpleasant but real: whoever puts the first number on the table is in charge —unless the process neutralizes it. And neutralizing it is not a matter of "being careful" (the appraisers were being careful): it is structural —independent estimates, in writing, before anyone sees the other side's number; ranges instead of single figures; discussion afterwards.
The sunk cost. Barry Staw (1976) had people make investment decisions in a business case with two divisions. When the division the participant had chosen himself performed badly, he assigned it US$13.1 million out of a US$20 million fund in the following round —against roughly US$9 million when the initial choice had been made by someone else. This is escalation of commitment: the more responsible I am for the course that failed, the more resources I throw at it rather than admit the error. The visible reasoning is always the same —"we have already invested too much to walk away"— and it is exactly the argument decision theory disqualifies: the sunk cost does not come back, whatever you decide; the only thing that counts is the future of each new dollar. Andy Grove described how Intel escaped the trap in 1985, when memories —the founding business— were bleeding: he asked Gordon Moore, "if we got kicked out and the board brought in a new CEO, what would he do?". The answer was obvious —get out of memories— and the question made it sayable. The successor test is still the cheapest de-anchoring device there is.
The double standard. Loss aversion —losses weigh, on average, about twice as much as equivalent gains; among the most replicated findings of Kahneman and Tversky's prospect theory— produces inside organizations a paradox Kahneman and Dan Lovallo labelled "timid choices and bold forecasts" (1993). In the plan, the company is optimistic (inside view, WYSIATI). In the portfolio, it is timid: each manager, looking at his project in isolation and at his own skin, turns down risks the company as a whole should want. Richard Thaler told it with a scene that became a classic: he asked 25 division managers of the same group whether they would accept a coin flip —a 50% chance of winning two million, a 50% chance of losing one; expected value clearly positive. Three accepted. The CEO, in the room, wanted all of them: at portfolio level, 25 coin flips like that are close to a sure gain. The problem was not risk; it was narrow framing —each manager was deciding his own coin, not the portfolio— and the system of rewards and punishments that keeps it in place.
The yardstick applied afterwards. A perverse complement: outcome bias. Jonathan Baron and John Hershey (1988) showed that we judge the quality of a decision by how it turned out, not by how it was made —the same decision, with the same information, is rated competent if it worked out and reckless if it didn't. An organization that rewards and punishes on outcome alone teaches its managers two things: never take any coin flip, and dress up the process afterwards. Assessing the quality of the decision process —what was known, which alternatives were examined, which risks were made explicit— is not leniency: it is the only way for the system to learn judgment instead of luck.
The outside view (the antidotes that work)
If biases could be cured with intelligence or good intentions, there would be no 0.5%. The bad news from the field is unanimous: knowing about biases does not immunize you against biases —Kahneman was the first to admit that, after fifty years of studying them, he kept falling for them. The good news is just as firm: what cannot be fixed in the head can be fixed in the process. Four tools with evidence behind them, from the cheapest to the most structural:
1. The outside view. Before arguing about why this project is different, answer a boring question: how did the last twenty projects of this class do —ours and other people's? That number (the base rate) is the correct anchor; you adjust for the specifics afterwards. It is the method Flyvbjerg turned into a standard (reference class forecasting, now required for major projects by the UK Treasury, among others): forecast from the statistics of the class, not from the movie of the case. It hurts less than a +447%.
2. The premortem. Gary Klein's tool, praised by Kahneman himself as his favourite: with the team assembled and before approval, the facilitator announces "it is a year from now; the project has failed spectacularly", and each person writes down, alone and in two minutes, the reasons for the failure. The trick is not rhetorical: "prospective hindsight" —imagining the event as already having happened rather than as merely possible— raises by roughly 30% the number of concrete causes people generate (Mitchell, Russo and Pennington, 1989). And it does something more valuable still: it legitimizes doubt —in the approval meeting the sceptic is disloyal; in the premortem he is the one playing best.
3. De-anchor by design. Every estimate that matters —costs, schedules, synergies, valuations— is produced first independently and in writing, and only then discussed. Ranges with probabilities instead of single numbers ("between 14 and 22 months, with 80% confidence" forces you to think through the scenarios a single number hides). And a record: who estimated what, so it can be calibrated over time (Part 3 comes back to this, because calibration can be trained).
4. A process with teeth. Dan Lovallo and Olivier Sibony studied 1,048 real strategic decisions —investments, launches, M&A— recording both the quality of the analysis and the quality of the process (were the risks explicitly discussed? was there genuine dissent at the table? were perspectives contrary to the sponsor's heard? was the case compared against alternatives?). The result is among the most cited in management of the past decade: process explained six times more of decision performance than analysis did. Not because analysis is superfluous —but because great analysis inside a bad process never even gets heard. Moving from the bottom to the top quartile of process quality was associated with +6.9 percentage points of ROI. For anyone who approves other people's decisions, Kahneman, Lovallo and Sibony distilled all this into a 12-question checklist (Before You Make That Big Decision, 2011) —is there self-interest in the recommendation? has the team fallen in love with it? where is the outside view?— which works precisely because it does not ask the decision-maker to de-bias himself: it asks him to audit the process of whoever is proposing.
| Bias | How it sounds in the room | Antidote with evidence |
|---|---|---|
| Overconfidence / planning fallacy | "The base case is conservative already" | Outside view: base rate for the class of project, then adjust |
| Anchoring | The first number frames the entire discussion | Independent written estimates before any discussion; ranges, not point figures |
| Escalation of commitment | "We have invested too much to stop now" | Exit rules written down before going in; the successor test |
| Loss aversion + narrow framing | Every manager turns down his coin flip; the portfolio loses | Decide at portfolio level; explicit risk appetite; reward the good bet whatever the outcome |
| Outcome bias | "It went wrong, someone will pay" | Assess process quality, not just results; documented premortem |
The bias, how it shows up in the committee, and its antidote.
Questions for Monday
This part's self-diagnosis, in board-level form. For your most recent big decision —an investment, a launch, an acquisition:
Seven questions for Monday
- Did the business case include the statistics of comparable projects —yours and other people's— or only the internal story of why this one was different?
- Who put the first number on the table —and what would have changed if the first number had been another one?
- Did anyone formally play the part of "this has already failed: these were the causes" —or did the sceptic end up looking disloyal?
- Were the critical estimates produced independently and in writing before the meeting —or "built" during the discussion, after everyone had heard the boss?
- Are there written exit rules —what signal, by what date, triggers what decision— for the projects under way, or is continuity defended with what has already been spent?
- If a new management team took over tomorrow, with no history and no commitments, would it stand behind your three biggest current bets?
- Has your organization ever rewarded a well-made decision that turned out badly —or punished a badly made one that got away with it?
If several of those answers are uncomfortable, the problem is not your people: it is that your decision process still trusts the mind to correct itself. It doesn't —not yours, not anyone's.
Half the problem remains, however. Biases explain why the apparatus fails; they do not explain why the same method that works brilliantly on one decision fails on the next. Deep analysis is the right answer when the problem is complicated —and a fatal trap when it is complex, where analysis cannot stand in for trial. That is what Part 2 is about: the four territories of decision-making, how to recognize which one you are standing on, and why two of the century's best decision-makers —Warren Buffett and Elon Musk— can say opposite things about strategy and both be right.
Is your organization about to make a big decision?
32sur works on decision quality as part of its management processes: business cases with an outside view and reference classes, documented premortems, independent estimates in investment approvals, exit rules set ex ante, and evaluation on process quality. We don't sell infallibility: we install the process that makes biases —yours and ours— cost less. If your organization is about to make a big decision, let's talk.
References
- Kahneman, D., Thinking, Fast and Slow, Farrar, Straus and Giroux, 2011 — the two systems, WYSIATI, the planning fallacy, the educational-textbook anecdote, the premortem as his favourite tool.
- Englich, B., Mussweiler, T. and Strack, F., "Playing Dice with Criminal Sentences: The Influence of Irrelevant Anchors on Experts' Judicial Decision Making", Personality and Social Psychology Bulletin, 32(2), 2006 — loaded dice and sentences: 5.28 vs. 7.81 months; the effect also holds for professionals with 10+ years of experience.
- Northcraft, G. B. and Neale, M. A., "Experts, Amateurs, and Real Estate: An Anchoring-and-Adjustment Perspective on Property Pricing Decisions", Organizational Behavior and Human Decision Processes, 39(1), 1987 — professional appraisers anchored by the listing price (US$67,811 vs. US$75,190); only 24% acknowledged it.
- Buehler, R., Griffin, D. and Ross, M., "Exploring the 'Planning Fallacy': Why People Underestimate Their Task Completion Times", Journal of Personality and Social Psychology, 67(3), 1994 — 33.9 days estimated vs. 55.5 actual; 29.7% met their date; not even the worst case sufficed (48.7%).
- Flyvbjerg, B. and Gardner, D., How Big Things Get Done, Currency, 2023 (and the Flyvbjerg project database, Oxford, 16,000+ cases) — 47.9% on budget; 8.5% on budget and schedule; 0.5% with benefits; overruns by type and fat tails (IT: +447% in the tail).
- Ben-David, I., Graham, J. R. and Harvey, C. R., "Managerial Miscalibration", Quarterly Journal of Economics, 128(4), 2013 — 13,300 CFO forecasts (2001–2011): the 80% confidence intervals captured reality 36.3% of the time.
- Malmendier, U. and Tate, G., "Who Makes Acquisitions? CEO Overconfidence and the Market's Reaction", Journal of Financial Economics, 89(1), 2008 — overconfident CEOs: 65% higher odds of acquiring; market reaction −90 bps vs. −12 bps. Complemented by Roll, R., "The Hubris Hypothesis of Corporate Takeovers", Journal of Business, 59(2), 1986.
- Staw, B. M., "Knee-Deep in the Big Muddy: A Study of Escalating Commitment to a Chosen Course of Action", Organizational Behavior and Human Performance, 16(1), 1976 — US$13.1M to the failing course chosen by the participant himself vs. ~US$9M when someone else had chosen it (laboratory study using business cases).
- Kahneman, D. and Lovallo, D., "Timid Choices and Bold Forecasts: A Cognitive Perspective on Risk Taking", Management Science, 39(1), 1993; and Lovallo, D. and Kahneman, D., "Delusions of Success", Harvard Business Review, July 2003 — inside and outside view; optimism in the plan, timidity in the portfolio; the Thaler anecdote with the 25 managers appears in Kahneman (2011), ch. 31.
- Baron, J. and Hershey, J. C., "Outcome Bias in Decision Evaluation", Journal of Personality and Social Psychology, 54(4), 1988 — the same decision judged by how it turned out.
- Klein, G., "Performing a Project Premortem", Harvard Business Review, September 2007; and Mitchell, D. J., Russo, J. E. and Pennington, N., "Back to the Future: Temporal Perspective in the Explanation of Events", Journal of Behavioral Decision Making, 2(1), 1989 — prospective hindsight generates ~30% more concrete causes.
- Lovallo, D. and Sibony, O., "The Case for Behavioral Strategy", McKinsey Quarterly, March 2010; and Kahneman, D., Lovallo, D. and Sibony, O., "Before You Make That Big Decision", Harvard Business Review, June 2011 — 1,048 decisions: process explained 6× more than analysis; +6.9 pp of ROI from the bottom to the top process quartile; the 12-question checklist.
- Grove, A., Only the Paranoid Survive, Currency, 1996 — the question to Gordon Moore and Intel's exit from the memory business (the "successor test").
- On the replication crisis: Klein, R. A. et al., "Investigating Variation in Replicability: A 'Many Labs' Replication Project", Social Psychology, 45(3), 2014 (anchoring among the most robust effects); and D. Kahneman's public comment (2017) on the priming chapter: "I placed too much faith in underpowered studies".