Every light green
Monthly management meeting. On the screen, the dashboard: thirty-odd indicators, almost all green, several with a little arrow pointing up. Sales above target, productivity on target, support tickets closed in record time. Applause, coffee, "good month." Nobody leaves worried.
Three months later a problem erupts that nobody saw coming: the big customers are leaving, the margin has collapsed, cash is at the limit. What is baffling is that none of those three disasters contradicted the dashboard. Sales were rising — with discounts that eroded the margin. Tickets were closing fast — without resolving, and so they came back. Productivity was on target — producing stock nobody had asked for. Every light was green, and the business was sinking in silence, because the numbers that were green were not the numbers that mattered.
This article is about that distance —between measuring a lot and measuring well— that sinks companies with the dashboard smiling. About why management's most repeated phrase is dangerously incomplete, about a law that explains how metrics corrupt, and about what to measure so the dashboard, at last, changes decisions instead of decorating them.
"What gets measured gets managed": the phrase almost no one quoted in full
It is, probably, the most repeated phrase in management: "what gets measured gets managed." It is almost always attributed to Peter Drucker, who most likely never said it that way, and sometimes to W. Edwards Deming, who in fact argued the opposite: that some of the most important things in an organization are unmeasurable, and that managing only by what can be counted is a costly mistake.
The phrase is true — but a half-truth, and half-truths in management are the most expensive. Yes: what gets measured receives attention. The problem is that this includes measuring the wrong thing, which then gets all the attention while what truly matters, and goes unmeasured, deteriorates with no witnesses. The journalist Simon Caulkin completed it with irony, glossing an old academic finding: "what gets measured gets managed — even when it is pointless to measure and manage it, and even if it harms the organization to do so." That long version is not a productivity slogan: it is a warning.
The dysfunction of measuring is seventy years old
The warning is not new. In 1956 —nearly seventy years ago— a researcher named V. F. Ridgway published in Administrative Science Quarterly a brief and devastating article: "Dysfunctional Consequences of Performance Measurements." His thesis: any quantitative measurement of performance, used with excessive confidence, distorts the very behavior it aims to improve.
Ridgway catalogued the forms. A single measure pushes people to optimize that number at the expense of everything else: the operator who maximizes units and neglects quality. Several measures shift the problem to the trade-offs, which get played in favor of what is measured and against what is not. And composite measures —the index that averages everything into a single number— hide the conflicts instead of resolving them. Ridgway's conclusion, in 1956, was still true in every dashboard presented this week: the problem is not the lack of metrics, it is the naive faith in them.
Goodhart's law (and its twin, Campbell's)
The precise mechanics of that distortion were named by a British economist. In 1975, Charles Goodhart, studying why monetary controls failed, formulated what now bears his name: "any observed statistical regularity tends to collapse once pressure is placed upon it for control purposes." The anthropologist Marilyn Strathern rewrote it in 1997 in the form that became a maxim: "when a measure becomes a target, it ceases to be a good measure."
What is remarkable is that the same law was discovered twice, in two sciences that did not talk to each other. While Goodhart looked at monetary policy, the social psychologist Donald Campbell was studying evaluations of educational and social programs, and arrived in 1976 at a twin formulation —the Campbell's law—: "the more any quantitative indicator is used for decision-making, the more subject it will be to corruption pressures and the more apt it will be to distort the processes it is meant to monitor." Two fields, two authors, the same regularity: the moment a number becomes the yardstick by which one is rewarded or punished, it starts to lie. That two independent disciplines stumbled onto the same thing should be enough to take it seriously.
The why is simple and universal. Almost no metric is the thing that truly matters; it is a proxy, a measurable approximation of something that is not. Sales are a proxy for a healthy business; handle time, a proxy for good service; closed tickets, a proxy for effective support. As long as no one rewards them, proxies work well as a thermometer. But the moment they are turned into a target —with a bonus, a ranking, a dashboard that must be kept green— people stop pursuing the end and start pursuing the number. And since the number is not the end, the number can be raised while destroying the end.
The list of examples is endless and always the same story:
| Metric | What it aims for | What it induces once it becomes a target |
|---|---|---|
| Salesperson revenue | Grow sales | Discounts that erode margin |
| Handle time | Support efficiency | Rushing customers off without resolving |
| Tickets closed per day | Fast support | Closing without resolving; the case reopens |
| Units per shift | Productivity | Stock nobody asked for; quality drops |
None of those behaviors belong to dishonest people: it is rational people responding to the number the organization decided to reward. The metric did not measure wrong; it measured exactly what it was asked to — and that is why it went wrong.
You reward A while hoping for B
Before the economists and the accountants, a management researcher had already put his finger on the sore spot, with a perfect title. In 1975, Steven Kerr published "On the Folly of Rewarding A, While Hoping for B," and his thesis is as simple as it is uncomfortable: organizations reward, all the time, something different from what they say they want — and then are surprised to get what they rewarded, not what they hoped for. Teamwork is proclaimed and individual performance is rewarded; quality is asked for and quantity is paid for; a long-term outlook is demanded and the quarter's result is bonused. People, Kerr says, make an elementary and rational calculation: they figure out what is rewarded and do that, "to the near-total exclusion of what is not rewarded."
What is valuable about Kerr is that he explains why we fall into the same trap again and again: out of fascination with the "objective" and measurable, out of overvaluing the visible, and out of a dose of institutional hypocrisy —proclaiming B looks good; rewarding A is easier to administer—. The bridge to everything above is direct: the metric that is rewarded is the "A." If it is not exactly what one wants to achieve, the company has just hired, through its own incentive system, its entire staff to pursue the wrong thing.
Surrogation: when the metric eats the end
Goodhart and Campbell describe what happens; behavioral accounting explains why it happens inside the head, and gave it a precise name: surrogation. The researchers Willie Choi, Gary Hecht and William Tayler defined it and proved it in experiments: it is the tendency to lose sight of the end a metric represents and to treat the metric as if it were the end. The salesperson does not think "I want healthy, profitable customers" and use sales to measure it; after a while staring at the number, they think, plainly and simply, "I want more sales." The proxy ate the end. And once that happens in the mind, destroying the end to raise the proxy does not feel like cheating: it feels like doing the job well.
What the research adds is doubly useful for anyone in charge. First, surrogation gets worse with incentives: the more money is tied to a number, the more the end is replaced by the metric (Choi, Hecht and Tayler, 2012). It is the experimental confirmation of this whole article's suspicion — a bonus on a single number is not neutral; it manufactures blindness. Second, and more hopeful: surrogation drops when people take part in defining what is measured and why (2013). Whoever helped choose the strategy does not so easily confuse the thermometer with the fever. The defense against surrogation is not a prettier dashboard: it is that the team understands —and has debated— which end each number pursues.
A real case: the 3.5 million accounts
Everything above sounds abstract until it costs billions. Wells Fargo, one of the largest banks in the United States, had a star metric —cross-selling: how many products per customer— and placed it at the heart of its bonuses and its sales pressure. The internal slogan was "eight is great" (eight products per household). The end that metric was meant to represent —customers with a deeper, more satisfied relationship with the bank— disappeared behind the number.
The rest is the textbook story, at industrial scale. Between 2009 and 2016, employees under pressure opened up to 3.5 million accounts and cards that customers never asked for —with forged signatures, diverted emails, invented PINs— just to hit quota. When it broke, in 2016, the regulator imposed a fine of 185 million dollars, thousands of firings followed and so did the top brass, and the reputational damage was counted in billions more. The metric hit its target almost every quarter, until the end. The business it was supposed to measure —trust— was destroyed along the way. It took no evil genius: a rewarded metric, pressure and surrogation at scale were enough. Exactly what Ridgway, Goodhart and Campbell had been warning about for half a century.
The signature of gaming: the spike at the threshold
When a rewarded metric becomes a threshold that must be crossed, it leaves an unmistakable statistical trace — and sometimes you can see it, literally, on a chart. The most studied case is the English public health system, which Gwyn Bevan and Christopher Hood analyzed in 2006 under a memorable name: targets and terror. The government set hard goals —that no patient wait more than four hours in the emergency room, for example— and tied them to serious consequences for hospitals and executives. The number improved almost magically. But looking at the distribution of waiting times, the signature of gaming appeared: a huge spike of patients "seen" just before the four hours, plus a battery of maneuvers —ambulances queuing outside so as not to start the clock, gurneys rebranded as "beds"— so the clock would read what the target demanded, without the patient necessarily being any better.
That spike hugging the threshold is the X-ray of Campbell's law in action, and it is not exclusive to public health: you see it in the sales shifted to land on the good side of the close, in the grades inflated to cross the line, in any number with a reward and a boundary. When you see a distribution that piles up suspiciously right on the convenient side, you are not looking at good performance: you are looking at people crossing a line.
The eight ways to cheat
That a metric corrupts is not a single thing: it is a family of behaviors, and it is worth knowing them to recognize them in time. In 1995, the economist Peter Smith catalogued eight unintended consequences of measuring performance —studying the English public sector, but they hold for any dashboard—. They are the concrete ways a number stops telling the truth:
| Trap | What it is | How it looks in the company |
|---|---|---|
| Tunnel vision | Looking only at what is measured | Everything not on the dashboard is neglected |
| Sub-optimization | Optimizing the part at the expense of the whole | One area meets its number and breaks another's |
| Myopia | The short term buries the long | Maintenance, brand and training are sacrificed |
| Measure fixation | Pursuing the indicator, not the end | "We closed the ticket" even if the problem persists |
| Misrepresentation | Reporting creatively | The number is dressed up before the meeting |
| Manipulation (gaming) | Altering behavior to inflate the number | Sales are shifted from one month to the next to make it |
| Ossification | Paralysis: not innovating so as not to risk the metric | No one tries anything that might lower the number |
| Complacency | Green lulls you to sleep | Improvement stops because "we're on target" |
Lined up, something uncomfortable stands out: almost all of them are rational responses to a poorly built measurement system. People are not cheats; the dashboard invites them. Jerry Muller gathered this entire library of disasters in a book whose title says it all —The Tyranny of Metrics— and its moral is that of this series: the problem is not measuring, it is measuring without criteria and rewarding without thinking.
The dark side of goals
Someone might object that none of this is the fault of goals themselves — that setting ambitious goals is, in fact, one of the best-proven practices in management. And they are half right: the goal-setting theory of Edwin Locke and Gary Latham, backed by hundreds of studies, shows that specific and challenging goals improve performance over a vague "do your best." Goals work. The problem is what comes with them when they are used carelessly.
In 2009, a group of researchers —Ordóñez, Schweitzer, Galinsky and Bazerman— gathered the flip side in an article titled "Goals Gone Wild." Their conclusion: goals are like a potent medication — they cure and, badly prescribed, they poison. The systematic side effects they document are, one by one, the ills of this article: narrow focus that neglects everything not in the goal, greater willingness toward unethical behavior, distorted risk appetite, inhibited learning, corroded culture and intrinsic motivation in decline. And when the goal is a budget tied to the bonus, Michael Jensen titled it without euphemism —"Paying People to Lie"—: a target linked to compensation induces lying twice, when setting the budget (to negotiate a low, achievable goal) and when meeting it (to inflate the result that gets reported).
The moral is not to set no goals: it is to prescribe them as what they are —a potent drug—, with dosage, contraindications and monitoring. A goal with no counterweight, tied to a bonus and with no one watching the side effects, is the most common recipe for measured disaster.
More is not better: few, on time, counterbalanced
The instinctive reaction to a dashboard that misleads is to add metrics: if the green lied, let's measure ten more things. It is the opposite and complementary error. A dashboard of sixty indicators does not inform: it buries. A leadership team's attention is finite, and spread across sixty numbers it is not enough to look at any one of them seriously. More metrics are not more control; past a certain point they are less, because the noise drowns the signal and paralysis replaces decision.
Between the blindness of measuring little and the noise of measuring everything there is a point where the dashboard governs. You get there with five rules.
1. Few and on time. Five or six numbers per area that arrive fast, rather than an encyclopedic dashboard that arrives late. As holds for all of management: the perfect indicator that arrives on the 25th loses to the reasonable one that arrives on Monday. A number that arrives late is not data: it is history.
2. Counterbalanced, so they cannot be gamed. The best defense against Goodhart is to leave no number on its own. To each metric, its counterweight: sales with margin, speed with quality, output with inventory turnover, growth with cash. When each indicator has a twin that suffers if it is forced, the gaming stops paying.
3. Leading, not only lagging. The month's revenue is a lagging indicator: it reports what already happened, when nothing can be done anymore. The pipeline, the quotes, satisfaction, staff turnover are leading: they warn beforehand. A dashboard made only of results is a rearview mirror — useful for knowing where you came from, not for choosing where to go.
4. Every metric is a proxy, with an owner and an expiry date. Remembering that no number is the thing itself protects you from believing it too much. Each indicator needs someone accountable and a review date: when a metric starts to rise while the business does not improve, it is a sign that it became a target and corrupted — and it must be changed, not celebrated.
5. Measure to decide, not to report. The question that orders any dashboard is brutal and liberating: what decision does this number change? If the answer is "none," the indicator is not information: it is decoration, and it takes up an attention that another one lacks. Any number that does not hang from a decision can be deleted at no cost — and with relief.
On the point about incentives, a warning that closes the circle with Goodhart: tying a bonus blindly to a single indicator is building the trap with your own hands. Incentives have to align with the metric —Eccles was already asking for it in 1991— but with the metric counterbalanced, not with the loose number anyone learns to inflate.
The step almost no one takes: validate
One final discipline remains, the most ignored and perhaps the most profitable: checking that the metric really predicts what you believe it does. We choose non-financial indicators —customer satisfaction, internal climate, quality— convinced they anticipate the results. But do they? Charles Ittner and David Larcker went to look, and their finding, published in Harvard Business Review in 2003, is sobering: only around 23% of companies had taken the trouble to build and verify a causal model linking their metrics to results. Those that did not only had more honest dashboards: they earned, on average, a return on assets almost 3 points higher and a return on equity 5 points higher than those that measured on faith.
To validate is to ask the data what we almost never ask it: did the metric we raised last quarter later translate into the result we expected? When the answer is no, there are two possible culprits —either the metric was never a good proxy, or it corrupted upon becoming a target— and both call for the same thing: changing it. There is, besides, a natural wear: metrics lose their edge over time, as everyone learns to meet them —what Meyer and Gupta called the performance paradox—. Without validation, a dashboard is a hypothesis no one tested, held up by habit. With it, it is an instrument. The difference, according to Ittner and Larcker, is paid in points of profitability.
The difficult conversation
"We need more data." Almost never. What is almost always needed is less, better chosen and actually used. An excess of data is a comfortable hiding place: while you argue over what other dashboard to build, you avoid the uncomfortable question of what decision is not being made with the data you already have.
"If it's green, we're fine." Green means the target was met, not that the business is healthy — and less so if the target is a proxy that can be inflated. The green light of a corrupted metric is more dangerous than the red of an honest one, because it anesthetizes. The question is not whether it is green, but whether that green would survive looking at the number next to it.
"What isn't measured can't be managed." Some of the most important things in a company —a customer's trust, a middle manager's judgment, a team's morale— resist measurement, and are managed all the same. Giving up on what matters because it does not fit in a spreadsheet is the most elegant way to end up optimizing the trivial with decimal precision.
Thirteen questions for Monday
The self-diagnosis, boardroom edition:
- Of the dashboard's indicators, how many changed a decision in the last three months?
- Does each key number have a counterweight that suffers if it is forced — or does it live alone?
- How many of our indicators are leading, and how many only report what already happened?
- Is there any number that has been rising for months while the business does not improve? (That is where it became a target.)
- Are bonuses tied to a single indicator that someone could inflate without the company winning?
- Does the dashboard arrive in time to decide, or does it arrive to report the irreversible?
- If we deleted half the indicators, would we lose control or gain focus?
- What important thing do we not measure — and so no one manages?
- Our last "good month," would it survive looking at margin, cash and customer turnover next to sales?
- Has any of our numbers already become an end in itself — do we pursue it even though the goal it represented has been lost from view? (That is surrogation.)
- Did the people who pursue each metric take part in defining which end it pursues — or did they just receive the number and the bonus?
- Do our bonuses reward exactly the behavior we want — or do they reward A while we hope for B?
- Have we ever verified that our non-financial metrics (satisfaction, climate, quality) really anticipate results — or do we take it for granted?
If several answers are uncomfortable, the good news is that no new dashboard costs what a bad one costs. Measuring better is not measuring more: it is measuring few things, on time, counterbalanced and tied to decisions — and having the courage to switch off the lights that only serve to let you sleep soundly.
A dashboard is not a boardroom ornament nor an exam to be passed by leaving everything green. It is an instrument for deciding better. When it stops serving that —when it measures the easy instead of what matters, when it rewards the number instead of the end— it is not neutral: it lies with the reassuring authority of a datum. And a company can sink with every light green.
Is your dashboard always green and you still don't sleep soundly?
32sur designs management dashboards and indicators within its advisory and its professionalization processes: few numbers, on time, counterbalanced and tied to decisions — not one more dashboard to glance at out of the corner of your eye.
References
- Ridgway, V. F., "Dysfunctional Consequences of Performance Measurements", Administrative Science Quarterly, vol. 1, 1956, pp. 240–247.
- Kerr, S., "On the Folly of Rewarding A, While Hoping for B", Academy of Management Journal, 18(4), 1975, pp. 769–783 — organizations systematically reward a behavior different from the one they expect.
- Goodhart, C. A. E., "Problems of Monetary Management: The U.K. Experience", 1975 — Goodhart's law; formulation by Marilyn Strathern (1997): "when a measure becomes a target, it ceases to be a good measure".
- Campbell, D. T., "Assessing the Impact of Planned Social Change", Occasional Paper Series, 1976 — Campbell's law, twin of Goodhart's in the terrain of social and educational policy.
- Smith, P., "On the Unintended Consequences of Publishing Performance Data in the Public Sector", International Journal of Public Administration, 18(2–3), 1995, pp. 277–310 — the eight dysfunctional consequences of measuring (tunnel vision, sub-optimization, myopia, fixation, misrepresentation, manipulation, ossification, complacency).
- Bevan, G. and Hood, C., "What's Measured Is What Matters: Targets and Gaming in the English Public Health Care System", Public Administration, 84(3), 2006, pp. 517–538 — "targets and terror"; the signature of gaming at the threshold.
- Choi, J., Hecht, G. and Tayler, W. B., "Lost in Translation: The Effects of Incentive Compensation on Strategy Surrogation", The Accounting Review, 87(4), 2012, pp. 1135–1163; and "Strategy Selection, Surrogation, and Strategic Performance Measurement Systems", Journal of Accounting Research, 51(1), 2013, pp. 105–133 — surrogation: the metric replaces the end; it gets worse with incentives, better with participation.
- Ordóñez, L. D., Schweitzer, M. E., Galinsky, A. D. and Bazerman, M. H., "Goals Gone Wild: The Systematic Side Effects of Overprescribing Goal Setting", Academy of Management Perspectives, 23(1), 2009, pp. 6–16 — the systematic side effects of setting goals.
- Locke, E. A. and Latham, G. P., A Theory of Goal Setting and Task Performance, Prentice Hall, 1990; and "Building a Practically Useful Theory of Goal Setting and Task Motivation", American Psychologist, 57(9), 2002 — specific and challenging goals improve performance.
- Jensen, M. C., "Paying People to Lie: The Truth about the Budgeting Process", European Financial Management, 9(3), 2003, pp. 379–406 — budgets tied to bonuses induce lying when setting them and when meeting them.
- Meyer, M. W. and Gupta, V., "The Performance Paradox", Research in Organizational Behavior, vol. 16, 1994 — metrics tend to lose discriminating power over time.
- Eccles, R. G., "The Performance Measurement Manifesto", Harvard Business Review, January–February 1991, pp. 131–137 — from financials as the sole foundation to one among many; aligning incentives.
- Kaplan, R. S. and Norton, D. P., "The Balanced Scorecard—Measures That Drive Performance", Harvard Business Review, January–February 1992 — indicators counterbalanced, leading and lagging.
- Ittner, C. D. and Larcker, D. F., "Coming Up Short on Nonfinancial Performance Measurement", Harvard Business Review, November 2003, pp. 88–95 — only ~23% of companies validate that their non-financial metrics predict results; those that do earn more.
- Ries, E., The Lean Startup, Crown Business, 2011 — vanity metrics versus actionable metrics.
- Muller, J. Z., The Tyranny of Metrics, Princeton University Press, 2018 — a synthesis of the dysfunctions of measurement in education, health, business and government.
- On the Wells Fargo case: Consumer Financial Protection Bureau (CFPB), September 2016 action — up to 3.5 million unauthorized accounts (2009–2016) driven by cross-selling targets; joint fine of US$185 million (CFPB, OCC and the City of Los Angeles).
- On "what gets measured gets managed": the phrase is usually attributed to Peter Drucker with no documentary evidence; W. E. Deming argued, in fact, that the most important figures for managing a company are often unknown or impossible to know.