Analytics

Goodhart's law

Goodhart's law says a measure stops being a reliable guide once people are pushed to hit it, because they start working on the number instead of the result it was meant to track.

In short

Goodhart's law is the observation that a measure stops tracking what you care about once it becomes a target. People then optimize the number itself, by gaming it or by neglecting what it leaves out. Charles Goodhart's 1975 version concerned monetary policy; the popular wording, 'When a measure becomes a target, it ceases to be a good measure', is Marilyn Strathern's from 1997.

Origin
Charles Goodhart (original, on monetary policy); Marilyn Strathern (popular wording), 1975; 1997
Level
301 · Advanced
Fits
Small and mid-size, Scale-up, Enterprise
Time to apply
About an hour per metric for a first review
What you need
the list of metrics that carry targets, bonuses or board attention · the real outcome each metric is meant to stand for, in one sentence · someone from the front line who knows how the number is produced

Goodhart’s law is the observation that a measure stops tracking what you care about once it becomes a target. Once people are rewarded for the number, they raise the number, and the link between the number and the real result loosens. It was named after the British economist Charles Goodhart, and the same pattern has since been described in schools, banks, police forces and AI training.

A line chart with time after the target is set on the horizontal axis. A black line for the reported metric keeps rising. A blue line for the real goal rises with it at first, peaks, then falls, leaving a widening gap marked at the right.
The metric and the goal move together until pressure is applied, then they part.

Which wording is Goodhart’s, and which is not?

Goodhart’s own sentence is narrower than the one on most slides. Three related statements are often blended, so here they are side by side.

Statement Author and date What it says
Goodhart’s law Charles Goodhart, 1975 (UK monetary policy) “Any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes.”
The popular wording Marilyn Strathern, 1997 “When a measure becomes a target, it ceases to be a good measure.”
Campbell’s law Donald Campbell, paper dated 1976, journal version 1979 The more a quantitative social indicator is used for decisions, the more it attracts “corruption pressures” and distorts the process it monitors

The first is the original. Goodhart’s own later account cites the wording to his paper “Problems of Monetary Management: The U.K. Experience”, which he quotes in a 2021 comment and which Manheim and Garrabrant also cite. It appeared in a Reserve Bank of Australia volume from a July 1975 conference; the RBA’s bibliography dates the volumes to 1976, so you will see both years. He was writing about money-supply targets: a statistical relationship that held while nobody tried to control it broke when policymakers leaned on it. That is close in spirit to the Lucas critique, which argued in 1976 that any change in policy will alter the structure of the models used to evaluate it.

The second is a paraphrase. Strathern’s 1997 paper on university audit contains the sentence on page 308, followed by a remark that Hoskin describes this as Goodhart’s law. So the sentence is hers, and the label comes through Keith Hoskin’s 1996 chapter, which we have not read. A reader who wants the original should cite Goodhart, and one who wants the short version should cite Strathern.

The third is a sibling. Campbell’s paper, reprinted by the Journal of MultiDisciplinary Evaluation, states it for social indicators and backs it with cases. The earliest wording we could open is dated 1976, a year after Goodhart’s talk, although Manheim and Garrabrant note that Campbell’s law “arguably has scholarly precedence”, so sources disagree on priority.

Four ways a measure fails

A measure can fail for different reasons, and the fix depends on which one. Manheim and Garrabrant’s 2018 taxonomy separates four. They define a Goodhart effect as optimization causing the statistical link between the goal and the proxy to collapse.

Four boxes in a two-by-two grid labelled Regressional, Extremal, Causal and Adversarial, each with a short description. Adversarial, people with other goals gaming the metric, is blue.
Manheim and Garrabrant's four ways a measure can fail once it is optimized.

Regressional failure comes from noise. Pick the top 5% by a score and you also pick everyone whose score was inflated by luck, so the group does worse than the score promised. Extremal failure comes from range: a relationship learned on ordinary values may not hold at extreme ones. Causal failure happens when the metric is a symptom, not a cause, so pushing it does nothing for the goal. Adversarial failure needs an actor with different aims who plays against the metric, and it is the kind the cases below show.

What it looks like in practice

The documented cases follow the same pattern: a number is attached to consequences, and behavior shifts to the number.

Case The measure What happened
A US school experiment in Texarkana Pupil test-score gains, which set contractor payment Campbell reported contractors in a Texarkana experiment teaching the answers to specific test items
Police in some US jurisdictions Crime clearance rates Campbell described under-recording of complaints and plea bargains that traded lighter sentences for confessions to unsolved crimes
Sears auto repair, early 1990s A goal of $147 sales per hour Staff overcharged and did unnecessary repairs, and the chairman later admitted the goal-setting had created an environment where mistakes occurred, according to Ordóñez and colleagues
English police, 2014 Targets built on recorded crime A UK parliamentary committee called the evidence that such targets distort recording incontrovertible
Wells Fargo, 2016 Cross-selling targets and bonuses More than 2 million accounts that may not have been authorized, per the CFPB

Bevan and Hood’s study of English health care tests the assumption that measurement problems and gaming do not matter against evidence from the health service, and ends with suggestions for making target-based governance more resilient. Jerry Muller’s book on metrics surveys the same dynamic in education, medicine, business and government.

AI research arrived at the same problem independently. DeepMind’s specification gaming examples include a boat-race agent that circled to hit reward blocks instead of finishing the race. Gao, Schulman and Hilton measured how pushing a proxy reward model too hard worsens true quality, and Skalse and colleagues gave a formal definition of reward hacking.

Why it happens inside companies

The mechanism has a name in accounting research: surrogation. Choi, Hecht and Tayler define it as managers treating a performance measure as if it were the strategy it represents. In two experiments, surrogation was more likely when pay depended on a single measure and less likely with several. Harris and Tayler make the practical case in Harvard Business Review: strategy is abstract, metrics are concrete, and the concrete thing wins attention.

Not everyone puts the blame on metrics. Michael Schrage argues that leadership, not measurement, is the problem, and that better, adaptive KPIs are the answer. The two views agree on one point: the damage comes from how a number is used.

How to defend a metric

Kaplan and Norton’s 1992 line, “What you measure is what you get”, is the starting point, and their balanced scorecard answers it with several measures instead of one. In practice that means pairing each target with a check metric, as in the steps above, and building the pairs into a KPI tree where every top-line goal shows the inputs that drive it. A North Star metric concentrates attention, so it needs the closest guarding. A leading indicator stands in for a result that arrives later, so pair it with the lagging number it is meant to predict. The balanced scorecard page covers the multi-measure approach in more detail.

Separate the metric from the consequence where you can. Google’s OKR guide says OKRs are not synonymous with performance evaluation and expects an average score of 0.6 to 0.7, which leaves room for ambitious goals without punishing a miss. If you are building a metric system from scratch, a Growth Lab plan starts from the outcome you want and works back to the numbers.

How to apply Goodhart's law, step by step

  1. Write the goal behind the metric. For each targeted metric, state in one sentence what you actually want: fewer lost customers, safer patients, healthier cash. The metric is a stand-in for that sentence. Result: a pair of lines, metric and real goal, for every target.
  2. Ask how to hit the number without the goal. Gather two or three people who handle the process and ask how they would raise the figure if they only cared about the figure. Cheap shortcuts, skipped steps and easy cases are the answers. Result: a short list of ways the metric can be gamed, written before anyone gets paid on it.
  3. Pair the target with a check metric. Choose a second measure that falls when the shortcut is used: reopened tickets next to tickets closed, refunds next to sales, readmissions next to discharges. Report both together. Result: no target travels alone.
  4. Separate the number from the consequence. Decide which metrics steer decisions and which trigger pay, ranking or discipline. The harder the consequences, the stronger the pressure to game. Result: a smaller set of metrics that carry personal stakes, each with a check metric.
  5. Audit how the number is produced. Sample a few records behind each reported figure and compare them with the source: the call, the chart, the contract. Result: a measured sense of how far the reported number sits from reality.
  6. Set a review date and retire what has decayed. Put a date on each target. At the review, ask whether the gap between metric and goal has grown, and replace or drop the metric if it has. Result: targets with an expiry date instead of permanent fixtures.

Examples

A fintech support team

Illustrative. A payments company sets a goal of 60 tickets closed per agent per day. Within a month the figure is met, but customers reopen 1 ticket in 4 because agents close cases after a canned reply. The real goal was problems solved, and the closing count never measured that. The fix is a pair: tickets closed, reported alongside the share reopened within seven days, with bonuses tied to neither alone.

A clinic's waiting-time target

Illustrative. A clinic promises that patients wait under 15 minutes, counted from check-in. The front desk starts checking patients in only when a doctor is nearly free, so the recorded wait looks fine while patients sit in the corridor. A check metric counted from the appointment time, plus a short patient survey, exposes the gap.

Wells Fargo's cross-selling targets

The US Consumer Financial Protection Bureau found that Wells Fargo employees opened unauthorized deposit and credit card accounts to hit sales targets and earn bonuses. The bank's own analysis counted millions of accounts that may not have been authorized, and the bureau fined it.

When to use it

Use it as a design check whenever a metric is about to carry a target, a bonus, a ranking or a board commitment. It fits especially well when a team reports one headline number, when a metric is tied to pay, or when a number has been improving for a long time without the business outcome following.

When not to use it

Do not treat it as a reason to avoid measuring. The law describes what happens when a measure is targeted without safeguards, and metrics that only inform decisions are much less exposed. It also does not tell you which metric to choose: for that, build a KPI tree or pick a North Star metric first, then apply the law to the result.

Common mistakes

  • Quoting Strathern's sentence as Goodhart's own. The 1975 original is about statistical regularities in monetary policy, and the popular line is a later paraphrase.
  • Blaming the people who game the number. A target that pays for the number invites it, as the cases on this page show.
  • Fixing a gamed metric by adding one more target. Each target without a check metric opens a new shortcut.
  • Tying pay to the only metric you can measure well. Choi, Hecht and Tayler found managers confuse the measure with the strategy most when pay depends on a single measure.
  • Never auditing the source records. A reported number can look healthy for years while the process behind it drifts.

FAQ

What is Goodhart's law in simple words?

When you reward people for hitting a number, they find ways to hit the number that do not help the real goal, and the number stops telling you the truth. Test scores that rise because teachers drill the test, or tickets closed fast because agents skip the fix, are everyday examples.

What did Goodhart actually say?

Goodhart's 1975 paper, as he later cited it himself, says that any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes. He was describing money-supply targets used for UK monetary policy, not office KPIs. The wording about measures and targets is Strathern's, from 1997.

What is the difference between Goodhart's law and Campbell's law?

Donald Campbell, a social scientist, wrote that the more a quantitative social indicator is used for decision-making, the more it will attract corruption pressures and distort the process it monitors. Goodhart's concerned statistical regularities in policy; Campbell's concerned social indicators. Later writers treat them as close relatives, and sources differ on which came first.

Who said 'when a measure becomes a target, it ceases to be a good measure'?

Anthropologist Marilyn Strathern wrote it in a 1997 paper on audit in British universities (European Review, p. 308). She credited the label 'Goodhart's law' to Keith Hoskin. Many websites attribute the sentence to Goodhart himself, which is wrong, so check the date and the author before quoting it.

Does Goodhart's law apply to OKRs?

Yes, any measure attached to consequences is exposed. Google's public OKR guide says OKRs are not synonymous with performance evaluation and expects an average grade of 0.6 to 0.7. Keeping key results separate from pay and ranking reduces the pressure to inflate them.

Sources

  1. David Manheim, Scott Garrabrant, Categorizing Variants of Goodhart's Law, arXiv:1803.04585, 2018 (v4, 2019)
  2. Donald T. Campbell, Assessing the Impact of Planned Social Change, Occasional Paper 8, Dartmouth Public Affairs Center, 1976 (reprint, Journal of MultiDisciplinary Evaluation, 2011)
  3. Marilyn Strathern, 'Improving ratings': audit in the British University system, European Review 5(3), 1997
  4. Charles Goodhart, Marglin and Money: Comments on Chapter 13, Just Money symposium, 2021
  5. Reserve Bank of Australia, RDP 9013, Select Bibliography of Published Research 1969-1990: Conference Volumes
  6. Robert E. Lucas Jr., Econometric policy evaluation: a critique, Carnegie-Rochester Conference Series on Public Policy 1, 1976
  7. Consumer Financial Protection Bureau, CFPB fines Wells Fargo $100 million for widespread illegal practice of secretly opening unauthorized accounts, 2016
  8. Lisa Ordóñez, Maurice Schweitzer, Adam Galinsky, Max Bazerman, Goals Gone Wild, Harvard Business School Working Paper 09-083, 2009
  9. Gwyn Bevan, Christopher Hood, What's measured is what matters: targets and gaming in the English public health care system, Public Administration 84(3), 2006
  10. UK Government, Response to the Public Administration Select Committee report Caught red-handed (HC 760), Cm 8910, 2014
  11. Jerry Z. Muller, The Tyranny of Metrics, Princeton University Press, 2018
  12. Jongwoon Choi, Gary Hecht, William B. Tayler, Lost in Translation: The Effects of Incentive Compensation on Strategy Surrogation, The Accounting Review 87(4), 2012
  13. Michael Harris, Bill Tayler, Don't Let Metrics Undermine Your Business, Harvard Business Review, September-October 2019
  14. Michael Schrage, Don't Let Metrics Critics Undermine Your Business, MIT Sloan Management Review, 2019
  15. Robert S. Kaplan, David P. Norton, The Balanced Scorecard: Measures That Drive Performance, Harvard Business Review, January-February 1992
  16. Google re:Work, Set goals with OKRs
  17. Victoria Krakovna and colleagues, Specification gaming: the flip side of AI ingenuity, Google DeepMind, 2020
  18. Leo Gao, John Schulman, Jacob Hilton, Scaling Laws for Reward Model Overoptimization, arXiv:2210.10760, 2022
  19. Joar Skalse, Nikolaus Howe, Dmitrii Krasheninnikov, David Krueger, Defining and Characterizing Reward Hacking, arXiv:2209.13085, 2022
  20. Simon Zhuang, Dylan Hadfield-Menell, Consequences of Misaligned AI, NeurIPS 2020, arXiv:2102.03896

Last updated Oct 9, 2026

Ilia PushinFounder, PUSHERS & COO Fintech ServiceIlia builds operating systems for growing companies in fintech and healthcare. Since 2021 he has run cross-border payments at ARBI Exchange, a licensed currency exchange in Thailand, including KYC and AML and the move into new jurisdictions.About the authorLinkedIn
Related frameworks
More frameworks
Want Goodhart's law running inside your company?Request an operations audit