The call center hits average handle time for the third month running. Leadership celebrates. Customers leave longer surveys that say the same thing: “They rushed me.” Agents learned the game. Transfer early. Soft-close and reopen. Skip the hard cases until after measurement windows. The metric is green. The experience is red. Nobody “cheated” in a cartoon-villain way. They did what the system paid them to do.
This is Part 5 of Metrics that matter. Parts 1 through 4 covered the goal-to-metric chain, leading and lagging indicators, metric specs, and the North Star versus team scorecard stack. Now the uncomfortable half: how people break metrics, why incentives warp measurement, and how to design so gaming is harder and honesty is cheaper. If your stack is still fuzzy, pair this with Part 4. Foundations still apply: know the decision before you weaponize a number (Analytics foundations). Bad data quality makes gaming harder to spot (Data quality). Pretty charts of a gamed metric are still a lie (data-viz). Path map: Learn.
What you’ll learn
- Goodhart’s law and Campbell’s law in plain English
- Common gaming patterns at work (not just call centers)
- How to pressure-test a metric before you attach a bonus
- A worked example with a “design against gaming” table
- Mistakes that invite distortion, and a practice checklist for your next KPI review
The law in one sentence (and the longer truth)
The version most people know comes from anthropologist Marilyn Strathern, summarizing a line of thought tied to economist Charles Goodhart:
When a measure becomes a target, it ceases to be a good measure.
Goodhart’s original insight, from 1970s UK monetary policy, was closer to: any observed statistical regularity tends to collapse once pressure is placed upon it for control purposes. In other words, the relationship you saw in the wild stops behaving once you use it as a steering wheel.
Donald Campbell made a related point in social measurement: the more a quantitative social indicator is used for decision-making, the more subject it becomes to corruption pressures, and the more it can distort the process it was meant to monitor. Education systems that teach only to the test are the classic example. Business has the same physics with different costumes.
These are not arguments against metrics. They are arguments against naive metrics: single numbers with high stakes, weak definitions, no guardrails, and no curiosity about side effects.
Why smart people game numbers
Gaming is not always malice. It is often rational adaptation inside a broken scoreboard.
- The metric is easier than the mission. Closing tickets is easier than resolving problems. Shipping tickets is easier than shipping outcomes.
- The reward is sharp and the cost is soft. Bonus is cash this quarter. Brand damage is next year’s problem.
- The definition has loopholes. If “active user” means any login, bots and password resets become growth hacks.
- Leaders ask for green. Culture that punishes bad news produces creative accounting, not better operations.
- Teams compete on a leaderboard. Relative ranking without shared guardrails turns colleagues into adversaries with spreadsheets.
If you design metrics as if people were robots who cannot read incentives, you will be surprised every quarter. Design as if people are clever, busy, and responsive to what gets measured and celebrated.
A map of how metrics break
The diagram below is a simple map from pressure to distortion. Use it in design reviews: start at the left, walk right, and ask where your KPI is vulnerable.

Pattern 1: Narrowing
People optimize the measured slice and abandon the unmeasured whole. Teachers drill tested topics. Sales push products that close fast, not products customers need. Analysts report only the segments that look good. The metric improves. Capability shrinks.
Pattern 2: Threshold theater
Behavior clusters just above the pass line. Support answers at 4 minutes 50 seconds when the SLA is 5:00. Engineers merge trivial PRs to hit “ship count.” Nobody aims for excellence. Everyone aims for not-red.
Pattern 3: Definition shopping
If the formula is vague, teams pick the version that flatters them. Is churn logo or revenue? Is a trial “converted” at signup or at paid invoice? Without a Part 3 spec, you get three truths and a political fight.
Pattern 4: Timing games
Pull revenue into this quarter. Delay refunds into next. Pause outbound at month end so “response time” looks fine. Push releases after the measurement freeze. Time is a dimension you must specify, or people will choose the flattering clock.
Pattern 5: Outright fabrication
Rarer in healthy cultures, real under extreme pressure: fake surveys, discarded samples, altered logs. This is a people and controls problem, not only a KPI design problem. Still, extreme single-metric stakes make fabrication more likely.
Pattern 6: Proxy collapse
You chose a proxy because the real outcome was hard to measure. Then you managed the proxy as if it were the outcome. Clicks for engagement. Lines of code for productivity. Training hours for skill. Proxies drift. When the proxy becomes the target, drift accelerates.
Incentives: the quiet co-author of every KPI
A metric on a dashboard is a suggestion. A metric tied to compensation, promotion, vendor scorecards, or public ranking is a behavioral program. Before you attach stakes, ask:
- What is the cheapest way to move this number without creating real value?
- Who can change the inputs without changing the customer outcome?
- Which guardrail would catch the cheap path?
- Can someone lose money for doing the right long-term thing?
- Is the window short enough to encourage timing games?
Sometimes the fix is not a cleverer formula. It is lower stakes on the proxy, higher stakes on a balanced set, or delaying bonuses until quality shows up in a lagging outcome. Part 2’s leading and lagging pair is an anti-gaming tool: pay on a mix, or review leading weekly and settle lagging quarterly.
Design moves that make gaming harder
Example:

When the number climbs for the wrong reason:

You will not eliminate Goodhart. You can raise the cost of empty optimization and lower the cost of telling the truth.
- Pair metrics. Never attach high stakes to a lone volume metric. Pair volume with quality, speed with rework, acquisition with retention.
- Write the loophole list. Before launch, brainstorm “how we hit the target while failing the mission.” Turn the best attacks into guardrails or definition clauses.
- Spec like you mean it. Owner, grain, formula, filters, time window, known caveats (Part 3). Publish the spec where the number appears.
- Separate learning metrics from scoring metrics. Experiment metrics should be allowed to look messy. Scorecards for pay should be stabler and audited.
- Sample the qualitative. Read tickets, listen to calls, review a random set of “closed” cases. Numbers without contact with reality drift.
- Rotate and retire. If a proxy is being drained dry, change the mix. Stale KPIs invite expert gaming.
- Reward exception honesty. Celebrate the team that flags a broken definition early. If only green is safe, green will be manufactured.
- Watch data quality. Gaming often shows up as weird distributions, sudden definition changes, or missing events. Quality checks from the data-quality series are part of metric hygiene.
Worked example: “resolved tickets” goes wrong
Northwind Support (same fictional company as Part 4) sets a target: 90% of tickets resolved within 24 hours. Bonus for team leads if the target holds for the quarter. Within six weeks:
- Agents mark tickets resolved and ask customers to “reply to reopen” if still stuck
- Hard tickets transfer to a backline queue that is measured differently
- Customers with incomplete data get closed as “info needed” without a real path back
- CSAT is not on the scorecard, so speed wins every trade-off
The 24-hour resolve rate hits 93%. Repeat-contact rate climbs. Churn risk rises in the mid-market segment that hits complex issues. Leadership is confused because “support is crushing it.”
A redesign against gaming might look like this:
| Element | Old design | Failure mode | Redesign |
|---|---|---|---|
| Primary metric | % resolved in 24h | Premature close, transfers | % resolved in 24h and not reopened within 7 days |
| Quality pair | None on scorecard | Speed over help | CSAT or ticket-level quality sample; reopen rate as hard guardrail |
| Definition | “Resolved” left to agent judgment | Definition shopping | Spec: resolved means customer-confirmed or policy checklist complete; auto-reopen rules |
| Incentives | Bonus on speed alone | Rational gaming | Bonus on composite: speed band + reopen under cap + quality sample pass |
| Segmentation | Company average only | Hide hard queues | Report by severity and channel; no cherry-picking average |
| Audit | None | Silent fabrication risk | Monthly random audit of 20 closed tickets with shared notes |
| Learning space | All tickets scored | Fear of experiments | Pilot queue excluded from bonus during process tests |
A compact metric contract for the redesign (the kind of text you store next to the dashboard) might read:
METRIC: durable_24h_resolve_rate
GRAIN: support ticket
FORMULA:
tickets_durably_resolved_in_24h / tickets_created
WHERE durable = resolved_at within 24h of created_at
AND no_reopen within 7d of resolved_at
AND resolve_reason NOT IN ('info_needed_auto', 'transferred_out')
FILTERS: exclude spam; exclude Sev-1 incident children (tracked elsewhere)
OWNER: Head of Support (definition), Analytics (pipeline)
GUARDRAILS:
reopen_rate_7d <= 12%
csat_trailing_28d >= 4.2
transfer_rate stable vs baseline
STAKE: team bonus uses composite scorecard, not this rate alone
KNOWN LOOPHOLES:
- soft resolve then chat offline (mitigate with reopen + audit)
- severity downgrades (mitigate with QA sample)
REVIEW: weekly ops; monthly audit; quarterly definition freezeThis is not bureaucracy for its own sake. It is the minimum writing that keeps a high-stakes number from eating the mission. Analysts who can draft this contract are more valuable than analysts who only can plot the old, gameable rate in a prettier chart.
Where analysts and managers each own the problem
Managers own incentives, ritual, and what happens when a number is red. If red always means punishment, expect theater. If red means “what did we learn and what will we try?”, metrics stay informative longer.
Analysts own clarity, instrumentation, and early warning. Sudden shape changes in a distribution, odd spikes at period end, or a definition that cannot be reproduced from the warehouse are smoke. Use the same skepticism you bring to joins and grain in the Python and data-quality work: if the number cannot be rebuilt, it cannot be trusted under pressure.
Both own the loophole brainstorm. The best anti-gaming sessions are mixed rooms: ops people who know the shortcuts, analysts who know the data path, and a leader who can change stakes.
Common mistakes
- One metric, big bonus, no guardrails. The purest form of Goodhart bait.
- Blaming individuals for system design. If half the team “games,” redesign the game.
- Secret definitions. Ambiguity is a gaming subsidy.
- Average-only reporting. Averages hide sacrificed segments and queues.
- Never visiting the work. Metrics without occasional ground truth rot.
- Changing the formula mid-quarter without a note. Looks like fraud even when it is an innocent fix. Version your specs.
- Confusing motion with progress. Activity metrics (emails sent, tickets touched) are easy to inflate. Prefer outcome pairs.
How to practice this week
- Pick one high-visibility metric on your scorecard.
- Write five ways a clever teammate could improve it without improving the real goal.
- For each way, add a guardrail, a definition clause, or a paired metric.
- Check whether compensation or ranking touches that metric. If yes, raise the design bar.
- Schedule one qualitative sample (calls, tickets, deals) before the next review.
- Read Part 6 next: review cadence, so your anti-gaming design actually shows up in the meeting, not only in a document nobody opens.
Quick recap
- When a measure becomes a target, it stops being a neutral mirror. That is Goodhart’s territory; Campbell adds that high-stakes indicators invite corruption and distortion.
- People game through narrowing, thresholds, definition shopping, timing, proxies, and sometimes fabrication.
- Incentives co-author your metrics. Design pairs, specs, audits, and honest culture.
- Brainstorm loopholes on purpose. Turn them into guardrails before you attach pay.
- Analysts and managers share the job: clear measurement plus sane stakes.
Next, and finally in this series: how to run weekly and monthly metric meetings that use this stack without becoming a theater of green dots.
Sources
Research and further reading used for this article:
- Wikipedia overview of Goodhart’s law (history and Strathern phrasing): https://en.wikipedia.org/wiki/Goodhart%27s_law
- Marilyn Strathern, “Improving ratings: audit in the British university system,” European Review (1997): commonly cited source of the “when a measure becomes a target…” phrasing; journal page via Cambridge: https://www.cambridge.org/core/journals/european-review/article/abs/improving-ratings-audit-in-the-british-university-system/FC2EE640C0C44E3DB87C29FB666E9AAB
- PMC discussion of Goodhart’s law in measurement contexts: https://pmc.ncbi.nlm.nih.gov/articles/PMC7901608/
- Wikipedia overview of Campbell’s law: https://en.wikipedia.org/wiki/Campbell%27s_law
- Nielsen Norman Group on Campbell’s law and metric manipulation: https://www.nngroup.com/articles/campbells-law/
- Amplitude North Star Framework (context for product metrics that still need guardrails): https://amplitude.com/books/north-star/about-north-star-framework
- Analytics Made Simple, Learn hub: https://analyticsmadesimple.com/learn/
- Analytics Made Simple, Analytics foundations: https://analyticsmadesimple.com/series/analytics-foundations/
- Analytics Made Simple, Data quality for people who ship numbers: https://analyticsmadesimple.com/series/data-quality/
- Analytics Made Simple, Charts that make sense: https://analyticsmadesimple.com/series/data-viz/
- Analytics Made Simple, Python for analytics: https://analyticsmadesimple.com/series/python/
