to tell the truth
Data · Metrics · OKR — No.47-goodharts-law

The Map Is Not the Territory

Sharper tools don't fix it: what school rankings and Google's OKRs prove about metrics
Questions this piece answers
  • Why did the No Child Left Behind Act lead to test-score cheating scandals?
  • Does Google's OKR framework fall into the same trap as KPIs?
  • Why are brand metrics harder to measure than other business metrics?
  • What questions keep a metric from being mistaken for the goal itself?
Written forLeaders who run teams on OKRs or KPIs, and marketers designing brand metrics

We assumed a sharper tool would get us out of this trap. Sophistication doesn't remove the trap. It just makes it harder to see.

Every case we've looked at so far shared one trait: the metric was simple, and the shortcut to hitting it was obvious. Which raises the natural next question. What if we designed a smarter metric — one that spelled out, more precisely, the relationship between the goal and the way we measure it? Wouldn't that close the gap?

The place that answered this question first, and at the largest scale, wasn't business. It was education. And the answer was no — sharpening the tool alone isn't enough.

What Test Scores Taught American Schools, the Hard Way

For a long time, there was no objective way to compare how schools were doing. Budget decisions, teacher evaluations, school-to-school comparisons — without a shared yardstick, all of it risked becoming arbitrary. Standardized test scores stepped into that gap. Suddenly a number could show which schools taught better. It was a well-intentioned fix. On paper, hard to argue with.

In 2001, the No Child Left Behind Act (NCLB) turned that logic into federal policy. Every school was assigned an annual target — Adequate Yearly Progress — and penalized for missing it. The underlying premise, that every student deserves to reach a baseline level of achievement, was hard to object to.

But once test scores determined whether a school survived, classrooms changed shape. Teachers taught to the test and quietly cut whatever wasn't on it. Art, physical education, social-emotional development — anything that didn't show up in the score slid out of the curriculum. Scores went up. Whether the thing the score was supposed to stand for — a genuinely good education — was actually happening became a question the score itself could no longer answer.

The most extreme outcome surfaced in 2009, in Atlanta, Georgia. A wide-reaching cheating scandal came to light — dozens of teachers and principals were indicted for altering test answers. When a school's survival hinges on a single score, faking that score becomes the most rational choice available inside the building.

When a school's survival depends on a single score, gaming that score becomes the most rational choice inside the organization.

The same structure repeats in conference rooms. The moment a marketer facing a quarterly deadline asks "does this hit this month's number" before asking "does this actually work" — that's the same mechanism as a teacher, under NCLB pressure, teaching only to the test. The more a metric determines survival, the more the easiest way to hit it becomes the organization's rational choice.

Does a Smarter Framework Change Anything

This is where many companies believed they'd found the answer: OKRs, Objectives and Key Results. Born at Intel and popularized by Google, the framework is more explicit than a plain KPI about the relationship between what a goal means and how it's measured. The Objective is a qualitative direction; the Key Result is the quantitative proof that you're moving toward it. The design intent was to keep the number tethered to the reason for the number. Andy Grove, who created the framework, put the point in a single sentence.

OKRs aren't for measurement. They're for focus.
OKR is not for measurement. It's for focus.
— Andy Grove

The design philosophy is sound. Inside Google, in the early days, that intent was reportedly well understood. But over time, at plenty of organizations, OKRs hardened into a quarterly ritual of filling in numbers and grading them. And once you watch how the framework actually operates on the ground, a familiar pattern resurfaces: the Key Result becomes the goal again. Say the Objective is "get customers to use the product more," and the Key Result is "grow monthly active users by 20%." The team starts running straight at MAU itself. If the fastest way to move MAU is to send more push notifications, the team sends more notifications. MAU climbs. Whether the Objective was "customers using the product more" or "the MAU number going up" stops being a distinction anyone tracks. Notification fatigue, muted alerts, quiet resentment toward the brand — none of that shows up immediately on the MAU chart.

The paradox sharpens in brand work. Set the Objective as "raise brand awareness" and the Key Result as "lift the awareness-survey score by five points," and the team goes looking for ways to move that survey score. Narrow the survey sample to people who recently saw an ad, and the score climbs easily. The gap opens between the Objective — more people meaningfully remembering the brand — and the Key Result — a survey score from a sample that's been filtered a particular way.

This deserves a name. Call it the sophistication illusion: the belief that a smarter tool earns smarter trust automatically. But how refined a tool is, and what that tool is actually measuring, are two separate questions. The fact that "we manage by OKRs" doesn't prove that "we're measuring the right thing." A more sophisticated tool doesn't remove the trap. It just gets caught in it more elegantly.

The dilemma built into OKRs is structural. Leave the Objective purely qualitative, and you have no way to track progress. Make the Key Result precisely quantitative, and that number becomes the goal all over again. No framework resolves this dilemma on its own. What it actually takes is a culture where the people who set the metric keep reminding themselves: this number is a window onto the situation, not the destination itself. Changing the tool is easier than changing the posture people take toward the tool. The second is harder, and it's the one that matters.

Why Brand Is the Harder Case

Ordinary business metrics — revenue, cost, conversion rate — are vulnerable to this trap too. But for brand metrics, the trap runs deeper. That's because the value a brand creates is, by nature, hard to measure in the first place.

At its core, a brand lives in associations, trust, emotional connection inside a consumer's head. You can ask "how much do you like this brand" in a survey. But how tightly that score connects to actual purchase behavior, willingness to pay a premium, or the loyalty that holds when a competitor shows up — none of that is simple. Awareness studies, preference surveys, brand-equity indices show one facet of brand strength. They aren't the brand itself.

Then there's the problem of time. Unlike performance ads, where today's campaign shows up in tomorrow's sales, brand effects arrive late and scattered. Isolating how much today's brand campaign contributed to sales six months from now — separate from every other marketing activity running at the same time — is extraordinarily hard. Multi-touch attribution models only credit the touchpoints they can measure: a search click, an email open. But those touchpoints are often just the last leg of a journey for someone who already cared about the brand before that. What the brand quietly built up in someone's memory, over time, sits outside that model's field of view.

Because it's hard to measure, brand loses arguments too. Performance marketing shows up to the budget meeting with this quarter's ROAS. Brand marketing shows up with a concept: "long-term equity." In a resource-allocation meeting, a number beats a concept, almost by default. The result is that investment flows toward whatever is easy to measure, and brand equity — hard to measure — gets systematically underfunded.

Firms like Interbrand and Brand Finance, with their methodologies for converting brand value into a dollar figure, are attempts to correct for that disadvantage. They help put brand on the table as a number, in a room where numbers win. But the moment that number becomes a target to be managed toward, Goodhart's Law reasserts itself right on schedule. The urge to compress everything into a single figure is exactly the material this trap is made of.

A more fundamental response is to give up the urge to express brand impact as one single number at all. A brand health scorecard — tracking awareness, preference, trust, differentiation, and purchase consideration simultaneously, as separate dimensions — gives a fuller picture than any single metric can. It looks less clean. It doesn't fit into one cell on a slide. But looking less clean is, in this case, closer to the truth.

Use the Map, Not the Territory

So what's actually available to do. Dropping metrics isn't the answer. Staying conscious, continuously, of the relationship between a metric and the real value it stands in for — that's closer to an answer.

Start from the question, not the metric. Before deciding what to measure, decide what state you actually want to exist. Only after that do you ask how you'd know that state exists, and which metric shows it best. Starting from "what can we measure" and reverse-engineering a goal to fit it is the most common road into the trap.

Before adopting a new metric, ask two questions first. When this number rises, can we confirm that what we actually want is happening? And among the ways to make this number rise, is there one that moves the metric up while moving the real goal further away? A usable metric answers yes to the first and no to the second. If either answer is uncertain, or if the second answer is yes, the metric is already fragile.

Don't rely on one metric. But adding more metrics isn't, by itself, the fix. Twenty metrics aren't clearer than five. They just multiply the number of things you don't know which one to trust. What matters is reading the relationship between metrics. If NPS is climbing while repeat-purchase rate is falling, that's a signal that one of the two is closer to reality than the other. Chasing down that contradiction, instead of ignoring it, is what it means to use metrics well. Watch short-term and long-term metrics side by side, too. Track this quarter's new customer acquisition against that same cohort's repeat-purchase rate six months out, and you can see how the quality of what you acquired reveals itself over time.

There's also a way to catch, early, the moment a metric splits from the real goal. Ask the team to list every way they can imagine hitting the metric. If even one item on that list would raise the number while eroding real value, the metric is already fragile. Another warning sign: the metric looks fine but the ground truth doesn't. If the dashboard is green while complaints to customer service are climbing, that mismatch is an early signal that something in the metric's design has gone wrong.

Amazon handles this balance by separating leading and lagging indicators. Revenue and profit are lagging indicators — they show a result that's already happened. Customer experience quality, breadth of selection, and price competitiveness are leading indicators — they forecast what's coming. Translated to marketing, brand awareness and unprompted mentions within a category are leading indicators. They won't show up in this quarter's revenue, but they forecast the revenue of the quarters after that. Organizations that watch only lagging indicators are always a beat behind.

And put an expiration date on every metric. Once adopted, a metric develops inertia. Reporting structures and incentive systems build on top of it, and even once it's clear the metric no longer stands in well for the goal, changing it becomes hard. Setting a review date — "revisit this in a year" — at the moment you adopt a metric is how you catch the moment it starts losing its life, before that moment slips by unnoticed.

Don't fix the review cycle to a single cadence, either. Watching only monthly numbers means managing only what's visible within that one month. Holding weekly, quarterly, and annual metrics together lets you see, in layers, whether whatever's improving the short-term number is quietly eating the long-term one. That's more work. But that extra work is usually smaller than the risk that oversimplification creates.

A metric is a map. Reading a map well and mistaking the map for the ground are two different things.
The map is useful. Mistaking it for the ground is the trap.

Neither the teachers in Atlanta, nor the team chasing MAU, nor the marketer narrowing the survey sample were bad people. Each of them was moving diligently, in their own seat, toward the number they'd been handed. The problem was never the sophistication of the tool. It was that no one kept measuring the distance between where the tool pointed and where they actually meant to go. Building a sharper map is still worth doing. Just don't forget, map in hand, that the ground beneath your feet is not the map.

A better map still isn't the ground you're standing on.