Home / Blog
Measurement

KPIs That Survive the Boardroom: Validating Metrics Against Bias and Vanity

The tests I run on any proposed metric before it earns a place on a dashboard, and why most metrics fail at least one of them.

Every organization has more metrics than it has decisions. Dashboards accumulate, reports multiply, and by the time a KPI reaches the boardroom it has usually survived three rounds of "what should we track" workshops without anyone asking whether it's actually useful. The result is a slide deck full of numbers that go up, which is a different thing from a slide deck full of numbers that tell you something.

A KPI that survives a boardroom presentation is not necessarily a good KPI. A KPI that survives the questions a rigorous executive will ask is. Those are different tests, and the second one eliminates most of what makes it onto the first list.

The difference between a metric and a KPI

A metric is any measurement. Sessions, revenue, error rate, headcount: all metrics. A Key Performance Indicator is a metric that meets a higher bar: it measures progress toward a specific objective, it's owned by someone who can act on it, and its movement matters to a decision that would otherwise be made differently.

The distinction matters because the word "KPI" has been diluted to the point of meaning "any number on a dashboard." When everything is a KPI, nothing is key. The practical consequence is that teams spend time explaining movements in metrics that nobody can act on, in meetings where the question "so what do we do?" never gets a concrete answer.

The five tests below are how I separate metrics from KPIs before either one reaches a dashboard.

Test 1: measurability

Can this be calculated consistently from available data, with a definition that would produce the same number if two analysts ran it independently? Measurability fails more often than it should on the definition part: "active users" defined as any user who logged in last month will produce a different number from "active users" defined as any user who completed at least one transaction last month. Both are defensible definitions. Neither is wrong. But they can't coexist as the same KPI, because a trend in one tells a different story from a trend in the other.

Before a metric earns its name, write down its definition in one sentence, its data source, its calculation logic, and the edge cases (new users, refunded transactions, deleted accounts). Two analysts, the same definition, the same data, the same number. Anything less is an approximation that will be argued about in exactly the meeting where you need clarity.

Test 2: actionability

If this metric moves in the wrong direction, is there a concrete action someone can take in response? Actionability is where most aggregate metrics fail. Total revenue is not actionable: it's an outcome of dozens of decisions made across the organization, and knowing it went down doesn't tell you what to do. Revenue by segment, by channel, by cohort: those start to be actionable, because they point to where to look and what lever to pull.

The actionability test has a useful corollary: if the same response would follow regardless of the direction of movement, the metric isn't driving decisions. It's being reported. Reporting and measurement are different activities, and confusing them fills dashboards with numbers that inform nobody.

Test 3: relevance

Does this metric measure what the underlying objective actually requires, or does it measure something correlated with it that's easier to count? Relevance failures are common in product and marketing contexts: measuring page views as a proxy for engagement, measuring calls closed as a proxy for support quality, measuring training completions as a proxy for capability development. The proxy is easier to count, and it may correlate with the real thing, but when incentives shift, the proxy and the real thing decouple. Goodhart's Law: when a measure becomes a target, it ceases to be a good measure. Teams optimize for the number, not the outcome, and the number goes up while the outcome stays flat or gets worse.

The relevance test asks: what would have to be true about this metric's relationship to the objective for it to be a valid indicator? Then asks whether that relationship is stable, or whether it can be broken by a change in behavior that the metric itself would reward.

Test 4: unambiguity

Does an increase in this metric unambiguously mean the situation improved? Many metrics fail this test in context-dependent ways. Customer contacts per day: high is bad if contacts are complaints, high is good if contacts are sales inquiries. Inventory turnover: high is usually good, but not if it means you're running out of stock. Support ticket volume: a spike could be a service failure, or it could be the consequence of a new product launch that's going well.

The ambiguity doesn't make the metric useless: it means the metric needs context, either a counter-metric that resolves the direction, or a segmentation that shows which type of movement is happening. A metric that requires contextual interpretation in every boardroom meeting is doing more work than it should.

The counter-metric requirement

Every KPI should have a counter-metric: a measurement that guards against optimizing the primary metric in a way that damages something else. Response time without a quality counter-metric rewards speed at the cost of accuracy. Cost reduction without a customer satisfaction counter-metric rewards efficiency at the cost of experience. Conversion rate without a retention counter-metric rewards acquisition at the cost of fit.

The counter-metric doesn't have to be reported alongside the primary metric in every context. It has to exist, be monitored by someone, and be visible to the person who owns the primary metric. Its absence is how organizations end up with a team that hit every target and still made things worse.

A KPI without a counter-metric is an invitation to game it. It isn't that the people tracking it are dishonest. It's that optimization is what metrics do to behavior: they focus attention, and focused attention finds the shortest path to the number, not always the right one. The counter-metric is what makes the shortest path and the right path the same path.

Validating against vanity

A vanity metric is one that makes the organization feel good about itself without providing information useful for making decisions. High absolute numbers, monotonically increasing curves, metrics that only move in the direction that reflects well on the team reporting them: all warning signs.

The test is simple: would you report this metric if it were going down? If the answer is no, or if the answer is "yes, but we'd add context explaining why the number going down is actually good," the metric is probably serving a political function rather than an analytical one. Political functions are legitimate: stakeholders need confidence, boards need reassurance, and not every number on a slide is a decision-making tool. But political metrics should be called what they are, and they should not occupy the slot where the decision-driving KPI belongs.

The validation checklist I use

Before any proposed KPI goes on a dashboard or into a report, I run through six questions. A metric that can't answer all six is not a KPI yet.

A boardroom that can answer all six of those questions for every KPI on the slide is a boardroom making decisions. A boardroom looking at a dashboard where those questions were never asked is doing something else. It might be important. It might be reassuring. It's probably not the thing that moves the business.

← Previous
BPMN in Practice