Regis Tremblay

Writing about work: who does it, on what terms, and how the claims made about it compare with what has been measured.

Measurement ยท 5.2

What happens to a measure once it is a target

What happens to a measure once it is a target. What is actually the case, and how it compares with what is repeated.

The observation exists in at least three versions, formulated independently within a few years of each other, which is usually a sign that something real was being noticed.

The three

An economist writing about monetary policy in 1975 observed that any statistical regularity tends to collapse once pressure is placed on it for control purposes.

A social scientist writing in 1979 about quantitative social indicators observed that the more an indicator is used for decision making, the more it will be subject to corruption pressures and the more it will distort the processes it was meant to monitor.

An anthropologist in 1997 produced the compact version everybody quotes: when a measure becomes a target, it ceases to be a good measure.

The four mechanisms

Effort reallocation. Attention moves to the measured dimension from the unmeasured ones. Nothing dishonest occurs.

Selection. Change who is measured rather than what they do: discharge the patients who will recover, admit the students who will pass, decline the difficult cases.

Gaming. Satisfy the definition without the substance. A waiting time target met by not starting the clock. A resolution target met by closing and reopening tickets.

Falsification. The rarest and the one everybody imagines first.

The first three require no bad faith and produce most of the damage, which is why treating this as an integrity problem misdiagnoses it.

Surrogation

The subtler failure: people stop treating the measure as a proxy and begin treating it as the goal itself. A manager who set out to improve care and now genuinely wants the score has not become cynical; the substitution happened without anybody noticing.

It is documented experimentally and it is more common than gaming, because it requires nothing but repetition.

Why targets are used anyway

Because the alternative on offer is usually no accountability at all, and organisations without measurement drift in ways that are worse and harder to see.

The law is not an argument against measuring. It is an argument against measuring one thing and attaching consequences to it.

What reduces the damage

Several measures rather than one, chosen so that gaming one degrades another. Measuring the things you do not want to lose, not only the things you want to improve. Rotating measures so that adaptation cannot settle. And keeping a qualitative channel, because the people doing the work know exactly how the target is being met and will say so if asked.

Separating measurement used for learning from measurement used for consequences is the strongest single intervention, and it is the one most often abandoned when budgets tighten.

The signature of a gamed measure

A distribution with a spike just above the threshold, and an improvement that appears in the measure and nowhere else. Both are checkable and neither requires accusing anybody.

If a metric improved sharply and nothing that the metric was standing in for improved, the metric was the thing that improved.

The version worth carrying

Any measure attached to consequences will be optimised, by decent people, in ways nobody intended. Design as though that is certain, because it is.

The measure that survives

Some measures resist gaming because the only way to move them is to do the thing. Cash in the bank. Whether the aircraft arrived. Whether the patient is alive at thirty days, with the caveat that even this one can be gamed by selecting patients.

The characteristic they share is that the measure and the outcome are the same object rather than one standing for the other, and where such a measure exists it should be preferred to any composite.

Two consequences worth separating

A measure can stop reflecting the underlying thing while the underlying thing improves. Or it can stop reflecting it while the underlying thing gets worse because effort was diverted from it. Both are Goodhart effects and only the second is a loss.

Distinguishing them requires an independent read on the underlying thing, which is exactly what an organisation running on a single metric has given up.

The audit paradox

Introducing measurement to establish trustworthiness signals that trust is absent, and the effort spent demonstrating performance is subtracted from performance. That is a well-described dynamic in the literature on audit cultures and it is not an argument against accountability, only against its unlimited expansion.

Where it was first noticed

In monetary policy, when central banks began targeting money supply aggregates that had previously tracked inflation reliably and found the relationship dissolving as soon as it was used for control. The regularity had existed because nobody was acting on it.

What this rests on

  1. The 1975 formulation appears in a paper on monetary policy; the 1979 formulation concerns quantitative social indicators; the 1997 compact restatement is by an anthropologist writing about audit.
  2. Surrogation, the substitution of a measure for the construct it represents, has been demonstrated experimentally in accounting research.
  3. Documented instances of target gaming in health, education and service operations are extensive and published by auditors and regulators.

For broader context, consult Bank of England note on Goodhart's law.