Regis Tremblay

Writing about work: who does it, on what terms, and how the claims made about it compare with what has been measured.

Measurement ยท 5.6

Counting output instead of hours

Counting output instead of hours. What is actually the case, and how it compares with what is repeated.

If hours are a bad measure of contribution, count output instead. The proposal is old, it works in a narrow set of cases, and the cases are the ones already using piece rates.

Where it works

Where output is countable, attributable to an individual or a small team, uniform in difficulty, and where quality is separately verifiable. The entry on piece rates covers this and the conditions are the same because it is the same idea.

Outside those conditions, counting output means counting a proxy, and the entry on Goodhart's law describes what happens next.

The attribution problem

Most valuable work in organisations is produced by several people whose contributions are not separable. Asking who produced a given output frequently has no answer, and any allocation is a convention.

Schemes that allocate anyway reward whoever is best placed to claim, which is a real selection effect and is usually visible to everybody except the people designing the scheme.

Objectives and key results

A widely adopted framework for setting and tracking goals, intended to be ambitious, transparent and explicitly not tied to compensation.

In practice it is frequently tied to compensation, at which point the ambition disappears, because nobody sets a stretching target they will be paid against. The framework's own documentation warns about this and it happens anyway, which is a good illustration that a warning in a document does not survive contact with a bonus.

Results-only arrangements

Schemes removing schedule requirements entirely and judging on output alone have been trialled at several large employers, with mixed institutional outcomes.

The better-evidenced relative is a workplace intervention that increased schedule control and manager support without removing structure altogether. Randomised evaluation found improvements in wellbeing and reductions in turnover intention, which is a more useful result than the more radical version produced.

Where output measurement shifts risk

Onto the worker, as piece rates do. If output falls because a supplier was late, a system was down, or a colleague was ill, and pay follows output, the worker has absorbed a variance they did not create and cannot control.

Which is why output-based pay is normally paired with a floor, and why schemes without one are transferring more than they intend.

What can be counted well at team level

Throughput of standardised units. Error and rework rates. Predictability of delivery, meaning the spread between promised and actual. Time from request to completion.

All four are meaningful, all four resist individual attribution, and all four are more useful for improving a process than for evaluating a person. That distinction is the practical conclusion of this part.

The honest position on hours

Hours are a poor measure of contribution and an excellent measure of availability, and availability is genuinely what many jobs require. Coverage roles need somebody present, and for them the hour is not a proxy for anything: it is the thing being bought.

The critique of hours applies to work where presence is not the service, which returns to the composition argument that opens this site.

What to do

Count outputs where the four conditions hold. Count process measures at team level everywhere else. Do not attach individual consequences to either unless attribution is genuinely clean, and be honest that it rarely is.

Counting for improvement rather than judgement

The same number behaves differently depending on what it is attached to. A defect rate reviewed by a team looking for causes is a diagnostic; the identical rate attached to individual bonuses becomes a target and stops being informative within a quarter.

Nothing about the metric changed. What changed was who benefits from its value.

The estimate that becomes a commitment

Ask a team how long something will take and use the answer as a target and you have converted an estimate into a promise. Estimates then inflate, which is rational, and the organisation concludes that the team has slowed down.

This is the single most common instance of the previous entry inside technical work, and it is almost always attributed to the people rather than to the mechanism.

Story points and the same mistake

Relative estimation units used in software teams, intended as a planning aid and explicitly not a productivity measure. Where they are aggregated upward and compared between teams they become a target, and they inflate, because the unit is defined by the team using it.

Comparing velocity between teams is comparing two different currencies, and it is done routinely.

What this rests on

  1. The objectives and key results framework is documented by its originators and by subsequent practitioner literature, including explicit warnings against linking it to pay.
  2. The randomised workplace intervention increasing schedule control and supervisor support is published with its outcome measures.
  3. Results-only arrangements at large employers and their subsequent history are matters of public record.

For broader context, consult BLS productivity data.