Regis Tremblay

Writing about work: who does it, on what terms, and how the claims made about it compare with what has been measured.

Measurement ยท 5.4

Performance reviews, and their reliability

Performance reviews, and their reliability. What is actually the case, and how it compares with what is repeated.

An annual rating is treated as a measurement of a person. Research on where the variance in ratings actually comes from suggests it is closer to a measurement of the rater.

The finding

Studies decomposing the variance in multi-rater performance ratings have repeatedly found that the largest single component is idiosyncratic rater effect, in the region of half of the total variance, and that the component attributable to the ratee's actual performance is considerably smaller.

In plain terms: knowing who did the rating tells you more about the score than knowing who was rated.

That result has been replicated and it is the central fact about performance ratings. Everything else in this entry follows from it.

The familiar distortions

Recency, in which the last two months dominate the year. Halo, in which one strong impression colours every dimension. Central tendency, in which everybody receives the middle box. Leniency, which varies systematically by manager and is why calibration meetings exist.

None of these is news to anybody who has conducted a review, and knowing about them does not remove them, which is the standard finding about cognitive biases in judgement.

Forced distribution

Requiring a fixed proportion in each category. Adopted widely, associated with damage to collaboration for exactly the reason a zero-sum scheme would predict, and abandoned by several of its most prominent users.

It does address leniency. It addresses it by imposing an assumption about the distribution of performance within every team, which is false in any team small enough to matter.

Calibration

Managers comparing ratings across teams before finalising. It reduces leniency differences and introduces a different problem, which is that the outcome now depends on who argues well in a room.

It is a genuine improvement over unmoderated rating and it is not a solution.

The two purposes that conflict

Development requires candour, which requires that saying something difficult is safe. Pay allocation requires a defensible number, which makes candour expensive for both parties.

Running both through one conversation guarantees that the pay purpose wins, because it has consequences attached. Organisations that separated them report better development conversations, and the separation is cheap.

The abolition wave

A number of large employers removed annual ratings during the twenty-tens, replacing them with more frequent, less formal check-ins. Several later reintroduced a rating in some form.

The reason given for the reversal is consistent: pay and promotion decisions still had to be made, and in the absence of a rating they were made on less visible grounds. Removing the instrument did not remove the decision.

What improves ratings

Frequent specific feedback close to the event, which is the only intervention with a good evidence base. Rating behaviours rather than traits. Multiple raters, which averages out some of the rater effect. And asking raters what they would do rather than what they think, since predictions of one's own behaviour are better calibrated than judgements of somebody's quality.

What to do with a rating you receive

Read it as partly information about your manager. Ask for specific instances rather than general assessments, because instances can be discussed and assessments cannot.

And note that the distribution of ratings across the organisation is frequently published or discoverable, which tells you whether you are reading a judgement or a quota.

What a rating is actually used for

In most organisations, to allocate a fixed pay pool. The rating is the justification produced after the allocation constraint is known, and managers describe the process this way privately with some consistency.

Recognising that would improve the conversation: an honest statement that the pool is limited is easier to hear than a contested judgement about performance.

The self-assessment

Asking employees to rate themselves first is standard and it anchors the manager's rating, which is either useful or a contamination depending on what the exercise is for.

It also produces a systematic asymmetry: confident people rate themselves higher and receive higher ratings as a result, which converts a personality difference into a pay difference.

The nine-box grid

Plotting performance against potential and placing people in one of nine cells. Potential is even less reliably assessed than performance, has no agreed definition, and correlates strongly with how much the assessor resembles the assessed.

The grid is used for succession planning in a large number of organisations and its predictive record has, as far as we can find, never been published by anybody.

The written record

A rating becomes a document that follows a person through promotion decisions, redundancy selection and references for years. Its reliability does not improve with age and its authority does.

What this rests on

  1. Variance decomposition studies of multi-rater performance ratings, finding rater effects to be the largest component, are published and have been replicated.
  2. Forced distribution systems and their adoption and abandonment by large employers are documented publicly.
  3. The reintroduction of ratings after abolition has been described by the organisations concerned.

For broader context, consult CIPD performance-appraisal guidance.