August 13, 2026 · 5 min read · Performance

Your review process was built for a world where output was hard

Chief People Officer

A head of engineering said something to me in the spring that I've been chewing on ever since. He was looking at his team's half-year numbers and he said, almost apologetically, "everyone's a top performer now and I don't believe it."

He wasn't being unkind. The numbers had genuinely gone up — throughput, tickets, merged work, all of it — across the whole team, including the people he'd have told you six months earlier were struggling. His review process was going to hand out uniformly strong ratings, his budget wouldn't cover uniformly strong outcomes, and he had no defensible basis on which to differentiate.

That's the problem in miniature, and it is arriving in a lot of organisations at once. Most performance systems, whatever they claim in the framework document, ultimately reward visible output. That was a reasonable proxy for as long as producing output was the expensive part. It isn't any more, and a proxy that everyone can now satisfy has stopped being a measurement.

The two bad reflexes

I've seen two instinctive responses and I'd argue against both.

The first is to raise the bar quietly. If everyone's doing twice as much, expect twice as much, and carry on with the same scale. This is what happens by default, because it requires no decision and no announcement. It's also how you get a workforce that received a genuinely useful tool and experienced it as an increase in workload with no acknowledgement — which is a fast route to the sort of resentment that doesn't show up in an engagement survey until it's structural. If you are going to reset expectations, reset them once, deliberately, out loud, with the reasoning attached. Drift is the cruellest way to do it.

The second reflex is to start measuring AI use itself. Put "adoption of AI tools" in the objectives, track licence usage, reward the enthusiastic. I understand the impulse — leadership wants to see the investment landing — but you are measuring an input, and the moment you attach a rating to an input you get theatre. People will use the tool for things it isn't good for, and the person who correctly judged that their work didn't need it gets marked down for good judgement. I've already watched this happen at one organisation. Six months of impressive dashboards and no discernible change in anything that mattered.

What's actually left to distinguish people by

If output is now cheap and roughly equal, what's expensive and unequal?

Judgement, mostly. Which is a soft word for a set of quite hard, specific things.

Choosing the right problem. When it takes an afternoon rather than a fortnight to produce a thing, the cost of producing the wrong thing drops too — so more wrong things get produced, and the person who works out what's actually worth doing becomes disproportionately valuable. That has always been true. It's just no longer masked by the effort involved in execution.

Knowing when the answer is wrong. This is the one I'd weight most heavily, and it's the hardest to see. The engineer who noticed that the generated code handled the timezone case incorrectly, the analyst who spotted that the summarised figures had double-counted a region — they prevented something, and prevention is invisible in every performance system I've ever read. If you don't deliberately go looking for it, you will rate the person who shipped fast above the person who stopped them shipping wrong.

Deciding not to. Restraint has never been rewarded in performance reviews and it's more valuable now than it's ever been.

How to evidence that without it becoming a personality contest

Here's my real worry, and I want to name it plainly rather than let it sit under the surface.

When you shift assessment from countable output to judgement, you shift it towards things that are argued for rather than demonstrated. And the people who are good at arguing for their own contribution are not a random sample. They skew towards the confident, the fluent, the socially central, the ones who look like the person doing the assessing. Every organisation I've worked in has had at least one quiet, excellent person whose value was obvious to their team and invisible to the process. Making the criteria more subjective makes that worse, systematically, and it makes it worse for exactly the groups your inclusion data is already telling you about.

So the criteria have to get more subjective and the evidence has to get more concrete, at the same time. That's the needle.

Practically, what I'd change. Ask for decisions, not volume — three things you decided this half, what the alternatives were, and how it turned out. Ask managers to record catches as they happen, in a line or two, in the moment, because nobody reconstructs them at review time. Ask peers a specific question rather than a general one: not "rate this person's collaboration," but "when did this person save you from something?" And separate the two questions that reviews currently conflate — did the organisation get value from this area of work, which is a team-level question, and did this individual work well, which is not the same thing and increasingly isn't even correlated.

One more, and it applies to anyone with a promotion ladder. Most ladders assume that doing the junior work well is how you demonstrate readiness for the senior work. If the junior work is now largely automated, that evidence path has quietly closed, and you will find yourself with a cohort of people who have no way to prove they're ready. That's a design problem in the ladder, not a shortcoming in the people, and it needs fixing before the first promotion round where it bites.

None of this is a small edit to the form. It's a change to what you believe good work is, which is why most organisations will do the quiet-bar-raise instead. I'd just note that the ones who take it seriously will be able to answer the question my engineering friend couldn't, and the ones who don't will spend the next few years handing out ratings they don't believe.

Chief People Officer

Bhavna leads people at Partech Systems. She spent two decades inside large, high-stakes technology organisations — across TCS and the National Stock Exchange of India — before concluding that the hardest part of shipping software was never the software; it was the people shipping it. She writes about how teams actually absorb new technology: the tradeoffs nobody likes talking about, the anxieties, and the unglamorous work of making a change stick.

Everything by Bhavna →