The Merry-Go-Round

Picture five team leaders in a room near the end of the year, working out an unwritten agreement: each of us will sacrifice one of our own people into the bottom performance rating, we will take turns so that nobody’s team bleeds twice in a row, and the deal extends upward too, because each year one of the five leaders themselves has to wear the bottom rating, regardless of how well their team performed. My senior leader at the time proposed it himself, openly, as the sensible way to run the year-end process. We called it the merry-go-round.

If that sounds absurd, hold the thought, because the absurdity is the point. The merry-go-round was the rational response to the system we were given, and once you see why, you start recognising the same machinery in every forced ranking exercise you’ve ever sat through, including the ones dressed up as redundancy selection.

Here’s the system. End-of-year appraisal, rank and stack. Every team sorts its people into buckets, call them A, B and C: exceptional, meets expectations, below expectations. And a fixed percentage of every team must land in C. Say 10%, and that percentage is compulsory. The spreadsheet does not ask whether your team actually has poor performers. It asks you to produce them.

Now suppose your team is genuinely high performing. Suppose you’ve spent 2 years building it, the delivery record shows it, and the team has been beating every target set for it. The quota doesn’t care. Someone in that room is going to carry a below-expectations rating home to their family, and your job as their leader is to choose who, knowing the choice has nothing to do with what they did this year.

The standard defence, and I heard it every year, is that the C ratings get compared across teams later, so the process is fair in aggregate. That comparison has a name: recalibration, and I watched how it actually ran. All the leaders come together in a room and rank everybody’s proposed ratings against each other, with HR coordinating the process. After that, the senior leaders hold their own round and rank their direct reports, the leaders themselves, the same way we had just ranked our teams. I’d guess there were further rounds above that, with rules I never got to see. And here is where the whole thing quietly falls apart, because the people settling the ratings are 2 or 3 levels removed from the actual work. They have never seen your team member deliver anything. The session runs on questions like “who is so-and-so?” and “what has so-and-so done this year?”, and a person’s rating, bonus and reputation get settled on the strength of whatever answer happens to be in the room.

Which means the real variable in your rating is your leader’s debating skill. If your team leader is articulate, if they can run logos, pathos and ethos across a conference table and argue a proposed C back up to a B, power to you. If your leader can’t debate, or won’t fight for you, then I’m sorry mate, that’s it. The rating follows the advocacy, the advocacy follows the personality of your leader, and none of it has much connection to your work. There was no shared benchmark for what performance meant across teams, so every leader rated arbitrarily, and recalibration didn’t remove the arbitrariness, it just decided whose arbitrariness won.

Once you understand that, the merry-go-round stops looking crazy and starts looking like game theory. If the quota is fixed before anyone looks at the work, the only question the system leaves open is who absorbs the damage. And once that’s the question, rationing the damage fairly, taking turns, spreading the pain across teams and across years, is exactly what reasonable people do. The five leaders in that room were queuing politely for a punishment the system insisted on handing out. They even queued themselves into it, since one leader per year took the C rating personally. You can call that integrity of a sort. The system asked for sacrifices; they organised a fair roster of sacrifices.

There’s a quieter cost underneath the game theory, and it’s the one the case studies keep naming. Forced distribution isolates people from their own performance. The rating a person receives stops being information about their work and becomes information about the quota, the roster and the room, yet it lands on them as if it were a verdict on the work. I’ve sat across the table from a team member who met every expectation we agreed at the start of the year and told them the organisation had rated them below expectations, and we both knew why, and neither of us could say it out loud. You do that to someone once and something doesn’t come back: the engagement goes first, and the trust goes with it. Multiply it across the team and you get the second effect: people stop helping each other. When the buckets are fixed, your teammate’s good year raises the odds that the C lands on you, so collaboration quietly turns into competition, inside the same team, among people whose work depends on each other. The system doesn’t announce this. It just prices it in.

None of this was invented locally. The apparatus was imported, mostly from the United States. Jack Welch ran it at GE as the vitality curve: celebrate the top 20%, keep the middle 70%, remove the bottom 10%, every year, forever. For a couple of decades the big corporates copied it as best practice, Microsoft among them. Microsoft finally abandoned stack ranking in 2013, after years of it being cited, internally and in business school case studies, as a system that made employees compete against their own teammates and that people experienced as demoralising and unfair. The verdict was in long before most organisations stopped. It kept travelling anyway, because forced distribution looks rigorous on a slide, and because it spares senior leaders the much harder job of actually knowing the work well enough to evaluate it.

And now the part that still makes me laugh, in the way you laugh at things that cost people real money. Suppose the organisation finally reads the case studies and decides to drop the system. Whoever’s turn it was on the merry-go-round in that final year is now carrying a below-expectations rating on their permanent record, for a rotation that no longer exists. Conned is the polite word for it. I know, because it happened to me. I objected, I brought the delivery evidence, and I had argued against the rotation itself in the calibration room more than once. My leader insisted it was my turn that year. Maybe my time was up for arguing. The wheel handed me the C, the system got retired, and the record stayed. I left not long after.

Redundancy runs on the same maths

Performance appraisal is the annual version of this machine. Redundancy is the crisis version, and it runs on the same maths: the number comes first, the evaluation comes after, if it comes at all. When an organisation decides it needs to remove some percentage of a division, the selection is another rank and stack, done quickly, by leaders who often have no better method for deciding who should stay than they had for deciding who was a C.

So the selection follows patterns that anyone who has watched a few rounds can recite. If you’re in the leader’s bad books, you’re a target. If you’re a mediocre performer, you’re a target. If you’re a poor performer, obviously. And if you’re a top performer, you may be a target as well, because nobody loves the most hardworking person in the team, and in New Zealand the tall poppy syndrome makes standing out its own kind of exposure. Good luck to the tall poppies.

The narrative wrapped around the process will talk about skills matrices, future capability requirements and fair selection criteria. Half of it is bullsh*t. The other half, depending on the organisation, is who you know and who you hang out with. That’s a blunt thing to write, and I’ve sat on both sides of these processes over the years, as the leader filling in the buckets and as a name on somebody else’s list, so I’ll stand behind it.

So what do you do with this, other than get cynical? Back in 2008 I read a book called Bulletproof Your Job and wrote about it on this site. I won’t repeat the whole playbook here, but the people I’ve watched survive these systems, myself included on the good years, did a few things consistently. They understood that the rating system is a game running on top of the work, with its own rules about quotas, turns and advocacy, and they never assumed the work would speak for itself. They paid attention to whether their leader could and would argue for them, because in a recalibration room your leader’s voice is your performance. They kept their own record of what they delivered, with numbers, because “what has so-and-so done?” gets answered with whatever evidence happens to be at hand. And when someone proposed an unwritten agreement where you volunteer for the bottom bucket because it’s your turn, they remembered that the wheel can stop at any time, and that it stops on whoever is in the worst seat.

That’s what I observed. Consider what applies to your situation.

I’ve taken my ride on the merry-go-round. Once was enough.