How to compare candidates for the same role without comparing impressions
Four conditions that have to hold before two interviews can sit side by side, and what to do with the table once they finally can.
Method · Comparison and decision
Published on · 6 min read
Two interviews are comparable only if they measured the same thing with the same instrument: the same items, the same scale, the same version of the script and the same evidence available to check them against. Once those four conditions hold, the comparison is read by competency rather than by candidate, and ties are broken by reading the fragments, not by trusting the impression each conversation left.
Three people interviewed the eight candidates for a role and now a decision has to be made. On the table there are eight sets of notes, each with its own idea of what mattered that day. What follows is not a comparison: it is a negotiation between three memories, and the most confident memory wins it.
Comparing candidates is a problem of instruments before it is a problem of judgement. What follows is what has to be true before two interviews can sit side by side, and what to do with the table once they finally can.
Four conditions, and all four are necessary
Two interviews are comparable when they measured the same thing with the same instrument. It sounds obvious and it almost never holds, because each of the four conditions breaks on its own and none of them announces that it broke.
- The same items. If one person was asked about a conflict with a supplier and another about a conflict with their manager, what you have is evidence about two different situations. Both can be good and still cannot be subtracted from each other.
- The same scale. A level 3 has to mean the same conduct in both assessments. If the scale is written in terms of observable conduct that holds on its own; if it says “handles conflict well”, every reading reinterprets it.
- The same version of the script. An item tuned mid-process turns the earlier and later candidates into two different populations, even though the table keeps showing them together.
- The same evidence available. An assessment you can go and check and one you cannot do not weigh the same, even when both say 3. The second is an opinion formatted as data.
Miss any one of the four and what is left is a table sorting numbers produced by different instruments. Sorting is still possible; what is lost is the right to say the order means anything. The issue on interview design covers the decisions that make all four hold from the start.
A single average answers a question nobody asked
The temptation, as soon as there is a table, is to collapse it: add the competency levels, divide, sort by the result. A final column appears with one number per person and the meeting is over in ten minutes.
The problem is not that the average is imprecise. It is that it hides a decision nobody made. Averaging amounts to declaring that every competency weighs the same, and that is almost never true for a specific role. An illustrative example, with fictional data: for a maintenance supervisor, someone at level 4 in plant safety and level 2 in communicating upward averages the same as someone at 3 and 3, and they are not interchangeable — the first can learn to write a report, and the second cannot learn not to have an accident.
There are two honest ways out and only one bad one. The bad one is averaging in silence. The honest ones: declare the weights in writing before seeing any results, and then sort; or do not sort at all and keep the full profile visible, which is what a decision actually needs. The second works better than it sounds: hiring managers rarely want the best average, they want the person with no hole in the competency that cannot fail.
If there are going to be weights, they are declared before the first interview and versioned like the rest of the script. A weight chosen after seeing the table is not a criterion: it is the technical way of justifying the person somebody had already picked.
A comparison view is read across rows
The reflex is to read down the columns — candidate by candidate — because that is how we arrive at the meeting, with one person in mind. But the new information is in the rows: a row is a competency, and it tells you how the whole pool is distributed on it.
| How you read it | What question it answers | When it helps |
|---|---|---|
| By row | How does the pool stand on this competency? | First. It is where problems with the item and with the market show up. |
| By column | What profile does this person have? | Later, and over two or three people, not eight. |
| By cell | What exactly did they say to earn this level? | Last, and only where the decision is actually close. |
A whole row sitting at level 2 almost never means all eight candidates are weak on that competency. It means the item did not get what was needed, or the scale asks for something that market does not have. Both are findings, and both get fixed before the next role instead of by rejecting eight people.
Ties are broken by reading, not by remembering
Two people with the same profile is where impression comes back in through the side door. “I just had a better feeling about her” is a sentence that always turns up and that cannot be written into any record.
The tie is in the level, not in what they said. Two level-3 answers can be very different. An illustrative example with fictional answers, on the same competency:
When the supplier missed the delivery I got procurement and the plant in a room the same day, documented the impact on the line, and brought two costed alternatives before escalating to the director.
I called the supplier that same afternoon, he gave me a new date, and I let the plant know so they could adjust the schedule. We did not lose output, but I did not leave anything in writing either.
Both act the same day and both avoid the damage, so with a reasonable scale both are level 3. But for a role where the recurring problem is that nothing ever gets documented, the first says something the second does not. That nuance lives in the fragment and vanishes in the number: which is why a tie is broken by reading the fragments of the heaviest competency side by side, and not by arguing about the table again.
And if the tie survives the reading, the tie is real and has to be settled with what is not in the table: availability, expectations, fit with the team. That is legitimate too. What is not legitimate is dressing it up as a result of the assessment.
A comparison expires if the script moves without a trace
A comparison table is a photograph of an instrument at a moment. If the instrument changes and the photograph does not record it, six months later nobody can tell whether the March candidate and the May one answered the same thing.
Which is why versioning the script is not a process formality but a condition of the comparison. The issue on evidence and traceability covers exactly what has to be kept for that photograph to be developed again, and the security page covers how long it is kept.
What the comparison does not do
It sorts evidence, not people. It says who demonstrated what against the same questions, which is far more than most processes have, and it does not say who will do the job well. That second question needs the conversation with the team, the reference check, and the judgement of someone who knows the role from inside.
Nor does it fix a badly defined role. If nobody is clear on what is being looked for, the comparison will stay perfect and stay useless: eight people assessed precisely against the wrong criterion. Structure makes answers comparable; what to ask is still a human judgement, and it is made first.
Questions about this issue
So should candidates never be ranked?
Ranking is legitimate when the weights are declared in writing before any results are seen. What is not legitimate is an order that appears on its own, produced by an average nobody chose, and is then defended as if it were a finding.
How do you compare someone interviewed three months ago with someone from this week?
Only if both interviews ran on the same version of the script, and that can only be asserted if the script is versioned and each interview remembers which version it ran with. Without that, the comparison can be made but not defended.
What if the hiring team disagrees with an assessment?
That is the best possible use of cited evidence: you go to the fragment and argue about what the person said, not about the level. Those arguments usually produce a better-written scale, which is what was missing in the first place.