Criterio Talent
Issue 01 · September 2025

Competency matrices and rubrics: how to write levels two people read the same way

Most competency matrices are a list of nouns with a one-to-five scale. That is not an instrument, it is a form. The instrument is the descriptors.

Method · Assessment instruments

Published on · 7 min read

In short

You write it the other way around from how it is usually done: first the descriptor for each level in terms of observable conduct, and only then the number that orders them. A level is well written when two people on the team read the same answer separately and assign the same thing; while they disagree, the problem is in the rubric, not in whoever applies it.

Ask any HR team for their competency matrix and you will almost always get the same object: a column of nouns — communication, leadership, results orientation — and a one-to-five scale beside it. Nothing else. That is not an assessment instrument; it is a form for recording what each person had already decided on their own.

The instrument is the descriptors: the sentence that says what a person had to do for their answer to be worth a three rather than a two. Writing them takes a few hours, and it is the work that decides whether everything downstream means anything. This issue is about how to write them.

A number without a descriptor informs nobody

“Communication: 3 of 5” tells three readers three different things. To the interviewer it may mean the person explained a complicated process well; to whoever reads the report, that they speak correctly but without spark; to the hiring manager, that they have no idea whether that is enough to stand in front of a committee.

The number is a compression of a judgement. Without the descriptor, the compression cannot be undone: nobody can reconstruct what was measured. And because it cannot be reconstructed, it cannot be argued with, corrected or defended either. The practical consequence shows up months later, when someone asks why a person advanced and all that exists is a table of numbers nobody knows how to read any more.

The quick test of whether a matrix works: cover the numbers on a report and leave only the descriptors. If a reader still understands what was concluded, the matrix works. If nothing survives, the numbers were all there ever was.

Levels describe conduct, not people

“Is proactive” is not a level, it is a label. Nobody can point to the sentence in a transcript where the person was proactive, because proactivity is not said, it is inferred. And everything inferred without being quotable is exactly what cannot be defended afterwards.

“Got ahead of the contract renewal and brought two quotes before being asked” is a level. It describes something that happened, that the person can recount, and that survives in the text of what they said. The rule for writing descriptors is that one: if a level cannot be quoted, it is not a level.

There is also a reason that is about risk rather than method. Attributes — “has presence”, “inspires confidence”, “is a fit” — are doors through which prejudice walks in unnoticed, because they never force anyone to name a concrete behaviour. The issue on bias takes that apart properly, but it is worth knowing while writing the rubric: a vague descriptor is not merely imprecise, it is permissive.

A complete rubric, as an example

What follows is a fictional rubric, written for this article, on supplier management in an industrial operation. Four levels, and for each one what has to appear in the answer to assign it. Note that no level mentions a trait of the person.

A fictional rubric, written as an example. No company uses it as it stands: levels are written with the words and the thresholds of each operation.
LevelWhat the person didWhat has to appear in the answer
1 of 4Describes the supplier problem, but no action of their own.Tells the story in the third person or in the plural. No first-person verb with a decision behind it.
2 of 4Acts, but only once the problem had already escalated on its own.An action of their own appears, with a date or a sequence placing it after somebody else raised it.
3 of 4Acts on their own initiative, brings in the affected areas and quantifies the impact.Their own action, who they coordinated it with, and a magnitude: days of downtime, units, cost, or a concrete breach.
4 of 4All of the above, and leaves a mechanism so it does not happen again.Beyond level 3, something that stayed in place: a clause, an indicator, a second source, a recurring review.

With that rubric in hand, an example answer like this one places itself:

When the packaging supplier missed the delivery, I sat down with planning the same day, we worked out it was two days of downtime, and I proposed moving the order to the secondary source we already had approved.

That is a three: there is an action of their own, there is coordination, and there is a magnitude. It is not a four because nothing appears that stayed in place afterwards. And the point is not that you agree with me, but that any two readers can argue about it while pointing at the same sentence.

The calibration test: two readers, the same answer

A rubric is not finished when it is written. It is finished when it is calibrated, and the procedure fits into an afternoon: gather five answers — from earlier processes, or written by hand for the test — have two people on the team read them separately and assign a level without talking to each other. Then compare.

The interesting part is not the agreements but the disagreements: each one points at exactly which boundary between two levels is badly written. If an answer gets a two and a three, the line between two and three does not say enough. You fix the descriptor, not the reader.

This test is worth running before the first role goes out, and again whenever the rubric changes. It is also the strongest argument to make in front of any automated system: if two people on the team cannot agree using the rubric, no model is going to produce a defensible result with it. On what the system does with a rubric once it is written, the platform page describes the whole path.

The four failure modes that keep recurring

  • Levels that overlap. Two descriptors that both fit the same answer. Calibration finds it, and the fix is moving a condition from one level to the other, not adding adjectives.
  • Attributes dressed as conduct. “Communicates clearly” looks like conduct and is not: it does not say what the person did. “Explained the problem to a team unfamiliar with the process and checked that they had understood it” does.
  • Five-level scales where only three get used. If nobody ever assigns the one or the five in practice, the real scale has three levels and it is better written that way. A scale with decorative extremes compresses everything into the middle.
  • Too many competencies. A matrix of twelve competencies is not an instrument, it is a wish list. Nobody can explore twelve things in one conversation, and what comes out is twelve unsupported assessments instead of four solid ones.

How many competencies actually fit

Three to five per short round. That is not a number from a textbook: it is what fits once each competency needs a behavioural item, its follow-up, and the time for a person to tell something with a beginning and an end. If the role asks for more, they get spread across the rounds of the process; the arithmetic is in the issue on what fits into one round.

How you choose them matters too. Competencies that separate candidates are useful; the ones everybody meets, or nobody meets, are not. If everyone reaching the first round communicates equally well, “communication” is not measuring anything, it is decorating the report.

What a rubric does not solve

It does not define the role. If the hiring team and HR do not agree on what is being looked for, the rubric will formalise the disagreement with great precision: every candidate will be assessed with the same confusion, and now it will be documented.

Nor does it remove judgement. Somebody has to decide which level is enough for this role and this market, and that decision is human, arguable and context-dependent. What the rubric does is move the argument onto ground where it can actually be had: about what the person said they did, rather than about the impression they left.

Questions about this issue

Is it worth reusing a purchased matrix or a generic framework?

As a starting point, yes; as an instrument, almost never. Generic frameworks bring the nouns, which is the easy part, and descriptors written for any company at all, which is the part that has to be rewritten with the thresholds and vocabulary of your own operation.

How often should a rubric be revised?

When calibration starts failing, when the role changes, or when a level stops being assigned at all. And every time it is touched it has to be versioned, or earlier comparisons stop being comparisons.

To keep reading on this site

Solutions

How it is configured for high volume, technical profiles, leadership, and multi-round processes.

Platform

How it runs the interview, follows up, and cites the evidence behind each conclusion.

Integrations

How your ATS requests the interview and receives the report, with nobody retyping anything.

Contact

A 30-minute demo on a real role of yours.

Other issues
Next step

See it with a role of yours on the table.

Thirty minutes: an interview is defined from your job post, walked through the way the candidate sees it, and a report is read with its evidence.