Criterio Talent
Issue 12 · August 2026

How to design a structured AI interview that leaves useful evidence and keeps the decision human

Eight design decisions, in order. The first is the one almost everyone skips: you do not ask a model to judge, you ask it to document.

Method · Interview design

Published on · 7 min read

In short

You design it by inverting the brief: the system does not score, it documents. First you write the scale for each competency, then the behavioural items that explore it, and you require every conclusion to quote the verbatim sentence supporting it. The decision to advance or reject is taken by an identified person, inside the tool, and it is recorded.

Almost every AI interview implementation I have seen starts in the same place: the conversation is handed to a model and a score is requested. The score arrives, it looks professional, and the first time somebody asks why this candidate got a 7 and that one an 8, there is nothing to show. That is the moment the project dies, even if it takes a few months to find out.

What follows are eight design decisions, in the order they have to be made. None of them is about the model. All of them are about what you ask it for, what you let it see, and what you keep of what it produced.

The system’s job is to document, not to judge

The inversion that changes everything: the main output is not a number, it is a record. A score without the sentence supporting it cannot be defended to the hiring team, to the candidate, or to an internal review. A sentence with its context defends itself, even when the score is argued about.

In practice this means the instruction is not “assess this candidate”, but “for each competency, find in this transcript what the person said that is relevant, quote it, and only then assign a level”. The order matters: if the level is decided before the quote, the quote gets chosen to justify it.

The scale before the question

A competency with no described levels produces a number that means whatever the reader wants. “Communication: 3 of 5” informs nobody. Before writing a single item you have to write what separates one level from the next, and it has to be written in terms of observable conduct, not attributes of the person.

An example, for conflict resolution with suppliers. Level 1: describes the problem but no action of their own. Level 2: acts, but after the problem escalated. Level 3: acts the same day, involves the affected parties and quantifies the impact. Level 4: all of that, and leaves a mechanism so it does not happen again.

Testing whether a scale is well written is simple: two people from the team read the same answer separately and assign the same level. If they disagree, the problem is not the model, it is the scale — and no system is going to fix what the team is not clear about.

An item is a question that can only be answered with a story

“How do you handle conflict?” collects opinions about oneself. “Tell me about a time a supplier failed to deliver and you had to respond” collects conduct. The difference is not stylistic: it is the difference between data you can verify and data you cannot.

A complete answer to a behavioural item has three parts: the situation, what the person did — them, not “the team” — and what happened afterwards. Those three parts are what the scale needs to assign a level, and they are also what has to be defined in writing so the interviewer knows when an answer is incomplete.

And this is where the follow-up comes in, as part of the item and not as decoration. If a person answers this:

Yes, we once had a serious problem with a supplier and we sorted it out between all of us.

there is nothing to assess. There is a situation with no action of their own and no outcome. A well-designed script says explicitly what that answer is missing and what to ask to get it, before moving to the next item. A system that does not follow up produces long transcripts and empty records.

The same question for everyone, or there is no comparison

Structured means exactly this: the same items, in the same order, assessed with the same scale, for every candidate for a role. The moment one candidate is asked something another was not, what you have is not two comparable assessments: it is two different interviews with a table on top.

From which follows a technical consequence that is almost always forgotten: the script has to be versioned. If somebody tunes an item mid-process and earlier interviews do not remember which version they were run with, comparing the March candidate to the May one is a coincidence. Versioning is not bureaucracy: it is what keeps a table meaning something six months later.

Evidence is quoted, not summarised

A summary is an interpretation with fewer words. What holds a conclusion up is the verbatim sentence and its position: which minute, which turn, inside which item. With that, whoever reads the report can go from the verdict to the sentence and from the sentence to the recording, and check for themselves whether the conclusion is sound.

The rule worth adopting is hard and simple: if a conclusion cannot cite, it is not a conclusion, it is an impression — and it comes out of the report. Losing a few assessments costs something. Defending one nobody can trace costs far more.

Architectural consequence: the transcript has to be stored turn by turn and timestamped, not merely processed in flight. A system that transcribes in order to assess and then throws the text away can produce the score, but it will never be able to support it again.

What the system must not look at

The technical temptation, once there is video, is to analyse it: gesture, tone, pace, “confidence”. It is worth resisting for two separate reasons, either of which is enough on its own.

  • The first is validity. For the vast majority of roles, how a person looks or sounds does not predict how they will work; what they say they did does, and it can be verified on top of that.
  • The second is risk. The face, the accent and the rhythm of speech are fairly direct proxies for someone’s origin, age and condition. A system that looks at them is discriminating even if nobody intended it, and there is no way to demonstrate that it did not.

The way to hold this line is not a written policy: it is the architecture. If the component that assesses receives only the text of what was said, it cannot fall to the temptation, and that restriction can be shown. A policy is signed; a data input is checked.

Where the system stops

The system produces an assessment and its evidence. Advancing or rejecting is a human decision, and it has to be taken inside the tool by an identified person, not in a parallel email thread or a separate spreadsheet.

If the decision happens outside, the trail breaks exactly where it matters: the detail of how someone was assessed is kept, and the fact of who decided is lost. Months later, the uncomfortable question is not “how was he assessed?”, it is “who said no?”.

What you must be able to answer six months later

This is the final test of the design. With the process closed and the team already changed, somebody asks about one specific candidate. The system has to be able to answer, without anyone reconstructing anything from memory:

  1. What they were asked, and with which version of the script.
  2. What they answered, verbatim and with its timestamp.
  3. What was concluded for each competency, and which fragment supports it.
  4. Who made the decision, and at what time.
  5. Until when that material is kept, and what happens afterwards.

If any of those five cannot be answered, the structure was decorative. And it is better to find that out at design time than the first time somebody asks in earnest.

What this does not solve

A well-designed structured interview makes whatever stage it is applied to comparable — screening, the technical round, or the conversation with the director — and leaves it documented. What it does not do is fix a badly defined role: if nobody is clear on what is being looked for, structure will only ensure every candidate is assessed with the same confusion.

What it does do, and it is not small, is put an end to the worst outcome of a corporate process: a decision nobody can explain, about people who took the trouble to answer.

Questions about this issue

Does a structured interview make the conversation rigid?

The structure lives in the items and the scale, not in the pace. The follow-up is part of the design precisely for that: the script defines what has to be obtained from each item, and the conversation gets there whichever way the person takes it.

How many competencies fit into a first round?

Fewer than you would like. Four to six items allow depth; a list of twelve produces headline answers that support no level at all. On how much fits into a short interview, the November issue measures it.

Does anything change for an executive role?

The script and the depth change, not the method. An executive search has fewer candidates, but the decision carries more weight and gets reviewed harder: having all three conversations explore the same ground, and having every conclusion able to quote where it was said, is exactly what makes a decision that size defensible.

To keep reading on this site

Solutions

How it is configured for high volume, technical profiles, leadership, and multi-round processes.

Platform

How it runs the interview, follows up, and cites the evidence behind each conclusion.

Integrations

How your ATS requests the interview and receives the report, with nobody retyping anything.

Contact

A 30-minute demo on a real role of yours.

Other issues
Next step

See it with a role of yours on the table.

Thirty minutes: an interview is defined from your job post, walked through the way the candidate sees it, and a report is read with its evidence.