# Bias in AI interviews: what can be reduced — Criterio Talent

> Which bias a structured interview reduces, which bias the model adds, why not to look at face or accent, and how to measure adverse impact without promising fairness.

URL: https://criteriotalent.com/en/recursos/sesgo-en-entrevistas-con-ia/

---

**Issue 06 · February 2026**

# Bias in AI interviews: what can be reduced and what cannot be promised

Structure reduces one concrete, measurable source of bias. The model brings its own and no instruction removes it. Telling the two apart is the work.

Judgement · Bias and fairness

Published on 29 August 2026 · 6 min read

**In short**

You can reduce variance between interviewers — different people asking different things and measuring differently — and you can keep the system from looking at signals that work as proxies for someone’s origin, age or condition. You cannot eliminate bias: the model brings its own from training, the script reflects the judgement of whoever wrote it, and the decision is still made by a person. The honest position is to reduce what is reducible and measure what is left.

Almost every AI interview platform promises, somewhere in its material, to reduce bias. Some promise to eliminate it. The second promise is false, and the first is true only if it says which bias it means, because there are several and they do not behave alike.

This issue separates the ones that can be attacked from the ones that cannot, and ends where it has to end: at what cannot be promised without lying.

## What is genuinely reduced: variance between interviewers

In a corporate process, much of the unfairness comes not from bad intent but from spread. Three people interview and ask different things, probe to different depths, and judge by criteria they never wrote down. Two equivalent candidates get different outcomes depending on who they drew, and that result is indistinguishable from bias even when nobody held a preference.

That is the bias structure attacks head-on, and it does so by construction: the same items, in the same order, assessed with the same scale written in terms of observable conduct. It requires trusting nobody, it can be checked by reading the script, and it is measurable — if two people on the team assign different levels to the same answer, the scale is not ready yet.

It also flattens order and fatigue effects: item four is the same at nine in the morning as at six in the evening, and it does not get shortened because the interviewer already has six conversations behind them.

## What the model brings with it

A language model learned from text written by people, and along with it learned the associations that text contains. Some are useful for assessing an answer and others are prejudice with good grammar. There is no way to separate them from the inside, and no instruction switches them off: asking a model to “be impartial” produces text that asserts impartiality, which is not the same thing.

What can be done is narrowing its room. A model asked for a general opinion about a candidate has all the space in the world to bring what it brings. A model asked to locate, in a transcript, the evidence relevant to a competency with described levels, quote it, and only then assign a level, has far less: the task is almost one of reading, and the result can be checked against the quoted text.

That shrinks the surface, it does not remove it. Choosing which fragment counts as “relevant” is still its call, and a preference fits in there. Which is why the quote matters so much: it is what lets a person see what it chose and disagree.

## The face, the accent and the rhythm of speech

Once there is video, the technical temptation is to analyse it: gesture, tone, pace, “confidence”. There are two reasons not to, and either is sufficient on its own.

- Validity. For the vast majority of roles, how a person looks or sounds says nothing useful about how they will work. What they say they did does, and it can be checked with a reference.
- Risk. The face, the accent and the rhythm of speech are fairly direct proxies for someone’s origin, age and condition. A system that looks at them is using those variables even if nobody intended it, and whoever operates it has no way to demonstrate that it did not.

The second point gets underrated. The question that arrives in a review is not “did you use the person’s origin to decide?”, it is “can you show that you did not?”. A system analysing audiovisual signals cannot answer no, because the protected variable sits inside the signal it did look at.

## Restricting the input is architecture, not policy

The usual way of answering all that is a document: a policy declaring that faces are not analysed. It helps little, because a policy describes an intention and has to be believed.

The version that holds is to make the component that assesses receive only the text of what was said. If the video never reaches that component, the restriction stops depending on anyone’s discipline and becomes a property of the system — something that can be shown, reviewed and checked. It is the difference between a policy that gets signed and a data input that gets inspected. [The platform page](https://criteriotalent.com/en/plataforma/) describes how it is applied here, and [the security page](https://criteriotalent.com/en/seguridad/) covers what each provider involved receives.

The video recording can still exist for other reasons — so the person can check what was recorded, so an assessment can be verified — and that is compatible with it never entering the analysis. Storing and analysing are different decisions and are better taken separately.

## How you find out whether there is adverse impact

Everything above is a design argument. The only way to know what is actually happening is to look at outcomes: compare advancement rates between groups, sustained over time and across enough processes, not over one eight-person role where any difference is noise.

And here comes the discomfort nobody mentions: you cannot measure what you do not collect. Comparing rates requires demographic information about people, which many companies do not hold and which carries its own privacy implications and its own legal framework. Deciding to measure is a real decision with a real cost; what is not honest is asserting fairness having never measured it.

When it is measured, what you are looking for is not a one-off difference but a sustained one, and the threshold at which a difference matters depends on the jurisdiction and on each company’s legal team. What an assessment system can contribute to that review is raw material: what was asked, what was answered, and which fragment supported each level — which is what lets a real pattern be told apart from a coincidence.

## The human decision is the last control, and a source of bias too

That the decision is taken by an identified person is a safeguard, and it is worth not overstating. People bring their preferences too, and a well-supported assessment can be ignored as easily as a sloppy one.

What changes is that now it shows. With cited evidence and a log, departing from what the record says leaves a trace: somebody advanced the person with the weakest profile on the competency that mattered most, and that is visible. It does not prevent it — sometimes there are good reasons to depart — but it forces the departure to be a conscious decision rather than a reflex. Reviewing those cases is, in practice, one of the most useful things to do with the log described in [the issue on traceability](https://criteriotalent.com/en/recursos/evidencia-por-turno-trazabilidad/).

## What cannot be promised

Eliminating bias cannot be promised. Nor can a model being impartial, nor a scale written by a team not reflecting that team’s judgement, nor a human decision being neutral. Any vendor promising it is selling reassurance, not a system.

What can be stood behind is more modest and considerably more useful: that every candidate for a role answered the same thing, that the assessment was built on what they said rather than on how they look, that every conclusion can be traced back to the sentence it came from, and that the decision has a name and a time on it. That does not make a process fair on its own. It makes it auditable, which is the condition for being able to argue about whether it is fair.

## Questions about this issue

### Is an AI interview more or less biased than a human panel?

It depends which panel you compare it to and how it is designed. Against an unscripted panel, structure removes a source of spread that is real and large. Against a panel trained on the same rubric, the advantage is consistency and traceability, not impartiality.

### Does removing the candidate’s name before assessment help?

It helps against one specific signal, and it is not sufficient: the text of an answer can itself carry signals of origin or generation. It reduces, like everything else on this list; it does not close the matter.

### How would an outsider audit this?

By asking for three things: the script with its versions, exactly what input the assessing component receives, and a sample of assessments with the quote supporting each. With those you can form your own judgement; without them, all that is left is belief.

**To keep reading on this site**

### [Solutions](https://criteriotalent.com/en/soluciones/)

How it is configured for high volume, technical profiles, leadership, and multi-round processes.

### [Platform](https://criteriotalent.com/en/plataforma/)

How it runs the interview, follows up, and cites the evidence behind each conclusion.

### [Integrations](https://criteriotalent.com/en/integraciones/)

How your ATS requests the interview and receives the report, with nobody retyping anything.

### [Contact](https://criteriotalent.com/en/contacto/)

A 30-minute demo on a real role of yours.

**Other issues**

**Method · Comparison and decision**

### [How to compare candidates for the same role without comparing impressions](https://criteriotalent.com/en/recursos/comparar-candidatos-sin-comparar-impresiones/)

Four conditions that have to hold before two interviews can sit side by side, and what to do with the table once they finally can.

**Method · Candidate experience**

### [How to design an AI interview the person does not experience as an insult](https://criteriotalent.com/en/recursos/experiencia-del-candidato-con-ia/)

Almost everything that makes an automated interview hateful is a design decision, not a technical limit. Seven decisions, with the uncomfortable one last.

**Method · Assessment instruments**

### [Competency matrices and rubrics: how to write levels two people read the same way](https://criteriotalent.com/en/recursos/matriz-de-competencias-y-rubricas/)

Most competency matrices are a list of nouns with a one-to-five scale. That is not an instrument, it is a form. The instrument is the descriptors.

**Next step**

## See it with a role of yours on the table.

Thirty minutes: an interview is defined from your job post, walked through the way the candidate sees it, and a report is read with its evidence.
