# Turn-level evidence: what to keep from an interview — Criterio Talent

> Turn-by-turn transcript versus summary, the verbatim quote as the minimum unit, script versioning, the access log, and the six-month test.

URL: https://criteriotalent.com/en/recursos/evidencia-por-turno-trazabilidad/

---

**Issue 05 · January 2026**

# Turn-level evidence: what to keep from an interview so you can still defend it six months later

The uncomfortable question does not arrive on decision day: it arrives half a year later, once the team has changed. This is what has to be stored to answer it.

Method · Evidence and traceability

Published on 29 August 2026 · 6 min read

**In short**

Four things have to be kept, and all four are useless on their own: the turn-by-turn transcript with its timestamps, the verbatim quote supporting each level assigned, the exact version of the script that interview ran with, and the log of who published, opened, assessed and downloaded what. A summary substitutes for none of them.

The uncomfortable question almost never arrives on decision day. It arrives months later, once the process has closed, the recruiter who ran it has changed jobs, and somebody — the hiring team, the person who did not advance, an internal review — wants to know why what happened happened.

At that moment only what was stored exists. The team’s memory does not count any more, and loose notes do not either: nobody can tell whether that line was a quote or an interpretation. What follows is what has to be kept for that conversation to be had with facts, and why no piece can stand in for the one next to it.

## A summary is not evidence

The most common way of losing evidence is not losing it: it is summarising it. The system transcribes, assesses, writes three paragraphs with the essentials and discards the long text because it takes up space. From then on the record holds an interpretation of the conversation, not the conversation.

The difference shows up precisely when somebody disagrees. You cannot disagree with a summary: you can only believe or disbelieve whoever wrote it. You can disagree with a verbatim sentence, and something comes out of that argument — sometimes that the assessment was wrong, and sometimes that the scale was badly written, which is the better finding.

This is an architectural decision, not a policy one. A system that transcribes in order to assess and then throws the text away can produce the score once, but it can never support it again. The moment to decide it is before the first interview, not when the first question arrives.

## The minimum unit is the sentence with its address

A quote without an address is worth half. “She said she got procurement in a room the same day” beats a summary, but it still asks you to believe that this was said. With the turn and the minute, the reader can go to the recording and hear it, and the evidence stops depending on anyone.

That is why the transcript is kept as turns and not as one continuous block. A turn is the natural unit of a conversation: it has a speaker, a moment and a content, and it maps onto the item that prompted it. Storing the whole text in a single field is storing the same material in a shape that can no longer be cited.

The operating rule that follows is the one that stings at first: if a conclusion cannot cite, it is not a conclusion, it is an impression, and it comes out of the report. You lose assessments. You lose fewer than you fear, and the ones you lose were exactly the ones that were never going to survive scrutiny.

## The script is evidence too

Hardly anyone counts it among the things to keep, and it is the piece that breaks most quietly. An item gets tuned in March because it was not producing what was needed. The change is an improvement. And from that afternoon on, February’s candidates and April’s answered different questions.

If each interview remembers which version of the script it ran with, that is a fact and it can be read carefully. If it does not, comparing one group with the other is a coincidence in the shape of a table — the problem [the issue on comparison](https://criteriotalent.com/en/recursos/comparar-candidatos-sin-comparar-impresiones/) describes from the other side.

Versioning here means little more than the obvious: that editing creates a new version instead of overwriting the previous one, and that the interview records which one it pointed at. It is cheap on day one and impossible to reconstruct afterwards, because it cannot be reconstructed at all.

## Who did what, and when

The evidence of the conversation answers how somebody was assessed. It does not answer who decided, and that is usually the real question. A log of access and actions — who published the role, who opened the report, who marked advance or reject, who downloaded a recording and at what time — is what closes the other half of the record.

It has one condition to be worth anything: the decision has to be taken inside the tool. If advancing or rejecting happens in an email or a separate spreadsheet, the trail preserves in detail how someone was assessed and loses the fact of who resolved it, which is exactly backwards from what is needed.

And it has a consequence worth accepting head-on: a serious log also records whoever consults it. That is sometimes uncomfortable internally, and it is precisely the property that makes it worth having.

## The six-month test

The way to find out whether an evidence design works is to put to it the conversation it will have to survive. Somebody asks about one specific candidate half a year later. This is what has to be answerable without anyone reconstructing anything from memory:

Six questions, six different places. If any of them has nowhere to live, the structure was decorative.

| The question that arrives later | Where the answer lives | What was she asked? | The version of the script that interview ran with. | What did she answer? | The turn-by-turn transcript, timestamped, with the recording behind it. | Why that level and not another? | The verbatim quote attached to that competency’s assessment. | What was she compared against? | The other interviews for the same role, run on the same version. | Who decided? | The log, with the name and the time of the action. | How long does this exist? | The retention policy, and the record of the purge when it expires.

The last row is the one most often forgotten at design time and the first one raised in a privacy review. Keeping things forever is not the rigorous version of keeping things: it is the lazy one. What makes a record defensible is that it exists while it has a reason to exist and disappears when it stops having one.

## Keeping has an expiry date

Everything above pushes towards keeping more, and that force needs an explicit counterweight or it ends in an eternal archive of recorded conversations with people who were never even hired.

The counterweight is a written period and a purge that runs unattended, leaves a record of what it deleted, and does not depend on anyone remembering. How that fits together — and why the storage expiry has to be longer than the promised period, not shorter — is in [the issue on retention and data governance](https://criteriotalent.com/en/recursos/retencion-y-gobierno-de-datos/) and, applied to this platform, in [the security page](https://criteriotalent.com/en/seguridad/).

## What this does not solve

A complete record does not turn a bad assessment into a good one. If the scale is badly written, traceability only lets you reconstruct precisely how a wrong conclusion was reached — still better than not being able to reconstruct it, but not the same as assessing well. The quality of the instrument is decided earlier, in [the competency matrix](https://criteriotalent.com/en/recursos/matriz-de-competencias-y-rubricas/).

Nor does it turn a debatable decision into a correct one. What it does is more modest and more useful: it guarantees the argument can be had over facts, and that the answer to “why did this person advance?” is something more than the memory of whoever still works here.

## Questions about this issue

### Is keeping the full recording excessive when there is already a transcript?

The transcript is the working instrument; the recording is what lets you check that the transcript says what was said. They hold each other up, which is why they share a retention period and expire together.

### What if the person asks for their interview to be deleted before the period ends?

That right is exercised with the company that interviewed them, which is the data controller, and the system has to be able to execute it and record that it did. A purge that leaves no trace of having run cannot be told apart from a purge that never ran.

**To keep reading on this site**

### [Solutions](https://criteriotalent.com/en/soluciones/)

How it is configured for high volume, technical profiles, leadership, and multi-round processes.

### [Platform](https://criteriotalent.com/en/plataforma/)

How it runs the interview, follows up, and cites the evidence behind each conclusion.

### [Integrations](https://criteriotalent.com/en/integraciones/)

How your ATS requests the interview and receives the report, with nobody retyping anything.

### [Contact](https://criteriotalent.com/en/contacto/)

A 30-minute demo on a real role of yours.

**Other issues**

**Method · Interview design**

### [How to design a structured AI interview that leaves useful evidence and keeps the decision human](https://criteriotalent.com/en/recursos/entrevista-estructurada-con-ia/)

Eight design decisions, in order. The first is the one almost everyone skips: you do not ask a model to judge, you ask it to document.

**Method · Retention and data governance**

### [Retention of interview recordings: how long to keep them, where, and how to actually delete them](https://criteriotalent.com/en/recursos/retencion-y-gobierno-de-datos/)

The period is the easy part. The hard part is deletion that actually happens, leaves a record, and does not block itself — which is exactly what happens when storage expires before the promise does.

**Method · Comparison and decision**

### [How to compare candidates for the same role without comparing impressions](https://criteriotalent.com/en/recursos/comparar-candidatos-sin-comparar-impresiones/)

Four conditions that have to hold before two interviews can sit side by side, and what to do with the table once they finally can.

**Next step**

## See it with a role of yours on the table.

Thirty minutes: an interview is defined from your job post, walked through the way the candidate sees it, and a report is read with its evidence.
