Clean verbatim vs full verbatim: what each one removes, and which one AI actually gives you
Every transcription order form asks the same question, and most people ordering for the first time guess. “Clean verbatim or full verbatim?” sounds like a quality setting. It is not. It is a decision about what counts as the record: what the person said, or what the person meant to say. Get it wrong one way and you pay for a transcript nobody can read; get it wrong the other way and you have quietly edited a witness.
Here is the same fifteen seconds of speech at each level, what each level removes, who genuinely needs which, and what you are actually holding when an AI hands you a transcript.
One sentence, three transcripts
Someone in a planning meeting says this:
Full verbatim:
I, um, I think we - we should, you know, ship it Tuesday. [laughs] Or, uh, Wednesday. Wednesday’s fine.
Clean verbatim:
I think we should ship it Tuesday. Or Wednesday. Wednesday’s fine.
Edited (sometimes sold as “intelligent verbatim”):
I think we should ship on Tuesday or Wednesday.
The first is a recording in text form. The second is what the speaker would sign off on. The third is a paraphrase - shorter and smoother, and no longer a quote. Only the first two are transcripts.
What each level removes
The industry’s definitions are consistent enough to tabulate. Full verbatim, as Rev defines it, means “every single word spoken on your audio file is written down word for word”; clean verbatim means the file has been “lightly edited for easy readability” with no paraphrasing. Checked September 2026.
| Removed? | Full verbatim | Clean verbatim | Edited |
|---|---|---|---|
| Hesitation sounds: um, uh, er | kept | removed | removed |
| Crutch words: you know, like, sort of | kept | removed | removed |
| False starts: “we - we should” | kept | removed | removed |
| Stutters and word repeats | kept | removed | removed |
| Self-corrections: “Tuesday. Or Wednesday” | kept | kept | merged |
| Other speakers’ “yeah”, “uh-huh” | kept | usually removed | removed |
| Non-speech: [laughs], [cough], [pause] | kept, tagged | removed | removed |
| Slang and grammar | as spoken | as spoken | corrected |
| Sentence structure | as spoken | as spoken | rewritten |
Two rows decide most arguments. Self-corrections stay in clean verbatim: “Or Wednesday. Wednesday’s fine” is the speaker changing their mind on the record, and cutting it changes the meaning. And grammar stays as spoken: clean verbatim removes noise, it does not fix English. The moment a transcript starts fixing grammar it has crossed into editing, and it can no longer be quoted with quotation marks around it.
Who needs full verbatim, and what it costs
Full verbatim is for anyone whose analysis depends on how something was said, not just what.
- Legal proceedings. A deposition transcript that drops a witness’s “uh, I - I don’t recall” has removed the hesitation a lawyer will later argue about. Court reporters write it all down for exactly that reason.
- Qualitative research that codes speech itself. Conversation analysis, discourse analysis, some psychology and linguistics work: the pause and the false start are the data.
- Oral history. An archive is a source. Cleaning it is the same as retouching a photograph in an evidence file.
The cost is real and it has been measured. A 2022 study in Emerging Themes in Epidemiology, drawing on two community health trials in Ghana, found that verbatim transcription took 6 to 10 hours for each hour of interview, against 1 to 2 hours for expanded interviewer notes. Their conclusion was not “verbatim is always necessary”. It was that for relatively simple applied questions, well-supervised notes were enough - and that for focus groups, where the interaction itself carried the meaning, notes “failed to capture the interaction and richness” and verbatim was still required. That is the honest shape of the answer: verbatim where the texture is the finding, not everywhere.
Who needs clean verbatim
Almost everyone else, and this is the default most services quote.
- Podcast transcripts and show notes. Nobody reads a podcast page for the ums.
- Interviews for articles. You will quote from this, and clean verbatim is what a quote is supposed to be: the speaker’s words, minus the noise the speaker would remove themselves. (What you are allowed to cut inside a quote is a separate question with its own rules, covered in what you can cut from a quote.)
- Meeting notes, lectures, study material. The point is the content, and every filler is a speed bump.
- Subtitles. Full verbatim captions fail the reading-speed rule twice over: the fillers eat characters, and the false starts become their own micro-cues, which is one of the failures in our subtitle timing rules.
What an AI transcript actually is
This is the part order forms never tell you, so here it is plainly.
Raw AI output is neither. Speech models are trained largely on captioned and subtitled audio, and captions leave the ums out. So the models learned to leave them out too. A raw transcript from any modern engine will keep most of the words and drop an unpredictable share of the hesitations, stutters and repeats - more than full verbatim allows, fewer than clean verbatim requires. If a legal or research standard says “full verbatim”, an unreviewed AI transcript does not meet it, whatever the marketing page says. Someone has to check it against the audio.
Clean verbatim is where AI is genuinely good, with one condition. Speecho does it in two places, and they work differently on purpose.
The filler word counter removes fillers by list: a known set of hesitation sounds and crutch words for each of 12 languages, plus any you add, plus immediate word repeats. It is free, it runs in your browser, and it highlights every single match before it removes anything, so you can see that “like” in “I like the plan” is about to go and put it back. It does not touch false starts or self-corrections, because a word list cannot tell “we - we should” from a deliberate repetition.
Inside Speecho, the Clean text (no filler words) notes style does the fuller job with a language model: fillers, false starts, stutters and immediate self-corrections (keeping the corrected version), with punctuation repaired where the removals require it. It works paragraph by paragraph and is instructed not to summarise, shorten, reorder or rephrase anything else. That instruction is the whole difference between clean verbatim and editing, and it is why the output should still be read once before it is published: a model following a rule is not a person checking a quote.
Neither should ever produce the third column. If a tool hands you back shorter, smoother sentences than the speaker said, it has edited, not cleaned, and you cannot put quotation marks around the result.
The rule of thumb
Ask one question: will anyone ever need to argue about how this was said? If yes - a court, a review board, a future historian - keep full verbatim, treat the AI transcript as a draft, and have it checked against the audio. If no, clean verbatim is the right record, and it is the one AI produces well: transcribe, run the clean pass, read it once, publish.
Either way, the raw transcript is the one to keep. You can always clean a verbatim file. You cannot un-clean a clean one.