Interview transcription for journalists: a workflow that survives fact-checking
The oldest ratio in journalism school still holds: an hour of recorded interview costs three to four hours to type up by hand. On a profile with three sources, that’s a full working day gone before you’ve written a word — and the piece is due Thursday.
AI transcription collapses that day into minutes, and at $0.83–$1.25 per audio hour it costs less than the coffee you’d drink while typing. But journalism has a constraint most industries don’t: a quote with one wrong word is not 97% correct, it’s wrong, and it’s wrong in print with someone’s name attached to it. So the interesting question isn’t whether to use AI transcription — it’s how to use it so nothing unverified reaches the page.
The three ways, honestly compared
Typing it yourself costs nothing and forces a close re-listen — some reporters genuinely value that pass. It’s also 3–4 hours per interview hour, which stops scaling the week you have five interviews.
Human transcription services deliver 99%+ accuracy for legal-grade needs, at roughly $1–2 per audio minute and one to several days of turnaround. For a $60–$120 invoice per interview, you get precision most news work doesn’t need — because you’re going to verify the quotes yourself anyway.
AI transcription lands in minutes at 90–97% accuracy on clean audio. The catch is inside that number: a one-hour interview runs 7,000–9,000 words, so even 96% accuracy means a few hundred small errors — dropped words, a mangled name, “can” heard as “can’t”. Fine for a searchable draft. Unpublishable without checking.
That last sentence is the whole method.
A workflow that survives the fact-checker
1. Record like the transcript depends on it — it does. Accuracy is decided at recording time, not transcription time. Phone on the table plus a backup recorder, close to the source, café hum behind you if you can manage it. A crisp recording transcribes at 96–97%; a noisy phone line can fall below 90%, and every lost point is minutes of your evening back.
2. Upload and let speakers separate. Turn on speaker identification for anything with more than one voice. “Speaker 1 / Speaker 2” becomes “Me / Minister” with two renames, and suddenly the transcript reads like a script instead of a wall.
3. Work from the transcript, verify from the audio. Search the transcript to find the moment; then play the audio at that timestamp before a single quoted word goes in the piece. This is the rule that makes AI transcription safe for journalism: the transcript locates the quote, the recording confirms it. Timestamped paragraphs make the round-trip a few seconds per quote — far faster than scrubbing a raw recording, and exactly as rigorous.
4. Keep the tape until publication — then decide what survives. Your transcript is your searchable archive; the audio is your defence if a quote is challenged. Keep both until the piece is out and settled, then apply your own retention judgement — more on that below.
The source-protection questions to ask any tool
Journalists upload conversations other people would like to read. Before any tool touches interview audio, get clear answers to three questions:
- Is the audio used to train AI models? The answer must be no. (Speecho’s is no.)
- What is stored, and for how long? With Speecho, source files are deleted after processing and transcripts auto-delete after 90 days; you can delete anything earlier, and should, once a sensitive piece has shipped.
- What leaves your machine? For video interviews, Speecho extracts the audio in your browser — the footage never uploads at all.
And one rule no tool can apply for you: if part of a conversation was truly off the record, cut it before uploading. Trim the file and transcribe the rest. The strongest protection for something that must never appear in text is for it never to enter the pipeline.
Interviews in other languages
Whisper-based transcription handles 99 languages, which changes fieldwork: interview your source in the language they actually think in, transcribe in that language, then translate the transcript into yours for drafting — quoting from the original when precision matters. Speecho translates finished transcripts into 18 languages, so the original-language transcript and your working-language copy come from the same upload.
What it costs a working freelancer
At $0.83–$1.25 per audio hour, a heavy month — say twelve hours of interviews — costs about $10–$15, prepaid credits, no subscription ticking in the months between commissions. The same month at a human service is $700–$1,400; typed by hand, it’s a week of unpaid evenings. That arithmetic is why the verify-against-audio workflow has quietly become the norm on newsdesks: the machine does the typing, the journalist does the journalism.
Try it on your next interview: upload the file, get 15 free minutes on email confirmation — no card, and no bot anywhere near your sources.