Speech to text

Turn spoken words into text. Speak now or upload a recording.

Speecho listens in 99 languages and writes it down: live subtitles while you talk, or an accurate transcript from a file you already have. Speaker labels, timestamps, subtitles and a summary come with it.

Start as a guest · 15 free minutes when you confirm your email · no subscription

The words appear while you are still speaking

Interim words show up dim and settle as the sentence finishes. Share a link and the room reads along on their own phones.

Spoken language Auto-detect
0:00 Stop
Listening…

so the first thing i want to

So the first thing I want to show you is the dashboard.

it pulls the numbers every

It pulls the numbers every morning, before anyone is awake.

and if something breaks you get

And if something breaks, you get a message rather than a surprise.

Three ways in, one kind of text out

Whether the words are being spoken right now or were recorded last week, what comes back is the same transcript.

Everything that comes with the transcript

The text is the start. What makes it usable is what sits around it.

Who said what

Speaker identification splits the transcript by voice, so an interview reads as an interview rather than a wall of text.

Timestamps and subtitles

Every segment carries its time. Export .srt or .vtt for video, or word-level timing for tighter editing.

A summary you can send

Key points, decisions and action items, written from the transcript rather than guessed at. Chapters too, for long recordings.

Translation into 18 languages

The transcript in another language, next to the original, from the same upload and the same balance.

Punctuation and paragraphs

Sentences end, questions get their mark and the text is broken into paragraphs, so it reads without a pass by hand.

Ask it questions

Chat with the recording: what was said about pricing, what was agreed, who owns the next step.

How accurate, honestly

Most of the industry prints 99%. On clean audio the real number is 90 to 97% of words, and we would rather say so than let you find out.

What pulls it down

  • Background noise, or a room with an echo
  • People talking over each other
  • A microphone across the table, or across the hall
  • A heavy accent or a strong regional dialect
  • Switching between two languages mid-sentence

What pulls it up

  • A microphone near the speaker, or audio taken from the desk
  • Setting the language instead of leaving it to detection
  • Adding names and jargon to your custom vocabulary
  • Splitting a very long recording into its real parts

Nobody reaches 100%, and neither do we. A transcript is a first draft you edit in the app rather than retype from scratch.

99 languages, recognised without being told

Upload and the language is detected on its own. Set it yourself when the recording is mixed or the first minute is quiet.

When to set the language yourself

Detection reads the opening of the recording. A bilingual meeting, an interpreter speaking over the original, or a long silence before the first word are the cases where a hint is worth the two seconds it takes.

Dictation, or transcription?

They sound like the same thing and they are not. If you only want to type by speaking, you may not need us at all.

Your computer already dictates

  • Free, and built into Windows, macOS, iOS, Android and Google Docs
  • Good for writing a message or a note as you speak
  • One speaker, one language, into the box you are looking at
  • Nothing to keep afterwards: no file, no timestamps, no subtitles

Speecho is for recordings and rooms

  • A meeting, an interview or a lecture that already exists as a file
  • Several speakers, separated and labelled
  • Live subtitles other people can read on their own phones
  • Subtitles, summaries, translation and a transcript that stays searchable

If your honest answer is "I just want to talk instead of typing", open your system settings first. We would rather tell you that than take the credits.

Pay only for what you use

Top up once. Credits never expire. No subscription, and the same balance also covers text to speech, scans and captions.

Confirm your email and get 15 free minutes of audio - enough to transcribe something real before you decide.

$5
2,400 credits

  • ≈ 4 hours of audio
  • 99 languages
  • Speakers, subtitles, summary
  • Translation into 18 languages
  • Live subtitles · 90-day history
Transcribe your first recording →
$25
15,000 credits
+1,800 bonus credits · ≈ 3 hours
  • ≈ 28 hours of audio
  • 99 languages
  • Speakers, subtitles, summary
  • Translation into 18 languages
  • Live subtitles · 90-day history
Transcribe your first recording →
$50
30,000 credits
+6,000 bonus credits · ≈ 10 hours
  • ≈ 60 hours of audio
  • 99 languages
  • Speakers, subtitles, summary
  • Translation into 18 languages
  • Live subtitles · 90-day history
Transcribe your first recording →

Estimates assume one credit per 6 seconds of audio, which is what transcription costs. Summaries, translation and text to speech draw on the same balance and are priced separately. The exact cost is shown before each job runs.

Frequently asked questions

Is there a free version?

You can start as a guest without an account, and confirming your email adds 15 free minutes of audio. After that you top up a balance: no monthly fee, and credits never expire.

How accurate is it?

90 to 97% of words on clean audio. Noise, crosstalk and a distant microphone lower it; setting the language and adding custom vocabulary raise it. Every transcript is editable in the app, so fixing a name takes seconds.

Which languages does it understand?

99 languages for uploaded recordings, detected automatically. Live sessions recognise 56 and can translate into another language while you speak.

Live, or after the recording?

Both, from the same account. Live subtitles appear while you talk and can be shared as a link that needs no app; uploaded files come back in minutes.

Does it add punctuation and paragraphs?

Yes. Sentences are punctuated, questions get their mark, and the text is broken into paragraphs and timed segments, so it reads without a pass by hand.

Can it tell speakers apart?

Yes. Speaker identification labels each turn, which is what makes an interview or a meeting readable. It works best when people are not talking over each other.

What happens to my audio?

It is streamed to the AI provider for inference and is not stored on our servers. Transcripts auto-delete after 90 days, and you can delete one sooner at any time.

Do I need to install anything?

No. It runs in the browser on desktop and mobile, and there is no meeting bot joining your calls: you record or share your own audio.

Say it, record it or upload it.

The text comes back the same way. The first 15 minutes are free.

Start transcribing →