Speech to text is live today, and it does one thing: turn a recording into a plain transcript and a duration. It does not produce captions, subtitles, SRT or VTT files, timestamps, word-level timing, or speaker labels — just text.
If you came here expecting subtitles, this isn't that tool. This post explains exactly what you do get, why that's genuinely useful on its own, and where the numbers land: file limits, supported formats, and how it's billed.
What You Actually Get
Upload an audio or video file and you get back two things: a plain-text transcript of everything said, and the file's duration. That's the whole output. There are no timestamps next to individual words, no speaker A / speaker B labels, and no formatting for a caption track — just the words, as one continuous piece of text.
No captions: This does not produce SRT, VTT, timestamps, word-level timing, or speaker labels. If your goal is subtitles for a video, this tool doesn't make them — it makes a transcript you can read, search, or paste somewhere else.
File Limits and Formats
Every plan can transcribe, but the ceiling on a single file depends on your tier.
| Plan | Max file length |
|---|---|
| Free | 20 minutes |
| Lite | 1 hour |
| Pro | 3 hours |
| Max | 6 hours |
Anything past roughly 18 minutes runs as an asynchronous job rather than a spinner you sit and watch. You upload the file, leave the page, and come back once it's finished. Accepted formats cover most of what people actually have sitting around: mp3, wav, m4a, ogg, flac, webm, mp4, mov, mkv and avi — audio or video, either works.
How It Works, Start to Finish
Upload your file
Drop in an audio or video file in a supported format — mp3, wav, m4a, ogg, flac, webm, mp4, mov, mkv or avi.
Wait if it's long
Files past roughly 18 minutes process as an async job. Shorter files return in far less time.
Read the transcript
You get back plain text and the file's duration — copy it, search it, or hand it to a voice for the next step.
Language
30 languages are recognized. The language picker is a hint, not a requirement. Leave it on auto-detect and the transcript engine identifies the language itself, which is the right fallback for anything you're not sure about.
Before You Upload
- Check your plan's file-length ceiling — 20 minutes on Free, up to 6 hours on Max.
- Use a supported format — mp3, wav, m4a, ogg, flac, webm, mp4, mov, mkv or avi.
- Leave the language on auto-detect if you're not certain, or set it as a hint if you are.
How It's Billed
Transcription is priced by duration, not by characters, because there's no text going in to count. It costs 15 points per minute of audio: a 20-minute file costs 300 points, a full hour costs 900. That's a different rate from generation, which is priced from the script you paste, and it's worth keeping the two separate in your head.
Turn a recording into a transcript
Upload audio or video and get back plain text and a duration — nothing more, nothing less.
Try speech to textAbout the SRT Posts on This Site
If you've read anywhere on this site about SRT export or word-level subtitles, that's real, and it's worth being precise about which direction it runs. Those posts, including SRT export from text to speech, describe generating audio from a script: you supply the text, the engine returns word timing alongside the audio, and an SRT file is built from that timing.
Speech to text runs the opposite way. You supply a recording and get a transcript back — no timing, no caption file, nothing built from it. Captions come out of text going in, not out of audio going in.
In short: Text in, captions possible. Audio in, transcript only. The two features share a voice library but not a direction.
What You Can Do With a Transcript
Search a recording
Find the moment someone said a specific phrase without scrubbing through the whole file.
Repurpose spoken content
Turn a meeting, interview or podcast episode into a written draft you can edit and publish.
Build captions yourself
Use the transcript as a starting point and add timing manually in a caption editor, since none is generated here.
Where This Fits
Speech to text
Upload a recording and get a plain transcript back.
Audio to text
Turn any spoken audio file into readable, searchable text.
Long recordings
Transcribe files that run past 18 minutes as an async job.
Meeting recordings
Turn a recorded meeting into a transcript you can search and share.
Podcast episodes
Get a text version of an episode for show notes or a written recap.
Interview recordings
Transcribe an interview so you can quote it accurately without re-listening.
Video files
Upload a video file directly — audio and video formats both work.
Once You Have a Transcript, These Voices Read It Back
Speech to text and text to speech run in opposite directions, but they connect in one useful way: once you've got a clean transcript, you can hand it back to a voice and generate narration from it — a recap, a spoken version of your own notes, a written interview turned back into audio. Here are six voices worth auditioning for that second step. Browse the full voice library for the rest.
| Voice | Accent | Tone | Rating |
|---|---|---|---|
| Articulate Male Explainer Voice | Neutral American | Articulate, precise | ★★★★★ |
| Lucid Male Explainer Voice | Neutral American | Clear, measured | ★★★★★ |
| Professional Clear American Male | Neutral American | Professional, clear | ★★★★★ |
| Warm Natural Female Explainer | Neutral American | Warm, natural | ★★★★★ |
| Polished Approachable British Female | British | Polished, approachable | ★★★★★ |
| Curious Female Explainer Voice | Neutral American | Curious, fast-paced | ★★★★☆ |
Frequently Asked Questions
Does speech to text generate captions or subtitles?▾
No. It returns a plain transcript and a duration — no SRT, VTT, timestamps, word timing, or speaker labels.
What's the longest file I can transcribe?▾
It depends on your plan: 20 minutes on Free, 1 hour on Lite, 3 hours on Pro, and 6 hours on Max, per file.
What happens with longer files?▾
Files past roughly 18 minutes run as an asynchronous job. You upload the file, leave the page, and come back once the transcript is ready.
What file types are supported?▾
mp3, wav, m4a, ogg, flac, webm, mp4, mov, mkv and avi — audio or video files both work.
How is transcription billed?▾
15 points per minute of audio, billed by duration rather than character count, since there's no script to count characters from.
Is this the same as the SRT export mentioned elsewhere on this site?▾
No, and the direction matters. SRT export comes from text to speech: you supply the text and get audio with timing back. Speech to text runs the other way and returns a transcript only.
Turn any recording into a transcript
Upload audio or video and get plain text back, no captions attached.









