Speech to Text
Upload audio and get back a plain text transcript with the duration — no captions or SRT. Auto language detection; free up to 20 minutes per file.
Step-by-step guide
How to Use Speech to Text
Start for free — no credit card, no software to install.
Upload your audio
Add an MP3, WAV or other common audio file from your device — no recording tool built in, just a file you already have.
Let language auto-detect, or set it
Leave language on auto-detect for most recordings, or choose the language yourself when a file mixes accents or technical terms.
Get transcript and duration back
Once the job finishes you receive plain text and the clip's duration — copy it out or download it, nothing more, nothing timed.
Why creators choose this tool
Speech to Text — Features & Capabilities
Everything you need to create professional voiceovers — in one tool.
Plain text output
You get the words as text and the file's duration — nothing timed, no subtitle format.
Automatic language detection
The engine identifies the spoken language on its own, with an optional manual override.
Duration limits by plan
Free files run up to 20 minutes; Lite 1 hour, Pro 3 hours, Max 6 hours per upload.
Asynchronous processing for long files
Submit a long recording and come back — it processes as a background job instead of holding a page open.
Use cases
What Can You Create with Speech to Text?
From TikTok to podcasts — here's where creators use this tool most.
Turning a recording into readable notes
Convert a voice memo, lecture or call into a text file you can skim, search and paste into a document.
Building a searchable archive of audio
Run a backlog of recordings through the tool so their content shows up in a text search instead of staying locked inside audio files.
Drafting article text from spoken content
Start from a rough transcript and edit it down into publishable copy, rather than typing up a recording from scratch.
Checking what was actually said
Get an exact text record of a recorded conversation when memory or shorthand notes aren't reliable enough.
vs Standard TTS
Vocallab vs Standard Text-to-Speech
See why creators switch from generic TTS tools to Vocallab.
About this tool
Everything you need to know about Speech to Text
Speech to text on VocalLab takes an audio file and gives back exactly two things: a plain text transcript and the file's duration. There are no captions, no SRT or VTT files, and no word-by-word timestamps — if you need the words, not a timed subtitle track, this is built for that. Language is detected automatically, or you can point it at a specific language if the recording mixes accents or terms. How long a single file can run depends on your plan: 20 minutes free, 1 hour on Lite, 3 hours on Pro, and up to 6 hours on Max. Longer files are queued as a background job rather than making you wait on a spinner.
People also search for
FAQ
Speech to Text — FAQ
Common questions users ask about the Speech to Text voice. If you need additional help, please contact us via the contact form or email us at support@vocallab.ai.
Does speech to text produce captions or subtitles?
No — the output is a plain text transcript and a duration, with no timing data, so it can't be used directly as an SRT or VTT file.
How long can a file be?
It depends on your plan: 20 minutes on Free, 1 hour on Lite, 3 hours on Pro and 6 hours on Max, per uploaded file.
Does it label who is speaking?
No, there are no speaker labels — a recording with several voices comes back as one continuous block of text.
What happens with a very long recording?
It runs as a background job rather than a live wait, so a multi-hour file can take a few minutes to finish processing after you submit it.
Do I need to set the language manually?
Usually not — language is detected automatically, though you can specify one yourself if a recording mixes languages or has technical vocabulary.
More in this collection
Related AI Voice Tools
Free Transcription Tool
Transcribe up to 20 minutes of audio free, no card needed — a plain text transcript and duration, the same output as every paid plan.
TranscriptionPodcast Transcription
Turn a recorded episode into a plain text transcript for show notes and quotes — fits Lite's 1-hour limit, no speaker labels included.
TranscriptionInterview Transcription
Get a verbatim transcript of a recorded interview — but know first: no speaker labels, so both voices return as one continuous text block.
TranscriptionAudio to Text
Upload an existing audio file and get a plain text transcript with its duration back — no timestamps, no caption file, just the words.
TranscriptionMeeting Transcription
Upload a recorded meeting or standup and get a plain text transcript with duration back — no Zoom integration, no speaker labels.
TranscriptionVideo to Text
Pull the audio from your video first, then upload it here for a plain text transcript and duration — no captions or subtitle file generated.


