VocalLab Speech to Text turns spoken audio into accurate text in seconds — across 25+ languages.
Transcribe audio
- Open Speech to Text in the Studio sidebar, under Create.
- Drag in an audio file, or click Record audio to use your microphone.
- Click Transcribe — your text appears within seconds.
Supported formats: MP3, WAV, M4A, OGG, FLAC and WebM, up to 16 MB (~18 minutes).
After transcribing
- Copy the transcript or download it as a
.txtfile. - Click Use in Text to Speech to send it straight to the TTS editor.
- Every transcription is saved to your History (the Speech to Text tab).
Languages
The language is detected automatically across 25+ languages — including English, Spanish, French, German, Italian, Portuguese, Russian, Arabic, Hindi, Chinese, Japanese and Korean. You don't need to pick one.
For the best accuracy, use a clean recording with minimal background noise and one speaker at a time.
Cost
Transcription costs points based on audio length — about 15 points per minute. You only pay for the audio you transcribe.


