ElevenLabs Speech-to-Text: Where Scribe Fits in a Voice Workflow
Speech recognition is the listening half of many voice systems. Here is what to test before you build on it.
- Scribe v2 covers recorded transcription; Scribe v2 Realtime targets live use.
- Current documentation lists 90+ languages.
- Test with your real acoustics and terminology.
What Scribe does
ElevenLabs’ speech-to-text product is called Scribe. The current documentation describes Scribe v2 and Scribe v2 Realtime for recorded and live transcription, with support for more than 90 languages.
Features that matter in real workflows
- Word-level timing for subtitles and editing.
- Speaker diarization for multi-person audio.
- Keyterm prompting for names and specialized vocabulary.
- Entity detection for structured extraction in API workflows.
- Realtime transcription for agents and live applications.
Recorded versus realtime
Use recorded transcription when you have a completed file and can prioritize accuracy and editing. Use realtime transcription when the transcript has to drive a live agent, captions, or another streaming application. Realtime adds latency and streaming behavior to the evaluation.
How to test
Use real audio with your microphones, accents, room noise, overlapping speakers, and domain terminology. Public benchmark claims can be useful context, but your own failure cases determine whether the system is production-ready.
Where STT connects to the rest of the stack
Transcription can feed captioning, editing, search, summaries, analytics, or an agent’s reasoning layer. That connection is why a broad audio platform can be operationally useful even when TTS is the original reason you considered it.
Sources checked
Product facts and pricing can change. These sources were checked on September 10, 2026.