ElevenLabs Scribe icon
audio-speech

ElevenLabs Scribe

Fast batch speech-to-text with word-level metadata, strongest on jargon and weaker on overlap/code-switching.

Visit ElevenLabs Scribe
Word timestampsSpeaker labels3 scored runsFast batch
TL;DR — our verdictUpdated September 2026 · 11 test artifacts

Strong default batch STT API

Where it wins
  • You need batch speech-to-text from a single multipart POST.
  • You need word-level timestamps, confidence values, and speaker labels in the transcript payload.
  • You need strong accuracy on jargon-heavy English narration.
Main limitation
  • You need validated speaker attribution correctness on overlapping conversations.
Pricing (verified plans)
Free / Pay-as-you-go $0.22/hrStarter $6/monthCreator $22/monthPro $99/month
Strongest test artifacts

Our take

Tested here as `scribe_v1`, ElevenLabs Scribe looks like a strong batch speech-to-text API for developers: every scored run returned word-level timestamps, confidence values, and speaker labels, and the medical-jargon sample came back at 3.01% WER. The tradeoff is clear on harder audio: crosstalk produced heavy insertions and over-segmentation, and the bilingual run lost Spanish tokens, so it is best when you need a fast, metadata-rich transcript and can tolerate weaker overlap and code-switching.

Tutorial recording of the ElevenLabs Scribe benchmark workflow and result review.

In-Depth Review

Our detailed analysis of ElevenLabs Scribe — features, performance, and real-world testing.

AD
AI Demos Team
Expert Reviewer
Verified Review

Feature-by-Feature Breakdown

Batch Speech-to-Text Transcription
Strong on jargon, weak on overlap/code-switching
Test Summary
Feature tested: Batch Speech-to-Text Transcription
Result: Partial — Strong on jargon, weak on overlap/code-switching

Feature tested: Batch Speech-to-Text Transcription

Result: Partial

Verdict: Strong on jargon, weak on overlap/code-switching

Expected behavior: ElevenLabs Scribe performs one-shot batch transcription over long audio via a single multipart POST. The benchmark exercised it on jargon-heavy narration, crosstalk, and bilingual code-switching, showing the core transcript engine works across these audio inputs but with varying quality.

Test case: Audio file → Image

Input type: Audio file

Input used: Input artifact (Audio file): INPUT - Overlapping speech / crosstalk audio (crosstalk.wav), 65.39 MB, 2,142.709 s; batch multipart speech-to-text with diarization and word timestamps enabled. — crosstalk.wav

Observed output: Output artifact (Image): On the crosstalk sample, ElevenLabs Scribe scored 26.67% WER with 856 substitutions, 782 deletions, and 383 insertions. The highlighted divergence shows a large missing span and the run over-segmented speakers. — 04-transcript-detail-input-1.png

Input artifact: Input artifact (Audio file): INPUT - Overlapping speech / crosstalk audio (crosstalk.wav), 65.39 MB, 2,142.709 s; batch multipart speech-to-text with diarization and word timestamps enabled. — crosstalk.wav

Output artifact: Output artifact (Image): On the crosstalk sample, ElevenLabs Scribe scored 26.67% WER with 856 substitutions, 782 deletions, and 383 insertions. The highlighted divergence shows a large missing span and the run over-segmented speakers. — 04-transcript-detail-input-1.png

What changed: Audio file transformed into Image

Test case: Audio file → Image

Input type: Audio file

Input used: Input artifact (Audio file): INPUT - Medical jargon narration (medical_terms.mp3), 8.58 MB, 1,123.944 s; batch multipart speech-to-text with diarization and word timestamps enabled. — medical_terms.mp3

Observed output: Output artifact (Image): On the medical-jargon sample, ElevenLabs Scribe scored 3.01% WER and stayed close to the reference, with only a narrow substitution span in the highlighted diff. — 04-transcript-detail-input-2.png

Input artifact: Input artifact (Audio file): INPUT - Medical jargon narration (medical_terms.mp3), 8.58 MB, 1,123.944 s; batch multipart speech-to-text with diarization and word timestamps enabled. — medical_terms.mp3

Output artifact: Output artifact (Image): On the medical-jargon sample, ElevenLabs Scribe scored 3.01% WER and stayed close to the reference, with only a narrow substitution span in the highlighted diff. — 04-transcript-detail-input-2.png

What changed: Audio file transformed into Image

Test case: Audio file → Image

Input type: Audio file

Input used: Input artifact (Audio file): INPUT - Bilingual code-switching audio (mix_language.mp3), 22.19 MB, 1,938.495 s; batch multipart speech-to-text with diarization and word timestamps enabled. — mix_language.mp3

Observed output: Output artifact (Image): On the bilingual sample, ElevenLabs Scribe scored 28.57% WER; the highlighted error site drops the Spanish token 'ahora' and the run leaves Spanish recall at 57.5%. — 04-transcript-detail-input-3.png

Input artifact: Input artifact (Audio file): INPUT - Bilingual code-switching audio (mix_language.mp3), 22.19 MB, 1,938.495 s; batch multipart speech-to-text with diarization and word timestamps enabled. — mix_language.mp3

Output artifact: Output artifact (Image): On the bilingual sample, ElevenLabs Scribe scored 28.57% WER; the highlighted error site drops the Spanish token 'ahora' and the run leaves Spanish recall at 57.5%. — 04-transcript-detail-input-3.png

What changed: Audio file transformed into Image

Why it matters / Conclusion: The core transcript engine is very input-sensitive: it can be excellent on technical English, but hard conversational audio and code-switching both degrade sharply.

ElevenLabs Scribe performs one-shot batch transcription over long audio via a single multipart POST. The benchmark exercised it on jargon-heavy narration, crosstalk, and bilingual code-switching, showing the core transcript engine works across these audio inputs but with varying quality.

audio
0:00 / 0:00
Loading audio...
INPUT - Overlapping speech / crosstalk audio (crosstalk.wav), 65.39 MB, 2,142.709 s; batch multipart speech-to-text with diarization and word timestamps enabled.
image
Output artifact for "Batch Speech-to-Text Transcription" test: On the crosstalk sample, ElevenLabs Scribe scored 26.67% WER with 856 substitutions, 782 deletions, and 383 insertions. The highlighted divergence shows a large missing span and the run over-segmented speakers., 04-transcript-detail-input-1.png
On the crosstalk sample, ElevenLabs Scribe scored 26.67% WER with 856 substitutions, 782 deletions, and 383 insertions. The highlighted divergence shows a large missing span and the run over-segmented speakers.
audio
0:00 / 0:00
Loading audio...
INPUT - Medical jargon narration (medical_terms.mp3), 8.58 MB, 1,123.944 s; batch multipart speech-to-text with diarization and word timestamps enabled.
image
Output artifact for "Batch Speech-to-Text Transcription" test: On the medical-jargon sample, ElevenLabs Scribe scored 3.01% WER and stayed close to the reference, with only a narrow substitution span in the highlighted diff., 04-transcript-detail-input-2.png
On the medical-jargon sample, ElevenLabs Scribe scored 3.01% WER and stayed close to the reference, with only a narrow substitution span in the highlighted diff.
audio
0:00 / 0:00
Loading audio...
INPUT - Bilingual code-switching audio (mix_language.mp3), 22.19 MB, 1,938.495 s; batch multipart speech-to-text with diarization and word timestamps enabled.
image
Output artifact for "Batch Speech-to-Text Transcription" test: On the bilingual sample, ElevenLabs Scribe scored 28.57% WER; the highlighted error site drops the Spanish token 'ahora' and the run leaves Spanish recall at 57.5%., 04-transcript-detail-input-3.png
On the bilingual sample, ElevenLabs Scribe scored 28.57% WER; the highlighted error site drops the Spanish token 'ahora' and the run leaves Spanish recall at 57.5%.
Bottom Line
The core transcript engine is very input-sensitive: it can be excellent on technical English, but hard conversational audio and code-switching both degrade sharply.
From our researchTranscribe Audio Accurately — Speech-to-Text Engine Benchmarkearlier research
Transcript Metadata Output
Consistent developer payload
Test Summary
Feature tested: Transcript Metadata Output
Result: Passed — Consistent developer payload

Feature tested: Transcript Metadata Output

Result: Passed

Verdict: Consistent developer payload

Expected behavior: The API returns structured transcript payloads with developer-facing metadata such as word-level timestamps, confidence/logprob values, speaker IDs, and stable JSON shape. The evidence came from transcript outputs used for captioning, search, and speaker-aware downstream workflows.

Test case: Audio file → Image

Input type: Audio file

Input used: Input artifact (Audio file): INPUT - Overlapping speech / crosstalk audio (crosstalk.wav), 65.39 MB, 2,142.709 s. — crosstalk.wav

Observed output: Output artifact (Image): Raw response preview showing `word_timestamps yes`, `confidence yes`, `speaker_labels yes`, 14,506 timed tokens, and 5 distinct speakers. — 02-response-raw-input-1.png

Input artifact: Input artifact (Audio file): INPUT - Overlapping speech / crosstalk audio (crosstalk.wav), 65.39 MB, 2,142.709 s. — crosstalk.wav

Output artifact: Output artifact (Image): Raw response preview showing `word_timestamps yes`, `confidence yes`, `speaker_labels yes`, 14,506 timed tokens, and 5 distinct speakers. — 02-response-raw-input-1.png

What changed: Audio file transformed into Image

Test case: Audio file → Image

Input type: Audio file

Input used: Input artifact (Audio file): INPUT - Medical jargon narration (medical_terms.mp3), 8.58 MB, 1,123.944 s. — medical_terms.mp3

Observed output: Output artifact (Image): Raw response preview showing `word_timestamps yes`, `confidence yes`, `speaker_labels yes`, 5,448 timed tokens, and 1 distinct speaker. — 02-response-raw-input-2.png

Input artifact: Input artifact (Audio file): INPUT - Medical jargon narration (medical_terms.mp3), 8.58 MB, 1,123.944 s. — medical_terms.mp3

Output artifact: Output artifact (Image): Raw response preview showing `word_timestamps yes`, `confidence yes`, `speaker_labels yes`, 5,448 timed tokens, and 1 distinct speaker. — 02-response-raw-input-2.png

What changed: Audio file transformed into Image

Test case: Audio file → Image

Input type: Audio file

Input used: Input artifact (Audio file): INPUT - Bilingual code-switching audio (mix_language.mp3), 22.19 MB, 1,938.495 s. — mix_language.mp3

Observed output: Output artifact (Image): Raw response preview showing `word_timestamps yes`, `confidence yes`, `speaker_labels yes`, 12,162 timed tokens, and 3 distinct speakers. — 02-response-raw-input-3.png

Input artifact: Input artifact (Audio file): INPUT - Bilingual code-switching audio (mix_language.mp3), 22.19 MB, 1,938.495 s. — mix_language.mp3

Output artifact: Output artifact (Image): Raw response preview showing `word_timestamps yes`, `confidence yes`, `speaker_labels yes`, 12,162 timed tokens, and 3 distinct speakers. — 02-response-raw-input-3.png

What changed: Audio file transformed into Image

Why it matters / Conclusion: The payload shape is stable across runs and includes the metadata a downstream developer would expect from a production STT API.

The API returns structured transcript payloads with developer-facing metadata such as word-level timestamps, confidence/logprob values, speaker IDs, and stable JSON shape. The evidence came from transcript outputs used for captioning, search, and speaker-aware downstream workflows.

audio
0:00 / 0:00
Loading audio...
INPUT - Overlapping speech / crosstalk audio (crosstalk.wav), 65.39 MB, 2,142.709 s.
image
Output artifact for "Transcript Metadata Output" test: Raw response preview showing `word_timestamps yes`, `confidence yes`, `speaker_labels yes`, 14,506 timed tokens, and 5 distinct speakers., 02-response-raw-input-1.png
Raw response preview showing `word_timestamps yes`, `confidence yes`, `speaker_labels yes`, 14,506 timed tokens, and 5 distinct speakers.
audio
0:00 / 0:00
Loading audio...
INPUT - Medical jargon narration (medical_terms.mp3), 8.58 MB, 1,123.944 s.
image
Output artifact for "Transcript Metadata Output" test: Raw response preview showing `word_timestamps yes`, `confidence yes`, `speaker_labels yes`, 5,448 timed tokens, and 1 distinct speaker., 02-response-raw-input-2.png
Raw response preview showing `word_timestamps yes`, `confidence yes`, `speaker_labels yes`, 5,448 timed tokens, and 1 distinct speaker.
audio
0:00 / 0:00
Loading audio...
INPUT - Bilingual code-switching audio (mix_language.mp3), 22.19 MB, 1,938.495 s.
image
Output artifact for "Transcript Metadata Output" test: Raw response preview showing `word_timestamps yes`, `confidence yes`, `speaker_labels yes`, 12,162 timed tokens, and 3 distinct speakers., 02-response-raw-input-3.png
Raw response preview showing `word_timestamps yes`, `confidence yes`, `speaker_labels yes`, 12,162 timed tokens, and 3 distinct speakers.
Bottom Line
The payload shape is stable across runs and includes the metadata a downstream developer would expect from a production STT API.
From our researchTranscribe Audio Accurately — Speech-to-Text Engine Benchmarkearlier research
Speaker Diarization
Present, but attribution correctness unvalidated
Test Summary
Feature tested: Speaker Diarization
Result: Partial — Present, but attribution correctness unvalidated

Feature tested: Speaker Diarization

Result: Partial

Verdict: Present, but attribution correctness unvalidated

Expected behavior: The engine emits speaker labels and speaker counts in the transcript output, enabling rough diarization and inspection of how many speakers appear in the audio. The benchmark only verified label presence and count, not correct attribution.

Test case: Audio file → Image

Input type: Audio file

Input used: Input artifact (Audio file): INPUT - Overlapping speech / crosstalk audio (crosstalk.wav), 65.39 MB, 2,142.709 s; 4-way meeting overlap. — crosstalk.wav

Observed output: Output artifact (Image): Run metrics report `diarization detected` and `distinct_speaker_labels 5` on the crosstalk sample, while the reference meeting had 4 participants. — 03-terminal-metrics-input-1.png

Input artifact: Input artifact (Audio file): INPUT - Overlapping speech / crosstalk audio (crosstalk.wav), 65.39 MB, 2,142.709 s; 4-way meeting overlap. — crosstalk.wav

Output artifact: Output artifact (Image): Run metrics report `diarization detected` and `distinct_speaker_labels 5` on the crosstalk sample, while the reference meeting had 4 participants. — 03-terminal-metrics-input-1.png

What changed: Audio file transformed into Image

Test case: Audio file → Image

Input type: Audio file

Input used: Input artifact (Audio file): INPUT - Medical jargon narration (medical_terms.mp3), 8.58 MB, 1,123.944 s; single-speaker narration. — medical_terms.mp3

Observed output: Output artifact (Image): Run metrics report `diarization detected` and `distinct_speaker_labels 1` on the medical-jargon sample, matching the single-speaker setup at the label-count level. — 03-terminal-metrics-input-2.png

Input artifact: Input artifact (Audio file): INPUT - Medical jargon narration (medical_terms.mp3), 8.58 MB, 1,123.944 s; single-speaker narration. — medical_terms.mp3

Output artifact: Output artifact (Image): Run metrics report `diarization detected` and `distinct_speaker_labels 1` on the medical-jargon sample, matching the single-speaker setup at the label-count level. — 03-terminal-metrics-input-2.png

What changed: Audio file transformed into Image

Test case: Audio file → Image

Input type: Audio file

Input used: Input artifact (Audio file): INPUT - Bilingual code-switching audio (mix_language.mp3), 22.19 MB, 1,938.495 s; spontaneous two-language conversation. — mix_language.mp3

Observed output: Output artifact (Image): Run metrics report `diarization detected` and `distinct_speaker_labels 3` on the bilingual sample, showing that speaker labels are returned even in mixed-language audio. — 03-terminal-metrics-input-3.png

Input artifact: Input artifact (Audio file): INPUT - Bilingual code-switching audio (mix_language.mp3), 22.19 MB, 1,938.495 s; spontaneous two-language conversation. — mix_language.mp3

Output artifact: Output artifact (Image): Run metrics report `diarization detected` and `distinct_speaker_labels 3` on the bilingual sample, showing that speaker labels are returned even in mixed-language audio. — 03-terminal-metrics-input-3.png

What changed: Audio file transformed into Image

Why it matters / Conclusion: Speaker labels are there, but the benchmark only proves label presence and count, not correct attribution.

The engine emits speaker labels and speaker counts in the transcript output, enabling rough diarization and inspection of how many speakers appear in the audio. The benchmark only verified label presence and count, not correct attribution.

audio
0:00 / 0:00
Loading audio...
INPUT - Overlapping speech / crosstalk audio (crosstalk.wav), 65.39 MB, 2,142.709 s; 4-way meeting overlap.
image
Output artifact for "Speaker Diarization" test: Run metrics report `diarization detected` and `distinct_speaker_labels 5` on the crosstalk sample, while the reference meeting had 4 participants., 03-terminal-metrics-input-1.png
Run metrics report `diarization detected` and `distinct_speaker_labels 5` on the crosstalk sample, while the reference meeting had 4 participants.
audio
0:00 / 0:00
Loading audio...
INPUT - Medical jargon narration (medical_terms.mp3), 8.58 MB, 1,123.944 s; single-speaker narration.
image
Output artifact for "Speaker Diarization" test: Run metrics report `diarization detected` and `distinct_speaker_labels 1` on the medical-jargon sample, matching the single-speaker setup at the label-count level., 03-terminal-metrics-input-2.png
Run metrics report `diarization detected` and `distinct_speaker_labels 1` on the medical-jargon sample, matching the single-speaker setup at the label-count level.
audio
0:00 / 0:00
Loading audio...
INPUT - Bilingual code-switching audio (mix_language.mp3), 22.19 MB, 1,938.495 s; spontaneous two-language conversation.
image
Output artifact for "Speaker Diarization" test: Run metrics report `diarization detected` and `distinct_speaker_labels 3` on the bilingual sample, showing that speaker labels are returned even in mixed-language audio., 03-terminal-metrics-input-3.png
Run metrics report `diarization detected` and `distinct_speaker_labels 3` on the bilingual sample, showing that speaker labels are returned even in mixed-language audio.
Bottom Line
Speaker labels are there, but the benchmark only proves label presence and count, not correct attribution.
From our researchearlier researchTranscribe Audio Accurately — Speech-to-Text Engine Benchmark
Multilingual Transcription
Test Summary
Feature tested: Multilingual Transcription
Result: Partial

Feature tested: Multilingual Transcription

Result: Partial

Expected behavior: ElevenLabs Scribe can transcribe mixed Spanish-English or code-switched audio without special configuration. The sampled evidence shows usable mixed-language transcription, though Spanish recall is weaker than in clean-English narration.

Test case: Audio file → Image

Input type: Audio file

Input used: Input artifact (Audio file): INPUT - Bilingual code-switching audio (mix_language.mp3), 22.19 MB, 1,938.495 s; mixed Spanish-English conversation. — mix_language.mp3

Observed output: Output artifact (Image): The metrics view reports 28.57% WER and Spanish recall of 57.5% (46/80 types) on the bilingual run, which is the clearest summary of multilingual quality. — 03-terminal-metrics-input-3.png

Input artifact: Input artifact (Audio file): INPUT - Bilingual code-switching audio (mix_language.mp3), 22.19 MB, 1,938.495 s; mixed Spanish-English conversation. — mix_language.mp3

Output artifact: Output artifact (Image): The metrics view reports 28.57% WER and Spanish recall of 57.5% (46/80 types) on the bilingual run, which is the clearest summary of multilingual quality. — 03-terminal-metrics-input-3.png

What changed: Audio file transformed into Image

Test case: Audio file → Image

Input type: Audio file

Input used: Input artifact (Audio file): INPUT - Bilingual code-switching audio (mix_language.mp3), 22.19 MB, 1,938.495 s; mixed Spanish-English conversation. — mix_language.mp3

Observed output: Output artifact (Image): The transcript-detail view shows the specific bilingual failure site: the reference has `ahora`, while the engine wrote `oh now` instead. — 04-transcript-detail-input-3.png

Input artifact: Input artifact (Audio file): INPUT - Bilingual code-switching audio (mix_language.mp3), 22.19 MB, 1,938.495 s; mixed Spanish-English conversation. — mix_language.mp3

Output artifact: Output artifact (Image): The transcript-detail view shows the specific bilingual failure site: the reference has `ahora`, while the engine wrote `oh now` instead. — 04-transcript-detail-input-3.png

What changed: Audio file transformed into Image

Why it matters / Conclusion: It can transcribe mixed-language audio, but Spanish quality is materially weaker than the clean-English medical sample.

ElevenLabs Scribe can transcribe mixed Spanish-English or code-switched audio without special configuration. The sampled evidence shows usable mixed-language transcription, though Spanish recall is weaker than in clean-English narration.

audio
0:00 / 0:00
Loading audio...
INPUT - Bilingual code-switching audio (mix_language.mp3), 22.19 MB, 1,938.495 s; mixed Spanish-English conversation.
image
Output artifact for "Multilingual Transcription" test: The metrics view reports 28.57% WER and Spanish recall of 57.5% (46/80 types) on the bilingual run, which is the clearest summary of multilingual quality., 03-terminal-metrics-input-3.png
The metrics view reports 28.57% WER and Spanish recall of 57.5% (46/80 types) on the bilingual run, which is the clearest summary of multilingual quality.
audio
0:00 / 0:00
Loading audio...
INPUT - Bilingual code-switching audio (mix_language.mp3), 22.19 MB, 1,938.495 s; mixed Spanish-English conversation.
image
Output artifact for "Multilingual Transcription" test: The transcript-detail view shows the specific bilingual failure site: the reference has `ahora`, while the engine wrote `oh now` instead., 04-transcript-detail-input-3.png
The transcript-detail view shows the specific bilingual failure site: the reference has `ahora`, while the engine wrote `oh now` instead.
Bottom Line
It can transcribe mixed-language audio, but Spanish quality is materially weaker than the clean-English medical sample.
From our researchTranscribe Audio Accurately — Speech-to-Text Engine Benchmark

How it scored on the research's own criteria

The 3 evaluation dimensions from our hands-on research on ElevenLabs Scribe, each judged from recorded runs on 3 test inputs — the same verdicts the ranking page ranks on.

held up  partial  failed  not exercised by this input

CriterionVerdictWhat the runs showedPer inputProof
Output qualityMixed3/5It was excellent on the dense anatomy narration, but the other two inputs were much rougher: one struggled with overlapping meeting speech and the other lost a lot of Spanish in the code-switched conversation. That makes the overall transcript quality clearly mixed rather than consistently strong.open proof ↗
Automation levelStrong5/5Each run was just one uploaded file and one returned transcript, with no manual cleanup or extra turns needed. That is about as hands-off as a transcription workflow gets, so the automation is top-tier.
Input handlingStrong5/5It took all three audio files in the same request shape, finished each one quickly, and kept costs low, so the input path looks robust and efficient. The only caveat is that this benchmark does not measure streaming, but that does not weaken the batch processing performance shown here.open proof ↗

Verdicts come verbatim from the study's recorded observations, never re-derived at render; a criterion with no recorded run shows Not exercised — this section cannot invent a score.

Official pricing

Current Scribe v2 list pricing from ElevenLabs' pricing page; the benchmark run itself used `scribe_v1`.

Free / Pay-as-you-go
$0.22/hr
4 h 30 min of Scribe v2 included; Scribe v2 Realtime is $0.39/hr with 2 h 30 min included.
Starter
$6/month
$0.22/hr with 27 h of Scribe v2 included; 15 h of Realtime included.
Creator
$22/month (first month $11)
$0.22/hr with 100 h of Scribe v2 included; 56 h of Realtime included.
Pro
$99/month
$0.22/hr with 450 h of Scribe v2 included; 254 h of Realtime included.
Scale
$299/month
$0.22/hr with 1,359 h of Scribe v2 included; 767 h of Realtime included.
Business
$990/month
$0.22/hr with 4,500 h of Scribe v2 included; 2,538 h of Realtime included.
Enterprise
Custom
Custom DPA/SLA, SSO, and HIPAA BAA; requires sales contact.
Startup Grants Program
Free for 12 months
33,000,000 characters; application required.

The per-hour Scribe rate is flat across plans at $0.22/hr; paid tiers mainly increase included hours and concurrency. Add-ons are priced separately: entity detection +$0.070/hr and keyterm prompting +$0.050/hr. Prices exclude taxes.

✓ Use This If
You need batch speech-to-text from a single multipart POST.
You need word-level timestamps, confidence values, and speaker labels in the transcript payload.
You need strong accuracy on jargon-heavy English narration.
You need fast batch turnaround at about $0.2202 per audio-hour.
✕ Skip This If
You need validated speaker attribution correctness on overlapping conversations.
You need overlap-heavy audio to stay faithful without insertions and over-segmentation.
You need bilingual Spanish-English audio to stay near the clean-English result.
You need streaming-latency evidence from this benchmark.
audio-speechaudio-to-textspeechOther
On the crosstalk sample it scored 26.67% WER, with 856 substitutions, 782 deletions, and 383 insertions. The transcript-detail view shows a large missing span, and the metrics report says it detected 5 speaker labels for a 4-participant meeting.
It was strongest on the medical-jargon sample: 3.01% WER, 51 substitutions, 8 deletions, and 23 insertions. The benchmark also reports 100.0% jargon recall on that input.
It handled mixed-language audio, but the bilingual run was much weaker than the medical sample: 28.57% WER, 57.5% Spanish token recall, and the highlighted diff dropped the Spanish token `ahora`.
Yes. Every scored run showed `word_timestamps yes`, `confidence yes`, and `speaker_labels yes`, with payload depth 3/3.
No. The benchmark measured label presence and label count, but not whether the labels were attributed to the correct speakers.
The run-level list-price cost came out to $0.13106, $0.06875, and $0.11857 across the three inputs, with a normalized rate of $0.2202 per audio-hour. Latency was 41.76s, 16.33s, and 6.07s, with RTFs of 0.01949, 0.01453, and 0.00313.
No. This research is a batch benchmark; it does not provide streaming-latency evidence.
No. The report says the Scribe v2 per-hour rate stays at $0.22/hr across plans; the plans mainly change included hours and concurrency.

Banner Preview

How the embed badge will look on your site

ElevenLabs Scribe featured on AI Demos

Embed HTML

Copy this code to your website source

<a target="_blank" href="https://aidemos.com/tools/elevenlabs-scribe?utm_source=elevenlabs-scribe_embed" style="width: 250px; height: 80px; border-radius:4px;" width="250" height="80"> <img src="https://aidemos-website-images.s3.amazonaws.com/featured.png" alt="ElevenLabs Scribe | Featured on AI Demos" style="width: 250px; height: 80px; border-radius:4px;" width="250" height="80"> </a>

Quick Integration Guide

  • 1Copy the HTML code block above.
  • 2Paste it into your site's HTML or CMS editor.
  • 3Banner appears instantly on your page.
  • 4Links back to your tool profile here.
Similar Tools

Similar Tools

Discover more AI tools like ElevenLabs Scribe to enhance your workflow.

Comments (0)

Please Log in to join the discussion.

Built by FutureSmart AI — the team behind AI Demos

Need a custom AI solution for this use case?

If you are looking to build a custom speech-to-text transcription, audio transcription, or subtitle generation system for your business or internal workflow, email us at contact@futuresmart.ai.

Get a custom build

Found something inaccurate or missing? We try to keep our AI research accurate and useful. If you found outdated information, an issue, or have a suggestion, email us at collaborate@aidemos.com.

Back to Top