Transcript accuracy is weak on crosstalk, with WER 26.67% and 856 substitutions, 782 deletions, and 383 insertions over a 7,579-word reference.

⚠ Struggledinput + output shown26.67Test date not recordedElevenLabs Scribe
What was measured
Output quality

How accurately the returned transcript matches the human reference transcript, measured by WER.

decisive for this rankingtransformation

WER-based transcript accuracy is the core thing this ranking is about; it directly measures how well the tool transcribes audio. (3 of 3 judges)

What was given, what came back

Test input: Overlapping meeting speech with cross-talk · audio · group: speech-to-text-benchmark
Input — what we sent
0:00 / 0:00
Loading audio...
Overlapping meeting speech with cross-talk

A long AMI meeting audio file with multiple speakers talking over one another, background room noise, and crosstalk. It was used to test how well an STT system handles noisy multi-speaker conversational audio and speaker separation.

Why this input is hard
  • · overlapping speech
  • · background noise robustness
  • · multi-speaker separation
  • · speaker diarization accuracy
  • · long-form audio handling
Output — unretouched
Output 1
Output 1
Output 2
Output 2
Provenance
Observation
667df61a-e337-48ae-b65c-6d910e2ccf83
Evidence run
469de0c2-d727-4f8f-a60e-e3a5bf8e8588
Study
Transcribe Audio Accurately — Speech-to-Text Engine Benchmark
Research task
86baxegpu
Tested at
not recorded
Source
first-party
Evidence state
observed
Proof shown
input + output shown
Cost / latency
not captured
Repeat run
not captured
Tester
not captured

The last three rows are honest blanks, not placeholders — our capture has no field for them yet.

Query this
get_evidence({
  tool: "elevenlabs-scribe",
  scenario: "speech-to-text-benchmark"
})
MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 9 other tools
measured on Output quality
AssemblyAI⚠ StruggledAccuracy is weak on overlapping speech: 33.16% WER with 461 substitutions, 1,976 deletions, and 76 insertions against a 7,579-word reference, ranking 4th of 8 on this input.AWS Transcribe⚠ Struggled33.88Transcript quality is weak at 33.88% WER, with 363 substitutions, 2162 deletions, and 43 insertions against 7579 reference words.Deepgram⚠ Struggled36.27On overlapping meeting crosstalk, the transcript quality is weak: WER is 36.27% with 1096 substitutions, 1143 deletions, and 510 insertions against a 7579-word reference, returning 6946 words overall.Gladia⚠ Struggled37.35On overlapping crosstalk, it still scored the run but WER was 37.35% with 529 substitutions, 2,213 deletions, and 89 insertions against a 7,579-word reference, leaving 5,455 words in the hypothesis.Google Cloud Speech-to-Text⚠ Struggled43.5On overlapping crosstalk, the transcript quality is poor at 43.50% WER, with 633 substitutions, 2614 deletions, and 50 insertions against 7579 reference words, yielding 5015 hypothesis words.GroqCloud◐ MixedProduces no transcript on the oversized file, so WER is unmeasured here; the report explicitly treats accuracy for this input as untested rather than poor.OpenAI Speech-to-Text◐ MixedTranscript accuracy is untested because no transcript was produced; the run stopped at the size rejection before any WER could be measured.Rev AI⚠ Struggled28.33On the overlapping meeting audio, it reached WER 28.33% against a 7579-word reference, with 601 substitutions, 1423 deletions and 123 insertions; it returned 6279 words, so accuracy degraded substantially under crosstalk.Speechmatics⚠ Struggled26.63The transcript is only moderately accurate on overlapping speech, with 26.63% WER (507 substitutions, 1430 deletions, 81 insertions over 7579 reference words) and diarization over-segmenting the meeting into 5 speaker labels.
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com