Transcript accuracy is weak on code-switching speech, with WER 28.57% and 752 substitutions, 778 deletions, and 332 insertions over a 6,517-word reference.
What was measured
Output quality
How accurately the returned transcript matches the human reference transcript, measured by WER.
decisive for this rankingtransformation
WER-based transcript accuracy is the core thing this ranking is about; it directly measures how well the tool transcribes audio. (3 of 3 judges)
What was given, what came back
Test input: Bilingual Spanish-English code-switching speech · audio · group: speech-to-text-benchmark
Input — what we sent
0:00 / 0:00
Loading audio...
Bilingual Spanish-English code-switching speech
A Bangor Miami bilingual corpus recording with mid-sentence switches between Spanish and English. It was used to test multilingual recognition, code-switch detection, and preservation of words across language transitions.
Why this input is hard
- · code-switching detection
- · multilingual language ID
- · mid-sentence language transitions
- · word preservation across language flips
- · hallucination resistance in bilingual speech
Output — unretouched


Also checked on this input — same tool, 1 other criterion
Provenance
- Observation
- dc0d4b3a-f02b-47bb-9bdb-7fafae6e8e3a
- Evidence run
- 469de0c2-d727-4f8f-a60e-e3a5bf8e8588
- Study
- Transcribe Audio Accurately — Speech-to-Text Engine Benchmark
- Research task
- 86baxegpu
- Tested at
- not recorded
- Source
- first-party
- Evidence state
- observed
- Proof shown
- input + output shown
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "elevenlabs-scribe",
scenario: "speech-to-text-benchmark"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 9 other tools
measured on Output quality
AssemblyAI◐ MixedAccuracy is middling on code-switching speech: 21.04% WER with 589 substitutions, 652 deletions, and 130 insertions against a 6,517-word reference; Spanish token recall is 72.5% (58/80 types), and the run is best of 10 on this input.AWS Transcribe◐ Mixed23.06Transcript quality is middling at 23.06% WER, with 673 substitutions, 665 deletions, and 165 insertions against 6517 reference words; Spanish recall is only 30.0% (24/80 types).Deepgram✗ Failed38.13Breaks down on code-switching speech: WER is 38.13% with 793 substitutions, 1259 deletions, and 433 insertions against a 6517-word reference, and Spanish token recall is only 3.8% (3/80 types).Gladia✗ Failed88.45On bilingual code-switching, the response duplicated across 2 channels, inflating output to 10,765 words against a 6,517-word reference and making the reported 88.45% WER an artefact of duplication rather than a usable accuracy score.Google Cloud Speech-to-Text✗ Failed56.13On code-switching speech, the transcript quality is very poor at 56.13% WER, with 716 substitutions, 2884 deletions, and 58 insertions against 6517 reference words, and only 5.0% Spanish token recall (4/80 types).GroqCloud⚠ Struggled27.94Performs much worse on code-switching speech than on medical narration: WER 27.94% with 636 substitutions, 1052 deletions, and 133 insertions over 6517 reference words, plus only 53.8% Spanish token recall.OpenAI Speech-to-Text⚠ Struggled26.12Transcription quality degrades sharply on code-switching: WER 26.12% with 623 substitutions, 947 deletions, and 132 insertions against 6517 reference words; Spanish token recall is only 46.2% (37/80).Rev AI⚠ Struggled25.16On the code-switching sample, it reached WER 25.16% against a 6517-word reference, with 691 substitutions, 793 deletions and 156 insertions; Spanish token recall was only 8.8% (7/80 types), so it dropped or anglicised most Spanish material.Speechmatics⚠ Struggled25.06The transcript struggles on code-switching speech, with 25.06% WER (560 substitutions, 981 deletions, 92 insertions over 6517 reference words) and only 12.5% Spanish token recall (10/80 types).
This evidence is published in
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com