On bilingual code-switching, the response duplicated across 2 channels, inflating output to 10,765 words against a 6,517-word reference and making the reported 88.45% WER an artefact of duplication rather than a usable accuracy score.
What was measured
Output quality
How accurately the returned transcript matches the human reference transcript, measured by WER.
decisive for this rankingtransformation
WER-based transcript accuracy is the core thing this ranking is about; it directly measures how well the tool transcribes audio. (3 of 3 judges)
What was given, what came back
Test input: Bilingual Spanish-English code-switching speech · audio · group: speech-to-text-benchmark
Input — what we sent
0:00 / 0:00
Loading audio...
52600606357c41998634fe876a6f214d.mp3
0:00 / 0:00
Loading audio...
Bilingual Spanish-English code-switching speech
A Bangor Miami bilingual corpus recording with mid-sentence switches between Spanish and English. It was used to test multilingual recognition, code-switch detection, and preservation of words across language transitions.
Why this input is hard
- · code-switching detection
- · multilingual language ID
- · mid-sentence language transitions
- · word preservation across language flips
- · hallucination resistance in bilingual speech
Output — unretouched



Also checked on this input — same tool, 2 other criteria
Automation level✓ WorkedThe workflow completes end to end without operator input, using the vendor's documented three-call sequence (upload, submit audio_url to /pre-recorded, then poll result_url until done); this run was multi-stage and did not instrument per-call timings.Export✓ WorkedReturns a rich transcript payload rather than plain text: payload depth is 3/3, with word-level timing, confidence, speaker labels, 26,166 timed tokens, and 7-level JSON nesting across 26,174 objects.
Provenance
- Observation
- fcccec3d-9bf1-46cc-8e80-e79e347f3665
- Evidence run
- 469de0c2-d727-4f8f-a60e-e3a5bf8e8588
- Study
- Transcribe Audio Accurately — Speech-to-Text Engine Benchmark
- Research task
- 86baxegpu
- Tested at
- not recorded
- Source
- first-party
- Evidence state
- verified
- Proof shown
- input + output shown
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "gladia",
scenario: "speech-to-text-benchmark"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 9 other tools
measured on Output quality
AssemblyAI◐ MixedAccuracy is middling on code-switching speech: 21.04% WER with 589 substitutions, 652 deletions, and 130 insertions against a 6,517-word reference; Spanish token recall is 72.5% (58/80 types), and the run is best of 10 on this input.AWS Transcribe◐ Mixed23.06Transcript quality is middling at 23.06% WER, with 673 substitutions, 665 deletions, and 165 insertions against 6517 reference words; Spanish recall is only 30.0% (24/80 types).Deepgram✗ Failed38.13Breaks down on code-switching speech: WER is 38.13% with 793 substitutions, 1259 deletions, and 433 insertions against a 6517-word reference, and Spanish token recall is only 3.8% (3/80 types).ElevenLabs Scribe⚠ Struggled28.57Transcript accuracy is weak on code-switching speech, with WER 28.57% and 752 substitutions, 778 deletions, and 332 insertions over a 6,517-word reference.Google Cloud Speech-to-Text✗ Failed56.13On code-switching speech, the transcript quality is very poor at 56.13% WER, with 716 substitutions, 2884 deletions, and 58 insertions against 6517 reference words, and only 5.0% Spanish token recall (4/80 types).GroqCloud⚠ Struggled27.94Performs much worse on code-switching speech than on medical narration: WER 27.94% with 636 substitutions, 1052 deletions, and 133 insertions over 6517 reference words, plus only 53.8% Spanish token recall.OpenAI Speech-to-Text⚠ Struggled26.12Transcription quality degrades sharply on code-switching: WER 26.12% with 623 substitutions, 947 deletions, and 132 insertions against 6517 reference words; Spanish token recall is only 46.2% (37/80).Rev AI⚠ Struggled25.16On the code-switching sample, it reached WER 25.16% against a 6517-word reference, with 691 substitutions, 793 deletions and 156 insertions; Spanish token recall was only 8.8% (7/80 types), so it dropped or anglicised most Spanish material.Speechmatics⚠ Struggled25.06The transcript struggles on code-switching speech, with 25.06% WER (560 substitutions, 981 deletions, 92 insertions over 6517 reference words) and only 12.5% Spanish token recall (10/80 types).
This evidence is published in
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com