Accuracy is middling on code-switching speech: 21.04% WER with 589 substitutions, 652 deletions, and 130 insertions against a 6,517-word reference; Spanish token recall is 72.5% (58/80 types), and the run is best of 10 on this input.

◐ Mixed🧾 artifact-verifiedinput + output shownTest date not recordedAssemblyAI
What was measured
Output quality

How accurately the returned transcript matches the human reference transcript, measured by WER.

decisive for this rankingtransformation

WER-based transcript accuracy is the core thing this ranking is about; it directly measures how well the tool transcribes audio. (3 of 3 judges)

What was given, what came back

Test input: Bilingual Spanish-English code-switching speech · audio · group: speech-to-text-benchmark
Input — what we sent
Input file 1 — as supplied
e319813d52304d89a559933409e837ec.png
e319813d52304d89a559933409e837ec.png
Input file 2 — as supplied
0:00 / 0:00
Loading audio...
Bilingual Spanish-English code-switching speech

A Bangor Miami bilingual corpus recording with mid-sentence switches between Spanish and English. It was used to test multilingual recognition, code-switch detection, and preservation of words across language transitions.

Why this input is hard
  • · code-switching detection
  • · multilingual language ID
  • · mid-sentence language transitions
  • · word preservation across language flips
  • · hallucination resistance in bilingual speech
Output — unretouched
Output 1
Output 1
Output 2
Output 2
Provenance
Observation
9d4f8145-0a65-45a9-bc6d-0299eef9e293
Evidence run
469de0c2-d727-4f8f-a60e-e3a5bf8e8588
Study
Transcribe Audio Accurately — Speech-to-Text Engine Benchmark
Research task
86baxegpu
Tested at
not recorded
Source
first-party
Evidence state
verified
Proof shown
input + output shown
Cost / latency
not captured
Repeat run
not captured
Tester
not captured

The last three rows are honest blanks, not placeholders — our capture has no field for them yet.

Query this
get_evidence({
  tool: "assemblyai-speech-to-text",
  scenario: "speech-to-text-benchmark"
})
MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 9 other tools
measured on Output quality
AWS Transcribe◐ Mixed23.06Transcript quality is middling at 23.06% WER, with 673 substitutions, 665 deletions, and 165 insertions against 6517 reference words; Spanish recall is only 30.0% (24/80 types).Deepgram✗ Failed38.13Breaks down on code-switching speech: WER is 38.13% with 793 substitutions, 1259 deletions, and 433 insertions against a 6517-word reference, and Spanish token recall is only 3.8% (3/80 types).ElevenLabs Scribe⚠ Struggled28.57Transcript accuracy is weak on code-switching speech, with WER 28.57% and 752 substitutions, 778 deletions, and 332 insertions over a 6,517-word reference.Gladia✗ Failed88.45On bilingual code-switching, the response duplicated across 2 channels, inflating output to 10,765 words against a 6,517-word reference and making the reported 88.45% WER an artefact of duplication rather than a usable accuracy score.Google Cloud Speech-to-Text✗ Failed56.13On code-switching speech, the transcript quality is very poor at 56.13% WER, with 716 substitutions, 2884 deletions, and 58 insertions against 6517 reference words, and only 5.0% Spanish token recall (4/80 types).GroqCloud⚠ Struggled27.94Performs much worse on code-switching speech than on medical narration: WER 27.94% with 636 substitutions, 1052 deletions, and 133 insertions over 6517 reference words, plus only 53.8% Spanish token recall.OpenAI Speech-to-Text⚠ Struggled26.12Transcription quality degrades sharply on code-switching: WER 26.12% with 623 substitutions, 947 deletions, and 132 insertions against 6517 reference words; Spanish token recall is only 46.2% (37/80).Rev AI⚠ Struggled25.16On the code-switching sample, it reached WER 25.16% against a 6517-word reference, with 691 substitutions, 793 deletions and 156 insertions; Spanish token recall was only 8.8% (7/80 types), so it dropped or anglicised most Spanish material.Speechmatics⚠ Struggled25.06The transcript struggles on code-switching speech, with 25.06% WER (560 substitutions, 981 deletions, 92 insertions over 6517 reference words) and only 12.5% Spanish token recall (10/80 types).
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com