On dense medical narration, the transcript reaches 13.09% WER with 228 substitutions, 49 deletions, and 80 insertions against 2728 reference words, and it recalls 100.0% of scored jargon terms.
◐ Mixed🧾 artifact-verifiedinput + output shown13.09Test date not recordedGoogle Cloud Speech-to-Text →
What was measured
Output quality
How accurately the returned transcript matches the human reference transcript, measured by WER.
decisive for this rankingtransformation
WER-based transcript accuracy is the core thing this ranking is about; it directly measures how well the tool transcribes audio. (3 of 3 judges)
What was given, what came back
Test input: Medical anatomy narration with dense jargon · audio · group: speech-to-text-benchmark
Input — what we sent
0:00 / 0:00
Loading audio...
Medical anatomy narration with dense jargon
A long narrated excerpt from Henry Gray's Anatomy of the Human Body containing dense medical terminology and accented articulation. It was used to test lexical accuracy on domain-specific jargon and spelling of technical terms.
Why this input is hard
- · domain-specific vocabulary recognition
- · medical term spelling accuracy
- · accented speech robustness
- · phoneme-to-grapheme precision
- · long-form audio handling
Output — unretouched


Also checked on this input — same tool, 2 other criteria
Automation level◐ MixedThe recorded workflow submits the audio as a POST to the v2 endpoint and reaches a scored result without operator input, but the trace explicitly says per-call timings were not instrumented, so no measured call count is claimed.Export◐ MixedReturns a mid-depth transcript payload: payload depth 2/3 with 2750 word-level timed tokens, confidence present, and no speaker labels.
Provenance
- Observation
- 491d0a47-7871-458c-8beb-234875efd947
- Evidence run
- 469de0c2-d727-4f8f-a60e-e3a5bf8e8588
- Study
- Transcribe Audio Accurately — Speech-to-Text Engine Benchmark
- Research task
- 86baxegpu
- Tested at
- not recorded
- Source
- first-party
- Evidence state
- verified
- Proof shown
- input + output shown
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "google-cloud-speech-to-text",
scenario: "speech-to-text-benchmark"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 9 other tools
measured on Output quality
AssemblyAI✓ WorkedAccuracy is strong on dense medical jargon: 3.78% WER with 71 substitutions, 14 deletions, and 18 insertions against a 2,728-word reference, and the run reports 100.0% jargon recall.AWS Transcribe✓ Worked3.63Transcript quality is strong at 3.63% WER, with 75 substitutions, 12 deletions, and 12 insertions against 2728 reference words.Deepgram✓ Worked5.43Keeps lexical accuracy high on dense medical narration: WER is 5.43% with 82 substitutions, 34 deletions, and 32 insertions against a 2728-word reference, and jargon recall is 100.0% (9/9 scored terms).ElevenLabs Scribe✓ Worked3.01Transcript accuracy is strong on dense medical narration, with WER 3.01% and 51 substitutions, 8 deletions, and 23 insertions over a 2,728-word reference.Gladia✓ Worked4.07On dense medical jargon, WER was 4.07% with 73 substitutions, 14 deletions, and 24 insertions against a 2,728-word reference; jargon recall was 100.0% (9/9 scored terms).GroqCloud✓ Worked3.15Delivers high lexical accuracy on dense medical narration: WER 3.15% with 58 substitutions, 13 deletions, and 15 insertions over 2728 reference words, returning 2730 hypothesis words.OpenAI Speech-to-Text✓ Worked5.17Achieves low transcription error on dense jargon: WER 5.17% with 54 substitutions, 76 deletions, and 11 insertions against 2728 reference words; it returned 2663 words and missed only 'trabeculae'.Rev AI✓ Worked9.79On the dense medical narration, it held WER to 9.79% against a 2728-word reference, with 195 substitutions, 7 deletions and 65 insertions; jargon recall was 77.8% (7/9 scored terms), so it handled the terminology reasonably well but still missed specific terms such as cancellous and trabeculae.Speechmatics✓ Worked3.01The transcript is highly accurate on dense medical narration, with 3.01% WER (49 substitutions, 17 deletions, 16 insertions over 2728 reference words) and 100.0% jargon recall.
This evidence is published in
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com