Produces no transcript on the oversized file, so WER is unmeasured here; the report explicitly treats accuracy for this input as untested rather than poor.
What was measured
Output quality
How accurately the returned transcript matches the human reference transcript, measured by WER.
decisive for this rankingtransformation
WER-based transcript accuracy is the core thing this ranking is about; it directly measures how well the tool transcribes audio. (3 of 3 judges)
What was given, what came back
Test input: Overlapping meeting speech with cross-talk · audio · group: speech-to-text-benchmark
Input — what we sent

a1e0fc3436944f299d84aaff8f30b14d.png
0:00 / 0:00
Loading audio...
Overlapping meeting speech with cross-talk
A long AMI meeting audio file with multiple speakers talking over one another, background room noise, and crosstalk. It was used to test how well an STT system handles noisy multi-speaker conversational audio and speaker separation.
Why this input is hard
- · overlapping speech
- · background noise robustness
- · multi-speaker separation
- · speaker diarization accuracy
- · long-form audio handling
Output — unretouched


Also checked on this input — same tool, 2 other criteria
Automation level◐ MixedUses a single multipart POST workflow, but this run stops at API rejection and never completes transcription end to end, so the automation is one-step but unsuccessful on this input.Export✗ FailedReturns no transcript payload at all on the oversized upload; the only returned content is an error object, so export richness is effectively zero.
Provenance
- Observation
- c48a4462-c6e1-48ae-8339-455ac16704ed
- Evidence run
- 469de0c2-d727-4f8f-a60e-e3a5bf8e8588
- Study
- Transcribe Audio Accurately — Speech-to-Text Engine Benchmark
- Research task
- 86baxegpu
- Tested at
- not recorded
- Source
- first-party
- Evidence state
- verified
- Proof shown
- input + output shown
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "groqcloud",
scenario: "speech-to-text-benchmark"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 9 other tools
measured on Output quality
AssemblyAI⚠ StruggledAccuracy is weak on overlapping speech: 33.16% WER with 461 substitutions, 1,976 deletions, and 76 insertions against a 7,579-word reference, ranking 4th of 8 on this input.AWS Transcribe⚠ Struggled33.88Transcript quality is weak at 33.88% WER, with 363 substitutions, 2162 deletions, and 43 insertions against 7579 reference words.Deepgram⚠ Struggled36.27On overlapping meeting crosstalk, the transcript quality is weak: WER is 36.27% with 1096 substitutions, 1143 deletions, and 510 insertions against a 7579-word reference, returning 6946 words overall.ElevenLabs Scribe⚠ Struggled26.67Transcript accuracy is weak on crosstalk, with WER 26.67% and 856 substitutions, 782 deletions, and 383 insertions over a 7,579-word reference.Gladia⚠ Struggled37.35On overlapping crosstalk, it still scored the run but WER was 37.35% with 529 substitutions, 2,213 deletions, and 89 insertions against a 7,579-word reference, leaving 5,455 words in the hypothesis.Google Cloud Speech-to-Text⚠ Struggled43.5On overlapping crosstalk, the transcript quality is poor at 43.50% WER, with 633 substitutions, 2614 deletions, and 50 insertions against 7579 reference words, yielding 5015 hypothesis words.OpenAI Speech-to-Text◐ MixedTranscript accuracy is untested because no transcript was produced; the run stopped at the size rejection before any WER could be measured.Rev AI⚠ Struggled28.33On the overlapping meeting audio, it reached WER 28.33% against a 7579-word reference, with 601 substitutions, 1423 deletions and 123 insertions; it returned 6279 words, so accuracy degraded substantially under crosstalk.Speechmatics⚠ Struggled26.63The transcript is only moderately accurate on overlapping speech, with 26.63% WER (507 substitutions, 1430 deletions, 81 insertions over 7579 reference words) and diarization over-segmenting the meeting into 5 speaker labels.
This evidence is published in
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com