Accepts a 65.39 MB, 2142.709 s WAV and processes it through without objection.
What was measured
Input handling
Whether the tool accepts the benchmark audio as provided and processes it end to end without objection, including the file size/duration it can handle.
context, not decisivecapability
Being able to accept the benchmark audio and handle its size/duration is necessary to test the tool, but it is an operational constraint rather than the transcription quality being ranked. (3 of 3 judges)
What was given, what came back
Test input: Overlapping meeting speech with cross-talk · audio · group: speech-to-text-benchmark
Input — what we sent



0:00 / 0:00
Loading audio...
Overlapping meeting speech with cross-talk
A long AMI meeting audio file with multiple speakers talking over one another, background room noise, and crosstalk. It was used to test how well an STT system handles noisy multi-speaker conversational audio and speaker separation.
Why this input is hard
- · overlapping speech
- · background noise robustness
- · multi-speaker separation
- · speaker diarization accuracy
- · long-form audio handling
Output — unretouched



Also checked on this input — same tool, 2 other criteria
Export✓ WorkedReturns a full developer payload with payload depth 3/3, 5912 word-level timed tokens, confidence values, and speaker labels; the raw JSON walk spans 7 levels and 19457 objects.Output quality⚠ StruggledTranscript quality is weak at 33.88% WER, with 363 substitutions, 2162 deletions, and 43 insertions against 7579 reference words.
Provenance
- Observation
- 311e6d75-7711-42c7-9249-985257561481
- Evidence run
- 469de0c2-d727-4f8f-a60e-e3a5bf8e8588
- Study
- Transcribe Audio Accurately — Speech-to-Text Engine Benchmark
- Research task
- 86baxegpu
- Tested at
- not recorded
- Source
- first-party
- Evidence state
- verified
- Proof shown
- input + output shown
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "aws-transcribe",
scenario: "speech-to-text-benchmark"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 1 other tool
measured on Input handling
This evidence is published in
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com