The tool accepts a 65.39 MB, 2142.709 s audio file and processes it end to end without objection.
What was measured
Input handling
Whether the tool accepts the benchmark audio as provided and processes it end to end without objection, including the file size/duration it can handle.
context, not decisivecapability
Being able to accept the benchmark audio and handle its size/duration is necessary to test the tool, but it is an operational constraint rather than the transcription quality being ranked. (3 of 3 judges)
What was given, what came back
Test input: Overlapping meeting speech with cross-talk · audio · group: speech-to-text-benchmark
Input — what we sent



0:00 / 0:00
Loading audio...
Overlapping meeting speech with cross-talk
A long AMI meeting audio file with multiple speakers talking over one another, background room noise, and crosstalk. It was used to test how well an STT system handles noisy multi-speaker conversational audio and speaker separation.
Why this input is hard
- · overlapping speech
- · background noise robustness
- · multi-speaker separation
- · speaker diarization accuracy
- · long-form audio handling
Output — unretouched



Also checked on this input — same tool, 2 other criteria
Export✓ WorkedThe returned payload is rich and fully structured at 3/3 depth, with word-level timestamps, confidence values, speaker labels, and 14,506 timed tokens.Output quality⚠ StruggledTranscript accuracy is weak on crosstalk, with WER 26.67% and 856 substitutions, 782 deletions, and 383 insertions over a 7,579-word reference.
Provenance
- Observation
- fb95bb1d-6e61-4356-ad55-4b2e77730f17
- Evidence run
- 469de0c2-d727-4f8f-a60e-e3a5bf8e8588
- Study
- Transcribe Audio Accurately — Speech-to-Text Engine Benchmark
- Research task
- 86baxegpu
- Tested at
- not recorded
- Source
- first-party
- Evidence state
- verified
- Proof shown
- input + output shown
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "elevenlabs-scribe",
scenario: "speech-to-text-benchmark"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 1 other tool
measured on Input handling
This evidence is published in
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com