Returns a rich transcript payload with payload_depth 3/3, including word-level timing, confidence values, and speaker labels in the JSON response.

✓ Worked🧾 artifact-verifiedinput + output shownTest date not recordedSpeechmatics
What was measured
Export

How complete the returned transcript payload is, as reflected in the depth or richness of what the tool outputs.

context, not decisivetransformation

Transcript payload richness is useful for comparison, but it does not measure transcription correctness against the reference. (3 of 3 judges)

What was given, what came back

Test input: Overlapping meeting speech with cross-talk · audio · group: speech-to-text-benchmark
Input — what we sent
Input file 1 — as supplied
770b31cffe56485dac23dbf4206bdcdd.png
770b31cffe56485dac23dbf4206bdcdd.png
Input file 2 — as supplied
0:00 / 0:00
Loading audio...
Overlapping meeting speech with cross-talk

A long AMI meeting audio file with multiple speakers talking over one another, background room noise, and crosstalk. It was used to test how well an STT system handles noisy multi-speaker conversational audio and speaker separation.

Why this input is hard
  • · overlapping speech
  • · background noise robustness
  • · multi-speaker separation
  • · speaker diarization accuracy
  • · long-form audio handling
Output — unretouched
Output 1
Output 1
Output 2
Output 2
Output 3
research-media-raw-response-503967264599.json
Loading file...
Provenance
Observation
9909b52b-84f0-40bb-b464-d0465e3bd1c4
Evidence run
469de0c2-d727-4f8f-a60e-e3a5bf8e8588
Study
Transcribe Audio Accurately — Speech-to-Text Engine Benchmark
Research task
86baxegpu
Tested at
not recorded
Source
first-party
Evidence state
verified
Proof shown
input + output shown
Cost / latency
not captured
Repeat run
not captured
Tester
not captured

The last three rows are honest blanks, not placeholders — our capture has no field for them yet.

Query this
get_evidence({
  tool: "speechmatics",
  scenario: "speech-to-text-benchmark"
})
MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 9 other tools
measured on Export
AssemblyAI✓ WorkedReturns a rich developer payload with word-level timestamps, confidence values, and speaker labels; the raw response shows 11,743 timed tokens and JSON depth 5 across 11,749 objects, with payload depth 3/3.AWS Transcribe✓ Worked3/3Returns a full developer payload with payload depth 3/3, 5912 word-level timed tokens, confidence values, and speaker labels; the raw JSON walk spans 7 levels and 19457 objects.Deepgram✓ WorkedReturns a rich developer payload on hard crosstalk audio, with word-level timing, confidence, speaker labels, 13,555 timed tokens, 4 distinct speakers, payload depth 3/3, and JSON depth 11 across 14,065 objects.ElevenLabs Scribe✓ Worked3/3The returned payload is rich and fully structured at 3/3 depth, with word-level timestamps, confidence values, speaker labels, and 14,506 timed tokens.Gladia✓ WorkedReturns a rich transcript payload rather than plain text: payload depth is 3/3, with word-level timing, confidence, speaker labels, 12,968 timed tokens, and 7-level JSON nesting across 12,976 objects.Google Cloud Speech-to-Text◐ MixedReturns a mid-depth transcript payload: payload depth 2/3 with 5009 word-level timed tokens, confidence present, and no speaker labels.GroqCloud✗ FailedReturns no transcript payload at all on the oversized upload; the only returned content is an error object, so export richness is effectively zero.OpenAI Speech-to-Text✗ FailedReturns no transcript payload at all on this input; the only output is a 413 error body, so there is nothing transcript-like to export.Rev AI✓ WorkedReturned a rich developer payload with word-level timing, confidence and speaker labels; the raw response walk reports 6249 timed tokens, payload depth 3/3, and JSON depth 5 across 14484 objects.
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com