Speechmatics icon
audio-speech

Speechmatics

Strong batch STT for hard English audio, but weak on code-switching as configured.

Visit Speechmatics
Batch STTWord timingsSpeaker labelsWeak code-switching
TL;DR — our verdictUpdated September 2026 · 22 test artifacts

Mixed but genuinely useful for the right audio

Where it wins
  • You need batch speech-to-text that returns word timings, confidence values, and speaker labels.
  • You are comparing engines on hard English audio such as overlapping speakers or jargon-heavy narration.
  • You can accept batch processing rather than live streaming.
Main limitation
  • You need proven balanced multilingual or code-switching performance; this run was mostly English and Spanish recall was poor.
Pricing (verified plans)
Free $0 — $100 in creditPro from $0.129/hrEnterprise Custom
Strongest test artifacts

Our take

Speechmatics is a strong batch STT choice for hard English audio: it won the overlapping-speech test, stayed near the top on medical jargon, and returned structured JSON with word timings, confidence, and speaker labels. The configured bilingual run was still mostly English and Spanish recall was very low, so multilingual robustness is not proven here. The benchmark also used the Enhanced batch rate, so the quality needs to justify that price against cheaper engines.

Screen recording of the STT benchmark folder in Finder, then Terminal running Speechmatics (Ursa/Enhanced) across the three scored inputs and printing benchmark progress and metrics.

In-Depth Review

Our detailed analysis of Speechmatics — features, performance, and real-world testing.

AD
AI Demos Team
Expert Reviewer
Verified Review

Feature-by-Feature Breakdown

Asynchronous Batch Transcription
Reliable batch execution
Test Summary
Feature tested: Asynchronous Batch Transcription
Result: Partial — Reliable batch execution

Feature tested: Asynchronous Batch Transcription

Result: Partial

Verdict: Reliable batch execution

Expected behavior: Submits audio as an asynchronous batch job and returns a completed transcript without manual intervention. The member cards exercised crosstalk meeting audio, a jargon-heavy lecture, a medical narration, and a Spanish-English conversation through the documented batch flow.

Test case: Audio file → Image

Input type: Audio file

Input used: Input artifact (Audio file): Input-1: Overlapping Speech / Crosstalk batch run on crosstalk.wav (long-form meeting audio with heavy overlap). — crosstalk.wav

Observed output: Output artifact (Image): Scored batch run for the crosstalk input: status scored, 46.71s latency, RTF 0.0218, cost $0.23809, and 6230 returned words against the reference transcript. — 07-automation-trace.png

Input artifact: Input artifact (Audio file): Input-1: Overlapping Speech / Crosstalk batch run on crosstalk.wav (long-form meeting audio with heavy overlap). — crosstalk.wav

Output artifact: Output artifact (Image): Scored batch run for the crosstalk input: status scored, 46.71s latency, RTF 0.0218, cost $0.23809, and 6230 returned words against the reference transcript. — 07-automation-trace.png

What changed: Audio file transformed into Image

Test case: Audio file → Image

Input type: Audio file

Input used: Input artifact (Audio file): Input-2: Medical Jargon batch run on medical_terms.mp3 (single-speaker anatomical narration). — medical_terms.mp3

Observed output: Output artifact (Image): Scored batch run for the medical-jargon input: status scored, 21.71s latency, RTF 0.0193, cost $0.12489, and 2727 returned words against the reference transcript. — 07-automation-trace-2.png

Input artifact: Input artifact (Audio file): Input-2: Medical Jargon batch run on medical_terms.mp3 (single-speaker anatomical narration). — medical_terms.mp3

Output artifact: Output artifact (Image): Scored batch run for the medical-jargon input: status scored, 21.71s latency, RTF 0.0193, cost $0.12489, and 2727 returned words against the reference transcript. — 07-automation-trace-2.png

What changed: Audio file transformed into Image

Test case: Audio file → Image

Input type: Audio file

Input used: Input artifact (Audio file): Input-3: Bilingual Code-Switching batch run on mix_language.mp3 (Spanish-English conversation configured with English language routing). — mix_language.mp3

Observed output: Output artifact (Image): Scored batch run for the bilingual input: status scored, 69.22s latency, RTF 0.03571, cost $0.2154, and 5628 returned words against the reference transcript. — 07-automation-trace-3.png

Input artifact: Input artifact (Audio file): Input-3: Bilingual Code-Switching batch run on mix_language.mp3 (Spanish-English conversation configured with English language routing). — mix_language.mp3

Output artifact: Output artifact (Image): Scored batch run for the bilingual input: status scored, 69.22s latency, RTF 0.03571, cost $0.2154, and 5628 returned words against the reference transcript. — 07-automation-trace-3.png

What changed: Audio file transformed into Image

Test case: Audio file → Image

Input type: Audio file

Input used: Input artifact (Audio file): Medical Jargon — dense anatomical vocabulary from Gray's Anatomy. — medical_terms.mp3

Observed output: Output artifact (Image): The medical-jargon sample scored 3.01% WER and 100.0% jargon recall; the transcript preserved the scored anatomy terms rather than dropping them. — 04-transcript-detail-2.png

Input artifact: Input artifact (Audio file): Medical Jargon — dense anatomical vocabulary from Gray's Anatomy. — medical_terms.mp3

Output artifact: Output artifact (Image): The medical-jargon sample scored 3.01% WER and 100.0% jargon recall; the transcript preserved the scored anatomy terms rather than dropping them. — 04-transcript-detail-2.png

What changed: Audio file transformed into Image

Test case: Audio file → Text/code file

Input type: Audio file

Input used: Input artifact (Audio file): INPUT — crosstalk.wav: 35:43 four-way overlapping meeting audio (AMI EN2002a). — crosstalk.wav

Observed output: Output artifact (Text/code file): Batch transcript returned for the crosstalk run; WER 26.63% and wall-clock 46.71s. — raw-response.json

Input artifact: Input artifact (Audio file): INPUT — crosstalk.wav: 35:43 four-way overlapping meeting audio (AMI EN2002a). — crosstalk.wav

Output artifact: Output artifact (Text/code file): Batch transcript returned for the crosstalk run; WER 26.63% and wall-clock 46.71s. — raw-response.json

What changed: Audio file transformed into Text/code file

Test case: Audio file → Text/code file

Input type: Audio file

Input used: Input artifact (Audio file): INPUT — medical_terms.mp3: 18:44 clean single-speaker medical narration (Gray's Anatomy via LibriVox). — medical_terms.mp3

Observed output: Output artifact (Text/code file): Batch transcript returned for the medical jargon run; WER 3.01% and wall-clock 21.71s. — raw-response-2.json

Input artifact: Input artifact (Audio file): INPUT — medical_terms.mp3: 18:44 clean single-speaker medical narration (Gray's Anatomy via LibriVox). — medical_terms.mp3

Output artifact: Output artifact (Text/code file): Batch transcript returned for the medical jargon run; WER 3.01% and wall-clock 21.71s. — raw-response-2.json

What changed: Audio file transformed into Text/code file

Test case: Audio file → Text/code file

Input type: Audio file

Input used: Input artifact (Audio file): INPUT — mix_language.mp3: 32:18 Spanish-English conversation, mostly English. — mix_language.mp3

Observed output: Output artifact (Text/code file): Batch transcript returned for the bilingual run; WER 25.06% and wall-clock 69.22s. — raw-response-3.json

Input artifact: Input artifact (Audio file): INPUT — mix_language.mp3: 32:18 Spanish-English conversation, mostly English. — mix_language.mp3

Output artifact: Output artifact (Text/code file): Batch transcript returned for the bilingual run; WER 25.06% and wall-clock 69.22s. — raw-response-3.json

What changed: Audio file transformed into Text/code file

Test case: Audio file → Image

Input type: Audio file

Input used: Input artifact (Audio file): Input-1: Overlapping Speech / Crosstalk — crosstalk.wav, 35:43, four-way overlapping meeting audio. — crosstalk.wav

Observed output: Output artifact (Image): WER 26.63% on the crosstalk run, with 507 substitutions, 1430 deletions, and 81 insertions against a 7579-word reference. It was the best of 8 engines on this input, but the largest divergence omitted a long overlapping span. — 03-terminal-metrics-input-1.png

Input artifact: Input artifact (Audio file): Input-1: Overlapping Speech / Crosstalk — crosstalk.wav, 35:43, four-way overlapping meeting audio. — crosstalk.wav

Output artifact: Output artifact (Image): WER 26.63% on the crosstalk run, with 507 substitutions, 1430 deletions, and 81 insertions against a 7579-word reference. It was the best of 8 engines on this input, but the largest divergence omitted a long overlapping span. — 03-terminal-metrics-input-1.png

What changed: Audio file transformed into Image

Test case: Audio file → Image

Input type: Audio file

Input used: Input artifact (Audio file): Input-2: Medical Jargon — medical_terms.mp3, 18:44, Gray's Anatomy narration dense with anatomical terms. — medical_terms.mp3

Observed output: Output artifact (Image): WER 3.01% on the medical-jargon run, with 49 substitutions, 17 deletions, and 16 insertions against a 2728-word reference. It was second of 10 and kept jargon recall at 100%. — 03-terminal-metrics-input-2.png

Input artifact: Input artifact (Audio file): Input-2: Medical Jargon — medical_terms.mp3, 18:44, Gray's Anatomy narration dense with anatomical terms. — medical_terms.mp3

Output artifact: Output artifact (Image): WER 3.01% on the medical-jargon run, with 49 substitutions, 17 deletions, and 16 insertions against a 2728-word reference. It was second of 10 and kept jargon recall at 100%. — 03-terminal-metrics-input-2.png

What changed: Audio file transformed into Image

Test case: Audio file → Image

Input type: Audio file

Input used: Input artifact (Audio file): Input-3: Bilingual Code-Switching — mix_language.mp3, 32:18, spontaneous Spanish-English conversation. — mix_language.mp3

Observed output: Output artifact (Image): WER 25.06% on the bilingual run, with 560 substitutions, 981 deletions, and 92 insertions against a 6517-word reference. Spanish token recall was only 12.5%, so this run does not support balanced code-switching claims. — 03-terminal-metrics-input-3.png

Input artifact: Input artifact (Audio file): Input-3: Bilingual Code-Switching — mix_language.mp3, 32:18, spontaneous Spanish-English conversation. — mix_language.mp3

Output artifact: Output artifact (Image): WER 25.06% on the bilingual run, with 560 substitutions, 981 deletions, and 92 insertions against a 6517-word reference. Spanish token recall was only 12.5%, so this run does not support balanced code-switching claims. — 03-terminal-metrics-input-3.png

What changed: Audio file transformed into Image

Why it matters / Conclusion: Stable batch execution across all three inputs. The run completed end to end, but per-call timings were not instrumented, so the exact HTTP call count is documented rather than measured here.

Submits audio as an asynchronous batch job and returns a completed transcript without manual intervention. The member cards exercised crosstalk meeting audio, a jargon-heavy lecture, a medical narration, and a Spanish-English conversation through the documented batch flow.

audio
0:00 / 0:00
Loading audio...
Input-1: Overlapping Speech / Crosstalk batch run on crosstalk.wav (long-form meeting audio with heavy overlap).
image
Output artifact for "Asynchronous Batch Transcription" test: Scored batch run for the crosstalk input: status scored, 46.71s latency, RTF 0.0218, cost $0.23809, and 6230 returned words against the reference transcript., 07-automation-trace.png
Scored batch run for the crosstalk input: status scored, 46.71s latency, RTF 0.0218, cost $0.23809, and 6230 returned words against the reference transcript.
audio
0:00 / 0:00
Loading audio...
Input-2: Medical Jargon batch run on medical_terms.mp3 (single-speaker anatomical narration).
image
Output artifact for "Asynchronous Batch Transcription" test: Scored batch run for the medical-jargon input: status scored, 21.71s latency, RTF 0.0193, cost $0.12489, and 2727 returned words against the reference transcript., 07-automation-trace-2.png
Scored batch run for the medical-jargon input: status scored, 21.71s latency, RTF 0.0193, cost $0.12489, and 2727 returned words against the reference transcript.
audio
0:00 / 0:00
Loading audio...
Input-3: Bilingual Code-Switching batch run on mix_language.mp3 (Spanish-English conversation configured with English language routing).
image
Output artifact for "Asynchronous Batch Transcription" test: Scored batch run for the bilingual input: status scored, 69.22s latency, RTF 0.03571, cost $0.2154, and 5628 returned words against the reference transcript., 07-automation-trace-3.png
Scored batch run for the bilingual input: status scored, 69.22s latency, RTF 0.03571, cost $0.2154, and 5628 returned words against the reference transcript.
audio
0:00 / 0:00
Loading audio...
Medical Jargon — dense anatomical vocabulary from Gray's Anatomy.
image
Output artifact for "Asynchronous Batch Transcription" test: The medical-jargon sample scored 3.01% WER and 100.0% jargon recall; the transcript preserved the scored anatomy terms rather than dropping them., 04-transcript-detail-2.png
The medical-jargon sample scored 3.01% WER and 100.0% jargon recall; the transcript preserved the scored anatomy terms rather than dropping them.
audio
0:00 / 0:00
Loading audio...
INPUT — crosstalk.wav: 35:43 four-way overlapping meeting audio (AMI EN2002a).
OUTPUT
raw-response.json
Loading file...
Batch transcript returned for the crosstalk run; WER 26.63% and wall-clock 46.71s.
audio
0:00 / 0:00
Loading audio...
INPUT — medical_terms.mp3: 18:44 clean single-speaker medical narration (Gray's Anatomy via LibriVox).
OUTPUT
raw-response-2.json
Loading file...
Batch transcript returned for the medical jargon run; WER 3.01% and wall-clock 21.71s.
audio
0:00 / 0:00
Loading audio...
INPUT — mix_language.mp3: 32:18 Spanish-English conversation, mostly English.
OUTPUT
raw-response-3.json
Loading file...
Batch transcript returned for the bilingual run; WER 25.06% and wall-clock 69.22s.
audio
0:00 / 0:00
Loading audio...
Input-1: Overlapping Speech / Crosstalk — crosstalk.wav, 35:43, four-way overlapping meeting audio.
image
Output artifact for "Asynchronous Batch Transcription" test: WER 26.63% on the crosstalk run, with 507 substitutions, 1430 deletions, and 81 insertions against a 7579-word reference. It was the best of 8 engines on this input, but the largest divergence omitted a long overlapping span., 03-terminal-metrics-input-1.png
WER 26.63% on the crosstalk run, with 507 substitutions, 1430 deletions, and 81 insertions against a 7579-word reference. It was the best of 8 engines on this input, but the largest divergence omitted a long overlapping span.
audio
0:00 / 0:00
Loading audio...
Input-2: Medical Jargon — medical_terms.mp3, 18:44, Gray's Anatomy narration dense with anatomical terms.
image
Output artifact for "Asynchronous Batch Transcription" test: WER 3.01% on the medical-jargon run, with 49 substitutions, 17 deletions, and 16 insertions against a 2728-word reference. It was second of 10 and kept jargon recall at 100%., 03-terminal-metrics-input-2.png
WER 3.01% on the medical-jargon run, with 49 substitutions, 17 deletions, and 16 insertions against a 2728-word reference. It was second of 10 and kept jargon recall at 100%.
audio
0:00 / 0:00
Loading audio...
Input-3: Bilingual Code-Switching — mix_language.mp3, 32:18, spontaneous Spanish-English conversation.
image
Output artifact for "Asynchronous Batch Transcription" test: WER 25.06% on the bilingual run, with 560 substitutions, 981 deletions, and 92 insertions against a 6517-word reference. Spanish token recall was only 12.5%, so this run does not support balanced code-switching claims., 03-terminal-metrics-input-3.png
WER 25.06% on the bilingual run, with 560 substitutions, 981 deletions, and 92 insertions against a 6517-word reference. Spanish token recall was only 12.5%, so this run does not support balanced code-switching claims.
Bottom Line
Stable batch execution across all three inputs. The run completed end to end, but per-call timings were not instrumented, so the exact HTTP call count is documented rather than measured here.
From our researchTranscribe Audio Accurately — Speech-to-Text Engine Benchmarkearlier research
Transcript Metadata Export
Consistently present across all three runs.
Test Summary
Feature tested: Transcript Metadata Export
Result: Partial — Consistently present across all three runs.

Feature tested: Transcript Metadata Export

Result: Partial

Verdict: Consistently present across all three runs.

Expected behavior: Returns transcript payloads with developer-facing metadata such as word-level timestamps, confidence values, speaker labels, transcription_config, language_pack_info, and timed_tokens. The member cards show those structured fields present across the observed runs.

Test case: Audio file → Text/code file

Input type: Audio file

Input used: Input artifact (Audio file): INPUT — crosstalk.wav: 35:43 four-way overlapping meeting audio (AMI EN2002a). — crosstalk.wav

Observed output: Output artifact (Text/code file): JSON payload includes timed tokens, confidence, and speaker labels; 7,812 timed tokens on the crosstalk run. — raw-response.json

Input artifact: Input artifact (Audio file): INPUT — crosstalk.wav: 35:43 four-way overlapping meeting audio (AMI EN2002a). — crosstalk.wav

Output artifact: Output artifact (Text/code file): JSON payload includes timed tokens, confidence, and speaker labels; 7,812 timed tokens on the crosstalk run. — raw-response.json

What changed: Audio file transformed into Text/code file

Test case: Audio file → Text/code file

Input type: Audio file

Input used: Input artifact (Audio file): INPUT — medical_terms.mp3: 18:44 clean single-speaker medical narration (Gray's Anatomy via LibriVox). — medical_terms.mp3

Observed output: Output artifact (Text/code file): JSON payload includes the same metadata on the medical run; 3,056 timed tokens. — raw-response-2.json

Input artifact: Input artifact (Audio file): INPUT — medical_terms.mp3: 18:44 clean single-speaker medical narration (Gray's Anatomy via LibriVox). — medical_terms.mp3

Output artifact: Output artifact (Text/code file): JSON payload includes the same metadata on the medical run; 3,056 timed tokens. — raw-response-2.json

What changed: Audio file transformed into Text/code file

Test case: Audio file → Text/code file

Input type: Audio file

Input used: Input artifact (Audio file): INPUT — mix_language.mp3: 32:18 Spanish-English conversation, mostly English. — mix_language.mp3

Observed output: Output artifact (Text/code file): JSON payload includes the same metadata on the bilingual run; 6,907 timed tokens. — raw-response-3.json

Input artifact: Input artifact (Audio file): INPUT — mix_language.mp3: 32:18 Spanish-English conversation, mostly English. — mix_language.mp3

Output artifact: Output artifact (Text/code file): JSON payload includes the same metadata on the bilingual run; 6,907 timed tokens. — raw-response-3.json

What changed: Audio file transformed into Text/code file

Test case: Audio file → Image

Input type: Audio file

Input used: Input artifact (Audio file): Input-1: Overlapping Speech / Crosstalk — crosstalk.wav, 35:43, four-way overlapping meeting audio. — crosstalk.wav

Observed output: Output artifact (Image): Raw API response showing transcription_config with operating_point enhanced, diarization speaker, language en, and word-level output with confidence and speaker fields. — 02-response-raw-input-1.png

Input artifact: Input artifact (Audio file): Input-1: Overlapping Speech / Crosstalk — crosstalk.wav, 35:43, four-way overlapping meeting audio. — crosstalk.wav

Output artifact: Output artifact (Image): Raw API response showing transcription_config with operating_point enhanced, diarization speaker, language en, and word-level output with confidence and speaker fields. — 02-response-raw-input-1.png

What changed: Audio file transformed into Image

Test case: Audio file → Image

Input type: Audio file

Input used: Input artifact (Audio file): Input-2: Medical Jargon — medical_terms.mp3, 18:44, Gray's Anatomy narration dense with anatomical terms. — medical_terms.mp3

Observed output: Output artifact (Image): Raw API response showing the same transcription configuration and payload structure, including confidence and speaker fields in the JSON transcript. — 02-response-raw-input-2.png

Input artifact: Input artifact (Audio file): Input-2: Medical Jargon — medical_terms.mp3, 18:44, Gray's Anatomy narration dense with anatomical terms. — medical_terms.mp3

Output artifact: Output artifact (Image): Raw API response showing the same transcription configuration and payload structure, including confidence and speaker fields in the JSON transcript. — 02-response-raw-input-2.png

What changed: Audio file transformed into Image

Test case: Audio file → Image

Input type: Audio file

Input used: Input artifact (Audio file): Input-3: Bilingual Code-Switching — mix_language.mp3, 32:18, spontaneous Spanish-English conversation. — mix_language.mp3

Observed output: Output artifact (Image): Raw API response showing timed_tokens 6907, confidence yes, speaker_labels yes, and distinct speakers 3 in the developer-feature summary. — 02-response-raw-input-3.png

Input artifact: Input artifact (Audio file): Input-3: Bilingual Code-Switching — mix_language.mp3, 32:18, spontaneous Spanish-English conversation. — mix_language.mp3

Output artifact: Output artifact (Image): Raw API response showing timed_tokens 6907, confidence yes, speaker_labels yes, and distinct speakers 3 in the developer-feature summary. — 02-response-raw-input-3.png

What changed: Audio file transformed into Image

Why it matters / Conclusion: Timing metadata was consistently present across all three runs, even when transcription quality varied.

Returns transcript payloads with developer-facing metadata such as word-level timestamps, confidence values, speaker labels, transcription_config, language_pack_info, and timed_tokens. The member cards show those structured fields present across the observed runs.

audio
0:00 / 0:00
Loading audio...
INPUT — crosstalk.wav: 35:43 four-way overlapping meeting audio (AMI EN2002a).
OUTPUT
raw-response.json
Loading file...
JSON payload includes timed tokens, confidence, and speaker labels; 7,812 timed tokens on the crosstalk run.
audio
0:00 / 0:00
Loading audio...
INPUT — medical_terms.mp3: 18:44 clean single-speaker medical narration (Gray's Anatomy via LibriVox).
OUTPUT
raw-response-2.json
Loading file...
JSON payload includes the same metadata on the medical run; 3,056 timed tokens.
audio
0:00 / 0:00
Loading audio...
INPUT — mix_language.mp3: 32:18 Spanish-English conversation, mostly English.
OUTPUT
raw-response-3.json
Loading file...
JSON payload includes the same metadata on the bilingual run; 6,907 timed tokens.
audio
0:00 / 0:00
Loading audio...
Input-1: Overlapping Speech / Crosstalk — crosstalk.wav, 35:43, four-way overlapping meeting audio.
image
Output artifact for "Transcript Metadata Export" test: Raw API response showing transcription_config with operating_point enhanced, diarization speaker, language en, and word-level output with confidence and speaker fields., 02-response-raw-input-1.png
Raw API response showing transcription_config with operating_point enhanced, diarization speaker, language en, and word-level output with confidence and speaker fields.
audio
0:00 / 0:00
Loading audio...
Input-2: Medical Jargon — medical_terms.mp3, 18:44, Gray's Anatomy narration dense with anatomical terms.
image
Output artifact for "Transcript Metadata Export" test: Raw API response showing the same transcription configuration and payload structure, including confidence and speaker fields in the JSON transcript., 02-response-raw-input-2.png
Raw API response showing the same transcription configuration and payload structure, including confidence and speaker fields in the JSON transcript.
audio
0:00 / 0:00
Loading audio...
Input-3: Bilingual Code-Switching — mix_language.mp3, 32:18, spontaneous Spanish-English conversation.
image
Output artifact for "Transcript Metadata Export" test: Raw API response showing timed_tokens 6907, confidence yes, speaker_labels yes, and distinct speakers 3 in the developer-feature summary., 02-response-raw-input-3.png
Raw API response showing timed_tokens 6907, confidence yes, speaker_labels yes, and distinct speakers 3 in the developer-feature summary.
Bottom Line
Timing metadata was consistently present across all three runs, even when transcription quality varied.
From our researchTranscribe Audio Accurately — Speech-to-Text Engine Benchmarkearlier research
Speaker Diarization
Labels present; attribution unverified
Test Summary
Feature tested: Speaker Diarization
Result: Partial — Labels present; attribution unverified

Feature tested: Speaker Diarization

Result: Partial

Verdict: Labels present; attribution unverified

Expected behavior: Labels speakers in the transcript output. The member cards exercised crosstalk meeting audio and observed speaker tags/labels across runs, including a multi-speaker conversation.

Test case: Audio file → Image

Input type: Audio file

Input used: Input artifact (Audio file): Input-1: Overlapping Speech / Crosstalk — crosstalk.wav, 35:43, four-way overlapping meeting audio. — crosstalk.wav

Observed output: Output artifact (Image): The transcript detail panel reports diarization detected 5 labels, while the reference meeting has 4 participants, so the run is over-segmented even though attribution correctness was not scored. — 04-transcript-detail-input-1.png

Input artifact: Input artifact (Audio file): Input-1: Overlapping Speech / Crosstalk — crosstalk.wav, 35:43, four-way overlapping meeting audio. — crosstalk.wav

Output artifact: Output artifact (Image): The transcript detail panel reports diarization detected 5 labels, while the reference meeting has 4 participants, so the run is over-segmented even though attribution correctness was not scored. — 04-transcript-detail-input-1.png

What changed: Audio file transformed into Image

Test case: Audio file → Image

Input type: Audio file

Input used: Input artifact (Audio file): Input-2: Medical Jargon — medical_terms.mp3, 18:44, Gray's Anatomy narration dense with anatomical terms. — medical_terms.mp3

Observed output: Output artifact (Image): The raw JSON response includes speaker fields in the returned alternatives, confirming that speaker labels are available in the payload. — 02-response-raw-input-2.png

Input artifact: Input artifact (Audio file): Input-2: Medical Jargon — medical_terms.mp3, 18:44, Gray's Anatomy narration dense with anatomical terms. — medical_terms.mp3

Output artifact: Output artifact (Image): The raw JSON response includes speaker fields in the returned alternatives, confirming that speaker labels are available in the payload. — 02-response-raw-input-2.png

What changed: Audio file transformed into Image

Test case: Audio file → Image

Input type: Audio file

Input used: Input artifact (Audio file): Input-3: Bilingual Code-Switching — mix_language.mp3, 32:18, spontaneous Spanish-English conversation. — mix_language.mp3

Observed output: Output artifact (Image): The raw API response's developer-feature summary shows speaker_labels yes and distinct speakers 3, confirming diarization metadata is present in the output. — 02-response-raw-input-3.png

Input artifact: Input artifact (Audio file): Input-3: Bilingual Code-Switching — mix_language.mp3, 32:18, spontaneous Spanish-English conversation. — mix_language.mp3

Output artifact: Output artifact (Image): The raw API response's developer-feature summary shows speaker_labels yes and distinct speakers 3, confirming diarization metadata is present in the output. — 02-response-raw-input-3.png

What changed: Audio file transformed into Image

Test case: Audio file → Image

Input type: Audio file

Input used: Input artifact (Audio file): Input-2: Medical Jargon diarization test. — medical_terms.mp3

Observed output: Output artifact (Image): Raw response preview reports one distinct speaker label on the medical-jargon narration. — 02-response-raw-2.png

Input artifact: Input artifact (Audio file): Input-2: Medical Jargon diarization test. — medical_terms.mp3

Output artifact: Output artifact (Image): Raw response preview reports one distinct speaker label on the medical-jargon narration. — 02-response-raw-2.png

What changed: Audio file transformed into Image

Test case: Audio file → Text/code file

Input type: Audio file

Input used: Input artifact (Audio file): INPUT-2: Medical Jargon — medical_terms.mp3, 8.58 MB, 1123.944 s, mp3 22.05 kHz mono. — medical_terms.mp3

Observed output: Output artifact (Text/code file): The transcript metadata exposed one speaker label on the single-speaker narration input. — raw-response-2.json

Input artifact: Input artifact (Audio file): INPUT-2: Medical Jargon — medical_terms.mp3, 8.58 MB, 1123.944 s, mp3 22.05 kHz mono. — medical_terms.mp3

Output artifact: Output artifact (Text/code file): The transcript metadata exposed one speaker label on the single-speaker narration input. — raw-response-2.json

What changed: Audio file transformed into Text/code file

Why it matters / Conclusion: Useful for downstream speaker-aware workflows, but this benchmark did not verify true diarization accuracy.

Labels speakers in the transcript output. The member cards exercised crosstalk meeting audio and observed speaker tags/labels across runs, including a multi-speaker conversation.

audio
0:00 / 0:00
Loading audio...
Input-1: Overlapping Speech / Crosstalk — crosstalk.wav, 35:43, four-way overlapping meeting audio.
image
Output artifact for "Speaker Diarization" test: The transcript detail panel reports diarization detected 5 labels, while the reference meeting has 4 participants, so the run is over-segmented even though attribution correctness was not scored., 04-transcript-detail-input-1.png
The transcript detail panel reports diarization detected 5 labels, while the reference meeting has 4 participants, so the run is over-segmented even though attribution correctness was not scored.
audio
0:00 / 0:00
Loading audio...
Input-2: Medical Jargon — medical_terms.mp3, 18:44, Gray's Anatomy narration dense with anatomical terms.
image
Output artifact for "Speaker Diarization" test: The raw JSON response includes speaker fields in the returned alternatives, confirming that speaker labels are available in the payload., 02-response-raw-input-2.png
The raw JSON response includes speaker fields in the returned alternatives, confirming that speaker labels are available in the payload.
audio
0:00 / 0:00
Loading audio...
Input-3: Bilingual Code-Switching — mix_language.mp3, 32:18, spontaneous Spanish-English conversation.
image
Output artifact for "Speaker Diarization" test: The raw API response's developer-feature summary shows speaker_labels yes and distinct speakers 3, confirming diarization metadata is present in the output., 02-response-raw-input-3.png
The raw API response's developer-feature summary shows speaker_labels yes and distinct speakers 3, confirming diarization metadata is present in the output.
audio
0:00 / 0:00
Loading audio...
Input-2: Medical Jargon diarization test.
image
Output artifact for "Speaker Diarization" test: Raw response preview reports one distinct speaker label on the medical-jargon narration., 02-response-raw-2.png
Raw response preview reports one distinct speaker label on the medical-jargon narration.
audio
0:00 / 0:00
Loading audio...
INPUT-2: Medical Jargon — medical_terms.mp3, 8.58 MB, 1123.944 s, mp3 22.05 kHz mono.
OUTPUT
raw-response-2.json
Loading file...
The transcript metadata exposed one speaker label on the single-speaker narration input.
Bottom Line
Useful for downstream speaker-aware workflows, but this benchmark did not verify true diarization accuracy.
From our researchTranscribe Audio Accurately — Speech-to-Text Engine Benchmark
Audio Transcription
Weak as configured
Test Summary
Feature tested: Audio Transcription
Result: Partial — Weak as configured

Feature tested: Audio Transcription

Result: Partial

Verdict: Weak as configured

Expected behavior: Transcribes speech audio into a transcript, including hard English narration and Spanish-English mixed-language audio. The member cards cover a hard-English overlap test, a jargon-heavy lecture, and a bilingual code-switching sample.

Test case: Audio file → Image

Input type: Audio file

Input used: Input artifact (Audio file): Input-3: Bilingual Code-Switching — mix_language.mp3, 32:18, spontaneous Spanish-English conversation. — mix_language.mp3

Observed output: Output artifact (Image): The largest error site dropped the Spanish token 'ahora'; Spanish token recall was only 12.5% (10/80), so this is not balanced code-switching evidence. — 04-transcript-detail-input-3.png

Input artifact: Input artifact (Audio file): Input-3: Bilingual Code-Switching — mix_language.mp3, 32:18, spontaneous Spanish-English conversation. — mix_language.mp3

Output artifact: Output artifact (Image): The largest error site dropped the Spanish token 'ahora'; Spanish token recall was only 12.5% (10/80), so this is not balanced code-switching evidence. — 04-transcript-detail-input-3.png

What changed: Audio file transformed into Image

Why it matters / Conclusion: Do not treat this run as proof of balanced multilingual performance; it is still mostly English and performs poorly on Spanish tokens.

Transcribes speech audio into a transcript, including hard English narration and Spanish-English mixed-language audio. The member cards cover a hard-English overlap test, a jargon-heavy lecture, and a bilingual code-switching sample.

audio
0:00 / 0:00
Loading audio...
Input-3: Bilingual Code-Switching — mix_language.mp3, 32:18, spontaneous Spanish-English conversation.
image
Output artifact for "Audio Transcription" test: The largest error site dropped the Spanish token 'ahora'; Spanish token recall was only 12.5% (10/80), so this is not balanced code-switching evidence., 04-transcript-detail-input-3.png
The largest error site dropped the Spanish token 'ahora'; Spanish token recall was only 12.5% (10/80), so this is not balanced code-switching evidence.
Bottom Line
Do not treat this run as proof of balanced multilingual performance; it is still mostly English and performs poorly on Spanish tokens.
From our researchTranscribe Audio Accurately — Speech-to-Text Engine Benchmark

How it scored on the research's own criteria

The 3 evaluation dimensions from our hands-on research on Speechmatics, each judged from recorded runs on 3 test inputs — the same verdicts the ranking page ranks on.

held up  partial  failed  not exercised by this input

CriterionVerdictWhat the runs showedPer inputProof
Output qualityMixed3/5Quality is uneven rather than consistently good or bad: it is excellent on dense medical narration, but it loses a lot of content on crosstalk and especially on Spanish-English code-switching. That mix lands in the middle overall, with one strong result offset by two clearly weaker ones.open proof ↗
Automation levelStrong5/5The workflow ran from upload to finished result on every test without manual intervention. Even though the vendor protocol is multi-stage, the observed runs were hands-off and completed successfully each time.
Input handlingStrong4/5It handled all three uploads cleanly and processed them quickly, so the core input path is solid. I did not give it a 5 because the runs are consistently on the Enhanced tier, so the workflow is fast but not especially cheap compared with the lowest batch option.open proof ↗

Verdicts come verbatim from the study's recorded observations, never re-derived at render; a criterion with no recorded run shows Not exercised — this section cannot invent a score.

Official pricing

Pro includes model-tiered batch rates; the benchmarked configuration used Batch Enhanced.

Free
$0 — $100 in credit
No credit card required; 2 concurrent real-time sessions; 1 batch job/sec.
TESTED
Pro
from $0.129/hr
Batch is tiered by model. The benchmarked configuration used Batch Enhanced at $0.40/hr ($0.006667/min), while Batch Melia 1 is the headline $0.129/hr batch price.
Enterprise
Custom
No rate limits; custom models; on-prem/container/virtual appliance/on-device; requires sales contact.

Official pricing page last updated 31 July 2026. The benchmarked run used the Enhanced batch rate, not the cheapest Batch Melia 1 headline rate. The page also lists model-training opt-in and volume discounts; the figures here are list prices.

✓ Use This If
You need batch speech-to-text that returns word timings, confidence values, and speaker labels.
You are comparing engines on hard English audio such as overlapping speakers or jargon-heavy narration.
You can accept batch processing rather than live streaming.
✕ Skip This If
You need proven balanced multilingual or code-switching performance; this run was mostly English and Spanish recall was poor.
You need measured speaker-attribution accuracy rather than just detected labels and counts.
You need streaming latency results; this benchmark was batch-only.
audio-speechother-audio-speechtextOther
On the crosstalk input, Speechmatics scored WER 26.63% with 507 substitutions, 1430 deletions, and 81 insertions against a 7579-word reference. That was the best score among the engines scored on that input, but it still missed a long overlapping span and over-segmented the meeting into 5 speaker labels for a 4-participant reference.
Very well. On the medical-jargon input it scored WER 3.01% with 49 substitutions, 17 deletions, and 16 insertions against a 2728-word reference, and it recalled all scored jargon terms for 100% jargon recall.
Poorly as configured. The bilingual run had WER 25.06%, but the benchmark notes that the audio was still about 95.5% English overall, and Spanish token recall was only 12.5% (10 of 80 types).
Yes. The raw response and developer-feature summaries show word-level timing, confidence fields, and speaker labels in the payload.
No. It only verified that speaker labels were present and counted them. The report explicitly says attribution correctness was not measured, so diarization accuracy remains unverified.
No. This was a batch benchmark. It measured wall-clock latency and real-time factor for batch jobs, but not live streaming latency.
The benchmarked configuration used the Enhanced batch rate at $0.40 per audio-hour ($0.006667/min), not the cheapest Batch Melia 1 headline rate of $0.129/hr. The official pricing page also lists Free and Enterprise tiers, plus model-tiered batch rates.

Banner Preview

How the embed badge will look on your site

Speechmatics featured on AI Demos

Embed HTML

Copy this code to your website source

<a target="_blank" href="https://aidemos.com/tools/speechmatics?utm_source=speechmatics_embed" style="width: 250px; height: 80px; border-radius:4px;" width="250" height="80"> <img src="https://aidemos-website-images.s3.amazonaws.com/featured.png" alt="Speechmatics | Featured on AI Demos" style="width: 250px; height: 80px; border-radius:4px;" width="250" height="80"> </a>

Quick Integration Guide

  • 1Copy the HTML code block above.
  • 2Paste it into your site's HTML or CMS editor.
  • 3Banner appears instantly on your page.
  • 4Links back to your tool profile here.
Similar Tools

Similar Tools

Discover more AI tools like Speechmatics to enhance your workflow.

Comments (0)

Please Log in to join the discussion.

Built by FutureSmart AI — the team behind AI Demos

Need a custom AI solution for this use case?

If you are looking to build a custom speech-to-text transcription, audio indexing, or speech analytics system for your business or internal workflow, email us at contact@futuresmart.ai.

Get a custom build

Found something inaccurate or missing? We try to keep our AI research accurate and useful. If you found outdated information, an issue, or have a suggestion, email us at collaborate@aidemos.com.

Back to Top