Deepgram  icon
audio-speech

Deepgram

Batch speech-to-text with word-level metadata and speaker labels, but weak on crosstalk and code-switching.

Visit Deepgram
Word timestampsSpeaker labelsJargon recall 100%Spanish recall 3.8%
TL;DR — our verdictUpdated September 2026 · 9 test artifacts

Good batch STT metadata, uneven hard-case accuracy

Where it wins
  • you need a batch STT API that returns word-level timing, confidence, and speaker labels
  • you mainly transcribe single-language technical narration and care about jargon recall
  • you care more about batch throughput than live streaming
Main limitation
  • you need reliable overlapping-speech handling for crosstalk-heavy meetings
Pricing (verified plans)
Pay As You Go $200 free credit, then pay-as-you-goGrowth $4,000+ / year prepaid creditsEnterprise Requires sales contact
Strongest test artifacts

Our take

Deepgram Nova-3 kept a stable developer payload and was excellent on the medical-jargon clip, but crosstalk and bilingual code-switching both produced high WER and lots of insertions and deletions. As configured here, it looks like a solid batch transcription API for mostly monolingual technical audio, not a safe default for overlapping meetings or Spanish-English conversation.

Tutorial recording from the benchmark task

In-Depth Review

Our detailed analysis of Deepgram — features, performance, and real-world testing.

AD
AI Demos Team
Expert Reviewer
Verified Review

Feature-by-Feature Breakdown

Batch Audio Transcription
Mixed: strong on jargon, weak on overlap and code-switching.
Test Summary
Feature tested: Batch Audio Transcription
Result: Failed — Mixed: strong on jargon, weak on overlap and code-switching.

Feature tested: Batch Audio Transcription

Result: Failed

Verdict: Mixed: strong on jargon, weak on overlap and code-switching.

Expected behavior: Converts uploaded pre-recorded audio into text in a single batch request. The member cards exercised it on overlapping meetings, medical-jargon narration, Spanish-English code-switching, and other prerecorded clips.

Test case: Audio file → Text/code file

Input type: Audio file

Input used: Input artifact (Audio file): Overlapping Speech / Crosstalk input audio — crosstalk.wav

Observed output: Output artifact (Text/code file): Returned the crosstalk transcript for the 65.39 MB meeting clip; the run scored 36.27% WER, returned 6,946 words against 7,579 reference words, and exposed word-level timing, confidence, and speaker labels. — raw-response.json

Input artifact: Input artifact (Audio file): Overlapping Speech / Crosstalk input audio — crosstalk.wav

Output artifact: Output artifact (Text/code file): Returned the crosstalk transcript for the 65.39 MB meeting clip; the run scored 36.27% WER, returned 6,946 words against 7,579 reference words, and exposed word-level timing, confidence, and speaker labels. — raw-response.json

What changed: Audio file transformed into Text/code file

Test case: Audio file → Text/code file

Input type: Audio file

Input used: Input artifact (Audio file): Medical Jargon input audio — medical_terms.mp3

Observed output: Output artifact (Text/code file): Returned the medical-jargon transcript for the 8.58 MB narration clip; the run scored 5.43% WER, returned 2,726 words against 2,728 reference words, and achieved 100.0% jargon recall. — raw-response-2.json

Input artifact: Input artifact (Audio file): Medical Jargon input audio — medical_terms.mp3

Output artifact: Output artifact (Text/code file): Returned the medical-jargon transcript for the 8.58 MB narration clip; the run scored 5.43% WER, returned 2,726 words against 2,728 reference words, and achieved 100.0% jargon recall. — raw-response-2.json

What changed: Audio file transformed into Text/code file

Test case: Audio file → Text/code file

Input type: Audio file

Input used: Input artifact (Audio file): Bilingual Code-Switching input audio — mix_language.mp3

Observed output: Output artifact (Text/code file): Returned the bilingual code-switching transcript; the run scored 38.13% WER, returned 5,691 words against 6,517 reference words, and only recalled 3 of 80 Spanish token types. — raw-response-3.json

Input artifact: Input artifact (Audio file): Bilingual Code-Switching input audio — mix_language.mp3

Output artifact: Output artifact (Text/code file): Returned the bilingual code-switching transcript; the run scored 38.13% WER, returned 5,691 words against 6,517 reference words, and only recalled 3 of 80 Spanish token types. — raw-response-3.json

What changed: Audio file transformed into Text/code file

Why it matters / Conclusion: Good on the medical narration clip, but not dependable on crosstalk or Spanish-English code-switching as configured here.

Converts uploaded pre-recorded audio into text in a single batch request. The member cards exercised it on overlapping meetings, medical-jargon narration, Spanish-English code-switching, and other prerecorded clips.

audio
0:00 / 0:00
Loading audio...
Overlapping Speech / Crosstalk input audio
file
raw-response.json
Loading file...
Returned the crosstalk transcript for the 65.39 MB meeting clip; the run scored 36.27% WER, returned 6,946 words against 7,579 reference words, and exposed word-level timing, confidence, and speaker labels.
audio
0:00 / 0:00
Loading audio...
Medical Jargon input audio
file
raw-response-2.json
Loading file...
Returned the medical-jargon transcript for the 8.58 MB narration clip; the run scored 5.43% WER, returned 2,726 words against 2,728 reference words, and achieved 100.0% jargon recall.
audio
0:00 / 0:00
Loading audio...
Bilingual Code-Switching input audio
file
raw-response-3.json
Loading file...
Returned the bilingual code-switching transcript; the run scored 38.13% WER, returned 5,691 words against 6,517 reference words, and only recalled 3 of 80 Spanish token types.
Bottom Line
Good on the medical narration clip, but not dependable on crosstalk or Spanish-English code-switching as configured here.
From our researchTranscribe Audio Accurately — Speech-to-Text Engine Benchmarkearlier research
Structured Transcript Output
Consistent developer payload across all three runs.
Test Summary
Feature tested: Structured Transcript Output
Result: Passed — Consistent developer payload across all three runs.

Feature tested: Structured Transcript Output

Result: Passed

Verdict: Consistent developer payload across all three runs.

Expected behavior: Returns transcript results as structured JSON with fields like word-level timestamps, confidence values, punctuation, and speaker labels. The member cards exercised this across multiple runs and options such as smart_format, diarize, punctuate, and utterance.

Test case: Audio file → Image

Input type: Audio file

Input used: Input artifact (Audio file): Overlapping Speech / Crosstalk input audio — crosstalk.wav

Observed output: Output artifact (Image): Raw response preview shows word_timestamps yes, confidence yes, speaker_labels yes, 13,555 timed tokens, and 4 distinct speakers. — 02-response-raw-input-1.png

Input artifact: Input artifact (Audio file): Overlapping Speech / Crosstalk input audio — crosstalk.wav

Output artifact: Output artifact (Image): Raw response preview shows word_timestamps yes, confidence yes, speaker_labels yes, 13,555 timed tokens, and 4 distinct speakers. — 02-response-raw-input-1.png

What changed: Audio file transformed into Image

Test case: Audio file → Image

Input type: Audio file

Input used: Input artifact (Audio file): Medical Jargon input audio — medical_terms.mp3

Observed output: Output artifact (Image): Raw response preview shows the same metadata shape on the medical clip, including word_timestamps, confidence, and speaker_labels. — 02-response-raw-input-2.png

Input artifact: Input artifact (Audio file): Medical Jargon input audio — medical_terms.mp3

Output artifact: Output artifact (Image): Raw response preview shows the same metadata shape on the medical clip, including word_timestamps, confidence, and speaker_labels. — 02-response-raw-input-2.png

What changed: Audio file transformed into Image

Test case: Audio file → Image

Input type: Audio file

Input used: Input artifact (Audio file): Bilingual Code-Switching input audio — mix_language.mp3

Observed output: Output artifact (Image): Raw response preview shows the same metadata shape on the bilingual clip, including word_timestamps, confidence, speaker_labels, and 11,684 timed tokens. — 02-response-raw-input-3.png

Input artifact: Input artifact (Audio file): Bilingual Code-Switching input audio — mix_language.mp3

Output artifact: Output artifact (Image): Raw response preview shows the same metadata shape on the bilingual clip, including word_timestamps, confidence, speaker_labels, and 11,684 timed tokens. — 02-response-raw-input-3.png

What changed: Audio file transformed into Image

Why it matters / Conclusion: This is the most consistent part of the product: the JSON shape stayed stable across all three inputs.

Returns transcript results as structured JSON with fields like word-level timestamps, confidence values, punctuation, and speaker labels. The member cards exercised this across multiple runs and options such as smart_format, diarize, punctuate, and utterance.

audio
0:00 / 0:00
Loading audio...
Overlapping Speech / Crosstalk input audio
image
Output artifact for "Structured Transcript Output" test: Raw response preview shows word_timestamps yes, confidence yes, speaker_labels yes, 13,555 timed tokens, and 4 distinct speakers., 02-response-raw-input-1.png
Raw response preview shows word_timestamps yes, confidence yes, speaker_labels yes, 13,555 timed tokens, and 4 distinct speakers.
audio
0:00 / 0:00
Loading audio...
Medical Jargon input audio
image
Output artifact for "Structured Transcript Output" test: Raw response preview shows the same metadata shape on the medical clip, including word_timestamps, confidence, and speaker_labels., 02-response-raw-input-2.png
Raw response preview shows the same metadata shape on the medical clip, including word_timestamps, confidence, and speaker_labels.
audio
0:00 / 0:00
Loading audio...
Bilingual Code-Switching input audio
image
Output artifact for "Structured Transcript Output" test: Raw response preview shows the same metadata shape on the bilingual clip, including word_timestamps, confidence, speaker_labels, and 11,684 timed tokens., 02-response-raw-input-3.png
Raw response preview shows the same metadata shape on the bilingual clip, including word_timestamps, confidence, speaker_labels, and 11,684 timed tokens.
Bottom Line
This is the most consistent part of the product: the JSON shape stayed stable across all three inputs.
From our researchTranscribe Audio Accurately — Speech-to-Text Engine Benchmarkearlier research
Speaker Diarization
Useful for speaker counts, but attribution correctness is unproven.
Test Summary
Feature tested: Speaker Diarization
Result: Partial — Useful for speaker counts, but attribution correctness is unproven.

Feature tested: Speaker Diarization

Result: Partial

Verdict: Useful for speaker counts, but attribution correctness is unproven.

Expected behavior: Detects multiple speakers and includes speaker labels in the transcript payload. It was exercised on crosstalk and bilingual audio, where the responses surfaced multiple speaker labels.

Test case: Audio file → Image

Input type: Audio file

Input used: Input artifact (Audio file): Overlapping Speech / Crosstalk input audio — crosstalk.wav

Observed output: Output artifact (Image): Transcript detail shows the largest overlap divergence and confirms 4 diarization labels on the crosstalk case. — 04-transcript-detail-input-1.png

Input artifact: Input artifact (Audio file): Overlapping Speech / Crosstalk input audio — crosstalk.wav

Output artifact: Output artifact (Image): Transcript detail shows the largest overlap divergence and confirms 4 diarization labels on the crosstalk case. — 04-transcript-detail-input-1.png

What changed: Audio file transformed into Image

Test case: Audio file → Image

Input type: Audio file

Input used: Input artifact (Audio file): Medical Jargon input audio — medical_terms.mp3

Observed output: Output artifact (Image): Raw response preview shows a single-speaker medical narration run with 1 distinct speaker label. — 02-response-raw-input-2.png

Input artifact: Input artifact (Audio file): Medical Jargon input audio — medical_terms.mp3

Output artifact: Output artifact (Image): Raw response preview shows a single-speaker medical narration run with 1 distinct speaker label. — 02-response-raw-input-2.png

What changed: Audio file transformed into Image

Test case: Audio file → Image

Input type: Audio file

Input used: Input artifact (Audio file): Bilingual Code-Switching input audio — mix_language.mp3

Observed output: Output artifact (Image): Transcript detail shows speaker labels in the code-switching transcript and highlights the dropped Spanish token 'ahora'. — 04-transcript-detail-input-3.png

Input artifact: Input artifact (Audio file): Bilingual Code-Switching input audio — mix_language.mp3

Output artifact: Output artifact (Image): Transcript detail shows speaker labels in the code-switching transcript and highlights the dropped Spanish token 'ahora'. — 04-transcript-detail-input-3.png

What changed: Audio file transformed into Image

Why it matters / Conclusion: Useful for speaker-count metadata, but the benchmark does not prove speaker-attribution accuracy.

Detects multiple speakers and includes speaker labels in the transcript payload. It was exercised on crosstalk and bilingual audio, where the responses surfaced multiple speaker labels.

audio
0:00 / 0:00
Loading audio...
Overlapping Speech / Crosstalk input audio
image
Output artifact for "Speaker Diarization" test: Transcript detail shows the largest overlap divergence and confirms 4 diarization labels on the crosstalk case., 04-transcript-detail-input-1.png
Transcript detail shows the largest overlap divergence and confirms 4 diarization labels on the crosstalk case.
audio
0:00 / 0:00
Loading audio...
Medical Jargon input audio
image
Output artifact for "Speaker Diarization" test: Raw response preview shows a single-speaker medical narration run with 1 distinct speaker label., 02-response-raw-input-2.png
Raw response preview shows a single-speaker medical narration run with 1 distinct speaker label.
audio
0:00 / 0:00
Loading audio...
Bilingual Code-Switching input audio
image
Output artifact for "Speaker Diarization" test: Transcript detail shows speaker labels in the code-switching transcript and highlights the dropped Spanish token 'ahora'., 04-transcript-detail-input-3.png
Transcript detail shows speaker labels in the code-switching transcript and highlights the dropped Spanish token 'ahora'.
Bottom Line
Useful for speaker-count metadata, but the benchmark does not prove speaker-attribution accuracy.
From our researchearlier researchTranscribe Audio Accurately — Speech-to-Text Engine Benchmark

How it scored on the research's own criteria

The 4 evaluation dimensions from our hands-on research on Deepgram , each judged from recorded runs on 3 test inputs — the same verdicts the ranking page ranks on.

held up  partial  failed  not exercised by this input

CriterionVerdictWhat the runs showedPer inputProof
Output qualityWeak2/5It is clearly strong on dense medical narration, but the other two tests show major accuracy trouble: crosstalk sits in the mid-30% WER range, and code-switching collapses with Spanish recall near zero. That pattern makes the transcript quality unreliable outside cleaner monolingual speech.open proof ↗
Automation levelStrong5/5It finishes each test in one POST with no operator touch, so the workflow is fully hands-off. The only caveat is that the trace records the vendor's single-call design rather than counting calls live, but that does not change the fact that the runs complete end to end without intervention.
ExportStrong5/5Every run returns the full developer-facing package: word timing, confidence, and speaker labels together, plus a deep structured payload. That is the kind of output you can plug into downstream tooling without needing to reconstruct transcript metadata yourself.open proof ↗
Input handlingStrong5/5All three uploads were accepted at their full sizes and durations, and each run completed cleanly. That shows it can take the provided long-form audio in batch form without rejecting the file or stopping short.

Verdicts come verbatim from the study's recorded observations, never re-derived at render; a criterion with no recorded run shows Not exercised — this section cannot invent a score.

Reported vendor tiers

The report notes that the served pricing page did not fully confirm the pre-recorded Nova-3 tab, so treat model-specific rates as unverified.

Pay As You Go
$200 free credit, then pay-as-you-go
No minimums, no expiration, no credit card required. STT concurrency up to 50 REST, 150 WSS, 5 Whisper Cloud.
Growth
$4,000+ / year prepaid credits
Save up to 20%; 10% overage fee. STT concurrency up to 50 REST, 225 WSS, 5 Whisper Cloud.
Enterprise
Requires sales contact
Large volume, data/deployment requirements, support; custom models, self-hosted/VPC, SLAs.

Benchmark cost was calculated from list price × measured duration; the task header also reports $0.0063/min ($0.378/audio-hour) for the tested configuration.

✓ Use This If
you need a batch STT API that returns word-level timing, confidence, and speaker labels
you mainly transcribe single-language technical narration and care about jargon recall
you care more about batch throughput than live streaming
✕ Skip This If
you need reliable overlapping-speech handling for crosstalk-heavy meetings
you need strong Spanish-English code-switching or broader multilingual transcription
you need validated diarization attribution rather than just speaker counts
you need measured streaming latency for a live benchmark
audio-speechaudio-to-texttextOther
The runs used a single raw-body POST with audio in the request body and query-string options such as model, smart_format, diarize, punctuate, and utterances. The report says there was no form wrapper, and the vendor docs describe this as a single-call protocol.
Yes. The raw response previews show word_timestamps, confidence, and speaker_labels detected on all three inputs, and the payload depth stayed at 3/3 each time.
On the crosstalk input, it scored 36.27% WER, returned 6,946 words against a 7,579-word reference, and produced 510 insertions and 1,143 deletions. It did return 4 speaker labels, but the report does not prove speaker-attribution correctness.
It did well on the medical-jargon clip: 5.43% WER, 2,726 words returned against 2,728 reference words, and 100.0% jargon recall on the scored terms.
Poorly. On the bilingual clip it scored 38.13% WER, returned 5,691 words against 6,517 reference words, and Spanish token recall was only 3.8% (3 of 80 types).
The benchmark reported $0.22498 on crosstalk, $0.11802 on the medical clip, and $0.20354 on the bilingual clip. Latency ranged from 12.02s to 63.87s, with RTF from 0.01069 to 0.02981; streaming latency was not measured because this was a batch benchmark.

Banner Preview

How the embed badge will look on your site

Deepgram  featured on AI Demos

Embed HTML

Copy this code to your website source

<a target="_blank" href="https://aidemos.com/tools/deepgram?utm_source=deepgram_embed" style="width: 250px; height: 80px; border-radius:4px;" width="250" height="80"> <img src="https://aidemos-website-images.s3.amazonaws.com/featured.png" alt="Deepgram | Featured on AI Demos" style="width: 250px; height: 80px; border-radius:4px;" width="250" height="80"> </a>

Quick Integration Guide

  • 1Copy the HTML code block above.
  • 2Paste it into your site's HTML or CMS editor.
  • 3Banner appears instantly on your page.
  • 4Links back to your tool profile here.
Similar Tools

Similar Tools

Discover more AI tools like Deepgram to enhance your workflow.

🤖
Gladia
Batch STT with rich word-level metadata and strong jargon recall, but overlap handling is weak and bilingual WER needs a mono-downmixed rerun.
AI Tool
🤖
AssemblyAI
Fast batch speech-to-text with rich metadata, strong jargon and mixed-language results, but overlap-heavy meetings can still lose too much.
AI Tool
🤖
Speechmatics
Strong batch STT for hard English audio, but weak on code-switching as configured.
AI Tool
🤖
OpenAI
Batch speech-to-text with word timestamps, but a strict upload cap and weak multilingual performance make it a mixed fit for hard audio.
AI Tool
🤖
Google Cloud Speech-to-Text
Timed batch transcripts for mostly English, jargon-heavy audio — but not for diarization or code-switching.
AI Tool
🤖
AWS Transcribe
Batch speech-to-text with word-level metadata, but accuracy drops on overlap and code-switching.
AI Tool
🤖
ElevenLabs
Natural-sounding voice cloning and narration, but with only approximate voice identity.
AI Tool
🤖
Rev AI
Low-cost batch speech-to-text with structured JSON, word timestamps, and speaker labels, but mixed accuracy on crosstalk and code-switching.
AI Tool
🤖
Groq
AI Tool
🤖
OpenAI Whisper-1
AI Tool
🤖
Google Cloud STT v2
AI Tool
🤖
ElevenLabs Scribe
Fast batch speech-to-text with word-level metadata, strongest on jargon and weaker on overlap/code-switching.
AI Tool
🤖
GroqCloud Whisper Large-v3
AI Tool
🤖
Gladia Solaria
AI Tool
🤖
AssemblyAI Universal
AI Tool
🤖
Speechmatics Ursa Enhanced
AI Tool
🤖
Google Cloud Speech-to-Text Chirp 2
AI Tool
🤖
Microsoft Azure Speech Services
AI Tool

Comments (0)

Please Log in to join the discussion.

Built by FutureSmart AI — the team behind AI Demos

Need a custom AI solution for this use case?

If you are looking to build a custom speech-to-text, audio transcription, or speaker diarization system for your business or internal workflow, email us at contact@futuresmart.ai.

Get a custom build

Found something inaccurate or missing? We try to keep our AI research accurate and useful. If you found outdated information, an issue, or have a suggestion, email us at collaborate@aidemos.com.

Back to Top