AssemblyAI
It accepted the 65.39 MB meeting recording and produced a scored transcript, but the overlap was hard on it and the result landed at 33.16% WER with a large amount of missing speech.
87aa31cbdbfb4c4997ed3805ad8fd206.png
Developers choosing a speech-to-text engine need more than clean-audio demos: they need to know which system holds up on overlapping speakers, technical terms, and code-switching, while still returning timestamps, speaker labels, low latency, and sensible cost. We benchmarked 10 engines on the same three long real-world recordings and compared WER, diarization, timestamp payload depth, runtime, and price.
The strongest overall pick if you need a long-form transcription engine that stays competitive on the hardest inputs and still returns speaker labels and word timestamps.
Accuracy swings from weak on overlapping speech to strong on medical narration, with code-switching landing in the middle, so the overall picture is genuinely mixed rather than consistently good or bad.
We rank on the 1 check that decide whether a tool does this job: Output quality. A check only carries a score when we recorded a finding for it, and a tool has to be measured on all of them to take the top spot. We also checked Automation level, Export, Input handling — compared for you, but not part of the ranking.
Columns, left to right: Output quality
Pick the tools you care about, then compare what they returned or how they scored.
It accepted the 65.39 MB meeting recording and produced a scored transcript, but the overlap was hard on it and the result landed at 33.16% WER with a large amount of missing speech.
87aa31cbdbfb4c4997ed3805ad8fd206.png
It ran the crosstalk file end to end, accepted the large upload, and returned a full transcript payload, but the meeting was only moderately accurate and the speaker count came back inflated.
538de469ee1f4ec98dbd4b93c3d9a27c.png
It took the meeting recording in one upload and returned a full transcript package, but the wording drifted a lot in the noisy overlap, so the result is usable but not tight.
e44c577a896a41ed81b79e2986cb72f3.png
It accepted the long crosstalk WAV, ran the batch job all the way through, and returned a full timed transcript with speaker labels. The transcript itself was shaky, with a high error rate and lots of dropped speech, so the run is useful but not strong overall.
f7e03bf80a8e4884aef9515ac9d05739.png
It rejected the 65.39 MB meeting recording with HTTP 413 before transcription could start, so nothing was returned except the error. The run was fast and explicit about the size cap, but this input was not processed.
fe907e170abe427bb355b1fef95ebc59.png
It refused the 65.39 MB WAV with HTTP 413 after about 50 seconds, so nothing was transcribed for this meeting recording.
3093c5df674d41179041b5987186ba6d.png
It accepted the long meeting recording, ran it hands-off, and returned a rich transcript package, but the transcript was noisy in the overlapping sections and the main diff shows substantial drift.
60060b3cf56047408d1e8220255cb68e.png
It accepted the 65.39 MB crosstalk file, ran on its own, and returned a detailed transcript package, but the transcription was only partly faithful under heavy overlap.
1e38730bc27c4fa4a44124570255061f.png
It accepted the 65.39 MB meeting recording and finished the run with a rich transcript payload, but the transcript was rough on overlapping speech, with 28.33% WER and extra speaker splitting.
6ce27357d73c4f53bde1e1d6610b0ea4.png
It accepted the 65.39 MB, 2142.709 s crosstalk file and finished automatically, but the transcript quality was poor at 43.50% WER and it did not return speaker labels.
ef993bd0e519426dadfbd740e61797cb.png
All 4 recorded checks per tool. Open a tool to inspect every finding.
Accuracy swings from weak on overlapping speech to strong on medical narration, with code-switching landing in the middle, so the overall picture is genuinely mixed rather than consistently good or bad.
Accuracy is weak on overlapping speech: 33.16% WER with 461 substitutions, 1,976 deletions, and 76 insertions against a 7,579-word reference, ranking 4th of 8 on this input.
permalink to this finding →Accuracy is strong on dense medical jargon: 3.78% WER with 71 substitutions, 14 deletions, and 18 insertions against a 2,728-word reference, and the run reports 100.0% jargon recall.
permalink to this finding →Accuracy is middling on code-switching speech: 21.04% WER with 589 substitutions, 652 deletions, and 130 insertions against a 6,517-word reference; Spanish token recall is 72.5% (58/80 types), and the run is best of 10 on this input.
permalink to this finding →It was strong on dense medical jargon, but performance dropped on overlapping speech and was middling on bilingual Spanish-English code-switching.
permalink to this finding →No final take available yet.
The tools we tested for this use case — each card opens its full tested review.
If you are looking to build a custom speech-to-text transcription, speaker diarization, or audio transcription system for your business or internal workflow, email us at contact@futuresmart.ai.
Found something inaccurate or missing? We try to keep our AI research accurate and useful. If you found outdated information, an issue, or have a suggestion, email us at collaborate@aidemos.com.
Comments (0)