Evidence · first-party tested/Best AI Meeting Notetakers for Accurate Transcripts, Summaries, and Action Items
Generated a timestamped transcript view, and the report says the full transcript was very accurate: nearly all names, tools, jargon, and numbers were captured correctly with no significant misheard terms or hallucinations.
What was measured
Transcription Accuracy
Word accuracy on the shared call, especially names, tools, numbers, and jargon.
decisive for this rankingtransformation
If the transcript gets names, numbers, and jargon wrong, the note-taker has failed at the core job of capturing the call accurately. (3 of 3 judges)
What was given, what came back
Test input: AI Demos Daily Standup — 31 July 2026 · image · group: ai-meeting-notetaker
Input — what we sent

AI Demos Daily Standup — 31 July 2026
A real 25-minute technical engineering daily standup with 14 attendees and about 10 active speakers, used as the single parallel-capture meeting for evaluating AI meeting notetakers on transcription, diarization, summaries, action items, search/chat, and collaboration features.
Why this input is hard
- · Transcription accuracy for real names, tool names, numbers, and technical jargon
- · Speaker diarization across multiple active speakers
- · Robustness to overlapping speech, crosstalk, and rapid turn-taking
- · Join reliability for bot-based and botless capture
- · Summary quality on identical source material
- · Action-item extraction with correct owners and commitments
- · Topic segmentation of standup updates
- · Search and chat grounded in the meeting content
- · Sharing, API, MCP, integrations, plan limits, languages, and privacy feature coverage
Output — unretouched


Also checked on this input — same tool, 7 other criteria
Action-Item Extraction✓ WorkedGrouped action items by owner, attributed them to the correct team member, and exposed a clickable source timestamp (19:13) for at least one item.Chat with Notes / Ask Questions✓ WorkedAskFred answered a natural-language question with a specific grounded response ('August 6th') and relevant context, with no hallucination reported in the tested query.Join Method & Reliability✓ WorkedUsed a bot-based Google Meet join: the Fireflies notetaker appeared in the People panel, and the recording player showed it still present in a 23:40 capture, matching the report’s claim of uninterrupted full-call capture.Search Across Notes✓ WorkedTranscript search worked with exact-match retrieval: a Ctrl+F query for 'API' returned 1/1 match at 20:42 with the hit highlighted and a clickable timestamp.Speaker Diarization✓ WorkedAttributed consecutive turns to distinct speakers in the transcript, and the report says speaker identification was almost complete with only minor attribution errors.Summary Quality✓ WorkedProduced a structured notes summary with a named header ('Task Status and Issue Resolution') rather than a blob, and the report says the full summary was multi-section and did not drop important points.Topic Segmentation✓ WorkedBroke the meeting notes into named sections with descriptive headers and short recap paragraphs, making the output skimmable instead of one undifferentiated block.
Provenance
- Observation
- bac222c7-f7cf-4041-8449-df5aa22c88ac
- Evidence run
- ace58582-3d1e-48ee-996c-9b3cd03f27a2
- Study
- AI Meeting Notetakers — Capture Accurate Transcripts, Summaries & Action Items From Live Calls
- Research task
- 86baxegnv
- Tested at
- not recorded
- Source
- first-party
- Evidence state
- verified
- Proof shown
- input + output shown
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "fireflies-ai",
scenario: "ai-meeting-notetaker"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 7 other tools
measured on Transcription Accuracy
Fathom◐ MixedFathom's transcript mostly preserves the meeting's names and technical content, but the report records one confirmed name-level error: "Mahreen" was rendered as "Meryl."Fellow✓ WorkedThe transcript was near-clean: the tool captured nearly all names, technical jargon, and numbers correctly, with no significant misheard terms or hallucinations observed in the tested meeting.Granola✗ FailedOn this 25-minute, multi-speaker standup, Granola’s transcript quality is unreliable: the published excerpt shows garbled phrasing and mistranscribed wording, and the report says the mishearing pattern recurs across early, middle, and late sections rather than being isolated to one moment.HappyScribe◐ MixedOn this ~25-minute multi-speaker standup, HappyScribe captured the vast majority of names, tools, and jargon correctly, and the report records only 1–2 misheard words.MeetGeek✓ WorkedIt transcribes a normal ~25-minute, ~10-active-speaker engineering standup mostly accurately, with only minor proper-noun/term drift noted in the report; one example given is "Madin" being misheard for "Mahreen".Notta✓ WorkedNotta’s transcript capture was accurate on the evaluated standup: the report says it correctly captured names, tool names, numbers, and engineering jargon with no significant word-level errors, silent hallucinations, or misheard terms.Otter.ai✓ WorkedOtter generated a full transcript for the standup and, per the report, captured names, tool names, jargon, and numbers correctly with minimal errors, making the transcript reliable for reference.
This evidence is published in
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com