Broke the meeting notes into named sections with descriptive headers and short recap paragraphs, making the output skimmable instead of one undifferentiated block.

✓ Worked🧾 artifact-verifiedinput + output shownTest date not recordedFireflies.ai
What was measured
Topic Segmentation

Breaks long multi-topic meetings into useful sections instead of one blob.

context, not decisivetransformation

Breaking long meetings into sections makes notes easier to use, but a tool can still succeed at core note-taking without perfect segmentation. (3 of 3 judges)

What was given, what came back

Test input: AI Demos Daily Standup — 31 July 2026 · image · group: ai-meeting-notetaker
Input — what we sent
AI Demos Daily Standup — 31 July 2026
AI Demos Daily Standup — 31 July 2026

A real 25-minute technical engineering daily standup with 14 attendees and about 10 active speakers, used as the single parallel-capture meeting for evaluating AI meeting notetakers on transcription, diarization, summaries, action items, search/chat, and collaboration features.

Why this input is hard
  • · Transcription accuracy for real names, tool names, numbers, and technical jargon
  • · Speaker diarization across multiple active speakers
  • · Robustness to overlapping speech, crosstalk, and rapid turn-taking
  • · Join reliability for bot-based and botless capture
  • · Summary quality on identical source material
  • · Action-item extraction with correct owners and commitments
  • · Topic segmentation of standup updates
  • · Search and chat grounded in the meeting content
  • · Sharing, API, MCP, integrations, plan limits, languages, and privacy feature coverage
Output — unretouched
image
Also checked on this input — same tool, 7 other criteria
Action-Item Extraction✓ WorkedGrouped action items by owner, attributed them to the correct team member, and exposed a clickable source timestamp (19:13) for at least one item.Chat with Notes / Ask Questions✓ WorkedAskFred answered a natural-language question with a specific grounded response ('August 6th') and relevant context, with no hallucination reported in the tested query.Join Method & Reliability✓ WorkedUsed a bot-based Google Meet join: the Fireflies notetaker appeared in the People panel, and the recording player showed it still present in a 23:40 capture, matching the report’s claim of uninterrupted full-call capture.Search Across Notes✓ WorkedTranscript search worked with exact-match retrieval: a Ctrl+F query for 'API' returned 1/1 match at 20:42 with the hit highlighted and a clickable timestamp.Speaker Diarization✓ WorkedAttributed consecutive turns to distinct speakers in the transcript, and the report says speaker identification was almost complete with only minor attribution errors.Summary Quality✓ WorkedProduced a structured notes summary with a named header ('Task Status and Issue Resolution') rather than a blob, and the report says the full summary was multi-section and did not drop important points.Transcription Accuracy✓ WorkedGenerated a timestamped transcript view, and the report says the full transcript was very accurate: nearly all names, tools, jargon, and numbers were captured correctly with no significant misheard terms or hallucinations.
Provenance
Observation
cc75e1cb-8718-41a2-9065-0ad7df27d332
Evidence run
ace58582-3d1e-48ee-996c-9b3cd03f27a2
Study
AI Meeting Notetakers — Capture Accurate Transcripts, Summaries & Action Items From Live Calls
Research task
86baxegnv
Tested at
not recorded
Source
first-party
Evidence state
verified
Proof shown
input + output shown
Cost / latency
not captured
Repeat run
not captured
Tester
not captured

The last three rows are honest blanks, not placeholders — our capture has no field for them yet.

Query this
get_evidence({
  tool: "fireflies-ai",
  scenario: "ai-meeting-notetaker"
})
MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 7 other tools
measured on Topic Segmentation
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com