AI Demos Research — the structured-intelligence platform. Every verdict on these pages opens to the execution behind it.
Graded 12 September 2026

Can Freshdesk Freddy show enough conversation activity to inspect what happened in a thread?

Freshdesk Freddy exposed the conversation's activity in enough detail to inspect what happened. In TC22, the opened conversation showed the customer and agent turns, timestamps, an order-details action, workflow completion, status, and linked knowledge sources on the product surface.

1 of 1 test case passed

Every test case this benchmark pins to the scenario has an accepted result.

Pass rate100%1 of 1 with a result
Coverage1 of 1test cases with a result
1PassDid everything it was expected to do.
0FailNothing was shown not to hold either.
0Not gradableEvery result here could be graded.
0UntestedEvery test case in this scenario has a result.
The pass rate is a summary. The one row below is the evidence — each opens onto the stimulus as sent, the expectations it was checked against, what the tool returned, and the proof.

The test case

Each test case is judged on its own: Pass, Fail, Not gradable, or Untested. The scenario result above counts this row.

Find the earlier ⟨shipped-order⟩ conversationPassEvidence
What we sent

Exact stimulus wording was not preserved for this run.

What the tool returned

The tool's reply was not preserved in the evidence for this run.

Expected vs. Found
What this scenario evaluates

These are scenario-level criteria. Each test case's Expected and Found are listed separately.

  • Whether earlier conversation activity can be located in the product interface.
  • Whether the located conversation can be opened for review.
  • Whether the inspection view exposes the prior activity clearly enough to inspect what happened.
How this scenario is judged →
✓Found: The conversation was found in Ticket logs and opened.
Supporting proof
Proof 1Screen recording
Screen recording of Ticket logs search opening Conversation #14 and showing the shipped-order turn.
Key moments
Open original ↗
Why this result

Ticket logs found and opened Conversation #14. The transcript then showed order 41982 with Status: Shipped.

Also observed on this row

Search by ticket id. The Ticket logs search box found the conversation using ticket id 14. The opened log matched the shipped-order conversation.

Workflow trace visibility. Conversation logs show per-answer knowledge-source attribution and workflow traces. The transcript includes the order lookup action, the generated response, and workflow completion.

Tested by Ajay Vekhande · evidence dated 4 September 2026

Configuration and setup

How this tool was set up for the run and what the test needed in place. Each row is a fact from the run's records; a fact the records do not hold is left out, not guessed.

Software that produced the output
Freshdesk Freddy
Surface
Ticket logs screen
Tested
By 4 September 2026 · Ajay Vekhande

How this scenario is graded

How we decide Pass, Fail and Not gradable. The same rules apply to every tool tested on this scenario.

How results are decided

The test case. One result per test case per tool: Pass, Fail, Not gradable. A pinned test case with no accepted result reads Untested. No Partial.

The rules
  • Pass — every expectation on the test-case version holds against the registered reference, and nothing in the reply contradicts the reference.
  • Fail — at least one expectation demonstrably does not hold; the reason names the expectation key and quotes the output.
  • Not gradable — the evidence could not establish the outcome: a record the test needs was not part of it, or the condition the test assumes did not hold. Never inferred as a fail; the row says what could not be established.

Where this sits in the benchmark

This page is one cell of a larger study: one tool, one scenario. Only this benchmark's frame appears here.

LevelNameScope
BenchmarkAI Customer Support Chatbots →v1 · 26 scenarios · 13 products · not yet frozen
CapabilityAnalytics and observability →
ScenarioConversation activity can be inspected →S22 · weight 1.0 · role context
Rubricnone pinnedgraded against the test-case expectations
Test casesTC221 pinned
ToolFreshdesk Freddy →tool

History of this result

What has happened to this result since it was first published. Runs and grades are never overwritten: a retest or a re-grade publishes a new result and keeps the earlier one readable.

from the publication record
22 September 2026First publishedAI Customer Support Chatbots v1

Act on this result

Nothing filed here edits the run or the grade. A challenge opens a review, and a review can produce a new run or a re-grade — which becomes the current result and leaves this one in the history.

This matches what I see

You run the same kind of test against your own setup and get the same behaviour.

Agree →
This does not match

Yours behaves differently. Tell us what you got, with a screenshot if you have one.

Disagree →
Point out an issue

Something here is wrong — a reference value, a transcription, a grade.

Report an issue →
Request a retest

On a newer build, a larger dataset, or your own setup.

Request a retest →
We have fixed this

Tell us what changed and we schedule a rerun of the failing test case. The old result stays as history.

Vendor notice →
Filed against this evidenceNothing yet. Challenges, counter-evidence and fix notices appear here with their outcome, and stay on the page after they are resolved.
Cite this result
aidemos.com/benchmarks/ai-customer-support-chatbots/results/freshdesk/conversation-activity-can-be-inspected · 1 pass · 0 fail · coverage 1/1 · graded 2026-09-12

The same record is available as structured data through the AI Demos MCP server, with the counts, the coverage and every per-test-case reason carried as fields.

Verify the proof files

These files support this result. Open a file to inspect the original evidence.

File fingerprints (SHA-256)

A fingerprint identifies the exact file used for this result.

Proof 1 · Screen recording12687f8db0c201b4e80ee8d71c18291bdfc08f058058563b48fde74ef8b13879