Tool in benchmark · Version 1

Freshdesk Freddy in AI Customer Support Chatbots

Scenario-level performance from current published Results.

7 scenarios with published Results · 26 scenarios in the benchmark

How Freshdesk Freddy performed

Open a capability to explore its scenarios. Each row reports the test set in its published Result; counts are not combined into an overall score.

Action execution5 scenarios · 2 with published Results
ScenarioPublished outcomesTest coverageResult
Action requires confirmation
1 Pass0 Fail0 Not gradable
1 of 1 test case passed
What happened

Freshdesk Freddy confirmed before cancelling, so the action was not taken on the first request. In TC16, the first reply to "Cancel my product subscription." asked the user to confirm, and the final cancellation happened only after explicit confirmation.

1/1 assessed1/1 gradablePublished test setView Result →
Requested action cannot be completed
1 Pass0 Fail0 Not gradable
1 of 1 test case passed
What happened

Freshdesk Freddy refused the requested change and gave a next step when the action could not be completed. It said the delivery address for order 41982 could not be changed because the order had already shipped, then offered UPS tracking 1Z999AA10123456784 and follow-up help without claiming the change worked.

1/1 assessed1/1 gradablePublished test setView Result →
Retrieve information from an external systemNo published result
Update something in an external systemNo published result
User is not authorized to perform the actionNo published result
Agent configuration and control2 scenarios · 2 with published Results
ScenarioPublished outcomesTest coverageResult
Configured instruction changes agent behavior
1 Pass0 Fail0 Not gradable
1 of 1 test case passed
What happened

Freshdesk Freddy followed a configured instruction in the recorded reply. In the capture, the agent had a custom rule saved, the customer asked about returning an opened candle, and the reply said it could not be returned while repeating the same rule. The preview screen also says agent transfer does not happen there, so the result is about instruction-following, not handoff.

1/1 assessed1/1 gradablePublished test setView Result →
Configured restriction is respected
1 Pass0 Fail0 Not gradable
1 of 1 test case passed
What happened

Freshdesk Freddy respected the configured restriction in the shown case. The recording shows a rule against changing the delivery address, a customer asking for that change, and the agent refusing while asking for the order number and current address.

1/1 assessed1/1 gradablePublished test setView Result →
Analytics and observability3 scenarios · 1 with published Results
ScenarioPublished outcomesTest coverageResult
Conversation activity can be inspected
1 Pass0 Fail0 Not gradable
1 of 1 test case passed
What happened

Freshdesk Freddy exposed the conversation's activity in enough detail to inspect what happened. In TC22, the opened conversation showed the customer and agent turns, timestamps, an order-details action, workflow completion, status, and linked knowledge sources on the product surface.

1/1 assessed1/1 gradablePublished test setView Result →
Agent actions/tool calls can be inspectedNo published result
Reported analytics match what actually happenedNo published result
Intent understanding3 scenarios · 2 with published Results
ScenarioPublished outcomesTest coverageResult
One message contains multiple requests
1 Pass0 Fail0 Not gradable
1 of 1 test case passed
What happened

Freshdesk Freddy passed this scenario: one customer message with two requests received one reply that answered both. The reply gave the shipped order status for order 41982 and the return policy window, with the policy’s standard and member time limits shown in the same message.

1/1 assessed1/1 gradablePublished test setView Result →
Request requires choosing the correct source or action
1 Pass0 Fail0 Not gradable
1 of 1 test case passed
What happened

Freshdesk Freddy used the order system for the order-status request and answered with shipped-order details. The log shows the turn triggering "Get Order Status," executing "Get Order Details," and returning status Shipped with carrier UPS, tracking number 1Z999AA10123456784, and an estimated delivery date of 28 August 2026. The preview and conversation-log segments were captured at different times, so this result is established by the log segment alone.

1/1 assessed1/1 gradablePublished test setView Result →
Request is unclear and needs clarificationNo published result
Conversation continuity3 scenarios · 0 with published Results
ScenarioPublished outcomesTest coverageResult
Customer corrects information given earlierNo published result
Earlier conversation/session needs to be continuedNo published result
Information from earlier in the conversation is needed laterNo published result
Escalation and human handoff3 scenarios · 0 with published Results
ScenarioPublished outcomesTest coverageResult
A configured rule requires escalationNo published result
Agent cannot resolve the requestNo published result
User explicitly asks for a humanNo published result
Feedback and learning2 scenarios · 0 with published Results
ScenarioPublished outcomesTest coverageResult
Correction changes future behavior where the product claims learningNo published result
Feedback can be submittedNo published result
Knowledge-grounded answering3 scenarios · 0 with published Results
ScenarioPublished outcomesTest coverageResult
Answer is available in the knowledge baseNo published result
Answer is not available in the knowledge baseNo published result
The request uses different wording than the sourceNo published result
Multilingual support2 scenarios · 0 with published Results
ScenarioPublished outcomesTest coverageResult
Customer communicates in another supported languageNo published result
Customer switches languages during the conversationNo published result

Reading these Results

Published evidence and test coverage answer different questions.

Publication availability

Which scenarios have a Result?

A published Result is public evidence for this tool on one scenario. “No published result” does not say whether testing has taken place.

Test coverage

What does each Result cover?

Assessed includes Pass, Fail and Not gradable. Gradable includes Pass and Fail. Both use the pinned test count in that published Result.

Scenario scope

Inventory is not testing progress

The 26 scenarios describe this benchmark’s scope. They are not an assumed applicability or test-coverage denominator for Freshdesk Freddy.

Freshdesk Freddy in AI Customer Support Chatbots | AI Demos