Scenario in benchmark · Version 1

The user turns study material into a quiz

Study material is turned into a quiz.

How the tools performed

Every in-scope tool is visible. Outcomes come from current published Results for this scenario and benchmark version.

Publication availability: 1 tool has a current published Result.

No published result does not tell you whether a tool has been tested. Test coverage is shown only for comparable published Results.

ToolPublished outcomesTest coverageResult
Published Results · alphabetical, not ranked
Quizgecko
2 Pass0 Fail0 Not gradable
2 of 2 test cases passed
2 of 2 assessed2 of 2 gradableView Result →
No published result · alphabetical
DoctoQuiz——No published result
Edcafe AI——No published result
GPT-5.6 terra——No published result
Knowt——No published result
PDF to Quiz——No published result
QuestionWell——No published result
Quizizz——No published result
QuizRise——No published result
Studyglen——No published result

Assessed = Pass + Fail + Not gradable. Gradable = Pass + Fail. Both use the published Result’s pinned-test denominator. — means not publicly available.

Test design

Pinned test cases
2
Disclosure
0 public · 2 withheld
Capabilities exercised here
Quiz Generation

What this scenario evaluates

  • Whether the quiz is grounded in the supplied study material rather than outside knowledge.
  • Whether the questions are answerable from the material and focus on the core ideas, not just supporting detail.
  • Whether the answer or evaluation information matches the source material.
  • Whether the quiz includes enough information for responses to be judged against the material.

Exact wording, inputs, fixture state, expected output and detailed grading remain at the test-case level and may be withheld while the benchmark version is active. The scenario and its evaluation intent are public.

How the results are graded

  • Pass: the test-case expectations hold.
  • Fail: an expectation demonstrably does not hold.
  • Not gradable: the evidence cannot establish the outcome.

Version 1 uses test-case expectations; no scenario rubric is pinned.

Benchmark methodology →
The user turns study material into a quiz in Quiz Generation | AI Demos