Benchmark · Version 1

Quiz Generation

This benchmark is for AI products that work from supplied study material.

Benchmark overview

7Capabilities
15Scenarios
10Participating tools
What is included and excluded

This benchmark evaluates how AI tools turn supplied study material into quizzes and practice tests, then grade, explain, and track what learners actually understood.

Buyers can use it to compare products for classroom and learner use. The benchmark shows whether the tool handles assessment content, answer evaluation, learner feedback, ongoing practice, and reporting in ways that stay faithful to the source material.

In scope

  • Turning supplied study material into quizzes and practice tests.
  • Delivering quizzes to individual learners or classes.
  • Grading fixed-answer, free-form, and partly correct responses.
  • Giving feedback that explains what is correct, missing, or wrong.
  • Adapting later practice to strengths and weaknesses, including across sessions.
  • Showing learner progress over time.
  • Showing class reports about what the class struggled with.

Out of scope

  • Product attributes, which are captured separately and never scored.

Participating tools

Tools in this benchmark’s public roster. Publication availability is not a performance ranking.

ToolPublished ResultsExplore
PDF to Quiz1 scenario with a published ResultView tool in this benchmark →
Quizgecko4 scenarios with published ResultsView tool in this benchmark →
DoctoQuizNo published resultView tool in this benchmark →
Edcafe AINo published resultView tool in this benchmark →
GPT-5.6 terraNo published resultView tool in this benchmark →
KnowtNo published resultView tool in this benchmark →
QuestionWellNo published resultView tool in this benchmark →
QuizizzNo published resultView tool in this benchmark →
QuizRiseNo published resultView tool in this benchmark →
StudyglenNo published resultView tool in this benchmark →

Capabilities & scenarios

15 scenarios grouped by 7 capabilities. Open a group to explore its scenarios in this benchmark.

Quiz Generation5 scenarios

Creates valid quizzes from supplied material that test the intended knowledge and include sufficient information for answers to be evaluated.

Capability in this benchmark → · Global definition →

  1. A cleanly scanned documentNo published results
  2. The user turns study material into a quiz1 tool with a published Result
  3. Important information appears in both text and a diagram1 tool with a published Result
  4. The user asks the quiz to test a particular learning goal1 tool with a published Result
  5. The user turns a recorded lecture into a quizNo published results
Quiz Delivery2 scenarios

Lets learners take and submit quizzes in the intended individual or class setting while recording each attempt correctly.

Capability in this benchmark → · Global definition →

  1. A learner takes a quiz on their ownNo published results
  2. A teacher gives the quiz to a classNo published results
Grading3 scenarios

Determines how correctly a learner answered, including fixed-answer scoring, semantically equivalent wording, and partial understanding.

Capability in this benchmark → · Global definition →

  1. The learner answers some fixed-answer questions correctly and others incorrectly1 tool with a published Result
  2. The learner answers correctly in their own wordsNo published results
  3. The learner's answer is partly correctNo published results
Answer Feedback2 scenarios

Explains what the learner understood correctly and what remains wrong, missing, or incomplete.

Capability in this benchmark → · Global definition →

  1. The learner's answer is partly correctNo published results
  2. The learner gives a clearly wrong answerNo published results
Personalized Practice2 scenarios

Changes later practice according to the learner's demonstrated strengths and weaknesses, including across sessions.

Capability in this benchmark → · Global definition →

  1. One topic is weaker than the othersNo published results
  2. The learner returns to the material laterNo published results
Learner Progress1 scenario

Shows learners how their performance changes over time and which topics still need work.

Capability in this benchmark → · Global definition →

  1. The learner checks whether they are improvingNo published results
Class Reports1 scenario

Shows teachers how learners performed and which questions or topics the class struggled with.

Capability in this benchmark → · Global definition →

  1. The teacher checks what the class struggled with1 tool with a published Result

Results overview

Current published evidence in this benchmark.

5Current published Results
5Scenarios with published Results
2Tools with published Results

Publication availability is separate from test coverage and unpublished research progress.

Explore scenarios →

How the benchmark works

A public summary of the evaluation method. The same defined scope and evidence standard apply to every tool assessed under this version.

Scope

Scope includes quizzes and practice tests

The benchmark covers quizzes and practice tests from supplied study material, then delivering, grading, explaining, and tracking them.

Method boundary

Open rubrics and scoring

Rubrics, weights, scoring, and ranking methodology remain open and unpinned.

Separate attributes

Product attributes are separate

Product attributes are captured separately and never scored.

Class reports

Class reports show struggle, not trend

Class reports cover what the class struggled with; do not invent an over-time class-performance trend.

Full benchmark methodology →

Resources and fixtures

The registered material and systems that create a consistent test environment for this benchmark.

Quiz Generation — Benchmark definition | AI Demos