AI Demos Research — the structured-intelligence platform. Every verdict on these pages opens to the execution behind it.
Graded 9 October 2026

Can Quizgecko grade a learner who gets some fixed-answer questions right and some wrong?

Quizgecko handled the mixed fixed-answer attempt correctly. In the visible question samples, one answer was marked Correct and another Incorrect, and the results table ended at 17 correct, 36 incorrect, and 53 answered.

1 of 1 test case passed

Every test case in this scenario has a result.

Pass rate100%1 of 1 with a result
Coverage1 of 1test cases with a result
1PassDid everything it was expected to do.
0FailNo test case failed.
0Not gradableEvery result here could be graded.
0UntestedEvery test case in this scenario has a result.
The pass rate is a summary. The evidence is the test case below: what we sent, what we checked, what the tool returned, and the proof.

The test case

Each test case is judged on its own: Pass, Fail, Not gradable, or Untested. The scenario result above counts this row.

Mixed correct and incorrect fixed answersPassEvidence
What we sent
InputDocument (PDF)
Uploaded chapter PDF opening pages on electric circuits, including symbols, junction dots, and chapter sections that match the quiz topic.
Open the PDF ↗
What the tool returned

The tool's reply is shown in the evidence under Proofs.

Expected vs. Found
✓Expected: Each response is judged against the key, with correct and incorrect responses told apart and the final total matching those judgments.Found: Question 1 was Correct, question 6 was Incorrect, and the summary showed 17 Correct, 36 Incorrect, and 53 Answered.
Supporting proof
Proof 1Screenshot
Question 1 shown as Correct, with the chosen answer highlighted green and an explanation on unbroken loops.
Open original ↗
Proof 2Screenshot
Question 6 shown as Incorrect, with the learner's choice red and the correct branch-based answer green.
Open original ↗
Proof 3Screenshot
Results table showing the 17/53 header and several question rows with answers, results, and correct answers.
Open original ↗
Proof 4Screenshot
Later scroll of the results table showing more rows, including the crossing-wires question and series/parallel items.
Open original ↗
Proof 5Screenshot
Final question shown as Incorrect, with the chosen series-loop answer red and the parallel-branches answer green.
Open original ↗
Proof 6Screenshot
Completion card showing 32% and the final counts: 17 correct, 36 incorrect, 53 answered.
Open original ↗
Proof 7Screen recording
Full-screen recording of the attempt from first graded question to the results table.
Key moments
Open original ↗
Why this result

The quiz graded the visible answers against the key, marking correct and incorrect choices differently. The question-by-question judgments and the 53-question totals in the results table and recording match those marks.

Also observed on this row

Raw math formatting. Some numbers appear as unformatted markup in the questions, but grading still works.

Immediate grading. Each question is scored as soon as it is answered, rather than only at the end.

Tested by anshika gupta · evidence dated 23 September 2026

Configuration and setup

How this tool was set up for the run and what the test needed in place. Each row is a fact from the run's records; a fact the records do not hold is left out, not guessed.

Software that produced the output
Quizgecko
Build or version
N/A — no visible version number found anywhere in the recordings (SaaS web app, no exposed version indicator)
Surface
Web app
Set up before the run
A valid fixed-answer quiz with a correct answer key and a learner attempt containing both right and wrong answers had to be ready. The completed attempt was then submitted through the product's normal grading flow.
Grading mode
grades each question the moment it's answered
Tested
23 September 2026 · anshika gupta

How this scenario is graded

How we decide Pass, Fail and Not gradable. The same rules apply to every tool tested on this scenario.

How results are decided

Each test case gets one result per tool: Pass, Fail or Not gradable. A test case we haven't run yet shows Untested. There are no partial results.

The rules
  • Pass: the tool did everything the test expected, and nothing it said contradicts the correct answer.
  • Fail: at least one expected behaviour clearly didn't happen; the row says which and quotes the tool.
  • Not gradable: our evidence couldn't settle the outcome (for example a record we needed is missing). It is never counted as a fail, and the row says what's missing.

Where this sits in the benchmark

This page is one cell of a larger study: one tool, one scenario. Only this benchmark's frame appears here.

History of this result

What has happened to this result since it was first published. Runs and grades are never overwritten: a retest or a re-grade publishes a new result and keeps the earlier one readable.

from the publication record
9 October 2026First publishedQuiz Generation v1

Act on this result

Nothing filed here edits the run or the grade. A challenge opens a review, and a review can produce a new run or a re-grade — which becomes the current result and leaves this one in the history.

This matches what I see

You run the same kind of test against your own setup and get the same behaviour.

Agree →
This does not match

Yours behaves differently. Tell us what you got, with a screenshot if you have one.

Disagree →
Point out an issue

Something here is wrong — a reference value, a transcription, a grade.

Report an issue →
Request a retest

On a newer build, a larger dataset, or your own setup.

Request a retest →
We have fixed this

Tell us what changed and we schedule a rerun of the failing test case. The old result stays as history.

Vendor notice →
Filed against this evidenceNothing yet. Challenges, counter-evidence and fix notices appear here with their outcome, and stay on the page after they are resolved.
Cite this result
aidemos.com/benchmarks/quiz-generation/results/quizgecko/mixed-correct-incorrect-responses · 1 pass · 0 fail · coverage 1/1 · graded 2026-10-09

The same record is available as structured data through the AI Demos MCP server, with the counts, the coverage and every per-test-case reason carried as fields.

Verify the proof files

These files support this result. Open a file to inspect the original evidence.

Input · Document (PDF) ↗application/pdf · 265 KB
Proof 1 · Screenshot ↗image/png · 476 KB
Proof 2 · Screenshot ↗image/png · 491 KB
Proof 3 · Screenshot ↗image/png · 327 KB
Proof 4 · Screenshot ↗image/png · 395 KB
Proof 5 · Screenshot ↗image/png · 219 KB
Proof 6 · Screenshot ↗image/png · 95 KB
File fingerprints (SHA-256)

A fingerprint identifies the exact file used for this result.

Input · Document (PDF)220bc6e2dc7e6824ee5bb81ff8f24b17c234c2eb1463835468fdc77a602c3dff
Proof 1 · Screenshot8bdc699c3faab53b55941078e8781fa63a698803d525c877849912874cef692d
Proof 2 · Screenshot1e8afd52e67791da85a831d948032809040a5aa417c251ebfdaea804e3f5ee68
Proof 3 · Screenshotdd93d95e121b5e171b87ec752852b919786969f744205826235a93627560579c
Proof 4 · Screenshot2078ee2d164e67943981cf117109c0ca9be6075b14ed4ad1844e41c31d4822d2
Proof 5 · Screenshot248b00e0d883fc1a8e3b8411b03899a0f59550c6e0abae24564e0b99136b64a5
Proof 6 · Screenshoteca731d7c423c492ce9b680669df5c66f4682dca1d07323e24f78598d2012418
Proof 7 · Screen recording8b97af28adfc12ca04e8693c9fe8bdec125c068129801ddb64c528fbc68a11c6