How we use Jev to answer quiz questions
QuizPilot is a Chrome extension that reads practice questions on a web page and suggests answers. Since TypeSafe released Jev, our Fast mode sends every multiple-choice and true/false question to it. Here is how we map quiz questions onto Jev, and what a 90-question benchmark showed about its accuracy, speed and confidence.
The short version
- Accuracy: Jev answered 85–86 of 90 questions correctly across two runs. A lightweight LLM got 89 and a frontier reasoning LLM got all 90.
- Speed: one Jev request answered all 60 basic questions in 0.17–0.44 s. The lightweight LLM took 2.6 s, the reasoning LLM 8.5 s.
- Where it slips: every true/false and multiple-select answer was right; all of its mistakes were single-choice questions that need several steps of arithmetic.
- Confidence is useful: 7 of Jev's 9 wrong answers came back with a confidence below 0.6, against an average of 0.93 for its right answers.
Why a decision model for quiz questions
A multiple-choice question already lists every possible answer. An LLM still has to write its reply as text, and we then parse the letter back out of it. Jev skips that: you give it a state and typed questions (pick one of these options, or yes/no), and it returns the choice with a probability. There is no prose to parse and no answer outside the options.
That fits quizzes well. Students want the answer quickly while the page is open, and a practice page often has 20 to 50 questions at once.
How a quiz question becomes a Jev question
QuizPilot reads the questions from the page, then translates each kind:
-
Single choice becomes one Jev
choicequestion whose criteria are the options. Jev's own confidence is shown to the student. -
True/false becomes one yes/no (
noul) question: "Is this statement true?" - Multiple select becomes one yes/no question per option: "Is option C one of the correct answers?" Every option at 0.5 or above is ticked; if none is, the likeliest one is.
- Fill-in, short answer and picture questions go to an LLM instead. Jev does not write text or read images.
A trimmed request for one true/false question looks like this:
{
"model": "typesafe/jev-1.13",
"state": "Questions from a quiz or exercise on a web page. Each question is independent. Judge by factual correctness, not by wording or position of the options.",
"questions": {
"q0": {
"type": "noul",
"instructions": { "statement": "The harmonic series 1 + 1/2 + 1/3 + … converges.", "task": "Is this statement true?" },
"criteria": { "true": "The statement is correct", "false": "The statement is incorrect" }
}
}
}
A yes/no answer is a single probability, so we turn it into a confidence by its distance
from 0.5: |2p − 1|. A multiple-select question takes the lowest confidence
among its options, so one uncertain option marks the whole question as uncertain. A whole
page goes out in one request, split only when it nears Jev's input limit.
The adapter is open source: QuizPilot on GitHub.
The benchmark
We wrote 90 questions with known answers, half in English and half in Chinese. None come from a real exam.
- Basic set (60): 20 single choice, 20 multiple select, 20 true/false, half easy general knowledge and half harder (probability, algorithms, physics).
- Hard set (30): mostly multi-step problems, such as the trailing zeros of 100!, 2100 mod 7, and inclusion–exclusion counting.
Each model was called the way QuizPilot calls it in production. For comparison we used a lightweight general LLM with thinking off (the kind QuizPilot uses for fill-in questions) and a frontier reasoning LLM at medium effort (the kind behind our Accurate mode). Jev and the lightweight LLM ran twice to check consistency.
| Model | Basic (60) | Hard (30) | Time, basic set |
|---|---|---|---|
| Jev | 59, 59 | 26, 27 | 0.17–0.44 s |
| Lightweight LLM | 60 | 29, 29 | 2.6 s |
| Reasoning LLM | 60 | 30 | 8.5 s |
Time is wall-clock for the whole set. Jev and the lightweight LLM answered it in one request; the reasoning LLM in batches of 8 sent in parallel. On the hard set the times were 0.18–0.25 s, about 3 s and 7.4 s.
By question type (Jev, both runs, both sets)
| Type | Correct |
|---|---|
| True/false | 60 / 60 |
| Multiple select | 50 / 50 |
| Single choice | 61 / 70 |
Where Jev gets it wrong
All nine mistakes were single-choice questions whose answer has to be computed, not recognised:
- Trailing zeros of 100! (answer 24; Jev picked 25, both runs).
- 2100 mod 7 (answer 2; Jev picked 4, both runs).
- Integers from 1 to 100 divisible by neither 2 nor 3 (answer 33; Jev picked 34, both runs).
- How many real solutions x³ − 3x = 1 has (answer 3; Jev picked 2, both runs).
- Sum of the interior angles of a 12-sided polygon (wrong in one run of two).
This matches what Jev is built for. It judges; it does not work through a calculation step by step. Questions that a person would answer by recalling or checking a fact were answered correctly, including harder ones such as "every group of prime order is cyclic".
Confidence tells you when to ask again
Across 180 Jev answers, 9 were wrong. Seven of them had a confidence between 0.26 and 0.49. The two exceptions were the same question (trailing zeros of 100!), at 0.69 and 0.71. Right answers averaged 0.93.
QuizPilot has an optional "review low-confidence answers" setting that sends anything below 0.6 to the lightweight LLM. In this benchmark it would have re-asked 14 of 180 answers (8%). Using the LLM's answers from the same benchmark, the combination scores 60 of 60 on the basic set and 28 of 30 on the hard set, up from 59 and 26–27.
This combined score is worked out from the separate runs, not from a live run with review switched on.
How QuizPilot uses this
- Fast mode: choice and true/false questions go to Jev, so a full page comes back in well under a second. Every answer shows its confidence, so a student can see which ones to double-check.
- Low-confidence review: optional, for students who want an LLM to look again at the answers Jev is unsure of.
- Accurate mode: for maths and science pages full of multi-step problems, every question goes to an advanced reasoning model instead. It is slower, and it got every question in this benchmark right.
Limits of this test
- 90 questions written by us is enough to see large gaps, not small ones. Treat the numbers as a sketch, not a leaderboard.
- The questions are text only; picture questions never go to Jev.
- We tested Jev 1.13 through OpenRouter on October 6, 2026.
Try it
QuizPilot's Free mode lets you plug in your own OpenRouter or TypeSafe key and route choice questions to Jev yourself. Paid mode uses Jev without any setup and charges per answered question. Use it for practice, homework and self-tests where study aids are allowed, not in exams that prohibit them.
More about the extension on the QuizPilot home page, or see the privacy policy.