JevMade Sign in
← Back to guides

JevMade field notes / Quiz adapter and author-reported evaluation

Check which quiz answers need a second look

QuizPilot's jack explains how practice questions become Jev judgments and reports two runs on 90 self-authored questions, with arithmetic mistakes and limits to its confidence-based review rule.

Original by jack (@jack9999)Evaluation

Listen to this guide

JevMade’s plain-English explanation

0:00 /

Our summary

QuizPilot is a browser extension that suggests answers to practice questions. Its author, jack, describes testing Jev on 90 questions written for the experiment, half in English and half in Chinese. Jev reportedly answered 85 or 86 correctly per run; the errors involved single-choice calculation questions.

A single-choice question asks Jev to pick an option. A true-or-false question asks for the chance that a statement is true, while multiple-select questions check each option separately. QuizPilot turns those chances into its own certainty measure and can send uncertain answers to another model for review.

Certainty is not proof: two wrong answers escaped the proposed review cutoff. The reported combined score was reconstructed from separate model runs, not measured with live review enabled. The original and its mirror also disagree about one question-type total. Use this small study to plan checks, not to promise reliable exam answers.

Key takeaways

  1. Use one choice for a single answer and separate yes-or-no checks when several options may be correct.
  2. QuizPilot's certainty measure for yes-or-no answers is its own calculation, not an extra confidence value supplied by Jev.
  3. Check arithmetic answers even when confidence looks strong. Test the actual review workflow before claiming the reconstructed improvement.

The original reports 59 of 60 basic answers and 26 or 27 of 30 hard answers right per run. Across both runs it lists true-or-false at 60 of 60, multiple-select at 50 of 50 and single-choice at 61 of 70. Its DEV mirror instead says 40 of 40 true-or-false, leaving only 160 answers. One error example also differs: counting integers divisible by neither two nor three versus summing primes. We use the original account; the question set was not audited. Another example is two to the hundredth power modulo seven, not the number 2,100. For yes-or-no answers, QuizPilot doubles the probability, subtracts one and ignores the minus sign to measure certainty; multiple-select uses the least certain option. A cutoff below 0.6 would flag 14 of 180 answers and seven of nine errors, missing errors at 0.69 and 0.71. The combined 60 of 60 basic and 28 of 30 hard scores came from separate runs, not live review. We read the adapter, but did not run it or verify the wider review workflow. Use only where study aids are allowed. The article was published October 6, 2026, not the repository's September 25 creation date.

QuizPilot · Original published

Read the original guide Opens the author’s site in a new tab.

Keep this for later

Sign in to bookmark experiments, guides and videos, and keep notes only you can see.

Continue to sign in

We’ll bring you back to this listing.