JevMade hello@JevMade.com
← Back to guides

JevMade field notes / Model construction and benchmark

Testing a free AI tool that chooses from fixed options

This guide explores a free AI tool that picks from a list of options instead of writing text. You will learn how the author tested its accuracy and speed against a commercial version using hidden questions.

Original by allebeeEvaluation

Listen to this guide

JevMade’s plain-English explanation

0:00 /

AI narration

Credits

“JevK5” by allebee. Read the original source.

This expanded guide is an AI-narrated adaptation prepared by JevMade. It expands the source’s essential ideas, examples and caveats in JevMade’s own words and is not a word-for-word reading. The synthetic voice does not imitate the author or imply their endorsement.

Our summary

Jev is an AI tool that chooses from options rather than writing an answer, which is helpful for sorting support tickets or checking rules. This project introduces a free version you can run on your own computer and explains how the author built and tested it.

The software reads a document and calculates the chance that each possible answer is correct. If a question has more than sixteen choices, the tool breaks them into smaller groups to find the best matches, then runs a final round to pick the overall winner.

This guide is useful for people testing AI decision software. The author kept a set of hidden questions separate to ensure a fair final test. Long documents take much more time to process, and the tool struggles to recognize when none of the options are correct.

Key takeaways

  1. Keep a set of fresh, hidden questions to test the software fairly at the very end.
  2. Measure response times separately for long documents because they take much longer than short questions.
  3. Group large lists of choices into smaller batches to help the software process them accurately.

An independent test showed the free tool scored 33.1 percent on hidden questions, while the commercial version scored 36.7 percent. The author's specific speed measurements and extra experiments were not independently verified.

GitHub repository · Source reviewed

Read the original guide Opens the author’s site in a new tab.