JevMade

Sign in
← Back to experiments

Benchmarks & research

LLM-as-Jev

LLM-as-Jev argues that an ordinary chat model can already work as a Jev-style decision model: number the answers, then read how likely the model is to say each number.

Bookmark: LLM-as-Jev Keep this in your collection.
Leave a noteWhat would you try with this? : LLM-as-Jev

Only you can see your notes.

Source screenshot of LLM-as-Jev
SOURCE SCREENSHOTFull screenshot ↗

What it does

Without any extra training, Qwen3.5-4B scored about as well as community decision models built on the same base on public JevBench items, and it could judge images too. Fine-tuning helped the smaller 0.6B model and long option lists most. It is a research paper by Yinheng Li and Justin Wagle, not a product.

How you can use it

If you already run a chat model on your own computer, this paper suggests a cheap experiment before building anything new: write the possible answers as a numbered list and let the model's own odds for each number pick the answer. The authors tested only two Qwen models, so check the results on your own questions.

Keep this for later

Sign in to bookmark experiments, guides and videos, and keep notes only you can see.

Continue to sign in

We’ll bring you back to this listing.