JevMade Sign in
← Back to guides

JevMade field notes / Written guide

Test Jev’s choices in a search-and-answer system

Akshat Kumar compares Jev with local models for choosing a search route, ranking passages and deciding whether there is enough evidence to answer. The results show why each decision needs its own test.

Original by Akshat Kumar, LyzrEvaluation

Listen to this guide

JevMade’s plain-English explanation

0:00 /

Our summary

A search-and-answer system must choose where to look, which passages to keep and when it has enough information. Akshat Kumar’s Lyzr article tests Jev on these jobs using public datasets. Jev supplies bounded judgments; a separate language model would still write the answer.

The author reports that Jev led the untrained routing comparisons and improved passage rankings in these tests. However, a small classifier trained with ten examples per category beat it on one banking task. For questions needing two pieces of evidence, raising the answer threshold reduced early answers but caused more unnecessary searches.

These are the author’s measurements, not independently repeated results. The comparisons exclude language-model judges, use small samples and run local models on a shared laptop processor. Possible training overlap is unknown. Use the method to test your own documents and decision thresholds, rather than treating one score as a general guarantee.

Key takeaways

  1. Test each job separately: choosing a route, ordering passages and deciding whether the evidence supports an answer can fail in different ways.
  2. Measure both kinds of threshold mistake. At a sufficiency cutoff of 0.5, Jev would answer 31% of the tested missing-evidence cases; a higher cutoff also causes extra searches.
  3. Compare with a simple trained classifier when you have labelled examples. In the author’s banking test, ten examples per category beat Jev’s untrained routing accuracy.

Author-reported evaluation of Jev 1.13.0: 21,314 calls across six public tasks. The reported evidence-sufficiency AUROCs are 0.898 on SQuAD v2 and 0.897 on HotpotQA; these describe score separation, not correctness at every cutoff. No language-model baseline, executable benchmark repository or downloadable result data was verified. Single-run samples contain 200–800 cases, some routing labels follow dataset identity, and local timing used a shared laptop GPU. Public-dataset training overlap is unknown. Recommendations, calibration and thresholds are specific to these samples.

Lyzr · Original published

Read the original guide Opens the author’s site in a new tab.

Keep this for later

Sign in to bookmark experiments, guides and videos, and keep notes only you can see.

Continue to sign in

We’ll bring you back to this listing.