JevMade hello@JevMade.com
← Back to experiments

Benchmarks & research

JEV Reranking Comparisons

Reranks fixed MS MARCO candidate sets by asking Jev one relevance question per query and document pair, then scores the runs with trec_eval.

Source screenshot of JEV Reranking Comparisons
SOURCE SCREENSHOTFull screenshot ↗

What it does

It compares against a local monoBERT reranker and published baselines on the same candidates. A complete Jev run cost $0.762991 where monoBERT cost nothing but compute.

How you can use it

If your app has search results, you can test different tools that reorder them. Give each tool the exact same starting list of text. Then, check if the new order puts the best answers at the top. This keeps the comparison fair for each tool.

A developer can review the saved test data and commands to get started. They can then adapt this testing approach for your project. This code only scores an existing list of answers. It is not a complete search system on its own.

Maker-reported (not independently measured by JevMade): JEV matched passage text: nDCG@10 0.6825, MAP 0.4748, P@10 0.6116, recip_rank 0.8594, $0.762991 per run · monoBERT (our local implementation): nDCG@10 0.7177, MAP 0.4488, $0 hosted API · BM25, no reranking (bm25base_p): nDCG@10 0.5058, MAP 0.3013, recip_rank 0.7036

Primitives
noul
Platform
Python
Added
Project created