JevMade hello@JevMade.com
← Back to guides

JevMade field notes / Benchmark report

Jev as a reranker: an honest negative

Jason Zhu reports a 164-query bilingual reranking study: Jev alone did not clearly beat bge-m3, judge choice reversed the conclusion, while reciprocal-rank fusion improved consistently.

Original by Jason ZhuRetrievalX postOriginal published Source reviewed

Before you dive in

What you’ll find in the original

  1. Reranking bge-m3’s top 30 with Jev added only 0.012 NDCG@10 and the confidence interval crossed zero, despite better first-result measures.
  2. The apparent gain changed from +0.053 with Jev judging to +0.012 with combined judges and −0.028 with Haiku alone, exposing circular evaluator bias.
  3. RRF fusion reached 0.864 NDCG@10 and beat vector-only ranking under all three judges, while poor initial recall remained impossible for reranking to repair.
Worth knowing

These are author-reported results on Agent Skills Hub’s dataset. The post says data and per-query scores are open, but JevMade did not reproduce the benchmark.