Before you dive in
What you’ll find in the original
- Reranking bge-m3’s top 30 with Jev added only 0.012 NDCG@10 and the confidence interval crossed zero, despite better first-result measures.
- The apparent gain changed from +0.053 with Jev judging to +0.012 with combined judges and −0.028 with Haiku alone, exposing circular evaluator bias.
- RRF fusion reached 0.864 NDCG@10 and beat vector-only ranking under all three judges, while poor initial recall remained impossible for reranking to repair.
Worth knowing
These are author-reported results on Agent Skills Hub’s dataset. The post says data and per-query scores are open, but JevMade did not reproduce the benchmark.