An evidence-organized survey of core and peripheral Jev studies, open implementations, calibration, selective prediction, routing, robustness, judging, and structured inference.
Original by EurekaleoEvaluationRepository README and surveySource reviewed
Before you dive in
What you’ll find in the original
Treat every number as reported by its original author or vendor, not as a result from one unified benchmark.
Separate core Jev studies, peripheral evidence, and open implementations before drawing conclusions across them.
Use calibration, selective prediction, decision utility, robustness, judging, and routing literature to test claims from more than one conceptual angle.
Worth knowing
This is a curated survey, not a unified leaderboard or independent replication of the cited results.