What it does
Its headline finding is that Jev ties the best of the six on the yes/no task at roughly a twenty-eighth of the price, and that no model clearly beats guessing on the preference-scoring task.
Benchmarks & research
The open data behind an independent benchmark that puts Jev and six LLMs through the same typed questions on PubMedQA, Banking77 and HelpSteer2.
Only you can see your notes.
Screenshot unavailable. Open the experiment ↗
Its headline finding is that Jev ties the best of the six on the yes/no task at roughly a twenty-eighth of the price, and that no model clearly beats guessing on the preference-scoring task.