What it does
It preserves raw API responses and reports bootstrap intervals for each performance gap.
Benchmarks & research
This benchmark compares Jev with Cohere, ZeroEntropy, and a chat model on 14 reranking datasets.
Screenshot unavailable. Open the experiment ↗
It preserves raw API responses and reports bootstrap intervals for each performance gap.
You can compare how accurately different AI services pick the best search results. The app sends thirty text passages to an outside AI service like TypeSafe. It can also test AI software running on your own equipment. The tool then asks the AI to score each passage.
Your developer starts by adding the new AI service to the project settings. Testing an external cloud service requires an access key that connects the tool to that provider. The total test cost depends on how much text the tool sends out.