What it does
This research notebook pairs runnable Jev demos with a close reading of the model's public claims.
Benchmarks & research
Screenshot unavailable. Open the experiment ↗
This research notebook pairs runnable Jev demos with a close reading of the model's public claims.
Before using AI scores to sort important messages, ask your developer to compare them with answers people have checked. Start with the saved tests here, which you can study without an AI connection. Do answers scored near 90% really turn out right nine times out of ten? Check that on your own messages; results from a different task cannot settle it.