In the maker's TRIVIA+ factual-answer test, adjusting confidence estimates reportedly reduced confidence-estimation error by 68%, not improved answer accuracy. Default pass/fail classification worsened. The specific test runner and saved predictions were not found, so those numbers were not reconstructed. Demonstrations using made-up data are separate from live tests.
How you can use it
Start with the refund example: pair a customer's question and the proposed thirty-day answer with the actual refund policy. A developer can use the toolkit to ask Jev whether the answer follows that evidence. Compare its judgments with human labels before choosing what counts as passing.
For a proposed action, provide the action details and a written policy before it runs. The model's judgment can help flag a questionable request, but is not permission to execute it. The application must still enforce access rights and action limits independently, including when the model is wrong.