What it does
An escape option that lets the model abstain works as an uncertainty filter: items Jev accepted were 78.8% correct, while items it had rejected scored 35.5% when forced.
Benchmarks & research
Benchmarks Jev on Brazil's ENEM 2025 exam against open-LLM baselines, contrasting a raw transcribed state with a structured one and measuring calibration.
Only you can see your notes.
Screenshot unavailable. Open the experiment ↗
An escape option that lets the model abstain works as an uncertainty filter: items Jev accepted were 78.8% correct, while items it had rejected scored 35.5% when forced.
Gather multiple-choice questions with known answer keys, such as test questions. Have a developer build an app using Jev to select between the answers. Add an explicit pass option so the system can abstain instead of guessing when a question is ambiguous or difficult.
Your developer will need an access key that connects the app to TypeSafe AI. While this setup helps filter out uncertain guesses, passing does not guarantee that the answers accepted by the model are always correct.