JevMade Sign in
← Back to experiments

Benchmarks & research

Jev no ENEM

Benchmarks Jev on Brazil's ENEM 2025 exam against open-LLM baselines, contrasting a raw transcribed state with a structured one and measuring calibration.

Bookmark: Jev no ENEM Keep this in your collection.
Leave a noteWhat would you try with this? : Jev no ENEM

Only you can see your notes.

Source screenshot of Jev no ENEM
SOURCE SCREENSHOTFull screenshot ↗

What it does

An escape option that lets the model abstain works as an uncertainty filter: items Jev accepted were 78.8% correct, while items it had rejected scored 35.5% when forced.

How you can use it

Gather multiple-choice questions with known answer keys, such as test questions. Have a developer build an app using Jev to select between the answers. Add an explicit pass option so the system can abstain instead of guessing when a question is ambiguous or difficult.

Your developer will need an access key that connects the app to TypeSafe AI. While this setup helps filter out uncertain guesses, passing does not guarantee that the answers accepted by the model are always correct.

Keep this for later

Sign in to bookmark experiments, guides and videos, and keep notes only you can see.

Continue to sign in

We’ll bring you back to this listing.