JevMade Sign in
← Back to experiments

Benchmarks & research

explore-typesafe-ai

Five hospital tasks, a thousand synthetic patients, and Jev answering every one, with Claude writing the scenarios and reviewing the results.

Source screenshot of explore-typesafe-ai
SOURCE SCREENSHOTFull screenshot ↗

What it does

The label sets were fixed before the first model run, so the answers could not shape the test.

How you can use it

For a healthcare project, borrow the way this study fixes expected answers before testing AI. Start with its made-up patient cases, not real records. A developer can compare new answers with the saved ones. Results on made-up patients are not proof of safety in care; qualified people must still make real clinical decisions.

Maker-reported (not independently measured by JevMade): 1,000 synthetic FHIR patients · 4,722 requests on jev-1.13.0 · 300 queries (and 2,922 single-note requests) in scenario 4

Primitives
choice, score, noul
Platform
Python
Added
Project created

Keep this for later

Sign in to bookmark experiments, guides and videos, and keep notes only you can see.

Continue to sign in

We’ll bring you back to this listing.