JevMade hello@JevMade.com
← Back to experiments

Benchmarks & research

Scientific Decision Evaluation

Tests Jev on scientific judgement questions, then follows the consequences: what those choices do to the results that depend on them.

Source screenshot of Scientific Decision Evaluation
SOURCE SCREENSHOTFull screenshot ↗

What it does

Code computes the downstream counts and claim labels, so a right answer, a right outcome and a right final label are three separate scores.

How you can use it

Start by writing science questions with clear choices and known correct answers. Map out how each choice changes later totals or conclusions. A developer can write basic code that updates these counts whenever an AI picks an answer.

Your developer can run the project's tests to score choices separately from final results. Keep in mind that passing these prepared test questions does not guarantee the model will make good decisions on brand new studies.

Maker-reported (not independently measured by JevMade): 20 groups, 40 scientific Choices and one engineering Choice · 650 saved S2 requests reparsed; 13 models analysed · 5 repetitions per model, concurrency 1, timeout 300s, request size limit 98304 bytes

Primitives
choice
Platform
Python
Added
Project created