JevMade hello@JevMade.com
← Back to experiments

Benchmarks & research

Scientific Decision Evaluation

Tests Jev on scientific judgement questions, then follows the consequences: what those choices do to the results that depend on them.

Source screenshot of Scientific Decision Evaluation
SOURCE SCREENSHOT · source ↗ · captured 2026-09-24Full screenshot ↗

What it does

Code computes the downstream counts and claim labels, so a right answer, a right outcome and a right final label are three separate scores.

Maker-reported (not independently measured by JevMade): 20 groups, 40 scientific Choices and one engineering Choice · 650 saved S2 requests reparsed; 13 models analysed · 5 repetitions per model, concurrency 1, timeout 300s, request size limit 98304 bytes

Primitives
choice
Platform
Python
Added
Project created

Source checked 2026-09-24 — opened the GitHub repository directly.