What it does
Code computes the downstream counts and claim labels, so a right answer, a right outcome and a right final label are three separate scores.
Benchmarks & research
Tests Jev on scientific judgement questions, then follows the consequences: what those choices do to the results that depend on them.
Screenshot unavailable. Open the experiment ↗
Code computes the downstream counts and claim labels, so a right answer, a right outcome and a right final label are three separate scores.
Start by writing science questions with clear choices and known correct answers. Map out how each choice changes later totals or conclusions. A developer can write basic code that updates these counts whenever an AI picks an answer.
Your developer can run the project's tests to score choices separately from final results. Keep in mind that passing these prepared test questions does not guarantee the model will make good decisions on brand new studies.