Our summary
When checking something becomes cheaper, organizations may check more things rather than simply spend less. Richard Hill uses Jev to explore that possibility. His paper distinguishes applying a scoring rule from deciding which evidence matters, what the rule should measure and who has authority to act on the result.
Extra checks can share the same missing evidence or repeat the same mistake. A model cannot recover an important warning that disappeared from its input. Hill asks whether cheaper checks could move the harder work toward choosing rules, preserving evidence, understanding disagreement and deciding which exceptions need a person's attention.
The paper proposes questions for research, not measured Jev adoption or organizational savings. Demand would grow only if useful needs remain and other costs do not outweigh cheaper checking. Repeated checks do not automatically provide independent evidence, and a high score does not itself give permission for an important action.
Key takeaways
- Keep scoring against a rule separate from choosing the rule and authorizing what happens next.
- Count the cost of gathering evidence, reviewing exceptions and correcting errors, not just the price of a model request.
- Treat agreement cautiously when evaluators share evidence or errors; more answers need not mean more independent support.