Our summary
Software makers need a way to check if their AI assistants are giving good answers. The authors test a new grading tool called Jev. Unlike standard AI that writes out text, Jev simply chooses from set options to score how well an assistant answered a question.
The authors saved five answers from a weather assistant and had a person grade them. They then sent these exact same answers to Jev and three other AI judges one hundred times each. This test checked if the grading software gave the same score every time.
Jev matched the human grades perfectly and gave more consistent scores than the other models in this test. However, the test only looked at five specific weather answers. Software testers should check if these results hold up when grading different topics or more complex mistakes.
Key takeaways
- Test grading tools on saved answers to see if they give the same score every time.
- Jev matched human grades on all five tested answers, but this does not prove broad accuracy.
- The authors report lower costs and faster speeds, which could help teams check software more often.