Our summary
A product often repeats a small judgment, such as accepting an audience question or sending a support message to a team. Paweł Huryn describes using Jev for these choices. In his AskOne app, a confidence value of at least 0.8 allows an automatic moderation decision; lower values wait for a host or moderator.
His linked tests show why the request needs the actual company policy. On twenty-four invented messages, Jev followed written house rules in all twenty-four cases, but only five when those rules were absent. A separate comparison used fifty invented documents and six models. These were the author's one-run examples, not tests on real customers.
Start with a varied set of cases whose expected answers you can check, then inspect mistakes before changing the product. Huryn used another AI model to judge answers, so those labels also need scrutiny. The public article explains his method, but paid setup instructions and templates are unavailable here. Neither the examples nor the app were run by JevMade.
Key takeaways
- Put the actual policy into the question. A strong answer cannot reveal a rule the model was never given.
- Keep reported scores attached to their sample and conditions. One successful run on invented examples does not predict accuracy on real product traffic.
- Check common, difficult and ambiguous cases, including attempts to make the model ignore instructions. A second AI judge is not an observed real-world outcome.