JevMade Sign in
← Back to guides

JevMade field notes / Public article sections and linked experiment notes

Test one repeated product decision with Jev

Paweł Huryn uses audience moderation and small document tests to explain where Jev might help a product, why rules must be supplied, and why reported examples do not establish production accuracy.

Original by Paweł HurynEvaluation

Listen to this guide

JevMade’s plain-English explanation

0:00 /

Our summary

A product often repeats a small judgment, such as accepting an audience question or sending a support message to a team. Paweł Huryn describes using Jev for these choices. In his AskOne app, a confidence value of at least 0.8 allows an automatic moderation decision; lower values wait for a host or moderator.

His linked tests show why the request needs the actual company policy. On twenty-four invented messages, Jev followed written house rules in all twenty-four cases, but only five when those rules were absent. A separate comparison used fifty invented documents and six models. These were the author's one-run examples, not tests on real customers.

Start with a varied set of cases whose expected answers you can check, then inspect mistakes before changing the product. Huryn used another AI model to judge answers, so those labels also need scrutiny. The public article explains his method, but paid setup instructions and templates are unavailable here. Neither the examples nor the app were run by JevMade.

Key takeaways

  1. Put the actual policy into the question. A strong answer cannot reveal a rule the model was never given.
  2. Keep reported scores attached to their sample and conditions. One successful run on invented examples does not predict accuracy on real product traffic.
  3. Check common, difficult and ambiguous cases, including attempts to make the model ignore instructions. A second AI judge is not an observed real-world outcome.

The publisher dates this article October 5, 2026. Only public sections and the linked repository were inspected; paid sections were not accessed or bypassed. Implementation and saved results are pinned to the source version inspected on October 6. Saved records contain selected results, not full provider responses. The main document run made several calls at once; quoted response times come from an earlier run made one call at a time. Price comparisons depend on the chosen rival, and no benchmark or source program was executed.

The Product Compass · Original published

Read the original guide Opens the author’s site in a new tab.

Keep this for later

Sign in to bookmark experiments, guides and videos, and keep notes only you can see.

Continue to sign in

We’ll bring you back to this listing.