Agentel reports four illustrative tests: routing a duplicate-charge request, checking support for an answer, scoring evidence for a product claim, and trying an ambiguous refund message. The last case illustrates why a confident label should not trigger an action on its own.
Original by AgentelEvaluationBeginner2 min 45 secPublished
Define the allowed categories before asking Jev to choose a route; the first example reports a refund label for a duplicate charge.
Distinguish a yes-or-no evidence check from an ordered strength score; the examples ask different questions about supplied material.
Keep the workflow in code and retain clarification or escalation: the ambiguous message still received a reported 100% refund answer.
Worth knowing
These are four creator-reported examples, not an accuracy or calibration benchmark. The claimed real API connection, 172 ms response and 100% outputs were not independently confirmed. No runnable code or raw test receipts are linked in the description. A probability is not authorization to refund money or a security boundary.