O JEV chegou. Veja como usar na prática (coloquei 14 modelos pra brigar)
Eli Rigobeli explains TypeSafe AI's Jev decision engine, contrasts schema enforcement with answer accuracy, and integrates it into a GTD inbox classifier. He then runs a 100-request benchmark against 14 LLMs, showing Jev used as a specialized, low-cost classifier alongside generative LLMs rather than replacing them.
Original by Eli Rigobeli - IAEvaluationIntermediate42 min 17 secPublished Source reviewed
Before you press play
What you’ll find in the video
The explainer distinguishes a valid yes-or-no, categorical, or scoring response from a correct answer.
The GTD inbox example combines state with explicit questions to assign a destination and priority.
The creator’s 100-prompt comparison shows ambiguous criteria affecting several models, emphasizing the need to evaluate the task definition as well as the model.
Worth knowing
Gemini-assisted video/transcript review. Results are from an informal 100-prompt custom test suite rather than a standardized, independently audited benchmark, and the creator noted English prompts adhere better than Portuguese ones.