Codebasics and AI research engineer Siddhant Pandey demonstrate Jev's API and LangChain integration. They explain how Jev functions as a decision and classification model using primitive types like Noul, score, and choice, testing it on insurance processing, model routing, and safety guardrails.
Original by codebasicsEvaluationIntermediate17 min 6 secPublished
Jev evaluates state against yes-or-no propositions, supplied categories, or ordered scores instead of generating text.
The presenters describe TypeSafe’s reinforcement learning for calibrated decisions (RLCD) approach and contrast decision outputs with autoregressive text generation.
Their guardrail tests include unsafe requests that Jev approves, illustrating why structured outputs still need task-specific evaluation.
Worth knowing
Speed and cost figures cited from founders (20x-200x speedups) are unverified claims; in their Colab test against GPT-3.5, Jev was only around 2x faster, and guardrailing tests exhibited failures.