JevMade hello@JevMade.com
← Back to evaluation videos

JevMade field notes / Video guide

Jev's 444x Claim: What TypeSafe's Own Benchmark Actually Says

This video examines TypeSafe AI's decision model Jev, which returns structured probabilities instead of text. It investigates TypeSafe's advertised speed and cost claims, unpacking the self-disclosed evaluation caveats and analyzing the significant trade-off of losing inspectable chain-of-thought reasoning.

Original by ClearboxEvaluationIntermediate9 min 20 sec Published Source reviewed

Before you press play

What you’ll find in the video

  1. Jev avoids text generation entirely, instead returning pre-defined typed values and probabilities in parallel for high-volume automated routing or scoring.
  2. TypeSafe's advertised benchmark gains are self-tested best-case figures scored against frontier model consensus rather than established ground truth.
  3. Eliminating natural language output removes inspectable reasoning trails, returning black-box floats where bias detection and thresholding fall entirely on developers.
Worth knowing

Gemini-assisted video/transcript review. Benchmark speedups of 193x and 444x are self-tested best-case estimates evaluated against model averages; independent tests showed narrower speedups (5–18x), and typed outputs provide no explanatory rationale.

Jev's 444x Claim: What TypeSafe's Own Benchmark Actually Says