JevMade hello@JevMade.com
← Back to experiments

Benchmarks & research

phoenix

A Phoenix Evals example pits Jev against a small grounded-versus-hallucinated answer benchmark.

Source screenshot of phoenix
SOURCE SCREENSHOT · source ↗ · captured 2026-09-26Full screenshot ↗

What it does

For each case it records Jev's label and probability, then builds correctness and latency tables. Running the example requires a TypeSafe key.

Primitives
choice
Platform
TypeScript
Added
Project created

Source checked 2026-09-26 — opened the primary source directly.