JevMade hello@JevMade.com
← Back to experiments

Benchmarks & research

phoenix

A Phoenix Evals example pits Jev against a small grounded-versus-hallucinated answer benchmark.

Source screenshot of phoenix
SOURCE SCREENSHOTFull screenshot ↗

What it does

For each case it records Jev's label and probability, then builds correctness and latency tables. Running the example requires a TypeSafe key.

How you can use it

Start by collecting sample questions with their source documents. Add answers you have already marked by hand as truthful or made up. A developer can write a script that sends these texts to TypeSafe AI, which judges whether each answer matches the facts.

The script then compares those AI decisions and response speeds directly against your original notes. Running this check requires an access key from TypeSafe. Because an automated evaluation can still make mistakes, a person should check important results.