JevMade hello@JevMade.com
← Back to experiments

Benchmarks & research

Using TypeSafe's Jev for evals in Datadog Agent Observability

Datadog's Agent Observability team shows Jev doing double duty in their eval stack: notebooks score production agent spans as they arrive through online evals and run the same focused questions offline, mapping typed answers onto Datadog metrics.

Source screenshot of Using TypeSafe's Jev for evals in Datadog Agent Observability
SOURCE SCREENSHOT · source ↗ · captured 2026-09-29Full screenshot ↗

What it does

The questions stay deliberately narrow — whether claims are supported, what failure occurred, how severe the impact is — and the post is candid about what Jev shouldn't be asked, like arithmetic or dates, where a parser belongs. Code ships as online and offline notebooks in DataDog's llm-observability repo. It's judge-style evaluation moving into the live path rather than after execution.

Primitives
choice, score, noul
Added
Project created

Source checked 2026-09-29 — opened the primary source directly.