JevMade

Sign in
← Back to experiments

Benchmarks & research

Span-01 (Respan)

Respan's Span-01 reads a long AI chat log and tells you, for each behaviour you describe in plain words, whether it is present, absent or impossible to judge from the log.

Source screenshot of Span-01 (Respan)
SOURCE SCREENSHOTFull screenshot ↗

What it does

The weights are closed, and a free Lite version sits beside the paid model. Respan also turned Span-01 into a judge and ranked 11 decision models with it, Jev first. Both rankings come from Respan alone, and its public benchmark was labelled by AI models agreeing with each other, not by people.

How you can use it

A team running a chatbot could use Span-01 to flag conversations where a user got frustrated, asked for a refund or tried to trick the assistant, across every log rather than a sample. The 'not observable' answer helps: it says when a conversation simply does not show enough to decide. Respan's figures are its own, so check its flags against real cases.

Keep this for later

Sign in to bookmark experiments, guides and videos, and keep notes only you can see.

Continue to sign in

We’ll bring you back to this listing.