Respan's Span-01 reads a long AI chat log and tells you, for each behaviour you describe in plain words, whether it is present, absent or impossible to judge from the log.
The weights are closed, and a free Lite version sits beside the paid model. Respan also turned Span-01 into a judge and ranked 11 decision models with it, Jev first. Both rankings come from Respan alone, and its public benchmark was labelled by AI models agreeing with each other, not by people.
How you can use it
A team running a chatbot could use Span-01 to flag conversations where a user got frustrated, asked for a refund or tried to trick the assistant, across every log rather than a sample. The 'not observable' answer helps: it says when a conversation simply does not show enough to decide. Respan's figures are its own, so check its flags against real cases.