What it does
Supports hosted TypeSafe Jev, other System One-compatible providers, and local decision models; remote providers require explicit opt-in. Its eight-ticket model comparison is a smoke test, not a general benchmark.
Integrations
An early DuckDB extension for evaluating natural-language predicates and defined answer sets from SQL.
Only you can see your notes.
Screenshot unavailable. Open the experiment ↗
Supports hosted TypeSafe Jev, other System One-compatible providers, and local decision models; remote providers require explicit opt-in. Its eight-ticket model comparison is a smoke test, not a general benchmark.
Use the support-ticket example to explore asking questions about rows in a database: does a message request a refund, and should billing, defect support or another team handle it? The example starts with a dummy model that returns fixed answers, so it demonstrates the workflow without judging real messages.
For actual judgments, a developer connects a decision model and compares its answers with known ticket labels. Hosted models receive the text being judged and require explicit permission to send it. The eight-ticket comparison is only a smoke test; it cannot establish accuracy on your own customer messages.