What it does
Preview one call with its full distribution, run a question over hundreds of items, then score your wording against labeled examples to find where the threshold should actually sit.
Agent tooling
MCP tools for testing your Jev questions before they matter.
Only you can see your notes.
Screenshot unavailable. Open the experiment ↗
Preview one call with its full distribution, run a question over hundreds of items, then score your wording against labeled examples to find where the threshold should actually sit.
Start by collecting sample records with their correct answers, such as support tickets marked urgent or normal. A developer can link this tool to your TypeSafe account. They will test different question wordings against your list to find which phrasing produces the most reliable results.
Once you pick the best question, your developer can run it across large batches to sort or filter incoming items. Your results depend on how well your test samples match everyday tasks. Retest your questions whenever the underlying AI service updates, since new versions can shift previous scores.