What it does
Uncertain answers can stop for human review, but the dashboard cannot approve those requests. The maker’s comparison uses 36 hand-labelled examples, one pass and mock tools; confident mistakes still occur on ambiguous messages.
Agent tooling
Describe a workflow as a decision tree: Jev chooses branches and checks conditions, while code runs tools and a local viewer shows what happened.
Only you can see your notes.
Screenshot unavailable. Open the experiment ↗
Uncertain answers can stop for human review, but the dashboard cannot approve those requests. The maker’s comparison uses 36 hand-labelled examples, one pass and mock tools; confident mistakes still occur on ambiguous messages.
Borrow the decision-tree idea for a process such as support triage, where a small judgment chooses the next step. Ask a developer to keep tool actions in code and give each decision only the information it needs. Use a separate path when the answer is uncertain or needs a person’s judgment.
The local viewer shows which steps were taken and how certain the answers were, but cannot approve requests for human help. The examples can use stand-in answers without connecting to an AI service. The maker’s small comparison used pretend tools, so test ambiguous requests before letting the program take real actions.