What it does
Harness Judge labels each agent step as acceptable, worth retrying, needing escalation or ready to stop.
Playable demos & bots
Screenshot unavailable. Open the experiment ↗
Harness Judge labels each agent step as acceptable, worth retrying, needing escalation or ready to stop.
To start, collect a few recorded actions from your AI helper, including the overall goal and the text log of what it just did. You can test these logs directly in the online demo to see how it reviews each action.
A developer can connect this idea to your automated workflow so each action gets checked right away. The judge decides whether the system should keep going, try the step again, ask a person for help, or stop entirely before mistakes repeat.