JevMade Sign in
← Back to guides

JevMade field notes / Agent supervision implementation and study

Check an AI coding assistant's decisions without mistaking activity for progress

Kristiyan Stoyanov explains an add-on for the Pi coding assistant that asks Jev to assess selected decisions, assumptions, progress and claims that work is finished. Ordinary code decides whether the assistant should pause or continue.

Original by Kristiyan StoyanovAgent workflows

Listen to this guide

JevMade’s plain-English explanation

0:00 /

Our summary

An AI coding assistant can build working software while quietly changing what it was asked to do. Kristiyan Stoyanov explores whether a second model can catch that drift. His add-on for the Pi coding assistant gathers selected evidence for Jev to judge, then uses ordinary code to decide whether the assistant should continue or pause.

The add-on checks decisions, assumptions, useful information, progress and claims that work is finished. It keeps recorded tool results separate from the assistant's statements and ties reported test results to a recorded version of the code. Decision and assumption checks require the assistant to report them before work that relies on them; progress and completion also have automatic checks.

The small paired pilot gave both versions 14 out of 15 acceptance checks, while supervision used more time and input tokens. Control tests show pauses can be delivered, not that coding improves. The simple file-protection example is not a sandbox, and the full extension needs explicit spending limits because its default call counts are unlimited.

Key takeaways

  1. Keep a recommendation, a delivered pause and a better final result separate. More supervisor activity is not evidence of better software.
  2. Report an important decision in a separate check-in and wait before work that relies on it. Unreported decisions can escape the decision and assumption checks.
  3. Start with observation and explicit budgets on a small task. Observation still calls hosted Jev, and recorded passing claims still need independent acceptance checks.

The article and relevant parts of version 0.5.0-dev.1 were inspected, not run. The teaching rule blocks Pi's built-in write and edit tools on the requirements file, not shell commands, other tools or alternate paths through file links. The full add-on's pauses can block more tools. Reported test results are linked to executed output and a recorded code version, not independently checked against every requirement. Without configuration, all five checks are enabled and can act, with unlimited requests, assessments and interventions; extra attempts when the assistant tries to finish are separately capped at two. The three-task pilot reports one hundred and sixty-three thousand six hundred and fifty-six input tokens, or pieces of text, without supervision versus eight hundred and eighty-eight thousand one hundred and eighty-five with it. Time was six hundred and forty-four point three eight seconds versus one thousand and thirty-one point zero one seconds. Both scored 14 out of 15. Studies used changing versions, and complete run records are not bundled. Even a local coding assistant sends selected evidence to hosted Jev, by default through Vercel AI Gateway. Records with credentials removed can still contain sensitive project content.

DEV Community · Original published

Read the original guide Opens the author’s site in a new tab.

Keep this for later

Sign in to bookmark experiments, guides and videos, and keep notes only you can see.

Continue to sign in

We’ll bring you back to this listing.