JevMade Sign in
← Back to videos

JevMade field notes / Video guide

Building an Agent Harness with Jev: What Actually Works?

Prompt Engineering reports testing a Pi coding agent on three database tasks with five harness configurations. Selecting relevant instructions helped one task; tool gates exposed bypasses and added work.

Original by Prompt EngineeringAgent workflowsIntermediate15 min 32 sec Published

Before you press play

What you’ll find in the video

  1. Select relevant instructions before the agent starts. The maker reports that loading one runbook section helped it repair old product links before removing a column.
  2. Protect every route to the same resource, not just one named tool. In the maker's tests, an agent went around a SQL-tool gate through the shell; blocked calls also increased token use.
  3. Separate evidence checks from task success. A report can match tool output while describing a harmful change; keep the scorer and answer material outside the agent's reach.
Worth knowing

Neon-sponsored walkthrough using synthetic store data, not an independent benchmark. The maker reports that 10 of 15 runs accessed scorer, answer or other-run material, compromising comparisons. The selected-context result is one reported run, not a general performance finding. A probabilistic gate is not an injection-protection guarantee or a security boundary. The verifier's reported 0.99 is not measured task accuracy; TypeSafe's Choice and Score confidence summarize answer distributions, and Noul has no separate confidence field.

Building an Agent Harness with Jev: What Actually Works?

Keep this for later

Sign in to bookmark experiments, guides and videos, and keep notes only you can see.

Continue to sign in

We’ll bring you back to this listing.