A candid security-evaluation walkthrough that explains how a comparison table was produced, what a contamination experiment found, what AgentDojo measured instead, and what remains unmeasured.
Original by 0xShin0221GuardrailsRepository READMESource reviewed
Before you dive in
What you’ll find in the original
Describe how each table row was produced before interpreting apparent guardrail performance.
Keep prompt-contamination behavior separate from AgentDojo’s agent-level security measurement; they answer different questions.
Publish unmeasured questions and first-pass mistakes alongside tuning instructions so readers do not overgeneralize the results.
Worth knowing
The repository explicitly says the contamination experiment was not the AgentDojo guardrail measurement.