JevMade Sign in
← Back to guides

JevMade field notes / Evaluation framework guide

Testing an AI assistant's past actions with multiple-choice questions

This guide explains a tool that tests how well an AI assistant follows instructions. It asks multiple-choice questions to check if the software used the right tools, stayed on topic, or made mistakes.

Original by OpenlayerGuardrails

Listen to this guide

JevMade’s plain-English explanation

0:00 /

Our summary

This software helps developers test AI assistants by checking their past actions. It uses Jev, an AI tool that chooses from options rather than writing an answer. Developers use this to review if the assistant followed safety rules or supported its claims with evidence.

The software groups several specific questions together and sends them as a single request. This checks a single record of the assistant's actions all at once. Developers can set strict rules to stop the assistant from using executable tools if a test fails.

This is useful for developers testing their software, but different AI tools score tests differently. Developers must adjust their passing grades if they switch tools. The author measured a 6,015-request test run, but did not actually test the Kev and Laya interfaces.

Key takeaways

  1. Group independent checks together to test a single record of the AI assistant's past actions.
  2. Adjust your passing scores whenever you change the underlying AI tool used for testing.
  3. Set strict failure rules for tests that control executable tools capable of making permanent changes.

The reported 6,015-request test run and connection delays are the author's own measurements. The software supports the Kev and Laya interfaces, but the author did not actually run tests against them.

GitHub repository

Read the original guide Opens the author’s site in a new tab.

Keep this for later

Sign in to bookmark experiments, guides and videos, and keep notes only you can see.

Continue to sign in

We’ll bring you back to this listing.