Our summary
This software helps developers test AI assistants by checking their past actions. It uses Jev, an AI tool that chooses from options rather than writing an answer. Developers use this to review if the assistant followed safety rules or supported its claims with evidence.
The software groups several specific questions together and sends them as a single request. This checks a single record of the assistant's actions all at once. Developers can set strict rules to stop the assistant from using executable tools if a test fails.
This is useful for developers testing their software, but different AI tools score tests differently. Developers must adjust their passing grades if they switch tools. The author measured a 6,015-request test run, but did not actually test the Kev and Laya interfaces.
Key takeaways
- Group independent checks together to test a single record of the AI assistant's past actions.
- Adjust your passing scores whenever you change the underlying AI tool used for testing.
- Set strict failure rules for tests that control executable tools capable of making permanent changes.