Our summary
JevAny is an AI tool that chooses from a list of options rather than writing out an answer. People use this kind of software to make decisions in software tasks, like sorting customer support tickets or picking the next step for a robot, based on text, images, or video.
The author tests the software by making sure the test questions and media were never used during the tool's training. Instead of just reporting one overall score, the tests record specific failures and compare results against blank or mixed-up images to prove the software actually uses the visual information.
This guide is useful for developers who want to rigorously test AI decision software. In these tests, the software struggled with opposite actions and video scoring. The author also notes that experimental updates are kept out of the main release if independent tests do not confirm they actually improve performance.
Key takeaways
- Remove matching text and media from your test data to get a true measure of performance.
- Publish specific task failures and ranges of results instead of relying on one overall success score.
- Leave experimental updates out of the main release if independent tests do not confirm the improvement.