JevMade hello@JevMade.com
← Back to experiments

Benchmarks & research

haiku.rag

haiku.rag provides a reusable System One judge for benchmark verdicts such as whether two answers are equivalent.

Source screenshot of haiku.rag
SOURCE SCREENSHOTFull screenshot ↗

What it does

It packages the input, expected answer, and generated answer into one Noul rubric. Uncertain or failed verdicts can be handed to a configured fallback judge.

How you can use it

A developer can use this code to test if a chat tool gives correct answers. The software compares a new response against a known good answer. This checks if the two messages mean the same thing. The app groups the question and both answers together for an AI to review.

If the first review fails, the app sends the texts to a second reviewer. The default setup runs the AI directly on your own computer. It does not require an access key that connects the software to an outside service.