What it does
It mirrors TypeSafe's system-one-adapter and can return normalized probabilities.
Benchmarks & research
This Rust port runs Jev-style Choice, Score, and Noul evaluations through ordinary language models.
Only you can see your notes.
Screenshot unavailable. Open the experiment ↗
It mirrors TypeSafe's system-one-adapter and can return normalized probabilities.
A developer uses this tool to test how different AI services judge text. They can compare these new services against an outside AI service named TypeSafe. For example, your app might read a book review to decide if it is positive. Your developer connects a different AI service to answer the question.
Sometimes an AI formats its answer incorrectly. The software catches this mistake and asks the service to try again automatically. However, some outside services limit how much text they return. If a service reaches this limit, the tool stops. Your developer must raise that limit or ask fewer questions.