What it does
Calibre measures which model suits each part of a dataset, then routes requests using those results.
Benchmarks & research
Only you can see your notes.
Screenshot unavailable. Open the experiment ↗
Calibre measures which model suits each part of a dataset, then routes requests using those results.
To start, collect a few hundred sample customer sentences with their true answers, such as banking queries. Write short descriptions for every category you want to sort. Your developer can install the tool to test a fast AI model against a larger, more expensive one to see if choosing between them actually saves money.
If the report shows combining both models works, your app asks the smaller model first. When that model is unsure, your app sends the text to the larger model instead. The exact score where the smaller model should give up only applies to the specific examples and AI tools you measured.