Our summary
A search website needs to decide which results answer a question and which written sentences have enough evidence. Knowviq compares five models on 938 requests taken from its own work. The article asks a practical question: does changing the model preserve the answers visitors should actually see?
Knowviq sends the same inputs to each model, then applies its website's rules for showing results and sentences. Copying Jev's cutoffs makes Clef reject many relevant results. The team adjusts Clef's settings and compares correct answers, hidden useful answers, delays and actual bills, rather than looking only at the advertised price per token.
These are Knowviq's own reported tests, not an independent repeat. Some answer labels can miss a reworded fact or match words out of context; only 50 sentence pairs form a random sample. Settings were adjusted on the same data. The GLiDE update contains vendor claims and estimated costs, not another tested model.
Key takeaways
- Compare both mistakes and coverage: a model that hides almost everything can look cautious while withholding useful answers.
- Do not transfer a cutoff unchanged just because two services accept the same questions. Their probability values can behave differently.
- Compare bills for identical requests. Providers may count very different numbers of tokens, the pieces of input used for billing.