JevMade Sign in
← Back to guides

JevMade field notes / Production-request replay and provider comparison

Compare decision models without losing useful answers

Knowviq replays its search checks through Jev, Clef, Clef-flash, d1 and Perplexity, finding that similar rankings can still hide very different numbers of useful answers.

Original by KnowviqEvaluation

Listen to this guide

JevMade’s plain-English explanation

0:00 /

AI narration

Credits

“Decision models compared: Jev, Clef, d1 and Perplexity” by Knowviq. Read the original source.

This expanded guide is an AI-narrated adaptation prepared by JevMade. It expands the source’s essential ideas, examples and caveats in JevMade’s own words and is not a word-for-word reading. The synthetic voice does not imitate the author or imply their endorsement.

Our summary

A search website needs to decide which results answer a question and which written sentences have enough evidence. Knowviq compares five models on 938 requests taken from its own work. The article asks a practical question: does changing the model preserve the answers visitors should actually see?

Knowviq sends the same inputs to each model, then applies its website's rules for showing results and sentences. Copying Jev's cutoffs makes Clef reject many relevant results. The team adjusts Clef's settings and compares correct answers, hidden useful answers, delays and actual bills, rather than looking only at the advertised price per token.

These are Knowviq's own reported tests, not an independent repeat. Some answer labels can miss a reworded fact or match words out of context; only 50 sentence pairs form a random sample. Settings were adjusted on the same data. The GLiDE update contains vendor claims and estimated costs, not another tested model.

Key takeaways

  1. Compare both mistakes and coverage: a model that hides almost everything can look cautious while withholding useful answers.
  2. Do not transfer a cutoff unchanged just because two services accept the same questions. Their probability values can behave differently.
  3. Compare bills for identical requests. Providers may count very different numbers of tokens, the pieces of input used for billing.

The reported delays depend on location, large inputs and Clef's development proxy; they are not a fair comparison with vendor speed claims. False-show rates rely on 12 natural negatives plus 100 artificial mismatches. d1 used a rate-limited free tier. GLiDE's published prices conflict. No runnable raw evaluation artifact was inspected.

Knowviq · Original published

Read the original guide Opens the author’s site in a new tab.

Keep this for later

Sign in to bookmark experiments, guides and videos, and keep notes only you can see.

Continue to sign in

We’ll bring you back to this listing.