JevMade Sign in
← Back to experiments

Benchmarks & research

S1Rank

Tests Jev as a search reranker, and asks whether its confidence can tell you which queries are worth spending more compute on.

Bookmark: S1Rank Keep this in your collection.
Leave a noteWhat would you try with this? : S1Rank

Only you can see your notes.

Source screenshot of S1Rank
SOURCE SCREENSHOTFull screenshot ↗

What it does

The routing idea did not work. It also found that 52% of probabilities change across byte-identical requests.

How you can use it

Your search project could put useful documents first by asking a yes-or-no question about each: does it answer the search question? A developer can send the documents together and sort by the AI's probability of yes. Try searches with known answers, since repeating a request can change its scores.

Maker-reported (not independently measured by JevMade): BM25 nDCG@10 50.6 / 48.0 / 59.5 / 32.2 / 67.9 across DL19, DL20, TREC-COVID, NFCorpus, SciFact · Jev joint pointwise 73.7 / 71.7 / 86.5 / 38.3 / 80.5 at 1 call per query, $0.57 per 1k queries · 52% of probabilities change across byte-identical requests

Primitives
noul, choice, score
Platform
Python
Added
Project created

Keep this for later

Sign in to bookmark experiments, guides and videos, and keep notes only you can see.

Continue to sign in

We’ll bring you back to this listing.