JevMade Sign in
← Back to experiments

Benchmarks & research

jev-search-rerank-eval

Source screenshot of jev-search-rerank-eval
SOURCE SCREENSHOTFull screenshot ↗

What it does

Can Jev find better agent skills than embedding search? This evaluation tests both on Chinese and English queries.

How you can use it

Your developer can adapt the scoring method from this experiment to improve a search feature. When someone searches, your app sends the initial results to TypeSafe. TypeSafe is an outside AI service that scores these matches. Your app then combines these new scores with the original search order.

This approach only reorders the results it receives. If your initial search misses a useful item completely, the outside service cannot find it. To borrow this method, your developer needs an access key that connects your app to TypeSafe.

Maker-reported (not independently measured by JevMade): paired bootstrap 95 % confidence intervals over queries · 9,831 (query, skill) pairs, 51–60 per query. 6,718 pairs were proposed by a single

Primitives
choice, score
Platform
Python
Added
Project created
GitHub stars
3 (snapshot, not live)