JevMade

Sign in
← Back to experiments

Benchmarks & research

SearchJev + SearchDecision-Bench

SearchJev takes a search agent's quick calls, such as whether a page is relevant or enough evidence is in, and leaves planning and writing to a larger model.

Source screenshot of SearchJev + SearchDecision-Bench
SOURCE SCREENSHOTFull screenshot ↗

What it does

Inspired by Jev, the authors fine-tune 0.8B and 4B Qwen3.5 models to score allowed options directly, train them on uncertain soft labels, and release SearchDecision-Bench with six decision types. On BrowseComp-Plus they report 54% accuracy for the 4B dual-system agent, against 45% for the large model alone and 46% with Jev 1.13 in the same role.

Maker-reported (not independently measured by JevMade): Authors report decisions 5.2–5.3 times faster than same-size Qwen3.5 models writing JSON, with 41–74% lower average calibration error on SearchDecision-Bench. · Authors report that Jev 1.13 remained stronger on navigation decisions, while both SearchJev sizes improved sufficiency, routing, rewriting and verification.

Primitives
choice
Added

Keep this for later

Sign in to bookmark experiments, guides and videos, and keep notes only you can see.

Continue to sign in

We’ll bring you back to this listing.