What it does
Inspired by Jev, the authors fine-tune 0.8B and 4B Qwen3.5 models to score allowed options directly, train them on uncertain soft labels, and release SearchDecision-Bench with six decision types. On BrowseComp-Plus they report 54% accuracy for the 4B dual-system agent, against 45% for the large model alone and 46% with Jev 1.13 in the same role.