Tool and skill trials stopped because they saved very little time or cost. The project repeated one research test on the same data to tune its settings. General results remain inconclusive, and the author did not publish the raw measurement data.
How you can use it
Start by measuring Pi with its normal tools unchanged. That comparison run shows whether there is enough unnecessary material to justify filtering. Shadow mode records what the filter would remove without removing it; enforced mode actually changes the available tools. Neither mode proves the agent still completes its work correctly.
Use the included fictional research task to compare retained evidence and citations, not to claim better web search. The model estimates relevance; ordinary code preserves essential tools and restores the original tool set on missing answers, errors or timeouts. The published trials remain inconclusive, with settings tuned on the same research example.