Before you dive in
What you’ll find in the original
- A recurring architecture keeps a large model for planning or writing while Jev chooses browser actions, routes models, filters retrieval results, or reviews proposed tool calls.
- Typed output prevents invalid labels but not wrong valid labels; the post highlights one test where Jev found six of seven planted defects while Fable found seven.
- Early headline figures varied substantially by workload: the post contrasts a reported 200× claim with a separate 25× test and treats maker-reported costs and latencies as claims rather than universal rates.
Worth knowing
This is a secondary survey assembled from launch-week posts. Most figures are reports by project authors, not measurements reproduced by Van Horn, and several featured projects have fuller primary sources elsewhere.