What it does
This harness evaluates Jev Ultrafast on curated research-browser cases and generates a field report from each run.
Benchmarks & research
Screenshot unavailable. Open the experiment ↗
This harness evaluates Jev Ultrafast on curated research-browser cases and generates a field report from each run.
For a research assistant project, start with the saved test results. Your developer can turn them into a web page showing each step the assistant took. You can browse that page without starting new AI work. Borrow this way of presenting results for your own tests, so you can inspect what happened rather than judge success from the final answer alone.