Before you dive in
What you’ll find in the original
- Keep measured state, semantic intent, action choice, validated Cartesian control, physics, and task evaluation as separate loop stages.
- Pair policies on identical seeds and resume without deleting completed failures.
- Report Wilson intervals: ten successes in ten trials still supports a wide 72.2–100.0% interval.
Worth knowing
The results cover ten trials per task under privileged-state simulation, not vision or a physical robot. Several rule baselines match or beat JEV, and JevMade did not rerun the paid-policy trials.