JevMade hello@JevMade.com
← Back to experiments

Benchmarks & research

Jev RL

CartPole, MountainCar, Acrobot, and FrozenLake trained with reinforcement rewards that come from Jev's calibrated scores instead of a hand-written reward function.

Source screenshot of Jev RL
SOURCE SCREENSHOT · source ↗ · captured 2026-09-29Full screenshot ↗

What it does

The judge layer runs over OpenRouter or TypeSafe direct, so the same reward signal works through either gateway, and a small web app replays the training runs.

Maker-reported (not independently measured by JevMade): 36 training runs · 3 seeds · 3.24 million environment steps · Recorded JEV API cost ≈ $0.00241 — 129 judgments, cached and reused across 12 training runs

Primitives
score
Platform
Python
Added
Project created

Source checked 2026-09-29 — opened the primary source directly.