Read & learn
Written guides.
Understand Jev, one idea at a time. Walkthroughs, recipes, and writeups — our notes first, the original next.
Learn, build, and play with Jev — from deep technical guides to creative experiments.
Read & learn
Understand Jev, one idea at a time. Walkthroughs, recipes, and writeups — our notes first, the original next.
Watch & learn
See an idea take shape. Tutorials, demos, and deep dives, organized by topic and credited to their creators.
Make & explore
See what builders made with Jev. Games, tools, repositories, articles, and more — traced to their sources.

Every category, from games and tools to repositories and writeups, traced to a primary source.
Same ideas.
Different paths.
This is the full registry, not just the featured picks. Figures like speed, cost, and stars are maker-reported or captured snapshots, not JevMade measurements.
268 experiments · showing 121–180
Switchloom experiments with Jev-routed persistent coding agents and publishes pilots where routing failed to beat a single Astra model.
This experiment tests Jev on financial-market prediction and reports that its forecasts perform poorly.
jev-alpha-bench separates news comprehension from market recall to test whether Jev can predict stock returns.
This benchmark compares Jev, Gemini, and GPT on structured annotation of São Paulo court judgments.
This project experiments with Jev as an evaluator rather than a task-solving model.
A position paper arguing that Jev-style models need fuzzy and Hidden Markov primitives before they settle on crisp decisions.
Jev picked the winner in 64.5% of 10,984 randomized Upworthy headline tests.
This project tests whether Jev can resolve ambiguous Japanese addresses against Japan Post's KEN_ALL data.
jev-lab collects TypeScript experiments on Jev's behavior, accuracy, and response latency.
A phishing test that pits Jev against Claude Haiku 4.5 on 2,000 emails.
This reproducible MuJoCo pilot compares Jev, Claude Haiku, and reactive rules on pick-and-place control.
This game playground compares Jev with other evaluators using explicit states, legal moves, and observable outcomes.
HackSing puts Jev through 50 tests and publishes the findings in a 52-page Chinese-language report.
This experiment measures Jev as a lower-cost model router on the RouterArena benchmark.
Can Jev tell a real credential from an innocent string? This benchmark tests it on file snippets.
A test of whether Jev can tell genuine shadcn-ui lint problems from false alarms.
This notebook compares zero-shot spam judgments from Jev with conventional TF-IDF classifiers.
This classifier distinguishes agent-written from human-written pages in the collusion.wiki corpus.
This repository tracks ongoing research into Jev and System One models as a set of Markdown slides.
This app extracts data from a document according to a supplied JSON schema using parallel constrained decoding.
A Jev-inspired experiment that makes visual decisions on an iPhone using Qwen3-VL, not TypeSafe's service.
This study tests Jev as a fast gate for suspicious agent actions in SHADE-Arena.
This Rust port runs Jev-style Choice, Score, and Noul evaluations through ordinary language models.
system-one-gemma adds a scoring head to Gemma 3 270M for calibrated decisions without text generation.
These charts place Jev's Thai standardized-exam results alongside those of 110 other models.
A chess evaluation that gives Jev an explicit list of moves and measures how well it chooses.
typesafe-oracles tests when a typed Jev judgment is more useful than a conventional language-model call.
This experiment places Jev before each proposed agent action and blocks steps judged unsafe.
This experiment recreates Jev-like decisions with Qwen 3.8 27B running on Cerebras.
Juan Macías tries Jev on Spanish data-protection documents and shares where it succeeds.
See how Jev fares in solved poker spots—and how much the wording of the game state changes its play.
Jev and an OpenRouter model race to score on identical Snake boards.
Match Jev against a text model in chess, where Jev selects legal moves and its opponent must produce valid notation.
Follow a Jev-guided ROS 2 turtlesim route from an early control failure to a corrected four-waypoint run.
Compare Jev's driving choices when its road description comes from perfect simulator state or a vision model.
Replay a documented Jev-versus-Stockfish tournament with legal moves, raw responses, complete games, and provisional ratings.
A Flappy Bird test that grades Jev's decisions and lets you replay the results.
Play four-seat Texas Hold'em while comparing Jev's betting judgments with a rules coach and Monte Carlo equity estimate.
Play terminal Sudoku yourself or watch Jev choose cell-and-digit moves while the benchmark tracks whether it solves the grid.
Compare Jev with a conventional language model on route planning, Minesweeper, Hanoi, and a simulated drone.
Jev picks strategic preferences while a traditional engine searches for moves on a chess.com board.
Compare ways of showing Jev candidate moves in a head-to-head arena for two falling-block games.
Compare Jev’s Battleship shots with random, hunt-and-target, and placement-density strategies on seeded games.
Compare 21 ways of describing legal SameGame moves to Jev, with every raw request and response preserved.
Pit Jev against random and hand-tuned players in Tetris, Snake, and 2048 while logging confidence for every move.
Triage simulated warehouse-robot incidents with Jev and compare its decisions with a small local classifier.
Compare Jev with general language-model judges on one multilingual booking-inquiry routing task.
A Chinese-language experiment suite for Jev's context limits, bilingual judgments, batching and game decisions.
An original platformer and reinforcement-learning lab testing whether frozen Jev risk features help PPO learn faster.
Political Compass and Pew typology probes for Jev, alongside a deliberately awkward character-by-character text decoder.
Battleship where Jev returns a probability of a ship for every open cell, creating both the heat map and the shot.
A roguelike dungeon director where Jev sketches the next room's role, size, danger, enemies, loot and secrets.
Runs Jev and a frontier model down parallel lanes, answering the same timed traffic-hazard choices and comparing the scorecards.
A controlled city simulation comparing four kinds of institutional memory, with Jev scoring each term beside a deterministic scoreboard.
Turns Jev probability distributions into one plain-language sureness verdict, from certain through torn to clueless.
Run open-Jev decisions in your browser without an API key or inference server.
Check generated answers against several rubrics, with Jev judging them in batches.
Test what an answer means, not just which words it contains, with Jev assertions for Vitest and Jest.
Compare Jev’s typed classifications with classical ML pipelines across the maker’s eight datasets.
Replay six Jev chess experiments, from unaided legal moves to tactical filters and Stockfish-assisted variants.
Jev, briefly
A chat model generates text token by token and hopes you parse it. Jev never generates a word. Your code sends the state of the world plus typed questions; Jev returns every answer in one parallel pass — typed values with calibrated probabilities and a confidence score your code can trust. About 70–500 ms end to end, $0.042 per million input tokens, output free. Trained with what TypeSafe calls Reinforcement Learning for Calibrated Decisions.
You define the options. Jev returns the chosen option, a probability for every option, and a confidence. The workhorse of routing, agents, and games — the answer space is always legal, so there is nothing to parse and nothing to hallucinate.
You describe ordered levels — say trivial / normal / critical. Jev returns the level, the probability of each, and a confidence. Scores turn fuzzy judgment (“how severe is this log line?”) into a number-free decision your code can branch on.
You assert a statement; Jev returns the probability that it is true — a noul. Moderation, verification, guardrails, “does this diff actually fix the bug?”: one question, one calibrated probability.
Official material lives at typesafe.ai and docs.typesafe.ai. JevMade is an independent community registry — not affiliated with or endorsed by TypeSafe AI.
Recurring lessons
Games, drones, trading bots, and browser agents all converge on the same loop: serialize the state, ask one decisive question, act, repeat. Jev’s latency makes the loop feel instant — the model lives inside the control loop, not outside it.
The answer says what; the confidence says whether to act. The most reliable entries threshold on confidence to route edge cases to a slower model or a human — automation with an honest escape hatch.
Makers rarely ask Jev to pick from everything. Local tactics prune 225 gomoku moves to ~40; DOM filters turn a page into an element table; code narrows options, Jev judges within them.
A recurring split: Jev makes every decision cheaply and instantly, and a small LLM is only invoked when a human-facing string must actually be written. Decision and generation are separate budgets.
Send the link — repo, live demo, post, or video — plus a line on what it does and which primitives it uses. Every entry is verified against its primary source before it ships.
Submit your experiment → hello@JevMade.com