Read & learn
Written guides.
Understand Jev, one idea at a time. Walkthroughs, recipes, and writeups — our notes first, the original next.
Learn, build, and play with Jev — from deep technical guides to creative experiments.
Read & learn
Understand Jev, one idea at a time. Walkthroughs, recipes, and writeups — our notes first, the original next.
Watch & learn
See an idea take shape. Tutorials, demos, and deep dives, organized by topic and credited to their creators.
Make & explore
See what builders made with Jev. Games, tools, repositories, articles, and more — traced to their sources.

Every category, from games and tools to repositories and writeups, traced to a primary source.
Same ideas.
Different paths.
This is the full registry, not just the featured picks. Figures like speed, cost, and stars are maker-reported or captured snapshots, not JevMade measurements.
269 experiments · showing 241–269
TypeLLM adds reasoning before type-safe output as an open-source alternative to Jev-style direct decisions.
A Venice API demonstration classifies 24,000 Hacker News posts into 12 categories with Jev.
This analyser processes viral X posts with Jev and compares its throughput and cost with Claude Opus 5.
A WebMCP benchmark run pairs Jev with Mercury 2.5 and reports solving every task at sharply lower model cost.
A measured browser-control trial separates Jev HTTP time from the full interaction loop while noting that network time remains included.
Analyze thousands of personal X posts across eight questions to compare which topics, hooks, and teaching styles correlate with engagement.
Decision Index 0.2 places XOR eighth among Jev-like text models, with comparatively strong results in arts and human taste.
Let Jev select sparse-attention rates across MiniMax H3 layers to shorten video generation on an RTX 4070.
wmoto ran a local decision-model prototype and published a ten-second demonstration while noting that its speed still needed work.
Jacky Ko and collaborators introduced an Apache-2.0 contrastive language model that connects states to actions as an open System One model.
Four ways of asking Jev to play 2048, compared with random moves and fixed rules, with replays of the resulting games.
A Phoenix Evals example pits Jev against a small grounded-versus-hallucinated answer benchmark.
haiku.rag provides a reusable System One judge for benchmark verdicts such as whether two answers are equivalent.
Run Laya’s typed decisions natively on Apple Silicon, then watch its terminal Snake demo choose moves with safety corrections.
Two days after Jev launched, Cribl's AI Research team published a grounded early look at what decision models can and cannot do inside a telemetry pipeline, benchmarking Jev against the purpose-built classifiers and LLM judges they already operate.
Datadog's Agent Observability team shows Jev doing double duty in their eval stack: notebooks score production agent spans as they arrive through online evals and run the same focused questions offline, mapping typed answers onto Datadog metrics.
SREGym gave its incident-solving agent a hesitant reviewer: before any diagnosis gets submitted, Jev has to agree the evidence supports it, and when a test is proposed, Jev ranks which one is worth running first. Pass rate moved from 20 to 24 out of 50.
CartPole, MountainCar, Acrobot, and FrozenLake trained with reinforcement rewards that come from Jev's calibrated scores instead of a hand-written reward function.
An open battleground that feeds Jev and Laya byte-identical rounds — same wording, same option order — and publishes the report card.
Hassan tested Jev on 1,018 paper summaries, assigning each to one of 24 topics for a reported eight cents.
Puts Jev's confidence numbers to the test on a task that was generated after the model shipped.
What does a probability model do when the answer is genuinely random? These experiments found out, and the answer is not flattering.
The fan-out docs say batching bills the state once. As far as anyone could tell, nobody had checked, so this repo checks with real money.
Databases sort by numbers and Jev hands back probabilities, so sorting search results by them is tempting. This measures how far the temptation carries.
Freeze a small Qwen model, train a decision head smaller than a million parameters, and it answers typed questions with zero generated tokens.
A 0.8B vision-language model playing ten browser games straight from pixels: one frame in, one move out, 43 ms on an H200.
Every voice agent transcribes before it decides anything. Prosodia asks the audio directly, so a misheard word can no longer poison the judgment downstream.
Bartosz Mikulski turns drawings into SVG coordinates and asks text-only Jev to choose which object each one depicts.
Shaun Mataire compares Jev’s multiple-choice and yes-or-no questions for matching product listings and research papers.
Jev, briefly
A chat model generates text token by token and hopes you parse it. Jev never generates a word. Your code sends the state of the world plus typed questions; Jev returns every answer in one parallel pass — typed values with calibrated probabilities and a confidence score your code can trust. About 70–500 ms end to end, $0.042 per million input tokens, output free. Trained with what TypeSafe calls Reinforcement Learning for Calibrated Decisions.
You define the options. Jev returns the chosen option, a probability for every option, and a confidence. The workhorse of routing, agents, and games — the answer space is always legal, so there is nothing to parse and nothing to hallucinate.
You describe ordered levels — say trivial / normal / critical. Jev returns the level, the probability of each, and a confidence. Scores turn fuzzy judgment (“how severe is this log line?”) into a number-free decision your code can branch on.
You assert a statement; Jev returns the probability that it is true — a noul. Moderation, verification, guardrails, “does this diff actually fix the bug?”: one question, one calibrated probability.
Official material lives at typesafe.ai and docs.typesafe.ai. JevMade is an independent community registry — not affiliated with or endorsed by TypeSafe AI.
Recurring lessons
Games, drones, trading bots, and browser agents all converge on the same loop: serialize the state, ask one decisive question, act, repeat. Jev’s latency makes the loop feel instant — the model lives inside the control loop, not outside it.
The answer says what; the confidence says whether to act. The most reliable entries threshold on confidence to route edge cases to a slower model or a human — automation with an honest escape hatch.
Makers rarely ask Jev to pick from everything. Local tactics prune 225 gomoku moves to ~40; DOM filters turn a page into an element table; code narrows options, Jev judges within them.
A recurring split: Jev makes every decision cheaply and instantly, and a small LLM is only invoked when a human-facing string must actually be written. Decision and generation are separate budgets.
Send the link — repo, live demo, post, or video — plus a line on what it does and which primitives it uses. Every entry is verified against its primary source before it ships.
Submit your experiment → hello@JevMade.com