Read & learn
Written guides.
Understand Jev, one idea at a time. Walkthroughs, recipes, and writeups — our notes first, the original next.
A public registry · independent & community-run
Learn, build, and play with Jev — from deep technical guides to creative experiments.
Read & learn
Understand Jev, one idea at a time. Walkthroughs, recipes, and writeups — our notes first, the original next.
Watch & learn
See an idea take shape. Tutorials, demos, and deep dives, organized by topic and credited to their creators.
Make & explore
See what builders made with Jev. Games, tools, repositories, articles, and more — traced to their sources.

Every category, from games and tools to repositories and writeups, traced to a primary source.
Same ideas.
Different paths.
This is the full registry, not just the featured picks. Figures like speed, cost, and stars are maker-reported or captured snapshots, not JevMade measurements.
1613 experiments · showing 1261–1320
An MCP server that lets a coding agent drive a real Chrome browser through Jev: the model answers small typed questions and the tool acts when it is confident.
Puts the Traditional Chinese TMMLU+ exam to Jev, one four-way question per item, and scores it the way the official leaderboard does.
An agentic memory system where Jev makes the frequent memory decisions — what to store, how to type and connect it, how retrieval routes — and an LLM writes the answer.
A local playground that shows the request as you build it: a form on the left and the JSON that will be sent on the right.
Pairs a curated map of the Jev ecosystem with independent, reproducible benchmarks of the three primitives, confidence gating, fan-out latency and agent control.
A System One model plays Craftax while an LLM sets the goals, picking one macro option per step or one of seventeen primitive actions.
Reduces 845 preflop spots to a single fold, call or raise decision, asks Jev five times each, and compares the result with four deterministic reference styles.
Does Jev pick better stocks than a mechanical momentum rule? A four-year backtest of an O'Neil strategy where Jev chooses the entries and the exits.
An OpenAI- and Anthropic-compatible gateway where a small System One model picks which larger model should answer each prompt.
Turns Jev's probabilities into a routing threshold with a provable bound on how many queries get silently misrouted.
Lints Japanese prose by asking, per sentence, for the probability of a typo, a twisted subject and predicate, an over-long sentence or a repeated phrase.
Ten runnable examples plus more than 120 use cases, four composition patterns and a theory write-up for building with Jev.
Benchmarks Jev on Brazil's ENEM 2025 exam against open-LLM baselines, contrasting a raw transcribed state with a structured one and measuring calibration.
A Swift 6 bridge that turns strongly typed @Generable structs and enums into Jev questions and returns the calibrated probabilities as decisions.
A browser extension that turns YouTube into a focused learning feed: Jev answers one narrow question per video and anything uncertain stays hidden.
Paste a public URL and Jev judges its first screen: what wall greets a visitor, how concrete the promise is, whether the call to action is obvious and whether trust signals show.
Flappy Bird where you race Jev on the same seeded course, with its typed action and confidence shown live under its board.
A pip-installable benchmark for typed System One models that measures what the probabilities buy: calibration against a noise floor, wording sensitivity, selective prediction and cost.
Two System One models fight a real Doom deathmatch — open-weight Laya locally against hosted Jev — from the same compressed state and the same questions.
A mailbox triage proof of concept: a fake IMAP server feeds a listener that asks Jev one batched call per email and labels it on two independent axes.
Five hospital tasks, a thousand synthetic patients, and Jev answering every one, with Claude writing the scenarios and reviewing the results.
An MCP server and agent skill for using System One at design time: decompose a judgment into questions, lint them, calibrate thresholds on labelled rows and rerank.
A local agent harness that puts Jev in front of an ordinary LLM and pays for the LLM only when the work needs prose.
A Taboo-style browser game for design systems: describe a UI component without naming it while Jev re-ranks all 135 components and shows the whole distribution.
Browser tic-tac-toe where you play X and Jev plays O, with move probabilities, threat assessment and latency in a side panel.
Autonomous Tetris where a local analyzer enumerates and scores every legal landing and only the best twelve reach Jev, which picks exactly one.
A village where every villager asks Jev what to do next each tick, and a meter shows what those decisions cost against a frontier chat model.
A coding-agent CLI that splits the work: a decision model picks the next tool, scores progress and judges whether the goal is reached, while an LLM only fills in arguments and writes code.
Turns an inbox into a short action queue: seven typed questions per thread, with plain Python deciding whether it needs you and what the next step is.
A character-level language model built out of a classifier: the vocabulary becomes 28 Choice options and generation is an ordinary loop over the distribution.
Pre-flights a System One question before you put its number behind an if, reporting a decision flip rate rather than a confidence score.
A Chrome extension that scores each post in an X timeline on five dimensions — firsthand experience, self-promotion, engagement bait, technical depth and relevance.
A browser agent with no LLM in the loop: code turns the page into a closed set of actions and Jev decides the next one, its target, and whether the task is done.
Tests Jev as a search reranker, and asks whether its confidence can tell you which queries are worth spending more compute on.
Lets Jev choose bounded browser actions over Chrome CDP observations while a separate model writes form text.
Tests Jev pairwise rankings of FOMC statements against rate decisions, FedLock scores, and a Haiku comparison.
Classifies sampled Bluesky posts with eight Jev questions and sends uncertain judgments to a human-review lane.
Routes Home Assistant conversations, resolves entities, and selects device actions through TypeSafe System One decisions.
Gates agent tool calls through detectors, Jev judgments, and deterministic policy before allowing, blocking, or requesting approval.
Measures how Jev accuracy and calibration change when the same tasks are presented in English and Spanish.
Provides guarded macOS computer use with opaque targets, approval-bound mutations, fresh observations, and optional Jev workflow recommendations.
Reports a preregistered adversarial evaluation spanning calibration, batching, answerability, option counts, and out-of-domain logic.
Lets Claude delegate a browser goal while Jev repeatedly chooses the next click, field, or action.
Searches an Obsidian vault locally, then—only with approval—sends shortlisted excerpts to Jev for reranking.
Reads completed Claude Code or Codex sessions and assigns typed judgements to each turn for later analysis.
Wraps Jev with deterministic caching, confidence calibration, memory, and runtime guardrails.
Moderate community posts with per-category probabilities and thresholds exposed through bots, a CLI, libraries, HTTP and MCP.
Blur LinkedIn posts that match your low-value-content rubric while leaving uncertain or failed judgments visible and offering one-click reveal.
Put many MCP servers behind two tools while Jev discovers the best enabled tool and the calling agent supplies its schema-bound arguments.
A Rust CLI turns Jev judgments into stable JSON, resumable batches, assertions, and meaningful shell exit codes.
A small CLI ranks local Agent Skills against a request, then returns suggestions without invoking anything automatically.
Guide a World of Warcraft character through leveling decisions while reserving expensive generation for harder moments.
Add three Claude Code hooks that flag dangerous commands, failed tool calls, and incomplete work with Jev.
Call TypeSafe’s System One API from Go with typed requests, answers, retries, and validation.
Overlay Android screens with warnings when Jev identifies manipulative interface patterns in captured accessibility context.
Run Jev-compatible typed judgments locally in a browser or Node server using small natural-language inference models.
Triage customer messages, assess account health, and check outbound drafts through a shared typed decision pipeline.
Compare sustainability-report claims with their evidence and surface disagreements across GRI, ESRS, BRSR, and ISSB mappings.
Automatically mark new Feedbin articles as read when Jev matches them to unwanted categories such as ads or crypto.
Research web and on-chain questions from the terminal, then show claims only after Jev-backed checks.
Jev, briefly
A chat model generates text token by token and hopes you parse it. Jev never generates a word. Your code sends the state of the world plus typed questions; Jev returns every answer in one parallel pass — typed values with calibrated probabilities and a confidence score your code can trust. About 70–500 ms end to end, $0.042 per million input tokens, output free. Trained with what TypeSafe calls Reinforcement Learning for Calibrated Decisions.
You define the options. Jev returns the chosen option, a probability for every option, and a confidence. The workhorse of routing, agents, and games — the answer space is always legal, so there is nothing to parse and nothing to hallucinate.
You describe ordered levels — say trivial / normal / critical. Jev returns the level, the probability of each, and a confidence. Scores turn fuzzy judgment (“how severe is this log line?”) into a number-free decision your code can branch on.
You assert a statement; Jev returns the probability that it is true — a noul. Moderation, verification, guardrails, “does this diff actually fix the bug?”: one question, one calibrated probability.
Official material lives at typesafe.ai and docs.typesafe.ai. JevMade is an independent community registry — not affiliated with or endorsed by TypeSafe AI.
Recurring lessons
Games, drones, trading bots, and browser agents all converge on the same loop: serialize the state, ask one decisive question, act, repeat. Jev’s latency makes the loop feel instant — the model lives inside the control loop, not outside it.
The answer says what; the confidence says whether to act. The most reliable entries threshold on confidence to route edge cases to a slower model or a human — automation with an honest escape hatch.
Makers rarely ask Jev to pick from everything. Local tactics prune 225 gomoku moves to ~40; DOM filters turn a page into an element table; code narrows options, Jev judges within them.
A recurring split: Jev makes every decision cheaply and instantly, and a small LLM is only invoked when a human-facing string must actually be written. Decision and generation are separate budgets.
Send the link — repo, live demo, post, or video — plus a line on what it does and which primitives it uses. Every entry is verified against its primary source before it ships.
Submit your experiment → hello@JevMade.com