Read & learn
Written guides.
Understand Jev, one idea at a time. Walkthroughs, recipes, and writeups — our notes first, the original next.
Learn, build, and play with Jev — from deep technical guides to creative experiments.
Read & learn
Understand Jev, one idea at a time. Walkthroughs, recipes, and writeups — our notes first, the original next.
Watch & learn
See an idea take shape. Tutorials, demos, and deep dives, organized by topic and credited to their creators.
Make & explore
See what builders made with Jev. Games, tools, repositories, articles, and more — traced to their sources.

Every category, from games and tools to repositories and writeups, traced to a primary source.
Same ideas.
Different paths.
This is the full registry, not just the featured picks. Figures like speed, cost, and stars are maker-reported or captured snapshots, not JevMade measurements.
268 experiments · showing 61–120
Maps 1,000 recent AI papers into 24 topics with Jev, then compares a seeded sample with an LLM judge.
jevfire batches many Jev-style decisions over one shared context on CUDA language models.
LegalForecastBench provides an alpha benchmark and evaluation workflow for LegalForecast-MTD.
A word-level language model that drafts with n-grams and asks Jev to check chunks of the result.
typesafe-local asks a local MLX model typed questions and returns calibrated probabilities without generating or parsing prose.
Verdict is a 151M-parameter ModernBERT decision engine with calibrated uncertainty and no autoregressive generation.
Organizes independent Jev studies by calibration, consistency, prompt injection, abstention, and other failure modes.
Reimplements the Jev request contract over existing SGLang or vLLM deployments by scoring finite candidate tokens.
Tests whether filtering coding-agent tool results through Jev improves task success, context size, or cost against an unfiltered arm.
Split local Laya batches across Apple Neural Engine and GPU with a Jev-shaped API; it accelerates an independent model, not Jev.
Rev trains three Qwen decision models and compares their accuracy, latency, cost, scaling, and Snake play with hosted Jev.
Follow Eikos and Jev as they paper-trade the same Hyperliquid markets from identical snapshots and precommitted rules.
How well does Jev understand Korean, including medical text? This benchmark publishes questions, results, runtime and cost.
system-one-open recreates Jev-style calibrated decisions with small open Gemma models.
LitJev wraps ordinary language models in a Jev-like layer for typed choices, scores, and binary judgments.
jevcal derives confidence thresholds, checks calibration, and watches typed decision models for drift against an LLM teacher.
Surveys evidence on Jev and compatible decision models, separating calibration claims, selective control, and open implementations.
Test whether Jev can choose compiler optimisation passes from program features, with its decisions compared against measured outcomes.
Build and test Jev decision workflows in Python, then use them from the command line or an MCP client.
A closer look at Jev 1.13.0, with controlled prompts and the raw answers available to inspect.
jev-chat constructs chatbot replies by repeatedly turning Jev probabilities into hierarchical speculative decoding choices.
Can Jev find better agent skills than embedding search? This evaluation tests both on Chinese and English queries.
jev_stock experiments with short-term market-direction forecasts derived from structured financial data.
This open-weights Jev alternative returns typed, calibrated choices in one Hugging Face or vLLM forward pass.
Sort through training data with a Rust tool that asks Jev to score rows and filter out unwanted examples.
Split source code mechanically, ask Jev what each fragment is, and render the guesses as syntax highlighting.
Compare Jev with another judge in a blind arena of bounded answer choices, then reveal the models after voting.
Tests Jev and the open-weight Laya counterpart as typed judgment layers in a look-ahead-free quantitative research stack.
Ask typed questions directly about images through a shared Qwen3-VL encoding instead of first producing captions.
Calibre measures which model suits each part of a dataset, then routes requests using those results.
Test Jev on chess puzzles and on figuring out which game character you're talking to.
Eight small examples apply Jev to mechanical and electrical engineering decisions.
jev-gate experiments with assigning Claude Code tasks to different models according to Jev's judgment.
jev-harness adds confidence gates, shadow runs, reusable recipes, and evaluations around Jev decisions.
jevgpt coaxes a non-generative decision model into chatting by choosing the response one step at a time.
trade-jev backtests Jev as a buy, sell, or hold trader on Nasdaq futures order-book data.
jevify helps agents spot suitable Jev tasks, frame typed questions, and draw on recent community experiments.
Play chess against Jev or Stockfish, or let the two models play each other.
Turn NYC 311 complaint text into explorable maps of reporting volume and Jev-estimated impact.
Before building an RPG, this project tests whether Jev's characters can react to context and keep a secret.
Metask Jev-Lab trains open typed-decision models and publishes calibrated option probabilities alongside reproducible JevBench evaluations.
DecisionBridge gives existing language models an interface for explicit choices, scores, calibration, and human-review thresholds.
This benchmark compares Jev with a strong language model at attributing agent failures in the text subset of Who&When Pro.
This research notebook pairs runnable Jev demos with a close reading of the model's public claims.
Build a chat response one character at a time, with every character chosen by Jev.
Little Airways puts Jev's decision-making into a small browser-based flying demo.
This benchmark compares Jev with Cohere, ZeroEntropy, and a chat model on 14 reranking datasets.
This harness evaluates Jev Ultrafast on curated research-browser cases and generates a field report from each run.
jev-sec-bench tests Jev blindly on prompt-injection and vulnerable-code detection tasks.
This demo compares Jev's abstract-screening decisions with gold labels from the ASReview SYNERGY dataset.
This static dashboard compares Jev and two Luna settings on Japan's 2026 Common Test.
Explore a Monte Carlo tree search that uses Gemini to propose paths and Jev to judge them.
PadFlow contributes anonymized land-development decisions for testing confidence-aware models such as Jev.
qwen-rlcd trains Qwen3.5-0.8B to make Jev-style choices, scores, and binary judgments with calibrated confidence.
RISC-jeV asks how far state-in, decision-out inference can go by wiring Jev up as a RISC-V CPU.
These OpenJev experiments classify sentence pairs, including whether one statement follows from or contradicts another.
A practical guide and prototype for putting Jev triage in front of Herdr-managed coding agents.
Experiment with a Jev-powered Codenames player and a separate CLI for checking determinism, sensitivity and calibration.
An experimental pharmacy checker that asks Jev the same question five ways and looks for agreement.
A record of five attempts to find a worthwhile job for Jev—and the reasons each stayed out of production.
Jev, briefly
A chat model generates text token by token and hopes you parse it. Jev never generates a word. Your code sends the state of the world plus typed questions; Jev returns every answer in one parallel pass — typed values with calibrated probabilities and a confidence score your code can trust. About 70–500 ms end to end, $0.042 per million input tokens, output free. Trained with what TypeSafe calls Reinforcement Learning for Calibrated Decisions.
You define the options. Jev returns the chosen option, a probability for every option, and a confidence. The workhorse of routing, agents, and games — the answer space is always legal, so there is nothing to parse and nothing to hallucinate.
You describe ordered levels — say trivial / normal / critical. Jev returns the level, the probability of each, and a confidence. Scores turn fuzzy judgment (“how severe is this log line?”) into a number-free decision your code can branch on.
You assert a statement; Jev returns the probability that it is true — a noul. Moderation, verification, guardrails, “does this diff actually fix the bug?”: one question, one calibrated probability.
Official material lives at typesafe.ai and docs.typesafe.ai. JevMade is an independent community registry — not affiliated with or endorsed by TypeSafe AI.
Recurring lessons
Games, drones, trading bots, and browser agents all converge on the same loop: serialize the state, ask one decisive question, act, repeat. Jev’s latency makes the loop feel instant — the model lives inside the control loop, not outside it.
The answer says what; the confidence says whether to act. The most reliable entries threshold on confidence to route edge cases to a slower model or a human — automation with an honest escape hatch.
Makers rarely ask Jev to pick from everything. Local tactics prune 225 gomoku moves to ~40; DOM filters turn a page into an element table; code narrows options, Jev judges within them.
A recurring split: Jev makes every decision cheaply and instantly, and a small LLM is only invoked when a human-facing string must actually be written. Decision and generation are separate budgets.
Send the link — repo, live demo, post, or video — plus a line on what it does and which primitives it uses. Every entry is verified against its primary source before it ships.
Submit your experiment → hello@JevMade.com