Read & learn
Written guides.
Understand Jev, one idea at a time. Walkthroughs, recipes, and writeups — our notes first, the original next.
Learn, build, and play with Jev — from deep technical guides to creative experiments.
Read & learn
Understand Jev, one idea at a time. Walkthroughs, recipes, and writeups — our notes first, the original next.
Watch & learn
See an idea take shape. Tutorials, demos, and deep dives, organized by topic and credited to their creators.
Make & explore
See what builders made with Jev. Games, tools, repositories, articles, and more — traced to their sources.

Every category, from games and tools to repositories and writeups, traced to a primary source.
Same ideas.
Different paths.
This is the full registry, not just the featured picks. Figures like speed, cost, and stars are maker-reported or captured snapshots, not JevMade measurements.
268 experiments · showing 181–240
Can a cheap classifier label every agent failure in a training run? Applied Compute clusters a sample of traces, then has Jev do the labeling.
3,608 one-subject-swapped statements put to Jev as forced true/false picks, in three wordings and eleven languages, with every raw answer published.
Runs the CLASH cross-modal contradiction test against Jev with each image replaced by its COCO caption, then reports accuracy and modality bias.
Auto-marks 2,054 MR-GSM8K maths scripts against a teacher rubric with one Jev call each, then checks the marks against human annotators.
The open data behind an independent benchmark that puts Jev and six LLMs through the same typed questions on PubMedQA, Banking77 and HelpSteer2.
The 2026 Korean CSAT Korean-language paper turned into a JSON state and typed choice questions, with Jev's answers, probabilities and error analysis kept alongside.
An independent benchmark of Jev 1.13 against a cheap and a frontier LLM on SMS spam and Banking77, measuring accuracy, ECE, Brier, latency and cost.
A drop-in proxy that answers your app with Jev byte for byte while asking the same question of local doubles, then reports where they would have decided differently.
Measures what the live Jev API actually returns across eight realistic use cases, reporting cost, latency and response shape rather than rate-card arithmetic.
A tree with a place for every closed question a person or program could ask, filled with real questions answered by Jev, plus a 3D explorer where each question is a star.
Tests Jev on scientific judgement questions, then follows the consequences: what those choices do to the results that depend on them.
Reranks fixed MS MARCO candidate sets by asking Jev one relevance question per query and document pair, then scores the runs with trec_eval.
A first hands-on measurement of Jev: seven small scripts probing its documented jaggedness, latency scaling, calibration, negation consistency and behaviour when every option is wrong.
Asks whether a System One model can do the job of an LLM judge in an agent guardrail, with the hypotheses and decision rules fixed before any run.
Six tracks of experiments against the System One endpoint, from size limits and needle-in-a-haystack retrieval to prompt injection, option bias and a fifteen-job bake-off.
A head-to-head benchmark of open Laya against hosted Jev on byte-identical inputs in the same run, with question hashes verified identical before comparison.
Puts the Traditional Chinese TMMLU+ exam to Jev, one four-way question per item, and scores it the way the official leaderboard does.
Pairs a curated map of the Jev ecosystem with independent, reproducible benchmarks of the three primitives, confidence gating, fan-out latency and agent control.
Reduces 845 preflop spots to a single fold, call or raise decision, asks Jev five times each, and compares the result with four deterministic reference styles.
Turns Jev's probabilities into a routing threshold with a provable bound on how many queries get silently misrouted.
Benchmarks Jev on Brazil's ENEM 2025 exam against open-LLM baselines, contrasting a raw transcribed state with a structured one and measuring calibration.
A pip-installable benchmark for typed System One models that measures what the probabilities buy: calibration against a noise floor, wording sensitivity, selective prediction and cost.
Five hospital tasks, a thousand synthetic patients, and Jev answering every one, with Claude writing the scenarios and reviewing the results.
Pre-flights a System One question before you put its number behind an if, reporting a decision flip rate rather than a confidence score.
Tests Jev as a search reranker, and asks whether its confidence can tell you which queries are worth spending more compute on.
Tests Jev pairwise rankings of FOMC statements against rate decisions, FedLock scores, and a Haiku comparison.
Measures how Jev accuracy and calibration change when the same tasks are presented in English and Spanish.
Reports a preregistered adversarial evaluation spanning calibration, batching, answerability, option counts, and out-of-domain logic.
Run Jev-compatible typed judgments locally in a browser or Node server using small natural-language inference models.
Compare Jev with VIGILIA’s own labels on a sealed set of risky and harmless agent actions.
Petru placed a local classifier behind a Claude pre-tool-use hook as a Jev-free comparison.
Vincent put Jev inside a browser agent and reported faster steps and fewer model calls, with lower accuracy.
An early optimizer compares Jev, classic algorithms and an LLM agent on data-center load balancing and power management.
Nick Khami’s experimental endpoint applies inference engineering to make an open model behave like Jev.
A fraud pipeline lets Jev classify emails quickly, then sends uncertain cases to Kimi K3.
Florian Darroman compares Jev and Fable 5.1 on a post-scheduler build task and reports Jev finishing work faster.
WOLF backtests a Jev trading setup and reports better results after removing time data and most technical inputs.
Dorian Smiley tests whether Jev can predict a program’s next state from its current partial state.
Frank Chen built a small demonstration to probe how Jev behaves under prompt injection.
Ed Plese built a local, Jev-inspired model that returns probabilistic judgments about images and video frames.
echild tests Jev as a filter for selecting tools before an agent acts.
Alex Rivas pits Jev and Laya against each other over a chessboard, finding both still struggle with the game.
A small smoke test compares Jev and Laya on accuracy, calibration, latency, and per-decision cost.
Ram Vinjamuri compares Jev with fast Llama and DeepSeek models on the same labelled inputs.
Jason Alco compares Jev with Sol on a text-categorization step inside a platform he is building.
A local benchmark contrasts Laya’s on-device throughput with Jev’s API round trip.
WquGuru runs Jev and three open alternatives over 20 Chinese support tickets on a MacBook.
A Tetris match compares local Laya on a 16 GB MacBook Air with cloud-hosted Jev.
Laya makes typed decisions on a 12-year-old dual-core laptop without a GPU or cloud model.
M37 Labs demonstrates an encoder-only Laya-MLX architecture as a per-turn classifier for typed decisions.
A Mario-style browser test records Jev missing the first enemy and measures the full browser round trip separately from inference.
OpenMed tests Jev and Laya on prelabelled questions over four fictional clinical notes, explicitly framing it as a diagnostic rather than clinical accuracy.
Utkarsh Maheshwari compared the open-source Laya decision model with a Qwen reranker in his movie recommendation engine.
Hugging Models presented an open DeBERTa V3 Large classifier designed to return calibrated, typed decisions.
OpenRouter tested Jev as a judge with Ori Eval and reported it was more than five times faster than the next model.
scoreboar v8 predicts which of two X posts outperformed its account-size baseline, entirely in the browser.
An evolving evaluation suite tests System One models such as Jev through inexpensive, rapid iterations.
tab-jev combines a Jev-like model with a tabular foundation model for in-context learning over mixed text and tables.
Tev1 is a tiny Jev-like classifier demonstrated running locally through Ollama on a Mac.
TypeLLM adds reasoning before type-safe output as an open-source alternative to Jev-style direct decisions.
Jev, briefly
A chat model generates text token by token and hopes you parse it. Jev never generates a word. Your code sends the state of the world plus typed questions; Jev returns every answer in one parallel pass — typed values with calibrated probabilities and a confidence score your code can trust. About 70–500 ms end to end, $0.042 per million input tokens, output free. Trained with what TypeSafe calls Reinforcement Learning for Calibrated Decisions.
You define the options. Jev returns the chosen option, a probability for every option, and a confidence. The workhorse of routing, agents, and games — the answer space is always legal, so there is nothing to parse and nothing to hallucinate.
You describe ordered levels — say trivial / normal / critical. Jev returns the level, the probability of each, and a confidence. Scores turn fuzzy judgment (“how severe is this log line?”) into a number-free decision your code can branch on.
You assert a statement; Jev returns the probability that it is true — a noul. Moderation, verification, guardrails, “does this diff actually fix the bug?”: one question, one calibrated probability.
Official material lives at typesafe.ai and docs.typesafe.ai. JevMade is an independent community registry — not affiliated with or endorsed by TypeSafe AI.
Recurring lessons
Games, drones, trading bots, and browser agents all converge on the same loop: serialize the state, ask one decisive question, act, repeat. Jev’s latency makes the loop feel instant — the model lives inside the control loop, not outside it.
The answer says what; the confidence says whether to act. The most reliable entries threshold on confidence to route edge cases to a slower model or a human — automation with an honest escape hatch.
Makers rarely ask Jev to pick from everything. Local tactics prune 225 gomoku moves to ~40; DOM filters turn a page into an element table; code narrows options, Jev judges within them.
A recurring split: Jev makes every decision cheaply and instantly, and a small LLM is only invoked when a human-facing string must actually be written. Decision and generation are separate budgets.
Send the link — repo, live demo, post, or video — plus a line on what it does and which primitives it uses. Every entry is verified against its primary source before it ships.
Submit your experiment → hello@JevMade.com