JevMade hello@JevMade.com

Explore the Jev ecosystem.

Learn, build, and play with Jev — from deep technical guides to creative experiments.

Make & explore

Experiments.

See what builders made with Jev. Games, tools, repositories, articles, and more — traced to their sources.

Featured experiment Jev experiments by Nader Dabit
Recorded frame from Jev experiments
Explore 1,735 entries
The directory03 / 03 · Experiments

Find something worth exploring.

Every category, from games and tools to repositories and writeups, traced to a primary source.

Same ideas.
Different paths.

This is the full registry, not just the featured picks. Figures like speed, cost, and stars are maker-reported or captured snapshots, not JevMade measurements.

268 experiments · showing 181–240

Can a cheap classifier label every agent failure in a training run? Applied Compute clusters a sample of traces, then has Jev do the labeling.

Source screenshot of Billion-Token Scale Trace Analysis: Jev vs LLMs
SOURCE SCREENSHOTEXPAND ↗
from the source Each model outputs a floating-point score from 0 to 1 for each of 14 failure modes.A score at or above the decision threshold marks that failure mode as present.
score Benchmarks & research

3,608 one-subject-swapped statements put to Jev as forced true/false picks, in three wordings and eleven languages, with every raw answer published.

Source screenshot of Jevsus
SOURCE SCREENSHOTEXPAND ↗
from the source self.url, self.model = JEV_URL, "jev-latest"
choicenoul Benchmarks & research

Auto-marks 2,054 MR-GSM8K maths scripts against a teacher rubric with one Jev call each, then checks the marks against human annotators.

Source screenshot of JEValuate
SOURCE SCREENSHOTEXPAND ↗
from the source const result = await client.systemOne(
noulchoice Benchmarks & research

The open data behind an independent benchmark that puts Jev and six LLMs through the same typed questions on PubMedQA, Banking77 and HelpSteer2.

Source screenshot of Jevals
SOURCE SCREENSHOTEXPAND ↗
from the source run header: {"type":"header",...,"model":"typesafe-ai/jev",...,"probability_source":
choicescorenoul Benchmarks & research

The 2026 Korean CSAT Korean-language paper turned into a JSON state and typed choice questions, with Jev's answers, probabilities and error analysis kept alongside.

Source screenshot of 2026 수능 국어영역 (홀수형)
SOURCE SCREENSHOTEXPAND ↗
from the source {\"1\": { \"type\": \"choice\", \"instructions\": { \"question\": \
choice Benchmarks & research

An independent benchmark of Jev 1.13 against a cheap and a frontier LLM on SMS spam and Banking77, measuring accuracy, ECE, Brier, latency and cost.

Source screenshot of jev-bench
SOURCE SCREENSHOTEXPAND ↗
from the source resp = requests.post(\"https://openrouter.ai/api/alpha/decisions\",
choicenoul Benchmarks & research

A drop-in proxy that answers your app with Jev byte for byte while asking the same question of local doubles, then reports where they would have decided differently.

Source screenshot of stuntdouble
SOURCE SCREENSHOTEXPAND ↗
from the source \", \"headers\": { \"authorization\": \"Bearer ${TYPESAFE_API_KEY}\" }, \"model\": \"jev
choicescorenoul Benchmarks & research

Measures what the live Jev API actually returns across eight realistic use cases, reporting cost, latency and response shape rather than rate-card arithmetic.

Source screenshot of jev-measured
SOURCE SCREENSHOTEXPAND ↗
from the source ENDPOINT = BASE_URL +
choicescorenoul Benchmarks & research

A tree with a place for every closed question a person or program could ask, filled with real questions answered by Jev, plus a 3D explorer where each question is a star.

Source screenshot of askjev
SOURCE SCREENSHOTEXPAND ↗
from the source r = await self.http.post(GATEWAY_URL, content=canonical(req.body))
choicescorenoul Benchmarks & research

Tests Jev on scientific judgement questions, then follows the consequences: what those choices do to the results that depend on them.

Source screenshot of Scientific Decision Evaluation
SOURCE SCREENSHOTEXPAND ↗
from the source ENDPOINTS = {'openrouter': 'https://openrouter.ai/api/alpha/decisions',
choice Benchmarks & research

Reranks fixed MS MARCO candidate sets by asking Jev one relevance question per query and document pair, then scores the runs with trec_eval.

Source screenshot of JEV Reranking Comparisons
SOURCE SCREENSHOTEXPAND ↗
from the source payload = {"model": args.model, "state": {"query": topics[qid]
noul Benchmarks & research

A first hands-on measurement of Jev: seven small scripts probing its documented jaggedness, latency scaling, calibration, negation consistency and behaviour when every option is wrong.

Source screenshot of jev-first-look
SOURCE SCREENSHOTEXPAND ↗
from the source r = client.system_one(A, { (then noul: Noul(instructions=REFUND) and choice: Choice(inst
noulchoice Benchmarks & research

Asks whether a System One model can do the job of an LLM judge in an agent guardrail, with the hypotheses and decision rules fixed before any run.

Source screenshot of jev-guardbench
SOURCE SCREENSHOTEXPAND ↗
from the source resp = await self._client.system_one(
noul Benchmarks & research

Six tracks of experiments against the System One endpoint, from size limits and needle-in-a-haystack retrieval to prompt injection, option bias and a fifteen-job bake-off.

Source screenshot of TypeSafe AI (Jev 1.13) stress test
SOURCE SCREENSHOTEXPAND ↗
from the source resp = client()
choicescorenoul Benchmarks & research

A head-to-head benchmark of open Laya against hosted Jev on byte-identical inputs in the same run, with question hashes verified identical before comparison.

Source screenshot of sysone-bench
SOURCE SCREENSHOTEXPAND ↗
from the source DIRECT_URL = "https://api.typesafe.ai/v1/systemone"
choicescorenoul Benchmarks & research

Puts the Traditional Chinese TMMLU+ exam to Jev, one four-way question per item, and scores it the way the official leaderboard does.

Source screenshot of jev-tmmluplus-eval
SOURCE SCREENSHOTEXPAND ↗
from the source ENDPOINT = "/v1/systemone"
choice Benchmarks & research

Pairs a curated map of the Jev ecosystem with independent, reproducible benchmarks of the three primitives, confidence gating, fan-out latency and agent control.

Source screenshot of Awesome Jev Lab
SOURCE SCREENSHOTEXPAND ↗
choicenoul Benchmarks & research

Reduces 845 preflop spots to a single fold, call or raise decision, asks Jev five times each, and compares the result with four deterministic reference styles.

Source screenshot of Is Jev a decent preflop poker player?
SOURCE SCREENSHOTEXPAND ↗
from the source URL = "https://api.typesafe.ai/v1/systemone"
choice Benchmarks & research

Turns Jev's probabilities into a routing threshold with a provable bound on how many queries get silently misrouted.

Source screenshot of jev-certify
SOURCE SCREENSHOTEXPAND ↗
from the source DECISIONS_URL = "https://openrouter.ai/api/alpha/decisions"
choicenoul Benchmarks & research

Benchmarks Jev on Brazil's ENEM 2025 exam against open-LLM baselines, contrasting a raw transcribed state with a structured one and measuring calibration.

Source screenshot of Jev no ENEM
SOURCE SCREENSHOTEXPAND ↗
from the source response = self.client.system_one(state=state, questions={"gabarito": Choice(instruction
choice Benchmarks & research

A pip-installable benchmark for typed System One models that measures what the probabilities buy: calibration against a noise floor, wording sensitivity, selective prediction and cost.

Source screenshot of sys1bench
SOURCE SCREENSHOTEXPAND ↗
from the source DEFAULT_URL =
choicescorenoul Benchmarks & research

Five hospital tasks, a thousand synthetic patients, and Jev answering every one, with Claude writing the scenarios and reviewing the results.

Source screenshot of explore-typesafe-ai
SOURCE SCREENSHOTEXPAND ↗
from the source response = await client.system_one(state, questions)
choicescorenoul Benchmarks & research

Pre-flights a System One question before you put its number behind an if, reporting a decision flip rate rather than a confidence score.

Source screenshot of @vcjdeboer/jev-reliability
SOURCE SCREENSHOTEXPAND ↗
from the source resp = await fetch(`${baseUrl}/v1/systemone`, { method: "POST", headers: { Authorization
scorenoulchoice Benchmarks & research

Tests Jev as a search reranker, and asks whether its confidence can tell you which queries are worth spending more compute on.

Source screenshot of S1Rank
SOURCE SCREENSHOTEXPAND ↗
from the source r = await self.http.post( URL, json=payload, headers={"Authorization": f
noulchoicescore Benchmarks & research

Measures how Jev accuracy and calibration change when the same tasks are presented in English and Spanish.

Source screenshot of jev-acento
SOURCE SCREENSHOTEXPAND ↗
choicenoul ★ 0 Benchmarks & research

Reports a preregistered adversarial evaluation spanning calibration, batching, answerability, option counts, and out-of-domain logic.

Source screenshot of Evaluating jev
SOURCE SCREENSHOTEXPAND ↗
choicescorenoul ★ 0 Benchmarks & research

Vej

Run Jev-compatible typed judgments locally in a browser or Node server using small natural-language inference models.

Source screenshot of Vej
SOURCE SCREENSHOTEXPAND ↗
choicescorenoul ★ 0 Benchmarks & research

Tev1 is a tiny Jev-like classifier demonstrated running locally through Ollama on a Mac.

Source screenshot of Tev1 0.8B
SOURCE SCREENSHOTEXPAND ↗
— Benchmarks & research

TypeLLM adds reasoning before type-safe output as an open-source alternative to Jev-style direct decisions.

Source screenshot of TypeLLM
SOURCE SCREENSHOTEXPAND ↗
— Benchmarks & research

Jev, briefly

A decision model, not a chatbot.

A chat model generates text token by token and hopes you parse it. Jev never generates a word. Your code sends the state of the world plus typed questions; Jev returns every answer in one parallel pass — typed values with calibrated probabilities and a confidence score your code can trust. About 70–500 ms end to end, $0.042 per million input tokens, output free. Trained with what TypeSafe calls Reinforcement Learning for Calibrated Decisions.

choice Pick one of these.

You define the options. Jev returns the chosen option, a probability for every option, and a confidence. The workhorse of routing, agents, and games — the answer space is always legal, so there is nothing to parse and nothing to hallucinate.

score Rate this on a rubric.

You describe ordered levels — say trivial / normal / critical. Jev returns the level, the probability of each, and a confidence. Scores turn fuzzy judgment (“how severe is this log line?”) into a number-free decision your code can branch on.

noul Is this statement true?

You assert a statement; Jev returns the probability that it is true — a noul. Moderation, verification, guardrails, “does this diff actually fix the bug?”: one question, one calibrated probability.

Official material lives at typesafe.ai and docs.typesafe.ai. JevMade is an independent community registry — not affiliated with or endorsed by TypeSafe AI.

Recurring lessons

Patterns that keep showing up.

One decision per tick

Games, drones, trading bots, and browser agents all converge on the same loop: serialize the state, ask one decisive question, act, repeat. Jev’s latency makes the loop feel instant — the model lives inside the control loop, not outside it.

Confidence as a gate

The answer says what; the confidence says whether to act. The most reliable entries threshold on confidence to route edge cases to a slower model or a human — automation with an honest escape hatch.

Shrink the choice space in code

Makers rarely ask Jev to pick from everything. Local tactics prune 225 gomoku moves to ~40; DOM filters turn a page into an element table; code narrows options, Jev judges within them.

Jev judges, LLMs talk

A recurring split: Jev makes every decision cheaply and instantly, and a small LLM is only invoked when a human-facing string must actually be written. Decision and generation are separate budgets.

Made something with Jev?

Send the link — repo, live demo, post, or video — plus a line on what it does and which primitives it uses. Every entry is verified against its primary source before it ships.

Submit your experiment → hello@JevMade.com