JevMade hello@JevMade.com

Explore the Jev ecosystem.

Learn, build, and play with Jev — from deep technical guides to creative experiments.

Make & explore

Experiments.

See what builders made with Jev. Games, tools, repositories, articles, and more — traced to their sources.

Featured experiment Jev experiments by Nader Dabit
Recorded frame from Jev experiments
Explore 1,737 entries
The directory03 / 03 · Experiments

Find something worth exploring.

Every category, from games and tools to repositories and writeups, traced to a primary source.

Same ideas.
Different paths.

This is the full registry, not just the featured picks. Figures like speed, cost, and stars are maker-reported or captured snapshots, not JevMade measurements.

269 experiments · showing 241–269

TypeLLM adds reasoning before type-safe output as an open-source alternative to Jev-style direct decisions.

Source screenshot of TypeLLM
SOURCE SCREENSHOTEXPAND ↗
— Benchmarks & research

Analyze thousands of personal X posts across eight questions to compare which topics, hooks, and teaching styles correlate with engagement.

Recorded frame of X Post Analysis
ACTUAL RECORDING12 SEC ↗
— Benchmarks & research

wmoto ran a local decision-model prototype and published a ten-second demonstration while noting that its speed still needed work.

Source screenshot of Local Jev
SOURCE SCREENSHOTEXPAND ↗
— Benchmarks & research

Jacky Ko and collaborators introduced an Apache-2.0 contrastive language model that connects states to actions as an open System One model.

Recorded frame of CLM-8B
ACTUAL RECORDING12 SEC ↗
— Benchmarks & research

A Phoenix Evals example pits Jev against a small grounded-versus-hallucinated answer benchmark.

Source screenshot of phoenix
SOURCE SCREENSHOTEXPAND ↗
from the source const jev = typeSafeAi.evaluationModel("jev-latest");const jevEvaluator = createHallucinationEvaluator({ model: jev });const [jevResult, jevMs] = await timed(() => jevEvaluator.evaluate(example));
choice Benchmarks & research

haiku.rag provides a reusable System One judge for benchmark verdicts such as whether two answers are equivalent.

Source screenshot of haiku.rag
SOURCE SCREENSHOTEXPAND ↗
from the source response = await self.client.system_one( self._state(ctx), {name: self._question()}, model=self.model)p = response.nouls[name].noulverdict = EvaluationReason(value=p >= threshold, reason=f"system_one p={p:.3f}")
noul Benchmarks & research

Run Laya’s typed decisions natively on Apple Silicon, then watch its terminal Snake demo choose moves with safety corrections.

Recorded frame of Laya-MLX
ACTUAL RECORDING12 SEC ↗
choicescorenoul Benchmarks & research

SREGym gave its incident-solving agent a hesitant reviewer: before any diagnosis gets submitted, Jev has to agree the evidence supports it, and when a test is proposed, Jev ranks which one is worth running first. Pass rate moved from 20 to 24 out of 50.

Source screenshot of SREGym + Jev
SOURCE SCREENSHOTEXPAND ↗
choicescorenoul Benchmarks & research

CartPole, MountainCar, Acrobot, and FrozenLake trained with reinforcement rewards that come from Jev's calibrated scores instead of a hand-written reward function.

Source screenshot of Jev RL
SOURCE SCREENSHOTEXPAND ↗
score Benchmarks & research

An open battleground that feeds Jev and Laya byte-identical rounds — same wording, same option order — and publishes the report card.

Source screenshot of Jev vs Laya
SOURCE SCREENSHOTEXPAND ↗
choice Benchmarks & research

The fan-out docs say batching bills the state once. As far as anyone could tell, nobody had checked, so this repo checks with real money.

Source screenshot of jev-fanout-bench
SOURCE SCREENSHOTEXPAND ↗
choicescorenoul Benchmarks & research

Databases sort by numbers and Jev hands back probabilities, so sorting search results by them is tempting. This measures how far the temptation carries.

Source screenshot of jev-orderby-bench
SOURCE SCREENSHOTEXPAND ↗
noulscore Benchmarks & research

Freeze a small Qwen model, train a decision head smaller than a million parameters, and it answers typed questions with zero generated tokens.

Source screenshot of minojev
SOURCE SCREENSHOTEXPAND ↗
choicescore Benchmarks & research

A 0.8B vision-language model playing ten browser games straight from pixels: one frame in, one move out, 43 ms on an H200.

Source screenshot of PlayJev
SOURCE SCREENSHOTEXPAND ↗
choice Benchmarks & research

Every voice agent transcribes before it decides anything. Prosodia asks the audio directly, so a misheard word can no longer poison the judgment downstream.

Source screenshot of Prosodia
SOURCE SCREENSHOTEXPAND ↗
choicescorenoul Benchmarks & research

Jev, briefly

A decision model, not a chatbot.

A chat model generates text token by token and hopes you parse it. Jev never generates a word. Your code sends the state of the world plus typed questions; Jev returns every answer in one parallel pass — typed values with calibrated probabilities and a confidence score your code can trust. About 70–500 ms end to end, $0.042 per million input tokens, output free. Trained with what TypeSafe calls Reinforcement Learning for Calibrated Decisions.

choice Pick one of these.

You define the options. Jev returns the chosen option, a probability for every option, and a confidence. The workhorse of routing, agents, and games — the answer space is always legal, so there is nothing to parse and nothing to hallucinate.

score Rate this on a rubric.

You describe ordered levels — say trivial / normal / critical. Jev returns the level, the probability of each, and a confidence. Scores turn fuzzy judgment (“how severe is this log line?”) into a number-free decision your code can branch on.

noul Is this statement true?

You assert a statement; Jev returns the probability that it is true — a noul. Moderation, verification, guardrails, “does this diff actually fix the bug?”: one question, one calibrated probability.

Official material lives at typesafe.ai and docs.typesafe.ai. JevMade is an independent community registry — not affiliated with or endorsed by TypeSafe AI.

Recurring lessons

Patterns that keep showing up.

One decision per tick

Games, drones, trading bots, and browser agents all converge on the same loop: serialize the state, ask one decisive question, act, repeat. Jev’s latency makes the loop feel instant — the model lives inside the control loop, not outside it.

Confidence as a gate

The answer says what; the confidence says whether to act. The most reliable entries threshold on confidence to route edge cases to a slower model or a human — automation with an honest escape hatch.

Shrink the choice space in code

Makers rarely ask Jev to pick from everything. Local tactics prune 225 gomoku moves to ~40; DOM filters turn a page into an element table; code narrows options, Jev judges within them.

Jev judges, LLMs talk

A recurring split: Jev makes every decision cheaply and instantly, and a small LLM is only invoked when a human-facing string must actually be written. Decision and generation are separate budgets.

Made something with Jev?

Send the link — repo, live demo, post, or video — plus a line on what it does and which primitives it uses. Every entry is verified against its primary source before it ships.

Submit your experiment → hello@JevMade.com