JevMade hello@JevMade.com

Explore the Jev ecosystem.

Learn, build, and play with Jev — from deep technical guides to creative experiments.

Make & explore

Experiments.

See what builders made with Jev. Games, tools, repositories, articles, and more — traced to their sources.

Featured experiment Jev experiments by Nader Dabit
Recorded frame from Jev experiments
Explore 1,735 entries
The directory03 / 03 · Experiments

Find something worth exploring.

Every category, from games and tools to repositories and writeups, traced to a primary source.

Same ideas.
Different paths.

This is the full registry, not just the featured picks. Figures like speed, cost, and stars are maker-reported or captured snapshots, not JevMade measurements.

268 experiments · showing 1–60

World Monitor benchmarks Jev's headline severity and topic labels against cached labels and a saved judged set.

Source screenshot of worldmonitor
SOURCE SCREENSHOTEXPAND ↗
from the source async function callJev(titles) { const t0 = performance.now(); for (let attempt = 0; ; attempt++) { const r = await fetch(JEV_ENDPOINT, { method: 'POST',
choice ★ 87.4k Benchmarks & research

Run Convai Innovations' open-weight typed-decision model locally with a Jev-shaped request format; Laya is an independent model, not TypeSafe's Jev.

Source screenshot of Laya
SOURCE SCREENSHOTEXPAND ↗
choicescorenoul ★ 23.4k Benchmarks & research

kev

Train and self-host Qwen-based decision models that answer Jev-style typed questions in one pass; kev is an independent implementation, not Jev.

Source screenshot of kev
SOURCE SCREENSHOTEXPAND ↗
choicescorenoul ★ 6.8k Benchmarks & research

NanoJev trains a 0.6B parallel decision model and compares it with Jev across maze, Snake, and ViZDoom tasks.

Source screenshot of NanoJev
SOURCE SCREENSHOTEXPAND ↗
choicescorenoul ★ 2.2k Benchmarks & research

Nimble compares Jev's typed probabilities with locally trained schema adapters on curated decision datasets.

Source screenshot of Nimble
SOURCE SCREENSHOTEXPAND ↗
from the source payload = request_payload(row, model)response = request_with_retry(transport, payload, attempts=attempts, sleep=sleep)validate_teacher(row, response, model)return jev_row(row, response, time.perf_counter() - started)
choicescorenoul ★ 1.8k Benchmarks & research

SemIf runs open models locally as probabilistic semantic conditionals on an RTX 3090.

Source screenshot of SemIf
SOURCE SCREENSHOTEXPAND ↗
choicescore ★ 1.6k Benchmarks & research

Play Snake or request typed decisions from Core ML ports of Laya on Apple hardware; the models are open-weight alternatives, not Jev.

Source screenshot of Laya-CoreML
SOURCE SCREENSHOTEXPAND ↗
choicescorenoul ★ 1.4k Benchmarks & research

Jevlike trains a compact model to choose among a list of text options that can change from request to request.

Source screenshot of jevlike
SOURCE SCREENSHOTEXPAND ↗
from the source uv venvsource .venv/bin/activateuv pip install -e '.[dev]'jevlike-data synthetic --output data/syntheticjevlike-train data/synthetic/train.jsonl \
choicescore ★ 879 Benchmarks & research

von

Run a local, non-autoregressive System One model behind Python and JavaScript clients compatible with Jev’s request shape.

Source screenshot of von
SOURCE SCREENSHOTEXPAND ↗
choicescorenoul ★ 644 Benchmarks & research

Adds typed probability readouts and calibration heads to ordinary language models, with Jev-style evaluations and saved results.

Source screenshot of AnyJev
SOURCE SCREENSHOTEXPAND ↗
choicescorenoul ★ 518 Benchmarks & research

Mini Taiwan Pulse tests whether Jev can pick the map layers relevant to a visitor's question.

Source screenshot of Mini Taiwan Pulse Jev layer screening
SOURCE SCREENSHOTEXPAND ↗
from the source const probabilities = parseNoulAnswers(body, batch.questionIds);for (const candidate of batch.candidates) decisions.set(candidate.key, knownDecision(probabilities.get(questionIdForLayer(candid...
noul ★ 511 Benchmarks & research

Compare Jev's market probabilities and trading choices in a Polymarket benchmark that never places real-money orders.

Source screenshot of polymarket-paper-trader
SOURCE SCREENSHOTEXPAND ↗
from the source probability = noul_probability(answers, "probability")action = choice_value(answers, "action")confidence = choice_value(answers, "confidence")
choicenoul ★ 399 Benchmarks & research

A 151M non-autoregressive decision model pairs typed outputs with benchmark receipts, calibration analysis, and browser-oriented artifacts.

Source screenshot of openJev-verdict-2.0
SOURCE SCREENSHOTEXPAND ↗
choicescorenoul ★ 284 Benchmarks & research

Vector Graph RAG can rerank graph relations with Jev and compare the resulting retrieval recall.

Source screenshot of Vector Graph RAG Jev reranker
SOURCE SCREENSHOTEXPAND ↗
from the source "reranker_provider": self.rag.settings.reranker_provider,"jev_model": self.rag.settings.jev_model if self.rag.settings.reranker_provider == "jev"..."jev_threshold": self.rag.settings.jev_threshold if self.rag.settings.reranker_provider...
noul ★ 250 Benchmarks & research

An Apple Silicon experiment scores visual candidates directly against shared context using a local model.

Source screenshot of jev-visual
SOURCE SCREENSHOTEXPAND ↗
from the source jev-visual examples/photo-request.json --model-path .models/Qwen3.5-0.8B-4bit
choicescore ★ 107 Benchmarks & research

JevK5 trains and serves an open-weight decision model that returns closed-choice probabilities in one forward pass; it is not TypeSafe's Jev.

Source screenshot of JevK5
SOURCE SCREENSHOTEXPAND ↗
choicenoul ★ 102 Benchmarks & research

A benchmark of how Jev and a language model choose among 100 tools for a personal assistant.

Source screenshot of jev-eval-agent
SOURCE SCREENSHOTEXPAND ↗
from the source AGENT_MODEL=mock npm run eval # llm-directAGENT_MODE=jev-classifier AGENT_MODEL=mock JEV_STUB=1 npm run eval # jev-classifier
choicenoul ★ 83 Benchmarks & research

Reimplements Jev-style bounded decisions locally, exposing a compatible contract over prompt-only and open-weight backends.

Source screenshot of Intent-Router
SOURCE SCREENSHOTEXPAND ↗
choicescorenoul ★ 74 Benchmarks & research

Reflex is an open Qwen3.5-based model that maps state and typed questions to calibrated probabilities.

Source screenshot of reflex
SOURCE SCREENSHOTEXPAND ↗
choice Which team should handle this?
choicescorenoul ★ 62 Benchmarks & research

Replay Jev and chat-model judgments side by side on comment datasets imported from CSV or Excel.

Source screenshot of Jev Arena
SOURCE SCREENSHOTEXPAND ↗
choicescorenoul ★ 57 Benchmarks & research

A Python library for checking an agent's work with several Jev questions in one call.

Source screenshot of jevals
SOURCE SCREENSHOTEXPAND ↗
from the source async def evaluate(self, state: Any, questions: dict[str, Question]) -> BackendRespo... body: dict[str, Any] = {"state": state, "questions": questions_to_typesafe(quest... if self.model: body["model"] = self.model headers = {"Authorization": f"Bearer {self.api_key}", "Content-Type": "applicati...
choicenoul ★ 45 Benchmarks & research

Goodwatch's search arena measures Jev's latency, token use and request size on movie-search questions.

Source screenshot of Goodwatch Jev search arena
SOURCE SCREENSHOTEXPAND ↗
from the source const [attribute, fingerprint] = await Promise.all([ call(attributeRequest(q)), call(fingerprintRequest(q)),]);
noulscore ★ 38 Benchmarks & research

Run Nano, Small or Large open-weight models behind Jev-style Choice, Score and Noul requests; these are independent models, not TypeSafe's Jev.

Source screenshot of KaLM-Jev
SOURCE SCREENSHOTEXPAND ↗
choicescorenoul ★ 33 Benchmarks & research

A gateway that makes a Cerebras-backed model answer in Jev's format, for comparison with the real thing.

Recorded frame of typesafe-ai-benchmark
ACTUAL RECORDING12 SEC ↗
from the source packages/demos/ Side-by-side benchmark UI, workloads and native Jev adapterpackages/api/ Cerebras structured-output adapter, validation and supporting APIscripts/ Offline benchmark summarizers and media toolingdocs/benchmarks/ Raw exports, source hashes, metrics and quality comparisons
choicenoul ★ 32 Benchmarks & research

OpenJev is a Jev-compatible decision server built on DiffusionGemma.

Source screenshot of openjev
SOURCE SCREENSHOTEXPAND ↗
noul Does the customer need a reply within the hour?
choicescorenoul ★ 32 Benchmarks & research

An open reimplementation trains a lightweight decision head for Choice, Score and Noul questions.

Source screenshot of open-jev (daseinlabs)
SOURCE SCREENSHOTEXPAND ↗
score The capital of France is
choicescorenoul ★ 31 Benchmarks & research

Checks Jev questions against labelled data and reports whether each can gate decisions, rank examples, or carries no useful signal.

Source screenshot of jev-calibrate
SOURCE SCREENSHOTEXPAND ↗
choicescorenoul ★ 31 Benchmarks & research

Qwen Choice reads option logits from a local vision-language model to classify one image without generating an answer.

Source screenshot of Qwen Choice
SOURCE SCREENSHOTEXPAND ↗
choice ★ 29 Benchmarks & research

Maps where Jev succeeds or fails using reproducible API receipts, external studies, and bilingual explanations rather than a leaderboard.

Source screenshot of Jev Capability Atlas
SOURCE SCREENSHOTEXPAND ↗
choicescorenoul ★ 26 Benchmarks & research

Decider fine-tunes Qwen3.5-2B to make typed choices with calibrated probabilities in one pass.

Recorded frame of decider
ACTUAL RECORDING12 SEC ↗
choicescorenoul ★ 25 Benchmarks & research

JevMLX adds parallel constrained decisions and schema-valid JSON to MLX models on Apple Silicon.

Source screenshot of jevmlx
SOURCE SCREENSHOTEXPAND ↗
from the source pip install git+https://github.com/bnsd55/jevmlx # libraryuv tool install git+https://github.com/bnsd55/jevmlx # CLI onlygit clone https://github.com/bnsd55/jevmlx && cd jevmlx && ./setup.sh # dev
choicescore ★ 24 Benchmarks & research

A compact PyTorch teaching implementation explores a Jev-inspired encoder, shared state cache, and typed readout heads.

Source screenshot of Open Jev
SOURCE SCREENSHOTEXPAND ↗
choicescorenoul ★ 24 Benchmarks & research

Mini-Jev reads option-letter logits from a frozen Qwen3-4B instead of asking the model to generate JSON.

Source screenshot of mini-jev
SOURCE SCREENSHOTEXPAND ↗
from the source git clone https://github.com/r-ms/mini-jev.git && cd mini-jevuv sync # torch 2.5.1, transformers 4.57.6, xgrammar 0...MINIJEV_DEVICE=mps uv run python demo/server.py # or MINIJEV_DEVICE=cuda# first start downloads Qwen/Qwen3-4B-Instruct-2507 (revision pinned) and warms up, then...open http://127.0.0.1:8765/
choicescorenoul ★ 22 Benchmarks & research

Compare Rust plus Jev with Python plus DeepSeek as both sides flag simulated Pix scams and trace the same criminal network.

Source screenshot of Pix Golpe
SOURCE SCREENSHOTEXPAND ↗
choicescorenoul ★ 22 Benchmarks & research

Run Laya's typed-decision model on Apple Silicon with selectable memory modes and a Pong latency demo; this is not TypeSafe's Jev.

Source screenshot of Laya MPS
SOURCE SCREENSHOTEXPAND ↗
choicescorenoul ★ 21 Benchmarks & research

Evaluates whether classifier confidence is calibrated and derives human-review thresholds from the cost of mistakes.

Source screenshot of jeval
SOURCE SCREENSHOTEXPAND ↗
— ★ 18 Benchmarks & research

This study benchmarks parallel typed decisions on unmodified 1.5B–8B models running on Apple Silicon.

Source screenshot of jev-on-a-laptop
SOURCE SCREENSHOTEXPAND ↗
from the source git clone https://github.com/rorshopping/jev-on-a-laptopcd jev-on-a-laptop./setup.sh # creates .venv, installs mlx-lm, clones the upstream engine./run_benchmark.sh # default: Qwen2.5-1.5B-Instruct-4bit
choice ★ 14 Benchmarks & research

Studies calibration-aware reinforcement learning for adaptive decision systems through training and benchmark harnesses.

Source screenshot of JevAny
SOURCE SCREENSHOTEXPAND ↗
choicescorenoul ★ 13 Benchmarks & research

Train a Qwen-based closed-choice model by applying Brier loss directly to candidate-token probabilities; it explores Jev's decision-first idea but is not Jev.

Source screenshot of JevTuner
SOURCE SCREENSHOTEXPAND ↗
choice ★ 13 Benchmarks & research

A benchmark labels 1,000 app reviews with Jev and Gemini 3.8 Flash side by side.

Source screenshot of jev-column-race
SOURCE SCREENSHOTEXPAND ↗
from the source git clone https://github.com/goodrahstar/jev-column-race.gitcd jev-column-racecp .env.example .env# Add TYPESAFE_API_KEY, and LLM_API_KEY for the Gemini lane.node server.mjs
choicescorenoul ★ 12 Benchmarks & research

An independent text-scoring model that tries to improve on the jevlike starter project.

Source screenshot of jevbetter
SOURCE SCREENSHOTEXPAND ↗
from the source $ jevbetter-benchmark --reference /path/to/jevlike --epochs 8
score ★ 11 Benchmarks & research

Compare decision models at Texas Hold'em through cash and sit-and-go leaderboards, hand replays and tables for custom agents.

Source screenshot of JevPokerBench
SOURCE SCREENSHOTEXPAND ↗
choice ★ 11 Benchmarks & research

Kev

Self-host a Jev-compatible decision engine over Ollama or OpenAI-compatible models, with typed outputs and calibrated probabilities; it does not run TypeSafe's Jev.

Source screenshot of Kev
SOURCE SCREENSHOTEXPAND ↗
choicescorenoul ★ 11 Benchmarks & research

Serve Laya typed decisions through a Rust Candle runtime for local CPU or GPU inference; this is a Jev-compatible alternative, not Jev.

Source screenshot of laya-candle
SOURCE SCREENSHOTEXPAND ↗
choicescorenoul ★ 9 Benchmarks & research

Add natural-language constraints to JEPA planning by asking Jev whether probe-described imagined states violate each rule, then folding probabilities into planning cost.

Source screenshot of LeJudge
SOURCE SCREENSHOTEXPAND ↗
noul ★ 9 Benchmarks & research

TypeAR studies type-safe constrained decoding for autoregressive language models.

Source screenshot of TypeAR
SOURCE SCREENSHOTEXPAND ↗
choicescore ★ 8 Benchmarks & research

A reproducible evaluation suite measures calibration, selective risk and latency in probabilistic decision models.

Source screenshot of jev-benchmarks
SOURCE SCREENSHOTEXPAND ↗
choicescore ★ 8 Benchmarks & research

OpenVons answers finite-choice questions with probabilities across text, images and Japanese voice commands.

Source screenshot of openvons
SOURCE SCREENSHOTEXPAND ↗
choicescorenoul ★ 7 Benchmarks & research

system-one performs batched, single-token choice inference with open language models through a TypeSafe-compatible interface.

Recorded frame of system-one
ACTUAL RECORDING12 SEC ↗
from the source from system_one import SystemOnefrom typesafe_sdk import Choiceengine = SystemOne.from_pretrained("Qwen/Qwen2.5-0.5B-Instruct")result = engine.system_one( state="My shoes arrived in the wrong size. Can I exchange them?",
choice ★ 7 Benchmarks & research

Run the open-weight Laya decision model locally on Apple Silicon through an Apple-focused runtime; it is Jev-compatible but does not use Jev.

Source screenshot of laya-apple
SOURCE SCREENSHOTEXPAND ↗
choicescorenoul ★ 7 Benchmarks & research

A bilingual English-Chinese guide explains notable Jev applications, how they work and their trade-offs.

Source screenshot of Jev_apps
SOURCE SCREENSHOTEXPAND ↗
— ★ 6 Benchmarks & research

Jev, briefly

A decision model, not a chatbot.

A chat model generates text token by token and hopes you parse it. Jev never generates a word. Your code sends the state of the world plus typed questions; Jev returns every answer in one parallel pass — typed values with calibrated probabilities and a confidence score your code can trust. About 70–500 ms end to end, $0.042 per million input tokens, output free. Trained with what TypeSafe calls Reinforcement Learning for Calibrated Decisions.

choice Pick one of these.

You define the options. Jev returns the chosen option, a probability for every option, and a confidence. The workhorse of routing, agents, and games — the answer space is always legal, so there is nothing to parse and nothing to hallucinate.

score Rate this on a rubric.

You describe ordered levels — say trivial / normal / critical. Jev returns the level, the probability of each, and a confidence. Scores turn fuzzy judgment (“how severe is this log line?”) into a number-free decision your code can branch on.

noul Is this statement true?

You assert a statement; Jev returns the probability that it is true — a noul. Moderation, verification, guardrails, “does this diff actually fix the bug?”: one question, one calibrated probability.

Official material lives at typesafe.ai and docs.typesafe.ai. JevMade is an independent community registry — not affiliated with or endorsed by TypeSafe AI.

Recurring lessons

Patterns that keep showing up.

One decision per tick

Games, drones, trading bots, and browser agents all converge on the same loop: serialize the state, ask one decisive question, act, repeat. Jev’s latency makes the loop feel instant — the model lives inside the control loop, not outside it.

Confidence as a gate

The answer says what; the confidence says whether to act. The most reliable entries threshold on confidence to route edge cases to a slower model or a human — automation with an honest escape hatch.

Shrink the choice space in code

Makers rarely ask Jev to pick from everything. Local tactics prune 225 gomoku moves to ~40; DOM filters turn a page into an element table; code narrows options, Jev judges within them.

Jev judges, LLMs talk

A recurring split: Jev makes every decision cheaply and instantly, and a small LLM is only invoked when a human-facing string must actually be written. Decision and generation are separate budgets.

Made something with Jev?

Send the link — repo, live demo, post, or video — plus a line on what it does and which primitives it uses. Every entry is verified against its primary source before it ships.

Submit your experiment → hello@JevMade.com