JevMade hello@JevMade.com

Explore the Jev ecosystem.

Learn, build, and play with Jev — from deep technical guides to creative experiments.

Make & explore

Experiments.

See what builders made with Jev. Games, tools, repositories, articles, and more — traced to their sources.

Featured experiment Jev experiments by Nader Dabit
Recorded frame from Jev experiments
Explore 1,734 entries
The directory03 / 03 · Experiments

Find something worth exploring.

Every category, from games and tools to repositories and writeups, traced to a primary source.

Same ideas.
Different paths.

This is the full registry, not just the featured picks. Figures like speed, cost, and stars are maker-reported or captured snapshots, not JevMade measurements.

268 experiments · showing 61–120

Maps 1,000 recent AI papers into 24 topics with Jev, then compares a seeded sample with an LLM judge.

Source screenshot of jev-papers
SOURCE SCREENSHOTEXPAND ↗
choice ★ 6 Benchmarks & research

jevfire batches many Jev-style decisions over one shared context on CUDA language models.

Source screenshot of jevfire
SOURCE SCREENSHOTEXPAND ↗
from the source git clone https://github.com/kikoncuo/jevfire.gitcd jevfirepython -m venv .venvsource .venv/bin/activatepython -m pip install -e '.[dev]'
choicescore ★ 5 Benchmarks & research

A word-level language model that drafts with n-grams and asks Jev to check chunks of the result.

Source screenshot of jev-lm
SOURCE SCREENSHOTEXPAND ↗
from the source jev-lm probe "the little dog was" --top 5jev-lm verify "the capital of france is" "paris" "london"jev-lm gen "the little dog" --words 18 --tracejev-lm gen "she went to the door" --draft --tracejev-lm eval --file data/heldout_en.txt --positions 60
choicescorenoul ★ 5 Benchmarks & research

typesafe-local asks a local MLX model typed questions and returns calibrated probabilities without generating or parsing prose.

Source screenshot of typesafe-local
SOURCE SCREENSHOTEXPAND ↗
from the source { "model": "mlx-community/Qwen3-1.7B-bf16", "answers": { "refund_requested": {"type": "noul", "noul": 1.0, "raw_mass": 1.0}, "department": {"type": "choice", "choice": "billing",
choicescorenoul ★ 5 Benchmarks & research

Verdict is a 151M-parameter ModernBERT decision engine with calibrated uncertainty and no autoregressive generation.

Source screenshot of Verdict-open-jev
SOURCE SCREENSHOTEXPAND ↗
choicescorenoul ★ 5 Benchmarks & research

Tests whether filtering coding-agent tool results through Jev improves task success, context size, or cost against an unfiltered arm.

Source screenshot of jev-harness
SOURCE SCREENSHOTEXPAND ↗
noul ★ 5 Benchmarks & research

Split local Laya batches across Apple Neural Engine and GPU with a Jev-shaped API; it accelerates an independent model, not Jev.

Source screenshot of Laya Fast
SOURCE SCREENSHOTEXPAND ↗
choicescorenoul ★ 5 Benchmarks & research

Rev

Rev trains three Qwen decision models and compares their accuracy, latency, cost, scaling, and Snake play with hosted Jev.

Source screenshot of Rev
SOURCE SCREENSHOTEXPAND ↗
choice ★ 5 Benchmarks & research

Follow Eikos and Jev as they paper-trade the same Hyperliquid markets from identical snapshots and precommitted rules.

Source screenshot of Eikos Arena
SOURCE SCREENSHOTEXPAND ↗
choicenoul ★ 5 Benchmarks & research

How well does Jev understand Korean, including medical text? This benchmark publishes questions, results, runtime and cost.

Source screenshot of jev-korean-benchmark
SOURCE SCREENSHOTEXPAND ↗
from the source .venv/Scripts/python.exe -X utf8 -m jevbench.overview_figure
choicescore ★ 4 Benchmarks & research

system-one-open recreates Jev-style calibrated decisions with small open Gemma models.

Source screenshot of system-one-open
SOURCE SCREENSHOTEXPAND ↗
from the source modal run build_data.py # ~5 min, CPU onlymodal run typesafe_eval.py # secondsS1_GPU=H100 modal run train.py --name lite-270m --model unsloth/gemma-3-270m-it --tier l...S1_GPU=H100 modal run evaluate.py --run lite-270mS1_GPU=H100 modal run train.py --name e2b-full --model google/gemma-4-E2B-it --tier full...
choicescorenoul ★ 4 Benchmarks & research

LitJev wraps ordinary language models in a Jev-like layer for typed choices, scores, and binary judgments.

Source screenshot of LitJev
SOURCE SCREENSHOTEXPAND ↗
choice What does the customer want?
choicescorenoul ★ 4 Benchmarks & research

jevcal derives confidence thresholds, checks calibration, and watches typed decision models for drift against an LLM teacher.

Source screenshot of jevcal
SOURCE SCREENSHOTEXPAND ↗
from the source pip install "git+https://github.com/abhixhek/jevcal" # corepip install "jevcal[anthropic] @ git+https://github.com/abhixhek/jevcal" # + Claude a...
choicenoul ★ 4 Benchmarks & research

Surveys evidence on Jev and compatible decision models, separating calibration claims, selective control, and open implementations.

Source screenshot of Awesome Jev
SOURCE SCREENSHOTEXPAND ↗
— ★ 4 Benchmarks & research

Test whether Jev can choose compiler optimisation passes from program features, with its decisions compared against measured outcomes.

Source screenshot of Jevopt
SOURCE SCREENSHOTEXPAND ↗
— ★ 4 Benchmarks & research

Build and test Jev decision workflows in Python, then use them from the command line or an MCP client.

Source screenshot of daf-jev
SOURCE SCREENSHOTEXPAND ↗
choice What is the tone? · calm / frustrated
choicescorenoul ★ 3 Benchmarks & research

jev-chat constructs chatbot replies by repeatedly turning Jev probabilities into hierarchical speculative decoding choices.

Source screenshot of jev-chat
SOURCE SCREENSHOTEXPAND ↗
from the source git clone https://github.com/adhyaay-karnwal/jev-chatcd jev-chatuv sync --extra devcp .env.example .env # set TYPESAFE_API_KEY
choicescorenoul ★ 3 Benchmarks & research

Can Jev find better agent skills than embedding search? This evaluation tests both on Chinese and English queries.

Source screenshot of jev-search-rerank-eval
SOURCE SCREENSHOTEXPAND ↗
from the source uv venv --python 3.12 .venv && uv pip install --python .venv/bin/python -e ".[dev]".venv/bin/pytest -q # 28 tests.venv/bin/python -m jse robustness # re-scores cached runs, no API key...export OPENROUTER_API_KEY=... # env only; never written to disk.venv/bin/python -m jse pool # ~30 min first time (bge-m3 over 3...
choicescore ★ 3 Benchmarks & research

jev_stock experiments with short-term market-direction forecasts derived from structured financial data.

Source screenshot of jev_stock
SOURCE SCREENSHOTEXPAND ↗
choicescore ★ 3 Benchmarks & research

This open-weights Jev alternative returns typed, calibrated choices in one Hugging Face or vLLM forward pass.

Source screenshot of open-alternative-jev
SOURCE SCREENSHOTEXPAND ↗
from the source pip install open-alternative-jev # Hugging Face backendpip install "open-alternative-jev[vllm]" # + vLLM backendpip install "open-alternative-jev[quant]" # + bitsandbytes 8-bit / 4-bit loading
choice ★ 3 Benchmarks & research

Split source code mechanically, ask Jev what each fragment is, and render the guesses as syntax highlighting.

Source screenshot of jev-lexer
SOURCE SCREENSHOTEXPAND ↗
choice ★ 3 Benchmarks & research

Tests Jev and the open-weight Laya counterpart as typed judgment layers in a look-ahead-free quantitative research stack.

Source screenshot of jev-as-quant
SOURCE SCREENSHOTEXPAND ↗
choicescorenoul ★ 3 Benchmarks & research

Ask typed questions directly about images through a shared Qwen3-VL encoding instead of first producing captions.

Source screenshot of visual-jev
SOURCE SCREENSHOTEXPAND ↗
choicescorenoul ★ 3 Benchmarks & research

Calibre measures which model suits each part of a dataset, then routes requests using those results.

Source screenshot of calibre
SOURCE SCREENSHOTEXPAND ↗
— ★ 2 Benchmarks & research

Test Jev on chess puzzles and on figuring out which game character you're talking to.

Recorded frame of jev-benchmark
ACTUAL RECORDING12 SEC ↗
from the source jevcommon/client.py rate-limited async TypeSafe client; persists every raw request/...jevchess/ chess benchmark state.py board → state at three levels (fen / ascii / rich); move descri... experiments.py A–G: Choice, Score fan-out, hierarchical, perception, mate-in-1... engine.py positions.py Stockfish ground truth; position generation
choicescorenoul ★ 2 Benchmarks & research

jev-gate experiments with assigning Claude Code tasks to different models according to Jev's judgment.

Source screenshot of jev-gate
SOURCE SCREENSHOTEXPAND ↗
choicescore ★ 2 Benchmarks & research

jev-harness adds confidence gates, shadow runs, reusable recipes, and evaluations around Jev decisions.

Source screenshot of jev-harness
SOURCE SCREENSHOTEXPAND ↗
choicescorenoul ★ 2 Benchmarks & research

jevgpt coaxes a non-generative decision model into chatting by choosing the response one step at a time.

Source screenshot of jevgpt
SOURCE SCREENSHOTEXPAND ↗
choicescorenoul ★ 2 Benchmarks & research

trade-jev backtests Jev as a buy, sell, or hold trader on Nasdaq futures order-book data.

Source screenshot of trade-jev
SOURCE SCREENSHOTEXPAND ↗
from the source uv run python -m trade_jev.run --days 2026-06-23 # one...uv run python -m trade_jev.run --days 2026-06-08,2026-06-09,2026-06-10 # sev...uv run python -m trade_jev.run --all # all...uv run python -m trade_jev.run --days 2026-06-23 --policies hold,random,imbalance # bas...
— ★ 2 Benchmarks & research

jevify helps agents spot suitable Jev tasks, frame typed questions, and draw on recent community experiments.

Source screenshot of jevify
SOURCE SCREENSHOTEXPAND ↗
noul Does `passages.a` help answer `query` with a step, prerequisite, or constraint?
choicescorenoul ★ 2 Benchmarks & research

Before building an RPG, this project tests whether Jev's characters can react to context and keep a secret.

Source screenshot of rpg-jev
SOURCE SCREENSHOTEXPAND ↗
from the source // Retries are handled here, not in the SDK, so a recorded latency is one attempt.const client = new TypeSafeClient({ defaultModel: MODEL, retry: { maxRetries: 0 } });
choicenoul ★ 2 Benchmarks & research

Metask Jev-Lab trains open typed-decision models and publishes calibrated option probabilities alongside reproducible JevBench evaluations.

Source screenshot of Metask Jev-Lab
SOURCE SCREENSHOTEXPAND ↗
choicescore ★ 2 Benchmarks & research

DecisionBridge gives existing language models an interface for explicit choices, scores, calibration, and human-review thresholds.

Source screenshot of decisionbridge
SOURCE SCREENSHOTEXPAND ↗
from the source import osfrom decisionbridge import Choicefrom decisionbridge.backends.api import APIDecisionModelmodel = APIDecisionModel("openai", os.environ["OPENAI_MODEL"], mode="json")choice = Choice("What is the main intent?", {
choicescore ★ 1 Benchmarks & research

Build a chat response one character at a time, with every character chosen by Jev.

Source screenshot of jev-freeform
SOURCE SCREENSHOTEXPAND ↗
from the source git clone https://github.com/kesku/jev-freeform.gitcd jev-freeformnpm installexport TYPESAFE_API_KEY="your-key"npm start
choice ★ 1 Benchmarks & research

Little Airways puts Jev's decision-making into a small browser-based flying demo.

Recorded frame of jev-little-airways
ACTUAL RECORDING12 SEC ↗
from the source demo/ the runnable demo (three.js WebGPU via CDN, zero build step)research/ the Jev capability dossier (live-captured API behavior, pricing, patterns)resources/ dream-loop artifacts (target, prompts, rounds) and screenshotsoriginals/ the Astra single-file build (fully offline) and the GLM round-3 build
— ★ 1 Benchmarks & research

This benchmark compares Jev with Cohere, ZeroEntropy, and a chat model on 14 reranking datasets.

Source screenshot of jev-rerank-bench
SOURCE SCREENSHOTEXPAND ↗
from the source uv synccp .env.example .env # JEV_API_KEY, OPENROUTER_API_KEY, ZEROENTROPY_API_KEYuv run candidates/build.py # downloads the datasets, builds BM25 top-30 and the "ab...uv run candidates/build_nevir.pyuv run run.py --model jev-score-batch --dataset all --workers 16
choicescorenoul ★ 1 Benchmarks & research

This harness evaluates Jev Ultrafast on curated research-browser cases and generates a field report from each run.

Source screenshot of jev-research-eval
SOURCE SCREENSHOTEXPAND ↗
from the source export JEV_ULTRAFAST_ROOT=/path/to/jev-ultrafastpython scripts/run_suite.py \ --jev-root "$JEV_ULTRAFAST_ROOT" \ --cases cases/research_browser_v1.yaml \ --out results/latest
— ★ 1 Benchmarks & research

Explore a Monte Carlo tree search that uses Gemini to propose paths and Jev to judge them.

Source screenshot of mcts-agent
SOURCE SCREENSHOTEXPAND ↗
from the source TYPESAFE_API_KEY=your_typesafe_api_key_here# Optional model overrides:AGY_MODEL=gemini-3.8-flash-medium
choicescorenoul ★ 1 Benchmarks & research

qwen-rlcd trains Qwen3.5-0.8B to make Jev-style choices, scores, and binary judgments with calibrated confidence.

Source screenshot of qwen-rlcd
SOURCE SCREENSHOTEXPAND ↗
choice Which team should handle this?
choicescorenoul ★ 1 Benchmarks & research

RISC-jeV asks how far state-in, decision-out inference can go by wiring Jev up as a RISC-V CPU.

Source screenshot of RISC-jeV
SOURCE SCREENSHOTEXPAND ↗
— ★ 1 Benchmarks & research

These OpenJev experiments classify sentence pairs, including whether one statement follows from or contradicts another.

Source screenshot of openjev-experiments
SOURCE SCREENSHOTEXPAND ↗
from the source from classify import ClassifierFactory, OpenJevSettings, TextPairclassifier = ClassifierFactory.create(OpenJevSettings())predictions = classifier.classify([ TextPair("A man is playing a guitar.", "Someone is making music."), TextPair("A man is playing a guitar.", "A man is sleeping."),
choice ★ 1 Benchmarks & research

Experiment with a Jev-powered Codenames player and a separate CLI for checking determinism, sensitivity and calibration.

Source screenshot of jev-lab
SOURCE SCREENSHOTEXPAND ↗
from the source questions[word] = style === "compact" ? { type: "noul", instructions } : { type: "noul",
choicenoul ★ 1 Benchmarks & research

An experimental pharmacy checker that asks Jev the same question five ways and looks for agreement.

Source screenshot of jev-labs
SOURCE SCREENSHOTEXPAND ↗
from the source //! - **Not deterministic.** Identical requests returned 0.03, 0.03, 0.03,//! 0.04, 0.04. This is the entire reason the protocol needs a stability//! gate; the client must not paper over it with caching.//! - **Billing is `input_tokens` only**, with a ~281-token fixed overhead per//! call and state billed once regardless of question count. Batching many
noul ★ 1 Benchmarks & research

A record of five attempts to find a worthwhile job for Jev—and the reasons each stayed out of production.

Source screenshot of jev_playground
SOURCE SCREENSHOTEXPAND ↗
from the source const r = await client.systemOne({ state: { command }, questions: { harm: noul(question) }, ...(model ? { model } : {}), });
noulscore ★ 1 Benchmarks & research

Jev, briefly

A decision model, not a chatbot.

A chat model generates text token by token and hopes you parse it. Jev never generates a word. Your code sends the state of the world plus typed questions; Jev returns every answer in one parallel pass — typed values with calibrated probabilities and a confidence score your code can trust. About 70–500 ms end to end, $0.042 per million input tokens, output free. Trained with what TypeSafe calls Reinforcement Learning for Calibrated Decisions.

choice Pick one of these.

You define the options. Jev returns the chosen option, a probability for every option, and a confidence. The workhorse of routing, agents, and games — the answer space is always legal, so there is nothing to parse and nothing to hallucinate.

score Rate this on a rubric.

You describe ordered levels — say trivial / normal / critical. Jev returns the level, the probability of each, and a confidence. Scores turn fuzzy judgment (“how severe is this log line?”) into a number-free decision your code can branch on.

noul Is this statement true?

You assert a statement; Jev returns the probability that it is true — a noul. Moderation, verification, guardrails, “does this diff actually fix the bug?”: one question, one calibrated probability.

Official material lives at typesafe.ai and docs.typesafe.ai. JevMade is an independent community registry — not affiliated with or endorsed by TypeSafe AI.

Recurring lessons

Patterns that keep showing up.

One decision per tick

Games, drones, trading bots, and browser agents all converge on the same loop: serialize the state, ask one decisive question, act, repeat. Jev’s latency makes the loop feel instant — the model lives inside the control loop, not outside it.

Confidence as a gate

The answer says what; the confidence says whether to act. The most reliable entries threshold on confidence to route edge cases to a slower model or a human — automation with an honest escape hatch.

Shrink the choice space in code

Makers rarely ask Jev to pick from everything. Local tactics prune 225 gomoku moves to ~40; DOM filters turn a page into an element table; code narrows options, Jev judges within them.

Jev judges, LLMs talk

A recurring split: Jev makes every decision cheaply and instantly, and a small LLM is only invoked when a human-facing string must actually be written. Decision and generation are separate budgets.

Made something with Jev?

Send the link — repo, live demo, post, or video — plus a line on what it does and which primitives it uses. Every entry is verified against its primary source before it ships.

Submit your experiment → hello@JevMade.com