JevMade hello@JevMade.com

Explore the Jev ecosystem.

Learn, build, and play with Jev — from deep technical guides to creative experiments.

Make & explore

Experiments.

See what builders made with Jev. Games, tools, repositories, articles, and more — traced to their sources.

Featured experiment Jev experiments by Nader Dabit
Recorded frame from Jev experiments
Explore 1,735 entries
The directory03 / 03 · Experiments

Find something worth exploring.

Every category, from games and tools to repositories and writeups, traced to a primary source.

Same ideas.
Different paths.

This is the full registry, not just the featured picks. Figures like speed, cost, and stars are maker-reported or captured snapshots, not JevMade measurements.

268 experiments · showing 121–180

Switchloom experiments with Jev-routed persistent coding agents and publishes pilots where routing failed to beat a single Astra model.

Source screenshot of Switchloom
SOURCE SCREENSHOTEXPAND ↗
choicenoul ★ 1 Benchmarks & research

This benchmark compares Jev, Gemini, and GPT on structured annotation of São Paulo court judgments.

Source screenshot of jev-anotacao-sentencas
SOURCE SCREENSHOTEXPAND ↗
from the source uv syncuv run python scripts/01_coletar.pyuv run python scripts/02_anotar.py jev-latest jev-preview gemini-3.8-flash gpt-5.6-luna... ref-gpt-5.6-sol ref-gemini-3.1-prouv run python scripts/03_referencia.py # lista divergências e auditoria
— ★ 0 Benchmarks & research

A position paper arguing that Jev-style models need fuzzy and Hidden Markov primitives before they settle on crisp decisions.

Source screenshot of jev-deferred-crispification
SOURCE SCREENSHOTEXPAND ↗
from the source git clone https://github.com/dnakhoa/jev-deferred-crispificationcd jev-deferred-crispificationpython3 experiments/run_all.py # E1–E5 + lemma checks, CPU, ~2 min
choicescorenoul ★ 0 Benchmarks & research

This project tests whether Jev can resolve ambiguous Japanese addresses against Japan Post's KEN_ALL data.

Source screenshot of jev-jp-address
SOURCE SCREENSHOTEXPAND ↗
from the source npm installscripts/download-data.sh # data/utf_ken_all.csv, data/jigyosyo_utf8.csv を取得...export TYPESAFE_AI_API_KEY=... # Jev を使う場合のみnpm run build # dist/cli.js
choice ★ 0 Benchmarks & research

jev-lab collects TypeScript experiments on Jev's behavior, accuracy, and response latency.

Source screenshot of jev-lab
SOURCE SCREENSHOTEXPAND ↗
choicescorenoul ★ 0 Benchmarks & research

A phishing test that pits Jev against Claude Haiku 4.5 on 2,000 emails.

Source screenshot of jev-phishing-bench
SOURCE SCREENSHOTEXPAND ↗
from the source cp .env.example .env # then fill in the keys (baseline used: claude-haiku-4-5...uv run prepare_data.py # download + checksum + data/emails.jsonluv run net_floor.py # network round-trip floor to each API hostuv run run_jev.py --limit 10 # smoke test, prints raw answersuv run run_jev.py # pass 1, all 2 000 emails
choicenoul ★ 0 Benchmarks & research

This reproducible MuJoCo pilot compares Jev, Claude Haiku, and reactive rules on pick-and-place control.

Recorded frame of jev-pick-and-place-study
ACTUAL RECORDING12 SEC ↗
from the source .venv\Scripts\python.exe -m zipfile -e results\evaluation_traces.zip outputs\comparison.venv\Scripts\python.exe demo.py render outputs\comparison\eval_jev_ordinary_1001.jsonl
— ★ 0 Benchmarks & research

This game playground compares Jev with other evaluators using explicit states, legal moves, and observable outcomes.

Source screenshot of jev-playground (hegargarcia)
SOURCE SCREENSHOTEXPAND ↗
from the source flowchart LR state[Current state] --> actions[Legal actions] actions --> model[Model evaluation] model --> choice[Validated choice] choice --> next[Next state]
choicescore ★ 0 Benchmarks & research

HackSing puts Jev through 50 tests and publishes the findings in a 52-page Chinese-language report.

Source screenshot of jev-report
SOURCE SCREENSHOTEXPAND ↗
choicescorenoul ★ 0 Benchmarks & research

Can Jev tell a real credential from an innocent string? This benchmark tests it on file snippets.

Source screenshot of jev-secret-detection
SOURCE SCREENSHOTEXPAND ↗
from the source echo "TYPESAFE_API_KEY=..." > .envuv run python main.py
scorenoul ★ 0 Benchmarks & research

A test of whether Jev can tell genuine shadcn-ui lint problems from false alarms.

Source screenshot of jev-shadcn-lint-eval
SOURCE SCREENSHOTEXPAND ↗
from the source node extract-design-system.mjs --repo <lint checkout> # only when the lint repo change...export TYPESAFE_API_KEY=... # never commit itnode run-rule-cases.mjs rule-cases.jsonl --repo <lint checkout> [--limit n]
score ★ 0 Benchmarks & research

This notebook compares zero-shot spam judgments from Jev with conventional TF-IDF classifiers.

Source screenshot of jev-spam-eval
SOURCE SCREENSHOTEXPAND ↗
from the source # New output directory: a three-message smoke test, then the full comparison.JEV_CONTEXT_OUTPUT_DIR=experiments/jev-context/rerun \ uv run python experiments/jev-context/evaluate.py --limit 3JEV_CONTEXT_OUTPUT_DIR=experiments/jev-context/rerun \ uv run python experiments/jev-context/evaluate.py
choicenoul ★ 0 Benchmarks & research

This classifier distinguishes agent-written from human-written pages in the collusion.wiki corpus.

Source screenshot of jev-trace-classifier
SOURCE SCREENSHOTEXPAND ↗
from the source ground_truth.py # the per-page agent/human label join (the only "ground truth" logic)baseline_regex.py # content-only floor: regex tells -> shows text alone is blindjev_client.py # minimal TypeSafe Jev client (noul primitive, plain HTTP)qwen_client.py # local Qwen3.8-Flash-Next client (OpenAI-compatible)compare.py # the head-to-head harness + metrics table + JSONL dump
scorenoul ★ 0 Benchmarks & research

A Jev-inspired experiment that makes visual decisions on an iPhone using Qwen3-VL, not TypeSafe's service.

Source screenshot of PocketJev
SOURCE SCREENSHOTEXPAND ↗
choicescore ★ 0 Benchmarks & research

This study tests Jev as a fast gate for suspicious agent actions in SHADE-Arena.

Source screenshot of shade-arena-jev-monitor
SOURCE SCREENSHOTEXPAND ↗
from the source conda create -n shade-arena python=3.11conda activate shade-arenapip install -r requirements-jev.txt # minimal pinned deps; upstream requirements.t...conda env config vars set PYTHONUTF8=1 # task data is UTF-8; Windows defaults to cp125...conda deactivate && conda activate shade-arena
score ★ 0 Benchmarks & research

This Rust port runs Jev-style Choice, Score, and Noul evaluations through ordinary language models.

Source screenshot of system-one-adapter-rust
SOURCE SCREENSHOTEXPAND ↗
from the source use system_one_adapter::{ Noul, ProviderName, SystemOneAdapterClient, SystemOneArgs,};let client = SystemOneAdapterClient::new(true, system_one_adapter::AnswerMode::Probabili... .normalize_probabilities(true);
choicescorenoul ★ 0 Benchmarks & research

system-one-gemma adds a scoring head to Gemma 3 270M for calibrated decisions without text generation.

Source screenshot of system-one-gemma
SOURCE SCREENSHOTEXPAND ↗
from the source from infer import load_trained_model, scoretok, model = load_trained_model("./pretrained-scorer")probs = score(model, tok, state="Patient has chest pain radiating to left arm, ST elevation on ECG", question="What is the triage level?",
choicescorenoul ★ 0 Benchmarks & research

typesafe-oracles tests when a typed Jev judgment is more useful than a conventional language-model call.

Source screenshot of typesafe-oracles
SOURCE SCREENSHOTEXPAND ↗
from the source npm run dataset # build the binary corpus -> .sandboxes/episodes.jsonnpm run run # arm A (Jev) + arm B (Haiku via `claude -p`)npm run run:api # arm C (Haiku via the API — the honest baseline)npm run report # accuracy, confidence separation, calibration, costnpm run run:repeat # 5 independent runs of arm A
choicescorenoul ★ 0 Benchmarks & research

See how Jev fares in solved poker spots—and how much the wording of the game state changes its play.

Source screenshot of Jev is the fish at the poker table
SOURCE SCREENSHOTEXPAND ↗
choice Hero is on the turn and it is hero's turn to act. Which action from `decision.legal_actions` should hero take? · check / all_in
choicenoul Benchmarks & research

Match Jev against a text model in chess, where Jev selects legal moves and its opponent must produce valid notation.

Recorded frame of jev-chess
ACTUAL RECORDING12 SEC ↗
choice ★ 0 Benchmarks & research

Replay a documented Jev-versus-Stockfish tournament with legal moves, raw responses, complete games, and provisional ratings.

Source screenshot of jev-chess-bench
SOURCE SCREENSHOTEXPAND ↗
choice ★ 0 Benchmarks & research

Play terminal Sudoku yourself or watch Jev choose cell-and-digit moves while the benchmark tracks whether it solves the grid.

Source screenshot of Sudoku vs Jev
SOURCE SCREENSHOTEXPAND ↗
choice ★ 0 Benchmarks & research

Compare Jev with a conventional language model on route planning, Minesweeper, Hanoi, and a simulated drone.

Source screenshot of JEV Playground
SOURCE SCREENSHOTEXPAND ↗
choicescorenoul ★ 0 Benchmarks & research

Compare Jev’s Battleship shots with random, hunt-and-target, and placement-density strategies on seeded games.

Source screenshot of Battleship vs. Jev
SOURCE SCREENSHOTEXPAND ↗
choice ★ 0 Benchmarks & research

Pit Jev against random and hand-tuned players in Tetris, Snake, and 2048 while logging confidence for every move.

Source screenshot of jev-arcade
SOURCE SCREENSHOTEXPAND ↗
choice ★ 0 Benchmarks & research

A Chinese-language experiment suite for Jev's context limits, bilingual judgments, batching and game decisions.

Source screenshot of jev-study
SOURCE SCREENSHOTEXPAND ↗
choicescorenoul ★ 0 Benchmarks & research

An original platformer and reinforcement-learning lab testing whether frozen Jev risk features help PPO learn faster.

Source screenshot of mario-play
SOURCE SCREENSHOTEXPAND ↗
from the source API_URL = "https://api.typesafe.ai/v1/systemone"MODEL = "jev-1.13.0"ACTION_NAMES = tuple(action_names("simple"))ACTION_CRITERIA = { "noop": "Release all buttons; coast or wait, and release jump.",
choicenoul ★ 0 Benchmarks & research

Political Compass and Pew typology probes for Jev, alongside a deliberately awkward character-by-character text decoder.

Source screenshot of Jev political benchmarks
SOURCE SCREENSHOTEXPAND ↗
from the source def choice(instructions: str, criteria: dict[str, str | None]) -> dict: """A Choice question: pick one of up to 255 options.""" return {"type": "choice", "instructions": instructions, "criteria": criteria}
choicescorenoul ★ 0 Benchmarks & research

Battleship where Jev returns a probability of a ship for every open cell, creating both the heat map and the shot.

Source screenshot of jev-sonar
SOURCE SCREENSHOTEXPAND ↗
from the source def argmax_cell(probabilities: Mapping[str, float], cells: Sequence[Cell]) -> Cell: if not cells: raise ValueError("no cells to choose from") best_cell = cells[0] best_probability = float(probabilities.get(cell_id(best_cell), 0.0))
noulchoice ★ 0 Benchmarks & research

A roguelike dungeon director where Jev sketches the next room's role, size, danger, enemies, loot and secrets.

Source screenshot of ai-hack-poc
SOURCE SCREENSHOTEXPAND ↗
from the source "room_type": { "type": "choice", "choice": "chamber", "confidence": 0.8, "probabilities": {"chamber": 0.8, "room": 0.2},
choicescorenoul ★ 0 Benchmarks & research

Runs Jev and a frontier model down parallel lanes, answering the same timed traffic-hazard choices and comparing the scorecards.

Source screenshot of Jev Reflex
SOURCE SCREENSHOTEXPAND ↗
from the source export const CALL_CRITERIA = { left: "Swerve left. The hazard is real and the left side of the road is open.", right: "Swerve right. The hazard is real and the right side of the road is open.", hold: "Hold the line. Do not swerve. The markings disagree with the open side, neither...
choice ★ 0 Benchmarks & research

A controlled city simulation comparing four kinds of institutional memory, with Jev scoring each term beside a deterministic scoreboard.

Source screenshot of Four Mayors
SOURCE SCREENSHOTEXPAND ↗
from the source ENDPOINT = "https://api.typesafe.ai/v1/systemone"MODEL = os.getenv("TYPESAFE_MODEL", "jev-latest") "prosperity": { "type": "score", "instructions": "Prosperity over the whole term: jobs relative to population at...
score ★ 0 Benchmarks & research

Turns Jev probability distributions into one plain-language sureness verdict, from certain through torn to clueless.

Source screenshot of How sure is Jev?
SOURCE SCREENSHOTEXPAND ↗
from the source QUESTIONS = { "department": {"type": "choice", "instructions": "Which team should handle this mess... "severity": { "type": "score", "instructions": "How severe is the reported issue? If there is no issue, pick th...
choicescore ★ 0 Benchmarks & research

Run open-Jev decisions in your browser without an API key or inference server.

Source screenshot of jev-web
SOURCE SCREENSHOTEXPAND ↗
choicenoulscore Benchmarks & research

Compare Jev’s typed classifications with classical ML pipelines across the maker’s eight datasets.

Source screenshot of Jev vs. ML
SOURCE SCREENSHOTEXPAND ↗
choice Benchmarks & research

Replay six Jev chess experiments, from unaided legal moves to tactical filters and Stockfish-assisted variants.

Recorded frame of Jev Chess Lab
ACTUAL RECORDING12 SEC ↗
choice Benchmarks & research

Jev, briefly

A decision model, not a chatbot.

A chat model generates text token by token and hopes you parse it. Jev never generates a word. Your code sends the state of the world plus typed questions; Jev returns every answer in one parallel pass — typed values with calibrated probabilities and a confidence score your code can trust. About 70–500 ms end to end, $0.042 per million input tokens, output free. Trained with what TypeSafe calls Reinforcement Learning for Calibrated Decisions.

choice Pick one of these.

You define the options. Jev returns the chosen option, a probability for every option, and a confidence. The workhorse of routing, agents, and games — the answer space is always legal, so there is nothing to parse and nothing to hallucinate.

score Rate this on a rubric.

You describe ordered levels — say trivial / normal / critical. Jev returns the level, the probability of each, and a confidence. Scores turn fuzzy judgment (“how severe is this log line?”) into a number-free decision your code can branch on.

noul Is this statement true?

You assert a statement; Jev returns the probability that it is true — a noul. Moderation, verification, guardrails, “does this diff actually fix the bug?”: one question, one calibrated probability.

Official material lives at typesafe.ai and docs.typesafe.ai. JevMade is an independent community registry — not affiliated with or endorsed by TypeSafe AI.

Recurring lessons

Patterns that keep showing up.

One decision per tick

Games, drones, trading bots, and browser agents all converge on the same loop: serialize the state, ask one decisive question, act, repeat. Jev’s latency makes the loop feel instant — the model lives inside the control loop, not outside it.

Confidence as a gate

The answer says what; the confidence says whether to act. The most reliable entries threshold on confidence to route edge cases to a slower model or a human — automation with an honest escape hatch.

Shrink the choice space in code

Makers rarely ask Jev to pick from everything. Local tactics prune 225 gomoku moves to ~40; DOM filters turn a page into an element table; code narrows options, Jev judges within them.

Jev judges, LLMs talk

A recurring split: Jev makes every decision cheaply and instantly, and a small LLM is only invoked when a human-facing string must actually be written. Decision and generation are separate budgets.

Made something with Jev?

Send the link — repo, live demo, post, or video — plus a line on what it does and which primitives it uses. Every entry is verified against its primary source before it ships.

Submit your experiment → hello@JevMade.com