JevMade Sign in
← Back to experiments

Benchmarks & research

Jev public operations experiments

Explore how Jev picks a faulty service from public telemetry, and how changing the input changes its answer.

Source screenshot of Jev public operations experiments
SOURCE SCREENSHOTFull screenshot ↗

What it does

The controlled fault studies keep requests, responses and review thresholds visible. A 180-call replay failed its repeatability gate: identical inputs did not keep every choice or display decision stable.

How you can use it

For research on incident diagnosis, borrow the separation between evidence, model answers and known causes. Follow the instructions for the read-only app and restore the public evidence bundles to inspect saved predictions. Browsing that app needs no service access key and makes no model calls.

Compare simple ranking methods alongside Jev, keep the cutoff for showing answers fixed, and check repeated answers before making recommendations. These are controlled faults with known start times, not proof of production readiness. The replay failed its research gate. New model calls need a separate test plan and your own access key.

Keep this for later

Sign in to bookmark experiments, guides and videos, and keep notes only you can see.

Continue to sign in

We’ll bring you back to this listing.