JevMade Sign in
← Back to videos

JevMade field notes / Video guide

I Tested Jev vs 12 Local Decision Models. Here's What I’d Use...

Compare hosted Jev with local decision-model profiles using Jev Arena. The walkthrough examines task fit and input limits, and shows why short-input speed results can reverse with a long policy handbook.

Original by The AI AutomatorsEvaluationIntermediate12 min 18 sec Published

Before you press play

What you’ll find in the video

  1. Compare models on the same supported questions and separate label agreement from usable output. Inspect missed rare cases: a model can score well by always choosing the majority label.
  2. Count all required answer options and the full input before choosing a model. Keep unsupported requests, failed requests and wrong answers separate, and test the output contract your software needs.
  3. Measure response times with the state you will actually send. In the author's tests, local models were faster on short inputs but lost that advantage with a full policy handbook; retrieving shorter passages changed the comparison again.
Worth knowing

Author-run Windows RTX 5090 comparison, not independently reproduced. The 95.23% headline is label agreement on 4,635 shared reference cases, not accuracy over all 7,671 records per profile; teacher agreement is separate. The 13 profiles include related variants and controls. Linked results disclose a Winnow option-limit error, output-format sensitivity and ambiguous references. ABCD support tests are a separate workload, not extra cases in that headline. Local precision, runtimes and GPU hardware differ from hosted Jev, whose timings include network travel. The checkout omits raw run databases and third-party case text; automated audits do not establish human semantic validity.

I Tested Jev vs 12 Local Decision Models. Here's What I’d Use...

Keep this for later

Sign in to bookmark experiments, guides and videos, and keep notes only you can see.

Continue to sign in

We’ll bring you back to this listing.