Before you press play
What you’ll find in the video
- Compare models on the same supported questions and separate label agreement from usable output. Inspect missed rare cases: a model can score well by always choosing the majority label.
- Count all required answer options and the full input before choosing a model. Keep unsupported requests, failed requests and wrong answers separate, and test the output contract your software needs.
- Measure response times with the state you will actually send. In the author's tests, local models were faster on short inputs but lost that advantage with a full policy handbook; retrieving shorter passages changed the comparison again.
Author-run Windows RTX 5090 comparison, not independently reproduced. The 95.23% headline is label agreement on 4,635 shared reference cases, not accuracy over all 7,671 records per profile; teacher agreement is separate. The 13 profiles include related variants and controls. Linked results disclose a Winnow option-limit error, output-format sensitivity and ambiguous references. ABCD support tests are a separate workload, not extra cases in that headline. Local precision, runtimes and GPU hardware differ from hosted Jev, whose timings include network travel. The checkout omits raw run databases and third-party case text; automated audits do not establish human semantic validity.