Before you press play
What you’ll find in the video
- Check confidence as well as accuracy. Calibration asks whether groups of answers with a stated probability are right about that often; it is not a guarantee for one answer or a new task.
- Match the model to the task. Briggs uses GLiNER2.5 for intent and tool selection, and GLiGuard for guardrail checks. Similar jobs do not establish identical outputs or equally useful confidence scores.
- Include the deployment in a speed comparison. Local GLiNER on an M4 Pro avoids the network trip to hosted Jev; in Briggs's Doom deathmatch results, coaching adds much more to GLiNER than to Jev.
Creator-reported comparisons, not independently reproduced benchmarks. Server time and end-to-end response time are different measures, and local hardware is not a hosted API. The language-model comparison uses reasoning off. The Doom discussion reports only a small Jev coaching gain, not a blanket fourfold improvement. These examples do not prove universal latency, model equivalence, safe guardrails or reliable confidence on new data. Selecting a tool is not generating its arguments or executing it. No models or games were run for these notes.