JevMade Sign in
← Back to guides

JevMade field notes / Model architecture and benchmark explanation

Keep related AI answers consistent with shared rules

Mary Newhauser introduces GLiNER2.5-Decide, which chooses related answers together under supplied rules, and reports Fastino’s comparisons with other decision models.

Original by Mary NewhauserGuardrails

Listen to this guide

JevMade’s plain-English explanation

0:00 /

AI narration

Credits

“GLiNER2.5-Decide: An Open-Weight Model for Structured Decision Making” by Mary Newhauser. Read the original source.

This expanded guide is an AI-narrated adaptation prepared by JevMade. It expands the source’s essential ideas, examples and caveats in JevMade’s own words and is not a word-for-word reading. The synthetic voice does not imitate the author or imply their endorsement.

Our summary

Fastino’s GLiNER2.5-Decide answers questions with supplied choices and can check related answers together. Mary Newhauser describes a model small enough to run on your own computer without a graphics card. It is a competing model, not Jev; the article compares several alternatives using Fastino’s own test collection.

Suppose one question detects a harmful instruction while another labels the same message safe. You can supply a rule that detected harm requires an unsafe answer. The model looks for the highest-scoring combination allowed by that rule, then reports whether the combination satisfies the rules. Consistency does not prove the judgments correct.

The reported accuracy and speed depend on the tests, computers and text lengths used. One comparison model, JevK5, was made separately to reproduce Jev’s approach; it is not TypeSafe’s own model. Fastino used a different test collection. JevMade has not repeated the tests or verified running the model without an internet connection.

Key takeaways

  1. Write explicit rules when several answers must agree; checking them together differs from asking independent questions.
  2. Answers can obey every rule and still be wrong; test against examples whose correct answers are known.
  3. Reported results depend on which model, tests, computer and text lengths were used, not just the model’s name.

Fastino reports 60.1% average accuracy across 17 maker-created sets of test cases. Its speed tests used short text and specific machines, not typical consumer computers. These results were not reproduced. Answers that assign a category do not identify supporting passages; tasks that extract information can identify its exact location in the text.

Fastino Blog · Original published

Read the original guide Opens the author’s site in a new tab.

Keep this for later

Sign in to bookmark experiments, guides and videos, and keep notes only you can see.

Continue to sign in

We’ll bring you back to this listing.