AVB explains BEV, an independent decision model inspired by Jev. It adapts pretrained Qwen layers and adds a small head that compares choices, answers yes/no questions and scores ordered levels.
Original by Neural Breakdown with AVBClassificationAdvanced108 min 59 secPublished
Adapt a pretrained Qwen backbone with a decision head and optional LoRA adapters. This is not pretraining a language model from zero or revealing TypeSafe’s private Jev architecture.
Start each option at the same position IDs and stop it from reading other options in the backbone. The decision head then compares their embeddings; check near-ties under reduced precision.
Test dates, arithmetic, rule exceptions and longer inputs on fresh examples. The companion training code selects its best checkpoint using test accuracy, so keep a separate evaluation set.
Worth knowing
Maker-reported results; the companion code was read, not run. Weights are noncommercial (CC-BY-NC-4.0); bev-decider code is Apache-2.0, bev-train has no pinned license, and the dataset has mixed rights. Training uses 1,024 state tokens versus a 2,048-token released default and is mainly English. Test feedback and source-family overlap are evaluation risks, not proof of exact contamination.