JevMade Sign in
← Back to experiments

Benchmarks & research

system-one-gemma

system-one-gemma adds a scoring head to Gemma 3 270M for calibrated decisions without text generation.

Source screenshot of system-one-gemma
SOURCE SCREENSHOTFull screenshot ↗

What it does

It was trained on 12,913 questions spanning six tasks; accuracy is modest, but reported confidence closely tracks observed correctness.

How you can use it

Your developer can use this tool to sort incoming customer messages. You give the app a message and a list of departments. The tool does not write a reply. Instead, it assigns a percentage score to each department. This helps the app decide if a message belongs in technical support or billing.

A developer can start by testing the included code. You must first agree to the maker's rules on their website. Then, the developer uses an access key to download the files. The ready-made version restricts commercial work. Your developer must train the tool on your own data to remove this limit.

Maker-reported (not independently measured by JevMade): Trained on 12,913 questions across 6 tasks, evaluated on held-out data: · The accuracy is modest (it's a 270M model!) but the calibration is excellent — when it says 80% confidence, it's right ~80% of the time. That's the real value: you can trust the probabilities

Primitives
choice, score, noul
Platform
Python
Added
Project created
GitHub stars
0 (snapshot, not live)

Keep this for later

Sign in to bookmark experiments, guides and videos, and keep notes only you can see.

Continue to sign in

We’ll bring you back to this listing.