JevMade hello@JevMade.com
← Back to experiments

Benchmarks & research

padflow-jev-evals

PadFlow contributes anonymized land-development decisions for testing confidence-aware models such as Jev.

Source screenshot of padflow-jev-evals
SOURCE SCREENSHOTFull screenshot ↗

What it does

The repository includes task schemas, labeled rows, a runner, and calibration checks for document routing and transaction coding.

How you can use it

Your developer can borrow this testing script to measure different AI models. They can check how well a model sorts incoming documents or labels business expenses. The tool records how often the AI finds the correct answers. It also tracks how sure the AI is when making a choice.

To start the test, your developer runs the provided script. They will need an access key that connects the tool to a service like OpenRouter. The included public examples are very small. They help you compare models, but they cannot prove a system is fully ready.

Maker-reported (not independently measured by JevMade): model that is 90% accurate and always says 0.95 · n = 5 / 5 / 7 rows per decision (route_document / code_transaction /

Primitives
score
Platform
Python
Added
Project created
GitHub stars
1 (snapshot, not live)