JevMade field notes / Bilingual capability research
Jev Capability Atlas
A Chinese-first evidence map that combines the author's reproducible suites with external studies to examine bounded decisions, calibration failures, and task loss during decomposition.
Original by ZaiousEvaluationGitHub repositorySource reviewed
Before you dive in
What you’ll find in the original
Separate calibration across a population from correctness on one item; high confidence can still be wrong on overlapping labels.
After reformulating a task as bounded decisions, test whether the retained evidence can still complete the original job.
Test decomposition on hard benign cases because more atomic checks can increase false positives.
Worth knowing
This is a synthesis with linked receipts, not one controlled benchmark. It marks its own experiments separately from documentation and third-party reports.