Our summary
Checking every line of a game translation with a text-writing AI can take time and money. Morikawa kota places Jev before that review. It judges whether a line deserves closer attention, while the other model still explains problems and suggests changes.
The author tests fifty correct translations and fifty faulty ones. At the chosen cutoff, 42% of review calls are skipped while 96% of the other model’s findings are retained. Missing scores or service errors send the line to the longer review instead. Ordinary code checks placeholders and text length.
Retaining another model’s findings is not the same as finding every real error. Neither model catches terminology inconsistencies without the game’s glossary. The test is small, and no separate dataset verifies the chosen cutoff. Measure misses on your own translations before relying on this shortcut.
Key takeaways
- Check what the filter misses, not only how many calls it saves. The reported 96% measures retained AI findings, not correctness against all human-labelled defects.
- Provide the glossary needed to judge consistent terminology. Use ordinary code for missing placeholders and exact character limits rather than asking either model to count.
- Send a line to the longer review when Jev fails or returns no usable score. Check the actual response structure; the author encountered answers wrapped inside a result object.