JevMade Sign in
← Back to guides

JevMade field notes / Written guide

Filter translation rows before a longer AI review

A Loclint developer asks Jev which game translations need a longer AI review. A small test shows how changing the filtering threshold saves calls but can also miss problems.

Original by morikawa kotaEvaluation

Listen to this guide

JevMade’s plain-English explanation

0:00 /

Our summary

Checking every line of a game translation with a text-writing AI can take time and money. Morikawa kota places Jev before that review. It judges whether a line deserves closer attention, while the other model still explains problems and suggests changes.

The author tests fifty correct translations and fifty faulty ones. At the chosen cutoff, 42% of review calls are skipped while 96% of the other model’s findings are retained. Missing scores or service errors send the line to the longer review instead. Ordinary code checks placeholders and text length.

Retaining another model’s findings is not the same as finding every real error. Neither model catches terminology inconsistencies without the game’s glossary. The test is small, and no separate dataset verifies the chosen cutoff. Measure misses on your own translations before relying on this shortcut.

Key takeaways

  1. Check what the filter misses, not only how many calls it saves. The reported 96% measures retained AI findings, not correctness against all human-labelled defects.
  2. Provide the glossary needed to judge consistent terminology. Use ordinary code for missing placeholders and exact character limits rather than asking either model to count.
  3. Send a line to the longer review when Jev fails or returns no usable score. Check the actual response structure; the author encountered answers wrapped inside a result object.

Japanese article about the author’s Loclint translation-review workflow. In a 100-item self-evaluation, cutoff 0.25 retains 96% of the downstream model’s findings while skipping 42% of calls; cutoff 0.30 retains 92% while skipping 53%. The article describes ten error categories but its category table has nine rows. No held-out threshold validation, production telemetry or independent run was verified. Its Cloudflare example and nested-response handling were read, not executed. Loclint is a companion product link, not another catalogue entry.

Zenn / morikawa kota · Original published

Read the original guide Opens the author’s site in a new tab.

Keep this for later

Sign in to bookmark experiments, guides and videos, and keep notes only you can see.

Continue to sign in

We’ll bring you back to this listing.