This Japanese walkthrough compares Jev and Claude on 100 mock customer inquiries. It explains the question types, reports the creator’s timing and accuracy results, and demonstrates threshold-based routing to another model or a person.
Original by AIツールの達人【AI・Web・ChatGPT最新情報】EvaluationIntermediate14 min 45 secPublished Source reviewed
Before you press play
What you’ll find in the video
Jev returns binary probabilities, categorical choices, or ordinal scores in a structured response.
The creator reports competitive categorization accuracy and lower time and cost on this 100-inquiry test, not a general benchmark.
The routing example uses a confidence threshold to decide what to automate and what to escalate; the shown threshold needs validation for another task.
Worth knowing
Gemini-assisted video/transcript review. Official documentation notes diminished accuracy in Japanese/CJK compared to English, and reported benchmark figures are author experiments on synthetic test sets rather than controlled industry standards.