JevMade

Sign in
← Back to guides

JevMade field notes / Written guide

Check whether another model helps your scheduling assistant

Try Jev AI compared three Chinese scheduling workflows on 48 authored examples. Keeping request context together helped; adding Jev to judge the proposed answer introduced mistakes.

Original by Try Jev AIEvaluation

Listen to this guide

JevMade’s plain-English explanation

0:00 /

Our summary

Try Jev AI tested three workflows for a Chinese scheduling assistant. The old time check inspected isolated sentence fragments before the main model saw the request. The alternative kept relevant request context together. A third workflow gave Jev that alternative's proposed answer and the original Chinese request to judge.

The September 22 experiment used 48 authored, anonymized examples, not customer conversations. Keeping context together matched both the expected full answer and action in 46 of 48 cases. Adding Jev corrected none of its errors and introduced eight full-answer mismatches where that alternative had matched the expected answer.

This was a small test on text, not a live system connected to a real calendar. The integration lacked a safe way for the model to skip uncertain decisions, and the categories of actions overlapped. The underlying test software remains unpublished.

Key takeaways

  1. Fix missing evidence in your prompt before adding another model to check the results.
  2. Use ordinary software rules to check permission to change a calendar and calculate dates; a model's answer does not establish either.
  3. Require a safe path for the system to skip a decision when it is uncertain.

The authors used 48 examples they created, mostly once each, with expected answers based on their own scheduling rules. Test software and full example records are unpublished. The added check lacked a safe way to skip uncertain decisions and used overlapping answer categories. Eight mismatches do not mean eight dangerous calendar changes: no real calendar changes occurred. This is not a general model verdict. No individual author is named; credit remains with Try Jev AI.

tryjevai.com · Original published

Read the original guide Opens the author’s site in a new tab.

Keep this for later

Sign in to bookmark experiments, guides and videos, and keep notes only you can see.

Continue to sign in

We’ll bring you back to this listing.