Our summary
Try Jev AI tested three workflows for a Chinese scheduling assistant. The old time check inspected isolated sentence fragments before the main model saw the request. The alternative kept relevant request context together. A third workflow gave Jev that alternative's proposed answer and the original Chinese request to judge.
The September 22 experiment used 48 authored, anonymized examples, not customer conversations. Keeping context together matched both the expected full answer and action in 46 of 48 cases. Adding Jev corrected none of its errors and introduced eight full-answer mismatches where that alternative had matched the expected answer.
This was a small test on text, not a live system connected to a real calendar. The integration lacked a safe way for the model to skip uncertain decisions, and the categories of actions overlapped. The underlying test software remains unpublished.
Key takeaways
- Fix missing evidence in your prompt before adding another model to check the results.
- Use ordinary software rules to check permission to change a calendar and calculate dates; a model's answer does not establish either.
- Require a safe path for the system to skip a decision when it is uncertain.