JevMade hello@JevMade.com
← Back to guides

JevMade field notes / Research paper

Testing if AI assistants follow option names instead of their instructions

You will learn how researchers tested whether AI assistants choose answers based on familiar labels like yes or no instead of following the actual rules provided for those choices.

Original by Yu Sun, Junhao Xu, Jiajia Shi, and Zijin YangEvaluation

Listen to this guide

JevMade’s plain-English explanation

0:00 /

AI narration

Credits

“Type-Safe Is Not Error-Free: A Constrained Decision Head Follows the Option Name, Not the Rubric Bound to It” by Yu Sun, Junhao Xu, Jiajia Shi, and Zijin Yang. Read the original source.

This expanded guide is an AI-narrated adaptation prepared by JevMade. It expands the source’s essential ideas, examples and caveats in JevMade’s own words and is not a word-for-word reading. The synthetic voice does not imitate the author or imply their endorsement.

Our summary

Jev is an AI tool that chooses from options rather than writing an answer. People use these tools to sort incoming messages or make automated workflow decisions. This paper tests if these tools actually follow the provided rules or just guess based on familiar labels.

The authors swapped the definitions attached to labels like yes and no across 1,200 test questions. They checked if the software still picked the correct definition or if it blindly followed the normal meaning of the label. They compared these results against neutral labels like numbers.

This guide is useful for people designing automated choices. The study only covered English questions. The authors suggest randomizing labels during training to fix this issue, but they did not test that idea. Hosted tools act as hidden systems, so the exact internal causes remain unknown.

Key takeaways

  1. Swapping definitions for familiar labels caused the tested AI models to frequently change their final decisions.
  2. Using neutral labels like numbers or random letters helped the software follow the actual instructions better.
  3. The open Laya model's accuracy dropped sharply, while the hosted Jev model showed a smaller decline.

The authors report these results, but JevMade did not independently reproduce them. Hosted Jev is a closed system observed only through its final outputs, so its internal mechanics are not established here.

arXiv preprint · Source reviewed · Original published

Read the original guide Opens the author’s site in a new tab.