Hüseyin Babal wrote 1,000 trick questions, laced with sarcasm, hidden severity and planted instructions, and sent the same requests to Jev and three smaller local models.
Jev got the most right, while the local models struggled most with 1-to-10 scores. Every question, label and reply is public, so anyone can check a result or rerun it against another model. He wrote and labelled the questions himself, calls 42 labels debatable, and most of the set is customer messages.
How you can use it
If you are picking a decision model for customer messages, this gives you a thousand awkward examples to test it with. Look for the kinds of trick that match your own messages, such as sarcasm or a serious problem described politely, and see which model stays right, and how confident each one was when it was wrong.