Our summary
A daily stock-news brief needs readable writing and a way to decide what to read first. M Lakshmanan reports trying Jev alongside Haiku's existing writing and classification. His proposed division gives Haiku the prose and Jev the labels; the accompanying diagram says that clean separation is not fully implemented yet.
Jev chooses from a fixed list of event types, including Other when none fits. The author reports confident Other answers for a strike and a product launch. His gap detector only watched low-confidence answers, so it would have missed these missing categories. Checking the selected label matters as well as checking uncertainty.
The two models disagreed on at least one field in twenty-one of twenty-seven items from one day. That is a reason to investigate, not evidence that Jev was right. The author explicitly calls the comparison unfair: Jev had fixed choices while Haiku wrote its labels in a text response. The thread provides no public code to inspect.
Key takeaways
- Keep writing and categorisation separate when they need different tools.
- Review confident Other answers as well as low-confidence choices when looking for missing categories.
- Disagreement between models does not tell you which answer is correct, especially when their answer formats differ.