Our summary
Simon Lim and Mantas Lukauskas at Hostinger ask whether Jev can replace the models their team already uses. They test two jobs: spotting web addresses used for phishing, where criminals try to steal information, and sorting customer support chats by product and the customer’s intent.
For the phishing test, they choose a confidence cutoff using training examples, then freeze it before testing 781 other addresses. Jev gets 88.5% right, compared with 90.0% for their best model. It answers much faster. On 528 support chats, however, its product labels are less accurate than GPT-5 mini’s.
The authors consider the phishing tradeoff acceptable, but would not switch their product sorter today. These figures are the authors’ measurements, not independently repeated tests. Their lesson is to compare against the model you actually use, test on separate examples and count the work caused by wrong answers, not just the price of a call.
Key takeaways
- Choose a confidence cutoff on training examples, freeze it, then test on examples you did not use to choose it.
- Compare speed, price and accuracy with your current system. A saving can be worth a small accuracy loss on one task and unacceptable on another.
- Use ordinary code for exact checks. Break larger work into small questions, and leave writing and planning to a model suited to those jobs.