JevMade Sign in
← Back to guides

JevMade field notes / Written guide

Test Jev on phishing checks and support chats

Hostinger tests Jev on suspicious web addresses and support chats. Its results show why a cheaper, faster model can suit one task but lose too much accuracy on another.

Original by Simon Lim and Mantas LukauskasEvaluation

Listen to this guide

JevMade’s plain-English explanation

0:00 /

Our summary

Simon Lim and Mantas Lukauskas at Hostinger ask whether Jev can replace the models their team already uses. They test two jobs: spotting web addresses used for phishing, where criminals try to steal information, and sorting customer support chats by product and the customer’s intent.

For the phishing test, they choose a confidence cutoff using training examples, then freeze it before testing 781 other addresses. Jev gets 88.5% right, compared with 90.0% for their best model. It answers much faster. On 528 support chats, however, its product labels are less accurate than GPT-5 mini’s.

The authors consider the phishing tradeoff acceptable, but would not switch their product sorter today. These figures are the authors’ measurements, not independently repeated tests. Their lesson is to compare against the model you actually use, test on separate examples and count the work caused by wrong answers, not just the price of a call.

Key takeaways

  1. Choose a confidence cutoff on training examples, freeze it, then test on examples you did not use to choose it.
  2. Compare speed, price and accuracy with your current system. A saving can be worth a small accuracy loss on one task and unacceptable on another.
  3. Use ordinary code for exact checks. Break larger work into small questions, and leave writing and planning to a model suited to those jobs.

The authors report mean phishing-check times of 0.38 seconds for Jev, compared with 2.03 seconds for self-hosted Gemma 4 31B. Product accuracy on support chats is 79.3% with Jev, compared with 86.9% for GPT-5 mini; getting both product and intent right reaches 42.9% versus 42.4%. The article does not publish its datasets or an executable benchmark. Its partner-content workflow is a proposal, not another completed Jev deployment.

Hostinger Blog · Original published

Read the original guide Opens the author’s site in a new tab.

Keep this for later

Sign in to bookmark experiments, guides and videos, and keep notes only you can see.

Continue to sign in

We’ll bring you back to this listing.