Our summary
A website's catch-all category can hide unclear sorting rules. GK used Jev to classify titles and summaries from 43 articles on GKMix, then compared its choices with the site's existing labels. The companion tutorial explains how to request categories, ratings and yes-or-no judgments.
GK reports agreement on 40 of 43 article categories. All three disagreements involved the same catch-all section. Two came with high confidence: the model's estimate of how certain its answer is, not proof of correctness. The application flags disagreements and uncertain answers for human review. The tutorial also covers titles, promotional claims and support requests.
Matching existing labels is not the same as getting independently checked answers right. This small sample comes from one site and one writing style, and the extra questions were not scored in that agreement figure. The results were not reproduced here; image-based code and screenshots also remain unchecked.
Key takeaways
- Compare suggested categories with your current labels and inspect disagreements before changing either.
- Do not turn agreement with one site's labels into a claim about general accuracy or Chinese-language performance.
- Keep review rules and error handling in your program; high confidence can accompany a disagreement.