JevMade Sign in
← Back to videos

JevMade field notes / Video guide

Cloudflare Just Challenged Jev, So I Tested Their New Model (Clef)

Better Stack compares Jev and Clef on 200 channel comments using Claude as referee, then tests whether Clef's judgments of 225 thumbnails predict video performance.

Original by Better StackEvaluationIntermediate7 min 59 sec Published

Before you press play

What you’ll find in the video

  1. Read the comment scores as agreement with Claude, not accuracy against independently checked human labels.
  2. Inspect each question separately: Clef wins three of five reported comparisons but loses the combined score, with tone a particular weakness.
  3. Separate describing an image from predicting audience response. Only Clef receives images in the thumbnail test; the result does not establish image input for Jev.
Worth knowing

Creator-run tests, not independently reproduced benchmarks. Speed includes the presenter's network travel, and price claims depend on recording-time inputs and hosting terms. The thumbnail comparison does not control for video topic or content. A confident answer is not measured correctness. These notes use selected provider-generated media analysis, not fetched captions or a whole-video watch.

Cloudflare Just Challenged Jev, So I Tested Their New Model (Clef)

Keep this for later

Sign in to bookmark experiments, guides and videos, and keep notes only you can see.

Continue to sign in

We’ll bring you back to this listing.