JevMade hello@JevMade.com
← Back to guides

JevMade field notes / Multimodal benchmark report

JevAny

JevAny evaluates a local typed-decision model across language, image, video, and agent tasks with overlap checks, bootstrap intervals, checkpoint-equivalence records, and published negative results.

Original by weitianxinEvaluationGitHub repositorySource reviewed

Before you dive in

What you’ll find in the original

  1. Remove media, question-option, and source-ID overlap with training, calibration, and development data before reporting multimodal transfer.
  2. Publish paired bootstrap intervals and task-level failures instead of relying on one overall gain.
  3. Keep a development improvement out of the default release when independent calibration or transfer does not confirm it.
Worth knowing

The author reports low action-antonym accuracy and poor video calibration. Hosted Jev is a separate comparison system, not a weight-matched ablation, and benchmark media is not redistributed under upstream terms.