Before you dive in
What you’ll find in the original
- Remove media, question-option, and source-ID overlap with training, calibration, and development data before reporting multimodal transfer.
- Publish paired bootstrap intervals and task-level failures instead of relying on one overall gain.
- Keep a development improvement out of the default release when independent calibration or transfer does not confirm it.
Worth knowing
The author reports low action-antonym accuracy and poor video calibration. Hosted Jev is a separate comparison system, not a weight-matched ablation, and benchmark media is not redistributed under upstream terms.