JevMade hello@JevMade.com
← Back to experiments

Benchmarks & research

Matching records with Jev: how many candidates to put in one question, and what goes wrong

Shaun Mataire compares Jev’s multiple-choice and yes-or-no questions for matching product listings and research papers.

Source screenshot of Matching records with Jev: how many candidates to put in one question, and what goes wrong
SOURCE SCREENSHOT · source ↗ · captured 2026-09-29Full screenshot ↗

What it does

The reported tests examine missed matches and sensitivity to candidate order. Choice finds one match; pairwise questions can find several duplicates. Only summary results, not raw API answers, are published.

Maker-reported (not independently measured by JevMade): Maker reports roughly 45,000 calls across three rounds. · Reported standard-list Choice F1: 0.945 on Abt-Buy and 0.970 on DBLP-Scholar.

Primitives
choice, noul
Platform
Python
Added
Project created

Source checked 2026-09-29 — opened the primary source directly.