The authors’ small-graph panel reports perfect adjacency recognition but weaker structural skills. Heuristic suggestions improve Jev’s tour results, yet a fixed rule does better; the paper and public code use different retrieval pools, and raw Jev runs are not included.
How you can use it
If your app chooses routes or works with connected records, borrow the benchmark’s comparison method. Ask easy questions about direct links, then harder questions about the wider network. Keep correct answers separate from whether the AI returned an allowed answer.
Compare the AI with simple rules before relying on it to build a route. Test both unrestricted choices and choices suggested by those rules. The paper’s results are reported comparisons, not a promise for your app; some model settings use different amounts of reasoning.