ChatGPT is most accurate — and the ranking survives every judge choice
Panel accuracy is 2.40 for ChatGPT versus 2.18 for Claude and 1.86 for Gemini (0–3 scale). The order is identical under all eight alternative judge choices — each judge alone, every leave-one-out panel, with and without the adjudicator. Normalised per named business, ChatGPT and Claude are on par on hallucination (5.6 % vs 4.7 %) — Gemini is worst (8.4 %).