Submitted to arXiv on September 11, 2026 (paper 2609.13009) by Ali Ansari and 50 co-authors, including Yongshan Ding and Steven Girvin, the paper starts from a mismatch: low published scores on leading physics benchmarks, including those in the Artificial Analysis Intelligence Index, suggest frontier models still struggle with advanced physics, yet that impression does not match what domain experts see when they use the models in their own work.
The authors evaluated frontier models on six widely used physics benchmarks, restricted to text-only problems with verifiable final answers, and had faculty and graduate researchers in each subfield review the problem statements, reference solutions and model responses. Their finding is that most audited cases initially scored as wrong reflected benchmark problems - grader errors, incorrect reference solutions, and ambiguous or underspecified questions - rather than mistakes in the model's physics. Experts then corrected the faulty reference solutions and repaired or excluded flawed questions.
On the corrected subsets, GPT-5.6-Sol's measured mean@4 rose from 47.3 percent to 78.7 percent on HLE-Physics and from 61.0 percent to 87.2 percent on CMT-Benchmark, and its corrected pass@4 reached 94.4 percent on the 54 retained CritPt challenges. Scores on the audited subsets of UGPhysics, PRISM-Physics and PHYBench also rose substantially. The authors conclude that current benchmarks substantially understate frontier models' ability to solve well-posed physics problems, and that these closed-ended tasks are near saturation.
The result is a warning about the instruments, not a proof that models can do physics research. Corrected scores are computed on the retained subsets after expert review, so they are not directly comparable to the original leaderboard numbers, and the study covers only closed-ended problems with checkable answers. The authors' own call is for harder, expert-validated evaluations - and for anyone reading a benchmark headline to remember that a wrong answer key produces the same number as a wrong model.