[Submitted connected 11 Sep 2026]
Authors:Ali Ansari, Haoran Sun, Andy Zeyi Liu, Mark Jabbour, Yongshan Ding, Steven Girvin, Yu He, Sohrab Ismail-Beigi, Aleksander Kubica, Owen D. Miller, Corey O'Hern, Vidvuds Ozolins, David Poland, A. Douglas Stone, Frank C. van den Bosch, Logan Wright, Navid Akbari, Santanu Antu, Kangle Cai, Andrew Calabrese-Day, Mateo Cárdenes Wuttig, Meng Cheng, Barry T. Chiang, Ali Ghorashi, Shouzhen Gu, Haoyang Huang, Zhibo Kang, Lukas Kienesberger, Hantian Liu, Charles Lomba, Zhongling Lu, Wenchao Ma, Rohin E. McIntosh, Evan McKinney, Ivan Rojkov, Xulei Sun, Yarone Meir Tokayer, Naveen Balaji Umasankar, Mira Varma, Leda Wang, Qimin Wang, Tyler Wang, Haoyu Wei, Jinming Yang, Jinchen Zhao, Sherlock Tingrui Zhao, Qinyuan Zheng, Jay S. Zou, Lucas Baker, Arman Cohan, John Sous
View PDF HTML (experimental)
Abstract:Low reported scores connected starring physics benchmarks, including those featured successful the Artificial Analysis Intelligence Index (2026), propose that frontier connection models still struggle pinch precocious physics, a demanding trial of their technological reasoning and quantitative problem-solving abilities. Yet this belief does not ever align pinch domain experts' experiences utilizing these models successful their work. We revisit these reported findings by evaluating frontier models connected six wide utilized physics benchmarks and auditing them pinch experts, focusing connected text-only problems pinch verifiable last answers. For each subfield of physics, module and postgraduate researchers pinch applicable expertise cautiously reappraisal problem statements, reference solutions, and exemplary responses to separate genuine exemplary errors from grader errors, incorrect reference solutions, and ambiguous aliases underspecified questions. Most audited cases initially evaluated arsenic incorrect bespeak these benchmarking issues alternatively than errors successful the models' physics reasoning. We past inquire experts to reside these benchmarking issues by correcting erroneous reference solutions and repairing aliases excluding flawed questions. We find that GPT-5.6-Sol's measured mean@4 rises from 47.3% to 78.7% connected HLE-Physics and from 61.0% to 87.2% connected CMT-Benchmark, while its corrected pass@4 reaches 94.4% connected the 54 retained CritPt challenges. Corrected scores are computed connected the retained information subsets pursuing master review. Scores connected the audited subsets of UGPhysics, PRISM-Physics, and PHYBench besides emergence substantially aft correction. These findings propose that existent benchmarks substantially understate frontier models' expertise to lick well-posed physics problems. Near-saturation connected these closed-ended tasks highlights the request for much demanding, expert-validated evaluations.Submission history
From: Ali Ansari [view email]
[v1] Fri, 11 Sep 2026 16:06:50 UTC (178 KB)
English (US) ·
Indonesian (ID) ·