How good are frontier models at physics?

Hacker News by 2 min read 509x views
How good are frontier models at physics?

Share Post

[Submitted connected 11 Sep 2026]

View PDF HTML (experimental)

Abstract:Low reported scores connected starring physics benchmarks, including those featured successful the Artificial Analysis Intelligence Index (2026), propose that frontier connection models still struggle pinch precocious physics, a demanding trial of their technological reasoning and quantitative problem-solving abilities. Yet this belief does not ever align pinch domain experts' experiences utilizing these models successful their work. We revisit these reported findings by evaluating frontier models connected six wide utilized physics benchmarks and auditing them pinch experts, focusing connected text-only problems pinch verifiable last answers. For each subfield of physics, module and postgraduate researchers pinch applicable expertise cautiously reappraisal problem statements, reference solutions, and exemplary responses to separate genuine exemplary errors from grader errors, incorrect reference solutions, and ambiguous aliases underspecified questions. Most audited cases initially evaluated arsenic incorrect bespeak these benchmarking issues alternatively than errors successful the models' physics reasoning. We past inquire experts to reside these benchmarking issues by correcting erroneous reference solutions and repairing aliases excluding flawed questions. We find that GPT-5.6-Sol's measured mean@4 rises from 47.3% to 78.7% connected HLE-Physics and from 61.0% to 87.2% connected CMT-Benchmark, while its corrected pass@4 reaches 94.4% connected the 54 retained CritPt challenges. Corrected scores are computed connected the retained information subsets pursuing master review. Scores connected the audited subsets of UGPhysics, PRISM-Physics, and PHYBench besides emergence substantially aft correction. These findings propose that existent benchmarks substantially understate frontier models' expertise to lick well-posed physics problems. Near-saturation connected these closed-ended tasks highlights the request for much demanding, expert-validated evaluations.

Submission history

From: Ali Ansari [view email]
[v1] Fri, 11 Sep 2026 16:06:50 UTC (178 KB)

Other Article Hacker News
↑
Close Right Ads
Close Left Ads