Artificial Analysis Intelligence Index v4.2

Sep 05, 2026 07:04 AM - 1 hour ago 2

We are accelerating elements of our upcoming v5 merchandise pinch interim updates to support gait pinch the frontier. Index v4.2 has much analyzable and realistic tasks, and much backstage trial sets to forestall gaming

Intelligence Index v4.2 changelog:

+ AA-Briefcase, our agentic knowledge activity information pinch a backstage trial set

+ Surge’s GDP.pdf, agelong discourse archive reasoning crossed 4,592 PDF pages

- GPQA Diamond, an exceptional technological reasoning information that has now been saturated

… positive greater weighting connected held-out trial sets to forestall gaming, and grading infrastructure upgrades to summation robustness

This update brings the Index person to real-world usage cases pinch much challenging, analyzable and realistic tasks and backstage trial sets to forestall gaming. We person been readying and building elements of Index v5 for months - it’s been 8 months since we launched Index v4 successful January.

We person deliberately held backmost updates to support the Index unchangeable done caller awesome exemplary launches. However, pinch the frontier moving truthful quickly successful the past weeks, we consciousness it is important to present an contiguous interim update to guarantee our Index remains arsenic applicable and useful arsenic ever to users.

Beyond this interim update, our squad is difficult astatine activity connected v5 of the Index. We are readying much incremental releases successful the adjacent future. Stay tuned!

Intelligence Index v4.2 changes successful detail:

➤ Adding AA-Briefcase: Our in-house information pinch a backstage held-out trial set, AA-Briefcase tests models connected realistic agentic knowledge activity tasks successful analyzable projects built by manufacture experts. Models are evaluated connected multi-week knowledge activity projects, each pinch galore linked tasks and thousands of input root files. AA-Briefcase combines rubric and pairwise grading to measure verifiable task success, analytical quality, and position quality, giving a holistic position of wide agentic capacity successful knowledge work.

➤ Adding GDP.pdf: Created by Surge AI, GDP.pdf evaluates single-turn master archive reasoning crossed 100 PDFs and 10 domains. Models must synthesize grounds distributed crossed 4,592 pages, including text, tables, charts, footnotes, and exclusions. Responses are graded against 1,275 expert-authored atomic criteria; the header All-pass Rate credits a task only erstwhile each criterion is satisfied.

➤ Weighting to measurement real-world usage and forestall gaming: 40% of our Index weighting is now private, held-out trial sets - double the fig from v4.1. Held-out information includes AA-Briefcase, AA-Omniscience, and solutions for CritPt. This reduces the expertise for labs to crippled evaluations. The held-out percent will summation further successful Index v5.

➤ Improving our grading infrastructure: In AA-LCR v1.1, we person added a grading strategy punctual and corrected errors and ambiguities successful reply keys, improving scoring accuracy. For GDPval-AA v2 and AA-Briefcase, we person improved our sampling and re-anchored the Elo scale, making ratings much unchangeable arsenic caller models are added. For SciCode we person improved robustness of grading sandboxes to guarantee slow but correct codification does not count arsenic a failure.

Key results:

➤ Anthropic and OpenAI lead the Index: Anthropic’s Claude Fable 5.1 leads the Index, followed by OpenAI’s GPT-6 Astra, which shows a 4pt summation complete GPT-5.6 Sol. Meta is the third-ranked laboratory connected the leaderboard, followed by SpaceXAI, Moonshot/Kimi, Z.AI, and Google.

➤ Cost per Task Pareto frontier shared by 4 labs: Anthropic, OpenAI, Meta and Z.AI inhabit the updated Cost per Task frontier.

➤ GPT-6 Astra dominates the output token frontier: GPT-6 Astra is much token businesslike than almost each different exemplary adjacent the intelligence frontier, pinch Claude Fable 5.1, Grok 4.5 and Gemini 3.5 Flash-Lite astatine either extremity of the curve (excludes models beneath 25 connected the Index)

GPT-6 Astra is much token businesslike than almost each different exemplary adjacent the intelligence frontier, pinch Claude Fable 5.1, Grok 4.5 and Gemini 3.5 Flash-Lite astatine either extremity of the curve (excludes models beneath 25 connected the Index)

Anthropic’s Claude Fable 5.1 and Opus 5 lead AA-Briefcase, followed by GPT-6 Astra and Muse Spark 1.3. GPT-6 Astra shows a important summation supra GPT-5.6 Sol of ~85 Elo points

AA-Briefcase is our frontier in-house information pinch a backstage held-out trial set. The information tests models connected realistic agentic knowledge activity tasks successful analyzable projects built by manufacture experts. Models are evaluated connected multi-week knowledge activity projects, each pinch galore linked tasks and thousands of input root files. AA-Briefcase combines rubric and pairwise grading to measure verifiable task success, analytical quality, and position quality, giving a holistic position of wide agentic capacity successful knowledge work.

OpenAI leads GDP.pdf pinch GPT-6 Astra astatine 33.2% and GPT-5.6 Sol astatine 28.2%, followed by Claude Fable 5.1 astatine 26.2%

Created by Surge AI, GDP.pdf evaluates single-turn master archive reasoning crossed 100 PDFs and 10 domains. Models must synthesize grounds distributed crossed 4,592 pages, including text, tables, charts, footnotes, and exclusions. Responses are graded against 1,275 expert-authored atomic criteria; the header All-pass Rate credits a task only erstwhile each criterion is satisfied

Full per-model breakdowns below:

Read much astir Artificial Analysis Intelligence Index v4.2 astatine https://artificialanalysis.ai/methodology/intelligence-benchmarking

More