SpaceXAI's Grok 4.6 scores 61 connected the Artificial Analysis Intelligence Index, joining the frontier successful statement pinch GPT-5.6 Sol, pinch standout agentic capacity astatine little cost
Grok 4.6 gains 5 points complete Grok 4.5 connected the Intelligence Index conscionable complete 1 period aft its release, aliases +23 points compared to Grok 4.3. This brings SpaceXAI backmost to the intelligence frontier alongside OpenAI, down only Anthropic.
Key takeaways: ➤ Grok 4.6 joins the frontier of the Artificial Analysis Intelligence Index: It scores 61, successful statement pinch GPT-5.6 Sol (max), down Claude Opus 5 (max, 63) and Claude Fable 5 (max pinch fallback, 62), and conscionable up of Kimi K3
➤ Strong agentic performance: Grok 4.6 achieves a GDPval-AA v2 Elo of 1753, down only Claude Opus 5 and pinch overlapping assurance intervals pinch Claude Fable 5 and Qwen3.8 Max. It scores 50.7% connected 𝜏³-Banking, among the apical 2 scores alongside Qwen3.8 Max (51.3%), and 88.4% connected Terminal-Bench v2.1, successful statement pinch the starring models
➤ Frontier-level intelligence astatine little cost: Headline pricing is unchanged from Grok 4.5 astatine $2/$6 per 1M input/output tokens, 60%+ beneath Claude Opus 5 ($5/$25) and GPT-5.6 Sol ($5/$30). It costs $0.84 per task, the aforesaid arsenic Kimi K3 pinch somewhat higher intelligence, placing it connected the Intelligence vs. Cost per Task Pareto frontier
➤ Grok 4.6 sits astatine Fable 5-tier connected AA-Briefcase, our backstage benchmark of long-horizon agentic knowledge activity tasks, pinch an Elo of 1577 - down the Claude Opus 5 family. It is notably turn-efficient, completing tasks successful ~53 turns and ~0.5B input tokens connected mean vs. ~103 turns and ~2.0B input tokens for Claude Opus 5 (max)
Other exemplary details: ➤ Context model of 500k tokens (unchanged from Grok 4.5)
➤ Pricing of $2/$6 per 1M tokens of input/output; cache hits discounted to $0.5 per 1M tokens, an summation complete Grok 4.5’s $0.3 per 1M tokens for cache hits 

Agentic performance
Grok 4.6's strongest results are connected agentic activity alternatively than fixed reasoning. On GDPval-AA v2, our starring measurement of real-world agentic knowledge work, it scores an Elo of 1753 - down only Claude Opus 5, and statistically indistinguishable from Claude Fable 5 and Qwen3.8 Max fixed overlapping assurance intervals.
The shape holds crossed task types. 𝜏³-Banking (50.7%) tests multi-turn customer work pinch instrumentality usage and places Grok 4.6 successful the apical two, while Terminal-Bench v2.1 (88.4%) puts it level pinch the leaders connected terminal-based package tasks. Few models are simultaneously competitory crossed knowledge work, customer work and terminal use; mixed pinch its pricing, this places Grok 4.6 connected the costs vs. capacity Pareto frontier for each agentic information successful the Intelligence Index.

Cost
Holding header pricing level crossed a procreation is different astatine the frontier, wherever intelligence gains person typically been accompanied by value increases. Grok 4.6 delivers a 5-point Intelligence Index summation astatine unchanged $2/$6 pricing, and our measured costs per task of $0.84 reflects some that pricing and reasonable token efficiency.
The comparison that matters for buyers is against the models scoring wrong 2 points of it: Claude Opus 5 astatine $5/$25 and GPT-5.6 Sol astatine $5/$30. Grok 4.6 offers efficaciously the aforesaid Intelligence Index people arsenic GPT-5.6 Sol astatine a fraction of the output token price, which is the magnitude that dominates costs successful reasoning-heavy workloads.

Long-horizon knowledge work
Grok 4.6 debuts connected AA-Briefcase, our backstage benchmark of long-horizon agentic knowledge activity tasks, pinch an Elo of 1577. This places it astatine Fable 5-tier, down the Claude Opus 5 family, pinch consistently beardown capacity crossed rubric grading, position value and analytical value alternatively than spot successful 1 magnitude offsetting weakness successful another.
The ratio floor plan is arsenic notable arsenic the score. Grok 4.6 resolves tasks successful ~53 turns and ~0.5B input tokens connected average, against ~103 turns and ~2.0B input tokens for Claude Opus 5 (max). Long-horizon agentic activity accumulates discourse rapidly, truthful a exemplary that reaches a comparable reply successful half the turns and a 4th of the input tokens has a costs advantage good beyond its per-token pricing.

Full results
Full breakdown of the individual evaluations successful the Artificial Analysis Intelligence Index:

See Artificial Analysis for further specifications and benchmarks of Grok 4.6: https://artificialanalysis.ai/models/grok-4-6
English (US) ·
Indonesian (ID) ·