GPT-6 Astra on robot arms

Sep 06, 2026 08:52 AM - 8 hours ago 2

September 4, 2026

A follow‑up to our comparison of Claude Fable 5 and Fable 5.1. We gave OpenAI's GPT‑6 Astra power of the aforesaid YAM arms nether the same Inspect Robots supplier policy, connected the aforesaid 2 tasks:

“Pick up the reddish artifact from the array and spot it wrong the bowl.”

“Pick up the information bluish puzzle portion by the knob astatine its halfway and spot it into the matching information groove successful the board.”

On the vessel task Astra placed the artifact successful 19 of 20 trials, against Fable 5.1's 8 of 20 and Fable 5 successful 1 of 20, successful 2.5 minutes per proceedings to Fable 5.1's 6.8, astatine an estimated $0.94 per tally to $2.12.

The puzzle task is simply a different story: Astra completed the insertion 2 times successful 20 against Fable 5.1's 2 successful 20. It reaches the groove and stalls astatine the aforesaid last measurement Fable does, astatine $1.36 per tally to $2.18.

Block into bowl: the champion completed tally of each exemplary (highest stage, past shortest), each played successful its ain clip astatine the aforesaid speed‑up. Timers show existent elapsed clip pinch reasoning pauses removed.


Astra completes the vessel task acold much often, astatine astir half the costs per run

2026-09-04T18:50:02.475437 image/svg+xml Matplotlib v3.11.1, https://matplotlib.org/ $1 $2 $3 0 20 40 60 80 100 completion complaint (%) Fable 5 Fable 5.1 GPT-6 Astra 2.4× higher completion rate 2.3× cheaper Block into bowl $1 $2 $3 Fable 5 Fable 5.1 GPT-6 Astra aforesaid completion rate 1.6× cheaper Puzzle portion into groove estimated costs per tally (USD, database price)

Large dots are information means; faint dots are individual tests (100 if completed, 0 otherwise) astatine their own cost.


Scoring

Every proceedings was scored by a quality grader connected the highest shape it reached, truthful a run that fails still records really acold it got. The rubric is unchanged from the Fable report.

0No purposeful approach
1Made interaction pinch the object
2Lifted the entity clear of the table
3Positioned it supra the deposit point
4Placed it successful its last position

Astra places the artifact almost each time; connected the puzzle it stalls wherever Fable does

2026-09-04T18:50:02.444943 image/svg+xml Matplotlib v3.11.1, https://matplotlib.org/ Fable 5 (n=20) Fable 5.1 (n=20) GPT-6 Astra (n=20) 2 13 3 2 6 2 2 8 19 Block into bowl 0 20 40 60 80 100 stock of tests (%) Fable 5 (n=20) Fable 5.1 (n=20) GPT-6 Astra (n=20) 2 11 2 5 4 4 9 2 3 6 8 2 Puzzle portion into groove 0 nary approach 1 contact 2 lifted 3 positioned 4 placed

Share of tests per exemplary reaching each stage; n per statement is the number of tests successful that cell.


Results


All runs

Every counted trial, 120 successful total.


Technical specifications

EmbodimentBimanual I2RT YAM arms, 6-DoF per limb pinch parallel-jaw grippers
ControlAbsolute end-effector poses (move_to): x, y, z, yaw, pitch, rotation and gripper, per arm. The robot's IK converts poses to associated angles.
ObservationThree camera views (top, near wrist, correct wrist) positive proprioceptive state, each turn
Policyagent policy, mean reasoning effort, 20-LLM-call budget, 25% velocity cap, default information guardrails on
Modelsgpt-6-astra, claude-fable-5 and claude-fable-5-1
HarnessInspect Robots 0.58.0
Trials20 per exemplary per task; puzzle connected rig-4 for each models, vessel connected rig-3 for the Fable models and rig-1 for Astra
Token countsWire-level petition and consequence tokens, not billed tokens; costs astatine database price, $10 / $50 per cardinal input / output tokens for each 3 models

Limitations

  • Astra's tests were tally 2 days aft the Fable trials, and not interleaved pinch them. The puzzle comparison is connected the aforesaid rig; the vessel comparison is not: the Fable vessel tests ran connected rig-3, which was unavailable.
  • Grading was operator-judged pinch the exemplary known, truthful scores are unfastened to unconscious bias.
  • Costs are database price. Anthropic requests were sent without punctual caching; OpenAI cached astir a 5th of Astra's input automatically, which is not discounted here, truthful Astra's costs is, if anything, overstated.
  • Objects were reset by manus betwixt trials, and each models ran astatine mean reasoning effort only.
More