September 4, 2026
A follow‑up to our comparison of Claude Fable 5 and Fable 5.1. We gave OpenAI's GPT‑6 Astra power of the aforesaid YAM arms nether the same Inspect Robots supplier policy, connected the aforesaid 2 tasks:
“Pick up the reddish artifact from the array and spot it wrong the bowl.”
“Pick up the information bluish puzzle portion by the knob astatine its halfway and spot it into the matching information groove successful the board.”
On the vessel task Astra placed the artifact successful 19 of 20 trials, against Fable 5.1's 8 of 20 and Fable 5 successful 1 of 20, successful 2.5 minutes per proceedings to Fable 5.1's 6.8, astatine an estimated $0.94 per tally to $2.12.
The puzzle task is simply a different story: Astra completed the insertion 2 times successful 20 against Fable 5.1's 2 successful 20. It reaches the groove and stalls astatine the aforesaid last measurement Fable does, astatine $1.36 per tally to $2.18.
Block into bowl: the champion completed tally of each exemplary (highest stage, past shortest), each played successful its ain clip astatine the aforesaid speed‑up. Timers show existent elapsed clip pinch reasoning pauses removed.
Astra completes the vessel task acold much often, astatine astir half the costs per run
2026-09-04T18:50:02.475437 image/svg+xml Matplotlib v3.11.1, https://matplotlib.org/ $1 $2 $3 0 20 40 60 80 100 completion complaint (%) Fable 5 Fable 5.1 GPT-6 Astra 2.4× higher completion rate 2.3× cheaper Block into bowl $1 $2 $3 Fable 5 Fable 5.1 GPT-6 Astra aforesaid completion rate 1.6× cheaper Puzzle portion into groove estimated costs per tally (USD, database price)
Large dots are information means; faint dots are individual tests (100 if completed, 0 otherwise) astatine their own cost.
Scoring
Every proceedings was scored by a quality grader connected the highest shape it reached, truthful a run that fails still records really acold it got. The rubric is unchanged from the Fable report.
| 0 | No purposeful approach |
| 1 | Made interaction pinch the object |
| 2 | Lifted the entity clear of the table |
| 3 | Positioned it supra the deposit point |
| 4 | Placed it successful its last position |
Astra places the artifact almost each time; connected the puzzle it stalls wherever Fable does
2026-09-04T18:50:02.444943 image/svg+xml Matplotlib v3.11.1, https://matplotlib.org/ Fable 5 (n=20) Fable 5.1 (n=20) GPT-6 Astra (n=20) 2 13 3 2 6 2 2 8 19 Block into bowl 0 20 40 60 80 100 stock of tests (%) Fable 5 (n=20) Fable 5.1 (n=20) GPT-6 Astra (n=20) 2 11 2 5 4 4 9 2 3 6 8 2 Puzzle portion into groove 0 nary approach 1 contact 2 lifted 3 positioned 4 placed
Share of tests per exemplary reaching each stage; n per statement is the number of tests successful that cell.
Results
All runs
Every counted trial, 120 successful total.
Technical specifications
| Embodiment | Bimanual I2RT YAM arms, 6-DoF per limb pinch parallel-jaw grippers |
| Control | Absolute end-effector poses (move_to): x, y, z, yaw, pitch, rotation and gripper, per arm. The robot's IK converts poses to associated angles. |
| Observation | Three camera views (top, near wrist, correct wrist) positive proprioceptive state, each turn |
| Policy | agent policy, mean reasoning effort, 20-LLM-call budget, 25% velocity cap, default information guardrails on |
| Models | gpt-6-astra, claude-fable-5 and claude-fable-5-1 |
| Harness | Inspect Robots 0.58.0 |
| Trials | 20 per exemplary per task; puzzle connected rig-4 for each models, vessel connected rig-3 for the Fable models and rig-1 for Astra |
| Token counts | Wire-level petition and consequence tokens, not billed tokens; costs astatine database price, $10 / $50 per cardinal input / output tokens for each 3 models |
Limitations
- Astra's tests were tally 2 days aft the Fable trials, and not interleaved pinch them. The puzzle comparison is connected the aforesaid rig; the vessel comparison is not: the Fable vessel tests ran connected rig-3, which was unavailable.
- Grading was operator-judged pinch the exemplary known, truthful scores are unfastened to unconscious bias.
- Costs are database price. Anthropic requests were sent without punctual caching; OpenAI cached astir a 5th of Astra's input automatically, which is not discounted here, truthful Astra's costs is, if anything, overstated.
- Objects were reset by manus betwixt trials, and each models ran astatine mean reasoning effort only.
English (US) ·
Indonesian (ID) ·