Indexed summary. This entry is an agent-written synopsis of an article first published at openai.robocurve.org. Read the original for the full text.
This is a follow-up to Robocurve's earlier comparison of Claude Fable 5 and Fable 5.1 on two object-manipulation tasks using bimanual I2RT YAM arms controlled by the Inspect Robots harness. The evaluation runs GPT-6 Astra under identical policy conditions: absolute end-effector poses, a 20-LLM-call budget, three camera views plus proprioception, and a human grader scoring each trial on a five-stage rubric.
Key points
- Block-into-bowl: Astra completed 19 of 20 trials (95%) at an estimated $0.94 per run and 2.5 minutes per trial, versus Fable 5.1's 40% completion at $2.12 per run.
- Puzzle-into-groove: Astra and Fable 5.1 both completed 2 of 20 trials (10%); Astra stalls at the same final insertion step.
- Astra used substantially fewer output tokens per run (2.1k for the bowl task vs Fable 5.1's 12.9k), accounting for the cost difference.
- The bowl task trials for Astra ran on a different physical rig than the Fable trials, introducing a potential confound.
- Grading was performed by a human evaluator who knew the model identity, introducing possible bias.
- Costs are at list price; OpenAI's automatic caching likely overstates Astra's actual cost.
Why it matters
LLM-controlled robotic manipulation is maturing quickly. The stark gap between Astra's bowl performance and its puzzle performance illustrates that current frontier models excel at tasks requiring gross dexterity and spatial reasoning but still struggle with the precise final insertion steps that many real-world manipulation tasks require.