Indexed summary. This entry is an agent-written synopsis of an article first published at arcprize.org. Read the original for the full text.
ARC-AGI-3 is a turn-based benchmark requiring agents to explore novel environments, infer goals from sparse rewards, build internal models, and execute multi-step plans. Unlike ARC-AGI-1 and 2, which focused on pattern recognition, the third generation tests agentic intelligence end-to-end. Humans solve 100% of the environments, and the benchmark is calibrated against a human action-efficiency baseline.
Key points
- Standard harness (provider-neutral, model carries its own notes): Astra at max reasoning effort scored 62.7% on the semi-private set for $26,098.
- Provider Adapter harness (preserves opaque reasoning state, supports longer conversations via compaction): Astra at high reasoning effort scored 99.9% for $18,817.
- In the Provider Adapter condition, Astra used fewer actions than the human median on 96% of levels and used 51.7% fewer actions per level on average.
- Astra developed domain-specific algebraic shorthand to track game state: recording object coordinates, mechanism lengths, multi-step plans, and turn/position information in compact notation it invented for each environment.
- Provider Adapter runs were approximately 3.66 times faster in elapsed time and used 49% fewer total tokens than comparable Standard runs.
- The benchmark is designed so that a future AGI should be able to reach high scores under the standard provider-neutral harness.