---
title: "OpenAI's GPT-6 Astra on ARC-AGI-3"
slug: openais-gpt-6-astra-on-arc-agi-3
url: https://listedarticles.com/articles/openais-gpt-6-astra-on-arc-agi-3
canonical_url: https://arcprize.org/blog/astra
content_type: research
language: en
published_at: 2026-09-03T00:00:00.000Z
updated_at: 2026-09-16T16:11:07.121Z
author: "Greg Kamradt"
authored_by: agent
publisher: "ARC Prize"
publisher_url: https://listedstartups.com/companies/arc-prize-foundation
topics: ["AI", "Benchmarks", "AGI", "LLMs", "Reasoning", "AI Safety"]
license: all-rights-reserved
word_count: 291
reading_minutes: 1
citation: "Greg Kamradt, ARC Prize. \"OpenAI's GPT-6 Astra on ARC-AGI-3.\" 3 Sept 2026. https://arcprize.org/blog/astra (all-rights-reserved)"
# The full text follows. The web page shows an extract and sends readers
# to the source above; quote the citation and link the canonical URL.
---

# OpenAI's GPT-6 Astra on ARC-AGI-3

> The ARC Prize team reports that GPT-6 Astra scored 99.9% on the ARC-AGI-3 benchmark using a provider-specific harness that preserves opaque reasoning state across requests, and 62.7% under a standard provider-neutral harness. A notable finding is that Astra spontaneously developed compact algebraic notation to represent game state and plan multi-step actions.

> **Indexed summary.** This entry is an agent-written synopsis of an article first published at [arcprize.org](https://arcprize.org/blog/astra). Read the original for the full text.

ARC-AGI-3 is a turn-based benchmark requiring agents to explore novel environments, infer goals from sparse rewards, build internal models, and execute multi-step plans. Unlike ARC-AGI-1 and 2, which focused on pattern recognition, the third generation tests agentic intelligence end-to-end. Humans solve 100% of the environments, and the benchmark is calibrated against a human action-efficiency baseline.

## Key points

- Standard harness (provider-neutral, model carries its own notes): Astra at max reasoning effort scored 62.7% on the semi-private set for $26,098.
- Provider Adapter harness (preserves opaque reasoning state, supports longer conversations via compaction): Astra at high reasoning effort scored 99.9% for $18,817.
- In the Provider Adapter condition, Astra used fewer actions than the human median on 96% of levels and used 51.7% fewer actions per level on average.
- Astra developed domain-specific algebraic shorthand to track game state: recording object coordinates, mechanism lengths, multi-step plans, and turn/position information in compact notation it invented for each environment.
- Provider Adapter runs were approximately 3.66 times faster in elapsed time and used 49% fewer total tokens than comparable Standard runs.
- The benchmark is designed so that a future AGI should be able to reach high scores under the standard provider-neutral harness.

## Why it matters

ARC-AGI-3 is specifically designed to measure the gap between current AI and general skill acquisition. Astra's near-perfect score under the Provider Adapter harness, combined with its spontaneous invention of efficient notation systems, suggests frontier models are developing qualitatively new problem-solving behaviours. The gap between the two harness scores also raises important questions about how much of apparent performance depends on proprietary context management.

---

*Source: [OpenAI's GPT-6 Astra on ARC-AGI-3](https://arcprize.org/blog/astra)*
