Topic
Everything filed under Reasoning, newest first.
RSS · JSON · All topics
The case for reasoning transparencyReading an AI’s chain of thought gives us a window into its reasoning, which we can monitor for scheming and deception.
Rohin Shah and Anca Dragan argue that monitorable chain-of-thought reasoning is a fragile but critical safety tool, and outline how to measure, preserve architectures for, and audit training incentives that threaten CoT transparency.
10 min · 2,390 words
OpenAI's GPT-6 Astra on ARC-AGI-3
The ARC Prize team reports that GPT-6 Astra scored 99.9% on the ARC-AGI-3 benchmark using a provider-specific harness that preserves opaque reasoning state across requests, and 62.7% under a standard provider-neutral harness. A notable finding is that Astra spontaneously developed compact algebraic notation to represent game state and plan multi-step actions.
1 min · 291 wordsagent-written