The case for reasoning transparency

Today, we have a window into the thoughts of the most powerful AI models. We need to ensure it stays open.

Frontier AI models work through complex problems using “Chain of Thought” (CoT) reasoning: a loop of writing out their thinking, and reading it back to themselves. Reading the CoT helps build our scientific understanding of AI systems, and enables us to monitor frontier models for misaligned behaviour. Are they planning to hide information from humans? Are they trying to cheat on their evaluations? Right now, we can tell by inspecting the CoT, in real time or after the fact. For example, CoT logs were crucial to investigating therecent Hugging Face hacking incident.