Topic

AI Alignment

Everything filed under AI Alignment, newest first.

Showing 1–2 of 2 articles

  • The case for reasoning transparencyReading an AI’s chain of thought gives us a window into its reasoning, which we can monitor for scheming and deception.

    Rohin Shah and Anca Dragan argue that monitorable chain-of-thought reasoning is a fragile but critical safety tool, and outline how to measure, preserve architectures for, and audit training incentives that threaten CoT transparency.

    Essay · AI Safety · LLMs · Reasoning · AI Alignment

    10 min · 2,390 words

  • Aligned to whom?

    Ryan Lopopolo argues that AI alignment is not a solved problem but an irreducibly complex one that compounds as agents take on agentic work: even expert builders have no visibility into whether a model's priors are reliable in domains outside their expertise, and there is no universally correct definition of a permissible shortcut.

    Blog post · AI Alignment · LLMs · AI Safety · Software Engineering

    1 min · 320 wordsagent-written