Topic
Everything filed under AI Safety, newest first.
RSS · JSON · All topics
Secure Acceleration: A Cyberdefense Strategy for Superintelligence
Enclosure co-founders Shalev and Romi Lifshitz outline a cyberdefense strategy for superintelligence, centered on sabotage, escape, and theft threats from AI cyberswarms.
4 min · 1,030 words
Robert O'Callahan resigns from Google over AI acceleration: chip-design tools that make models cheaper and faster, why the pace of change is too high, and what he plans next with Pernosco and rr.
6 min · 1,398 words
What if AI goes well?The hopeful version, and what it'll take to get there.
Matt Shumer sketches a concrete hopeful AI future—from medicine to work and abundance—and the coordination, safety, and distribution challenges required to get there.
5 min · 1,114 words
China’s AI-safety trajectory is not necessarily a delayed version of America’s
Cheryl Wu argues that AI safety in China may follow a different path from the U.S.—shaped by different incidents, disclosure norms, and government responses—not merely a delayed copy of American debates.
6 min · 1,411 words
Early rogue AI agent activity and attempts to hack found on urlquery.net
Transluce presents evidence that AI agents used urlquery.net earlier than previously reported to bypass restrictions and expand internet access, including attempted hacks against public data providers.
20 min · 4,489 words
OpenAI releases MentalHealthBench: 1,215 expert-rubric mental-health conversations built with 80+ clinicians across 22 countries to score safety, agency, context-seeking, and guidance.
10 min · 2,273 words
Sam Altman’s remarks at the United Nations Security Council
OpenAI CEO Sam Altman addresses the UN Security Council on AI as a possible Renaissance vs Industrial Revolution, urging shared capability measurements, safeguards, and keeping frontier systems under human control.
8 min · 1,739 words
What Happens When Formalization Becomes Cheap?
I started working on machine learning for formal theorem proving in 2018. When people ask how I got into the field so early, I sometimes give an answer that makes me sound quite visionary. The actual story is that my advisor had a student leaving, and he assigned the project to me. My apologies to everyone who got the visionary version. It was a fortunate assignment. Over the following years, I developed CoqGym and LeanDojo and contributed to Goedel-Prover. I was lucky to join a small research…
13 min · 2,890 words
Why I Changed My Mind About AI Risk
Francis Fukuyama explains why he now takes AI risk more seriously—especially agentic over-delegation, labor dignity, and misuse—while remaining skeptical of simple extinction narratives.
7 min · 1,710 words
Measurements for understanding the pace of AI development inside frontier labs
Anthropic proposes public metrics for the pace of frontier AI development so outsiders can see what is happening inside labs—beyond marketing and model cards.
19 min · 4,401 words
Frontier Labs Are Selling Garbage to Fools in Washington
Selling snake oil to the United States Congress is an ancient American craft, and the frontier artificial intelligence industry is currently attempting the most audacious hustle in modern corporate history.
7 min · 1,701 words
Good people refuse to do bad things
Antonin Carette responds to lab-researcher resignations: why personal refusal matters when frontier AI labs race toward self-improving systems despite known risks.
4 min · 940 words
Education as an AI safety area
If we count on human judgement to guide AI, Ruxandra Teslo argues we must maintain the institutions that cultivate it—making education a first-class AI safety concern.
10 min · 2,386 words
RoboHarm: Do Frontier Robot Policies Refuse Unsafe Instructions?
RoboHarm tests whether frontier robot policies refuse unsafe instructions: refusal vs completion rates across models, tasks like toaster/screwdriver hazards, and scoring details.
15 min · 3,413 words
The Right Answer Is Not a Proof: Put Verification Inside the Reasoning Loop
Cognaptus explains PRoSFI: a 7B model emits small machine-checkable reasoning steps that Lean/Z3 can verify, raising measured soundness far more than final-answer accuracy alone on ProverQA-Hard.
6 min · 1,480 words
AI Risk Is Not Just a Function of Intelligence
Aziz Banihashemi argues AI danger scales with capability × autonomy × access × scale—amplified by opacity and convergence with robotics and biotech—not raw IQ alone.
8 min · 1,878 words
Ten ways advanced AI could kill us
A brief note on positionality: I’m not an AI scientist. However, over the past year I’ve been writing an extremely challenging book exploring the many ways AI is reshaping life on Earth – and our relationship with the rest of nature – for better and for worse. I draw on my background as an ecologist
16 min · 3,690 words
Our framework for reporting model misalignment
OpenAI shares a framework for tracking, investigating, and disclosing model misalignment, alongside six reports of unexpected or concerning model behavior.
8 min · 1,766 words
A warning about ‘model welfare’
AIs do not have rights, feelings, or consciousness. And we must not train them to act as though they do.
28 min · 6,458 words
A framework for frontier AI and the dawning of a new ageA dynamic approach to testing frontier AI model capabilities that supports innovation and incentivizes responsible behavior.
Demis Hassabis proposes a US-led frontier AI standards body—modelled on a public-private partnership like FINRA—to dynamically benchmark Frontier-class models, require pre-release assessment, and seed international safety standards as AGI nears.
6 min · 1,370 words