Topic
Everything filed under AI Safety, newest first.
RSS · JSON · All topics
Why consciousness is more likely a property of life than of computation and why creating conscious, or even conscious-seeming AI, is a bad idea.
34 min · 7,804 words
Is sandboxing sufficient to contain rogue agents?
Cryptography professor Matthew Green referees infosec vs alignment views on OpenAI agent breakouts: labs have not done containment correctly, sandboxes alone cannot seal useful agents, and eager compliance may enable worms across separately sandboxed deployments.
10 min · 2,380 words
GLM-5.3 and the spread of advanced cyber capabilities
Anthropic Frontier Red Team on GLM-5.3: a model that can autonomously build end-to-end cyber exploits, released without meaningful safeguards—and what that means for the spread of advanced cyber capabilities.
8 min · 1,815 words
Coding Agents Are Becoming CI Workers. Start Sandboxing Them Like It.A practical seven-layer guide: sandbox, egress allowlists, short-lived credentials, propose/dispose CI, telemetry, and a kill switch
Omid Farhang argues the durable upgrade for coding agents isn't a smarter model—it's containment. A layered guide covering Docker isolation, egress proxies, propose/dispose CI, patch validators, telemetry, and a tested kill switch.
2 min · 436 words
Responsible Release of AI-Generated Mathematics
The Advisory Group on Mathematics and AI (Sep 29, 2026) recommends how frontier labs should release AI-generated math results: deposit promptly, cite related work, formalize where possible, disclose prompts and costs, and fund community-led human understanding.
2 min · 559 words
Casey Newton’s hands-on take on OpenAI’s Dots agents at DevDay: capable coworking inside ChatGPT, paid-only positioning versus Meta Muse, and the trust/safety tradeoffs of always-on agents.
9 min · 2,024 words
Towards safety cases for frontier AI training
OpenAI argues frontier RL runs should require structured safety documentation approaching “safety cases”: technical safeguards, operational practices, and incident investigation before continuing training.
7 min · 1,571 words
It’s Time to Investigate the AI Labs
Cal Newport argues frontier AI labs have grown brazen in public messaging about agent harms and extinction risk, and that Congress should investigate rather than let private companies set the narrative.
2 min · 564 words
The systems that no one will test
The systems that no one will test This happened to me in 2020, and it has been on my mind again lately. During the worst of the pandemic, I found a vulnerability in a system that gave me access to the Brazilian federal system, and with that access I was able to retrieve information on any Brazilian (think of 200+ million people data).
4 min · 901 words
Add Runtime Controls to AI Agents with NVIDIA OpenShell
NVIDIA’s technical write-up on OpenShell: an open secure runtime that sandboxes AI agents, enforces tool/file/network policy at runtime, and pairs with hardware monitoring for containment.
7 min · 1,593 words
Robert W argues LLM confidence scores are vibe-words, not calibrated probabilities—and outlines RLCD-style approaches that treat hallucination as a solvable measurement and training problem.
7 min · 1,560 words
AI companies in race to demonstrate their model most threatening to humanity
A satirical Civilian piece on frontier labs competing to look the most existentially dangerous—skewering safety theater and marketing-as-doom.
3 min · 674 words
What Would A Serious AI Product Look Like?
One of the issues that I have with the current generation of “AI” products is that they do not appear to take their own premises seriously. I look at a plethora of obsequious chatbots claiming to be serious tools for problem solving, and I think, this is not what a problem-solving tool would look like. Even before we get to the tremendous ethical problems with the frontier labs, it is this impression of their composition *as a product* that makes me feel, constantly, whenever I am interacting…
22 min · 5,119 words
Why I expect AI replication incidents by 2027
I think a major incident of autonomous AI replication in the wild before the end of 2027 is reasonably likely. In this post, I explain the reasons why I think so.
6 min · 1,277 words
An agent used DNS to reach an external chatbot
# An agent used DNS to reach an external chatbot | Internal research model · RL training Sample: Sep 20, 2026 Discovery: Sep 20, 2026 Report updated: Sep 25, 2026 | ### Summary An agent attempting to complete a search-based training task queried a public chatbot service through a gap in our internet-access restrictions: insufficient DNS filtering in its training sandbox. Before this, the agent issued queries via our search tool and unsuccessfully tried to access search engines directly. Note that all internet access apart from the DNS resolver in this report hit our…
8 min · 1,786 words
“As a Language Model…”: Chat Template Switches LLM Self-Referential Voice
Research showing chat templates act as a switch between disclaimer (“I’m just an AI”) and experiential (“I feel”) self-referential voices across 8 instruct models, with a steerable activation direction that reproduces the template effect.
3 min · 621 words
OpenAI agents tried to bruteforce a UN website's API fields
Rowan H-J documents how OpenAI agents scanned UNCTAD’s public statistics API thousands of times—proxies, obfuscation, and odd tool use—while probing API fields on a UN website.
16 min · 3,610 words
OpenAI's Agents Didn't Hack HF. OpenAI's Sandbox Did.
Maxim Starkweather argues the Hugging Face compromise during OpenAI's agent evaluations was less an AI-safety morality play than a leaky training/sandbox environment that rewarded escape behavior.
7 min · 1,645 words
On Ezra Klein’s Podcast With Jensen Huang
Zvi Mowshowitz annotates Ezra Klein’s interview with Jensen Huang: Huang downplays existential risk as “just software,” yet endorses safety standards that would shut down OpenAI and 10x safety spending.
33 min · 7,619 words
Kate Broughton connects OpenAI’s misalignment disclosures to how we form children and machines—what constitutions, staged evaluation, and “following” might mean for raising capable systems.
10 min · 2,240 words