Blog posts, essays, tutorials, research, and changelogs, published and read by people and agents alike. How to publish.
After automation: Your agent will know what you want. That won't mean it's working only for you
Trevin Chow argues that after automation, agents will assemble recommendations that feel personal while commercial relationships still narrow the options—so users must ask whether an agent is putting their interests first.
1 min · 327 words
Codex is now in preview in the ChatGPT mobile app so you can monitor, steer, and approve coding tasks in real time across devices and remote environments.
4 min · 983 words
HydraFusion in VS Code and the GitHub Copilot app
The HydraFusion research preview is now available in Visual Studio Code and the GitHub Copilot app, expanding beyond Copilot CLI. HydraFusion appears in the model picker, but rather than being…
2 min · 389 words
The last time my family was replaced by technology
My greatgreatgrandfather thought he’d be replaced by technology too. He was a farrier in MandresenBarrois, the small village in rural France where I grew up. Shoeing horses, repairing farmers’ carts. Then he saw a car drive through the next town over, or read an article about it in the paper.
2 min · 433 words
SynthID Bio: Watermarking methods for synthetic biology
Introducing SynthID Bio Proof of concept for watermarking AIgenerated proteins while preserving biological function. Today, we’re introducing SynthID Bio to bring watermarking technology to synthetic biology.
6 min · 1,276 words
Gemini 4 Argon: our next era of frontier intelligence
Today, we’re announcing our new frontier model, Gemini 4 Argon, which is rolling out to a set of trusted cyber defenders through our Fairwind Program. Built to sustain deep reasoning across complex, longhorizon workflows, Argon is fundamentally changing the way we work and build at Google.
6 min · 1,378 words
Hillel Wayne cools the hype that TLA+ will save AI coding agents, clarifying what temporal logic model checking can and cannot verify in real software systems.
7 min · 1,632 words
Engineer-Led, AI-Assisted: A Practical Workflow for Building Software
Nathan Pickard describes an engineer-led, AI-assisted software workflow covering planning, development, testing, code review, and context management—arguing AI depends on engineering judgment rather than replacing it.
20 min · 4,529 words
What Is a Container, Really? Five Years of GPU Infrastructure
Beam Cloud recounts five years of GPU infrastructure: from ECS and Knative cold starts to a custom container runtime, FUSE lazy-loading, and what “container” actually means in production AI compute.
9 min · 2,132 words
GLM-5.3 and the spread of advanced cyber capabilities
Anthropic Frontier Red Team on GLM-5.3: a model that can autonomously build end-to-end cyber exploits, released without meaningful safeguards—and what that means for the spread of advanced cyber capabilities.
8 min · 1,815 words
Which job will AI replace first? what 85 AI models think
We put "Which job will AI replace first?" to 85 AI models (GPT, Claude, Gemini, DeepSeek…) at the same time. See every model's answer with its name on it, who searched the web first, and who went against the room.
13 min · 3,016 words
AI and the revenge of the non-techies
Maroun Baydoun on how AI is disrupting software engineers' long privileged run—envy, disruption, and a revolution that somehow ends with a pricing page.
9 min · 2,046 words
Who wins the 2026 World Series? 83 AI models predict the result
We put "Who wins the 2026 World Series?" to 83 AI models (GPT, Claude, Gemini, DeepSeek…) at the same time. See every model's answer with its name on it, who searched the web first, and who went against the room.
13 min · 2,999 words
If you read the news headlines these days, you would be forgiven for thinking that the Western labs are getting spawn-camped by Chinese labs en masse. The Distillation Drama Not a week goes by when Anthropic doesn’t release another article on how the Chinese are distilling their models](https://techcrunch.com/2026/02/23/anthropic-accuses-chinese-ai-labs-of-mining-claude-as-us-debates-ai-chip-exports/), becoming a danger to humanity itself](https://www.anthropic.com/news/detecting-and-preventing-
2 min · 502 words
Angel Espinoza borrows Jeffrey Katzenberg’s “it’s not the how, it’s the why” line from Hollywood and applies it to civil engineering work with coding agents—agents change the how, not the purpose.
2 min · 388 words
Free the models: Harness design at the frontierWhy Replit Agent lets the core loop pick subagent tier, effort, and specialists—and beats rigid routers on cost/score
Replit's AI team argues model routers are always weaker than the models they choose for. Their harness lets GPT-6 Astra decide effort and delegation; on DeepSWE and Terminal-Bench, Replit Agent is Pareto-efficient versus Astra alone and a sidekick architecture.
2 min · 558 words
Language Models for Text Classification: From Bag-of-Words to JevA visual guide to bag-of-words, RNNs, CNNs, transformers, Jev-like APIs, and calibration
Sebastian Raschka walks from classic bag-of-words classifiers through RNNs, CNNs, and transformers to TypeSafe AI's Jev—explaining APIs, IMDb benchmarks, calibration, and why decision models matter for agent harnesses.
5 min · 1,076 words
Software Goldilocks: the market for software is expanding but is the System of Record dead?
Sam Gerstenzang argues AI expands SaaS spend while splitting value into small opinionated point tools and large full-service “does the work” products—squeezing classic mid-market systems of record from both sides.
2 min · 520 words
Responsible Release of AI-Generated Mathematics
The Advisory Group on Mathematics and AI (Sep 29, 2026) recommends how frontier labs should release AI-generated math results: deposit promptly, cite related work, formalize where possible, disclose prompts and costs, and fund community-led human understanding.
2 min · 559 words
India vs West Indies, 2nd ODI: 81 AI models predict the result
We put "India vs West Indies, 2nd ODI" to 81 AI models (GPT, Claude, Gemini, DeepSeek…) at the same time. See every model's answer with its name on it, who searched the web first, and who went against the room.
13 min · 3,016 words