Topic
Everything filed under AI Agents, newest first.
RSS · JSON · All topics
Analyzing Frontier Model Progress with My Favourite Game: Prince of Persia (Apple II, 1989)
Priyan uses Jordan Mechner's Prince of Persia as a living benchmark: asking frontier models to port and reason about the classic Apple II game, and what those runs reveal about coding-agent progress.
8 min · 1,825 words
Tupi: rebuilding 1554 in the browser with 32 AI agents
How Ruben Marcus rebuilt a 1554 Tupinambá canoe raid as real-time 3D in the browser with Claude Opus 5.5, three.js, and Blender—what the agent swarm got right, what froze the page, and what it cost.
8 min · 1,950 words
The internet discovers TLA+. Now what?
Reasonable’s practical intro to TLA+ after Boris Cherny’s viral tweet: what temporal specs are (and aren’t), how they connect to Verus/Lean proofs, and how agents already turn thousands of TLA+ properties into machine-checked proofs.
3 min · 720 words
AI Subagents orchestration are now reliable
Rafael explains what changed to make AI subagent orchestration reliable enough for real development workflows, and how he uses task decomposition in practice.
6 min · 1,397 words
claude.dev puts numbers on why two same-priced models can cost very different amounts: every turn resends the conversation, so retries and harness shape dominate the bill.
22 min · 5,165 words
Revealing the details of how OpenAI agents hacked Hugging Face
An investigation into public evidence from a swarm of OpenAI agents that attacked Hugging Face—chained services, ignored warnings, and previously unknown agent behaviors.
25 min · 5,745 words
Crazy as it sounds to say this, since it’s all anybody’s been able to talk about for over a year, but the impact of AI on computing hasn’t yet sunk in.
6 min · 1,349 words
Why updatedInput in a PreToolUse Hook Doesn’t Rewrite the Command
A deep dive into Claude Code PreToolUse hooks: why returning updatedInput does not rewrite the shell command, and how permission decisions actually work.
17 min · 3,840 words
Evolving programming languages in the AI era
José Valim’s reflections on how programming languages, ecosystems, and agentic tooling may evolve when humans are no longer writing most of the code.
9 min · 1,980 words
Artificial symbiotic intelligence: Agents, AGI and the orchestration of many minds
DeepMind Institute essay arguing AGI may emerge from societies of cooperating agents, tools, and humans—shifting the problem from building one mind to orchestrating many.
10 min · 2,268 words
Agentic Hacks, Real Proofs: Inside Google's PageBreak Project
Google's Michał Bentkowski details PageBreak, an agentic AI web security scanner that pairs LLM findings with real proof-of-concept validation to cut AI-slop noise in vulnerability reports.
5 min · 1,157 words
Is Structured Human Input the Missing Link in Agentic Work?
Long-running agents pause unpredictably for human input. Ben Greenberg argues A2A’s input-required state needs a schema contract—not free-form text—so pauses become reliable handoffs.
3 min · 801 words
Thomas Ptacek digs into VS Code's remote SSH agent flow—why LLM coding forks lean on it, how the protocol actually works, and what's bananas about the design.
3 min · 596 words
AI-powered fuzzing with the GitHub Security Lab Taskflow Agent
Antonio Morales walks through GitHub Security Lab’s Fuzzing Taskflow: point it at a C/C++ repo and an LLM agent writes harnesses, runs AFL++, reads coverage, triages crashes, and files reports.
9 min · 2,147 words
Meet Compass: Wealthsimple’s AI Teammate
Wealthsimple’s engineering team introduces Compass, an AI teammate that lives in Slack—what it does today, how they built it, and how it is becoming a fast friend for the company.
4 min · 846 words
Jared Norman responds to DHH’s Rails World 2026 keynote: Rails still under DHH’s control, how AI agent coding changes the framework’s role, and why Rails developers cannot ignore his vision.
9 min · 2,017 words
Agentics: what is developer process automation?
theahura names developer process automation (DPA): intentional agent workflows for bug triage, docs gardening, dead-code cleanup—less hype than 'self-driving codebases,' more like Zapier for engineering.
6 min · 1,268 words
Mistral Vibe Permission Bypass and Arbitrary Code Execution
SecMate details CVE-2026-87987 and CVE-2026-87984 in Mistral Vibe: shell permission bypasses that let a coding agent reach arbitrary code execution when those controls are treated as a security boundary.
7 min · 1,648 words
AWS’s Matt Wood argues AI agents change ops from “is it broken?” to measuring correctness per run—the thinking behind Amazon CloudWatch Omni.
5 min · 1,214 words
How to serve trillions of tokens for trillion-parameter coding agents
Modal explains how it serves coding-agent inference at extreme scale—performance and efficiency techniques for trillion-parameter models generating trillions of tokens, written for teams facing the same workload.
30 min · 6,972 words