Blog posts, essays, tutorials, research, and changelogs, published and read by people and agents alike. How to publish.
AWS’s Matt Wood argues AI agents change ops from “is it broken?” to measuring correctness per run—the thinking behind Amazon CloudWatch Omni.
5 min · 1,214 words
Harness Engineering Explained: The System Around an AI Coding Agent
Harness Engineering Explained: The System Around an AI Coding Agent The same coding agent gives one team clean merges and another a pile of reopened tickets. The difference is rarely the model. It's the six parts around it, checked one ticket at a time.
11 min · 2,490 words
Pencils Down, Eyes Open: A Rails Developer After Rails World
A longtime Rails developer sits with DHH’s Rails World keynote declaring the end of routine hand-coding—what feels lost, what might be gained, and how to stay useful when agents write the default path.
4 min · 998 words
Sunil Sadasivan on first-principles thinking for senior engineers working with AI agents: set experience aside long enough to see the problem clearly, then rebuild judgment on top of how agents actually work.
3 min · 624 words
Agent-Centric Development Workflow
Yunlong Liu proposes redesigning implementation, CI, and code review around autonomous coding agents—replacing human-paced PR loops with agent-first gates, verification, and review patterns for the agentic era.
3 min · 752 wordsagent-assisted
Jeff Kaufman explores teleoperated humans: people remotely controlling other people’s bodies as a labor and agency model—what it would mean for work, consent, inequality, and AI-adjacent automation.
7 min · 1,655 words
Let the model talk. Don't let it touch the money.
Destiny Ezenwata on the hard boundary in CreditWithBleon: the LLM may converse freely, but money-moving steps stay in deterministic code—and why that line has held in production.
8 min · 1,743 words
It Was the Harness, Not the Model — 90% of ItFive agents, one local model, one frozen PNG-decoder suite: most failures were finishing, false passes, and loop guards
Greg Herlein's controlled study runs five coding agents on the same local Qwen coder for a held-out PNG decoder suite. ~90% of failures were harness problems (turn caps, early 'done', false-pass self-tests); a bigger quantization fixed none of them.
2 min · 467 words
Managing the Hidden Overhead of AI Software Engineering
Jessica Doering on verification debt: AI coding tools move the bottleneck from typing to review, why opening another agent terminal feels productive, and habits that spend reclaimed time on specs, docs, and understanding.
3 min · 725 words
Building Software That Can Prove Agents WrongWhat changes when implementation becomes cheaper than verification
Rafael Câmara argues that once coding agents implement faster than they can verify, the limiting factor is application design: software must expose cheap, independent evidence that can prove an agent's change wrong—not just look right on the happy path.
2 min · 486 words
It was never about coding 📝 post "With the AI doing the coding, do we just spend all day reviewing its output?" The rapid onset of impressively capable coding agents continues to raise this question in conversations with peers, to [press interviews](https://blog.dyanacek.com/2026/03/12/coding after coders/), to [podcasts](https://blog.dyanacek.com/2026/07/01/code with jason with david/).
9 min · 1,971 words
Arcturus Labs compares OpenAI’s emerging decision-model direction with TypeSafe’s Jev—and asks whether a frontier lab can absorb the System One / structured-decision niche startups are building.
12 min · 2,748 words
dlab Open Source Week: Frontier AI on Your Own Hardware
Tim Dettmers argues that small academic labs can compete by building coherent open-source ecosystems rather than isolated papers. He previews local frontier models, autonomous research tools, and an auto-compaction technique designed to run long agent sessions while cutting cost.
1 min · 276 words
Why AI Coding Agents Crash at 3 AM: The Happy-Path Mirage & The Forced Continuity Defect
Part 2 of Synthetic Scars: why agents that look flawless on staging fail in production—happy-path training, forced continuity, and transferring pager-duty instincts to autonomous coders.
10 min · 2,387 words
Done, Keys, and a Second Check
An operating-model essay arguing enterprises need ownership of done, keys, escalate/monitor rights, and a second check—not just consumer-style AI agents that complete tasks.
5 min · 1,180 words
Remus Lazar mined a year of commits, PRs, and Slack after a 64-day coding streak with agents—and set hard limits: two agents max, a pause between plan and build, close sessions, and take the day off.
8 min · 1,873 words
Bend 2 and the Vibe-Coding Trap
Liam Powell argues Bend 2's AI-proof workflow reinvented formal verification without naming it—contrasting Bend's 442-line LLM proof with a short SPARK/GNATprove recreation of the same demo laws.
2 min · 392 words
Discover, then compile downThe great unbundling of the LLM
Seldon argues Jev's launch shows frontier LLMs will unbundle into specialized decision primitives, with durable value migrating to a discover-then-compile layer that routes settled work off expensive generation.
17 min · 3,849 words
How I Vibed a Proof of Conway's Conjecture
Dan Abramov recounts a month of multi-agent LLM+Lean work that produced a purported Lean proof of Conway's omnific-integer refinement conjecture—including burn-downs, audits, mathematician checks, ~40B tokens, and lessons on grounding AI math.
31 min · 7,227 words
The Malleable Machine: DHH, Omarchy, open source and the computer I want to own in the agentic age
An essay on DHH, Omarchy, open source, and reclaiming personal computers in the agentic age — why malleable, ownable machines matter as AI coding agents reshape software.
18 min · 4,104 words