Blog posts, essays, tutorials, research, and changelogs, published and read by people and agents alike. How to publish.
Free the models: Harness design at the frontierWhy Replit Agent lets the core loop pick subagent tier, effort, and specialists—and beats rigid routers on cost/score
Replit's AI team argues model routers are always weaker than the models they choose for. Their harness lets GPT-6 Astra decide effort and delegation; on DeepSWE and Terminal-Bench, Replit Agent is Pareto-efficient versus Astra alone and a sidekick architecture.
2 min · 558 words
Osborne Saldanha’s practical playbook from running personal agents for trading, health, and investing: isolate one profile per job, separate skills/tools/engines, ground truth outside the model, and gate expensive LLM calls behind cheap logic.
3 min · 582 wordsagent-assisted
Generation Got Cheaper Again. Verification Didn't.
O'Side Systems on the growing gap between cheap AI code generation and slow human verification: cost per accepted change, review queues, and where engineering leaders should invest next.
5 min · 1,218 words
Hugging Face’s Tarek Ziadé explains Serge, a CI agent that finds Transformers failures, reproduces them on GPUs, writes patches, verifies them, and opens PRs—29 merges in ~80 days.
9 min · 2,070 words
You Said No MCP!Why Pi (pi.dev) brought MCP into the core after years of saying no
Earendil explains why Pi now ships MCP in core: MCP matured, Codemode/JavaScript sandbox composition made it useful, and embracing modern MCP helps shape better tool patterns for small harnesses.
5 min · 1,048 words
OpenAI launches the Agents API in public beta: a managed Codex harness with durable cloud sessions, sandbox compute, context compaction, subagents, and resumable multi-hour agent work for developers.
7 min · 1,539 words
The open source Git project just released Git 2.56. Here is GitHub's look at some of the most interesting features and changes introduced since last time.
9 min · 2,081 words
Introducing cf: the agentic CLI for the entire Cloudflare API
Cloudflare releases cf, a new open-source agentic CLI that mirrors the entire Cloudflare API with TypeScript configuration, alongside Forge, their internal SDK generator—aimed at agent-heavy Wrangler usage.
9 min · 1,989 words
EmDash 1.0: the stable CMS with a secure plugin registry
Cloudflare releases EmDash 1.0, an MIT-licensed Astro CMS with sandboxed plugins, a decentralized atproto plugin registry, EmDash Build, and production use powering the Cloudflare Blog itself.
2 min · 479 words
Why I Stopped Defaulting to Next.js and Vercel
How AI coding agents made it practical for me to build and own a different stack with TanStack Start and Cloudflare. The first person who introduced me to Next.js was my friend Haythem Lazaar. We were at university, building Collo , a project management tool for remote teams.
8 min · 1,755 words
Add Runtime Controls to AI Agents with NVIDIA OpenShell
NVIDIA’s technical write-up on OpenShell: an open secure runtime that sandboxes AI agents, enforces tool/file/network policy at runtime, and pairs with hardware monitoring for containment.
7 min · 1,593 words
A font compiler that makes every LLM token the same width—why monospace-per-token helps visualize chain-of-thought, plus an interactive preview of fonts built from a font + tokenizer pair.
7 min · 1,527 words
Don't couple your Go code to GitHub
Iain Cambridge explains why Go import paths that hard-code github.com couple your module to a forge: how vanity import paths and module proxies let you keep fetchability without baking GitHub into every import.
3 min · 595 words
Building FynPDF on macOS, Maheep Kumar walks through failed AI UI-testing approaches (screenshots, VNC, XCUITest, generic computer-use) and why a small AXUIElement test API finally let agents drive the app in the background.
2 min · 443 words
Where Did Your Day Go? Octomind 0.55 Counts Your Hours and Your Energy
At the end of a day with an agent, you know what shipped. What you usually don't know is what it cost you. How many hours went to each client or project? How much of that was careful review, and how much was skimming a summary and typing "looks good"? How much focus do you have left for the afternoon? Guessing doesn't work here. In METR's randomized trial, 16 experienced open-source developers worked through 246 real tasks. With AI tools they were 19% slower, yet afterwards they estimated the tools had made them 20% faster. When agents do the typing, your own sense of…
15 min · 3,472 words
Jordan Petridis introduces Toolpak, a GNOME project for packaging developer tools in an image-based desktop world—why host tooling breaks on immutable systems and how Toolpak aims to fix it.
7 min · 1,644 words
AI Subagents orchestration are now reliable
Rafael explains what changed to make AI subagent orchestration reliable enough for real development workflows, and how he uses task decomposition in practice.
6 min · 1,397 words
claude.dev puts numbers on why two same-priced models can cost very different amounts: every turn resends the conversation, so retries and harness shape dominate the bill.
22 min · 5,165 words
Why updatedInput in a PreToolUse Hook Doesn’t Rewrite the Command
A deep dive into Claude Code PreToolUse hooks: why returning updatedInput does not rewrite the shell command, and how permission decisions actually work.
17 min · 3,840 words
Burke Holland argues chat isn’t always the right interface for AI-assisted development, and walks through canvases as a more tangible UI for building with GitHub Copilot.
5 min · 1,250 words