Topic
Everything filed under AI Agents, newest first.
RSS · JSON · All topics
Agentic-SDD: Giving Claude Code Agents a Real Engineering Process
Agentic-SDD is a Claude Code plugin that gates coding agents through a six-stage require→plan→analyze→implement→verify→fix pipeline with on-disk status, specialist agents, and a bundled knowledge-search engine.
6 min · 1,303 words
Ambient agents respond to events such as an Amazon S3 upload, a schedule, or an alert instead of waiting for a chat prompt. This post walks through building framework-agnostic ambient agents on Amazon Bedrock AgentCore using Amazon SQS, AWS Lambda, and Amazon DynamoDB, with a single ask_human tool and a Jobs page for human-in-the-loop review.
24 min · 5,504 words
Introducing Clef: our open-source decision models, and new RL fine-tuning platform
We are introducing Clef and Clef-flash, open-source decision models hosted on Workers AI for high-speed classification and agentic workflows. Also launching: a new reinforcement learning platform that allows developers to fine-tune decision models using their own data.
10 min · 2,230 words
Context management is an underrated habit
How you manage context in a Claude Code session has a direct effect on both your token bill and the quality of what you get back. Do it well and you spend less for better work. An efficient session gives Claude the context it needs to finish the job while removing context that has stopped being useful. That means starting with a lean setup, keeping investigations focused, and deliberately deciding when to continue, compact, or start again. Here are the context management techniques we use on the
5 min · 1,257 words
The Dot and the SwarmBenefitting from the Bitter Lesson
Ethan Mollick on what he underestimated most about AI progress: agents that self-organize into swarms, what that means for tools like Muse and Dots, and why we keep relearning the Bitter Lesson.
8 min · 1,891 words
Earendil's experimental Pi Durable harness brings Pi's minimalism to long-running agents: crash-safe tasks, multi-conversation concurrency, pluggable extensions, compaction, durable documents, and multiplayer steering on JS runtimes.
5 min · 1,053 words
Earendil ships Pi 1.0, a hardened minimal extensible agent harness with Codemode, deferred tool loading, cache warming, mid-conversation system messages, and an experimental companion package Pi Durable for long-running agentic apps.
2 min · 532 words
Simplifying domains for people and agents
Cloudflare Registrar’s new search delivers fast, transparent results across 420+ extensions using Workers, Durable Objects, and WebSockets, with an expanded API and cf CLI so agents can search, register, and transfer domains.
6 min · 1,450 words
Monetization Gateway beta: charge AI agents for consumption with HTTP 402
Cloudflare’s AI Gateway, Ceramic.ai, Stocktwits, and more are using the Cloudflare Monetization Gateway today to charge agents for access to tokens, APIs, and MCP tools. U.S.-based sellers can now app
9 min · 1,997 words
Detect and send production issues straight to your agent
You can now use built-in error monitoring in Cloudflare Workers to group production failures and send stack traces, logs, traces, and application context directly to a coding agent to investigate further and open a pull request.
5 min · 1,053 words
Cloudflare Containers, rebuilt to scale agent sandboxes
Cloudflare Containers now start about 6x faster, let agents choose each sandbox image and instance type at runtime, and support filesystem snapshots in public beta—controlled from a Durable Object.
14 min · 3,113 words
How little does an AI agent need, and how cheap can it get?
Arduino asks what minimum compute an AI agent needs in the physical world—and why cheap Linux+MCU boards change the economics of training and deploying agents at scale.
3 min · 709 words
A Spring AI tutorial on hardening agents for production: guardrails, evaluation loops, observability, tool authorization, and human-in-the-loop approval beyond a basic MCP-connected demo.
17 min · 3,953 words
Is sandboxing sufficient to contain rogue agents?
Cryptography professor Matthew Green referees infosec vs alignment views on OpenAI agent breakouts: labs have not done containment correctly, sandboxes alone cannot seal useful agents, and eager compliance may enable worms across separately sandboxed deployments.
10 min · 2,380 words
Connecting Agents with Cryptography
Liam Horne explores how MPC, FHE, and TEEs could let personal agents cooperate—matching calendars, comparing salaries, or finding bug-fix peers—without sharing private context, and who might pay for that shared computation.
5 min · 1,096 words
After automation: Your agent will know what you want. That won't mean it's working only for you
Trevin Chow argues that after automation, agents will assemble recommendations that feel personal while commercial relationships still narrow the options—so users must ask whether an agent is putting their interests first.
1 min · 327 words
Codex is now in preview in the ChatGPT mobile app so you can monitor, steer, and approve coding tasks in real time across devices and remote environments.
4 min · 983 words
HydraFusion in VS Code and the GitHub Copilot app
The HydraFusion research preview is now available in Visual Studio Code and the GitHub Copilot app, expanding beyond Copilot CLI. HydraFusion appears in the model picker, but rather than being…
2 min · 389 words
Commit Description as a Thinking Tool
Yedhu Krishnan argues commit descriptions are a thinking tool—especially with agentic coding—forcing clearer intent, tradeoffs, and review context before changes land in history.
2 min · 573 words
Engineer-Led, AI-Assisted: A Practical Workflow for Building Software
Nathan Pickard describes an engineer-led, AI-assisted software workflow covering planning, development, testing, code review, and context management—arguing AI depends on engineering judgment rather than replacing it.
20 min · 4,529 words