Topic
Everything filed under AI Agents, newest first.
RSS · JSON · All topics
Arcturus Labs compares OpenAI’s emerging decision-model direction with TypeSafe’s Jev—and asks whether a frontier lab can absorb the System One / structured-decision niche startups are building.
12 min · 2,748 words
How Google Agent Substrate Works: 250 Agents on 8 Pods
A technical breakdown of Google’s Agent Substrate: how it multiplexes hundreds of stateful agent sessions onto a handful of Kubernetes pods with fast suspend/resume.
11 min · 2,511 words
Introducing Strands harness: frontier performance with 28% lower token cost
Arron Bailiss introduces Strands harness: a fully assembled, customizable local/cloud agent that aims for Claude Code/Codex-like “it just works” behavior with about 28% lower token cost.
5 min · 1,198 words
AI coding has made CI a bottleneck, so we reworked ours to keep up
Linear's Mufeez Amjad explains how agent-accelerated shipping made CI the bottleneck, and how they cut PR wait time and runner cost while test suites nearly quadrupled.
8 min · 1,907 words
dlab Open Source Week: Frontier AI on Your Own Hardware
Tim Dettmers argues that small academic labs can compete by building coherent open-source ecosystems rather than isolated papers. He previews local frontier models, autonomous research tools, and an auto-compaction technique designed to run long agent sessions while cutting cost.
1 min · 276 words
M5 Ultra Mac Studio Review: The Dream Mac for Local AI Agents
Federico Viticci reviews the M5 Ultra Mac Studio with 256 GB RAM—why it is a dream machine for local AI agents, what workloads it unlocks, and where it still falls short.
14 min · 3,293 words
Rebuilding Nym’s agent around Jev
How Nym rebuilt its agent stack around TypeSafe’s Jev for guardrails, browser actions, and tool selection—with benchmarks and a shopping demo.
10 min · 2,186 words
Language-model groups overstate consensus when replaying human deliberation on a reasoning task
LLM groups replaying human Wason discussions reach full consensus far more often than humans—partly because agents almost always speak up—cautioning against treating multi-agent agreement as truth.
2 min · 378 words
Multimodal Agents: From Perception to Action
Illustrated notes from Berkeley’s LLM Agents lecture 7: OSWorld outcome checks, AgentTrek trajectories, TACO tools, and Aguvis grounding—why reading a screen is not the same as finishing the task.
8 min · 1,915 words
Building AI Agents with Spring AI — Tool Calling, Memory, and Autonomous Workflows
A practical Java guide to Spring AI agents: tool calling, memory, and autonomous workflows that move beyond one-shot LLM calls into real multi-step agent loops.
12 min · 2,810 words
Why AI Coding Agents Crash at 3 AM: The Happy-Path Mirage & The Forced Continuity Defect
Part 2 of Synthetic Scars: why agents that look flawless on staging fail in production—happy-path training, forced continuity, and transferring pager-duty instincts to autonomous coders.
10 min · 2,387 words
Your AI Coding Agent Can Be Attacked by the Repository It Opens
Why opening an untrusted repo with an AI coding agent is a security boundary problem—prompt injection via files, tool abuse, and practical defenses for agent workflows.
4 min · 1,028 words
Done, Keys, and a Second Check
An operating-model essay arguing enterprises need ownership of done, keys, escalate/monitor rights, and a second check—not just consumer-style AI agents that complete tasks.
5 min · 1,180 words
Trying the Software Factory Pattern
One of the interesting challenges of the AI ecosystem in 2026 is that new, effective patterns emerge faster than I can adopt them. I’ll find a handful, get back to work, and realize a month later that I’d missed four or five more. The adoption cycle for Imprint this year has been something like: - April: local development is bottlenecked on checkout and worktree model, instead create ~10 local workspaces which each have an independent checkout of every repository, and operate at the workspace level, not at the repository level, so it can generate cross-repository pull requests across…
4 min · 871 words
How Instinct's memory works: a reverse-engineering teardown
A black-box teardown of Instinct
15 min · 3,416 words
Keva: Running Coding Agents On-Device on Unrooted Android
Simon Lin's technical paper on Keva—an on-device Android AI coding agent running Claude Code/Codex-style loops—covering architecture, failure modes, and systems lessons without rooting the phone.
33 min · 7,580 words
Model Context Protocol with Spring AI, Building MCP Clients and Servers in Java
Ayush Shrivastava walks through building MCP clients and servers with Spring AI in Java: tool discovery, protocol basics, and wiring MCP into agentic Spring applications beyond a basic demo.
13 min · 2,972 words
Remus Lazar mined a year of commits, PRs, and Slack after a 64-day coding streak with agents—and set hard limits: two agents max, a pause between plan and build, close sessions, and take the day off.
8 min · 1,873 words
Own the Agent, Rent the Intelligence: Building My Always-On AI Agent Server
James M explains why a Mac mini M6 became his always-on Hermes agent server—routing hard work to cheap cloud models like DeepSeek Flash and Claude Sonnet instead of owning local inference hardware.
24 min · 5,439 words
Ben Swerdlow benchmarks Codex, Claude, and Grok agents across 171 StarCraft: Brood War matches—leaderboards, APM, cost per game, and analysis showing none played beyond beginner while Codex Astra led.
9 min · 2,151 words