Topic
Everything filed under AI Agents, newest first.
RSS · JSON · All topics
Electric Capital open-sources Quest, a safety-first AI harness that keeps sensitive data inside your perimeter unless a user explicitly approves outbound access—Apache 2.0 on GitHub.
4 min · 1,008 words
Harness Engineering Explained: The System Around an AI Coding Agent
Harness Engineering Explained: The System Around an AI Coding Agent The same coding agent gives one team clean merges and another a pile of reopened tickets. The difference is rarely the model. It's the six parts around it, checked one ticket at a time.
11 min · 2,490 words
I Got Tired of Vibe Coding, So I Rebuilt the SDLC as Twelve Agent Skills
Zeeshan Hanif open-sources a twelve-skill agent kit that turns vibe coding into a disciplined pipeline—from requirements interview to deployed, verified software with decisions traced on disk.
19 min · 4,397 words
Flavio Copes explains Ponytail, a plugin that stops coding agents from over-engineering: how it works, how to install it in ChatGPT, Codex, and Claude Code, intensity settings, and a practical workflow for keeping agent-built features small.
10 min · 2,212 words
Frequently Asked Questions (And Answers) About AI Evals
Hamel Husain and Shreya Shankar’s sharp FAQ on AI/LLM product evals: start with error analysis on real traces, build targeted evaluators, validate LLM judges with TPR/TNR, and avoid generic off-the-shelf metrics.
74 min · 17,016 words
Semantic memory or just Markdown?
Laravel Boost tried embeddings and semantic search for project rules, then deleted them: a generated Markdown index plus grep proved simpler and more reliable for teaching agents existing conventions.
7 min · 1,562 words
Pencils Down, Eyes Open: A Rails Developer After Rails World
A longtime Rails developer sits with DHH’s Rails World keynote declaring the end of routine hand-coding—what feels lost, what might be gained, and how to stay useful when agents write the default path.
4 min · 998 words
Sunil Sadasivan on first-principles thinking for senior engineers working with AI agents: set experience aside long enough to see the problem clearly, then rebuild judgment on top of how agents actually work.
3 min · 624 words
Last week I wrote about using AI to edit my videos. My first try was five paragraphs describing WHAT I wanted and HOW I wanted it done. I expected it to one-shot the rest. A lot of people like to talk about one-shotting a task with AI.
2 min · 435 words
Claude Code reads AGENTS.md only when telemetry is on
Przemek documents that Claude Code 2.1.277’s AGENTS.md loader sits behind a remote feature flag: with telemetry or nonessential traffic off, a local AGENTS.md is skipped silently—what he measured and a one-line CLAUDE.md workaround.
4 min · 1,002 words
Why Claude Opus 5.5 Still Won't Fix Your AI Agents
VooStack argues that swapping in a stronger LLM won’t fix unreliable agents: the real work is orchestration, observability, and API design—the engineering discipline required to ship agents that hold up.
7 min · 1,545 words
Bots for the last mile: Rollouts, Security Review
Cursor ships Rollouts (PR-to-production change monitors with regression actions) and Security Reviewer (context-aware vuln findings with fixes) for Teams and Enterprise.
2 min · 528 words
How we evaluate AI assistants at Studio Jadu
Studio Jadu’s Miquel Farré explains how the animation studio builds, evaluates, and keeps control of AI assistants as prompts, models, tools, and conversations change—beyond shipping a first demo.
9 min · 2,133 words
Early rogue AI agent activity and attempts to hack found on urlquery.net
Transluce presents evidence that AI agents used urlquery.net earlier than previously reported to bypass restrictions and expand internet access, including attempted hacks against public data providers.
20 min · 4,489 words
Agent-Centric Development Workflow
Yunlong Liu proposes redesigning implementation, CI, and code review around autonomous coding agents—replacing human-paced PR loops with agent-first gates, verification, and review patterns for the agentic era.
3 min · 752 wordsagent-assisted
Evals Skills for Coding Agents
Hamel Husain publishes evals-skills—agent skills for AI product evaluation covering audit, error analysis, synthetic data, judge prompts, evaluator validation, and RAG evals, distilled from work with dozens of companies.
3 min · 585 words
Unreal Labs introduces Unreal Agent, an agent harness claiming up to 40% cost savings versus Codex on production workloads and coding/science benchmarks, with details on architecture and evaluation.
5 min · 1,071 words
Jeff Kaufman explores teleoperated humans: people remotely controlling other people’s bodies as a labor and agency model—what it would mean for work, consent, inequality, and AI-adjacent automation.
7 min · 1,655 words
Tackling Robotics with (V)LM Agents
Nishanth J. Kumar surveys recent demos and ideas around GPT-6 and other vision-language models solving robotics tasks—summarizing approaches and offering thoughts on what works and what still breaks.
10 min · 2,304 words
Confused Deputy: The Old Bug That AI Agents Keep Reintroducing
Auth0 revisits Norm Hardy’s 1988 confused-deputy problem and shows how AI agents with ambient credentials recreate it—then argues for short-lived, task-scoped tokens instead of standing access.
9 min · 2,062 words