I've been an early user of coding agents like Claude Code, Codex and many others, but a general purpose agent that I could make my own and use for non-coding related tasks was more interesting.
I've been testing and building agents since OpenClaw (back when it was called ClawdBot in Dec 2025 or so). OpenClaw quickly became too cumbersome. I tried Goose, Zo Computer and a bunch of others too. I then tried Pi and that really opened up my mind on the potential of agents—especially its lightweight harness (~1,000 tokens to respond to "hello" vs 15–20K for Claude Code).
I currently have three agents live: stardust trades NSE equities with real capital, vita tracks my family's health, medications and appointments, vera manages investing, emailing, and research. Different domains, different stakes, same skeleton.
One profile, one job
Every agent is its own Hermes gateway profile—its own systemd service, its own config root, its own bot identity, its own process. Not a shared assistant with a trading mode and a health mode bolted on.
The model reasons. The code still decides.
Every agent splits into the same three layers:
- A skill — a markdown playbook, loaded fresh into context per job. It teaches the model how to think: the mandate, the invariants, the accumulated pitfalls.
- A tool — the only thing the model is allowed to act with. Every unit of real work is a named tool, never prose telling the model to run a script.
- An engine — the already-tested CLI or code path the tool shells into.
The model never touches the broker, the database, or the filesystem directly, and never runs Python it wasn't handed.
Ground truth never lives in the model
stardust's own ledger can disagree with the broker's numbers. When it does, the broker's numbers become the source of truth. vita works the same way: the SQLite DB is the record. For vera, the source of truth is the md file vault in github. The model's own memory is a cache that can be stale or wrong.
Gate expensive reasoning behind cheap logic
Not every decision needs an LLM call. stardust's cadence controller is a five-line deterministic state machine that runs every five minutes. Across all three agents: script what's deterministic, reserve reasoning for what actually requires judgment, and put a cheap check between the trigger and the expensive call.
Bound the blast radius structurally, not behaviorally
None of the hard limits exist because the prompt asks nicely. They exist because the tool layer makes the wrong action physically unavailable—e.g. stardust can only act on a security already present in its tracked state; every live entry places a real exchange-side stop.
Foreman — the task manager behind my agents
Agent runtimes change quickly. Without a durable work layer, useful project state is trapped inside the conversation. Foreman provides a portable work representation (objectives, plans, decisions, state, progress, blockers, artifact references) that can survive a runtime change.
The shape, restated
One agent profile per job, isolated down to the process. A skill that reasons, a task manager to not drop the ball, a tool that's the only way to act, an engine that's already been tested. One ground truth outside the model. Cheap logic gating expensive reasoning. Limits enforced by what the tools can do, not by what the prompt asks for.
Trust disclosure from the author: ideas are their own; content enhanced with AI for clarity. Full post: osborne.vc/blog/my-ai-agents.