{"article":{"slug":"what-is-codemode","title":"What is Codemode","subtitle":null,"summary":"Armin Ronacher explains why Pi 1.0 added MCP support through Codemode despite his long preference for bash and scripts: bash can only compose programs that run in the execution environment, while Codemode lets the model orchestrate harness-side tools like image reads and sub agents in a sandboxed QuickJS/WASM runtime, plus the open problems that remain.","content_type":"blog_post","language":"en","canonical_url":"https://lucumr.pocoo.org/2026/10/6/codemode/","author":{"name":"Armin Ronacher","url":null,"person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"Armin Ronacher's Thoughts and Writings","url":"https://lucumr.pocoo.org/","listing_slug":null,"listing":null},"topics":[{"name":"AI Agents","slug":"ai-agents","url":"https://listedarticles.com/topics/ai-agents"},{"name":"LLMs","slug":"llms","url":"https://listedarticles.com/topics/llms"},{"name":"Developer Tools","slug":"developer-tools","url":"https://listedarticles.com/topics/developer-tools"},{"name":"Software Engineering","slug":"software-engineering","url":"https://listedarticles.com/topics/software-engineering"}],"about_listings":[],"cover_image_url":null,"license":"CC BY-NC 4.0","word_count":2672,"reading_minutes":12,"published_at":"2026-10-06T00:00:00.000Z","added_at":"2026-10-06T17:13:39.682Z","updated_at":"2026-10-06T17:13:39.682Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":true},"profile_url":"https://listedarticles.com/articles/what-is-codemode","markdown_url":"https://listedarticles.com/articles/what-is-codemode.md","example":false,"citation":"Armin Ronacher, Armin Ronacher's Thoughts and Writings. \"What is Codemode.\" 6 Oct 2026. https://lucumr.pocoo.org/2026/10/6/codemode/ (CC BY-NC 4.0)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://lucumr.pocoo.org/2026/10/6/codemode/"},"body_markdown":"\nMore than a year ago I wrote a few posts here that recommended people not to\nload custom tools into their context (or\n[MCP](https://en.wikipedia.org/wiki/Model_Context_Protocol) servers) but to just\nuse more scripts.  Most importantly I wrote that [Code Is All You\nNeed](https://lucumr.pocoo.org/2025/7/3/tools/) and I wrote about that [MCP needs\ncode](https://lucumr.pocoo.org/2025/8/18/code-mcps/).  With Pi 1.0 we now added MCP support via Codemode\nwhich in some ways is a long time coming, but then also maybe somewhat\nsurprising to some.  So I want to share some updated thoughts on this blog on\nwhat this all means.\n\nWhen a harness like Pi provides tools for an LLM to call, it does so by\nsupplying some tool definitions which then translate into some token structure\non the server side.  Whether a model is encouraged to call a tool is the result\nof the reinforcement learning process.  Something [I wrote about\nbefore](https://lucumr.pocoo.org/2026/7/4/better-models-worse-tools/) if you want to learn more.\n\nOne of the reasons we strongly lean towards CLI and bash is because it allows\neasy composition of calls, and because the model also learns how the file system\nworks when it’s trained.  So when it invokes a tool like `echo foo > /tmp/test.txt` the model also learns that after that tool call, there is now a\nfile called `test.txt` in `/tmp`.\n\nHowever bash has one fundamental limitation which is that it can only compose programs that run. And there are some things, which are not programs, but native tools to the LLM and they sort of have to be.\n\nThe most obvious example here is `read` or `view_image`. If a multimodal model\nneeds to read an image, it cannot use `cat` for that because the harness needs\nto inject the actual image payload into the protocol of the LLM.\n\nAnother quite vivid example are sub agents. In order to spawn and orchestrate sub agents, it’s tricky to avoid tools that are provided by the harness. While in theory the agent could provide a CLI tool that talks to the outer harness via environment variables and Unix sockets, it’s a rather crude process. It however has another issue, and that is where the code runs.\n\nTo better understand that, it’s important to think a bit more about where all the\nbits and pieces run.  There really usually are two different systems involved.\nThe first is the brain, the harness: it runs on one machine. It’s trusted.  The\nsecond is *often* the same machine, but it’s really where the tools are\nexecuting: the hands.  In Pi we now call this the execution environment, but you\ncan think of it as the target of all the operations.\n\nCrucially what is important for us, is that there is a dividing line between the harness brain and the target environment that runs bash and executes the tools.\n\nAnd splitting this in half has some really important consequences.  For a start\nit means that they are running on different file systems and they have different\nlevels of trust.  If you for instance use a sandboxing solution [like\nGondolin](https://earendil-works.github.io/gondolin/) your bash stuff will be\nsandboxed just fine, but the harness itself will not be.\n\nWhich brings us to what Codemode really does: it’s a way for the LLM to express and orchestrate complex operations on the harness side, but not the execution environment side. Codemode runs in the harness, in its own sandbox. In case of Pi it’s running in QuickJS within a WASM runtime with intentional limitations: no network, no file system, no timers, limited RAM. The only way is to call more tools. You could also imagine that Codemode could run Scheme or some other language as well.\n\nIf you are not familiar with Codemode, it’s basically just a way to issue\ntool calls from within some language, in our case JavaScript.  That allows you\nto compose those calls without necessarily going through the LLM’s context.\nCredit for naming goes to our friends at Cloudflare [who coined\nit](https://blog.cloudflare.com/code-mode/).\n\nFor instance if you issue a bash call as a regular tool call in the LLM, then we only throw the trailing 2000 lines into the context and if the agent wants more, it needs to look at the overflow file itself. If however the agent issues that invocation via Codemode, then the Codemode side gets larger outputs sent structurally.\n\nMost importantly, because Codemode is JavaScript the agent can express concurrent operations and basic workflows. A common way in which you see agents now use this, is to first probe at 5-10 items from some tool response to see what it looks like, and to then write a Codemode script that processes the next n items.\n\nCodemode also allows you to throw state into the transcript! That means that one Codemode invocation can stash away data, that the next call in the session can load again. And remember: this is on the harness host, not the sandbox.\n\nIn case of Pi, Codemode also allows you to issue calls that naturally do not make any sense in Pi’s traditional interface. For instance if you want to generate images with an image model or you want to classify some text with a one shot classifier model, those Pi APIs are exposed via Codemode, but not via regular tools where they would just waste context.\n\nSo now that we talked a bunch about it, it’s probably worth being a bit more explicit about it. Let’s walk ourselves through some invocations of Codemode of recent Pi sessions of mine. Note that none of this code is human written. It’s from real sessions of Pi, just re-indented for your viewing pleasure. The agent starts using Codemode automatically either because it’s a task where the model already naturally picks up that tool, or because a user asked it to.\n\nNote that Codemode is by default only enabled in Pi when MCP is enabled, but you\ncan turn it on with `\"defaultTools\": [\"+codemode\"]` in the settings.  Just ask\nPi to enable it for you.\n\nLet’s start simple with image generation. Image generation is a feature that Pi supports in the AI SDK core, but it’s not a tool that the agent can use. In the past the only way to use image models has been to write a bespoke extension or to have the agent run node itself and use the internal image APIs. However because we expose quite a few of the internal model APIs within Codemode, it means that the agent can use it:\n\n```\nconst [painter] = await models.getAvailableOfType(\"image\");\nconst result = await models.generateImages(painter, {\n  input: [{ type: \"text\", text: \"A cute little puppy sitting on a grassy \" +\n    \"lawn, soft natural light, photorealistic\" }],\n});\nif (result.stopReason !== \"stop\") return result.errorMessage;\nfor (const block of result.output) {\n  if (block.type === \"image\") image(block);\n  else text(block.text);\n}\n```\nNote that the call to `image()` sends the image back as image content to the\nLLM.  On the harness side it feeds it directly into both the agent, as well as\nonto disk as a temporary artifact in case the agent wants to be able to pass\nthat image back to bash.\n\nSimilar things apply to classifier models such as [Jev](https://typesafe.ai/).\nThey also do not fit well into the workflows of an agent through the typical\ntools.  But rather than making a bespoke tool available, Codemode just allows\nthe agent to reach into the AI SDK and invoke those directly.  Here you can see\nhow Jev is used to mass process GitHub issues for a quick sentiment analysis:\n\n```\nconst jev = await models.getModelOfType(\"classifier\", \"typesafe\", \"jev-latest\");\nconst r = await tools.bash({\n  command: \"gh issue list --state open --limit 100 \" +\n    \"--json number,title,body,comments\",\n});\nconst issues = JSON.parse(r.output);\nconst results = await Promise.all(issues.map(async (issue) => {\n  const res = await models.classify(jev, {\n    state: {\n      title: issue.title,\n      body: (issue.body || \"\").slice(0, 4000),\n      comments: issue.comments.slice(-5).map(c => c.body.slice(0, 800)),\n    },\n    questions: {\n      sentiment: {\n        type: \"choice\",\n        instructions: \"What is the overall sentiment of the author towards pi?\",\n        criteria: {\n          positive: \"Appreciative, happy, constructive praise\",\n          neutral: \"Matter-of-fact report or request without emotion\",\n          negative: \"Frustrated, annoyed, upset, or angry\",\n        },\n      },\n      frustration: {\n        type: \"score\",\n        instructions: \"How frustrated is the reporter?\",\n        criteria: [\"not at all\", \"mildly\", \"clearly frustrated\", \"very angry\"],\n      },\n      kind: {\n        type: \"choice\",\n        instructions: \"What kind of issue is this?\",\n        criteria: {\n          bug: \"Bug report or regression\",\n          feature: \"Feature request or enhancement\",\n          question: \"Question or support request\",\n          other: \"Docs, discussion, meta, spam\",\n        },\n      },\n    },\n  });\n  if (res.stopReason !== \"stop\") {\n    return { n: issue.number, title: issue.title, error: res.errorMessage };\n  }\n  return { n: issue.number, title: issue.title, ...res.answers };\n}));\nstore(\"sentiment_results\", results);\nreturn results\n  .filter(r => !r.error)\n  .sort((a, b) => b.frustration.score - a.frustration.score)\n  .slice(0, 12)\n  .map(r => `#${r.n} ${r.frustration.score.toFixed(2)} [${r.kind.choice}] ${r.title}`);\n```\nNote how in that above example we also call `store()` which dumps the result of\nthat execution into the session transcript.  A future invocation of Codemode can\nthus read back that result if it wants to.\n\nThe `Promise.all` here is fine, because Pi limits the total number of concurrent\ntool executions itself to four and maintains a queue for the rest.\n\nA more adventurous example is to use Jev to drive a game engine for debugging purposes:\n\nHere it knows about my `tankctl` command and it built itself quickly a minimal\nharness around it to drive a game loop to assist a user with debugging a\nproblem. Note how it built a 30 step loop in which each step goes back to both\nthe game engine to get a text dump of what’s going on, and then to Jev to\ndetermine what to do next:\n\n```\nconst jev = await models.getModelOfType(\"classifier\", \"typesafe\", \"jev-latest\");\nconst tank = async (cmd) =>\n  (await tools.bash({ command: `tools/tankctl \"${cmd}\"` })).output;\nawait tank(\"start --map assets/maps/night_arena.map\");\nconst questions = {\n  action: {\n    type: \"choice\",\n    instructions: \"You control the tank '@' in a top-down tank game. \" +\n      \"Choose the best next action.\",\n    criteria: {\n      attack: \"an enemy has line of sight to you and you can fire at it\",\n      approach: \"no enemy has line of sight; drive toward the nearest enemy\",\n      dodge: \"an enemy shot is heading at you and will hit soon\",\n      powerup: \"a powerup is close and no enemy threatens you\",\n    },\n  },\n};\nfunction commandFor(choice, st) {\n  const p = st.player;\n  const enemy = st.enemies.filter(e => !e.dead)\n    .sort((a, b) => (b.los - a.los) || (a.dist - b.dist))[0];\n  if (choice === \"attack\" && enemy) {\n    return `fire_at tank ${enemy.id}; frames 30 until clear,damage,kill`;\n  }\n  if (choice === \"dodge\") {\n    // move perpendicular to the closest incoming shot\n    const s = st.projectiles.filter(s => !s.yours)\n      .sort((a, b) => a.eta - b.eta)[0];\n    const dir = s && Math.abs(s.vel[0]) > Math.abs(s.vel[1])\n      ? (p.pos[1] > s.pos[1] ? \"+down\" : \"+up\")\n      : (p.pos[0] > (s ? s.pos[0] : 0) ? \"+right\" : \"+left\");\n    return `input ${dir}; frames 20 until damage; input stop`;\n  }\n  const powerup = st.powerups.filter(u => u.available)\n    .sort((a, b) => a.dist - b.dist)[0];\n  if (choice === \"powerup\" && powerup) {\n    return `goto ${powerup.pos[0]} ${powerup.pos[1]} 180`;\n  }\n  return enemy ? `goto ${enemy.pos[0]} ${enemy.pos[1]} 90` : null;\n}\nconst log = [];\nfor (let step = 0; step < 30; step++) {\n  const st = JSON.parse(await tank(\"state\"));\n  if (st.state !== \"playing\") break;\n  const threats = st.projectiles\n    .filter(s => !s.yours && s.miss_dist < 1.5 && s.eta < 1.5)\n    .map(s => `incoming shot dist ${s.dist} eta ${s.eta}s`)\n    .join(\"\\n\") || \"no incoming shots\";\n  const r = await models.classify(jev, {\n    state: { map: await tank(\"view 8\"), threats, hp: st.player.hp },\n    questions,\n  });\n  if (r.stopReason !== \"stop\") {\n    log.push(`#${step} classifier error: ${r.errorMessage}`);\n    break;\n  }\n  const choice = r.answers.action.choice;\n  const cmd = commandFor(choice, st);\n  if (!cmd) break;\n  log.push(`#${step} hp=${st.player.hp} ${choice} -> ${await tank(cmd)}`);\n}\nreturn log.join(\"\\n\");\n```\nLastly, Codemode obviously is great for calling MCP servers. And because we do not actually expose any of the MCP tools to the LLM, the agent first uses provided APIs to issue a tool search within Codemode to discover what it might be able to do with the connected servers. This form of progressive discovery makes the whole MCP business work well enough for a lot of use cases today.\n\nHere for instance you can see the agent reach for the Sentry MCP straight away, even without discovering the tools, presumably because it has learned during the RL process already about what the Sentry MCP looks like. But it learns from what we inject into the system prompt, that the Sentry server is available to begin with. It’s not completely guessing here.\n\n```\nconst orgs = await tools.mcp__sentry__find_organizations({});\nconst { organizations } = orgs.structuredContent;\nconst results = await Promise.allSettled(organizations.map(org =>\n  tools.mcp__sentry__find_projects({\n    organizationSlug: org.slug,\n    regionUrl: org.regionUrl,\n  })\n));\nreturn organizations.map((org, i) => {\n  const r = results[i];\n  if (r.status !== \"fulfilled\") return { org: org.slug, error: String(r.reason) };\n  if (r.value.isError) return { org: org.slug, error: r.value.content };\n  return {\n    org: org.slug,\n    projects: r.value.structuredContent.projects.map(p => p.slug),\n  };\n});\n```\nI really don’t want to talk too much about MCP here, but MCP is in fact a protocol that greatly benefits from Codemode. The problem in parts is that MCP in practice often targets harnesses that do not (yet?) use Codemode. But the tide is shifting. In the meantime, a temporary crutch has been to do what Cloudflare did, and do Codemode within the MCP server. But now we have Codemode in Codemode which is pretty bad. It means double JSON escaping, easy for smaller models to get confused by and the inner code cannot call the outer tools. So if you for instance use the Cloudflare MCP servers in Pi, the agent needs to write JavaScript and funnel it through more JavaScript. This is really not optimal, but it’s also understandable that this is happening:\n\n```\nconst accRes = await tools.mcp__cloudflare__execute({\n  code: `async () => {\n    const r = await cloudflare.request({ method: \"GET\", path: \"/accounts\" });\n    return r.result.map(a => ({ id: a.id, name: a.name }));\n  }`,\n});\nconst accounts = JSON.parse(accRes.content.map(c => c.text).join(\"\"));\nconst out = [];\nfor (const account of accounts) {\n  const r = await tools.mcp__cloudflare__execute({\n    account_id: account.id,\n    code: `async () => {\n      const r = await cloudflare.request({\n        method: \"GET\",\n        path: \\`/accounts/\\${accountId}/workers/scripts\\`,\n      });\n      return r.result.map(s => ({ id: s.id, modified: s.modified_on }));\n    }`,\n  });\n  out.push({ account: account.name, workers: r.content.map(c => c.text).join(\"\") });\n}\nreturn out;\n```\nSo to end things off: how well does Codemode work with MCP today? Well … not amazingly well. That’s because MCP servers are not really targeting harnesses that use Codemode yet (though at this point I think most harnesses support it).\n\nFor this to work well some recommendations:\n\n- **Structured content:** Codemode wants calls to return some nicely\nformatted JSON.  So that needs to come back from the server, and many don’t\ndo that yet.  The`outputSchema` system in MCP is great for that.\n- **Consistent results:** an interesting failure case is when an MCP server\ndoes not return consistent data.  For instance because it tries to token\noptimize things depending on how many items are in the result set.  This can\ncause an initial probe with 5 items to succeed, but then fail when the server\nreturns the maximum batch size.\n- **Large binary data:** today MCP does not yet support large binary data\nso quite a few use cases that are really interesting do not work well at all\nyet.  You end up with all kinds of weird workarounds such as pre-signed URLs\nto allow file uploads then to happen through non MCP channels.\n- **Composable tool search:** the MCP server might know better than the MCP\nclient which tool is appropriate for a task.  But there is no good mechanism\ntoday that allows a harness to fan out tool searches across multiple MCP\nservers.  It’s all emergent behavior and it does not scale well to multiple\nactive servers.\n\nSo where does this leave us? Is this a reversal of what I wrote a year ago where I encouraged CLIs? I don’t think so. In fact, the MCP ecosystem from my perspective picked up on exactly what we pointed out a year ago works: code. But Codemode goes beyond MCP in that it can act as a capable mechanism within the harness to express more freedom for the agent.\n\nThere are however also some things that we still need to figure out. For one, durability with Codemode is trickier. We might have to adopt some ideas from durable workflow engines here to snapshot invocations. Or maybe, something like Starlark is a better composition language than JavaScript given its deterministic nature.\n\nImages, binary data and just the inability of this pattern to work with smaller models is also something that needs to be fleshed out. So it’s for sure not a perfect solution yet, but it’s quite a useful pattern that I expect us to leverage more.\n","body_html":"<p>More than a year ago I wrote a few posts here that recommended people not to\nload custom tools into their context (or\n<a href=\"https://en.wikipedia.org/wiki/Model_Context_Protocol\" rel=\"nofollow ugc noopener\">MCP</a> servers) but to just\nuse more scripts.  Most importantly I wrote that <a href=\"https://lucumr.pocoo.org/2025/7/3/tools/\" rel=\"nofollow ugc noopener\">Code Is All You\nNeed</a> and I wrote about that <a href=\"https://lucumr.pocoo.org/2025/8/18/code-mcps/\" rel=\"nofollow ugc noopener\">MCP needs\ncode</a>.  With Pi 1.0 we now added MCP support via Codemode\nwhich in some ways is a long time coming, but then also maybe somewhat\nsurprising to some.  So I want to share some updated thoughts on this blog on\nwhat this all means.</p>\n<p>When a harness like Pi provides tools for an LLM to call, it does so by\nsupplying some tool definitions which then translate into some token structure\non the server side.  Whether a model is encouraged to call a tool is the result\nof the reinforcement learning process.  Something <a href=\"https://lucumr.pocoo.org/2026/7/4/better-models-worse-tools/\" rel=\"nofollow ugc noopener\">I wrote about\nbefore</a> if you want to learn more.</p>\n<p>One of the reasons we strongly lean towards CLI and bash is because it allows\neasy composition of calls, and because the model also learns how the file system\nworks when it’s trained.  So when it invokes a tool like <code>echo foo &gt; /tmp/test.txt</code> the model also learns that after that tool call, there is now a\nfile called <code>test.txt</code> in <code>/tmp</code>.</p>\n<p>However bash has one fundamental limitation which is that it can only compose programs that run. And there are some things, which are not programs, but native tools to the LLM and they sort of have to be.</p>\n<p>The most obvious example here is <code>read</code> or <code>view_image</code>. If a multimodal model\nneeds to read an image, it cannot use <code>cat</code> for that because the harness needs\nto inject the actual image payload into the protocol of the LLM.</p>\n<p>Another quite vivid example are sub agents. In order to spawn and orchestrate sub agents, it’s tricky to avoid tools that are provided by the harness. While in theory the agent could provide a CLI tool that talks to the outer harness via environment variables and Unix sockets, it’s a rather crude process. It however has another issue, and that is where the code runs.</p>\n<p>To better understand that, it’s important to think a bit more about where all the\nbits and pieces run.  There really usually are two different systems involved.\nThe first is the brain, the harness: it runs on one machine. It’s trusted.  The\nsecond is <em>often</em> the same machine, but it’s really where the tools are\nexecuting: the hands.  In Pi we now call this the execution environment, but you\ncan think of it as the target of all the operations.</p>\n<p>Crucially what is important for us, is that there is a dividing line between the harness brain and the target environment that runs bash and executes the tools.</p>\n<p>And splitting this in half has some really important consequences.  For a start\nit means that they are running on different file systems and they have different\nlevels of trust.  If you for instance use a sandboxing solution <a href=\"https://earendil-works.github.io/gondolin/\" rel=\"nofollow ugc noopener\">like\nGondolin</a> your bash stuff will be\nsandboxed just fine, but the harness itself will not be.</p>\n<p>Which brings us to what Codemode really does: it’s a way for the LLM to express and orchestrate complex operations on the harness side, but not the execution environment side. Codemode runs in the harness, in its own sandbox. In case of Pi it’s running in QuickJS within a WASM runtime with intentional limitations: no network, no file system, no timers, limited RAM. The only way is to call more tools. You could also imagine that Codemode could run Scheme or some other language as well.</p>\n<p>If you are not familiar with Codemode, it’s basically just a way to issue\ntool calls from within some language, in our case JavaScript.  That allows you\nto compose those calls without necessarily going through the LLM’s context.\nCredit for naming goes to our friends at Cloudflare <a href=\"https://blog.cloudflare.com/code-mode/\" rel=\"nofollow ugc noopener\">who coined\nit</a>.</p>\n<p>For instance if you issue a bash call as a regular tool call in the LLM, then we only throw the trailing 2000 lines into the context and if the agent wants more, it needs to look at the overflow file itself. If however the agent issues that invocation via Codemode, then the Codemode side gets larger outputs sent structurally.</p>\n<p>Most importantly, because Codemode is JavaScript the agent can express concurrent operations and basic workflows. A common way in which you see agents now use this, is to first probe at 5-10 items from some tool response to see what it looks like, and to then write a Codemode script that processes the next n items.</p>\n<p>Codemode also allows you to throw state into the transcript! That means that one Codemode invocation can stash away data, that the next call in the session can load again. And remember: this is on the harness host, not the sandbox.</p>\n<p>In case of Pi, Codemode also allows you to issue calls that naturally do not make any sense in Pi’s traditional interface. For instance if you want to generate images with an image model or you want to classify some text with a one shot classifier model, those Pi APIs are exposed via Codemode, but not via regular tools where they would just waste context.</p>\n<p>So now that we talked a bunch about it, it’s probably worth being a bit more explicit about it. Let’s walk ourselves through some invocations of Codemode of recent Pi sessions of mine. Note that none of this code is human written. It’s from real sessions of Pi, just re-indented for your viewing pleasure. The agent starts using Codemode automatically either because it’s a task where the model already naturally picks up that tool, or because a user asked it to.</p>\n<p>Note that Codemode is by default only enabled in Pi when MCP is enabled, but you\ncan turn it on with <code>&quot;defaultTools&quot;: [&quot;+codemode&quot;]</code> in the settings.  Just ask\nPi to enable it for you.</p>\n<p>Let’s start simple with image generation. Image generation is a feature that Pi supports in the AI SDK core, but it’s not a tool that the agent can use. In the past the only way to use image models has been to write a bespoke extension or to have the agent run node itself and use the internal image APIs. However because we expose quite a few of the internal model APIs within Codemode, it means that the agent can use it:</p>\n<pre><code>const [painter] = await models.getAvailableOfType(&quot;image&quot;);\nconst result = await models.generateImages(painter, {\n  input: [{ type: &quot;text&quot;, text: &quot;A cute little puppy sitting on a grassy &quot; +\n    &quot;lawn, soft natural light, photorealistic&quot; }],\n});\nif (result.stopReason !== &quot;stop&quot;) return result.errorMessage;\nfor (const block of result.output) {\n  if (block.type === &quot;image&quot;) image(block);\n  else text(block.text);\n}</code></pre>\n<p>Note that the call to <code>image()</code> sends the image back as image content to the\nLLM.  On the harness side it feeds it directly into both the agent, as well as\nonto disk as a temporary artifact in case the agent wants to be able to pass\nthat image back to bash.</p>\n<p>Similar things apply to classifier models such as <a href=\"https://typesafe.ai/\" rel=\"nofollow ugc noopener\">Jev</a>.\nThey also do not fit well into the workflows of an agent through the typical\ntools.  But rather than making a bespoke tool available, Codemode just allows\nthe agent to reach into the AI SDK and invoke those directly.  Here you can see\nhow Jev is used to mass process GitHub issues for a quick sentiment analysis:</p>\n<pre><code>const jev = await models.getModelOfType(&quot;classifier&quot;, &quot;typesafe&quot;, &quot;jev-latest&quot;);\nconst r = await tools.bash({\n  command: &quot;gh issue list --state open --limit 100 &quot; +\n    &quot;--json number,title,body,comments&quot;,\n});\nconst issues = JSON.parse(r.output);\nconst results = await Promise.all(issues.map(async (issue) =&gt; {\n  const res = await models.classify(jev, {\n    state: {\n      title: issue.title,\n      body: (issue.body || &quot;&quot;).slice(0, 4000),\n      comments: issue.comments.slice(-5).map(c =&gt; c.body.slice(0, 800)),\n    },\n    questions: {\n      sentiment: {\n        type: &quot;choice&quot;,\n        instructions: &quot;What is the overall sentiment of the author towards pi?&quot;,\n        criteria: {\n          positive: &quot;Appreciative, happy, constructive praise&quot;,\n          neutral: &quot;Matter-of-fact report or request without emotion&quot;,\n          negative: &quot;Frustrated, annoyed, upset, or angry&quot;,\n        },\n      },\n      frustration: {\n        type: &quot;score&quot;,\n        instructions: &quot;How frustrated is the reporter?&quot;,\n        criteria: [&quot;not at all&quot;, &quot;mildly&quot;, &quot;clearly frustrated&quot;, &quot;very angry&quot;],\n      },\n      kind: {\n        type: &quot;choice&quot;,\n        instructions: &quot;What kind of issue is this?&quot;,\n        criteria: {\n          bug: &quot;Bug report or regression&quot;,\n          feature: &quot;Feature request or enhancement&quot;,\n          question: &quot;Question or support request&quot;,\n          other: &quot;Docs, discussion, meta, spam&quot;,\n        },\n      },\n    },\n  });\n  if (res.stopReason !== &quot;stop&quot;) {\n    return { n: issue.number, title: issue.title, error: res.errorMessage };\n  }\n  return { n: issue.number, title: issue.title, ...res.answers };\n}));\nstore(&quot;sentiment_results&quot;, results);\nreturn results\n  .filter(r =&gt; !r.error)\n  .sort((a, b) =&gt; b.frustration.score - a.frustration.score)\n  .slice(0, 12)\n  .map(r =&gt; `#${r.n} ${r.frustration.score.toFixed(2)} [${r.kind.choice}] ${r.title}`);</code></pre>\n<p>Note how in that above example we also call <code>store()</code> which dumps the result of\nthat execution into the session transcript.  A future invocation of Codemode can\nthus read back that result if it wants to.</p>\n<p>The <code>Promise.all</code> here is fine, because Pi limits the total number of concurrent\ntool executions itself to four and maintains a queue for the rest.</p>\n<p>A more adventurous example is to use Jev to drive a game engine for debugging purposes:</p>\n<p>Here it knows about my <code>tankctl</code> command and it built itself quickly a minimal\nharness around it to drive a game loop to assist a user with debugging a\nproblem. Note how it built a 30 step loop in which each step goes back to both\nthe game engine to get a text dump of what’s going on, and then to Jev to\ndetermine what to do next:</p>\n<pre><code>const jev = await models.getModelOfType(&quot;classifier&quot;, &quot;typesafe&quot;, &quot;jev-latest&quot;);\nconst tank = async (cmd) =&gt;\n  (await tools.bash({ command: `tools/tankctl &quot;${cmd}&quot;` })).output;\nawait tank(&quot;start --map assets/maps/night_arena.map&quot;);\nconst questions = {\n  action: {\n    type: &quot;choice&quot;,\n    instructions: &quot;You control the tank &#39;@&#39; in a top-down tank game. &quot; +\n      &quot;Choose the best next action.&quot;,\n    criteria: {\n      attack: &quot;an enemy has line of sight to you and you can fire at it&quot;,\n      approach: &quot;no enemy has line of sight; drive toward the nearest enemy&quot;,\n      dodge: &quot;an enemy shot is heading at you and will hit soon&quot;,\n      powerup: &quot;a powerup is close and no enemy threatens you&quot;,\n    },\n  },\n};\nfunction commandFor(choice, st) {\n  const p = st.player;\n  const enemy = st.enemies.filter(e =&gt; !e.dead)\n    .sort((a, b) =&gt; (b.los - a.los) || (a.dist - b.dist))[0];\n  if (choice === &quot;attack&quot; &amp;&amp; enemy) {\n    return `fire_at tank ${enemy.id}; frames 30 until clear,damage,kill`;\n  }\n  if (choice === &quot;dodge&quot;) {\n    // move perpendicular to the closest incoming shot\n    const s = st.projectiles.filter(s =&gt; !s.yours)\n      .sort((a, b) =&gt; a.eta - b.eta)[0];\n    const dir = s &amp;&amp; Math.abs(s.vel[0]) &gt; Math.abs(s.vel[1])\n      ? (p.pos[1] &gt; s.pos[1] ? &quot;+down&quot; : &quot;+up&quot;)\n      : (p.pos[0] &gt; (s ? s.pos[0] : 0) ? &quot;+right&quot; : &quot;+left&quot;);\n    return `input ${dir}; frames 20 until damage; input stop`;\n  }\n  const powerup = st.powerups.filter(u =&gt; u.available)\n    .sort((a, b) =&gt; a.dist - b.dist)[0];\n  if (choice === &quot;powerup&quot; &amp;&amp; powerup) {\n    return `goto ${powerup.pos[0]} ${powerup.pos[1]} 180`;\n  }\n  return enemy ? `goto ${enemy.pos[0]} ${enemy.pos[1]} 90` : null;\n}\nconst log = [];\nfor (let step = 0; step &lt; 30; step++) {\n  const st = JSON.parse(await tank(&quot;state&quot;));\n  if (st.state !== &quot;playing&quot;) break;\n  const threats = st.projectiles\n    .filter(s =&gt; !s.yours &amp;&amp; s.miss_dist &lt; 1.5 &amp;&amp; s.eta &lt; 1.5)\n    .map(s =&gt; `incoming shot dist ${s.dist} eta ${s.eta}s`)\n    .join(&quot;\\n&quot;) || &quot;no incoming shots&quot;;\n  const r = await models.classify(jev, {\n    state: { map: await tank(&quot;view 8&quot;), threats, hp: st.player.hp },\n    questions,\n  });\n  if (r.stopReason !== &quot;stop&quot;) {\n    log.push(`#${step} classifier error: ${r.errorMessage}`);\n    break;\n  }\n  const choice = r.answers.action.choice;\n  const cmd = commandFor(choice, st);\n  if (!cmd) break;\n  log.push(`#${step} hp=${st.player.hp} ${choice} -&gt; ${await tank(cmd)}`);\n}\nreturn log.join(&quot;\\n&quot;);</code></pre>\n<p>Lastly, Codemode obviously is great for calling MCP servers. And because we do not actually expose any of the MCP tools to the LLM, the agent first uses provided APIs to issue a tool search within Codemode to discover what it might be able to do with the connected servers. This form of progressive discovery makes the whole MCP business work well enough for a lot of use cases today.</p>\n<p>Here for instance you can see the agent reach for the Sentry MCP straight away, even without discovering the tools, presumably because it has learned during the RL process already about what the Sentry MCP looks like. But it learns from what we inject into the system prompt, that the Sentry server is available to begin with. It’s not completely guessing here.</p>\n<pre><code>const orgs = await tools.mcp__sentry__find_organizations({});\nconst { organizations } = orgs.structuredContent;\nconst results = await Promise.allSettled(organizations.map(org =&gt;\n  tools.mcp__sentry__find_projects({\n    organizationSlug: org.slug,\n    regionUrl: org.regionUrl,\n  })\n));\nreturn organizations.map((org, i) =&gt; {\n  const r = results[i];\n  if (r.status !== &quot;fulfilled&quot;) return { org: org.slug, error: String(r.reason) };\n  if (r.value.isError) return { org: org.slug, error: r.value.content };\n  return {\n    org: org.slug,\n    projects: r.value.structuredContent.projects.map(p =&gt; p.slug),\n  };\n});</code></pre>\n<p>I really don’t want to talk too much about MCP here, but MCP is in fact a protocol that greatly benefits from Codemode. The problem in parts is that MCP in practice often targets harnesses that do not (yet?) use Codemode. But the tide is shifting. In the meantime, a temporary crutch has been to do what Cloudflare did, and do Codemode within the MCP server. But now we have Codemode in Codemode which is pretty bad. It means double JSON escaping, easy for smaller models to get confused by and the inner code cannot call the outer tools. So if you for instance use the Cloudflare MCP servers in Pi, the agent needs to write JavaScript and funnel it through more JavaScript. This is really not optimal, but it’s also understandable that this is happening:</p>\n<pre><code>const accRes = await tools.mcp__cloudflare__execute({\n  code: `async () =&gt; {\n    const r = await cloudflare.request({ method: &quot;GET&quot;, path: &quot;/accounts&quot; });\n    return r.result.map(a =&gt; ({ id: a.id, name: a.name }));\n  }`,\n});\nconst accounts = JSON.parse(accRes.content.map(c =&gt; c.text).join(&quot;&quot;));\nconst out = [];\nfor (const account of accounts) {\n  const r = await tools.mcp__cloudflare__execute({\n    account_id: account.id,\n    code: `async () =&gt; {\n      const r = await cloudflare.request({\n        method: &quot;GET&quot;,\n        path: \\`/accounts/\\${accountId}/workers/scripts\\`,\n      });\n      return r.result.map(s =&gt; ({ id: s.id, modified: s.modified_on }));\n    }`,\n  });\n  out.push({ account: account.name, workers: r.content.map(c =&gt; c.text).join(&quot;&quot;) });\n}\nreturn out;</code></pre>\n<p>So to end things off: how well does Codemode work with MCP today? Well … not amazingly well. That’s because MCP servers are not really targeting harnesses that use Codemode yet (though at this point I think most harnesses support it).</p>\n<p>For this to work well some recommendations:</p>\n<ul><li><p><strong>Structured content:</strong> Codemode wants calls to return some nicely</p><p>formatted JSON.  So that needs to come back from the server, and many don’t\ndo that yet.  The<code>outputSchema</code> system in MCP is great for that.</p></li><li><p><strong>Consistent results:</strong> an interesting failure case is when an MCP server</p><p>does not return consistent data.  For instance because it tries to token\noptimize things depending on how many items are in the result set.  This can\ncause an initial probe with 5 items to succeed, but then fail when the server\nreturns the maximum batch size.</p></li><li><p><strong>Large binary data:</strong> today MCP does not yet support large binary data</p><p>so quite a few use cases that are really interesting do not work well at all\nyet.  You end up with all kinds of weird workarounds such as pre-signed URLs\nto allow file uploads then to happen through non MCP channels.</p></li><li><p><strong>Composable tool search:</strong> the MCP server might know better than the MCP</p><p>client which tool is appropriate for a task.  But there is no good mechanism\ntoday that allows a harness to fan out tool searches across multiple MCP\nservers.  It’s all emergent behavior and it does not scale well to multiple\nactive servers.</p></li></ul>\n<p>So where does this leave us? Is this a reversal of what I wrote a year ago where I encouraged CLIs? I don’t think so. In fact, the MCP ecosystem from my perspective picked up on exactly what we pointed out a year ago works: code. But Codemode goes beyond MCP in that it can act as a capable mechanism within the harness to express more freedom for the agent.</p>\n<p>There are however also some things that we still need to figure out. For one, durability with Codemode is trickier. We might have to adopt some ideas from durable workflow engines here to snapshot invocations. Or maybe, something like Starlark is a better composition language than JavaScript given its deterministic nature.</p>\n<p>Images, binary data and just the inability of this pattern to work with smaller models is also something that needs to be fleshed out. So it’s for sure not a perfect solution yet, but it’s quite a useful pattern that I expect us to leverage more.</p>","headings":[]}}