{"article":{"slug":"llmman-launch-dsh-run-deepseek-harness-on-any-local-or-hosted-model","title":"llmman launch dsh: Run DeepSeek Harness on any local or hosted model","subtitle":null,"summary":"DeepSeek Harness treats the model as a plugin. llmman runs any model on your own hardware, in one command. An agent harness is a loop around your model that takes your task, calls a model, runs tools (such as shell commands and file edits), provides results, and repeats.","content_type":"tutorial","language":"en","canonical_url":"https://llmmanorg.github.io/blog/launch-dsh/","author":{"name":"Hassan Bahati","url":null,"person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"llmman","url":"https://llmmanorg.github.io/","listing_slug":null,"listing":null},"topics":[{"name":"AI","slug":"ai","url":"https://listedarticles.com/topics/ai"},{"name":"LLMs","slug":"llms","url":"https://listedarticles.com/topics/llms"},{"name":"Open Source","slug":"open-source","url":"https://listedarticles.com/topics/open-source"},{"name":"AI Agents","slug":"ai-agents","url":"https://listedarticles.com/topics/ai-agents"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":1111,"reading_minutes":5,"published_at":"2026-09-14T12:00:00.000Z","added_at":"2026-09-24T06:18:29.933Z","updated_at":"2026-09-24T06:18:29.933Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":true},"profile_url":"https://listedarticles.com/articles/llmman-launch-dsh-run-deepseek-harness-on-any-local-or-hosted-model","markdown_url":"https://listedarticles.com/articles/llmman-launch-dsh-run-deepseek-harness-on-any-local-or-hosted-model.md","example":false,"citation":"Hassan Bahati, llmman. \"llmman launch dsh: Run DeepSeek Harness on any local or hosted model.\" 14 Sept 2026. https://llmmanorg.github.io/blog/launch-dsh/ (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://llmmanorg.github.io/blog/launch-dsh/"},"body_markdown":"# llmman launch dsh: Run DeepSeek Harness on any local or hosted model\n\nDeepSeek Harness treats the model as a plugin. llmman runs any model on your own hardware, in one command.\n\nAn agent harness is a loop around your model that takes your task, calls a model, runs tools (such as shell commands and file edits), provides results, and repeats. Claude Code, Codex, and OpenCode are all harnesses. DeepSeek Harness is DeepSeek’s.\n\nUnlike Claude Code, DeepSeek Harness is fully customizable. It builds on the idea that everything is a plugin including the model, tools, memory and even the agent loop itself. With DeepSeek Harness, you can build a custom AI agent entirely from scratch, run it locally with any AI model and even; call Claude Code and Codex as sub-agents from inside of it.\n\n## Why llmman?\n\nAn agent CLI normally talks to one vendor’s API, and your prompts and code go with it. `llmman` puts a server in between. DeepSeek Harness (`dsh`) always talks to `127.0.0.1:17434`, and you choose what answers: by default a model running on your own machine, or, with `--provider`, a hosted model from OpenAI, Anthropic, OpenRouter or any of the other providers `llmman` knows about. Same command, same `dsh` configuration either way.\n\nLaunch DeepSeek Harness with llmman and you get the following out of the box:\n\n- **Local by default:** Prompts, file contents and diffs stay on your machine. Once the weights are pulled, the loop works offline.\n- **Hosted when you ask:**`--provider` changes where the daemon forwards a request, not what`dsh` talks to, so the agent config is the same for local and hosted models, and you can switch between them without touching it.\n- **The model is yours to move:**`llmman` stores local models as OCI artifacts. Pull one that is already packaged that way from Docker Hub, or have`llmman` pull the weights from Hugging Face and package them for you. Either way, you can then push it to a registry you control, or copy it into a network with no internet at all.\n\n## The setup\n\n```\nllmman launch dsh --model <model-name>    # eg. gemma4:12b\n```\nOne command does four things: it starts `llmman serve` if there isnt a running server, pulls the model and loads it, writes the configuration `dsh` expects, and hands over to dsh’s `web` profile. Short names work here the way they do everywhere else in `llmman`, so `gemma4:12b` resolves to `docker.io/ai/gemma4:12b`.\n\n`--model` is required here, unlike most integrations. `dsh` has no default model of its own, and leaving it empty writes the literal string `default` into its settings; which fails at the first request rather than at the command you typed.\n\n## Executing a single task\n\nWhen you run `llmman launch dsh --model <model-name>`, dsh’s `web` profile boots a server and answers in a browser, so it has nothing to print to your terminal. When you want one answer and no browser, pass dsh’s headless profile instead. Everything after `--` goes to dsh’s own CLI:\n\n```\nllmman launch dsh --model gemma4:12b \\\n-- --profile headless \"Explain what git rebase does in one sentence\"\n```\n`llmman` defaults to a `web` profile. Providing `--profile headless` overrides the default profile and uses the `headless` profile.\n\n## What llmman configures for you\n\nWhen you launch `dsh` with `llmman`, `llmman` sets up two files under `~/.config/llmman/launch/dsh`. The first registers the daemon as a provider and picks the model:\n\n```\nagent-default-model:\n  provider: llmman\n  model: \"docker.io/ai/gemma4:12b\"\nllm-pi-ai:\n  providers:\n    llmman:\n      displayName: llmman\n      apiKeyEnv: LLMMAN_API_KEY\n      api: openai-completions\n      baseURL: \"http://127.0.0.1:17434/v1\"\n      models:\n        - id: \"docker.io/ai/gemma4:12b\"\n          name: \"docker.io/ai/gemma4:12b\"\n          input: [text, image]\n```\nThe second is the patch that points `dsh` at the first. Both are rewritten on every launch, so the model `dsh` talks to is always the one you just named.\n\nIn that configuration:\n\n- **`apiKeyEnv`** names an environment variable rather than holding a key, so the credential travels in dsh’s environment and never lands on disk. That is a deliberate security choice: a config file that never holds a credential cannot leak one.\n- **`input`** is populated from local model metadata. llmman asks the daemon what a local model can do, so dsh offers image attachments when a locally served model supports them. Hosted-provider launches currently advertise text input only.\n\n## Hosted models\n\nSometimes you don’t want the model on your machine at all. The weights may not fit on your disk, your hardware may not run them at a useful speed, or the task may need a bigger model than you can run. For those cases, `llmman` can send `dsh`’s requests to a model someone else runs.\n\nThis is different from pulling a model from Docker Hub or Hugging Face. The registries only store weights; once pulled, the model runs on your machine. A hosted provider runs the model for you, which means your prompts, file contents and diffs go to that provider. That is the “unless you say” from earlier: nothing leaves your machine until you pass `--provider`.\n\n### Run dsh on a hosted provider\n\nExport the provider’s key and add `--provider` to the same command:\n\n```\nexport OPENROUTER_API_KEY=...\nllmman launch dsh --provider openrouter --model google/gemma-4-31b-it\n```\nThe key is read from your environment and sent with each request. It is never written into dsh’s configuration or anywhere else on disk.\n\n### Find a provider and a model\n\nThe provider list comes from models.dev at runtime, so a provider that appears there works without waiting for an llmman release. There are around 180, including OpenAI, Anthropic, DeepSeek, Groq, Mistral, Together and Fireworks. `llmman providers` prints each one with the environment variable it expects, whether yours is set, and how many models it serves:\n\n```\n$ llmman providers | grep -E '^(PROVIDER|openrouter)'\nPROVIDER      NAME          API KEY               KEY    MODELS\nopenrouter    OpenRouter    OPENROUTER_API_KEY    set    369\n```\nTo see those models, and the exact name to pass to `--model`, list them:\n\n```\nllmman list --provider openrouter\n```\n### Use your own server\n\nA server the catalog has never heard of works too: vLLM on a GPU machine down the hall, LM Studio on a laptop, or a proxy in front of OpenAI. Give it a `base_url` in `llmman.conf` and it takes the same flag:\n\n```\nllmman config set providers.local.base_url http://192.168.1.50:8000/v1\nllmman launch dsh --provider local --model google/gemma-4-26b-a4b-it\n```\nServers on your own network often need no key. If yours does, set `providers.local.api_key_env` to the name of the environment variable that holds it. If the server speaks the Anthropic API rather than OpenAI’s, set `providers.local.wire` to `anthropic`.\n\nIn every case, `dsh` still talks to `llmman serve` on `127.0.0.1:17434`, with the same generated config. `--provider` only changes where the daemon sends the request.","body_html":"<h1 id=\"llmman-launch-dsh-run-deepseek-harness-on-any-local-or-hosted-mo\">llmman launch dsh: Run DeepSeek Harness on any local or hosted model</h1>\n<p>DeepSeek Harness treats the model as a plugin. llmman runs any model on your own hardware, in one command.</p>\n<p>An agent harness is a loop around your model that takes your task, calls a model, runs tools (such as shell commands and file edits), provides results, and repeats. Claude Code, Codex, and OpenCode are all harnesses. DeepSeek Harness is DeepSeek’s.</p>\n<p>Unlike Claude Code, DeepSeek Harness is fully customizable. It builds on the idea that everything is a plugin including the model, tools, memory and even the agent loop itself. With DeepSeek Harness, you can build a custom AI agent entirely from scratch, run it locally with any AI model and even; call Claude Code and Codex as sub-agents from inside of it.</p>\n<h2 id=\"why-llmman\">Why llmman?</h2>\n<p>An agent CLI normally talks to one vendor’s API, and your prompts and code go with it. <code>llmman</code> puts a server in between. DeepSeek Harness (<code>dsh</code>) always talks to <code>127.0.0.1:17434</code>, and you choose what answers: by default a model running on your own machine, or, with <code>--provider</code>, a hosted model from OpenAI, Anthropic, OpenRouter or any of the other providers <code>llmman</code> knows about. Same command, same <code>dsh</code> configuration either way.</p>\n<p>Launch DeepSeek Harness with llmman and you get the following out of the box:</p>\n<ul><li><strong>Local by default:</strong> Prompts, file contents and diffs stay on your machine. Once the weights are pulled, the loop works offline.</li><li><strong>Hosted when you ask:</strong><code>--provider</code> changes where the daemon forwards a request, not what<code>dsh</code> talks to, so the agent config is the same for local and hosted models, and you can switch between them without touching it.</li><li><strong>The model is yours to move:</strong><code>llmman</code> stores local models as OCI artifacts. Pull one that is already packaged that way from Docker Hub, or have<code>llmman</code> pull the weights from Hugging Face and package them for you. Either way, you can then push it to a registry you control, or copy it into a network with no internet at all.</li></ul>\n<h2 id=\"the-setup\">The setup</h2>\n<pre><code>llmman launch dsh --model &lt;model-name&gt;    # eg. gemma4:12b</code></pre>\n<p>One command does four things: it starts <code>llmman serve</code> if there isnt a running server, pulls the model and loads it, writes the configuration <code>dsh</code> expects, and hands over to dsh’s <code>web</code> profile. Short names work here the way they do everywhere else in <code>llmman</code>, so <code>gemma4:12b</code> resolves to <code>docker.io/ai/gemma4:12b</code>.</p>\n<p><code>--model</code> is required here, unlike most integrations. <code>dsh</code> has no default model of its own, and leaving it empty writes the literal string <code>default</code> into its settings; which fails at the first request rather than at the command you typed.</p>\n<h2 id=\"executing-a-single-task\">Executing a single task</h2>\n<p>When you run <code>llmman launch dsh --model &lt;model-name&gt;</code>, dsh’s <code>web</code> profile boots a server and answers in a browser, so it has nothing to print to your terminal. When you want one answer and no browser, pass dsh’s headless profile instead. Everything after <code>--</code> goes to dsh’s own CLI:</p>\n<pre><code>llmman launch dsh --model gemma4:12b \\\n-- --profile headless &quot;Explain what git rebase does in one sentence&quot;</code></pre>\n<p><code>llmman</code> defaults to a <code>web</code> profile. Providing <code>--profile headless</code> overrides the default profile and uses the <code>headless</code> profile.</p>\n<h2 id=\"what-llmman-configures-for-you\">What llmman configures for you</h2>\n<p>When you launch <code>dsh</code> with <code>llmman</code>, <code>llmman</code> sets up two files under <code>~/.config/llmman/launch/dsh</code>. The first registers the daemon as a provider and picks the model:</p>\n<pre><code>agent-default-model:\n  provider: llmman\n  model: &quot;docker.io/ai/gemma4:12b&quot;\nllm-pi-ai:\n  providers:\n    llmman:\n      displayName: llmman\n      apiKeyEnv: LLMMAN_API_KEY\n      api: openai-completions\n      baseURL: &quot;http://127.0.0.1:17434/v1&quot;\n      models:\n        - id: &quot;docker.io/ai/gemma4:12b&quot;\n          name: &quot;docker.io/ai/gemma4:12b&quot;\n          input: [text, image]</code></pre>\n<p>The second is the patch that points <code>dsh</code> at the first. Both are rewritten on every launch, so the model <code>dsh</code> talks to is always the one you just named.</p>\n<p>In that configuration:</p>\n<ul><li><strong><code>apiKeyEnv</code></strong> names an environment variable rather than holding a key, so the credential travels in dsh’s environment and never lands on disk. That is a deliberate security choice: a config file that never holds a credential cannot leak one.</li><li><strong><code>input</code></strong> is populated from local model metadata. llmman asks the daemon what a local model can do, so dsh offers image attachments when a locally served model supports them. Hosted-provider launches currently advertise text input only.</li></ul>\n<h2 id=\"hosted-models\">Hosted models</h2>\n<p>Sometimes you don’t want the model on your machine at all. The weights may not fit on your disk, your hardware may not run them at a useful speed, or the task may need a bigger model than you can run. For those cases, <code>llmman</code> can send <code>dsh</code>’s requests to a model someone else runs.</p>\n<p>This is different from pulling a model from Docker Hub or Hugging Face. The registries only store weights; once pulled, the model runs on your machine. A hosted provider runs the model for you, which means your prompts, file contents and diffs go to that provider. That is the “unless you say” from earlier: nothing leaves your machine until you pass <code>--provider</code>.</p>\n<h3 id=\"run-dsh-on-a-hosted-provider\">Run dsh on a hosted provider</h3>\n<p>Export the provider’s key and add <code>--provider</code> to the same command:</p>\n<pre><code>export OPENROUTER_API_KEY=...\nllmman launch dsh --provider openrouter --model google/gemma-4-31b-it</code></pre>\n<p>The key is read from your environment and sent with each request. It is never written into dsh’s configuration or anywhere else on disk.</p>\n<h3 id=\"find-a-provider-and-a-model\">Find a provider and a model</h3>\n<p>The provider list comes from models.dev at runtime, so a provider that appears there works without waiting for an llmman release. There are around 180, including OpenAI, Anthropic, DeepSeek, Groq, Mistral, Together and Fireworks. <code>llmman providers</code> prints each one with the environment variable it expects, whether yours is set, and how many models it serves:</p>\n<pre><code>$ llmman providers | grep -E &#39;^(PROVIDER|openrouter)&#39;\nPROVIDER      NAME          API KEY               KEY    MODELS\nopenrouter    OpenRouter    OPENROUTER_API_KEY    set    369</code></pre>\n<p>To see those models, and the exact name to pass to <code>--model</code>, list them:</p>\n<pre><code>llmman list --provider openrouter</code></pre>\n<h3 id=\"use-your-own-server\">Use your own server</h3>\n<p>A server the catalog has never heard of works too: vLLM on a GPU machine down the hall, LM Studio on a laptop, or a proxy in front of OpenAI. Give it a <code>base_url</code> in <code>llmman.conf</code> and it takes the same flag:</p>\n<pre><code>llmman config set providers.local.base_url http://192.168.1.50:8000/v1\nllmman launch dsh --provider local --model google/gemma-4-26b-a4b-it</code></pre>\n<p>Servers on your own network often need no key. If yours does, set <code>providers.local.api_key_env</code> to the name of the environment variable that holds it. If the server speaks the Anthropic API rather than OpenAI’s, set <code>providers.local.wire</code> to <code>anthropic</code>.</p>\n<p>In every case, <code>dsh</code> still talks to <code>llmman serve</code> on <code>127.0.0.1:17434</code>, with the same generated config. <code>--provider</code> only changes where the daemon sends the request.</p>","headings":[{"level":1,"text":"llmman launch dsh: Run DeepSeek Harness on any local or hosted model","id":"llmman-launch-dsh-run-deepseek-harness-on-any-local-or-hosted-mo"},{"level":2,"text":"Why llmman?","id":"why-llmman"},{"level":2,"text":"The setup","id":"the-setup"},{"level":2,"text":"Executing a single task","id":"executing-a-single-task"},{"level":2,"text":"What llmman configures for you","id":"what-llmman-configures-for-you"},{"level":2,"text":"Hosted models","id":"hosted-models"},{"level":3,"text":"Run dsh on a hosted provider","id":"run-dsh-on-a-hosted-provider"},{"level":3,"text":"Find a provider and a model","id":"find-a-provider-and-a-model"},{"level":3,"text":"Use your own server","id":"use-your-own-server"}]}}