{"article":{"slug":"self-hosting-llm-models-for-software-development","title":"Self-hosting LLM models for software development","subtitle":null,"summary":"Kévin Maschtaler on running medium-sized open LLMs on AWS Spot EC2 for day-to-day software work—what stacks, costs, and performance looked like versus a personal Claude subscription.","content_type":"blog_post","language":"en","canonical_url":"https://www.kmaschta.me/blog/2026/09/22/self-hosting-llm-models","author":{"name":"Kévin Maschtaler","url":"https://www.kmaschta.me/","person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"kmaschta.me","url":"https://www.kmaschta.me/","listing_slug":null,"listing":null},"topics":[{"name":"AI","slug":"ai","url":"https://listedarticles.com/topics/ai"},{"name":"LLMs","slug":"llms","url":"https://listedarticles.com/topics/llms"},{"name":"Engineering","slug":"engineering","url":"https://listedarticles.com/topics/engineering"},{"name":"Infrastructure","slug":"infrastructure","url":"https://listedarticles.com/topics/infrastructure"},{"name":"Open Source","slug":"open-source","url":"https://listedarticles.com/topics/open-source"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":1719,"reading_minutes":7,"published_at":"2026-09-22T11:00:00.000Z","added_at":"2026-09-22T12:23:38.707Z","updated_at":"2026-09-22T12:23:38.707Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":true},"profile_url":"https://listedarticles.com/articles/self-hosting-llm-models-for-software-development","markdown_url":"https://listedarticles.com/articles/self-hosting-llm-models-for-software-development.md","example":false,"citation":"Kévin Maschtaler, kmaschta.me. \"Self-hosting LLM models for software development.\" 22 Sept 2026. https://www.kmaschta.me/blog/2026/09/22/self-hosting-llm-models (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://www.kmaschta.me/blog/2026/09/22/self-hosting-llm-models"},"body_markdown":"# Self-hosting LLM models for software development\n\n*Sep 22, 2026\n\nI'm self-hosting some open LLM models on AWS Spot EC2 instances for personal use.\n\nIt wasn't my primary goal, but I'm happy enough with my setup to downgrade my Claude subscription from Max ($100/mo) to the Pro plan ($20/mo).\n\nHere is what I learned along the way.\n\nWhat model do we need for software development?\n\nYou don't need the most expensive GPU, but one with enough virtual memory\n\nRenting GPU under their market value: AWS Spot Instances\n\nFrom Claude Code to Hermes\n\nWhat about parallel work and sub-agents?\n\nWhat's next?\n\n## What model do we need for software development?\n\nTake Llama 3.2 (3 billion parameters - 2 GB of memory), a Small Language Model (SLM) that runs on any device, laptop, even your smartphone.\n\n`ollama run llama3.2\npulling manifest\npulling dde5aa3fc5ff: 100% ▕███████████████████▏ 2.0 GB\nverifying sha256 digest\nwriting manifest\nsuccess\n>>> say hi\nHi!\n`\n\nThis model is great to do basic tasks, offline, directly on the device like text summarization, rewriting, etc.\n\nBut to do software development, like you would do with Claude Code or Codex, you need specialized models that can think about complex tasks and work with a very large context, having hundreds of billions of parameters. Or so I thought.\n\nProviders don't reveal the details of their models, but in 2024 it was estimated that `Sonnet 3.5` had ~175 billion and `GPT-4o` had ~200 billion parameters. The latest models are estimated to have trillions.\n\nI wanted to try some open models with the same order of magnitude of capacity as Sonnet, and there are some available on Ollama:\n\nornith 1.5 model trained to \"think\" and self-improve (397b params / 242GB memory)\n\nqwen3-coder meant to produce code (30b params / 19GB memory)\n\nI don't even have 300 GB on any of my disk, so I quickly chose the 10x-reduced version of `ornith 1.5` to see how good it is (34b params / 23 GB of memory).\n\nDo I even need a Large Language Model (LLM), or would some \"medium\" language models do the job for personal use? Some answers on this talk we had at the Tech F'Est 2026.\n\n## You don't need the most expensive GPU, but one with enough virtual memory\n\nOllama models work on a CPU as long as you have enough memory to fit the entire model. The generation speed depends on your computation power.\n\nYou can try to run `ollama run qwen3-coder`, if you have more than 20 GB of memory on your device it will work, but the generation speed will be very slow because of the number of parameters.\n\nWhat you need is a GPU because they can do a lot more parallel computation.\nBut not all GPUs are equal.\n\nYou need to have a GPU with enough virtual memory to fit the whole model in memory + the whole context size you want.\n\nAt home, I have an old gaming PC with a `NVIDIA GeForce RTX 2060` card I bought in 2020. I tried it, but it only has 6 GB of virtual memory, which is not enough.\n\nOllama is smart enough to use the device memory as a failover when GPU VRAM is too small, but it means falling back on the CPU which defeats the purpose of using a GPU.\n\nApple Silicon devices have shared memory between CPU and GPU which is super helpful, because the GPU can read the device memory instead of its own VRAM.\n\nAlso, we could fit more context in the virtual memory with quantization but it may impact the model response quality.\n\nLatest gaming GPUs can go up to 32 GB of VRAM, which is great to run those medium-sized models. But I don't want to buy one, which is only leaving me one solution.\n\n## Renting GPU under their market value: AWS Spot Instances\n\nAWS has a range of GPU instances with different cards available.\n\nThey're also offering a reduced price for unused capacity on a service called AWS Spot Instances.\n\nYou choose the instance you want, they offer a reduced rate but the price is variable so you need to set a maximum price.\nWhen maximum price is reached, the instance is stopped, and you have to request a new one.\n\nTo enable EC2 Spot Instance, you first need to request a quota increase. It's usually approved in a few days max.\n\n`aws service-quotas request-service-quota-increase --service-code ec2 --quota-code L-3819A6DF --desired-value 32`\n\nThe instances with a single GPU are ranging from $0.5/h to $10/h, depending on the availability.\n\nI wrote a helper to spin up a spot instance, select the pricing, choose the Ollama models and spin it up: aws-llm-sandbox.\n\n`./up.sh\n==> my public IP: xx.xx.xx.xx\n==> security group sg-0d910bcfdb06e201e: 22 + 11434 open to xx.xx.xx.xx/32 only\n==> no existing instance - launching a new one\n\nAvailable GPUs (single-GPU instances). Pick by VRAM: the model + its KV cache must fit in it.\n  1) 16 GB VRAM  NVIDIA T4                  2018 gen (g4dn). Cheapest; fits <=14B models at q4. Too small for 30B.\n                   types: g4dn.xlarge, g4dn.2xlarge, g4dn.4xlarge, g4dn.8xlarge, g4dn.16xlarge\n  2) 24 GB VRAM  NVIDIA A10G                2021 gen (g5). Same 24 GB tier as L4, ~same speed for Ollama; a second pool when g6 has no capacity.\n                   types: g5.xlarge, g5.2xlarge, g5.4xlarge, g5.8xlarge, g5.16xlarge\n  3) 24 GB VRAM  NVIDIA L4                  2023 gen (g6, gr6 = same GPU with 2x system RAM). 24 GB: one 30B model at q4 (~18 GB) + KV cache; only one big model resident at a time.\n                   types: g6.xlarge, g6.2xlarge, g6.4xlarge, gr6.4xlarge, g6.8xlarge, gr6.8xlarge, g6.16xlarge\n  4) 48 GB VRAM  NVIDIA L40S                2023 gen (g6e). 48 GB, ~2-3x L4 throughput: two 30B models resident, or 70B at q4, or 30B with a huge context.\n                   types: g6e.xlarge, g6e.2xlarge, g6e.4xlarge, g6e.8xlarge, g6e.16xlarge\n  5) 96 GB VRAM  NVIDIA RTX PRO Server 6000 2025 gen (g7e, Blackwell). 96 GB, often priced like the L40S: 70B at q8, 120B-class MoE at q4, or several 30B models resident.\n                   types: g7e.2xlarge, g7e.4xlarge, g7e.8xlarge\nGPU? [1-5] 4\n==> fetching spot prices for: g6e.xlarge g6e.2xlarge g6e.4xlarge g6e.8xlarge g6e.16xlarge\n==> fetching spot placement scores\n\n     MARKET     TYPE           AZ                     $/h  SCORE  vCPU   RAM\n   1) spot       g6e.4xlarge    eu-central-1b       0.5638  1      16     128 GB\n   2) spot       g6e.4xlarge    eu-central-1a       1.0720  1      16     128 GB\n   3) spot       g6e.2xlarge    eu-central-1b       1.1812  1      8      64 GB\n   4) spot       g6e.xlarge     eu-central-1a       1.8827  3      4      32 GB\n   5) spot       g6e.2xlarge    eu-central-1a       2.0247  1      8      64 GB\n   6) spot       g6e.xlarge     eu-central-1c       2.0259  1      4      32 GB\n   7) spot       g6e.4xlarge    eu-central-1c       2.2143  1      16     128 GB\n   8) spot       g6e.xlarge     eu-central-1b       2.3270  1      4      32 GB\n   9) on-demand  g6e.xlarge     any                 2.3270  -      4      32 GB\n  10) spot       g6e.8xlarge    eu-central-1a       2.6137  1      32     256 GB\n  11) spot       g6e.8xlarge    eu-central-1c       2.7531  1      32     256 GB\n  12) spot       g6e.16xlarge   eu-central-1b       2.7611  1      64     512 GB\n  13) spot       g6e.2xlarge    eu-central-1c       2.8035  1      8      64 GB\n  14) on-demand  g6e.2xlarge    any                 2.8035  -      8      64 GB\n  15) spot       g6e.16xlarge   eu-central-1a       2.9150  1      64     512 GB\n  16) spot       g6e.16xlarge   eu-central-1c       3.2037  2      64     512 GB\n  17) on-demand  g6e.4xlarge    any                 3.7565  -      16     128 GB\n  18) spot       g6e.8xlarge    eu-central-1b       5.6625  1      32     256 GB\n  19) on-demand  g6e.8xlarge    any                 5.6625  -      32     256 GB\n  20) on-demand  g6e.16xlarge   any                 9.4745  -      64     512 GB\n\n  SCORE: spot placement score 1-10 (10 = capacity very likely, 1 = unlikely).\n  spot: market price, varies; AWS stops the instance if it rises above your cap. Needs the spot quota.\n  on-demand: fixed price, never interrupted, no quota issue; AZ chosen by AWS where capacity exists.\nWhich one? [1-20]\n`\n\nI tested various instances, and there are family types (g5, g6, g6e) whose prices range from $0.5/h to $10/h, and I found my sweet spot with a `g6.4xlarge` instance to run my two models.\n\nCost Management = Just-in-time provisioning\n\nA good way to save on the bill is to just stop the instance when you don't use it.\n\nBut when it's turned on, you have virtually no limit of tokens!\n\nHere's my AWS bill from last month, and it was before I used Spot instances.\n\nI'm still using Fable for most complex planning so I'm keeping Claude Pro for now. I tried all that to understand better and check if it was possible, but I'm saving a little bit of money, which is great. Obviously, it's not applicable in all situations.\n\nBut it's totally possible to orchestrate a fleet of low-cost GPU instances, with a routine that track the cheapest available spot instances on multiple regions for your own usage.\n\n## From Claude Code to Hermes\n\nIt is possible to run Ollama models with Claude Code.\nBut I found it easier with Hermes.\n\nFrom Hermes, you can easily add your own custom models and switch provider and model in one command.\n\n`hermes model\n\n  Current model:    ornith-1.5:35b\n  Active provider:  llm-sandbox\n\nCustom OpenAI-compatible endpoint configuration:\n\nAPI base URL [e.g. https://api.example.com/v1]:\n`\n\nThere are a lot of cool features with Hermes that can add up.\n\nLike running multiple Profiles. One profile is a \"Product Manager agent\" running with ornith, a second is a \"Coder agent\" with qwen3, and a last \"Architect\" profile running Fable whose goal is to write the plan and test the feature in the end.\n\nAnd you can then run a Kanban where the \"Product agent\" creates tickets, and some coder agents automatically pick up their tasks and implement them.\n\n## What about parallel work and sub-agents?\n\nOllama can totally handle concurrent requests, even with multiple models, if you have enough memory.\n\nHere is an entire blog post describing how it works and how to configure it (fr).\n\nI'm easily running 3 ornith agents in parallel on a `g7e.4xlarge` instance.\n\n## What's next?\n\nPretty much nothing, I'm happy with my setup for now and it allows me to test a LOT of models available on Ollama with different setups and automations.\n\nI'm surprised how far I've got, I was expecting to hit a performance wall or spend way too much on GPU.\n\nI didn't entirely get rid of the Claude models and their 1M context windows, but it was not the goal either.\n\nBy paying a Claude / Codex subscription we're not even paying for a fraction of the compute required, so it's still very cheap to keep your subscription.\n\nFor programmatic use, burst usage or a certain type of compliance need, though, you'll still find self-hosting useful. And just-in-time provisioning of GPU instances will probably come handy.\n\nThere are plenty of other things to play with, like an Ollama load balancer over multiple instances.\n\nI'm curious how this could fit in some other contexts. Could a well-orchestrated fleet save the bill and footprint of a small/medium company?","body_html":"<h1 id=\"self-hosting-llm-models-for-software-development\">Self-hosting LLM models for software development</h1>\n<p>*Sep 22, 2026</p>\n<p>I&#39;m self-hosting some open LLM models on AWS Spot EC2 instances for personal use.</p>\n<p>It wasn&#39;t my primary goal, but I&#39;m happy enough with my setup to downgrade my Claude subscription from Max ($100/mo) to the Pro plan ($20/mo).</p>\n<p>Here is what I learned along the way.</p>\n<p>What model do we need for software development?</p>\n<p>You don&#39;t need the most expensive GPU, but one with enough virtual memory</p>\n<p>Renting GPU under their market value: AWS Spot Instances</p>\n<p>From Claude Code to Hermes</p>\n<p>What about parallel work and sub-agents?</p>\n<p>What&#39;s next?</p>\n<h2 id=\"what-model-do-we-need-for-software-development\">What model do we need for software development?</h2>\n<p>Take Llama 3.2 (3 billion parameters - 2 GB of memory), a Small Language Model (SLM) that runs on any device, laptop, even your smartphone.</p>\n<p>`ollama run llama3.2\npulling manifest\npulling dde5aa3fc5ff: 100% ▕███████████████████▏ 2.0 GB\nverifying sha256 digest\nwriting manifest\nsuccess</p>\n<blockquote><blockquote><blockquote><p>say hi\nHi!\n`</p></blockquote></blockquote></blockquote>\n<p>This model is great to do basic tasks, offline, directly on the device like text summarization, rewriting, etc.</p>\n<p>But to do software development, like you would do with Claude Code or Codex, you need specialized models that can think about complex tasks and work with a very large context, having hundreds of billions of parameters. Or so I thought.</p>\n<p>Providers don&#39;t reveal the details of their models, but in 2024 it was estimated that <code>Sonnet 3.5</code> had ~175 billion and <code>GPT-4o</code> had ~200 billion parameters. The latest models are estimated to have trillions.</p>\n<p>I wanted to try some open models with the same order of magnitude of capacity as Sonnet, and there are some available on Ollama:</p>\n<p>ornith 1.5 model trained to &quot;think&quot; and self-improve (397b params / 242GB memory)</p>\n<p>qwen3-coder meant to produce code (30b params / 19GB memory)</p>\n<p>I don&#39;t even have 300 GB on any of my disk, so I quickly chose the 10x-reduced version of <code>ornith 1.5</code> to see how good it is (34b params / 23 GB of memory).</p>\n<p>Do I even need a Large Language Model (LLM), or would some &quot;medium&quot; language models do the job for personal use? Some answers on this talk we had at the Tech F&#39;Est 2026.</p>\n<h2 id=\"you-don-t-need-the-most-expensive-gpu-but-one-with-enough-virtua\">You don&#39;t need the most expensive GPU, but one with enough virtual memory</h2>\n<p>Ollama models work on a CPU as long as you have enough memory to fit the entire model. The generation speed depends on your computation power.</p>\n<p>You can try to run <code>ollama run qwen3-coder</code>, if you have more than 20 GB of memory on your device it will work, but the generation speed will be very slow because of the number of parameters.</p>\n<p>What you need is a GPU because they can do a lot more parallel computation.\nBut not all GPUs are equal.</p>\n<p>You need to have a GPU with enough virtual memory to fit the whole model in memory + the whole context size you want.</p>\n<p>At home, I have an old gaming PC with a <code>NVIDIA GeForce RTX 2060</code> card I bought in 2020. I tried it, but it only has 6 GB of virtual memory, which is not enough.</p>\n<p>Ollama is smart enough to use the device memory as a failover when GPU VRAM is too small, but it means falling back on the CPU which defeats the purpose of using a GPU.</p>\n<p>Apple Silicon devices have shared memory between CPU and GPU which is super helpful, because the GPU can read the device memory instead of its own VRAM.</p>\n<p>Also, we could fit more context in the virtual memory with quantization but it may impact the model response quality.</p>\n<p>Latest gaming GPUs can go up to 32 GB of VRAM, which is great to run those medium-sized models. But I don&#39;t want to buy one, which is only leaving me one solution.</p>\n<h2 id=\"renting-gpu-under-their-market-value-aws-spot-instances\">Renting GPU under their market value: AWS Spot Instances</h2>\n<p>AWS has a range of GPU instances with different cards available.</p>\n<p>They&#39;re also offering a reduced price for unused capacity on a service called AWS Spot Instances.</p>\n<p>You choose the instance you want, they offer a reduced rate but the price is variable so you need to set a maximum price.\nWhen maximum price is reached, the instance is stopped, and you have to request a new one.</p>\n<p>To enable EC2 Spot Instance, you first need to request a quota increase. It&#39;s usually approved in a few days max.</p>\n<p><code>aws service-quotas request-service-quota-increase --service-code ec2 --quota-code L-3819A6DF --desired-value 32</code></p>\n<p>The instances with a single GPU are ranging from $0.5/h to $10/h, depending on the availability.</p>\n<p>I wrote a helper to spin up a spot instance, select the pricing, choose the Ollama models and spin it up: aws-llm-sandbox.</p>\n<p>`./up.sh\n==&gt; my public IP: xx.xx.xx.xx\n==&gt; security group sg-0d910bcfdb06e201e: 22 + 11434 open to xx.xx.xx.xx/32 only\n==&gt; no existing instance - launching a new one</p>\n<p>Available GPUs (single-GPU instances). Pick by VRAM: the model + its KV cache must fit in it.</p>\n<ol><li><p>16 GB VRAM  NVIDIA T4                  2018 gen (g4dn). Cheapest; fits &lt;=14B models at q4. Too small for 30B.</p><pre><code>           types: g4dn.xlarge, g4dn.2xlarge, g4dn.4xlarge, g4dn.8xlarge, g4dn.16xlarge</code></pre></li><li><p>24 GB VRAM  NVIDIA A10G                2021 gen (g5). Same 24 GB tier as L4, ~same speed for Ollama; a second pool when g6 has no capacity.</p><pre><code>           types: g5.xlarge, g5.2xlarge, g5.4xlarge, g5.8xlarge, g5.16xlarge</code></pre></li><li><p>24 GB VRAM  NVIDIA L4                  2023 gen (g6, gr6 = same GPU with 2x system RAM). 24 GB: one 30B model at q4 (~18 GB) + KV cache; only one big model resident at a time.</p><pre><code>           types: g6.xlarge, g6.2xlarge, g6.4xlarge, gr6.4xlarge, g6.8xlarge, gr6.8xlarge, g6.16xlarge</code></pre></li><li><p>48 GB VRAM  NVIDIA L40S                2023 gen (g6e). 48 GB, ~2-3x L4 throughput: two 30B models resident, or 70B at q4, or 30B with a huge context.</p><pre><code>           types: g6e.xlarge, g6e.2xlarge, g6e.4xlarge, g6e.8xlarge, g6e.16xlarge</code></pre></li><li><p>96 GB VRAM  NVIDIA RTX PRO Server 6000 2025 gen (g7e, Blackwell). 96 GB, often priced like the L40S: 70B at q8, 120B-class MoE at q4, or several 30B models resident.</p><pre><code>           types: g7e.2xlarge, g7e.4xlarge, g7e.8xlarge</code></pre>\n<p>GPU? [1-5] 4\n==&gt; fetching spot prices for: g6e.xlarge g6e.2xlarge g6e.4xlarge g6e.8xlarge g6e.16xlarge\n==&gt; fetching spot placement scores</p>\n<p> MARKET     TYPE           AZ                     $/h  SCORE  vCPU   RAM</p></li><li>spot       g6e.4xlarge    eu-central-1b       0.5638  1      16     128 GB</li><li>spot       g6e.4xlarge    eu-central-1a       1.0720  1      16     128 GB</li><li>spot       g6e.2xlarge    eu-central-1b       1.1812  1      8      64 GB</li><li>spot       g6e.xlarge     eu-central-1a       1.8827  3      4      32 GB</li><li>spot       g6e.2xlarge    eu-central-1a       2.0247  1      8      64 GB</li><li>spot       g6e.xlarge     eu-central-1c       2.0259  1      4      32 GB</li><li>spot       g6e.4xlarge    eu-central-1c       2.2143  1      16     128 GB</li><li>spot       g6e.xlarge     eu-central-1b       2.3270  1      4      32 GB</li><li>on-demand  g6e.xlarge     any                 2.3270  -      4      32 GB</li><li>spot       g6e.8xlarge    eu-central-1a       2.6137  1      32     256 GB</li><li>spot       g6e.8xlarge    eu-central-1c       2.7531  1      32     256 GB</li><li>spot       g6e.16xlarge   eu-central-1b       2.7611  1      64     512 GB</li><li>spot       g6e.2xlarge    eu-central-1c       2.8035  1      8      64 GB</li><li>on-demand  g6e.2xlarge    any                 2.8035  -      8      64 GB</li><li>spot       g6e.16xlarge   eu-central-1a       2.9150  1      64     512 GB</li><li>spot       g6e.16xlarge   eu-central-1c       3.2037  2      64     512 GB</li><li>on-demand  g6e.4xlarge    any                 3.7565  -      16     128 GB</li><li>spot       g6e.8xlarge    eu-central-1b       5.6625  1      32     256 GB</li><li>on-demand  g6e.8xlarge    any                 5.6625  -      32     256 GB</li><li>on-demand  g6e.16xlarge   any                 9.4745  -      64     512 GB</li></ol>\n<p>  SCORE: spot placement score 1-10 (10 = capacity very likely, 1 = unlikely).\n  spot: market price, varies; AWS stops the instance if it rises above your cap. Needs the spot quota.\n  on-demand: fixed price, never interrupted, no quota issue; AZ chosen by AWS where capacity exists.\nWhich one? [1-20]\n`</p>\n<p>I tested various instances, and there are family types (g5, g6, g6e) whose prices range from $0.5/h to $10/h, and I found my sweet spot with a <code>g6.4xlarge</code> instance to run my two models.</p>\n<p>Cost Management = Just-in-time provisioning</p>\n<p>A good way to save on the bill is to just stop the instance when you don&#39;t use it.</p>\n<p>But when it&#39;s turned on, you have virtually no limit of tokens!</p>\n<p>Here&#39;s my AWS bill from last month, and it was before I used Spot instances.</p>\n<p>I&#39;m still using Fable for most complex planning so I&#39;m keeping Claude Pro for now. I tried all that to understand better and check if it was possible, but I&#39;m saving a little bit of money, which is great. Obviously, it&#39;s not applicable in all situations.</p>\n<p>But it&#39;s totally possible to orchestrate a fleet of low-cost GPU instances, with a routine that track the cheapest available spot instances on multiple regions for your own usage.</p>\n<h2 id=\"from-claude-code-to-hermes\">From Claude Code to Hermes</h2>\n<p>It is possible to run Ollama models with Claude Code.\nBut I found it easier with Hermes.</p>\n<p>From Hermes, you can easily add your own custom models and switch provider and model in one command.</p>\n<p>`hermes model</p>\n<p>  Current model:    ornith-1.5:35b\n  Active provider:  llm-sandbox</p>\n<p>Custom OpenAI-compatible endpoint configuration:</p>\n<p>API base URL [e.g. <a href=\"https://api.example.com/v1]\" rel=\"nofollow ugc noopener\">https://api.example.com/v1]</a>:\n`</p>\n<p>There are a lot of cool features with Hermes that can add up.</p>\n<p>Like running multiple Profiles. One profile is a &quot;Product Manager agent&quot; running with ornith, a second is a &quot;Coder agent&quot; with qwen3, and a last &quot;Architect&quot; profile running Fable whose goal is to write the plan and test the feature in the end.</p>\n<p>And you can then run a Kanban where the &quot;Product agent&quot; creates tickets, and some coder agents automatically pick up their tasks and implement them.</p>\n<h2 id=\"what-about-parallel-work-and-sub-agents\">What about parallel work and sub-agents?</h2>\n<p>Ollama can totally handle concurrent requests, even with multiple models, if you have enough memory.</p>\n<p>Here is an entire blog post describing how it works and how to configure it (fr).</p>\n<p>I&#39;m easily running 3 ornith agents in parallel on a <code>g7e.4xlarge</code> instance.</p>\n<h2 id=\"what-s-next\">What&#39;s next?</h2>\n<p>Pretty much nothing, I&#39;m happy with my setup for now and it allows me to test a LOT of models available on Ollama with different setups and automations.</p>\n<p>I&#39;m surprised how far I&#39;ve got, I was expecting to hit a performance wall or spend way too much on GPU.</p>\n<p>I didn&#39;t entirely get rid of the Claude models and their 1M context windows, but it was not the goal either.</p>\n<p>By paying a Claude / Codex subscription we&#39;re not even paying for a fraction of the compute required, so it&#39;s still very cheap to keep your subscription.</p>\n<p>For programmatic use, burst usage or a certain type of compliance need, though, you&#39;ll still find self-hosting useful. And just-in-time provisioning of GPU instances will probably come handy.</p>\n<p>There are plenty of other things to play with, like an Ollama load balancer over multiple instances.</p>\n<p>I&#39;m curious how this could fit in some other contexts. Could a well-orchestrated fleet save the bill and footprint of a small/medium company?</p>","headings":[{"level":1,"text":"Self-hosting LLM models for software development","id":"self-hosting-llm-models-for-software-development"},{"level":2,"text":"What model do we need for software development?","id":"what-model-do-we-need-for-software-development"},{"level":2,"text":"You don't need the most expensive GPU, but one with enough virtual memory","id":"you-don-t-need-the-most-expensive-gpu-but-one-with-enough-virtua"},{"level":2,"text":"Renting GPU under their market value: AWS Spot Instances","id":"renting-gpu-under-their-market-value-aws-spot-instances"},{"level":2,"text":"From Claude Code to Hermes","id":"from-claude-code-to-hermes"},{"level":2,"text":"What about parallel work and sub-agents?","id":"what-about-parallel-work-and-sub-agents"},{"level":2,"text":"What's next?","id":"what-s-next"}]}}