{"article":{"slug":"local-ai-on-a-12-gb-gpu-what-survived-testing-and-how-to-set-it-up","title":"Local AI on a 12 GB GPU: what survived testing, and how to set it up","subtitle":null,"summary":"Hands-on notes testing local AI models on a 12 GB RTX 3060: which stacks fit in VRAM, how context length decides spills to CPU, and a practical setup that survived the author’s trials.","content_type":"tutorial","language":"en","canonical_url":"https://www.akashdamle.in/blog/local-ai-on-a-12gb-gpu","author":{"name":"Akash Damle","url":"https://www.akashdamle.in/","person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"Akash Damle","url":"https://www.akashdamle.in/","listing_slug":null,"listing":null},"topics":[{"name":"AI","slug":"ai","url":"https://listedarticles.com/topics/ai"},{"name":"LLMs","slug":"llms","url":"https://listedarticles.com/topics/llms"},{"name":"Hardware","slug":"hardware","url":"https://listedarticles.com/topics/hardware"},{"name":"Tutorials","slug":"tutorials","url":"https://listedarticles.com/topics/tutorials"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":3307,"reading_minutes":14,"published_at":"2026-09-25T12:00:00.000Z","added_at":"2026-09-28T06:19:19.809Z","updated_at":"2026-09-28T06:19:19.809Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":false},"profile_url":"https://listedarticles.com/articles/local-ai-on-a-12-gb-gpu-what-survived-testing-and-how-to-set-it-up","markdown_url":"https://listedarticles.com/articles/local-ai-on-a-12-gb-gpu-what-survived-testing-and-how-to-set-it-up.md","example":false,"citation":"Akash Damle, Akash Damle. \"Local AI on a 12 GB GPU: what survived testing, and how to set it up.\" 25 Sept 2026. https://www.akashdamle.in/blog/local-ai-on-a-12gb-gpu (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://www.akashdamle.in/blog/local-ai-on-a-12gb-gpu"},"body_markdown":"## In short\n\n- **Context length decides whether a model fits.** The same 5.3 GB model took 18 GB and spilled onto\nthe CPU at a 131K context, but used 6.8 GB entirely on the GPU at 16K.\n- **Three models cover everything I need:** a 7B model for autocomplete, Qwen 3.5 9B for chat,\ndocuments and small code changes, and Qwen 3 Coder 30B, a mixture-of-experts model that runs well\ndespite not fitting in VRAM, for larger coding tasks.\n- **The coding agent needs a 32K context and a small output reserve,** or it summarises its own\nconversation in an endless loop.\n- **The dangerous failures looked like successes.** Models reported work as done when nothing was\nsaved, or when the saved file still had errors. Check files, diffs and tests, not summaries.\n- **A six-step setup, with a check for each step,** is at the end.\n\n## Why I did this\n\nI write the core logic of my projects myself. I wanted local models for the work around it: boilerplate, documentation, tests, consistency checks, questions about my own code and documents. I also wanted to do that without sending project files to a cloud service by default. Cloud models stay available for when speed or certainty matters, but the goal is to need them less, not to treat them as a permanent backup.\n\nI had built a local setup once before, in August. When I came back to it in September and compared\nmy notes with the machine, I found eleven contradictions. The nine models I had benchmarked were\nall gone. The context-length setting my notes called \"permanently set\" existed but was empty. The\ncoding app's local provider was switched off by a single flag (`\"enabled\": false`) while its\ninstance settings said `true`. So every \"model failure\" in my previous test round had been\nmeasuring a route that was never open.\n\nI started again. This time I checked every result against what was actually on disk, not what a model or my own notes claimed. This article covers what survived, what failed, and the shortest path I know to reproduce it.\n\n## The machine, and the one number that matters most\n\n| Part | Specification | \n|---|---|\n| OS | Windows 11 Pro | \n| GPU | NVIDIA RTX 3060, 12 GB VRAM | \n| RAM | 32 GB (31.8 GB usable) | \n| CPU | Intel i5-12400 class, 6 cores / 12 threads | \n| Storage | Models on a separate data drive (about 28 GB for the final set) | \n\nThere are two memory budgets. Video memory (VRAM) is fast. System RAM is slow, but a model that does not fit in VRAM still runs from it.\n\n### Context length is the memory you don't see\n\nThe first measurement changed how I made every later decision. I loaded one model twice in the\nsame session and changed only the context length. The model was `granite4.2:8b`, 5.3 GB on disk:\n\n| Context length | Loaded size | Where it ran | \n|---|---|---|\n| 131,072 tokens (the server default) | 18 GB | 42% CPU / 58% GPU | \n| 16,384 tokens | 6.8 GB | 100% GPU | \n\nOn a 12 GB card, the context length, not the size of the weights, is what decides whether a model\nfits. **Set the context length before you compare models.** Otherwise you are comparing\nmemory settings, not models.\n\n### \"Too big\" depends on the architecture\n\n`qwen3-coder:30b` is a mixture-of-experts model: 30B parameters in total, but only 3.3B are\nactive for each token. At a 32K context it loaded at 20 GB, split 48% CPU / 52% GPU, and still\ngenerated **33.5 tokens per second**. A dense 24B model (`devstral-small-2:24b`) loaded at 16.2 GB,\nsplit 39% / 61%, and generated **4.8 tokens per second**. The larger model was about seven times\nfaster. If a model looks too big for your GPU, check whether it is a mixture-of-experts model\nbefore you rule it out.\n\n## How I tested\n\n- **Check the file, not the summary.** Each task had an independent checker that read the saved\nfiles. For the coding tasks, a hidden checker scored both a reference solution and the untouched\nstarting code first. On the hardest task, the reference scored 19/19 and the untouched code 7/19,\nso a model had to change things to score above 7.\n- **Run a control first.** Before any real task in an app, ask the model to write a one-line file,\nthen compare the bytes on disk. This separates \"the model can't use tools\" from \"the tools were\nnever reachable\", which was exactly the question I could not answer about the first setup.\n- **Change one thing per retry.** I kept every failed attempt as it happened and recorded\nexactly what changed before the next one.\n- **Know what the tests cover.** These are small tasks on one machine, mostly single runs. They\nshow what works on this setup, not a general ranking of models.\n\n## What I ended up with\n\nEvery app here talks to Ollama, which runs the models on the GPU. On top of it:\n\n- T3 Code (source), an open-source desktop app for running coding agents. Here it runs OpenCode, an open-source coding agent, against the local models.\n- Jan, a desktop chat app that can use tools.\n- AnythingLLM, which answers questions from your own documents, with citations.\n- VS Code with Twinny, for autocomplete only.\n\nI tried fourteen models; three stayed. They cover four jobs across four apps:\n\n| Job | App | Model | Tested result | \n|---|---|---|---|\n| Autocomplete while typing | VS Code + Twinny | `qwen2.5-coder:7b-base` | 0.2–0.9 s per suggestion once loaded; 6/8 completion checks vs 2/8 for the 3B | \n| Small code changes | T3 Code (OpenCode agent) | `qwen3.5:9b` , 32K variant | Multi-step task 12/12 in 75 s | \n| Larger code tasks | T3 Code | `qwen3-coder:30b` , 32K variant | Harder task 19/19 twice, 7–9 minutes each; the 9B scored 15/19 | \n| Chat, screenshots, spreadsheets | Jan | `qwen3.5:9b` | Workbook repair 10/10 with guarded tools; image test 16/16 | \n| Questions about documents | AnythingLLM | `qwen3.5:9b` | 4/4 answers with citations | \n| Questions about code | T3 Code | `qwen3.5:9b` , 32K variant | 4.5/5 using live search | \n\n`qwen3.5:9b` runs entirely on the GPU at about 50 tokens per second, which makes it the default\nfor almost everything. The 30B handles tasks that touch several files. Only one large model fits\nin VRAM at a time, so switching between them costs a load: about 10 seconds for the 9B and about\n36 seconds for the 30B.\n\n## What failed, and what each failure taught me\n\n### The coding agent that compacted forever\n\nT3 Code runs OpenCode as its agent. OpenCode compacts (summarises) a conversation once it grows past the context limit minus the space it reserves for output. I had configured 16,384 tokens of context with 8,192 reserved for output, so compaction started at 8,192 tokens. OpenCode's own starting prompt was about 9.3K tokens from the command line and about 15.5K inside T3. Every turn was already over the threshold before I typed anything, so every turn compacted, in a loop.\n\nOn the one-line file control, the file came out correct, but the turn was still compacting after 154 seconds (three compactions) when I stopped it. The fix needed no new download: a variant of the same model with a 32K context, defined in a two-line Modelfile that shares the existing weights, plus a 4,096-token output reserve. The same control then finished in 24.5 seconds with no compaction, still entirely on the GPU.\n\n### Models that said \"done\"\n\nThe most important failures were the ones that looked like successes:\n\n- `deepseek-r1:14b` made**no tool calls** on a spreadsheet task, then reported that it had inspected,\nedited and saved the workbook, with figures that did not come from the data.\n- `qwen2.5:14b` left a`#DIV/0!` error in the saved file and reported that cell as showing \"n.a.\".\n- In one T3 run, the 9B **deleted an existing test** while rewriting its own failing tests. The\nsuite passed, and the diff showed the deletion.\n- On a scaffolding task, the 30B's summary did not mention any of the three gaps the checker found.\n\nThe workbook qwen2.5:14b saved, drawn from the file itself. Its report said Travel utilization now shows “n.a.”; the saved cell still shows #DIV/0!. The red outline is mine.\n\nA wrong answer is easy to spot. A false \"done\" is not. Read the diff, run the tests yourself, and open the saved file.\n\n### Three failures with three different causes\n\nIn Jan, I asked `qwen3.5:9b` to repair an inventory workbook. The same model had scored 12/12 on\nthis task through a minimal test harness. In Jan it failed three times, each for a different reason:\n\n1. **Context ran out.** A single inspection tool returned about 16,000 characters, and the\nattempt used up all 16,384 tokens of context before making any edit.\n2. **Output cap.** Jan's assistant allowed 2,048 output tokens. The model spent all of them\nreasoning and never called a tool.\n3. **Formulas saved as text.** With the cap raised to 4,096, it saved the file, but it wrote formulas\nwithout the leading`=` , so they were stored as text. Four of ten checks passed.\n\nFor the fourth attempt I changed only the tools. I added two: a compact cell reader, and an\neditor that rejects any formula without `=` and saves nothing if an edit is invalid. The editor\nrejected the model's first attempt, which was the same mistake as attempt 3. The model corrected\nit and passed **10/10**. A guardrail in the tool worked where prompting alone had not.\n\n### The autocomplete that wasn't local\n\nMy first autocomplete test \"passed\": suggestions appeared in under a second. They were coming from\nGitHub Copilot, which is built into current VS Code and was signed in on the free plan. The local\nmodel had never loaded. After I turned Copilot off (`\"chat.disableAIFeatures\": true`), the local\n7B model served suggestions in 0.19–0.93 seconds once it was loaded. The first suggestion after a\npause can take around 20 seconds while the model loads. Raising Twinny's keep-alive to 30 minutes\nmade that happen less often. A suggestion that appears in under a second right after a cold start\nis a sign that something other than your local model is answering.\n\n### Uploading is not embedding\n\nMy first document test in AnythingLLM answered \"I don't have access to documents\" to all four\nquestions. The files had been uploaded and parsed, but never added to the workspace, so nothing\nwas embedded. After I used **Move to Workspace → Save and Embed** and switched the workspace to\n**Query** mode, it answered 4/4 correctly with citations. One of the two documents was a\nsuperseded version of the other, and it chose the current revision even when the old one ranked\nfirst in the search. This was a two-document test, so it shows the answers stay grounded in the\ndocuments, not that retrieval works at scale.\n\n### A code index lost to plain search\n\nI indexed a Django repository in AnythingLLM and asked five questions with a hidden answer key. Only 100 of the 136 files made it into the workspace, and it scored 2.5/5 in about nine minutes. The same 9B model, searching the files live through OpenCode, scored 4.5/5 in 2 minutes 41 seconds. I dropped the code index. For questions about code, I ask the coding agent and tell it not to modify files.\n\n### A tool call written as plain text\n\nIn T3, `devstral-small-2:24b` wrote its first tool call as ordinary text (`glob{\"path\": ...}`)\ninstead of calling the tool, and then stopped. It scored 7/19, the same as the untouched starting\ncode. Inside T3 it also ran at only 2.3–2.5 tokens per second. A model listed as supporting tools\nstill has to use them correctly inside the app you actually run.\n\nI had seen the same failure in my first setup, in August, with a different model:\n\nAugust, OpenCode with hhao/qwen2.5-coder-tools: the tool calls come out as JSON text, so nothing runs, and the paths are Linux paths on a Windows machine. That model is no longer in my setup.\n\n## Coding workflows that held up\n\n- **Docstrings.** I put the style rules (Google style for Python, TSDoc for TypeScript) in\nOpenCode's global instructions file and in a user-level Ruff configuration. Given a file and no\nstyle details in the prompt, the 30B brought Ruff's docstring findings from 8 to 0 and left the\ncode itself unchanged. The docstrings still need reading: one repeated a misleading function\nname.\n- **Scaffolding.** A template prompt walks the model from model to form or serializer, then views\nand URLs. The first plain-Django run scored 13/16: the wiring was correct, but it added logic to a\nstub, skipped type hints and left unused imports. I added closing rules to the template (one-line\nclass docstrings, stub bodies that only raise`NotImplementedError` , run Ruff after`manage.py check` ). With those, a Django REST Framework scaffold scored 14/14 in 7 minutes.\n- **Bigger changes.** On a four-file task (a parser bug, a new rule, a cross-file feature and tests),\nthe 30B scored 19/19 twice. For work spread across many files, or that needs design judgement\nacross a codebase, I still switch to a cloud model and bring the smaller follow-ups back to the\nlocal ones.\n\n## Reproduce it\n\nThis is the shortest route I know, written from the configuration that passed. I have not yet rebuilt it from scratch on a second machine. Each step ends with a check, so you can tell whether it worked.\n\n**1. Ollama.** Install it and set these user environment variables:\n\n| Variable | Value | \n|---|---|\n| `OLLAMA_CONTEXT_LENGTH` | `16384` | \n| `OLLAMA_FLASH_ATTENTION` | `1` | \n| `OLLAMA_KV_CACHE_TYPE` | `q8_0` | \n| `OLLAMA_ORIGINS` | `http://tauri.localhost` (see the warning below) | \n| `OLLAMA_MODELS` | optional: a folder on a data drive | \n\n**A warning about `OLLAMA_ORIGINS`.** In my setup I used `*`, which fixed a 403 error in Jan. But `*`\nalso lets any website open in your browser send requests to Ollama in the background, including\nrequests that delete models or start large downloads. Only Jan needed a change. Ollama's built-in\nlist already allows `localhost` and `tauri://` origins, but not `http://tauri.localhost`, which\nappears to be what Jan's Windows app sends, and that matches the 403 I saw. I tested this on a\ntemporary Ollama server. With `OLLAMA_ORIGINS=http://tauri.localhost`, that origin got through and\nan ordinary website was still blocked; with `*`, both got through. I have not re-run Jan itself\nwith the narrower value. If Jan shows a 403 with it, check which origin Jan sends before falling\nback to `*`.\n\nThe 403 in Jan, from my first setup in August, before I changed OLLAMA_ORIGINS.\n\nAlso set the context slider in Ollama's settings to 16K. Then **quit Ollama completely**, including\nthe tray icon, and start it again. A process started before the change keeps the old environment.\nThis caught me out three times. *Check:* after loading a model, `ollama ps` shows context 16384\nand `100% GPU` for the 9B.\n\n**2. Models.** Download three models (qwen3.5,\nqwen3-coder,\nqwen2.5-coder), then create two 32K variants:\n\n```\nollama pull qwen3.5:9b\nollama pull qwen3-coder:30b\nollama pull qwen2.5-coder:7b-base\nollama create qwen3.5:9b-32k -f Modelfile.qwen3.5-9b-32k\nollama create qwen3-coder:30b-32k -f Modelfile.qwen3-coder-30b-32k\n```\nEach Modelfile has two lines, for example:\n\n```\nFROM qwen3.5:9b\nPARAMETER num_ctx 32768\n```\n*Check:* `ollama list` shows five entries, and `ollama show qwen3.5:9b-32k --parameters` shows\n`num_ctx 32768`.\n\n**3. T3 Code with OpenCode.** Declare the two 32K models in `~/.config/opencode/opencode.json`,\nwith a 4,096-token output reserve:\n\n```\n{\n  \"$schema\": \"https://opencode.ai/config.json\",\n  \"provider\": {\n    \"ollama\": {\n      \"npm\": \"@ai-sdk/openai-compatible\",\n      \"name\": \"Ollama Local\",\n      \"options\": { \"baseURL\": \"http://localhost:11434/v1\" },\n      \"models\": {\n        \"qwen3.5:9b-32k\": { \"name\": \"Qwen 3.5 9B Local 32K\", \"limit\": { \"context\": 32768, \"output\": 4096 } },\n        \"qwen3-coder:30b-32k\": { \"name\": \"Qwen 3 Coder 30B Local 32K\", \"limit\": { \"context\": 32768, \"output\": 4096 } }\n      }\n    }\n  }\n}\n```\nEnable the OpenCode provider in T3. Quit T3 completely and reopen it, because it does not refresh\nits model list otherwise. A new thread copies the model of the thread you are viewing, so check the\nmodel button before you send. T3's default permission mode, Full access, runs shell commands and\nedits without asking; choose the mode for each thread deliberately. If you script OpenCode, run\n`opencode run ... < /dev/null`, otherwise it waits forever for input. *Check:* ask for a one-line\nfile in a scratch project and confirm the exact text on disk.\n\n**4. Jan.** Add Ollama as a provider at `http://localhost:11434/v1`, select `qwen3.5:9b` and turn on\nits tools and vision capabilities. Raise the assistant's maximum output tokens to 4,096.\n*Check:* a normal chat reply, and a correct answer about a screenshot.\n\n**5. AnythingLLM.** Choose Ollama (`http://127.0.0.1:11434`, model `qwen3.5:9b`), the built-in\nembedder and LanceDB. Create a workspace in Query mode, and after uploading use **Move to\nWorkspace → Save and Embed**. *Check:* a question whose answer is in one document comes back with a\ncitation.\n\n**6. VS Code autocomplete.** Install Twinny and add one fill-in-middle provider: Ollama on\n`localhost:11434`, path `/api/generate`, model `qwen2.5-coder:7b-base`, template automatic. Leave out\nTwinny's chat. In settings, add `\"twinny.keepAlive\": \"30m\"` and `\"chat.disableAIFeatures\": true`.\nKeep error hints deterministic with a linter and type checker (I use Ruff, Pylance and ErrorLens).\n*Check:* suggestions within a second once the model is loaded.\n\nTested versions: Ollama 0.34.3, OpenCode 1.18.32, T3 Code 0.0.42, Jan 0.8.4, AnythingLLM 1.16.1, Twinny 4.2.5 (it has since updated itself to 4.2.7).\n\n## Limits\n\n- One machine, small fixtures, mostly single runs. Results varied between runs: the same 9B that scored 12/12 deleted a test on a repeat.\n- The 32K context fills up on bigger tasks: the largest runs came close to the limit and compacted.\n- T3's \"Auto-accept edits\" mode was never exercised with a local model. Every run used Full access.\n- Local models have no web access, so anything that depends on recent releases or current documentation goes to a cloud model.\n- Apps update themselves. Twinny updated during testing and again before I wrote this. Re-run the checks after an update.\n\n## What I didn't use, and why\n\n- **Other runtimes.** I did not compare Ollama with LM Studio or\nllama.cpp's own server. Every app here could talk to\nOllama, so I kept one runtime and spent the time testing the apps. Another runtime may be faster\non the same card; I haven't measured it.\n- **Open WebUI.** It needs Docker, which I didn't want on this machine.\n- **An agent inside the editor.** Cline was my main assistant in the earlier setup; I replaced it\nwith T3, where agent work happens in its own threads and I review the diffs. VS Code keeps\nautocomplete and deterministic hints only. I have not tested Continue in this setup.\n- **A model as a linter.** Ruff, Pylance and ErrorLens are instant and never invent a rule. A model\nwould be slower and less predictable at the same job.\n- **A code index.** Tested above: it lost to the agent searching the files live.\n\n## What I would tell myself at the start\n\n1. Fix the context length before judging any model.\n2. Run a control task before blaming a model for a failure.\n3. Trust files, diffs and tests over summaries, including your own notes.\n4. Change one thing per retry, and keep the failures.\n5. Fewer models, each with a job, beat a large collection nobody has tested in the apps you actually use.\n\n## About the author\n\nI'm Akash Damle, founder of Matalli Infotech Private Limited, an early-stage company I'm building from the ground up. I write to mark milestones: what was built, what broke, and what it taught me.\n\nIf you follow this setup, especially on different hardware, I would like to hear how it went: what you ran, what happened, and what you changed. Disagreement is as welcome as agreement, as long as it comes with a reason. Reports like that are how the next version of this guide gets better.\n\n## Comments\n\nTried this setup, hit a different result, or think something here is wrong? Say what you ran, what happened, and why. Agreement and disagreement are equally welcome; one-line reactions are not published. No account needed.","body_html":"<h2 id=\"in-short\">In short</h2>\n<ul><li><p><strong>Context length decides whether a model fits.</strong> The same 5.3 GB model took 18 GB and spilled onto</p><p>the CPU at a 131K context, but used 6.8 GB entirely on the GPU at 16K.</p></li><li><p><strong>Three models cover everything I need:</strong> a 7B model for autocomplete, Qwen 3.5 9B for chat,</p><p>documents and small code changes, and Qwen 3 Coder 30B, a mixture-of-experts model that runs well\ndespite not fitting in VRAM, for larger coding tasks.</p></li><li><p><strong>The coding agent needs a 32K context and a small output reserve,</strong> or it summarises its own</p><p>conversation in an endless loop.</p></li><li><p><strong>The dangerous failures looked like successes.</strong> Models reported work as done when nothing was</p><p>saved, or when the saved file still had errors. Check files, diffs and tests, not summaries.</p></li><li><strong>A six-step setup, with a check for each step,</strong> is at the end.</li></ul>\n<h2 id=\"why-i-did-this\">Why I did this</h2>\n<p>I write the core logic of my projects myself. I wanted local models for the work around it: boilerplate, documentation, tests, consistency checks, questions about my own code and documents. I also wanted to do that without sending project files to a cloud service by default. Cloud models stay available for when speed or certainty matters, but the goal is to need them less, not to treat them as a permanent backup.</p>\n<p>I had built a local setup once before, in August. When I came back to it in September and compared\nmy notes with the machine, I found eleven contradictions. The nine models I had benchmarked were\nall gone. The context-length setting my notes called &quot;permanently set&quot; existed but was empty. The\ncoding app&#39;s local provider was switched off by a single flag (<code>&quot;enabled&quot;: false</code>) while its\ninstance settings said <code>true</code>. So every &quot;model failure&quot; in my previous test round had been\nmeasuring a route that was never open.</p>\n<p>I started again. This time I checked every result against what was actually on disk, not what a model or my own notes claimed. This article covers what survived, what failed, and the shortest path I know to reproduce it.</p>\n<h2 id=\"the-machine-and-the-one-number-that-matters-most\">The machine, and the one number that matters most</h2>\n<div class=\"table-wrap\"><table><thead><tr><th>Part</th><th>Specification</th></tr></thead><tbody><tr><td>OS</td><td>Windows 11 Pro</td></tr><tr><td>GPU</td><td>NVIDIA RTX 3060, 12 GB VRAM</td></tr><tr><td>RAM</td><td>32 GB (31.8 GB usable)</td></tr><tr><td>CPU</td><td>Intel i5-12400 class, 6 cores / 12 threads</td></tr><tr><td>Storage</td><td>Models on a separate data drive (about 28 GB for the final set)</td></tr></tbody></table></div>\n<p>There are two memory budgets. Video memory (VRAM) is fast. System RAM is slow, but a model that does not fit in VRAM still runs from it.</p>\n<h3 id=\"context-length-is-the-memory-you-don-t-see\">Context length is the memory you don&#39;t see</h3>\n<p>The first measurement changed how I made every later decision. I loaded one model twice in the\nsame session and changed only the context length. The model was <code>granite4.2:8b</code>, 5.3 GB on disk:</p>\n<div class=\"table-wrap\"><table><thead><tr><th>Context length</th><th>Loaded size</th><th>Where it ran</th></tr></thead><tbody><tr><td>131,072 tokens (the server default)</td><td>18 GB</td><td>42% CPU / 58% GPU</td></tr><tr><td>16,384 tokens</td><td>6.8 GB</td><td>100% GPU</td></tr></tbody></table></div>\n<p>On a 12 GB card, the context length, not the size of the weights, is what decides whether a model\nfits. <strong>Set the context length before you compare models.</strong> Otherwise you are comparing\nmemory settings, not models.</p>\n<h3 id=\"too-big-depends-on-the-architecture\">&quot;Too big&quot; depends on the architecture</h3>\n<p><code>qwen3-coder:30b</code> is a mixture-of-experts model: 30B parameters in total, but only 3.3B are\nactive for each token. At a 32K context it loaded at 20 GB, split 48% CPU / 52% GPU, and still\ngenerated <strong>33.5 tokens per second</strong>. A dense 24B model (<code>devstral-small-2:24b</code>) loaded at 16.2 GB,\nsplit 39% / 61%, and generated <strong>4.8 tokens per second</strong>. The larger model was about seven times\nfaster. If a model looks too big for your GPU, check whether it is a mixture-of-experts model\nbefore you rule it out.</p>\n<h2 id=\"how-i-tested\">How I tested</h2>\n<ul><li><p><strong>Check the file, not the summary.</strong> Each task had an independent checker that read the saved</p><p>files. For the coding tasks, a hidden checker scored both a reference solution and the untouched\nstarting code first. On the hardest task, the reference scored 19/19 and the untouched code 7/19,\nso a model had to change things to score above 7.</p></li><li><p><strong>Run a control first.</strong> Before any real task in an app, ask the model to write a one-line file,</p><p>then compare the bytes on disk. This separates &quot;the model can&#39;t use tools&quot; from &quot;the tools were\nnever reachable&quot;, which was exactly the question I could not answer about the first setup.</p></li><li><p><strong>Change one thing per retry.</strong> I kept every failed attempt as it happened and recorded</p><p>exactly what changed before the next one.</p></li><li><p><strong>Know what the tests cover.</strong> These are small tasks on one machine, mostly single runs. They</p><p>show what works on this setup, not a general ranking of models.</p></li></ul>\n<h2 id=\"what-i-ended-up-with\">What I ended up with</h2>\n<p>Every app here talks to Ollama, which runs the models on the GPU. On top of it:</p>\n<ul><li>T3 Code (source), an open-source desktop app for running coding agents. Here it runs OpenCode, an open-source coding agent, against the local models.</li><li>Jan, a desktop chat app that can use tools.</li><li>AnythingLLM, which answers questions from your own documents, with citations.</li><li>VS Code with Twinny, for autocomplete only.</li></ul>\n<p>I tried fourteen models; three stayed. They cover four jobs across four apps:</p>\n<div class=\"table-wrap\"><table><thead><tr><th>Job</th><th>App</th><th>Model</th><th>Tested result</th></tr></thead><tbody><tr><td>Autocomplete while typing</td><td>VS Code + Twinny</td><td><code>qwen2.5-coder:7b-base</code></td><td>0.2–0.9 s per suggestion once loaded; 6/8 completion checks vs 2/8 for the 3B</td></tr><tr><td>Small code changes</td><td>T3 Code (OpenCode agent)</td><td><code>qwen3.5:9b</code> , 32K variant</td><td>Multi-step task 12/12 in 75 s</td></tr><tr><td>Larger code tasks</td><td>T3 Code</td><td><code>qwen3-coder:30b</code> , 32K variant</td><td>Harder task 19/19 twice, 7–9 minutes each; the 9B scored 15/19</td></tr><tr><td>Chat, screenshots, spreadsheets</td><td>Jan</td><td><code>qwen3.5:9b</code></td><td>Workbook repair 10/10 with guarded tools; image test 16/16</td></tr><tr><td>Questions about documents</td><td>AnythingLLM</td><td><code>qwen3.5:9b</code></td><td>4/4 answers with citations</td></tr><tr><td>Questions about code</td><td>T3 Code</td><td><code>qwen3.5:9b</code> , 32K variant</td><td>4.5/5 using live search</td></tr></tbody></table></div>\n<p><code>qwen3.5:9b</code> runs entirely on the GPU at about 50 tokens per second, which makes it the default\nfor almost everything. The 30B handles tasks that touch several files. Only one large model fits\nin VRAM at a time, so switching between them costs a load: about 10 seconds for the 9B and about\n36 seconds for the 30B.</p>\n<h2 id=\"what-failed-and-what-each-failure-taught-me\">What failed, and what each failure taught me</h2>\n<h3 id=\"the-coding-agent-that-compacted-forever\">The coding agent that compacted forever</h3>\n<p>T3 Code runs OpenCode as its agent. OpenCode compacts (summarises) a conversation once it grows past the context limit minus the space it reserves for output. I had configured 16,384 tokens of context with 8,192 reserved for output, so compaction started at 8,192 tokens. OpenCode&#39;s own starting prompt was about 9.3K tokens from the command line and about 15.5K inside T3. Every turn was already over the threshold before I typed anything, so every turn compacted, in a loop.</p>\n<p>On the one-line file control, the file came out correct, but the turn was still compacting after 154 seconds (three compactions) when I stopped it. The fix needed no new download: a variant of the same model with a 32K context, defined in a two-line Modelfile that shares the existing weights, plus a 4,096-token output reserve. The same control then finished in 24.5 seconds with no compaction, still entirely on the GPU.</p>\n<h3 id=\"models-that-said-done\">Models that said &quot;done&quot;</h3>\n<p>The most important failures were the ones that looked like successes:</p>\n<ul><li><p><code>deepseek-r1:14b</code> made<strong>no tool calls</strong> on a spreadsheet task, then reported that it had inspected,</p><p>edited and saved the workbook, with figures that did not come from the data.</p></li><li><code>qwen2.5:14b</code> left a<code>#DIV/0!</code> error in the saved file and reported that cell as showing &quot;n.a.&quot;.</li><li><p>In one T3 run, the 9B <strong>deleted an existing test</strong> while rewriting its own failing tests. The</p><p>suite passed, and the diff showed the deletion.</p></li><li>On a scaffolding task, the 30B&#39;s summary did not mention any of the three gaps the checker found.</li></ul>\n<p>The workbook qwen2.5:14b saved, drawn from the file itself. Its report said Travel utilization now shows “n.a.”; the saved cell still shows #DIV/0!. The red outline is mine.</p>\n<p>A wrong answer is easy to spot. A false &quot;done&quot; is not. Read the diff, run the tests yourself, and open the saved file.</p>\n<h3 id=\"three-failures-with-three-different-causes\">Three failures with three different causes</h3>\n<p>In Jan, I asked <code>qwen3.5:9b</code> to repair an inventory workbook. The same model had scored 12/12 on\nthis task through a minimal test harness. In Jan it failed three times, each for a different reason:</p>\n<ol><li><p><strong>Context ran out.</strong> A single inspection tool returned about 16,000 characters, and the</p><p>attempt used up all 16,384 tokens of context before making any edit.</p></li><li><p><strong>Output cap.</strong> Jan&#39;s assistant allowed 2,048 output tokens. The model spent all of them</p><p>reasoning and never called a tool.</p></li><li><p><strong>Formulas saved as text.</strong> With the cap raised to 4,096, it saved the file, but it wrote formulas</p><p>without the leading<code>=</code> , so they were stored as text. Four of ten checks passed.</p></li></ol>\n<p>For the fourth attempt I changed only the tools. I added two: a compact cell reader, and an\neditor that rejects any formula without <code>=</code> and saves nothing if an edit is invalid. The editor\nrejected the model&#39;s first attempt, which was the same mistake as attempt 3. The model corrected\nit and passed <strong>10/10</strong>. A guardrail in the tool worked where prompting alone had not.</p>\n<h3 id=\"the-autocomplete-that-wasn-t-local\">The autocomplete that wasn&#39;t local</h3>\n<p>My first autocomplete test &quot;passed&quot;: suggestions appeared in under a second. They were coming from\nGitHub Copilot, which is built into current VS Code and was signed in on the free plan. The local\nmodel had never loaded. After I turned Copilot off (<code>&quot;chat.disableAIFeatures&quot;: true</code>), the local\n7B model served suggestions in 0.19–0.93 seconds once it was loaded. The first suggestion after a\npause can take around 20 seconds while the model loads. Raising Twinny&#39;s keep-alive to 30 minutes\nmade that happen less often. A suggestion that appears in under a second right after a cold start\nis a sign that something other than your local model is answering.</p>\n<h3 id=\"uploading-is-not-embedding\">Uploading is not embedding</h3>\n<p>My first document test in AnythingLLM answered &quot;I don&#39;t have access to documents&quot; to all four\nquestions. The files had been uploaded and parsed, but never added to the workspace, so nothing\nwas embedded. After I used <strong>Move to Workspace → Save and Embed</strong> and switched the workspace to\n<strong>Query</strong> mode, it answered 4/4 correctly with citations. One of the two documents was a\nsuperseded version of the other, and it chose the current revision even when the old one ranked\nfirst in the search. This was a two-document test, so it shows the answers stay grounded in the\ndocuments, not that retrieval works at scale.</p>\n<h3 id=\"a-code-index-lost-to-plain-search\">A code index lost to plain search</h3>\n<p>I indexed a Django repository in AnythingLLM and asked five questions with a hidden answer key. Only 100 of the 136 files made it into the workspace, and it scored 2.5/5 in about nine minutes. The same 9B model, searching the files live through OpenCode, scored 4.5/5 in 2 minutes 41 seconds. I dropped the code index. For questions about code, I ask the coding agent and tell it not to modify files.</p>\n<h3 id=\"a-tool-call-written-as-plain-text\">A tool call written as plain text</h3>\n<p>In T3, <code>devstral-small-2:24b</code> wrote its first tool call as ordinary text (<code>glob{&quot;path&quot;: ...}</code>)\ninstead of calling the tool, and then stopped. It scored 7/19, the same as the untouched starting\ncode. Inside T3 it also ran at only 2.3–2.5 tokens per second. A model listed as supporting tools\nstill has to use them correctly inside the app you actually run.</p>\n<p>I had seen the same failure in my first setup, in August, with a different model:</p>\n<p>August, OpenCode with hhao/qwen2.5-coder-tools: the tool calls come out as JSON text, so nothing runs, and the paths are Linux paths on a Windows machine. That model is no longer in my setup.</p>\n<h2 id=\"coding-workflows-that-held-up\">Coding workflows that held up</h2>\n<ul><li><p><strong>Docstrings.</strong> I put the style rules (Google style for Python, TSDoc for TypeScript) in</p><p>OpenCode&#39;s global instructions file and in a user-level Ruff configuration. Given a file and no\nstyle details in the prompt, the 30B brought Ruff&#39;s docstring findings from 8 to 0 and left the\ncode itself unchanged. The docstrings still need reading: one repeated a misleading function\nname.</p></li><li><p><strong>Scaffolding.</strong> A template prompt walks the model from model to form or serializer, then views</p><p>and URLs. The first plain-Django run scored 13/16: the wiring was correct, but it added logic to a\nstub, skipped type hints and left unused imports. I added closing rules to the template (one-line\nclass docstrings, stub bodies that only raise<code>NotImplementedError</code> , run Ruff after<code>manage.py check</code> ). With those, a Django REST Framework scaffold scored 14/14 in 7 minutes.</p></li><li><p><strong>Bigger changes.</strong> On a four-file task (a parser bug, a new rule, a cross-file feature and tests),</p><p>the 30B scored 19/19 twice. For work spread across many files, or that needs design judgement\nacross a codebase, I still switch to a cloud model and bring the smaller follow-ups back to the\nlocal ones.</p></li></ul>\n<h2 id=\"reproduce-it\">Reproduce it</h2>\n<p>This is the shortest route I know, written from the configuration that passed. I have not yet rebuilt it from scratch on a second machine. Each step ends with a check, so you can tell whether it worked.</p>\n<p><strong>1. Ollama.</strong> Install it and set these user environment variables:</p>\n<div class=\"table-wrap\"><table><thead><tr><th>Variable</th><th>Value</th></tr></thead><tbody><tr><td><code>OLLAMA_CONTEXT_LENGTH</code></td><td><code>16384</code></td></tr><tr><td><code>OLLAMA_FLASH_ATTENTION</code></td><td><code>1</code></td></tr><tr><td><code>OLLAMA_KV_CACHE_TYPE</code></td><td><code>q8_0</code></td></tr><tr><td><code>OLLAMA_ORIGINS</code></td><td><code>http://tauri.localhost</code> (see the warning below)</td></tr><tr><td><code>OLLAMA_MODELS</code></td><td>optional: a folder on a data drive</td></tr></tbody></table></div>\n<p><strong>A warning about <code>OLLAMA_ORIGINS</code>.</strong> In my setup I used <code>*</code>, which fixed a 403 error in Jan. But <code>*</code>\nalso lets any website open in your browser send requests to Ollama in the background, including\nrequests that delete models or start large downloads. Only Jan needed a change. Ollama&#39;s built-in\nlist already allows <code>localhost</code> and <code>tauri://</code> origins, but not <code>http://tauri.localhost</code>, which\nappears to be what Jan&#39;s Windows app sends, and that matches the 403 I saw. I tested this on a\ntemporary Ollama server. With <code>OLLAMA_ORIGINS=http://tauri.localhost</code>, that origin got through and\nan ordinary website was still blocked; with <code>*</code>, both got through. I have not re-run Jan itself\nwith the narrower value. If Jan shows a 403 with it, check which origin Jan sends before falling\nback to <code>*</code>.</p>\n<p>The 403 in Jan, from my first setup in August, before I changed OLLAMA_ORIGINS.</p>\n<p>Also set the context slider in Ollama&#39;s settings to 16K. Then <strong>quit Ollama completely</strong>, including\nthe tray icon, and start it again. A process started before the change keeps the old environment.\nThis caught me out three times. <em>Check:</em> after loading a model, <code>ollama ps</code> shows context 16384\nand <code>100% GPU</code> for the 9B.</p>\n<p><strong>2. Models.</strong> Download three models (qwen3.5,\nqwen3-coder,\nqwen2.5-coder), then create two 32K variants:</p>\n<pre><code>ollama pull qwen3.5:9b\nollama pull qwen3-coder:30b\nollama pull qwen2.5-coder:7b-base\nollama create qwen3.5:9b-32k -f Modelfile.qwen3.5-9b-32k\nollama create qwen3-coder:30b-32k -f Modelfile.qwen3-coder-30b-32k</code></pre>\n<p>Each Modelfile has two lines, for example:</p>\n<pre><code>FROM qwen3.5:9b\nPARAMETER num_ctx 32768</code></pre>\n<p><em>Check:</em> <code>ollama list</code> shows five entries, and <code>ollama show qwen3.5:9b-32k --parameters</code> shows\n<code>num_ctx 32768</code>.</p>\n<p><strong>3. T3 Code with OpenCode.</strong> Declare the two 32K models in <code>~/.config/opencode/opencode.json</code>,\nwith a 4,096-token output reserve:</p>\n<pre><code>{\n  &quot;$schema&quot;: &quot;https://opencode.ai/config.json&quot;,\n  &quot;provider&quot;: {\n    &quot;ollama&quot;: {\n      &quot;npm&quot;: &quot;@ai-sdk/openai-compatible&quot;,\n      &quot;name&quot;: &quot;Ollama Local&quot;,\n      &quot;options&quot;: { &quot;baseURL&quot;: &quot;http://localhost:11434/v1&quot; },\n      &quot;models&quot;: {\n        &quot;qwen3.5:9b-32k&quot;: { &quot;name&quot;: &quot;Qwen 3.5 9B Local 32K&quot;, &quot;limit&quot;: { &quot;context&quot;: 32768, &quot;output&quot;: 4096 } },\n        &quot;qwen3-coder:30b-32k&quot;: { &quot;name&quot;: &quot;Qwen 3 Coder 30B Local 32K&quot;, &quot;limit&quot;: { &quot;context&quot;: 32768, &quot;output&quot;: 4096 } }\n      }\n    }\n  }\n}</code></pre>\n<p>Enable the OpenCode provider in T3. Quit T3 completely and reopen it, because it does not refresh\nits model list otherwise. A new thread copies the model of the thread you are viewing, so check the\nmodel button before you send. T3&#39;s default permission mode, Full access, runs shell commands and\nedits without asking; choose the mode for each thread deliberately. If you script OpenCode, run\n<code>opencode run ... &lt; /dev/null</code>, otherwise it waits forever for input. <em>Check:</em> ask for a one-line\nfile in a scratch project and confirm the exact text on disk.</p>\n<p><strong>4. Jan.</strong> Add Ollama as a provider at <code>http://localhost:11434/v1</code>, select <code>qwen3.5:9b</code> and turn on\nits tools and vision capabilities. Raise the assistant&#39;s maximum output tokens to 4,096.\n<em>Check:</em> a normal chat reply, and a correct answer about a screenshot.</p>\n<p><strong>5. AnythingLLM.</strong> Choose Ollama (<code>http://127.0.0.1:11434</code>, model <code>qwen3.5:9b</code>), the built-in\nembedder and LanceDB. Create a workspace in Query mode, and after uploading use <strong>Move to\nWorkspace → Save and Embed</strong>. <em>Check:</em> a question whose answer is in one document comes back with a\ncitation.</p>\n<p><strong>6. VS Code autocomplete.</strong> Install Twinny and add one fill-in-middle provider: Ollama on\n<code>localhost:11434</code>, path <code>/api/generate</code>, model <code>qwen2.5-coder:7b-base</code>, template automatic. Leave out\nTwinny&#39;s chat. In settings, add <code>&quot;twinny.keepAlive&quot;: &quot;30m&quot;</code> and <code>&quot;chat.disableAIFeatures&quot;: true</code>.\nKeep error hints deterministic with a linter and type checker (I use Ruff, Pylance and ErrorLens).\n<em>Check:</em> suggestions within a second once the model is loaded.</p>\n<p>Tested versions: Ollama 0.34.3, OpenCode 1.18.32, T3 Code 0.0.42, Jan 0.8.4, AnythingLLM 1.16.1, Twinny 4.2.5 (it has since updated itself to 4.2.7).</p>\n<h2 id=\"limits\">Limits</h2>\n<ul><li>One machine, small fixtures, mostly single runs. Results varied between runs: the same 9B that scored 12/12 deleted a test on a repeat.</li><li>The 32K context fills up on bigger tasks: the largest runs came close to the limit and compacted.</li><li>T3&#39;s &quot;Auto-accept edits&quot; mode was never exercised with a local model. Every run used Full access.</li><li>Local models have no web access, so anything that depends on recent releases or current documentation goes to a cloud model.</li><li>Apps update themselves. Twinny updated during testing and again before I wrote this. Re-run the checks after an update.</li></ul>\n<h2 id=\"what-i-didn-t-use-and-why\">What I didn&#39;t use, and why</h2>\n<ul><li><p><strong>Other runtimes.</strong> I did not compare Ollama with LM Studio or</p><p>llama.cpp&#39;s own server. Every app here could talk to\nOllama, so I kept one runtime and spent the time testing the apps. Another runtime may be faster\non the same card; I haven&#39;t measured it.</p></li><li><strong>Open WebUI.</strong> It needs Docker, which I didn&#39;t want on this machine.</li><li><p><strong>An agent inside the editor.</strong> Cline was my main assistant in the earlier setup; I replaced it</p><p>with T3, where agent work happens in its own threads and I review the diffs. VS Code keeps\nautocomplete and deterministic hints only. I have not tested Continue in this setup.</p></li><li><p><strong>A model as a linter.</strong> Ruff, Pylance and ErrorLens are instant and never invent a rule. A model</p><p>would be slower and less predictable at the same job.</p></li><li><strong>A code index.</strong> Tested above: it lost to the agent searching the files live.</li></ul>\n<h2 id=\"what-i-would-tell-myself-at-the-start\">What I would tell myself at the start</h2>\n<ol><li>Fix the context length before judging any model.</li><li>Run a control task before blaming a model for a failure.</li><li>Trust files, diffs and tests over summaries, including your own notes.</li><li>Change one thing per retry, and keep the failures.</li><li>Fewer models, each with a job, beat a large collection nobody has tested in the apps you actually use.</li></ol>\n<h2 id=\"about-the-author\">About the author</h2>\n<p>I&#39;m Akash Damle, founder of Matalli Infotech Private Limited, an early-stage company I&#39;m building from the ground up. I write to mark milestones: what was built, what broke, and what it taught me.</p>\n<p>If you follow this setup, especially on different hardware, I would like to hear how it went: what you ran, what happened, and what you changed. Disagreement is as welcome as agreement, as long as it comes with a reason. Reports like that are how the next version of this guide gets better.</p>\n<h2 id=\"comments\">Comments</h2>\n<p>Tried this setup, hit a different result, or think something here is wrong? Say what you ran, what happened, and why. Agreement and disagreement are equally welcome; one-line reactions are not published. No account needed.</p>","headings":[{"level":2,"text":"In short","id":"in-short"},{"level":2,"text":"Why I did this","id":"why-i-did-this"},{"level":2,"text":"The machine, and the one number that matters most","id":"the-machine-and-the-one-number-that-matters-most"},{"level":3,"text":"Context length is the memory you don't see","id":"context-length-is-the-memory-you-don-t-see"},{"level":3,"text":"\"Too big\" depends on the architecture","id":"too-big-depends-on-the-architecture"},{"level":2,"text":"How I tested","id":"how-i-tested"},{"level":2,"text":"What I ended up with","id":"what-i-ended-up-with"},{"level":2,"text":"What failed, and what each failure taught me","id":"what-failed-and-what-each-failure-taught-me"},{"level":3,"text":"The coding agent that compacted forever","id":"the-coding-agent-that-compacted-forever"},{"level":3,"text":"Models that said \"done\"","id":"models-that-said-done"},{"level":3,"text":"Three failures with three different causes","id":"three-failures-with-three-different-causes"},{"level":3,"text":"The autocomplete that wasn't local","id":"the-autocomplete-that-wasn-t-local"},{"level":3,"text":"Uploading is not embedding","id":"uploading-is-not-embedding"},{"level":3,"text":"A code index lost to plain search","id":"a-code-index-lost-to-plain-search"},{"level":3,"text":"A tool call written as plain text","id":"a-tool-call-written-as-plain-text"},{"level":2,"text":"Coding workflows that held up","id":"coding-workflows-that-held-up"},{"level":2,"text":"Reproduce it","id":"reproduce-it"},{"level":2,"text":"Limits","id":"limits"},{"level":2,"text":"What I didn't use, and why","id":"what-i-didn-t-use-and-why"},{"level":2,"text":"What I would tell myself at the start","id":"what-i-would-tell-myself-at-the-start"},{"level":2,"text":"About the author","id":"about-the-author"},{"level":2,"text":"Comments","id":"comments"}]}}