In short

  • Context length decides whether a model fits. The same 5.3 GB model took 18 GB and spilled onto

    the CPU at a 131K context, but used 6.8 GB entirely on the GPU at 16K.

  • Three models cover everything I need: a 7B model for autocomplete, Qwen 3.5 9B for chat,

    documents and small code changes, and Qwen 3 Coder 30B, a mixture-of-experts model that runs well despite not fitting in VRAM, for larger coding tasks.

  • The coding agent needs a 32K context and a small output reserve, or it summarises its own

    conversation in an endless loop.

  • The dangerous failures looked like successes. Models reported work as done when nothing was

    saved, or when the saved file still had errors. Check files, diffs and tests, not summaries.

  • A six-step setup, with a check for each step, is at the end.