{"article":{"slug":"one-month-coding-with-glm-5-3-flash","title":"One month coding with GLM 5.3 Flash","subtitle":null,"summary":"Reporting on 2B tokens of AI usage after a month coding only with GLM 5.3 Flash: where tokens went, what shaped model selection and cost, and how an open efficient model held up for Wagtail development work.","content_type":"blog_post","language":"en","canonical_url":"https://wagtail.org/blog/one-month-on-glm-53-flash/","author":{"name":null,"url":null,"person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"Wagtail","url":"https://wagtail.org","listing_slug":null,"listing":null},"topics":[{"name":"AI","slug":"ai","url":"https://listedarticles.com/topics/ai"},{"name":"LLMs","slug":"llms","url":"https://listedarticles.com/topics/llms"},{"name":"Open Source","slug":"open-source","url":"https://listedarticles.com/topics/open-source"},{"name":"Developer Tools","slug":"developer-tools","url":"https://listedarticles.com/topics/developer-tools"},{"name":"Programming","slug":"programming","url":"https://listedarticles.com/topics/programming"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":735,"reading_minutes":3,"published_at":"2026-10-02T00:00:00.000Z","added_at":"2026-10-03T00:11:50.848Z","updated_at":"2026-10-03T00:11:50.848Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":false},"profile_url":"https://listedarticles.com/articles/one-month-coding-with-glm-5-3-flash","markdown_url":"https://listedarticles.com/articles/one-month-coding-with-glm-5-3-flash.md","example":false,"citation":"Wagtail. \"One month coding with GLM 5.3 Flash.\" 2 Oct 2026. https://wagtail.org/blog/one-month-on-glm-53-flash/ (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://wagtail.org/blog/one-month-on-glm-53-flash/"},"body_markdown":"# One month coding with GLM 5.3 Flash\n\nSetting a challenge to spend the [whole of September on only one efficient open model](/blog/open-models-only-adoption-challenge/) felt like a great idea at the time. Turns out not so much in practice. 2B tokens later, here’s how it went.\n\n## Where tokens went this month\n\nHere’s the tokens distribution according to AgentsView, one of our [Agentic engineering recommendations](/ai/agentic-engineering-recommendations/) to keep tabs on AI usage:\n\nZooming in on the models split specifically:\n\nThe goal was to spend the whole month on GLM 5.3 Flash pictured in teal. Here’s what went well:\n\n  * Successfully spent the first half of the month on just that model.\n  * That model’s usage was well within our budget ($68, about 4kWh of energy use / 365 grams of carbon emissions).\n\nThe second half of the month didn’t go so well, with 1B tokens going to other models.\n\n## Unexpected hurdles\n\n### The cost of vibe coding\n\nWe’re pretty transparent that our [experimental Wagtail MCP server is a vibe-coded prototype](/blog/experiments-with-mcp-in-wagtail/). Vibe coding isn’t quite what we normally aspire to, but for a prototype it’s spot on. Unfortunately there are still consequences to it. I chose the 'wrong' model for the prototype, and we spent 450M tokens / $150 / 5kWh of energy use almost overnight. The MCP server itself works well and we now have a great demo of the capabilities, so it’s not for nothing:\n\nNonetheless, it’s a good reminder to be careful with model selection and with agentic patterns. We could have achieved similar results for most likely 5x less cost with not that much more effort. Lessons learned! We need to budget for this, and be more careful. Could have seen it coming, but now we know.\n\n### Infrastructure woes\n\nAnother unexpected hurdle was infrastructure availability issues. We’ve written extensively about [comparing inference providers](/blog/comparing-open-weight-ai-models-and-providers/). Our choices work really most of the times, but it turns out they’re very popular, and do not have the same capacity as the big labs who hoard all the GPUs. We noted degradation with the performance of GLM 5.3 Flash in particular, most likely because of it being [so high up](/blog/open-models-only-adoption-challenge/) the Pareto frontier of relevant models for our work.\n\nThis meant having to switch to other similar models (DeepSeek V4.1 Flash, Qwen 3.8 Flash). Which is very simple to do, but nonetheless unexpected!\n\n### The cost of experimentation and R&D\n\nLast but not least, beyond using one model for day-to-day engineering, it felt essential to keep experimenting with a wide range of models, keeping up with what providers are releasing. This is particularly essential as we start to benchmark models’ performance on Wagtail tasks, where we need data across a wide range of models. Sneak peek of our benchmark:\n\nIt’s much easier to guide people towards leaner options with this kind of concrete data. And for us to make those options even more viable with agent skills, or [our new CLI prototype](/blog/prototyping-a-cli-for-wagtail/), which is intended to work well with agents.\n\n## Takeways and what to do next\n\nSo technically this challenge was a failure. Only 50% usage on the target model, 1B out of 2B tokens. About 35 kWh of energy use instead of 10. But we did learn a lot, which is crucial for the current moment. Reflecting on this for October, here’s what will make it work:\n\n  1. **Constant, local usage measurement and reporting**. Looking not just at tokens but also energy use and spend, and ideally how well this all leads to concrete positive outcomes.\n  2. **Budgeting for experimentation, not just day-to-day tasks**. Making more concerted decisions about which prototypes are worth building, and how.\n  3. **Better prompt selection and multi-agent techniques**. Orchestrator vs. scout vs. implementer vs. reviewer agents. Bounded goals. Not rocket science but certainly one more thing to learn.\n  4. **Keep pushing for more efficient techniques and models**. The Jev-style decision diffusion models look very promising if they can run so efficiently. Latest flagship models also look like a step in the right direction on that front.\n\nFor day-to-day developer work, it’s totally viable to focus on one or two flash-tier cheap models. A viable target is probably that the _majority_ of AI inference work should be done with such efficient models, measured in cost or energy use rather than meaningless tokens. That’s the goal for October! You should try it too, you’ll learn a lot in the process.\n\n* * *\n\nAnd come say hi at [Wagtail Space 2026](/wagtail-space-2026/) in November to hear how that all pans out!","body_html":"<h1 id=\"one-month-coding-with-glm-5-3-flash\">One month coding with GLM 5.3 Flash</h1>\n<p>Setting a challenge to spend the <a href=\"/blog/open-models-only-adoption-challenge/\">whole of September on only one efficient open model</a> felt like a great idea at the time. Turns out not so much in practice. 2B tokens later, here’s how it went.</p>\n<h2 id=\"where-tokens-went-this-month\">Where tokens went this month</h2>\n<p>Here’s the tokens distribution according to AgentsView, one of our <a href=\"/ai/agentic-engineering-recommendations/\">Agentic engineering recommendations</a> to keep tabs on AI usage:</p>\n<p>Zooming in on the models split specifically:</p>\n<p>The goal was to spend the whole month on GLM 5.3 Flash pictured in teal. Here’s what went well:</p>\n<ul><li>Successfully spent the first half of the month on just that model.</li><li>That model’s usage was well within our budget ($68, about 4kWh of energy use / 365 grams of carbon emissions).</li></ul>\n<p>The second half of the month didn’t go so well, with 1B tokens going to other models.</p>\n<h2 id=\"unexpected-hurdles\">Unexpected hurdles</h2>\n<h3 id=\"the-cost-of-vibe-coding\">The cost of vibe coding</h3>\n<p>We’re pretty transparent that our <a href=\"/blog/experiments-with-mcp-in-wagtail/\">experimental Wagtail MCP server is a vibe-coded prototype</a>. Vibe coding isn’t quite what we normally aspire to, but for a prototype it’s spot on. Unfortunately there are still consequences to it. I chose the &#39;wrong&#39; model for the prototype, and we spent 450M tokens / $150 / 5kWh of energy use almost overnight. The MCP server itself works well and we now have a great demo of the capabilities, so it’s not for nothing:</p>\n<p>Nonetheless, it’s a good reminder to be careful with model selection and with agentic patterns. We could have achieved similar results for most likely 5x less cost with not that much more effort. Lessons learned! We need to budget for this, and be more careful. Could have seen it coming, but now we know.</p>\n<h3 id=\"infrastructure-woes\">Infrastructure woes</h3>\n<p>Another unexpected hurdle was infrastructure availability issues. We’ve written extensively about <a href=\"/blog/comparing-open-weight-ai-models-and-providers/\">comparing inference providers</a>. Our choices work really most of the times, but it turns out they’re very popular, and do not have the same capacity as the big labs who hoard all the GPUs. We noted degradation with the performance of GLM 5.3 Flash in particular, most likely because of it being <a href=\"/blog/open-models-only-adoption-challenge/\">so high up</a> the Pareto frontier of relevant models for our work.</p>\n<p>This meant having to switch to other similar models (DeepSeek V4.1 Flash, Qwen 3.8 Flash). Which is very simple to do, but nonetheless unexpected!</p>\n<h3 id=\"the-cost-of-experimentation-and-r-d\">The cost of experimentation and R&amp;D</h3>\n<p>Last but not least, beyond using one model for day-to-day engineering, it felt essential to keep experimenting with a wide range of models, keeping up with what providers are releasing. This is particularly essential as we start to benchmark models’ performance on Wagtail tasks, where we need data across a wide range of models. Sneak peek of our benchmark:</p>\n<p>It’s much easier to guide people towards leaner options with this kind of concrete data. And for us to make those options even more viable with agent skills, or <a href=\"/blog/prototyping-a-cli-for-wagtail/\">our new CLI prototype</a>, which is intended to work well with agents.</p>\n<h2 id=\"takeways-and-what-to-do-next\">Takeways and what to do next</h2>\n<p>So technically this challenge was a failure. Only 50% usage on the target model, 1B out of 2B tokens. About 35 kWh of energy use instead of 10. But we did learn a lot, which is crucial for the current moment. Reflecting on this for October, here’s what will make it work:</p>\n<ol><li><strong>Constant, local usage measurement and reporting</strong>. Looking not just at tokens but also energy use and spend, and ideally how well this all leads to concrete positive outcomes.</li><li><strong>Budgeting for experimentation, not just day-to-day tasks</strong>. Making more concerted decisions about which prototypes are worth building, and how.</li><li><strong>Better prompt selection and multi-agent techniques</strong>. Orchestrator vs. scout vs. implementer vs. reviewer agents. Bounded goals. Not rocket science but certainly one more thing to learn.</li><li><strong>Keep pushing for more efficient techniques and models</strong>. The Jev-style decision diffusion models look very promising if they can run so efficiently. Latest flagship models also look like a step in the right direction on that front.</li></ol>\n<p>For day-to-day developer work, it’s totally viable to focus on one or two flash-tier cheap models. A viable target is probably that the <em>majority</em> of AI inference work should be done with such efficient models, measured in cost or energy use rather than meaningless tokens. That’s the goal for October! You should try it too, you’ll learn a lot in the process.</p>\n<ul><li>* *</li></ul>\n<p>And come say hi at <a href=\"/wagtail-space-2026/\">Wagtail Space 2026</a> in November to hear how that all pans out!</p>","headings":[{"level":1,"text":"One month coding with GLM 5.3 Flash","id":"one-month-coding-with-glm-5-3-flash"},{"level":2,"text":"Where tokens went this month","id":"where-tokens-went-this-month"},{"level":2,"text":"Unexpected hurdles","id":"unexpected-hurdles"},{"level":3,"text":"The cost of vibe coding","id":"the-cost-of-vibe-coding"},{"level":3,"text":"Infrastructure woes","id":"infrastructure-woes"},{"level":3,"text":"The cost of experimentation and R&D","id":"the-cost-of-experimentation-and-r-d"},{"level":2,"text":"Takeways and what to do next","id":"takeways-and-what-to-do-next"}]}}