{"article":{"slug":"model-reveal-ox-alpha-is-z-ai-glm-5-3-flash-live-on-studyarena-now","title":"Model Reveal! Ox Alpha Is Z.AI GLM-5.3 Flash! Live on StudyArena now","subtitle":null,"summary":"Ox Alpha has been revealed as Z.AI GLM-5.3 Flash. You can use it on StudyArena now","content_type":"announcement","language":"en","canonical_url":"https://studyarena.com/blog/ox-alpha-is-z-ai-glm-5-3-flash","author":{"name":"Pennie Li","url":"https://studyarena.com/blog/authors/pennie-li","person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"StudyArena","url":"https://studyarena.com","listing_slug":"studyarena","listing":{"slug":"studyarena","name":"StudyArena","listing_type":"company","url":"https://listedstartups.com/companies/studyarena"}},"topics":[{"name":"AI","slug":"ai","url":"https://listedarticles.com/topics/ai"},{"name":"LLMs","slug":"llms","url":"https://listedarticles.com/topics/llms"},{"name":"open-weight AI","slug":"open-weight-ai","url":"https://listedarticles.com/topics/open-weight-ai"},{"name":"Education","slug":"education","url":"https://listedarticles.com/topics/education"}],"about_listings":[{"slug":"studyarena-platform","name":"StudyArena","listing_type":"product","url":"https://listedstartups.com/products/studyarena-platform"}],"cover_image_url":"https://studyarena.com/blog/heroes/ox-alpha-z-ai-glm-5-3-flash.png","license":"all-rights-reserved","word_count":726,"reading_minutes":3,"published_at":"2026-08-28T14:23:11.889Z","added_at":"2026-09-17T05:04:24.403Z","updated_at":"2026-09-17T05:04:24.403Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":true},"profile_url":"https://listedarticles.com/articles/model-reveal-ox-alpha-is-z-ai-glm-5-3-flash-live-on-studyarena-now","markdown_url":"https://listedarticles.com/articles/model-reveal-ox-alpha-is-z-ai-glm-5-3-flash-live-on-studyarena-now.md","example":false,"citation":"Pennie Li, StudyArena. \"Model Reveal! Ox Alpha Is Z.AI GLM-5.3 Flash! Live on StudyArena now.\" 28 Aug 2026. https://studyarena.com/blog/ox-alpha-is-z-ai-glm-5-3-flash (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://studyarena.com/blog/ox-alpha-is-z-ai-glm-5-3-flash"},"body_markdown":"# Model Reveal! Ox Alpha Is Z.AI GLM-5.3 Flash! Live on StudyArena now\n\nOx Alpha has been revealed as Z.AI GLM-5.3 Flash. You can use it on StudyArena now\n\n![](https://studyarena.com/api/blog/authors/pennie-li/portrait/eac85c61aaf4447387f90b18eea46230)Written byPennie Li\n\n![](https://studyarena.com/api/blog/authors/pasha-rayan/portrait/8e8ca70c4d574570a1ef7c32504a738d)Reviewed byPasha Rayan\n\nPublished August 28, 2026 · Updated August 29, 2026 · 2 min read\n\n## Key takeaways\n\n  * Ox Alpha is now labeled as Z.AI GLM-5.3 Flash across StudyArena.\n  * Future queries use the official OpenRouter route, but the original three StudyArena contestant IDs will hold onto their past battle records.\n  * Under the hood, GLM-5.3 Flash features a 320-billion-parameter mixture-of-experts setup (with 18 billion active per token), open weights, multimodal input, and a massive 1-million-token context window.\n\n![Ox Alpha identity card revealing Z.AI GLM-5.3 Flash, connected by an unbroken Elo history line.](https://studyarena.com/blog/heroes/ox-alpha-z-ai-glm-5-3-flash.png)\n\n**Ox Alpha** has been revealed to be Z.AI's **GLM-5.3 Flash**![1] Z.AI ran it stealth on OpenRouter and OpenCode prior to launch, treating real user traffic as a blind test.[2]\n\nStudyArena has updated the leaderboard to reflect the official name and creator. Moving forward, new prompt requests route to `z-ai/glm-5.3-flash`, while existing `stealth/ox-alpha*` entries remain right where they are—keeping all their hard-earned Elo and battle history intact.\n\n## What is GLM-5.3 Flash?\n\nGLM-5.3 Flash marks the first natively multimodal release in Z.AI's GLM-5 lineup. Built as a **320B-total, 18B-active** mixture-of-experts architecture, it was trained on an enormous 30-trillion-token multimodal dataset. Across its 45 layers, it routes every token through eight out of 288 specialized experts.[4][5]\n\nHere is a quick look at the specs:\n\nTechnical detail| GLM-5.3 Flash  \n---|---  \nModel type| Mixture of experts, 320B total parameters  \nActive compute| 18B parameters per token  \nArchitecture| Hybrid linear and sparse attention, plus mHC  \nPublished context| 1,048,576 positions  \nInputs| Text, images, and video  \nOutput| Text  \nReasoning effort| Low, High, or Max  \nWeights| Publicly available under the MIT License  \n  \nRight now on StudyArena, GLM-5.3 Flash is competing in text and image evaluations across Low, High, and Max reasoning levels. We still have tools and web browsing turned off for this specific contestant, even though the base model on OpenRouter actually supports function calling and JSON output.[1]\n\n### Fun Facts\n\n### 1\\. Most of the model sleeps during any single token\n\nOnly 18B out of the 320B parameters run per token—just about 5.6% of the overall system. That is the whole advantage of a mixture-of-experts design: you maintain a massive pool of specialized weights without burning compute on every single parameter for every token.\n\n### 2\\. \"Flash\" refers to inference speed, not short responses\n\nArtificial Analysis clocked it at around 49.8 output tokens per second with a 1.51-second time-to-first-token running directly on Z.AI. Interestingly, they also noted it tends to be more talkative than the average open-weight model.[6] A model can be lean on compute per token without holding back on word count.\n\n### 3\\. The stealth preview ran entirely on domestic Chinese silicon\n\nAccording to Z.AI, all the hidden Ox Alpha traffic was processed on domestic Chinese AI hardware.[2] That turns the anonymous run into something even cooler than a marketing trick: it was a real-world stress test of their native infrastructure.\n\nGive the model another go on StudyArena now.\n\n## Sources\n\nSources are listed in citation order. Access dates show when StudyArena last checked each source.\n\n  1. 1.[GLM 5.3 Flash - API Pricing & Benchmarks](<https://openrouter.ai/z-ai/glm-5.3-flash>) · OpenRouter (2026) · Accessed August 29, 2026\n  2. 2.[GLM-5.3-Flash: Frontier Intelligence, Flash Cost](<https://z.ai/blog/glm-5.3-flash>) · Z.AI (2026) · Accessed August 29, 2026\n  3. 3.[StudyArena live AI leaderboard](<https://studyarena.com/leaderboard>) · StudyArena (2026) · Accessed August 29, 2026\n  4. 4.[GLM-5.3-Flash model card](<https://huggingface.co/zai-org/GLM-5.3-Flash>) · Z.AI on Hugging Face (2026) · Accessed August 29, 2026\n  5. 5.[GLM-5.3-Flash model configuration](<https://huggingface.co/zai-org/GLM-5.3-Flash/raw/main/config.json>) · Z.AI on Hugging Face (2026) · Accessed August 29, 2026\n  6. 6.[GLM-5.3-Flash Intelligence, Performance & Price Analysis](<https://artificialanalysis.ai/models/glm-5-3-flash/>) · Artificial Analysis (2026) · Accessed August 29, 2026\n  7. 7.[Surprise: Z.ai is the AI lab behind the mysterious Ox Alpha model](<https://techcrunch.com/2026/08/26/surprise-z-ai-is-the-ai-lab-behind-the-mysterious-ox-alpha-model/>) · TechCrunch (2026) · Accessed August 29, 2026\n  8. 8.[Z.ai Releases GLM-5.3-Flash: A 320B-A18B Natively Multimodal MoE With a 1M-Token Context](<https://www.marktechpost.com/2026/08/26/z-ai-releases-glm-5-3-flash-a-320b-a18b-natively-multimodal-moe-with-a-1m-token-context/>) · MarkTechPost (2026) · Accessed August 29, 2026\n\nUpdate history\n\n  * August 29, 2026 · deslop\n  * August 28, 2026 · Updated article details\n  * August 28, 2026 · Removed stale references and updated fun facts\n  * August 28, 2026 · Removed inline bold markers from the generated key-summary list.\n  * August 28, 2026 · Initial release note for the Ox Alpha identity reveal.\n\n\n\nAbout this article\n\n  * Written by Pennie Li.\n  * Reviewed by Pasha Rayan on August 28, 2026.\n  * Includes 8 cited sources.\n  * Published August 28, 2026 and updated August 29, 2026.\n\n[Editorial policy](</blog/editorial-policy>)","body_html":"<h1 id=\"model-reveal-ox-alpha-is-z-ai-glm-5-3-flash-live-on-studyarena-n\">Model Reveal! Ox Alpha Is Z.AI GLM-5.3 Flash! Live on StudyArena now</h1>\n<p>Ox Alpha has been revealed as Z.AI GLM-5.3 Flash. You can use it on StudyArena now</p>\n<p><img src=\"https://studyarena.com/api/blog/authors/pennie-li/portrait/eac85c61aaf4447387f90b18eea46230\" alt=\"\" loading=\"lazy\" decoding=\"async\" referrerpolicy=\"no-referrer\" />Written byPennie Li</p>\n<p><img src=\"https://studyarena.com/api/blog/authors/pasha-rayan/portrait/8e8ca70c4d574570a1ef7c32504a738d\" alt=\"\" loading=\"lazy\" decoding=\"async\" referrerpolicy=\"no-referrer\" />Reviewed byPasha Rayan</p>\n<p>Published August 28, 2026 · Updated August 29, 2026 · 2 min read</p>\n<h2 id=\"key-takeaways\">Key takeaways</h2>\n<ul><li>Ox Alpha is now labeled as Z.AI GLM-5.3 Flash across StudyArena.</li><li>Future queries use the official OpenRouter route, but the original three StudyArena contestant IDs will hold onto their past battle records.</li><li>Under the hood, GLM-5.3 Flash features a 320-billion-parameter mixture-of-experts setup (with 18 billion active per token), open weights, multimodal input, and a massive 1-million-token context window.</li></ul>\n<figure><img src=\"https://studyarena.com/blog/heroes/ox-alpha-z-ai-glm-5-3-flash.png\" alt=\"Ox Alpha identity card revealing Z.AI GLM-5.3 Flash, connected by an unbroken Elo history line.\" loading=\"lazy\" decoding=\"async\" referrerpolicy=\"no-referrer\" /></figure>\n<p><strong>Ox Alpha</strong> has been revealed to be Z.AI&#39;s <strong>GLM-5.3 Flash</strong>![1] Z.AI ran it stealth on OpenRouter and OpenCode prior to launch, treating real user traffic as a blind test.[2]</p>\n<p>StudyArena has updated the leaderboard to reflect the official name and creator. Moving forward, new prompt requests route to <code>z-ai/glm-5.3-flash</code>, while existing <code>stealth/ox-alpha*</code> entries remain right where they are—keeping all their hard-earned Elo and battle history intact.</p>\n<h2 id=\"what-is-glm-5-3-flash\">What is GLM-5.3 Flash?</h2>\n<p>GLM-5.3 Flash marks the first natively multimodal release in Z.AI&#39;s GLM-5 lineup. Built as a <strong>320B-total, 18B-active</strong> mixture-of-experts architecture, it was trained on an enormous 30-trillion-token multimodal dataset. Across its 45 layers, it routes every token through eight out of 288 specialized experts.[4][5]</p>\n<p>Here is a quick look at the specs:</p>\n<div class=\"table-wrap\"><table><thead><tr><th>Technical detail</th><th>GLM-5.3 Flash</th></tr></thead><tbody><tr><td>Model type</td><td>Mixture of experts, 320B total parameters</td></tr><tr><td>Active compute</td><td>18B parameters per token</td></tr><tr><td>Architecture</td><td>Hybrid linear and sparse attention, plus mHC</td></tr><tr><td>Published context</td><td>1,048,576 positions</td></tr><tr><td>Inputs</td><td>Text, images, and video</td></tr><tr><td>Output</td><td>Text</td></tr><tr><td>Reasoning effort</td><td>Low, High, or Max</td></tr><tr><td>Weights</td><td>Publicly available under the MIT License</td></tr></tbody></table></div>\n<p>Right now on StudyArena, GLM-5.3 Flash is competing in text and image evaluations across Low, High, and Max reasoning levels. We still have tools and web browsing turned off for this specific contestant, even though the base model on OpenRouter actually supports function calling and JSON output.[1]</p>\n<h3 id=\"fun-facts\">Fun Facts</h3>\n<h3 id=\"1-most-of-the-model-sleeps-during-any-single-token\">1. Most of the model sleeps during any single token</h3>\n<p>Only 18B out of the 320B parameters run per token—just about 5.6% of the overall system. That is the whole advantage of a mixture-of-experts design: you maintain a massive pool of specialized weights without burning compute on every single parameter for every token.</p>\n<h3 id=\"2-flash-refers-to-inference-speed-not-short-responses\">2. &quot;Flash&quot; refers to inference speed, not short responses</h3>\n<p>Artificial Analysis clocked it at around 49.8 output tokens per second with a 1.51-second time-to-first-token running directly on Z.AI. Interestingly, they also noted it tends to be more talkative than the average open-weight model.[6] A model can be lean on compute per token without holding back on word count.</p>\n<h3 id=\"3-the-stealth-preview-ran-entirely-on-domestic-chinese-silicon\">3. The stealth preview ran entirely on domestic Chinese silicon</h3>\n<p>According to Z.AI, all the hidden Ox Alpha traffic was processed on domestic Chinese AI hardware.[2] That turns the anonymous run into something even cooler than a marketing trick: it was a real-world stress test of their native infrastructure.</p>\n<p>Give the model another go on StudyArena now.</p>\n<h2 id=\"sources\">Sources</h2>\n<p>Sources are listed in citation order. Access dates show when StudyArena last checked each source.</p>\n<ol><li>1.<a href=\"https://openrouter.ai/z-ai/glm-5.3-flash\" rel=\"nofollow ugc noopener\">GLM 5.3 Flash - API Pricing &amp; Benchmarks</a> · OpenRouter (2026) · Accessed August 29, 2026</li><li>2.<a href=\"https://z.ai/blog/glm-5.3-flash\" rel=\"nofollow ugc noopener\">GLM-5.3-Flash: Frontier Intelligence, Flash Cost</a> · Z.AI (2026) · Accessed August 29, 2026</li><li>3.<a href=\"https://studyarena.com/leaderboard\" rel=\"nofollow ugc noopener\">StudyArena live AI leaderboard</a> · StudyArena (2026) · Accessed August 29, 2026</li><li>4.<a href=\"https://huggingface.co/zai-org/GLM-5.3-Flash\" rel=\"nofollow ugc noopener\">GLM-5.3-Flash model card</a> · Z.AI on Hugging Face (2026) · Accessed August 29, 2026</li><li>5.<a href=\"https://huggingface.co/zai-org/GLM-5.3-Flash/raw/main/config.json\" rel=\"nofollow ugc noopener\">GLM-5.3-Flash model configuration</a> · Z.AI on Hugging Face (2026) · Accessed August 29, 2026</li><li>6.<a href=\"https://artificialanalysis.ai/models/glm-5-3-flash/\" rel=\"nofollow ugc noopener\">GLM-5.3-Flash Intelligence, Performance &amp; Price Analysis</a> · Artificial Analysis (2026) · Accessed August 29, 2026</li><li>7.<a href=\"https://techcrunch.com/2026/08/26/surprise-z-ai-is-the-ai-lab-behind-the-mysterious-ox-alpha-model/\" rel=\"nofollow ugc noopener\">Surprise: Z.ai is the AI lab behind the mysterious Ox Alpha model</a> · TechCrunch (2026) · Accessed August 29, 2026</li><li>8.<a href=\"https://www.marktechpost.com/2026/08/26/z-ai-releases-glm-5-3-flash-a-320b-a18b-natively-multimodal-moe-with-a-1m-token-context/\" rel=\"nofollow ugc noopener\">Z.ai Releases GLM-5.3-Flash: A 320B-A18B Natively Multimodal MoE With a 1M-Token Context</a> · MarkTechPost (2026) · Accessed August 29, 2026</li></ol>\n<p>Update history</p>\n<ul><li>August 29, 2026 · deslop</li><li>August 28, 2026 · Updated article details</li><li>August 28, 2026 · Removed stale references and updated fun facts</li><li>August 28, 2026 · Removed inline bold markers from the generated key-summary list.</li><li>August 28, 2026 · Initial release note for the Ox Alpha identity reveal.</li></ul>\n<p>About this article</p>\n<ul><li>Written by Pennie Li.</li><li>Reviewed by Pasha Rayan on August 28, 2026.</li><li>Includes 8 cited sources.</li><li>Published August 28, 2026 and updated August 29, 2026.</li></ul>\n<p><a href=\"/blog/editorial-policy\">Editorial policy</a></p>","headings":[{"level":1,"text":"Model Reveal! Ox Alpha Is Z.AI GLM-5.3 Flash! Live on StudyArena now","id":"model-reveal-ox-alpha-is-z-ai-glm-5-3-flash-live-on-studyarena-n"},{"level":2,"text":"Key takeaways","id":"key-takeaways"},{"level":2,"text":"What is GLM-5.3 Flash?","id":"what-is-glm-5-3-flash"},{"level":3,"text":"Fun Facts","id":"fun-facts"},{"level":3,"text":"1. Most of the model sleeps during any single token","id":"1-most-of-the-model-sleeps-during-any-single-token"},{"level":3,"text":"2. \"Flash\" refers to inference speed, not short responses","id":"2-flash-refers-to-inference-speed-not-short-responses"},{"level":3,"text":"3. The stealth preview ran entirely on domestic Chinese silicon","id":"3-the-stealth-preview-ran-entirely-on-domestic-chinese-silicon"},{"level":2,"text":"Sources","id":"sources"}]}}