Ox Alpha has been revealed as Z.AI GLM-5.3 Flash. You can use it on StudyArena now
StudyArena has updated the leaderboard to reflect the official name and creator. Moving forward, new prompt requests route to z-ai/glm-5.3-flash, while existing stealth/ox-alpha* entries remain right where they are—keeping all their hard-earned Elo and battle history intact.
What is GLM-5.3 Flash?
GLM-5.3 Flash marks the first natively multimodal release in Z.AI's GLM-5 lineup. Built as a 320B-total, 18B-active mixture-of-experts architecture, it was trained on an enormous 30-trillion-token multimodal dataset. Across its 45 layers, it routes every token through eight out of 288 specialized experts.[4][5]
Here is a quick look at the specs:
| Technical detail | GLM-5.3 Flash |
|---|
| Model type | Mixture of experts, 320B total parameters |
| Active compute | 18B parameters per token |
| Architecture | Hybrid linear and sparse attention, plus mHC |
| Published context | 1,048,576 positions |
| Inputs | Text, images, and video |
| Output | Text |
| Reasoning effort | Low, High, or Max |
| Weights | Publicly available under the MIT License |
Right now on StudyArena, GLM-5.3 Flash is competing in text and image evaluations across Low, High, and Max reasoning levels. We still have tools and web browsing turned off for this specific contestant, even though the base model on OpenRouter actually supports function calling and JSON output.[1]
Fun Facts
1. Most of the model sleeps during any single token
Only 18B out of the 320B parameters run per token—just about 5.6% of the overall system. That is the whole advantage of a mixture-of-experts design: you maintain a massive pool of specialized weights without burning compute on every single parameter for every token.
2. "Flash" refers to inference speed, not short responses
Artificial Analysis clocked it at around 49.8 output tokens per second with a 1.51-second time-to-first-token running directly on Z.AI. Interestingly, they also noted it tends to be more talkative than the average open-weight model.[6] A model can be lean on compute per token without holding back on word count.
3. The stealth preview ran entirely on domestic Chinese silicon
According to Z.AI, all the hidden Ox Alpha traffic was processed on domestic Chinese AI hardware.[2] That turns the anonymous run into something even cooler than a marketing trick: it was a real-world stress test of their native infrastructure.
Give the model another go on StudyArena now.
Sources
Sources are listed in citation order. Access dates show when StudyArena last checked each source.
- 1.GLM 5.3 Flash - API Pricing & Benchmarks · OpenRouter (2026) · Accessed August 29, 2026
- 2.GLM-5.3-Flash: Frontier Intelligence, Flash Cost · Z.AI (2026) · Accessed August 29, 2026
- 3.StudyArena live AI leaderboard · StudyArena (2026) · Accessed August 29, 2026
- 4.GLM-5.3-Flash model card · Z.AI on Hugging Face (2026) · Accessed August 29, 2026
- 5.GLM-5.3-Flash model configuration · Z.AI on Hugging Face (2026) · Accessed August 29, 2026
- 6.GLM-5.3-Flash Intelligence, Performance & Price Analysis · Artificial Analysis (2026) · Accessed August 29, 2026
- 7.Surprise: Z.ai is the AI lab behind the mysterious Ox Alpha model · TechCrunch (2026) · Accessed August 29, 2026
- 8.Z.ai Releases GLM-5.3-Flash: A 320B-A18B Natively Multimodal MoE With a 1M-Token Context · MarkTechPost (2026) · Accessed August 29, 2026
Update history
- August 29, 2026 · deslop
- August 28, 2026 · Updated article details
- August 28, 2026 · Removed stale references and updated fun facts
- August 28, 2026 · Removed inline bold markers from the generated key-summary list.
- August 28, 2026 · Initial release note for the Ox Alpha identity reveal.
About this article
- Written by Pennie Li.
- Reviewed by Pasha Rayan on August 28, 2026.
- Includes 8 cited sources.
- Published August 28, 2026 and updated August 29, 2026.
Editorial policy