Five New AI Models Are Live on StudyArena
Gemini 3.8 Flash, Hy4 Preview, Muse Spark 1.3, Mercury 2.5 Preview, and Granite 4.2 8B are now available for blind comparisons on StudyArena.
Written byPasha Rayan
Reviewed byPennie Li
Published September 4, 2026 · 5 min read
Key takeaways
- Five new models from Google, Tencent, Meta, Inception, and IBM are available in StudyArena's custom model picker.
- Their strengths are genuinely different: multimodal work, million-token context, agent workflows, fast diffusion generation, and compact open-weight reasoning.
- StudyArena exposes supported reasoning levels as separate contestants, so you can test whether more thinking actually produces a better answer.
- Muse Spark 1.3 is available for named comparisons first; the other four can also appear in automatic arena lineups.

Five new AI models are now available to compare on StudyArena: Gemini 3.8 Flash , Hy4 Preview , Muse Spark 1.3 , Mercury 2.5 Preview , and Granite 4.2 8B.
This is an unusually varied release group. It includes a fast multimodal workhorse, two open-weight models at opposite ends of the size spectrum, a model designed for long-running agent work, and a diffusion language model built around speed. Instead of asking which one has the most impressive launch chart, you can now give them the same question and judge their answers blind.
Five models at a glance
| Model | What stands out | Published context | StudyArena choices |
|---|---|---|---|
| Gemini 3.8 Flash | Broad multimodal work and long-context tasks | 1M tokens | Low, Medium, High |
| Hy4 Preview | Huge open-weight mixture-of-experts model | 1M tokens | Non-thinking, Low, High |
| Muse Spark 1.3 | Long-running agent and coding workflows | 1M tokens | Minimal, Low, Medium, High, XHigh |
| Mercury 2.5 Preview | Parallel diffusion generation for fast iteration | 260K tokens | Non-thinking, Low, Medium, High |
| Granite 4.2 8B | Compact open-weight reasoning baseline | 128K tokens on the current route | Non-thinking, Low, High |
Gemini 3.8 Flash: a multimodal workhorse
Google positions Gemini 3.8 Flash as its general-purpose workhorse for complex, multi-step tasks. It accepts text, images, video, audio, and PDFs, with a 1,048,576-token context window. Google also exposes Low, Medium, and High reasoning levels, with Medium as the default.[1][2]
On StudyArena, those three reasoning levels compete separately. That makes Gemini especially interesting for questions where extra deliberation may or may not be worth it: interpreting a dense reading, planning an assignment, debugging from a screenshot, or connecting evidence across a long PDF. Run the same prompt in separate Low- and High-reasoning rounds and see whether the extra reasoning adds substance or just length.
Hy4 Preview: enormous and open
Tencent's Hy4 Preview is a 770-billion-parameter mixture-of-experts model that activates 49 billion parameters for each token. It is open-weight under Apache 2.0, text-only, and supports a million-token context window. Tencent highlights coding, tool-driven productivity, cross-document office work, game prototyping, and scientific reasoning as target workloads.[3][4]
The word Preview matters. Tencent says the model can reason longer than necessary and sometimes over-check its own work. That is exactly the sort of behavior blind comparison can reveal. Use Hy4 for a difficult coding or synthesis task, then compare its answer with a smaller model: did the extra deliberation catch something important, or did it bury the useful part?
Muse Spark 1.3: built for longer-running agent work
Meta says Muse Spark 1.3 is better at preserving detailed requirements, navigating messy context, and collaborating with the user. In comparisons by Meta engineers with Muse Spark 1.2, it also used about 20% fewer tool calls and 25% fewer tokens.[5] OpenRouter lists the available model as multimodal reasoning for long-running agentic, multi-agent, and coding workflows.[6]
The current OpenRouter route exposes a million-token context window and reasoning from Minimal through XHigh.[6] We have made all five reasoning levels available in the picker, but Muse Spark 1.3 will not enter automatic random lineups yet. Its route availability was less consistent when we added it, so named comparisons are the honest starting point while routing settles.
Mercury 2.5 Preview: generating text differently
Mercury 2.5 Preview is the architectural outlier in this group. Inception calls it a diffusion language model: instead of producing an answer strictly one token after another, it can generate and refine groups of tokens in parallel. The model supports tunable reasoning, structured output, parallel tool calls, and a 260K context window.[7][8]
That design makes Mercury an interesting contestant for latency-sensitive work such as iterative coding, search, support, and short agent loops. Vendor speed benchmarks are useful context, but they are not a promise about how fast every StudyArena answer will arrive; network and provider routing still matter. The useful test is practical: does the answer arrive quickly and hold up next to slower competitors?
Granite 4.2 8B: a compact open baseline
IBM Granite 4.2 8B is the smallest model in this update, and that is the point. It is a dense, Apache-2.0 open-weight model with full-thinking, low-effort, and non-thinking modes. IBM positions the 8B member of the family as a balanced option for reasoning, code, mathematics, tool use, and multilingual work.[9]
The OpenRouter route currently gives Granite 4.2 8B a 128K context window.[10] In StudyArena, you can compare Non-thinking, Low, and High reasoning versions. It is a useful reality check against the assumption that every good answer needs frontier-scale compute: for a focused explanation, rewrite, or structured task, a compact model may be all you need.
Why reasoning levels appear as separate contestants
When a provider supports multiple reasoning levels, StudyArena treats those settings as distinct contestants. The underlying model is the same, but the answer behavior can change: higher effort may improve planning or verification, while lower effort can be clearer and more direct.
Keeping the variants separate also makes the vote history useful. A vote for a High-reasoning answer should not silently become evidence about the Low-reasoning configuration. Over time, the arena can show not only which model students prefer, but which version of that model works best for a particular kind of question.
Try the update
Supporters can open the model lineup on the StudyArena home page, choose three or six contestants, and search for any of the five names above. Free-plan comparisons use automatic lineups; Gemini 3.8 Flash, Hy4 Preview, Mercury 2.5 Preview, and Granite 4.2 8B can now be drawn there too. Muse Spark 1.3 remains a named choice while its routing settles.
Ask one real question, read the answers without knowing which model wrote them, and vote for the one that helped. Launch claims tell you what a model was designed to do. A blind comparison tells you whether it did it for your work.
Sources
Sources are listed in citation order. Access dates show when StudyArena last checked each source.
- 1.Gemini 3.8 Flash model documentation · Google AI for Developers (2026) · Accessed September 3, 2026
- 2.Introducing Gemini 3.8 Flash and Gemini 3.8 Flash Cyber · Google (2026) · Accessed September 3, 2026
- 3.Tencent releases and open-sources Tencent Hy4 Preview · Tencent (2026) · Accessed September 3, 2026
- 4.Tencent Hy4 Preview model card · Tencent on Hugging Face (2026) · Accessed September 3, 2026
- 5.Introducing Muse Spark 1.3 · Meta AI Research (2026) · Accessed September 3, 2026
- 6.Muse Spark 1.3 · OpenRouter (2026) · Accessed September 3, 2026
- 7.Mercury models · Inception Labs (2026) · Accessed September 3, 2026
- 8.Mercury 2.5 Preview · OpenRouter (2026) · Accessed September 3, 2026
- 9.Granite 4.2 8B model card · IBM on Hugging Face (2026) · Accessed September 3, 2026
- 10.Granite 4.2 8B · OpenRouter (2026) · Accessed September 3, 2026
About this article
- Written by Pasha Rayan.
- Reviewed by Pennie Li on September 4, 2026.
- Includes 10 cited sources.
- Published September 4, 2026.