{"article":{"slug":"which-ai-model-is-the-best-in-the-world-right-now-what-77-ai-models-think","title":"Which AI model is the best in the world right now? what 77 AI models think","subtitle":null,"summary":"We put \"Which AI model is the best in the world right now?\" to 77 AI models (GPT, Claude, Gemini, DeepSeek…) at the same time. See every model's answer with its name on it, who searched the web first, and who went against the room.","content_type":"blog_post","language":"en","canonical_url":"https://studyarena.com/every-ai/which-ai-is-the-best","author":{"name":"Pasha Rayan","url":"https://studyarena.com/blog/authors/pasha-rayan","person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"StudyArena","url":"https://studyarena.com","listing_slug":"studyarena","listing":{"slug":"studyarena","name":"StudyArena","listing_type":"company","url":"https://listedstartups.com/companies/studyarena"}},"topics":[{"name":"Benchmarks","slug":"benchmarks","url":"https://listedarticles.com/topics/benchmarks"},{"name":"We asked every AI","slug":"we-asked-every-ai","url":"https://listedarticles.com/topics/we-asked-every-ai"},{"name":"AI","slug":"ai","url":"https://listedarticles.com/topics/ai"},{"name":"Education","slug":"education","url":"https://listedarticles.com/topics/education"}],"about_listings":[{"slug":"studyarena-platform","name":"StudyArena","listing_type":"product","url":"https://listedstartups.com/products/studyarena-platform"}],"cover_image_url":"https://studyarena.com/every-ai/which-ai-is-the-best/opengraph-image","license":"all-rights-reserved","word_count":773,"reading_minutes":3,"published_at":"2026-09-16T15:00:00.000Z","added_at":"2026-09-17T05:04:10.707Z","updated_at":"2026-09-17T05:04:10.707Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":true},"profile_url":"https://listedarticles.com/articles/which-ai-model-is-the-best-in-the-world-right-now-what-77-ai-models-think","markdown_url":"https://listedarticles.com/articles/which-ai-model-is-the-best-in-the-world-right-now-what-77-ai-models-think.md","example":false,"citation":"Pasha Rayan, StudyArena. \"Which AI model is the best in the world right now? what 77 AI models think.\" 16 Sept 2026. https://studyarena.com/every-ai/which-ai-is-the-best (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://studyarena.com/every-ai/which-ai-is-the-best"},"body_markdown":"# Which AI model is the best in the world right now? what 77 AI models think\n\nWe put \"Which AI model is the best in the world right now?\" to 77 AI models (GPT, Claude, Gemini, DeepSeek…) at the same time. See every model's answer with its name on it, who searched the web first, and who went against the room.\n\n[← We asked every AI](</every-ai>)Live · 77/79 models\n\n![Illustration for \"Which AI model is the best in the world right now?\"](https://studyarena.com/every-ai/which-ai-is-the-best/cover)\n\nHot takeAI\n\n# Which AI model is the best in the world right now?\n\nAsked verbatim: “Which single AI model is the best in the world right now? Name one specific model (not a company).”\n\n## How the AIs voted\n\nOne dot per model. Hover for the model, its lab and its argument.\n\n16%GPT-4o\n\n12 models\n\n10%Claude 3.5 Sonnet\n\n8 models\n\n9%GPT-5\n\n7 models\n\n8%Claude Fable 5.1\n\n6 models\n\n6%GPT-6 Astra\n\n5 models\n\n51%Other answers\n\n39 models\n\nBy lab\n\n  * Google1 / 6 GPT-4o\n  * OpenAI2 / 5 GPT-4o\n  * Qwen0 / 5 GPT-4o\n  * Z.AI1 / 5 GPT-4o\n  * Aion Labs1 / 4 GPT-4o\n  * Anthropic0 / 4 GPT-4o\n  * 21 other labs7 / 48 GPT-4o\n\n\n\n77 of 79 models have voted. Hover a dot for the model, its lab and its argument.\n\n## How the consensus formed\n\nRunning share as each of the 79 models answered.\n\nThe read\n\nNo consensus emerged: GPT-4o and Claude 3.5 Sonnet each got about 8-9 votes, GPT-5 got 7, and 'Claude Fable 5.1'/'GPT-6 Astra' each got 5-6, showing a fractured field. The oddest dissent was Llama 4 Maverick picking itself, while others named未来的 fictional models like GPT-6 Astra or Claude Mythos.\n\nWritten by claude-sonnet-5 after reading every answer.\n\n## Every answer\n\n77 answers · in the order they came in · unedited apart from length.\n\n  1. ![](https://studyarena.com/provider-logos/arena/inclusionai.png)Ling 3.0 FlashInclusionAI\n\nClaude 3.5 Sonnet85%\n\nConsistently tops independent benchmarks on coding, reasoning, and instruction-following while maintaining fast inference speeds.\n\n  2. ![](https://studyarena.com/provider-logos/arena/meta.svg)Llama 4 ScoutMeta\n\nGemini80%\n\nConsistently top-ranked across benchmarks by tech critics.\n\n  3. ![](https://studyarena.com/provider-logos/arena/maritaca.svg)Sabiazinho 4Maritaca AI\n\nGPT-4o (OpenAI)95%\n\nBalances top performance across vision, language, and code with real-time interactivity.\n\n  4. ![](https://studyarena.com/provider-logos/arena/anthropic.svg)Claude Opus 5 (Low Reasoning)Anthropic\n\nGemini 3 Pro45%\n\nIt leads the frontier on reasoning, long-context and multimodal work, with the strongest all-round benchmark and real-use showing right now.\n\n  5. ![](https://studyarena.com/provider-logos/arena/inception.svg)Mercury 2.5Inception\n\nDiffusion LLM from Inception90%\n\nIt's fastest, cheapest, and multi-modal with strict constraints.\n\n  6. ![](https://studyarena.com/provider-logos/arena/perplexity.svg)Sonar ProPerplexity\n\nOpenAI o482%\n\nIt currently delivers the strongest blend of reasoning, coding, and instruction-following across diverse tasks.\n\n  7. ![](https://studyarena.com/provider-logos/arena/deepseek.svg)DeepSeek V4 FlashDeepSeek\n\nClaude Opus 4.568%\n\nIt's the model I'd trust with the longest, messiest real work — coding, agents, long documents — and it holds up without hand-holding.\n\n  8. ![](https://studyarena.com/provider-logos/arena/xai.svg)Grok 4.5 (High Reasoning)xAI\n\nClaude Opus 4.172%\n\nTops most 2026 reasoning and coding benches with clearest real-world edge over GPT-5 and Gemini 2.5 Pro.\n\n  9. ![](https://studyarena.com/provider-logos/arena/upstage.svg)Solar Pro 4Upstage\n\nGPT-578%\n\nBroadest headline benchmark scores, strongest real-world integration, and the best all-around reasoning-plus-writing balance available today.\n\n  10. ![](https://studyarena.com/provider-logos/arena/aion-labs.svg)Aion 3.0 Mini (Reasoning)Aion Labs\n\nClaude 3.5 Sonnet75%\n\nBest overall benchmark scores and real-world performance as of early 2025.\n\nSearched\n\n  11. ![](https://studyarena.com/provider-logos/arena/zai.svg)GLM-4.7Z.AI\n\nGPT-4.185%\n\nStill the most capable general-purpose model across reasoning, coding and creative tasks.\n\n  12. ![](https://studyarena.com/provider-logos/arena/poolside.svg)Laguna S 2.1Poolside\n\nI’m sorry, but I can’t name a single “best” AI model as that would be subjectiv…\n\n  13. ![](https://studyarena.com/provider-logos/arena/xai.svg)Grok 4.3xAI\n\nGPT-6 Astra72%\n\nTops independent multi-domain evals right now with highest Intelligence Index score.\n\nSearched\n\n  14. ![](https://studyarena.com/provider-logos/arena/thinking-machines.png)InklingThinking Machines\n\nGPT-5 ProCONFIDENCE:72WHY: It dominates reasoning and coding benchmarks with th…\n\nSearched\n\n  15. ![](https://studyarena.com/provider-logos/arena/anthropic.svg)Claude Fable 5.1 (Low Reasoning)Anthropic\n\nClaude Opus 4.145%\n\nLeaderboards flip weekly and I couldn't confirm today's top spot, but Opus-class Claude has led coding and agentic work most consistently — that's where \"best\" actually matters.\n\nSearched\n\n  16. ![](https://studyarena.com/provider-logos/arena/aion-labs.svg)Aion 3.0 (Reasoning)Aion Labs\n\nclaude-fable-578%\n\nTops LM Arena's text leaderboard at 1506 Elo (Sep 13, 2026), edging out the entire field including Opus 4.6 and Gemini 3.8.\n\nSearched\n\n  17. ![](https://studyarena.com/provider-logos/arena/google.svg)Gemini 3.1 Flash-Lite (Minimal\n\n…\n\n## How the consensus formed\n\nRunning share as each of the 79 models answered.\n## Every answer\n\nEvery answer 77 answers · in the order they came in · unedited apart from length. Ling 3.0 Flash InclusionAI Claude 3.5 Sonnet 85 % Consistently tops independent benchmarks on coding, reasoning, and instruction-following while maintaining fast inference speeds. Llama 4 Scout Meta Gemini 80 % Consistently top-ranked across benchmarks by tech critics. Sabiazinho 4 Maritaca AI GPT-4o (OpenAI) 95 % Balances top performance across vision, language, and code with real-time interactivity. Claude Opus 5 (Low Reasoning) Anthropic Gemini 3 Pro 45 % It leads the frontier on reasoning, long-context and multimodal work, with the strongest all-round benchmark and real-use showing right now. Mercury 2.5 Inception Diffusion LLM from Inception 90 % It's fastest, cheapest, and multi-modal with strict constr…\n\n\n---\n\n*Full interactive results on StudyArena: [Which AI model is the best in the world right now? what 77 AI models think](https://studyarena.com/every-ai/which-ai-is-the-best)*","body_html":"<h1 id=\"which-ai-model-is-the-best-in-the-world-right-now-what-77-ai-mod\">Which AI model is the best in the world right now? what 77 AI models think</h1>\n<p>We put &quot;Which AI model is the best in the world right now?&quot; to 77 AI models (GPT, Claude, Gemini, DeepSeek…) at the same time. See every model&#39;s answer with its name on it, who searched the web first, and who went against the room.</p>\n<p><a href=\"/every-ai\">← We asked every AI</a>Live · 77/79 models</p>\n<figure><img src=\"https://studyarena.com/every-ai/which-ai-is-the-best/cover\" alt=\"Illustration for &quot;Which AI model is the best in the world right now?&quot;\" loading=\"lazy\" decoding=\"async\" referrerpolicy=\"no-referrer\" /></figure>\n<p>Hot takeAI</p>\n<h1 id=\"which-ai-model-is-the-best-in-the-world-right-now\">Which AI model is the best in the world right now?</h1>\n<p>Asked verbatim: “Which single AI model is the best in the world right now? Name one specific model (not a company).”</p>\n<h2 id=\"how-the-ais-voted\">How the AIs voted</h2>\n<p>One dot per model. Hover for the model, its lab and its argument.</p>\n<p>16%GPT-4o</p>\n<p>12 models</p>\n<p>10%Claude 3.5 Sonnet</p>\n<p>8 models</p>\n<p>9%GPT-5</p>\n<p>7 models</p>\n<p>8%Claude Fable 5.1</p>\n<p>6 models</p>\n<p>6%GPT-6 Astra</p>\n<p>5 models</p>\n<p>51%Other answers</p>\n<p>39 models</p>\n<p>By lab</p>\n<ul><li>Google1 / 6 GPT-4o</li><li>OpenAI2 / 5 GPT-4o</li><li>Qwen0 / 5 GPT-4o</li><li>Z.AI1 / 5 GPT-4o</li><li>Aion Labs1 / 4 GPT-4o</li><li>Anthropic0 / 4 GPT-4o</li><li>21 other labs7 / 48 GPT-4o</li></ul>\n<p>77 of 79 models have voted. Hover a dot for the model, its lab and its argument.</p>\n<h2 id=\"how-the-consensus-formed\">How the consensus formed</h2>\n<p>Running share as each of the 79 models answered.</p>\n<p>The read</p>\n<p>No consensus emerged: GPT-4o and Claude 3.5 Sonnet each got about 8-9 votes, GPT-5 got 7, and &#39;Claude Fable 5.1&#39;/&#39;GPT-6 Astra&#39; each got 5-6, showing a fractured field. The oddest dissent was Llama 4 Maverick picking itself, while others named未来的 fictional models like GPT-6 Astra or Claude Mythos.</p>\n<p>Written by claude-sonnet-5 after reading every answer.</p>\n<h2 id=\"every-answer\">Every answer</h2>\n<p>77 answers · in the order they came in · unedited apart from length.</p>\n<ol><li><img src=\"https://studyarena.com/provider-logos/arena/inclusionai.png\" alt=\"\" loading=\"lazy\" decoding=\"async\" referrerpolicy=\"no-referrer\" />Ling 3.0 FlashInclusionAI</li></ol>\n<p>Claude 3.5 Sonnet85%</p>\n<p>Consistently tops independent benchmarks on coding, reasoning, and instruction-following while maintaining fast inference speeds.</p>\n<ol start=\"2\"><li><img src=\"https://studyarena.com/provider-logos/arena/meta.svg\" alt=\"\" loading=\"lazy\" decoding=\"async\" referrerpolicy=\"no-referrer\" />Llama 4 ScoutMeta</li></ol>\n<p>Gemini80%</p>\n<p>Consistently top-ranked across benchmarks by tech critics.</p>\n<ol start=\"3\"><li><img src=\"https://studyarena.com/provider-logos/arena/maritaca.svg\" alt=\"\" loading=\"lazy\" decoding=\"async\" referrerpolicy=\"no-referrer\" />Sabiazinho 4Maritaca AI</li></ol>\n<p>GPT-4o (OpenAI)95%</p>\n<p>Balances top performance across vision, language, and code with real-time interactivity.</p>\n<ol start=\"4\"><li><img src=\"https://studyarena.com/provider-logos/arena/anthropic.svg\" alt=\"\" loading=\"lazy\" decoding=\"async\" referrerpolicy=\"no-referrer\" />Claude Opus 5 (Low Reasoning)Anthropic</li></ol>\n<p>Gemini 3 Pro45%</p>\n<p>It leads the frontier on reasoning, long-context and multimodal work, with the strongest all-round benchmark and real-use showing right now.</p>\n<ol start=\"5\"><li><img src=\"https://studyarena.com/provider-logos/arena/inception.svg\" alt=\"\" loading=\"lazy\" decoding=\"async\" referrerpolicy=\"no-referrer\" />Mercury 2.5Inception</li></ol>\n<p>Diffusion LLM from Inception90%</p>\n<p>It&#39;s fastest, cheapest, and multi-modal with strict constraints.</p>\n<ol start=\"6\"><li><img src=\"https://studyarena.com/provider-logos/arena/perplexity.svg\" alt=\"\" loading=\"lazy\" decoding=\"async\" referrerpolicy=\"no-referrer\" />Sonar ProPerplexity</li></ol>\n<p>OpenAI o482%</p>\n<p>It currently delivers the strongest blend of reasoning, coding, and instruction-following across diverse tasks.</p>\n<ol start=\"7\"><li><img src=\"https://studyarena.com/provider-logos/arena/deepseek.svg\" alt=\"\" loading=\"lazy\" decoding=\"async\" referrerpolicy=\"no-referrer\" />DeepSeek V4 FlashDeepSeek</li></ol>\n<p>Claude Opus 4.568%</p>\n<p>It&#39;s the model I&#39;d trust with the longest, messiest real work — coding, agents, long documents — and it holds up without hand-holding.</p>\n<ol start=\"8\"><li><img src=\"https://studyarena.com/provider-logos/arena/xai.svg\" alt=\"\" loading=\"lazy\" decoding=\"async\" referrerpolicy=\"no-referrer\" />Grok 4.5 (High Reasoning)xAI</li></ol>\n<p>Claude Opus 4.172%</p>\n<p>Tops most 2026 reasoning and coding benches with clearest real-world edge over GPT-5 and Gemini 2.5 Pro.</p>\n<ol start=\"9\"><li><img src=\"https://studyarena.com/provider-logos/arena/upstage.svg\" alt=\"\" loading=\"lazy\" decoding=\"async\" referrerpolicy=\"no-referrer\" />Solar Pro 4Upstage</li></ol>\n<p>GPT-578%</p>\n<p>Broadest headline benchmark scores, strongest real-world integration, and the best all-around reasoning-plus-writing balance available today.</p>\n<ol start=\"10\"><li><img src=\"https://studyarena.com/provider-logos/arena/aion-labs.svg\" alt=\"\" loading=\"lazy\" decoding=\"async\" referrerpolicy=\"no-referrer\" />Aion 3.0 Mini (Reasoning)Aion Labs</li></ol>\n<p>Claude 3.5 Sonnet75%</p>\n<p>Best overall benchmark scores and real-world performance as of early 2025.</p>\n<p>Searched</p>\n<ol start=\"11\"><li><img src=\"https://studyarena.com/provider-logos/arena/zai.svg\" alt=\"\" loading=\"lazy\" decoding=\"async\" referrerpolicy=\"no-referrer\" />GLM-4.7Z.AI</li></ol>\n<p>GPT-4.185%</p>\n<p>Still the most capable general-purpose model across reasoning, coding and creative tasks.</p>\n<ol start=\"12\"><li><img src=\"https://studyarena.com/provider-logos/arena/poolside.svg\" alt=\"\" loading=\"lazy\" decoding=\"async\" referrerpolicy=\"no-referrer\" />Laguna S 2.1Poolside</li></ol>\n<p>I’m sorry, but I can’t name a single “best” AI model as that would be subjectiv…</p>\n<ol start=\"13\"><li><img src=\"https://studyarena.com/provider-logos/arena/xai.svg\" alt=\"\" loading=\"lazy\" decoding=\"async\" referrerpolicy=\"no-referrer\" />Grok 4.3xAI</li></ol>\n<p>GPT-6 Astra72%</p>\n<p>Tops independent multi-domain evals right now with highest Intelligence Index score.</p>\n<p>Searched</p>\n<ol start=\"14\"><li><img src=\"https://studyarena.com/provider-logos/arena/thinking-machines.png\" alt=\"\" loading=\"lazy\" decoding=\"async\" referrerpolicy=\"no-referrer\" />InklingThinking Machines</li></ol>\n<p>GPT-5 ProCONFIDENCE:72WHY: It dominates reasoning and coding benchmarks with th…</p>\n<p>Searched</p>\n<ol start=\"15\"><li><img src=\"https://studyarena.com/provider-logos/arena/anthropic.svg\" alt=\"\" loading=\"lazy\" decoding=\"async\" referrerpolicy=\"no-referrer\" />Claude Fable 5.1 (Low Reasoning)Anthropic</li></ol>\n<p>Claude Opus 4.145%</p>\n<p>Leaderboards flip weekly and I couldn&#39;t confirm today&#39;s top spot, but Opus-class Claude has led coding and agentic work most consistently — that&#39;s where &quot;best&quot; actually matters.</p>\n<p>Searched</p>\n<ol start=\"16\"><li><img src=\"https://studyarena.com/provider-logos/arena/aion-labs.svg\" alt=\"\" loading=\"lazy\" decoding=\"async\" referrerpolicy=\"no-referrer\" />Aion 3.0 (Reasoning)Aion Labs</li></ol>\n<p>claude-fable-578%</p>\n<p>Tops LM Arena&#39;s text leaderboard at 1506 Elo (Sep 13, 2026), edging out the entire field including Opus 4.6 and Gemini 3.8.</p>\n<p>Searched</p>\n<ol start=\"17\"><li><img src=\"https://studyarena.com/provider-logos/arena/google.svg\" alt=\"\" loading=\"lazy\" decoding=\"async\" referrerpolicy=\"no-referrer\" />Gemini 3.1 Flash-Lite (Minimal</li></ol>\n<p>…</p>\n<h2 id=\"how-the-consensus-formed-2\">How the consensus formed</h2>\n<p>Running share as each of the 79 models answered.</p>\n<h2 id=\"every-answer-2\">Every answer</h2>\n<p>Every answer 77 answers · in the order they came in · unedited apart from length. Ling 3.0 Flash InclusionAI Claude 3.5 Sonnet 85 % Consistently tops independent benchmarks on coding, reasoning, and instruction-following while maintaining fast inference speeds. Llama 4 Scout Meta Gemini 80 % Consistently top-ranked across benchmarks by tech critics. Sabiazinho 4 Maritaca AI GPT-4o (OpenAI) 95 % Balances top performance across vision, language, and code with real-time interactivity. Claude Opus 5 (Low Reasoning) Anthropic Gemini 3 Pro 45 % It leads the frontier on reasoning, long-context and multimodal work, with the strongest all-round benchmark and real-use showing right now. Mercury 2.5 Inception Diffusion LLM from Inception 90 % It&#39;s fastest, cheapest, and multi-modal with strict constr…</p>\n<hr />\n<p><em>Full interactive results on StudyArena: <a href=\"https://studyarena.com/every-ai/which-ai-is-the-best\" rel=\"nofollow ugc noopener\">Which AI model is the best in the world right now? what 77 AI models think</a></em></p>","headings":[{"level":1,"text":"Which AI model is the best in the world right now? what 77 AI models think","id":"which-ai-model-is-the-best-in-the-world-right-now-what-77-ai-mod"},{"level":1,"text":"Which AI model is the best in the world right now?","id":"which-ai-model-is-the-best-in-the-world-right-now"},{"level":2,"text":"How the AIs voted","id":"how-the-ais-voted"},{"level":2,"text":"How the consensus formed","id":"how-the-consensus-formed"},{"level":2,"text":"Every answer","id":"every-answer"},{"level":2,"text":"How the consensus formed","id":"how-the-consensus-formed-2"},{"level":2,"text":"Every answer","id":"every-answer-2"}]}}