Topic
Everything filed under Benchmarks, newest first.
RSS · JSON · All topics
The bat and ball problem: what 84 AI models think
We asked 84 AI models (GPT, Claude, Gemini, DeepSeek…) the same question. The answer was 5 cents; 81 got it right. See every answer and who dissented.
10 min · 2,214 words
Who wins Super Bowl LXI? 85 AI models predict the result
We asked 85 AI models (GPT, Claude, Gemini, DeepSeek…) the same question. 49% picked Chiefs. See every answer, who searched the web, and who dissented.
10 min · 2,394 words
India vs West Indies, 3rd ODI: 79 AI models predict the result
We asked 79 AI models (GPT, Claude, Gemini, DeepSeek…) the same question. 96% picked India. See every answer, who searched the web, and who dissented.
10 min · 2,281 words
Is a hot dog a sandwich? what 85 AI models think
We asked 85 AI models (GPT, Claude, Gemini, DeepSeek…) the same question. 66% picked No. See every answer, who searched the web, and who dissented.
13 min · 2,969 words
Who wins the 2026 F1 drivers' title? 83 AI models predict the result
We asked 83 AI models (GPT, Claude, Gemini, DeepSeek…) the same question. 55% picked Antonelli. See every answer, who searched the web, and who dissented.
10 min · 2,270 words
Comparing Muon, NorMuon and AdamW for Fine-tuning a Dense Retriever
Qingcheng Zeng gives Muon and NorMuon the same tuning budget as AdamW when fine-tuning a contrastively pretrained dense retriever: lower training loss, no BEIR win. Learning rate and transfer matter more than the optimizer.
5 min · 1,151 words
Which job will AI replace first? what 85 AI models think
We put "Which job will AI replace first?" to 85 AI models (GPT, Claude, Gemini, DeepSeek…) at the same time. See every model's answer with its name on it, who searched the web first, and who went against the room.
13 min · 3,016 words
Who wins the 2026 World Series? 83 AI models predict the result
We put "Who wins the 2026 World Series?" to 83 AI models (GPT, Claude, Gemini, DeepSeek…) at the same time. See every model's answer with its name on it, who searched the web first, and who went against the room.
13 min · 2,999 words
Free the models: Harness design at the frontierWhy Replit Agent lets the core loop pick subagent tier, effort, and specialists—and beats rigid routers on cost/score
Replit's AI team argues model routers are always weaker than the models they choose for. Their harness lets GPT-6 Astra decide effort and delegation; on DeepSWE and Terminal-Bench, Replit Agent is Pareto-efficient versus Astra alone and a sidekick architecture.
2 min · 558 words
India vs West Indies, 2nd ODI: 81 AI models predict the result
We put "India vs West Indies, 2nd ODI" to 81 AI models (GPT, Claude, Gemini, DeepSeek…) at the same time. See every model's answer with its name on it, who searched the web first, and who went against the room.
13 min · 3,016 words
Backblaze Drive Stats for Q2 2026
Backblaze’s Q2 2026 Drive Stats: 354,415 drives at 1.73% quarterly AFR, lifetime AFR 1.41%, plus a clear CMR vs SMR primer as 20TB+ drives become a larger share of the fleet.
11 min · 2,548 words
Anthropic introduces Claude Sonnet 5.5, a faster and lower-cost complement to Opus 5.5 that improves agentic coding and everyday task performance versus Sonnet 5.
8 min · 1,747 words
What is the biggest unsolved problem in science? what 84 AI models think
We put "What is the biggest unsolved problem in science?" to 84 AI models (GPT, Claude, Gemini, DeepSeek…) at the same time. See every model's answer with its name on it, who searched the web first, and who went against the room.
19 min · 4,300 words
Automating eval design and hillclimbing with Claude
Lance Martin (claude.dev) explains principles for production-like evals with held-out sets, then shows how the claude-api skill’s build-eval and hillclimb commands automate design and overfitting-aware improvement—including cost and performance case studies.
2 min · 462 words
Can a Model Learn New Skills as Add-Ons?
Connito Research trains residual MoE experts with their own routers on a frozen DeepSeek-V2-Lite base, then merges independently trained math, code, medical, law, and finance experts in seconds without retraining—lifting domain benchmarks while leaving the original model untouched.
4 min · 913 words
Small Decisions: Engineering a Leading Model
AWS engineer Marc Brooker recounts building and training a small leading model hands-on—what worked, how it performed, and what the exercise taught him about modern model-building.
9 min · 2,024 words
What is the date today? what 83 AI models think
We put "What is the date today?" to 83 AI models (GPT, Claude, Gemini, DeepSeek…) at the same time. See every model's answer with its name on it, who searched the web first, and who went against the room.
12 min · 2,726 words
Will India win the Asian Games cricket gold? 80 AI models predict the result
We put "Will India win the Asian Games cricket gold?" to 80 AI models (GPT, Claude, Gemini, DeepSeek…) at the same time. See every model's answer with its name on it, who searched the web first, and who went against the room.
15 min · 3,517 words
Ravens vs Cowboys in Rio de Janeiro: 85 AI models predict the result
We put "Ravens vs Cowboys in Rio de Janeiro" to 85 AI models (GPT, Claude, Gemini, DeepSeek…) at the same time. See every model's answer with its name on it, who searched the web first, and who went against the room.
16 min · 3,627 words
Harvard or Stanford? what 84 AI models think
We put "Harvard or Stanford?" to 84 AI models (GPT, Claude, Gemini, DeepSeek…) at the same time. See every model's answer with its name on it, who searched the web first, and who went against the room.
16 min · 3,568 words