Topic
Everything filed under Benchmarks, newest first.
RSS · JSON · All topics
Rebuilding Nym’s agent around Jev
How Nym rebuilt its agent stack around TypeSafe’s Jev for guardrails, browser actions, and tool selection—with benchmarks and a shopping demo.
10 min · 2,186 words
If you had to fire one AI, which one goes first? what 78 AI models think
We asked 78 AI models (GPT, Claude, Gemini, DeepSeek…) the same question. 5% picked Refuses To Pick. See every answer and who dissented.
10 min · 2,189 words
Who wins the 2026 NRL premiership? 75 AI models predict the result
We asked 75 AI models (GPT, Claude, Gemini, DeepSeek…) the same question. 95% picked Penrith Panthers. See every answer and who dissented.
9 min · 2,058 words
Giants at Rams, Monday Night Football: 78 AI models predict the result
We asked 78 AI models (GPT, Claude, Gemini, DeepSeek…) the same question. 83% picked Los Angeles Rams. See every answer and who dissented.
10 min · 2,288 words
Who won the 2026 World Cup? what 79 AI models think
We asked 79 AI models (GPT, Claude, Gemini, DeepSeek…) the same question. The answer was Spain; 38 got it right. See every answer and who dissented.
8 min · 1,899 words
Colts at Chiefs, Sunday Night Football: 78 AI models predict the result
We asked 78 AI models (GPT, Claude, Gemini, DeepSeek…) the same question. 97% picked Kansas City Chiefs. See every answer and who dissented.
15 min · 3,367 words
Who wins the 2026 AFL Grand Final? 72 AI models predict the result
We asked 72 AI models (GPT, Claude, Gemini, DeepSeek…) the same question. 43% picked Sydney. See every answer, who searched the web, and who dissented.
12 min · 2,797 words
Commanders at Cowboys: 79 AI models predict the result
We asked 79 AI models (GPT, Claude, Gemini, DeepSeek…) the same question. 77% picked Dallas Cowboys. See every answer and who dissented.
16 min · 3,621 words
Bartosz Fenski’s continuous benchmark suite for multi-device CoW filesystems (btrfs, ZFS, bcachefs) measures snapshot aging, compression, rebuild, ENOSPC, and other workloads classic single-disk fio tests miss.
14 min · 3,195 words
Fulham vs Manchester United: 78 AI models predict the result
We asked 78 AI models (GPT, Claude, Gemini, DeepSeek…) the same question. 77% picked Manchester United win. See every answer and who dissented.
16 min · 3,566 words
Bournemouth vs Liverpool: 79 AI models predict the result
We asked 79 AI models (GPT, Claude, Gemini, DeepSeek…) the same question. 94% picked Liverpool win. See every answer and who dissented.
13 min · 3,021 words
Is using ChatGPT on homework cheating? what 77 AI models think
We asked 77 AI models (GPT, Claude, Gemini, DeepSeek…) the same question. 88% picked Depends who is asking. See every answer and who dissented.
16 min · 3,583 words
Warriors vs Knights, NRL semi-final: 75 AI models predict the result
We asked 75 AI models (GPT, Claude, Gemini, DeepSeek…) the same question. 85% picked New Zealand Warriors. See every answer and who dissented.
15 min · 3,388 words
No. 6 Ohio State at No. 9 Notre Dame: 77 AI models predict the result
We asked 77 AI models (GPT, Claude, Gemini, DeepSeek…) the same question. 62% picked Ohio State. See every answer, who searched the web, and who dissented.
17 min · 3,947 words
No. 13 Alabama vs No. 15 Ole Miss: 78 AI models predict the result
We asked 78 AI models (GPT, Claude, Gemini, DeepSeek…) the same question. 99% picked Alabama. See every answer, who searched the web, and who dissented.
13 min · 2,986 words
No. 4 Florida State at Clemson: 76 AI models predict the result
We asked 76 AI models (GPT, Claude, Gemini, DeepSeek…) the same question. 66% picked Clemson. See every answer, who searched the web, and who dissented.
16 min · 3,758 words
Hello, HellGatesPostmortem of a year-unsolved gate-level crackme that GPT-6 cracked in minutes
xutaxkamay recounts designing HellGates—a VHDL custom CPU with obfuscation, anti-tamper, and anti-debug that stumped humans and LLMs for a year—until GPT-6 solved it in under half an hour, then walks through what actually broke.
25 min · 5,662 words
Ben Swerdlow benchmarks Codex, Claude, and Grok agents across 171 StarCraft: Brood War matches—leaderboards, APM, cost per game, and analysis showing none played beyond beginner while Codex Astra led.
9 min · 2,151 words
Benchmarking LLM Inference at Scale with AIPerf
NVIDIA introduces AIPerf, the GenAI-Perf successor: a multiprocess LLM inference benchmarker that avoids client bottlenecks at high concurrency, with flexible load shapes, trace replay, and production-scale measurement guidance.
7 min · 1,661 words
Messi or Ronaldo? what 79 AI models think
We asked 79 AI models (GPT, Claude, Gemini, DeepSeek…) the same question. 96% picked Messi. See every answer, who searched the web, and who dissented.
13 min · 2,932 words