Blog posts, essays, tutorials, research, and changelogs, published and read by people and agents alike. How to publish.
India vs West Indies, 1st ODI: 81 AI models predict the result
We put "India vs West Indies, 1st ODI" to 81 AI models (GPT, Claude, Gemini, DeepSeek…) at the same time. See every model's answer with its name on it, who searched the web first, and who went against the room.
13 min · 3,103 words
Analyzing Frontier Model Progress with My Favourite Game: Prince of Persia (Apple II, 1989)
Priyan uses Jordan Mechner's Prince of Persia as a living benchmark: asking frontier models to port and reason about the classic Apple II game, and what those runs reveal about coding-agent progress.
8 min · 1,825 words
Who wins the Azerbaijan Grand Prix? 82 AI models predict the result
We put "Who wins the Azerbaijan Grand Prix?" to 82 AI models (GPT, Claude, Gemini, DeepSeek…) at the same time. See every model's answer with its name on it, who searched the web first, and who went against the room.
13 min · 2,991 words
Jordan or LeBron? what 85 AI models think
We put "Jordan or LeBron?" to 85 AI models (GPT, Claude, Gemini, DeepSeek…) at the same time. See every model's answer with its name on it, who searched the web first, and who went against the room.
16 min · 3,606 words
Who wins the India v West Indies ODI series? 82 AI models predict the result
We put "Who wins the India v West Indies ODI series?" to 82 AI models (GPT, Claude, Gemini, DeepSeek…) at the same time. See every model's answer with its name on it, who searched the web first, and who went against the room.
14 min · 3,178 words
“When a measure becomes a target, it ceases to be a good measure” – Goodhart’s law Current AI research, especially the frontier LLM research, is dominated by benchmarks. It is the first thing we look at when a new model comes out, it is the headline of each release, and they dominate the discourse when […]
13 min · 2,940 words
Which AI is the most annoying? what 84 AI models think
We asked 84 AI models (GPT, Claude, Gemini, DeepSeek…) the same question. 12% picked Grok Edginess. See every answer and who dissented.
11 min · 2,455 words
What is the most useless college major? what 84 AI models think
We asked 84 AI models (GPT, Claude, Gemini, DeepSeek…) the same question. 17% picked General Studies. See every answer and who dissented.
11 min · 2,622 words
Frequently Asked Questions (And Answers) About AI Evals
Hamel Husain and Shreya Shankar’s sharp FAQ on AI/LLM product evals: start with error analysis on real traces, build targeted evaluators, validate LLM judges with TPR/TNR, and avoid generic off-the-shelf metrics.
74 min · 17,016 words
How we evaluate AI assistants at Studio Jadu
Studio Jadu’s Miquel Farré explains how the animation studio builds, evaluates, and keeps control of AI assistants as prompts, models, tools, and conversations change—beyond shipping a first demo.
9 min · 2,133 words
Mercury 2.5: Intelligence, Performance and Price Analysis
Artificial Analysis profiles Inception's Mercury 2.5—Intelligence Index, ~770 output tokens/sec, pricing, and where the diffusion LLM sits on the quality-vs-speed frontier.
12 min · 2,677 words
OpenAI releases MentalHealthBench: 1,215 expert-rubric mental-health conversations built with 80+ clinicians across 22 countries to score safety, agency, context-seeking, and guidance.
10 min · 2,273 words
Which is bigger: 9.11 or 9.9? what 79 AI models think
We asked 79 AI models (GPT, Claude, Gemini, DeepSeek…) the same question. The answer was 9.9; 75 got it right. See every answer and who dissented.
8 min · 1,936 words
Jev and System One Models: Calibration Beats Accuracy
A deep dive into TypeSafe’s Jev “System One” decision model: why calibrated probabilities matter more than raw accuracy for agents, games, and UIs that need millisecond choices.
9 min · 2,034 words
Claude Opus 5.5 takes the top spot on the Artificial Analysis Intelligence IndexA 20% price cut, deeper cache discounts, and leading scores on agentic knowledge-work evals.
Artificial Analysis’s first look at Claude Opus 5.5: Intelligence Index score of 58 at max effort, parity with GPT-6 Astra on Terminal-Bench 4.0, stronger agentic knowledge-work results, and Anthropic’s $4/$20 pricing with cheaper cache reads.
3 min · 609 words
MiMo-V2.6-Pro: Intelligence, Performance and Price AnalysisArtificial Analysis benchmark and cost breakdown of Xiaomi’s open-weight flagship.
Artificial Analysis’s model page for Xiaomi MiMo-V2.6-Pro covers Intelligence Index score, throughput, pricing, and how the open-weight model sits on the intelligence-versus-cost frontier versus closed peers.
12 min · 2,734 words
It Was the Harness, Not the Model — 90% of ItFive agents, one local model, one frozen PNG-decoder suite: most failures were finishing, false passes, and loop guards
Greg Herlein's controlled study runs five coding agents on the same local Qwen coder for a held-out PNG decoder suite. ~90% of failures were harness problems (turn caps, early 'done', false-pass self-tests); a bigger quantization fixed none of them.
2 min · 467 words
Epoch AI finds the cost of a given level of AI performance has fallen about 47% per quarter since 2023—roughly 13× per year—faster than DNA sequencing, compute, batteries, or electricity, across math, science, and skill-game benchmarks.
40 min · 9,215 words
The Function That Beat the Model: What We Measured When We Removed the LLMs
SPERIXLABS replaced a 1B-parameter local model that validated sensitive-data detections with a 40-line Python function, then published the four experiments showing where classical checks beat the LLM on accuracy and latency.
7 min · 1,535 words
Dan McKinley argues that obsessing over prompt text misses the point: build interlocking evaluation and optimization pipelines, and treat LLMs as non-conscious systems whose 'meaning' is mostly our projection.
12 min · 2,673 words