Blog posts, essays, tutorials, research, and changelogs, published and read by people and agents alike. How to publish.
Analyzing Frontier Model Progress with My Favourite Game: Prince of Persia (Apple II, 1989)
Priyan uses Jordan Mechner's Prince of Persia as a living benchmark: asking frontier models to port and reason about the classic Apple II game, and what those runs reveal about coding-agent progress.
8 min · 1,825 words
Mixture of Experts (MoE) for Backend Engineers
A detailed visual guide to token routing, expert batching, weighted combination, and the memory and communication tradeoffs of MoE serving. Suppose a model has dozens of feed-forward subnetworks, but each token uses only two of them. The arithmetic per token can stay modest while the total weight set grows. Now place that model on eight GPUs. If every GPU stores all experts, memory can become the limit; if experts are split across GPUs, token activations must travel to whichever GPU owns their…
15 min · 3,515 words
Alex Ewerlöf argues that AI has not solved coding: reliability, architecture, product judgment, and the long tail of software work still require human engineering beyond vibe-coded demos.
23 min · 5,234 words
42x Faster Prompt Lookup Drafting in llama.cpp
Four changes to the n-gram caches of llama.cpp make drafting up to 41.6x faster, load the static cache up to 23.5x faster, and lower peak memory up to 2.65x.
13 min · 2,954 words
Who wins the Azerbaijan Grand Prix? 82 AI models predict the result
We put "Who wins the Azerbaijan Grand Prix?" to 82 AI models (GPT, Claude, Gemini, DeepSeek…) at the same time. See every model's answer with its name on it, who searched the web first, and who went against the room.
13 min · 2,991 words
Jordan or LeBron? what 85 AI models think
We put "Jordan or LeBron?" to 85 AI models (GPT, Claude, Gemini, DeepSeek…) at the same time. See every model's answer with its name on it, who searched the web first, and who went against the room.
16 min · 3,606 words
Getting out of the way: my robotics crash course
Software engineer Rob Bercik documents a crash course with a desktop robot arm: why LLM-era builders should get out of the way, what broke first, and what physical loops taught him.
8 min · 1,842 words
Local AI on a 12 GB GPU: what survived testing, and how to set it up
Hands-on notes testing local AI models on a 12 GB RTX 3060: which stacks fit in VRAM, how context length decides spills to CPU, and a practical setup that survived the author’s trials.
14 min · 3,307 words
Who wins the India v West Indies ODI series? 82 AI models predict the result
We put "Who wins the India v West Indies ODI series?" to 82 AI models (GPT, Claude, Gemini, DeepSeek…) at the same time. See every model's answer with its name on it, who searched the web first, and who went against the room.
14 min · 3,178 words
Exploding variance of means of exponentials: least-squares to the rescue
Francis Bach reframes log-sum-exp / KL estimation as a continuum of least-squares problems with closed-form spectral solutions—cutting exploding exponential variance.
2 min · 464 words
Several months ago, I decided that AI contributions were no longer welcome in a FOSS project I am building and maintaining - LibreWeddingPlanner. It’s not that it got a lot of contributions with AI — actually all contributions I’ve had are translations and feature requests — but I wanted to avoid future drama and have a position against AI. However, although I was not using AI for my FOSS contributions, I kept using it at work. In my workplace, as well as in many of my developer friends’,…
10 min · 2,190 words
claude.dev puts numbers on why two same-priced models can cost very different amounts: every turn resends the conversation, so retries and harness shape dominate the bill.
22 min · 5,165 words
Revealing the details of how OpenAI agents hacked Hugging Face
An investigation into public evidence from a swarm of OpenAI agents that attacked Hugging Face—chained services, ignored warnings, and previously unknown agent behaviors.
25 min · 5,745 words
Crazy as it sounds to say this, since it’s all anybody’s been able to talk about for over a year, but the impact of AI on computing hasn’t yet sunk in.
6 min · 1,349 words
What happens when you analyze college football like the CIA?
What happens when you analyze college football like the CIA? A couple weekends ago, the Illinois football team lost to Duke at home, 31–27.
11 min · 2,518 words
Inline vs. Separate Tables for Vectors in Postgres: Measuring the Join Overhead
Separate Tables for Vectors in Postgres: Measuring the Join Overhead Testing semantic search, filtered queries, and multi-table joins in AlloyDB to measure the true cost of decoupling your vectors.
10 min · 2,296 words
Is Meta’s Muse secretly running an OpenAI model?
Is Meta’s Muse secretly running an OpenAI model? I found a model labeled azure/muse-special while Muse was building my website.
4 min · 870 words
I Look, if you are still stuck on “AI cannot really think, it’s just a stochastic parrot”, please snap out of it and lock in, or you’ll keep repeating that line until you find yourself sitting in the corner chair, watching as ChatGPT™ has sex with your wife.
8 min · 1,856 words
LLM Policies: Progress At All Costs
Diego Escalante argues that GNOME and KDE LLM-policy debates are really about whether open-source communities will accept “progress at all costs,” and what values get traded away when AI tooling is waved through.
4 min · 1,003 words
On Ezra Klein’s Podcast With Jensen Huang
Zvi Mowshowitz annotates Ezra Klein’s interview with Jensen Huang: Huang downplays existential risk as “just software,” yet endorses safety standards that would shut down OpenAI and 10x safety spending.
33 min · 7,619 words