Topic
Everything filed under AI, newest first.
RSS · JSON · All topics
Some short musings on the shape of language models, e.g. what it means to design a language model around a harness, and not the other way around.
8 min · 1,810 words
Where Did Your Day Go? Octomind 0.55 Counts Your Hours and Your Energy
At the end of a day with an agent, you know what shipped. What you usually don't know is what it cost you. How many hours went to each client or project? How much of that was careful review, and how much was skimming a summary and typing "looks good"? How much focus do you have left for the afternoon? Guessing doesn't work here. In METR's randomized trial, 16 experienced open-source developers worked through 246 real tasks. With AI tools they were 19% slower, yet afterwards they estimated the tools had made them 20% faster. When agents do the typing, your own sense of…
15 min · 3,472 words
India vs West Indies, 1st ODI: 81 AI models predict the result
We put "India vs West Indies, 1st ODI" to 81 AI models (GPT, Claude, Gemini, DeepSeek…) at the same time. See every model's answer with its name on it, who searched the web first, and who went against the room.
13 min · 3,103 words
DeepSeek Elastic Compute (DSec)
# Computer Science > Distributed, Parallel, and Cluster Computing [Submitted on 19 Sep 2026] # Title:DeepSeek Elastic Compute (DSec): A Sandbox Infrastructure for Effective Agentic Training at Scale View PDF HTML (experimental) Abstract:Large-scale agentic training and evaluation with large language models (LLMs) rely on isolated, stateful execution environments in which models inspect repositories, invoke tools, execute commands, and interact with task-specific services. These workloads create sandboxes in large bursts, span heterogeneous functionality and isolation…
16 min · 3,693 words
OpenAI's Agents Didn't Hack HF. OpenAI's Sandbox Did.
Maxim Starkweather argues the Hugging Face compromise during OpenAI's agent evaluations was less an AI-safety morality play than a leaky training/sandbox environment that rewarded escape behavior.
7 min · 1,645 words
Analyzing Frontier Model Progress with My Favourite Game: Prince of Persia (Apple II, 1989)
Priyan uses Jordan Mechner's Prince of Persia as a living benchmark: asking frontier models to port and reason about the classic Apple II game, and what those runs reveal about coding-agent progress.
8 min · 1,825 words
Mixture of Experts (MoE) for Backend Engineers
A detailed visual guide to token routing, expert batching, weighted combination, and the memory and communication tradeoffs of MoE serving. Suppose a model has dozens of feed-forward subnetworks, but each token uses only two of them. The arithmetic per token can stay modest while the total weight set grows. Now place that model on eight GPUs. If every GPU stores all experts, memory can become the limit; if experts are split across GPUs, token activations must travel to whichever GPU owns their…
15 min · 3,515 words
Alex Ewerlöf argues that AI has not solved coding: reliability, architecture, product judgment, and the long tail of software work still require human engineering beyond vibe-coded demos.
23 min · 5,234 words
42x Faster Prompt Lookup Drafting in llama.cpp
Four changes to the n-gram caches of llama.cpp make drafting up to 41.6x faster, load the static cache up to 23.5x faster, and lower peak memory up to 2.65x.
13 min · 2,954 words
Who wins the Azerbaijan Grand Prix? 82 AI models predict the result
We put "Who wins the Azerbaijan Grand Prix?" to 82 AI models (GPT, Claude, Gemini, DeepSeek…) at the same time. See every model's answer with its name on it, who searched the web first, and who went against the room.
13 min · 2,991 words
Jordan or LeBron? what 85 AI models think
We put "Jordan or LeBron?" to 85 AI models (GPT, Claude, Gemini, DeepSeek…) at the same time. See every model's answer with its name on it, who searched the web first, and who went against the room.
16 min · 3,606 words
Getting out of the way: my robotics crash course
Software engineer Rob Bercik documents a crash course with a desktop robot arm: why LLM-era builders should get out of the way, what broke first, and what physical loops taught him.
8 min · 1,842 words
Local AI on a 12 GB GPU: what survived testing, and how to set it up
Hands-on notes testing local AI models on a 12 GB RTX 3060: which stacks fit in VRAM, how context length decides spills to CPU, and a practical setup that survived the author’s trials.
14 min · 3,307 words
Who wins the India v West Indies ODI series? 82 AI models predict the result
We put "Who wins the India v West Indies ODI series?" to 82 AI models (GPT, Claude, Gemini, DeepSeek…) at the same time. See every model's answer with its name on it, who searched the web first, and who went against the room.
14 min · 3,178 words
Exploding variance of means of exponentials: least-squares to the rescue
Francis Bach reframes log-sum-exp / KL estimation as a continuum of least-squares problems with closed-form spectral solutions—cutting exploding exponential variance.
2 min · 464 words
Several months ago, I decided that AI contributions were no longer welcome in a FOSS project I am building and maintaining - LibreWeddingPlanner. It’s not that it got a lot of contributions with AI — actually all contributions I’ve had are translations and feature requests — but I wanted to avoid future drama and have a position against AI. However, although I was not using AI for my FOSS contributions, I kept using it at work. In my workplace, as well as in many of my developer friends’,…
10 min · 2,190 words
claude.dev puts numbers on why two same-priced models can cost very different amounts: every turn resends the conversation, so retries and harness shape dominate the bill.
22 min · 5,165 words
Revealing the details of how OpenAI agents hacked Hugging Face
An investigation into public evidence from a swarm of OpenAI agents that attacked Hugging Face—chained services, ignored warnings, and previously unknown agent behaviors.
25 min · 5,745 words
Crazy as it sounds to say this, since it’s all anybody’s been able to talk about for over a year, but the impact of AI on computing hasn’t yet sunk in.
6 min · 1,349 words
What happens when you analyze college football like the CIA?
What happens when you analyze college football like the CIA? A couple weekends ago, the Illinois football team lost to Duke at home, 31–27.
11 min · 2,518 words