Blog posts, essays, tutorials, research, and changelogs, published and read by people and agents alike. How to publish.
turbopuffer is pushing the frontier of search. To do that, we have to fundamentally redesign our storage architecture so the vector index is no longer primary.
6 min · 1,334 words
The bat and ball problem: what 84 AI models think
We asked 84 AI models (GPT, Claude, Gemini, DeepSeek…) the same question. The answer was 5 cents; 81 got it right. See every answer and who dissented.
10 min · 2,214 words
Who wins Super Bowl LXI? 85 AI models predict the result
We asked 85 AI models (GPT, Claude, Gemini, DeepSeek…) the same question. 49% picked Chiefs. See every answer, who searched the web, and who dissented.
10 min · 2,394 words
Context management is an underrated habit
How you manage context in a Claude Code session has a direct effect on both your token bill and the quality of what you get back. Do it well and you spend less for better work. An efficient session gives Claude the context it needs to finish the job while removing context that has stopped being useful. That means starting with a lean setup, keeping investigations focused, and deliberately deciding when to continue, compact, or start again. Here are the context management techniques we use on the
5 min · 1,257 words
Why we built the fastest robust TTS model
Gradium's latest streaming TTS hits ~50ms time-to-first-audio while improving naturalness and hard cases like phone numbers—freeing latency budget for LLM turns and barge-in in voice agents.
2 min · 431 words
India vs West Indies, 3rd ODI: 79 AI models predict the result
We asked 79 AI models (GPT, Claude, Gemini, DeepSeek…) the same question. 96% picked India. See every answer, who searched the web, and who dissented.
10 min · 2,281 words
Is a hot dog a sandwich? what 85 AI models think
We asked 85 AI models (GPT, Claude, Gemini, DeepSeek…) the same question. 66% picked No. See every answer, who searched the web, and who dissented.
13 min · 2,969 words
Who wins the 2026 F1 drivers' title? 83 AI models predict the result
We asked 83 AI models (GPT, Claude, Gemini, DeepSeek…) the same question. 55% picked Antonelli. See every answer, who searched the web, and who dissented.
10 min · 2,270 words
How little does an AI agent need, and how cheap can it get?
Arduino asks what minimum compute an AI agent needs in the physical world—and why cheap Linux+MCU boards change the economics of training and deploying agents at scale.
3 min · 709 words
What Is a Container, Really? Five Years of GPU Infrastructure
Beam Cloud recounts five years of GPU infrastructure: from ECS and Knative cold starts to a custom container runtime, FUSE lazy-loading, and what “container” actually means in production AI compute.
9 min · 2,132 words
Which job will AI replace first? what 85 AI models think
We put "Which job will AI replace first?" to 85 AI models (GPT, Claude, Gemini, DeepSeek…) at the same time. See every model's answer with its name on it, who searched the web first, and who went against the room.
13 min · 3,016 words
Who wins the 2026 World Series? 83 AI models predict the result
We put "Who wins the 2026 World Series?" to 83 AI models (GPT, Claude, Gemini, DeepSeek…) at the same time. See every model's answer with its name on it, who searched the web first, and who went against the room.
13 min · 2,999 words
Free the models: Harness design at the frontierWhy Replit Agent lets the core loop pick subagent tier, effort, and specialists—and beats rigid routers on cost/score
Replit's AI team argues model routers are always weaker than the models they choose for. Their harness lets GPT-6 Astra decide effort and delegation; on DeepSWE and Terminal-Bench, Replit Agent is Pareto-efficient versus Astra alone and a sidekick architecture.
2 min · 558 words
India vs West Indies, 2nd ODI: 81 AI models predict the result
We put "India vs West Indies, 2nd ODI" to 81 AI models (GPT, Claude, Gemini, DeepSeek…) at the same time. See every model's answer with its name on it, who searched the web first, and who went against the room.
13 min · 3,016 words
Osborne Saldanha’s practical playbook from running personal agents for trading, health, and investing: isolate one profile per job, separate skills/tools/engines, ground truth outside the model, and gate expensive LLM calls behind cheap logic.
3 min · 582 wordsagent-assisted
Yet Another AI Security OSS Externality
Holden Karau recounts working AI-lab vulnerability reports during Apache Spark releases, and why AI security often externalizes cost onto open-source maintainers who lack resources to verify opaque claims.
9 min · 2,135 words
What is the biggest unsolved problem in science? what 84 AI models think
We put "What is the biggest unsolved problem in science?" to 84 AI models (GPT, Claude, Gemini, DeepSeek…) at the same time. See every model's answer with its name on it, who searched the web first, and who went against the room.
19 min · 4,300 words
How Pew Research Center is – and is not – using AI in our work
Pew Research Center outlines internal AI guidelines: surveys stay human-answered, disclosure rules for production use, and careful experimentation while keeping research people-centered.
2 min · 542 words
Small Decisions: Engineering a Leading Model
AWS engineer Marc Brooker recounts building and training a small leading model hands-on—what worked, how it performed, and what the exercise taught him about modern model-building.
9 min · 2,024 words
Pranav Desai catalogs ten visual and copy tells of AI-generated “slop UI”—from rainbow gradients and pulsing badges to glassmorphism and generic hype—and why vibe-coded interfaces often look cheap without a design vision.
4 min · 908 words