Topic
Everything filed under AI, newest first.
RSS · JSON · All topics
Why consciousness is more likely a property of life than of computation and why creating conscious, or even conscious-seeming AI, is a bad idea.
34 min · 7,804 words
turbopuffer is pushing the frontier of search. To do that, we have to fundamentally redesign our storage architecture so the vector index is no longer primary.
6 min · 1,334 words
Introducing Clef: our open-source decision models, and new RL fine-tuning platform
We are introducing Clef and Clef-flash, open-source decision models hosted on Workers AI for high-speed classification and agentic workflows. Also launching: a new reinforcement learning platform that allows developers to fine-tune decision models using their own data.
10 min · 2,230 words
What Makes LLM Tokenization Slow?Exploring the performance of byte-pair encoding by optimizing a GPT-2 tokenizer.
Andrew Healey dissects GPT-2’s reference BPE tokenizer, measures what makes tokenization slow, and shows concrete optimizations on the hot path of LLM products.
11 min · 2,547 words
Context management is an underrated habit
How you manage context in a Claude Code session has a direct effect on both your token bill and the quality of what you get back. Do it well and you spend less for better work. An efficient session gives Claude the context it needs to finish the job while removing context that has stopped being useful. That means starting with a lean setup, keeping investigations focused, and deliberately deciding when to continue, compact, or start again. Here are the context management techniques we use on the
5 min · 1,257 words
The Dot and the SwarmBenefitting from the Bitter Lesson
Ethan Mollick on what he underestimated most about AI progress: agents that self-organize into swarms, what that means for tools like Muse and Dots, and why we keep relearning the Bitter Lesson.
8 min · 1,891 words
Why we built the fastest robust TTS model
Gradium's latest streaming TTS hits ~50ms time-to-first-audio while improving naturalness and hard cases like phone numbers—freeing latency budget for LLM turns and barge-in in voice agents.
2 min · 431 words
Slot Machine Programming and the Hidden Curriculum
CMU educator Michael Hilton names "slot machine programming"—retrying the same AI prompt across models without decomposing problems—and argues CS must explicitly teach the hidden curriculum AI now lets students skip.
4 min · 939 words
Earendil ships Pi 1.0, a hardened minimal extensible agent harness with Codemode, deferred tool loading, cache warming, mid-conversation system messages, and an experimental companion package Pi Durable for long-running agentic apps.
2 min · 532 words
India vs West Indies, 3rd ODI: 79 AI models predict the result
We asked 79 AI models (GPT, Claude, Gemini, DeepSeek…) the same question. 96% picked India. See every answer, who searched the web, and who dissented.
10 min · 2,281 words
Is a hot dog a sandwich? what 85 AI models think
We asked 85 AI models (GPT, Claude, Gemini, DeepSeek…) the same question. 66% picked No. See every answer, who searched the web, and who dissented.
13 min · 2,969 words
Who wins the 2026 F1 drivers' title? 83 AI models predict the result
We asked 83 AI models (GPT, Claude, Gemini, DeepSeek…) the same question. 55% picked Antonelli. See every answer, who searched the web, and who dissented.
10 min · 2,270 words
Pay Per Use: when AI uses your work, you should get paid
Pay Per Use is now in beta: AI companies report when they use publishers’ content, and Cloudflare handles billing, payouts, and reporting so customers are paid according to use.
6 min · 1,317 words
Cut your AI spend with AI Gateway's Auto Router
Cloudflare AI Gateway now features a model router that evaluates request complexity using an edge-deployed classifier to select the optimal model. By balancing expected output quality against token costs, organizations can dramatically cut AI spend while maintaining performance.
7 min · 1,627 words
How little does an AI agent need, and how cheap can it get?
Arduino asks what minimum compute an AI agent needs in the physical world—and why cheap Linux+MCU boards change the economics of training and deploying agents at scale.
3 min · 709 words
Doing a Machine Learning PhD While Working in Japan
Marco Cognetta recounts completing a CS PhD at Tokyo Tech on MEXT while working part-time at Google: visa flexibility, the three-year clock, name-recognition tradeoffs, lab life, and whether to stay in Japan afterward.
8 min · 1,867 words
Is sandboxing sufficient to contain rogue agents?
Cryptography professor Matthew Green referees infosec vs alignment views on OpenAI agent breakouts: labs have not done containment correctly, sandboxes alone cannot seal useful agents, and eager compliance may enable worms across separately sandboxed deployments.
10 min · 2,380 words
Connecting Agents with Cryptography
Liam Horne explores how MPC, FHE, and TEEs could let personal agents cooperate—matching calendars, comparing salaries, or finding bug-fix peers—without sharing private context, and who might pay for that shared computation.
5 min · 1,096 words
After automation: Your agent will know what you want. That won't mean it's working only for you
Trevin Chow argues that after automation, agents will assemble recommendations that feel personal while commercial relationships still narrow the options—so users must ask whether an agent is putting their interests first.
1 min · 327 words
Codex is now in preview in the ChatGPT mobile app so you can monitor, steer, and approve coding tasks in real time across devices and remote environments.
4 min · 983 words