Topic
Everything filed under AI, newest first.
RSS · JSON · All topics
On turn-taking in voice AI: why silence, timing, and semantic VAD matter as much as low latency for natural conversational agents.
6 min · 1,380 words
Infinite-Parameter LLMs: Generating and Adapting Weights from Live Data
Research proposing infinite-parameter LLMs that generate and adapt weights from live data streams, rather than relying only on a fixed pretrained parameter set.
56 min · 12,974 words
Breaking the 1.58-bit Barrier for Ternary LLMs
Breaking the 1.58-bit Barrier for Ternary LLMs Abstract Ternary Large Language Models (LLM) store every weight as one of three symbols , so the cost of a ternary model is conventionally referenced to the information-theoretic bits per weight. The prevailing deployment format…
34 min · 7,811 words
Our framework for reporting model misalignment
OpenAI shares a framework for tracking, investigating, and disclosing model misalignment, alongside six reports of unexpected or concerning model behavior.
8 min · 1,766 words
How an AI moratorium can save AI bosses
- How an AI moratorium can save AI bosses: If you can't impose switching costs, just eliminate the competition. - Hey look at this: Delights to delectate. - Object permanence: Flash Worms; This Film is Not Yet Rated; Libdem copyright sabotage; Religion worth more than Big Tech; Geographic tubemap; Selective censorship resistance. - Upcoming appearances: Budapest, Edmonton, Boston, South Bend, Hudson, Calgary, Winnipeg, Vancouver, Victoria, Ottawa. - Recent appearances: Where I've been. - Latest books: You keep readin' em, I'll keep writin' 'em. - Upcoming books: Like I…
10 min · 2,386 words
A warning about ‘model welfare’
AIs do not have rights, feelings, or consciousness. And we must not train them to act as though they do.
28 min · 6,458 words
Will the Fed raise rates this week? 76 AI models predict the result
We put "Will the Fed raise rates this week?" to 76 AI models (GPT, Claude, Gemini, DeepSeek…) at the same time. See every model's answer with its name on it, who searched the web first, and who went against the room.
3 min · 803 words
A framework for frontier AI and the dawning of a new ageA dynamic approach to testing frontier AI model capabilities that supports innovation and incentivizes responsible behavior.
Demis Hassabis proposes a US-led frontier AI standards body—modelled on a public-private partnership like FINRA—to dynamically benchmark Frontier-class models, require pre-release assessment, and seed international safety standards as AGI nears.
6 min · 1,370 words
Principles for a new utopianismIf AGI is to transform society, we must decide what transformations we want.
Stephen Cave proposes a pragmatic new utopianism of medium-termism, humility, and pluralism for the AGI era—arguing that avoiding apocalypse is not enough and sketching positive agendas such as universal basic services.
18 min · 4,117 words
Economic policy for AGIEleven policies for managing potential economic disruption from advanced AI.
Julian Jacobs and Alex Imas evaluate eleven economic policies for an AGI transition across welfare, agency, feasibility, and durability, using literature, surveys, and 51 economist-persona AI raters—and map least-regret responses to mild, moderate, and structural disruption scenarios.
18 min · 4,253 words
Introducing the DeepMind InstituteAs we near AGI, we urgently need interdisciplinary thinking to better understand its profound implications for humanity.
Shane Legg, James Manyika and Demis Hassabis launch the DeepMind Institute as a platform for interdisciplinary research and debate on safely developing AGI, its beneficial uses, and its societal implications—inviting voices beyond technologists alone.
2 min · 532 words
Silvia De Toffoli and Eamon Duede argue OpenAI’s Navier–Stokes announcement is an answer, not yet a solution—and that AI forces math to choose whether success means certified answers or human understanding.
9 min · 2,136 words
Asking Authors About Their Own Papers
TMLR Editor-in-Chief Nihar B. Shah interviewed authors of 10 papers slated for desk rejection; many could not answer basic questions about their own submissions as desk-reject rates rose from ~6% to ~53%.
6 min · 1,438 words
Getting Started with MCP Apps in Node.js
Valeri Karpov (Mastering JS) walks through MCP Apps in Node.js: from a basic get-time tool to rendering interactive maps and widgets inside Claude.
8 min · 1,901 words
Ian Duncan traces how parts of the rationalist/EA AI-safety milieu incubated salvation narratives, abusive experiments, race science, and authoritarian affection—and why that history matters as alumni steer frontier labs.
48 min · 11,153 words
On learning programming in an age of LLMs
Mark Seemann answers a reader’s letter on learning to program in the age of LLMs: which fundamentals still matter, how to practice, and how to keep agency when models can generate working code.
9 min · 2,004 words
Dream-RSI: Recursive Self-Improvement through Evolving Worlds
Google researchers present Dream-RSI: treat discovery trees as exact replay simulators so agents can offline-evaluate exploration policies—cutting discovery cost up to 162× while leaving coding-model weights unchanged.
70 min · 16,101 words
Migrating the GitHub Copilot runtime to Rust, using Copilot
Stephen Toub recounts porting GitHub Copilot’s agent runtime from TypeScript/Node to 800k+ lines of production Rust with Copilot agents across 128 incremental PRs, and what the performance and process lessons were.
64 min · 14,699 words
You are the AI agent's harnessPreventing hallucinations upstream by treating the engineer as the harness.
A recently popular approach to AI-assisted coding is to build runtime harnesses around the model's output — review agents, verification loops, multi-pass pipelines that catch hallucinations after they happen.
14 min · 3,269 words
Training a 4B model to produce 81% faster query plans than Postgres
Leis et al. asked this exact question in 2015. Then, they asked it again 10 years later. Despite an enormous body of research spanning a decade since their original exploration, they found that query optimizers continue to leave much to be desired. I was surprised when I first learned about this. A Postgres database should know everything about the stuff that lives in its tables, no? How hard can it be?
41 min · 9,393 words