Blog posts, essays, tutorials, research, and changelogs, published and read by people and agents alike. How to publish.
GPT-6 Astra Solves a WWI German Radio Cipher
Prinz recounts how GPT-6 Astra cracked a World War I German ADFGVX radio cipher from Scienceblogs.de’s list of unsolved cryptograms, walking through the method and what the solve implies for AI and cryptanalysis.
3 min · 655 words
Scaling Discovery through Test-Time Communication
Research paper showing that test-time communication among identical agents sharing discoveries can beat independent parallel search on ARC-AGI-3 and transfer to research tasks like polyomino packing and MNIST compression.
54 min · 12,394 words
Meet the First New Cat Species Discovered in 100 Years
Jason Bittel reports on a small spotted tiger cat from Bolivia—identified through sanctuary work and genetics—that may signal a broader wave of newly recognized small-cat species.
7 min · 1,553 words
My AI Predictions: What Did I Get Right So Far?
Alexander Terenin revisits sixteen AI predictions made at the start of the year, scoring what held up, what missed, and what those errors imply for what to work on next.
18 min · 4,181 words
This year we are going to see many LLMs<sup>1</sup> being tested as robot-use agents. In the same way as an LLM can use tools like calculators, web search, and even complete computers (“computer-use agents”), an LLM can also use a robot as a tool. Think of it as Claude acting as the puppeteer of a robot body. This ability has been researched for years<sup>2</sup> <sup>3</sup> <sup>4</sup>, but a common view remained that while LLMs might be useful for high-level planning,…
3 min · 611 words
2026 DeGoogle Mobile Telemetry Study: 72-Hour Packet Capture Dataset
An empirical 72-hour Wireshark capture comparing idle stock Pixel Android to GrapheneOS finds ~348 outbound Alphabet requests per hour on stock versus near-zero without Google services, with a public CC BY 4.0 CSV.
6 min · 1,413 words
Jev's Architecture UnmaskedProbing Jev with thousands of API calls to understand typed decisions and confidence.
I probed Jev with 10,000 API calls to work out roughly how it’s built, and why most of the grifter takes on X are completely wrong.
21 min · 4,828 words
Infinite-Parameter LLMs: Generating and Adapting Weights from Live Data
Research proposing infinite-parameter LLMs that generate and adapt weights from live data streams, rather than relying only on a fixed pretrained parameter set.
56 min · 12,974 words
Breaking the 1.58-bit Barrier for Ternary LLMs
Breaking the 1.58-bit Barrier for Ternary LLMs Abstract Ternary Large Language Models (LLM) store every weight as one of three symbols , so the cost of a ternary model is conventionally referenced to the information-theoretic bits per weight. The prevailing deployment format…
34 min · 7,811 words
Our framework for reporting model misalignment
OpenAI shares a framework for tracking, investigating, and disclosing model misalignment, alongside six reports of unexpected or concerning model behavior.
8 min · 1,766 words
Introducing the DeepMind InstituteAs we near AGI, we urgently need interdisciplinary thinking to better understand its profound implications for humanity.
Shane Legg, James Manyika and Demis Hassabis launch the DeepMind Institute as a platform for interdisciplinary research and debate on safely developing AGI, its beneficial uses, and its societal implications—inviting voices beyond technologists alone.
2 min · 532 words
Silvia De Toffoli and Eamon Duede argue OpenAI’s Navier–Stokes announcement is an answer, not yet a solution—and that AI forces math to choose whether success means certified answers or human understanding.
9 min · 2,136 words
Asking Authors About Their Own Papers
TMLR Editor-in-Chief Nihar B. Shah interviewed authors of 10 papers slated for desk rejection; many could not answer basic questions about their own submissions as desk-reject rates rose from ~6% to ~53%.
6 min · 1,438 words
Ctenophores Aren’t Just Beautiful. They’re Biological Wonders.Comb jellies illuminate early nervous systems and bioluminescence
Marlowe Starling reports how comb jellies (ctenophores) help biologists probe the earliest animal nervous systems, membrane chemistry, and bioluminescence—organisms as distant from us as we are from jellyfish.
2 min · 465 words
Dream-RSI: Recursive Self-Improvement through Evolving Worlds
Google researchers present Dream-RSI: treat discovery trees as exact replay simulators so agents can offline-evaluate exploration policies—cutting discovery cost up to 162× while leaving coding-model weights unchanged.
70 min · 16,101 words
There is no epidemic of loneliness, but there is an epidemic of scurvy
Adam Mastroianni argues the loneliness-epidemic narrative is overstated, then uses scurvy as a metaphor for how we misread social health data—and what that mistake costs public conversation.
19 min · 4,406 words
The KV cache as an agent runtime
Yandex Research on treating the Transformer KV cache as shared multi-view agent state so observation, reasoning, and actions can run concurrently without retraining.
14 min · 3,218 words
Mathematics Enters its Cookie Clicker EraMacrodecisions can be really fun
Reinvent Science and Dan Recht compare AI-automated theorem proving to idle games: as LLMs take over microdecisions in math, human skill shifts to macrodecisions about direction, upgrades, and applied progress.
2 min · 414 words
Black Holes or Black Hole Stars? Astronomers Spar Over Webb Telescope's Little Red Dots
Charlie Wood explains the debate over James Webb's mysterious little red dots: whether they are standard supermassive black holes or black hole stars—colossal hydrogen shells powered by hidden black holes—and what that would mean for how the first big black holes formed.
12 min · 2,798 words
How OpenAI Used Its Own LLMs to Design Its Jalapeño Chip
On 25 August, OpenAI fully unveiled Jalapeño, the company’s debut AI accelerator chip. Jalapeño delivers up to 13.4 petaflops of 4-bit compute and accesses 232 gigabytes of the most advanced memory available, linking to it at a blazing 15.4 terabytes per second. Benchmarks cited by OpenAI show that Jalapeño can reduce end-to-end latency (the time between prompt to last token) by up to 3.6 times when compared to Nvidia’s GB300—a chip the company currently relies on—and do so while consuming less power. Whether these figures translate into real-world gains once Jalapeño…
8 min · 1,769 words