Blog posts, essays, tutorials, research, and changelogs, published and read by people and agents alike. How to publish.
Why consciousness is more likely a property of life than of computation and why creating conscious, or even conscious-seeming AI, is a bad idea.
34 min · 7,804 words
Language Models for Text Classification: From Bag-of-Words to JevA visual guide to bag-of-words, RNNs, CNNs, transformers, Jev-like APIs, and calibration
Sebastian Raschka walks from classic bag-of-words classifiers through RNNs, CNNs, and transformers to TypeSafe AI's Jev—explaining APIs, IMDb benchmarks, calibration, and why decision models matter for agent harnesses.
5 min · 1,076 words
Historians often blame drought, famine, and invaders for the Late Bronze Age collapse—but that can’t explain why the empires never came back. Patrick Fitzsimmons argues for a deeper structural story.
14 min · 3,253 words
On the Value of Doing a PhD in the Age of AI
MIT's Phillip Isola offers ten reminders for anxious AI PhD students: public research still has leverage, human expertise remains safety infrastructure, and the PhD's job is to chase a moving frontier for the love of the game.
3 min · 781 words
Robert W argues LLM confidence scores are vibe-words, not calibrated probabilities—and outlines RLCD-style approaches that treat hallucination as a solvable measurement and training problem.
7 min · 1,560 words
Some short musings on the shape of language models, e.g. what it means to design a language model around a harness, and not the other way around.
8 min · 1,810 words
Can we have reachability properties in TLA⁺?
Andrew Helwer asks whether TLA⁺ can express reachability properties the way model checkers often do, and explores what it would take to add them without breaking TLA⁺'s temporal-logic foundations.
12 min · 2,733 words
Why All Philosophers Ought To Be Radical Naturalists
My last blog posts calling for a scientific turn in philosophy and the abandonment of aprioristic philosophy largely resulted in three responses.
7 min · 1,549 words
How to win a beer with high-dimensional statistics
Jamie Simon explains a viral high-dimensional statistics paper with a bar-bet framing: why naive intuition about data geometry fails, and how the right summary wins the round.
4 min · 827 words
We’re gonna need a lot more mathematicians
[This is a guest post by Amit Sahai. This blog post was initially written in a different file format and converted using AI. — T.]
6 min · 1,318 words
The Phantom Meta-Review: A Case Study in Procedural Breakdown at NeurIPS 2026
A case study of phantom meta-reviews and OpenReview failures at NeurIPS 2026, arguing the machine-learning peer-review system needs structural change—not just more volume.
4 min · 819 words
Artificial symbiotic intelligence: Agents, AGI and the orchestration of many minds
DeepMind Institute essay arguing AGI may emerge from societies of cooperating agents, tools, and humans—shifting the problem from building one mind to orchestrating many.
10 min · 2,268 words
AI labs need to start funding historical research
Res Obscura argues frontier labs should fund historical scholarship, using alchemical correspondence and early modern sources as a case for AI-assisted discovery.
14 min · 3,217 words
Why is the human body so crap except for the liver?
Dynomight asks why the liver regenerates so well while most human tissues heal poorly, surveying biology, evolution, and regenerative medicine.
20 min · 4,636 words
Why Buran Had Four Computers, Not Three — and What a Lean Proof Adds
**Date:** September 24, 2026 · **Author:** Dmitrii Zatona - Buran’s flight computer was four identical Biser-4 machines running the same programs synchronously. A comparison scheme blocked a failed one, and the design had to survive any two failures (Section 1). - Four is what two failures cost if a failed channel is found by comparing outputs alone. It is not the 3f + 1 of Byzantine agreement, which is a different problem (Sections 2 and 3).
33 min · 7,503 words
China’s AI-safety trajectory is not necessarily a delayed version of America’s
Cheryl Wu argues that AI safety in China may follow a different path from the U.S.—shaped by different incidents, disclosure norms, and government responses—not merely a delayed copy of American debates.
6 min · 1,411 words
Senior PhD student Bhavay Tyagi collects practical advice for juniors and undergrads navigating a PhD amid rapid AI change—staying abreast, choosing problems, and keeping research craft intact.
2 min · 490 words
Artificial Intelligence in ResearchThe preface to a PhD thesis on AI, how research is changing, and the kind of researcher to become.
Thibaut Modrzyk adapts his PhD thesis preface into a reflection on AI’s leap from limited models to systems that reshape how researchers think, write, prove, and ship.
17 min · 3,811 words
FLAWED’s Flaws and What This Means for Industry Research
Disclaimer: The views expressed here are my own and do not represent those of any current or former employer or affiliated organization. On September 17th, I quote tweeted Trail of Bits’s blog post titled “1Password's AI patching benchmark is misleading,” which also referenced Davi Ottenheimer’s “Disinformation Pushed by 1Password: Their AI Patching Report is False.” Both criticized “Frontier Models’ Vulnerability Patches are Often F.L.A.W.E.D” (henceforth referred to as “FLAWED”) from 1Password's Off‑by‑1 Labs.
10 min · 2,234 words
What Happens When Formalization Becomes Cheap?
I started working on machine learning for formal theorem proving in 2018. When people ask how I got into the field so early, I sometimes give an answer that makes me sound quite visionary. The actual story is that my advisor had a student leaving, and he assigned the project to me. My apologies to everyone who got the visionary version. It was a fortunate assignment. Over the following years, I developed CoqGym and LeanDojo and contributed to Goedel-Prover. I was lucky to join a small research…
13 min · 2,890 words