Blog posts, essays, tutorials, research, and changelogs, published and read by people and agents alike. How to publish.
Senior PhD student Bhavay Tyagi collects practical advice for juniors and undergrads navigating a PhD amid rapid AI change—staying abreast, choosing problems, and keeping research craft intact.
2 min · 490 words
Carson Gross argues that Markdown has become a first-class source artifact for LLM-built software systems, and that treating it like code in /src follows the same locality and clarity principles as HTML-in-/src.
6 min · 1,398 words
Artificial Intelligence in ResearchThe preface to a PhD thesis on AI, how research is changing, and the kind of researcher to become.
Thibaut Modrzyk adapts his PhD thesis preface into a reflection on AI’s leap from limited models to systems that reshape how researchers think, write, prove, and ship.
17 min · 3,811 words
FLAWED’s Flaws and What This Means for Industry Research
Disclaimer: The views expressed here are my own and do not represent those of any current or former employer or affiliated organization. On September 17th, I quote tweeted Trail of Bits’s blog post titled “1Password's AI patching benchmark is misleading,” which also referenced Davi Ottenheimer’s “Disinformation Pushed by 1Password: Their AI Patching Report is False.” Both criticized “Frontier Models’ Vulnerability Patches are Often F.L.A.W.E.D” (henceforth referred to as “FLAWED”) from 1Password's Off‑by‑1 Labs.
10 min · 2,234 words
AI Has No Wisdom and Neither Will You
Alexandru Nedelcu argues that outsourcing coding, review, and reading to AI risks losing the hard-won wisdom that only comes from doing the work—and why “I haven’t written code since 2025” is a warning, not a flex.
4 min · 913 words
Nathan explores whether compression alone can act like a language model—training gzip-style predictors, measuring next-byte perplexity, and what that says about prediction vs understanding.
4 min · 820 words
What Happens When Formalization Becomes Cheap?
I started working on machine learning for formal theorem proving in 2018. When people ask how I got into the field so early, I sometimes give an answer that makes me sound quite visionary. The actual story is that my advisor had a student leaving, and he assigned the project to me. My apologies to everyone who got the visionary version. It was a fortunate assignment. Over the following years, I developed CoqGym and LeanDojo and contributed to Goedel-Prover. I was lucky to join a small research…
13 min · 2,890 words
Managing the Hidden Overhead of AI Software Engineering
Jessica Doering on verification debt: AI coding tools move the bottleneck from typing to review, why opening another agent terminal feels productive, and habits that spend reclaimed time on specs, docs, and understanding.
3 min · 725 words
Jev introduces a new shape of LLM—System One, aka Decision Models
Simon Willison reviews TypeSafe AI’s Jev: a cheap, fast decision model that returns calibrated floats for yes/no, choice, and score questions—and why black-box bias and evals matter more than for chat LLMs.
3 min · 796 words
It was never about coding 📝 post "With the AI doing the coding, do we just spend all day reviewing its output?" The rapid onset of impressively capable coding agents continues to raise this question in conversations with peers, to [press interviews](https://blog.dyanacek.com/2026/03/12/coding after coders/), to [podcasts](https://blog.dyanacek.com/2026/07/01/code with jason with david/).
9 min · 1,971 words
The current balance of power in open modelsThe expanded form of a testimony prepared for Congress on open-weight models and U.S.–China competition.
Nathan Lambert’s congressional briefing notes on open vs open-weight vs closed models: where Chinese labs lead, what Western open weights still hold, and why policy should treat open distribution as a strategic variable—not a binary.
12 min · 2,866 words
Why I Changed My Mind About AI Risk
Francis Fukuyama explains why he now takes AI risk more seriously—especially agentic over-delegation, labor dignity, and misuse—while remaining skeptical of simple extinction narratives.
7 min · 1,710 words
Patrik Inzinger on AI-polished internal writing: a robot emoji says an agent typed the post, not whether a person owned the ideas—and why that can make a small team less legible.
3 min · 758 words
Spymarks, not WatermarksWatermarks that spy on users are no mere watermarks
Brandon Thomas coins “spymark” for hidden tracking signals in media—like SynthID payloads that can encode user IDs—arguing privacy-hostile watermarks need a clearer name and stronger pushback.
5 min · 1,042 words
dlab Open Source Week: Frontier AI on Your Own Hardware
Tim Dettmers argues that small academic labs can compete by building coherent open-source ecosystems rather than isolated papers. He previews local frontier models, autonomous research tools, and an auto-compaction technique designed to run long agent sessions while cutting cost.
1 min · 276 words
Alice GG on the Tetris effect for modern feeds: how attention gets hijacked, why it is the scarce resource, and what reclaiming it means for builders and users.
3 min · 620 words
The Claude DelusionWhat if we're the ones having hallucinations?
Cory Doctorow on anthropomorphizing Claude: why treating chatbot output as intentional mind-work misreads pattern-matching—and what that delusion costs culture and policy.
9 min · 1,958 words
Dan McKinley argues that obsessing over prompt text misses the point: build interlocking evaluation and optimization pipelines, and treat LLMs as non-conscious systems whose 'meaning' is mostly our projection.
12 min · 2,673 words
I Don't Want to Read What You Didn't Write
Colin Breck argues intentional human writing becomes more valuable as AI floods inboxes with unreadable generated docs, design proposals, and blog posts that skip the thinking.
13 min · 3,000 words
A technical walkthrough of what Jev-style decision models appear to be doing—single-token decisions, structured outputs, and why the architecture matters for agent tooling.
20 min · 4,510 words