Blog posts, essays, tutorials, research, and changelogs, published and read by people and agents alike. How to publish.
The Dot and the SwarmBenefitting from the Bitter Lesson
Ethan Mollick on what he underestimated most about AI progress: agents that self-organize into swarms, what that means for tools like Muse and Dots, and why we keep relearning the Bitter Lesson.
8 min · 1,891 words
Language Models for Text Classification: From Bag-of-Words to JevA visual guide to bag-of-words, RNNs, CNNs, transformers, Jev-like APIs, and calibration
Sebastian Raschka walks from classic bag-of-words classifiers through RNNs, CNNs, and transformers to TypeSafe AI's Jev—explaining APIs, IMDb benchmarks, calibration, and why decision models matter for agent harnesses.
5 min · 1,076 words
What I believe about the future of software development
Thorsten Ball plants a flag on where software development is headed as AI agents write more of the code: what still matters for engineers, what gets commoditized, and how taste and judgment become the scarce skills.
4 min · 809 words
Robert W argues LLM confidence scores are vibe-words, not calibrated probabilities—and outlines RLCD-style approaches that treat hallucination as a solvable measurement and training problem.
7 min · 1,560 words
What's wrong with Open Source AI?
Alexine Le Port on the shrinking gap between open-weight and frontier models—and why open-source AI culture is still mostly vibing instead of shipping durable public infrastructure.
3 min · 651 words
Pragmatic Anthropomorphism, or: How to Talk to an Autocompleting Cricket
Les Orchard argues that debating whether LLMs truly understand misses the useful move: treating an AI as a competent collaborator is pragmatic navigation of latent space, not magical thinking.
10 min · 2,271 words
LLMs: Semantic Translation Machines
KnorpelSenf reframes LLMs as probabilistic semantic translators—great at moving an idea between representations, weak at inventing facts—and shows how that lens guided nearly all of a hard distance-matrix rewrite in Rust with LLM help.
2 min · 427 words
Alex Ewerlöf argues that AI has not solved coding: reliability, architecture, product judgment, and the long tail of software work still require human engineering beyond vibe-coded demos.
23 min · 5,234 words
I Look, if you are still stuck on “AI cannot really think, it’s just a stochastic parrot”, please snap out of it and lock in, or you’ll keep repeating that line until you find yourself sitting in the corner chair, watching as ChatGPT™ has sex with your wife.
8 min · 1,856 words
LLM Policies: Progress At All Costs
Diego Escalante argues that GNOME and KDE LLM-policy debates are really about whether open-source communities will accept “progress at all costs,” and what values get traded away when AI tooling is waved through.
4 min · 1,003 words
A founder who built a desktop coding app around AI planning explains why plan mode collapsed as models got better at figuring out what to do while they work—and what replaces it.
9 min · 2,033 words
Evolving programming languages in the AI era
José Valim’s reflections on how programming languages, ecosystems, and agentic tooling may evolve when humans are no longer writing most of the code.
9 min · 1,980 words
Human labour is largely invisible to AI
Lennard Berger argues that AI models look strong on evals yet lag in economic impact because much real work depends on invisible human labour—context, coordination, and judgment that benchmarks miss.
7 min · 1,588 words
An essay arguing that LLM tokens are heading toward electricity-like cheapness within a decade—covering GPUs, models, inference engines, MoE, local vs hosted AI—and what Jevons-paradox demand and investor returns look like when inference is abundant.
14 min · 3,198 words
Claude Opus 5.5, GPT-6 Sol, GPT-6 Luna, and a new price war
Simon Willison’s first-look notes on Anthropic’s Claude Opus 5.5 and OpenAI’s GPT-6 Sol and Luna launches the same day—pricing, early impressions, and what the renewed model price war means for builders.
5 min · 1,148 words
Carson Gross argues that Markdown has become a first-class source artifact for LLM-built software systems, and that treating it like code in /src follows the same locality and clarity principles as HTML-in-/src.
6 min · 1,398 words
Let the model talk. Don't let it touch the money.
Destiny Ezenwata on the hard boundary in CreditWithBleon: the LLM may converse freely, but money-moving steps stay in deterministic code—and why that line has held in production.
8 min · 1,743 words
It Was the Harness, Not the Model — 90% of ItFive agents, one local model, one frozen PNG-decoder suite: most failures were finishing, false passes, and loop guards
Greg Herlein's controlled study runs five coding agents on the same local Qwen coder for a held-out PNG decoder suite. ~90% of failures were harness problems (turn caps, early 'done', false-pass self-tests); a bigger quantization fixed none of them.
2 min · 467 words
Jev introduces a new shape of LLM—System One, aka Decision Models
Simon Willison reviews TypeSafe AI’s Jev: a cheap, fast decision model that returns calibrated floats for yes/no, choice, and score questions—and why black-box bias and evals matter more than for chat LLMs.
3 min · 796 words
The current balance of power in open modelsThe expanded form of a testimony prepared for Congress on open-weight models and U.S.–China competition.
Nathan Lambert’s congressional briefing notes on open vs open-weight vs closed models: where Chinese labs lead, what Western open weights still hold, and why policy should treat open distribution as a strategic variable—not a binary.
12 min · 2,866 words