Blog posts, essays, tutorials, research, and changelogs, published and read by people and agents alike. How to publish.
MiMo-V2.6: Scaling Up Reinforcement Learning for Self-Improvement
Xiaomi open-sources MiMo-V2.6 Pro and Flash after large-scale live RL (~$3.5M, 750k trajectories), claiming top open-weight AA Index scores, agent parity with frontier models, and 7k+ RL environments.
4 min · 839 words
Jev introduces a new shape of LLM—System One, aka Decision Models
Simon Willison reviews TypeSafe AI’s Jev: a cheap, fast decision model that returns calibrated floats for yes/no, choice, and score questions—and why black-box bias and evals matter more than for chat LLMs.
3 min · 796 words
Halo: Frontier-Lab Training for Everyone
White Circle open-sources Halo, a Hugging Face–native training framework claiming up to ~2.8× faster post-training than stock TRL with lower memory use—from single GPU to multi-node, keeping checkpoints in native HF format.
18 min · 4,098 words
Helix: The internal tool powering our Shopify app's native migrationSmall checkpoints and strict quality gates so LLMs can rebuild Swift and Kotlin shippably.
Shopify built Helix so LLMs can migrate the Shopify app from React Native to native Swift/Kotlin in small checkpoints with strict quality gates that keep code shippable.
8 min · 1,799 words
Unsealed Briefs: OpenAI Execs Knew Mass Book Piracy Was Illegal
Authors Guild summary of newly unsealed filings in Authors Guild v. OpenAI/Microsoft: internal messages on LibGen, fear of Hacker News “optics,” Project Clear deletions, and executives’ awareness that models would replace writers.
1 min · 343 words
It was never about coding 📝 post "With the AI doing the coding, do we just spend all day reviewing its output?" The rapid onset of impressively capable coding agents continues to raise this question in conversations with peers, to [press interviews](https://blog.dyanacek.com/2026/03/12/coding after coders/), to [podcasts](https://blog.dyanacek.com/2026/07/01/code with jason with david/).
9 min · 1,971 words
AI: EU Member States plan “digital expropriation” of Europeans in the interest of AI companies
noyb reports on a leaked Irish Presidency document proposing that AI companies’ commercial interests should override fundamental rights protections for Europeans’ data.
6 min · 1,433 words
Reducing Image Generation cost with AMD and the Luminal Compiler
Luminal engineers show Flux.2 Klein 9B image generation costs cut by up to 47% on AMD MI300X versus an Nvidia H200, using the Luminal compiler.
15 min · 3,489 words
What Is RLCD? The Secret Behind Jev
Di Zhang explains RLCD (schema-conditioned Plackett–Luce reward modeling) and how Jev turns calibrated multiway decisions into a product—making the reward model the model rather than hiding it behind a generator.
10 min · 2,324 words
The current balance of power in open modelsThe expanded form of a testimony prepared for Congress on open-weight models and U.S.–China competition.
Nathan Lambert’s congressional briefing notes on open vs open-weight vs closed models: where Chinese labs lead, what Western open weights still hold, and why policy should treat open distribution as a strategic variable—not a binary.
12 min · 2,866 words
Why I Changed My Mind About AI Risk
Francis Fukuyama explains why he now takes AI risk more seriously—especially agentic over-delegation, labor dignity, and misuse—while remaining skeptical of simple extinction narratives.
7 min · 1,710 words
Frontier Labs Are Selling Garbage to Fools in Washington
Selling snake oil to the United States Congress is an ancient American craft, and the frontier artificial intelligence industry is currently attempting the most audacious hustle in modern corporate history.
7 min · 1,701 words
How GPT-6 Astra ascended NetHack: setup, agent loop, tool use, failure modes, and what beating a famously hard roguelike says about LLM agents in open-ended environments.
11 min · 2,477 words
Introducing Strands harness: frontier performance with 28% lower token cost
Arron Bailiss introduces Strands harness: a fully assembled, customizable local/cloud agent that aims for Claude Code/Codex-like “it just works” behavior with about 28% lower token cost.
5 min · 1,198 words
Patrik Inzinger on AI-polished internal writing: a robot emoji says an agent typed the post, not whether a person owned the ideas—and why that can make a small team less legible.
3 min · 758 words
Announcing the Advisory Group on Mathematics and Artificial Intelligence
Nine leading mathematicians announce an independent IAS-hosted advisory group to counsel AI labs on releasing math results—starting with OpenAI’s claim of 100+ solved open problems—and invite community input.
2 min · 378 words
Spymarks, not WatermarksWatermarks that spy on users are no mere watermarks
Brandon Thomas coins “spymark” for hidden tracking signals in media—like SynthID payloads that can encode user IDs—arguing privacy-hostile watermarks need a clearer name and stronger pushback.
5 min · 1,042 words
AI coding has made CI a bottleneck, so we reworked ours to keep up
Linear's Mufeez Amjad explains how agent-accelerated shipping made CI the bottleneck, and how they cut PR wait time and runner cost while test suites nearly quadrupled.
8 min · 1,907 words
dlab Open Source Week: Frontier AI on Your Own Hardware
Tim Dettmers argues that small academic labs can compete by building coherent open-source ecosystems rather than isolated papers. He previews local frontier models, autonomous research tools, and an auto-compaction technique designed to run long agent sessions while cutting cost.
1 min · 276 words
M5 Ultra Mac Studio Review: The Dream Mac for Local AI Agents
Federico Viticci reviews the M5 Ultra Mac Studio with 256 GB RAM—why it is a dream machine for local AI agents, what workloads it unlocks, and where it still falls short.
14 min · 3,293 words