Topic
Everything filed under Machine Learning, newest first.
RSS · JSON · All topics
Gemini 3.8 text-to-speech says hello
Google introduces Gemini 3.8 Flash TTS and Flash-Lite TTS—more expressive audio models for custom character voices and scene dialogue across AI Studio, the Gemini API, Enterprise, Notebook, and Vids.
6 min · 1,376 words
SlopShape: Identifying AI-Generated Commercial Web Content
Research paper introducing SlopShape, a method for identifying AI-generated commercial web content from structural signals alone—motivation, method, and evaluation on web-scale data.
48 min · 11,021 words
Making the MiniMax H3 Video VAE 2x Faster
The MiniMax H3 video VAE now encodes up to ~2.2x faster and decodes ~1.4 2.7x faster. Together that can roughly halve the time a video workflow spends in the VAE: a 1344x768, 129 frame encode and decode round trip drops from 24.3 to 12.7 seconds. What changed, the technical details A fused encoder kernel, on by default. Between convolutions the encoder normalized each frame, applied an activation and padded the edges, each one a separate pass over hundreds of megabytes.
2 min · 454 words
Artificial Intelligence in ResearchThe preface to a PhD thesis on AI, how research is changing, and the kind of researcher to become.
Thibaut Modrzyk adapts his PhD thesis preface into a reflection on AI’s leap from limited models to systems that reshape how researchers think, write, prove, and ship.
17 min · 3,811 words
Jev and System One Models: Calibration Beats Accuracy
A deep dive into TypeSafe’s Jev “System One” decision model: why calibrated probabilities matter more than raw accuracy for agents, games, and UIs that need millisecond choices.
9 min · 2,034 words
Heretic tutorial: automatic censorship removal for language models
A hands-on tutorial for Heretic, an open-source tool that automatically removes refusal/censorship behaviors from language models—setup, workflow, and what to watch for.
7 min · 1,688 words
Nathan explores whether compression alone can act like a language model—training gzip-style predictors, measuring next-byte perplexity, and what that says about prediction vs understanding.
4 min · 820 words
Transformer Explainer: LLM Transformer Model Visually Explained
Georgia Tech’s Polo Club walks through GPT-2’s Transformer stack—embeddings, multi-head attention, MLP, sampling—with an interactive in-browser model for learning how next-token prediction works.
4 min · 827 words
What Happens When Formalization Becomes Cheap?
I started working on machine learning for formal theorem proving in 2018. When people ask how I got into the field so early, I sometimes give an answer that makes me sound quite visionary. The actual story is that my advisor had a student leaving, and he assigned the project to me. My apologies to everyone who got the visionary version. It was a fortunate assignment. Over the following years, I developed CoqGym and LeanDojo and contributed to Goedel-Prover. I was lucky to join a small research…
13 min · 2,890 words
MiMo-V2.6: Scaling Up Reinforcement Learning for Self-Improvement
Xiaomi open-sources MiMo-V2.6 Pro and Flash after large-scale live RL (~$3.5M, 750k trajectories), claiming top open-weight AA Index scores, agent parity with frontier models, and 7k+ RL environments.
4 min · 839 words
Jev introduces a new shape of LLM—System One, aka Decision Models
Simon Willison reviews TypeSafe AI’s Jev: a cheap, fast decision model that returns calibrated floats for yes/no, choice, and score questions—and why black-box bias and evals matter more than for chat LLMs.
3 min · 796 words
Halo: Frontier-Lab Training for Everyone
White Circle open-sources Halo, a Hugging Face–native training framework claiming up to ~2.8× faster post-training than stock TRL with lower memory use—from single GPU to multi-node, keeping checkpoints in native HF format.
18 min · 4,098 words
Reducing Image Generation cost with AMD and the Luminal Compiler
Luminal engineers show Flux.2 Klein 9B image generation costs cut by up to 47% on AMD MI300X versus an Nvidia H200, using the Luminal compiler.
15 min · 3,489 words
What Is RLCD? The Secret Behind Jev
Di Zhang explains RLCD (schema-conditioned Plackett–Luce reward modeling) and how Jev turns calibrated multiway decisions into a product—making the reward model the model rather than hiding it behind a generator.
10 min · 2,324 words
How GPT-6 Astra ascended NetHack: setup, agent loop, tool use, failure modes, and what beating a famously hard roguelike says about LLM agents in open-ended environments.
11 min · 2,477 words
Amit Shekhar walks through how LLM design moved from RNNs to attention, Transformers, scaling laws, Mixture of Experts, and the open problems still ahead.
25 min · 5,729 words
Dan McKinley argues that obsessing over prompt text misses the point: build interlocking evaluation and optimization pipelines, and treat LLMs as non-conscious systems whose 'meaning' is mostly our projection.
12 min · 2,673 words
Multimodal Agents: From Perception to Action
Illustrated notes from Berkeley’s LLM Agents lecture 7: OSWorld outcome checks, AgentTrek trajectories, TACO tools, and Aguvis grounding—why reading a screen is not the same as finishing the task.
8 min · 1,915 words
A technical walkthrough of what Jev-style decision models appear to be doing—single-token decisions, structured outputs, and why the architecture matters for agent tooling.
20 min · 4,510 words
The new season of Survivor airs next week. My favorite thing about watching Survivor is arguing about the strategy of the show: who should have won the season? Is it better to play under the radar? Or does making big moves set you up to win? Who is the greatest player of all time: Cirie? Tony? Boston Rob? I thought it’d be fun to build a machine learning model to help answer these questions.
13 min · 3,012 words