Topic
Everything filed under Machine Learning, newest first.
RSS · JSON · All topics
Why we built the fastest robust TTS model
Gradium's latest streaming TTS hits ~50ms time-to-first-audio while improving naturalness and hard cases like phone numbers—freeing latency budget for LLM turns and barge-in in voice agents.
2 min · 431 words
Surprisingly Complex Waves Reveal the Brain's Inner Workings
Unexpected patterns traveling across the human brain may be reorganizing its activity in real time.
9 min · 2,113 words
Functional ultrasound imaging from scratch
Functional ultrasound imaging (fUSI) is a brain imaging modality that is just starting to be demonstrated in human studies. I’m sure many will be familiar with structural ultrasound imaging, which forms static images using ultrasound, as in a sonogram1. Functional ultrasound instead captures movies of brain tissue very rapidly, tracking subtle changes caused by the blood flow that follows neural activity.
11 min · 2,451 words
Doing a Machine Learning PhD While Working in Japan
Marco Cognetta recounts completing a CS PhD at Tokyo Tech on MEXT while working part-time at Google: visa flexibility, the three-year clock, name-recognition tradeoffs, lab life, and whether to stay in Japan afterward.
8 min · 1,867 words
Gemini 4 Argon: our next era of frontier intelligence
Today, we’re announcing our new frontier model, Gemini 4 Argon, which is rolling out to a set of trusted cyber defenders through our Fairwind Program. Built to sustain deep reasoning across complex, longhorizon workflows, Argon is fundamentally changing the way we work and build at Google.
6 min · 1,378 words
Comparing Muon, NorMuon and AdamW for Fine-tuning a Dense Retriever
Qingcheng Zeng gives Muon and NorMuon the same tuning budget as AdamW when fine-tuning a contrastively pretrained dense retriever: lower training loss, no BEIR win. Learning rate and transfer matter more than the optimizer.
5 min · 1,151 words
Language Models for Text Classification: From Bag-of-Words to JevA visual guide to bag-of-words, RNNs, CNNs, transformers, Jev-like APIs, and calibration
Sebastian Raschka walks from classic bag-of-words classifiers through RNNs, CNNs, and transformers to TypeSafe AI's Jev—explaining APIs, IMDb benchmarks, calibration, and why decision models matter for agent harnesses.
5 min · 1,076 words
Hugging Face’s Tarek Ziadé explains Serge, a CI agent that finds Transformers failures, reproduces them on GPUs, writes patches, verifies them, and opens PRs—29 merges in ~80 days.
9 min · 2,070 words
Anthropic introduces Claude Sonnet 5.5, a faster and lower-cost complement to Opus 5.5 that improves agentic coding and everyday task performance versus Sonnet 5.
8 min · 1,747 words
Build an LLM Tokenizer and Attention from Scratch in TypeScript
SitePoint tutorial that implements BPE tokenization, cosine similarity vector search, and scaled dot-product attention in TypeScript—exposing LLM primitives as ordinary readable code.
25 min · 5,827 words
Meta FAIR introduces RL-XAR (Reinforcement Learning from eXpert-Aligned Rubrics): learn rubrics from the gap between expert writing and model output, then train models toward expert-level text generation to reduce AI slop.
16 min · 3,572 words
Can a Model Learn New Skills as Add-Ons?
Connito Research trains residual MoE experts with their own routers on a frozen DeepSeek-V2-Lite base, then merges independently trained math, code, medical, law, and finance experts in seconds without retraining—lifting domain benchmarks while leaving the original model untouched.
4 min · 913 words
On the Value of Doing a PhD in the Age of AI
MIT's Phillip Isola offers ten reminders for anxious AI PhD students: public research still has leverage, human expertise remains safety infrastructure, and the PhD's job is to chase a moving frontier for the love of the game.
3 min · 781 words
Small Decisions: Engineering a Leading Model
AWS engineer Marc Brooker recounts building and training a small leading model hands-on—what worked, how it performed, and what the exercise taught him about modern model-building.
9 min · 2,024 words
The systems that no one will test
The systems that no one will test This happened to me in 2020, and it has been on my mind again lately. During the worst of the pandemic, I found a vulnerability in a system that gave me access to the Brazilian federal system, and with that access I was able to retrieve information on any Brazilian (think of 200+ million people data).
4 min · 901 words
Throughout the history of AI, open research has played a critical role in driving progress. Today, many key details of frontier large language models (LLMs) remain proprietary, but open-weights model families— such as DeepSeek, Kimi, and MiMo —continue to provide a valuable window into the development process for modern LLMs.
7 min · 1,609 words
AMD to Acquire World Labs to Advance the Future of AI Compute
AMD announces a definitive agreement to acquire Fei-Fei Li’s World Labs, bringing spatial AI researchers and models in-house to shape future AI hardware, software, and systems.
4 min · 1,011 words
World Labs announces a definitive agreement to join AMD, arguing that scale, reach, and closer hardware partnership are needed to accelerate spatial and physical AI research.
1 min · 218 words
Robert W argues LLM confidence scores are vibe-words, not calibrated probabilities—and outlines RLCD-style approaches that treat hallucination as a solvable measurement and training problem.
7 min · 1,560 words
Turn GLM-5.3-Flash into a Jev-like System One model
Johannes Hötter walks through turning GLM-5.3-Flash into a fast Jev-like “System One” decision model—typed options with probabilities in a single forward pass, matching Jev’s accuracy and speed.
11 min · 2,426 words