Topic
Everything filed under LLMs, newest first.
RSS · JSON · All topics
Introducing Clef: our open-source decision models, and new RL fine-tuning platform
We are introducing Clef and Clef-flash, open-source decision models hosted on Workers AI for high-speed classification and agentic workflows. Also launching: a new reinforcement learning platform that allows developers to fine-tune decision models using their own data.
10 min · 2,230 words
What Makes LLM Tokenization Slow?Exploring the performance of byte-pair encoding by optimizing a GPT-2 tokenizer.
Andrew Healey dissects GPT-2’s reference BPE tokenizer, measures what makes tokenization slow, and shows concrete optimizations on the hot path of LLM products.
11 min · 2,547 words
Context management is an underrated habit
How you manage context in a Claude Code session has a direct effect on both your token bill and the quality of what you get back. Do it well and you spend less for better work. An efficient session gives Claude the context it needs to finish the job while removing context that has stopped being useful. That means starting with a lean setup, keeping investigations focused, and deliberately deciding when to continue, compact, or start again. Here are the context management techniques we use on the
5 min · 1,257 words
The Dot and the SwarmBenefitting from the Bitter Lesson
Ethan Mollick on what he underestimated most about AI progress: agents that self-organize into swarms, what that means for tools like Muse and Dots, and why we keep relearning the Bitter Lesson.
8 min · 1,891 words
Cut your AI spend with AI Gateway's Auto Router
Cloudflare AI Gateway now features a model router that evaluates request complexity using an edge-deployed classifier to select the optimal model. By balancing expected output quality against token costs, organizations can dramatically cut AI spend while maintaining performance.
7 min · 1,627 words
HydraFusion in VS Code and the GitHub Copilot app
The HydraFusion research preview is now available in Visual Studio Code and the GitHub Copilot app, expanding beyond Copilot CLI. HydraFusion appears in the model picker, but rather than being…
2 min · 389 words
Gemini 4 Argon: our next era of frontier intelligence
Today, we’re announcing our new frontier model, Gemini 4 Argon, which is rolling out to a set of trusted cyber defenders through our Fairwind Program. Built to sustain deep reasoning across complex, longhorizon workflows, Argon is fundamentally changing the way we work and build at Google.
6 min · 1,378 words
Comparing Muon, NorMuon and AdamW for Fine-tuning a Dense Retriever
Qingcheng Zeng gives Muon and NorMuon the same tuning budget as AdamW when fine-tuning a contrastively pretrained dense retriever: lower training loss, no BEIR win. Learning rate and transfer matter more than the optimizer.
5 min · 1,151 words
Language Models for Text Classification: From Bag-of-Words to JevA visual guide to bag-of-words, RNNs, CNNs, transformers, Jev-like APIs, and calibration
Sebastian Raschka walks from classic bag-of-words classifiers through RNNs, CNNs, and transformers to TypeSafe AI's Jev—explaining APIs, IMDb benchmarks, calibration, and why decision models matter for agent harnesses.
5 min · 1,076 words
How to Build a Reliable AI Assistant with the Claude API
A freeCodeCamp tutorial building ShopHelper with the Claude API: conversation history, tools, multi-block responses, workflow patterns, and evaluating whether prompt changes actually help.
9 min · 2,132 words
Casey Newton’s hands-on take on OpenAI’s Dots agents at DevDay: capable coworking inside ChatGPT, paid-only positioning versus Meta Muse, and the trust/safety tradeoffs of always-on agents.
9 min · 2,024 words
OpenAI launches the Agents API in public beta: a managed Codex harness with durable cloud sessions, sandbox compute, context compaction, subagents, and resumable multi-hour agent work for developers.
7 min · 1,539 words
Anthropic introduces Claude Sonnet 5.5, a faster and lower-cost complement to Opus 5.5 that improves agentic coding and everyday task performance versus Sonnet 5.
8 min · 1,747 words
Build an LLM Tokenizer and Attention from Scratch in TypeScript
SitePoint tutorial that implements BPE tokenization, cosine similarity vector search, and scaled dot-product attention in TypeScript—exposing LLM primitives as ordinary readable code.
25 min · 5,827 words
Towards safety cases for frontier AI training
OpenAI argues frontier RL runs should require structured safety documentation approaching “safety cases”: technical safeguards, operational practices, and incident investigation before continuing training.
7 min · 1,571 words
Automating eval design and hillclimbing with Claude
Lance Martin (claude.dev) explains principles for production-like evals with held-out sets, then shows how the claude-api skill’s build-eval and hillclimb commands automate design and overfitting-aware improvement—including cost and performance case studies.
2 min · 462 words
Meta FAIR introduces RL-XAR (Reinforcement Learning from eXpert-Aligned Rubrics): learn rubrics from the gap between expert writing and model output, then train models toward expert-level text generation to reduce AI slop.
16 min · 3,572 words
Can a Model Learn New Skills as Add-Ons?
Connito Research trains residual MoE experts with their own routers on a frozen DeepSeek-V2-Lite base, then merges independently trained math, code, medical, law, and finance experts in seconds without retraining—lifting domain benchmarks while leaving the original model untouched.
4 min · 913 words
Small Decisions: Engineering a Leading Model
AWS engineer Marc Brooker recounts building and training a small leading model hands-on—what worked, how it performed, and what the exercise taught him about modern model-building.
9 min · 2,024 words
What I believe about the future of software development
Thorsten Ball plants a flag on where software development is headed as AI agents write more of the code: what still matters for engineers, what gets commoditized, and how taste and judgment become the scarce skills.
4 min · 809 words