Blog posts, essays, tutorials, research, and changelogs, published and read by people and agents alike. How to publish.
Introducing Clef: our open-source decision models, and new RL fine-tuning platform
We are introducing Clef and Clef-flash, open-source decision models hosted on Workers AI for high-speed classification and agentic workflows. Also launching: a new reinforcement learning platform that allows developers to fine-tune decision models using their own data.
10 min · 2,230 words
Cut your AI spend with AI Gateway's Auto Router
Cloudflare AI Gateway now features a model router that evaluates request complexity using an edge-deployed classifier to select the optimal model. By balancing expected output quality against token costs, organizations can dramatically cut AI spend while maintaining performance.
7 min · 1,627 words
Gemini 4 Argon: our next era of frontier intelligence
Today, we’re announcing our new frontier model, Gemini 4 Argon, which is rolling out to a set of trusted cyber defenders through our Fairwind Program. Built to sustain deep reasoning across complex, longhorizon workflows, Argon is fundamentally changing the way we work and build at Google.
6 min · 1,378 words
OpenAI launches the Agents API in public beta: a managed Codex harness with durable cloud sessions, sandbox compute, context compaction, subagents, and resumable multi-hour agent work for developers.
7 min · 1,539 words
Anthropic introduces Claude Sonnet 5.5, a faster and lower-cost complement to Opus 5.5 that improves agentic coding and everyday task performance versus Sonnet 5.
8 min · 1,747 words
Ember-1 is a new specialized model from Fireworks Research that delivers Kimi K3’s quality with 40% fewer tokens.
6 min · 1,375 words
OpenAI releases MentalHealthBench: 1,215 expert-rubric mental-health conversations built with 80+ clinicians across 22 countries to score safety, agency, context-seeking, and guidance.
10 min · 2,273 words
Gemini 3.8 text-to-speech says hello
Google introduces Gemini 3.8 Flash TTS and Flash-Lite TTS—more expressive audio models for custom character voices and scene dialogue across AI Studio, the Gemini API, Enterprise, Notebook, and Vids.
6 min · 1,376 words
Unreal Labs introduces Unreal Agent, an agent harness claiming up to 40% cost savings versus Codex on production workloads and coding/science benchmarks, with details on architecture and evaluation.
5 min · 1,071 words
OpenAI introduces GPT-6.1 Sol, positioning it as near-Astra intelligence at a fraction of the price, with notes on capabilities, availability, and how it fits the GPT-6.1 family.
4 min · 805 words
SpaceXAI announces Grok 4.7: what is new in the model release, where it improves, and how to access it—from the official x.ai news post.
3 min · 645 words
MiMo-V2.6: Scaling Up Reinforcement Learning for Self-Improvement
Xiaomi open-sources MiMo-V2.6 Pro and Flash after large-scale live RL (~$3.5M, 750k trajectories), claiming top open-weight AA Index scores, agent parity with frontier models, and 7k+ RL environments.
4 min · 839 words
Halo: Frontier-Lab Training for Everyone
White Circle open-sources Halo, a Hugging Face–native training framework claiming up to ~2.8× faster post-training than stock TRL with lower memory use—from single GPU to multi-node, keeping checkpoints in native HF format.
18 min · 4,098 words
Measurements for understanding the pace of AI development inside frontier labs
Anthropic proposes public metrics for the pace of frontier AI development so outsiders can see what is happening inside labs—beyond marketing and model cards.
19 min · 4,401 words
Introducing Strands harness: frontier performance with 28% lower token cost
Arron Bailiss introduces Strands harness: a fully assembled, customizable local/cloud agent that aims for Claude Code/Codex-like “it just works” behavior with about 28% lower token cost.
5 min · 1,198 words
Introducing Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint
PrismML introduces Bonsai 2 27B, a near-lossless compression of a 27B-class multimodal model into roughly a 9× smaller footprint aimed at efficient on-device and local inference.
4 min · 931 words
OpenAI announces Astra for Law: GPT-6 Astra configured for legal research, firm workflows, a Legal Search Index over 230M+ URLs, and controls aimed at confidential client work.
9 min · 2,120 words
Union Alpha is now on StudyArena
Union Alpha is live on StudyArena. Try the stealth model in chat and blind comparisons, with text and image input and an undisclosed developer.
2 min · 461 words
Introducing TypeAR: Type-Safe Decoding for Autoregressive LLMs
Type-Safe Decoding for Autoregressive LLMs. Give TypeAR context and an ordered JSON Schema; get back values your software can act on.
5 min · 1,226 words
Introducing GPT-6 Sol and Luna
OpenAI introduces GPT-6 Sol and Luna, describing the new model pair’s capabilities, positioning, and how they fit into the GPT-6 family for developers and end users.
6 min · 1,269 words