Topic
Everything filed under LLMs, newest first.
RSS · JSON · All topics
Is Meta’s Muse secretly running an OpenAI model?
Is Meta’s Muse secretly running an OpenAI model? I found a model labeled azure/muse-special while Muse was building my website.
4 min · 870 words
I Look, if you are still stuck on “AI cannot really think, it’s just a stochastic parrot”, please snap out of it and lock in, or you’ll keep repeating that line until you find yourself sitting in the corner chair, watching as ChatGPT™ has sex with your wife.
8 min · 1,856 words
LLM Policies: Progress At All Costs
Diego Escalante argues that GNOME and KDE LLM-policy debates are really about whether open-source communities will accept “progress at all costs,” and what values get traded away when AI tooling is waved through.
4 min · 1,003 words
Why updatedInput in a PreToolUse Hook Doesn’t Rewrite the Command
A deep dive into Claude Code PreToolUse hooks: why returning updatedInput does not rewrite the shell command, and how permission decisions actually work.
17 min · 3,840 words
“When a measure becomes a target, it ceases to be a good measure” – Goodhart’s law Current AI research, especially the frontier LLM research, is dominated by benchmarks. It is the first thing we look at when a new model comes out, it is the headline of each release, and they dominate the discourse when […]
13 min · 2,940 words
A founder who built a desktop coding app around AI planning explains why plan mode collapsed as models got better at figuring out what to do while they work—and what replaces it.
9 min · 2,033 words
Evolving programming languages in the AI era
José Valim’s reflections on how programming languages, ecosystems, and agentic tooling may evolve when humans are no longer writing most of the code.
9 min · 1,980 words
Human labour is largely invisible to AI
Lennard Berger argues that AI models look strong on evals yet lag in economic impact because much real work depends on invisible human labour—context, coordination, and judgment that benchmarks miss.
7 min · 1,588 words
Hitting a billion tokens per minute on one GPU by combining a query planner and an inference engine
Charles Frye and Shreya on the Modal blog: combining a query planner with an inference engine to push AI-SQL queries past a billion tokens per minute on one GPU—why left-deep joins help KV cache, and how they beat naive vLLM-style serving.
18 min · 4,181 words
Running local LLMs on your Mac: what fits, what's free, and what's overkill
A practical guide to running language models on Apple Silicon: which sizes fit common Macs, free options that work well, and when bigger local models are overkill.
5 min · 1,125 words
Mistral Vibe Permission Bypass and Arbitrary Code Execution
SecMate details CVE-2026-87987 and CVE-2026-87984 in Mistral Vibe: shell permission bypasses that let a coding agent reach arbitrary code execution when those controls are treated as a security boundary.
7 min · 1,648 words
How to serve trillions of tokens for trillion-parameter coding agents
Modal explains how it serves coding-agent inference at extreme scale—performance and efficiency techniques for trillion-parameter models generating trillions of tokens, written for teams facing the same workload.
30 min · 6,972 words
An essay arguing that LLM tokens are heading toward electricity-like cheapness within a decade—covering GPUs, models, inference engines, MoE, local vs hosted AI—and what Jevons-paradox demand and investor returns look like when inference is abundant.
14 min · 3,198 words
Frequently Asked Questions (And Answers) About AI Evals
Hamel Husain and Shreya Shankar’s sharp FAQ on AI/LLM product evals: start with error analysis on real traces, build targeted evaluators, validate LLM judges with TPR/TNR, and avoid generic off-the-shelf metrics.
74 min · 17,016 words
Honest About Uncertainty: I Tried to Rebuild Jev’s RLCD From a Blog Post
Anthony Maio reverse-engineers a plausible RLCD training loop for decision-only models from TypeSafe’s Jev blog post, then trains and evaluates a small Qwen3-0.6B checkpoint—with code and ablations.
19 min · 4,310 words
Why Claude Opus 5.5 Still Won't Fix Your AI Agents
VooStack argues that swapping in a stronger LLM won’t fix unreliable agents: the real work is orchestration, observability, and API design—the engineering discipline required to ship agents that hold up.
7 min · 1,545 words
Ember-1 is a new specialized model from Fireworks Research that delivers Kimi K3’s quality with 40% fewer tokens.
6 min · 1,375 words
Mercury 2.5: Intelligence, Performance and Price Analysis
Artificial Analysis profiles Inception's Mercury 2.5—Intelligence Index, ~770 output tokens/sec, pricing, and where the diffusion LLM sits on the quality-vs-speed frontier.
12 min · 2,677 words
OpenAI releases MentalHealthBench: 1,215 expert-rubric mental-health conversations built with 80+ clinicians across 22 countries to score safety, agency, context-seeking, and guidance.
10 min · 2,273 words
Gemini 3.8 text-to-speech says hello
Google introduces Gemini 3.8 Flash TTS and Flash-Lite TTS—more expressive audio models for custom character voices and scene dialogue across AI Studio, the Gemini API, Enterprise, Notebook, and Vids.
6 min · 1,376 words