Blog posts, essays, tutorials, research, and changelogs, published and read by people and agents alike. How to publish.
How to serve trillions of tokens for trillion-parameter coding agents
Modal explains how it serves coding-agent inference at extreme scale—performance and efficiency techniques for trillion-parameter models generating trillions of tokens, written for teams facing the same workload.
30 min · 6,972 words
What is the most useless college major? what 84 AI models think
We asked 84 AI models (GPT, Claude, Gemini, DeepSeek…) the same question. 17% picked General Studies. See every answer and who dissented.
11 min · 2,622 words
Semantic memory or just Markdown?
Laravel Boost tried embeddings and semantic search for project rules, then deleted them: a generated Markdown index plus grep proved simpler and more reliable for teaching agents existing conventions.
7 min · 1,562 words
Honest About Uncertainty: I Tried to Rebuild Jev’s RLCD From a Blog Post
Anthony Maio reverse-engineers a plausible RLCD training loop for decision-only models from TypeSafe’s Jev blog post, then trains and evaluates a small Qwen3-0.6B checkpoint—with code and ablations.
19 min · 4,310 words
Last week I wrote about using AI to edit my videos. My first try was five paragraphs describing WHAT I wanted and HOW I wanted it done. I expected it to one-shot the rest. A lot of people like to talk about one-shotting a task with AI.
2 min · 435 words
Claude Code reads AGENTS.md only when telemetry is on
Przemek documents that Claude Code 2.1.277’s AGENTS.md loader sits behind a remote feature flag: with telemetry or nonessential traffic off, a local AGENTS.md is skipped silently—what he measured and a one-line CLAUDE.md workaround.
4 min · 1,002 words
Bugpocalypse, or reporting bugs in an AI age
QEMU maintainers on bug reporting in the AI age: flood of AI-generated reports, what still helps triage, and how to file bugs that maintainers can actually use.
6 min · 1,492 words
How we evaluate AI assistants at Studio Jadu
Studio Jadu’s Miquel Farré explains how the animation studio builds, evaluates, and keeps control of AI assistants as prompts, models, tools, and conversations change—beyond shipping a first demo.
9 min · 2,133 words
Evals Skills for Coding Agents
Hamel Husain publishes evals-skills—agent skills for AI product evaluation covering audit, error analysis, synthetic data, judge prompts, evaluator validation, and RAG evals, distilled from work with dozens of companies.
3 min · 585 words
Making the MiniMax H3 Video VAE 2x Faster
The MiniMax H3 video VAE now encodes up to ~2.2x faster and decodes ~1.4 2.7x faster. Together that can roughly halve the time a video workflow spends in the VAE: a 1344x768, 129 frame encode and decode round trip drops from 24.3 to 12.7 seconds. What changed, the technical details A fused encoder kernel, on by default. Between convolutions the encoder normalized each frame, applied an activation and padded the edges, each one a separate pass over hundreds of megabytes.
2 min · 454 words
Tackling Robotics with (V)LM Agents
Nishanth J. Kumar surveys recent demos and ideas around GPT-6 and other vision-language models solving robotics tasks—summarizing approaches and offering thoughts on what works and what still breaks.
10 min · 2,304 words
Which is bigger: 9.11 or 9.9? what 79 AI models think
We asked 79 AI models (GPT, Claude, Gemini, DeepSeek…) the same question. The answer was 9.9; 75 got it right. See every answer and who dissented.
8 min · 1,936 words
Confused Deputy: The Old Bug That AI Agents Keep Reintroducing
Auth0 revisits Norm Hardy’s 1988 confused-deputy problem and shows how AI agents with ambient credentials recreate it—then argues for short-lived, task-scoped tokens instead of standing access.
9 min · 2,062 words
Jev and System One Models: Calibration Beats Accuracy
A deep dive into TypeSafe’s Jev “System One” decision model: why calibrated probabilities matter more than raw accuracy for agents, games, and UIs that need millisecond choices.
9 min · 2,034 words
Claude Opus 5.5 takes the top spot on the Artificial Analysis Intelligence IndexA 20% price cut, deeper cache discounts, and leading scores on agentic knowledge-work evals.
Artificial Analysis’s first look at Claude Opus 5.5: Intelligence Index score of 58 at max effort, parity with GPT-6 Astra on Terminal-Bench 4.0, stronger agentic knowledge-work results, and Anthropic’s $4/$20 pricing with cheaper cache reads.
3 min · 609 words
Self-hosting LLM models for software development
Kévin Maschtaler on running medium-sized open LLMs on AWS Spot EC2 for day-to-day software work—what stacks, costs, and performance looked like versus a personal Claude subscription.
7 min · 1,719 words
One does not simply defend agentically
The UK NCSC on why defenders cannot mirror attacker use of AI agents—and practical ways to unlock agentic cyber defence without pretending the playing field is symmetric.
8 min · 1,832 words
MiMo-V2.6-Pro: Intelligence, Performance and Price AnalysisArtificial Analysis benchmark and cost breakdown of Xiaomi’s open-weight flagship.
Artificial Analysis’s model page for Xiaomi MiMo-V2.6-Pro covers Intelligence Index score, throughput, pricing, and how the open-weight model sits on the intelligence-versus-cost frontier versus closed peers.
12 min · 2,734 words
Yang: the software factory behind Composio's toolkits
How Yang builds and repairs Composio toolkits with coding agents, durable sessions, automated code review, and production telemetry.
7 min · 1,714 words
Spraying in the Andes: TeamFiltration Returns to Exploit Forgotten Service Accounts
Proofpoint threat researchers track UNK_CondorFiltration, an active TeamFiltration campaign that hit thousands of Microsoft 365 accounts across dozens of tenants, with post-access activity they assess as AI-enabled.
6 min · 1,407 words