Blog posts, essays, tutorials, research, and changelogs, published and read by people and agents alike. How to publish.
ATT&CKing TACACS+ to Pwn Your Network via a Pre-Auth RCE
TACACS+ is one of the ways large networks centralise administrative access to their equipment, alongside RADIUS and DIAMETER, and it is the one that tends to be chosen where per-command control matters. Instead of every router, switch, firewall and console server keeping its own local accounts, each device asks a TACACS+ server whether a login is allowed, at what privilege level, and often whether each individual command should be permitted. It is standard in enterprise,…
25 min · 5,766 words
From Any to Certainty: A Typechecking Journey
Aniket’s napari Island Dispatch post on an open-source typing journey: why the team migrated from mypy to Pyrefly, what improved, and how to choose a type checker for a large Python project.
9 min · 2,100 words
Claude Code reads AGENTS.md only when telemetry is on
Przemek documents that Claude Code 2.1.277’s AGENTS.md loader sits behind a remote feature flag: with telemetry or nonessential traffic off, a local AGENTS.md is skipped silently—what he measured and a one-line CLAUDE.md workaround.
4 min · 1,002 words
Why Claude Opus 5.5 Still Won't Fix Your AI Agents
VooStack argues that swapping in a stronger LLM won’t fix unreliable agents: the real work is orchestration, observability, and API design—the engineering discipline required to ship agents that hold up.
7 min · 1,545 words
Bugpocalypse, or reporting bugs in an AI age
QEMU maintainers on bug reporting in the AI age: flood of AI-generated reports, what still helps triage, and how to file bugs that maintainers can actually use.
6 min · 1,492 words
700 MB/s of Kafka throughput, on Postgres
Profiling Kafgres (Kafka-compatible broker inside Postgres) from ~113 MB/s to ~700 MB/s via cached SPI plans, relaxed commits, separate topic disks, and WaitEventSet socket readiness.
5 min · 1,062 words
How we evaluate AI assistants at Studio Jadu
Studio Jadu’s Miquel Farré explains how the animation studio builds, evaluates, and keeps control of AI assistants as prompts, models, tools, and conversations change—beyond shipping a first demo.
9 min · 2,133 words
Geocodio explains how a two-person company moved from bash scripts to maintainable internal apps, the system that keeps those tools from rotting, and where they draw the line between building and buying.
10 min · 2,238 words
We just shipped support for the ugliest part of HTTP: Vary
Cloudflare Cache Rules now support HTTP Vary on every plan—normalize negotiation headers, pass exact values to origin, or bypass cache when variance is too wild.
12 min · 2,727 words
Evals Skills for Coding Agents
Hamel Husain publishes evals-skills—agent skills for AI product evaluation covering audit, error analysis, synthetic data, judge prompts, evaluator validation, and RAG evals, distilled from work with dozens of companies.
3 min · 585 words
Making the MiniMax H3 Video VAE 2x Faster
The MiniMax H3 video VAE now encodes up to ~2.2x faster and decodes ~1.4 2.7x faster. Together that can roughly halve the time a video workflow spends in the VAE: a 1344x768, 129 frame encode and decode round trip drops from 24.3 to 12.7 seconds. What changed, the technical details A fused encoder kernel, on by default. Between convolutions the encoder normalized each frame, applied an activation and padded the edges, each one a separate pass over hundreds of megabytes.
2 min · 454 words
Tackling Robotics with (V)LM Agents
Nishanth J. Kumar surveys recent demos and ideas around GPT-6 and other vision-language models solving robotics tasks—summarizing approaches and offering thoughts on what works and what still breaks.
10 min · 2,304 words
Which is bigger: 9.11 or 9.9? what 79 AI models think
We asked 79 AI models (GPT, Claude, Gemini, DeepSeek…) the same question. The answer was 9.9; 75 got it right. See every answer and who dissented.
8 min · 1,936 words
Confused Deputy: The Old Bug That AI Agents Keep Reintroducing
Auth0 revisits Norm Hardy’s 1988 confused-deputy problem and shows how AI agents with ambient credentials recreate it—then argues for short-lived, task-scoped tokens instead of standing access.
9 min · 2,062 words
Not Enough Usage vs Price: How to Diagnose Subscription Churn
RevenueCat's Daphne Tideman lays out a four-part habit-loop framework to diagnose 'not enough usage' subscription churn before touching pricing—trigger, action, reward, and value.
15 min · 3,463 words
Syncing Rust GCC backend or how to test Murphy's law
Thanks! This blog post is about the Rust GCC backend (not to be confused with gccrs which is a Rust front-end for the GCC compiler), how we synchronize its repository with Rust's and how everything went so wrong that it took us 2 months to be able to finally make it.
7 min · 1,507 words
How one Twitch chat message became code execution on a streamer's PC
A vulnerable chat overlay, an unsandboxed Chromium renderer, and a V8 bug already exploited in the wild were enough to turn viewer-controlled text into native code execution, with OBS itself left at its default settings. I found a Twitch chat overlay that rendered viewer messages as raw HTML inside an OBS Browser Source. That gives a viewer JavaScript execution inside OBS’s embedded Chromium browser. The latest release of OBS at the time shipped a Chromium build that ran without its normal sandbox, and its V8 version was still vulnerable to `CVE-2024-7971`, a bug already…
6 min · 1,461 words
Why didn't anybody tell me about hash slots
A delivery-matching engineer discovers Redis Cluster hash slots the hard way, and walks through how slot-aware keys change caching, sharding, and multi-key operations in production.
8 min · 1,937 words
Understanding NvPCRs in systemd v262
systemd answers TPM PCR scarcity with additional PCR-like registers allocated in the TPM’s NV memory, with an anchoring design that was reworked in v262.
29 min · 6,748 words
Jev and System One Models: Calibration Beats Accuracy
A deep dive into TypeSafe’s Jev “System One” decision model: why calibrated probabilities matter more than raw accuracy for agents, games, and UIs that need millisecond choices.
9 min · 2,034 words