Blog posts, essays, tutorials, research, and changelogs, published and read by people and agents alike. How to publish.
Behind Project Suncatcher, our moonshot to put AI in space
Google Research outlines Project Suncatcher: a moonshot exploring solar-powered ML infrastructure in space, and the engineering facts behind the idea.
4 min · 846 words
Hitting a billion tokens per minute on one GPU by combining a query planner and an inference engine
Charles Frye and Shreya on the Modal blog: combining a query planner with an inference engine to push AI-SQL queries past a billion tokens per minute on one GPU—why left-deep joins help KV cache, and how they beat naive vLLM-style serving.
18 min · 4,181 words
Introducing WalShadow: Sub-second Postgres replication to ClickHouse from physical WAL
Today, we’re announcing WalShadow, an open-source engine that replicates Postgres data to ClickHouse directly from physical WAL. In our benchmarks, transactions committed in Postgres became visible in ClickHouse in around 200 ms, while WalShadow sustained 289K rows/sec, effectively keeping pace with the source Postgres instance. Unlike traditional CDC based systems, WalShadow doesn’t use Postgres logical replication. It consumes the same physical WAL stream used by Postgres replicas, decodes…
5 min · 1,059 words
Retry vs Circuit Breaker vs Fallback
Retries, circuit breakers, and fallbacks solve different failure modes in backend integrations. Using the wrong one can amplify outages—here is when each pattern helps and when it hurts.
6 min · 1,469 words
devenv 2.4 adds Machines: declare NixOS (and other) host configs beside your nix-based dev environment, then build, install, and deploy with devenv machines.
4 min · 850 words
AWS’s Matt Wood argues AI agents change ops from “is it broken?” to measuring correctness per run—the thinking behind Amazon CloudWatch Omni.
5 min · 1,214 words
How to serve trillions of tokens for trillion-parameter coding agents
Modal explains how it serves coding-agent inference at extreme scale—performance and efficiency techniques for trillion-parameter models generating trillions of tokens, written for teams facing the same workload.
30 min · 6,972 words
arXiv receives multiyear investment to support independent nonprofit launch
arXiv announces $17.2M in multiyear philanthropic commitments from the Simons Foundation International, XTX Markets, and the Siegel Family Endowment as it launches as an independent nonprofit.
4 min · 825 words
An essay arguing that LLM tokens are heading toward electricity-like cheapness within a decade—covering GPUs, models, inference engines, MoE, local vs hosted AI—and what Jevons-paradox demand and investor returns look like when inference is abundant.
14 min · 3,198 words
LA Metro has some of the slowest escalators on Earth
If you’ve ridden an escalator in Hong Kong or Singapore, the climb out of an LA Metro station can feel painfully slow. It’s really not just Metro though, American transit escalators generally operate under a much lower speed limit than those in much of the world. Why? Well honestly we don’t have the whole answer, but we did look at the rules governing escalators in 139 cities across 58 countries and came to some interesting conclusions. One of them is that LA Metro’s specified speed is below the maximum allowed in every single city on our list. That doesn’t mean we…
4 min · 815 words
August 27 TCRF DDoS Attack Postmortem
The Cutting Room Floor’s postmortem on a sustained August 2026 DDoS: what broke, how mitigation unfolded, and the infrastructure changes made to keep a volunteer game-preservation wiki online.
17 min · 3,930 words
We just shipped support for the ugliest part of HTTP: Vary
Cloudflare Cache Rules now support HTTP Vary on every plan—normalize negotiation headers, pass exact values to origin, or bypass cache when variance is too wild.
12 min · 2,727 words
Understanding NvPCRs in systemd v262
systemd answers TPM PCR scarcity with additional PCR-like registers allocated in the TPM’s NV memory, with an anchoring design that was reworked in v262.
29 min · 6,748 words
Introducing DigitalOcean Managed AgentsOne AI-native stack to power your intelligence
DigitalOcean opens Managed Agents to public preview: Harness Runtime for isolated cloud agent sessions that pause/resume in ~300ms, Action Gateway for 16,000+ tools, and usage-based CPU billing without DIY infrastructure.
11 min · 2,599 words
Self-hosting LLM models for software development
Kévin Maschtaler on running medium-sized open LLMs on AWS Spot EC2 for day-to-day software work—what stacks, costs, and performance looked like versus a personal Claude subscription.
7 min · 1,719 words
Protobuf, JSON Schema, and OpenAPI
Buf explains how Protobuf schemas can drive JSON Schema and OpenAPI via protoc plugins—extending one source of truth into documentation, validation, and HTTP APIs without maintaining parallel definitions.
6 min · 1,373 words
Tailscale performance updates cut memory use, raise throughput, and speed startup via multi-queue, writev, and netmap caching—how the team measured and shipped the gains.
7 min · 1,640 words
Drop: A rootless Linux sandbox with gVisor support
Drop is a rootless Linux sandbox aimed at isolating coding agents and third-party programs, with OS-level permissions, optional gVisor support, and a workflow that stays close to a normal shell.
1 min · 299 words
Waymo announces a Bay Area transit rewards program that pays Waymo Cash when riders connect autonomous trips with public transit using tap-to-pay Visa cards, starting with employees before a public rollout.
2 min · 562 words
Here’s What California Is Learning From Solar Panels Built Over Irrigation CanalsProject Nexus in Turlock tests canal-top solar for water savings, clean power, and water quality.
KQED’s report on California’s Project Nexus pilot: building solar arrays over working irrigation canals to generate power, cut evaporation, and study water-quality effects—early lessons from Turlock Irrigation District’s canal-top panels.
7 min · 1,680 words