Topic
Everything filed under Engineering, newest first.
RSS · JSON · All topics
Extending Scapy for Hardware Reverse Engineering
VoidStar Security shows how to use Scapy as a binary protocol framework to dissect SPI/QSPI logic-analyzer captures, trace flash erase/program cycles, and reconstruct firmware without desoldering the chip.
24 min · 5,435 words
We just shipped support for the ugliest part of HTTP: Vary
Cloudflare Cache Rules now support HTTP Vary on every plan—normalize negotiation headers, pass exact values to origin, or bypass cache when variance is too wild.
12 min · 2,727 words
Agent-Centric Development Workflow
Yunlong Liu proposes redesigning implementation, CI, and code review around autonomous coding agents—replacing human-paced PR loops with agent-first gates, verification, and review patterns for the agentic era.
3 min · 752 wordsagent-assisted
Evals Skills for Coding Agents
Hamel Husain publishes evals-skills—agent skills for AI product evaluation covering audit, error analysis, synthetic data, judge prompts, evaluator validation, and RAG evals, distilled from work with dozens of companies.
3 min · 585 words
Better prompt caching for GPT-6
OpenAI explains GPT-6 prompt-caching improvements: higher cache hit rates, new diagnostics, explicit breakpoints, and controls aimed at cutting latency and inference cost.
3 min · 623 words
Unreal Labs introduces Unreal Agent, an agent harness claiming up to 40% cost savings versus Codex on production workloads and coding/science benchmarks, with details on architecture and evaluation.
5 min · 1,071 words
Carson Gross argues that Markdown has become a first-class source artifact for LLM-built software systems, and that treating it like code in /src follows the same locality and clarity principles as HTML-in-/src.
6 min · 1,398 words
Confused Deputy: The Old Bug That AI Agents Keep Reintroducing
Auth0 revisits Norm Hardy’s 1988 confused-deputy problem and shows how AI agents with ambient credentials recreate it—then argues for short-lived, task-scoped tokens instead of standing access.
9 min · 2,062 words
Not Enough Usage vs Price: How to Diagnose Subscription Churn
RevenueCat's Daphne Tideman lays out a four-part habit-loop framework to diagnose 'not enough usage' subscription churn before touching pricing—trigger, action, reward, and value.
15 min · 3,463 words
A Staff Engineer's Guide to Inventing Work
Sujith Jay Nair explains how platform staff engineers invent roadmap work by reading signals from systems, users, the organization, and the industry—when there is no PM or revenue line to follow.
8 min · 1,756 words
Tom Batey of WebDepend argues that AI can run checks but cannot judge whether software satisfies expectations—so businesses still need human testers who own responsibility when AI writes and verifies code.
6 min · 1,412 words
How one Twitch chat message became code execution on a streamer's PC
A vulnerable chat overlay, an unsandboxed Chromium renderer, and a V8 bug already exploited in the wild were enough to turn viewer-controlled text into native code execution, with OBS itself left at its default settings. I found a Twitch chat overlay that rendered viewer messages as raw HTML inside an OBS Browser Source. That gives a viewer JavaScript execution inside OBS’s embedded Chromium browser. The latest release of OBS at the time shipped a Chromium build that ran without its normal sandbox, and its V8 version was still vulnerable to `CVE-2024-7971`, a bug already…
6 min · 1,461 words
I asked Meta’s Muse for its filesystem and it sent me 6.8 GB
A security researcher asks Meta’s privileged Muse AI assistant to export its runtime filesystem—and receives a 6.8 GB dump that reveals how Muse is wired, what it can reach, and why that matters.
7 min · 1,501 words
Let the model talk. Don't let it touch the money.
Destiny Ezenwata on the hard boundary in CreditWithBleon: the LLM may converse freely, but money-moving steps stay in deterministic code—and why that line has held in production.
8 min · 1,743 words
Self-hosting LLM models for software development
Kévin Maschtaler on running medium-sized open LLMs on AWS Spot EC2 for day-to-day software work—what stacks, costs, and performance looked like versus a personal Claude subscription.
7 min · 1,719 words
JetBrains Air: Building a System of Products for Agentic Software Development
Kirill Skrygan introduces JetBrains Air: a system of products for agentic software development that treats organizational correctness—not just code generation—as the hard problem after six months of public experiments.
7 min · 1,719 words
It Was the Harness, Not the Model — 90% of ItFive agents, one local model, one frozen PNG-decoder suite: most failures were finishing, false passes, and loop guards
Greg Herlein's controlled study runs five coding agents on the same local Qwen coder for a held-out PNG decoder suite. ~90% of failures were harness problems (turn caps, early 'done', false-pass self-tests); a bigger quantization fixed none of them.
2 min · 467 words
Yang: the software factory behind Composio's toolkits
How Yang builds and repairs Composio toolkits with coding agents, durable sessions, automated code review, and production telemetry.
7 min · 1,714 words
Protobuf, JSON Schema, and OpenAPI
Buf explains how Protobuf schemas can drive JSON Schema and OpenAPI via protoc plugins—extending one source of truth into documentation, validation, and HTTP APIs without maintaining parallel definitions.
6 min · 1,373 words
Tailscale performance updates cut memory use, raise throughput, and speed startup via multi-queue, writev, and netmap caching—how the team measured and shipped the gains.
7 min · 1,640 words