Topic
Everything filed under Engineering, newest first.
RSS · JSON · All topics
Agentics: what is developer process automation?
theahura names developer process automation (DPA): intentional agent workflows for bug triage, docs gardening, dead-code cleanup—less hype than 'self-driving codebases,' more like Zapier for engineering.
6 min · 1,268 words
Ten years of tmux, and the 1,495 lines of zsh it cost me
Yogesh Lonkar measured a tmux status bar burning about 15% of a CPU core, then rewrote a decade of forking shell scripts into a leaner setup—and documents what 1,495 lines of zsh had been doing the whole time.
13 min · 2,974 words
How to serve trillions of tokens for trillion-parameter coding agents
Modal explains how it serves coding-agent inference at extreme scale—performance and efficiency techniques for trillion-parameter models generating trillions of tokens, written for teams facing the same workload.
30 min · 6,972 words
Harness Engineering Explained: The System Around an AI Coding Agent
Harness Engineering Explained: The System Around an AI Coding Agent The same coding agent gives one team clean merges and another a pile of reopened tickets. The difference is rarely the model. It's the six parts around it, checked one ticket at a time.
11 min · 2,490 words
Michael Heap on why leaders asking for “reasonable explanations” can stall change: good stories aren’t fixes, process theater replaces outcomes, and trust depends on asking the right question instead of drowning in details.
4 min · 845 words
Evading Machine Learning Based Detections
Companion post to an x33fcon talk on packer/loader architecture and how machine-learning-based detections work—plus practical ML-evasion techniques and RustPack 1.7 features that aim to bypass those detectors by default.
13 min · 2,880 words
Frequently Asked Questions (And Answers) About AI Evals
Hamel Husain and Shreya Shankar’s sharp FAQ on AI/LLM product evals: start with error analysis on real traces, build targeted evaluators, validate LLM judges with TPR/TNR, and avoid generic off-the-shelf metrics.
74 min · 17,016 words
Semantic memory or just Markdown?
Laravel Boost tried embeddings and semantic search for project rules, then deleted them: a generated Markdown index plus grep proved simpler and more reliable for teaching agents existing conventions.
7 min · 1,562 words
LA Metro has some of the slowest escalators on Earth
If you’ve ridden an escalator in Hong Kong or Singapore, the climb out of an LA Metro station can feel painfully slow. It’s really not just Metro though, American transit escalators generally operate under a much lower speed limit than those in much of the world. Why? Well honestly we don’t have the whole answer, but we did look at the rules governing escalators in 139 cities across 58 countries and came to some interesting conclusions. One of them is that LA Metro’s specified speed is below the maximum allowed in every single city on our list. That doesn’t mean we…
4 min · 815 words
Sunil Sadasivan on first-principles thinking for senior engineers working with AI agents: set experience aside long enough to see the problem clearly, then rebuild judgment on top of how agents actually work.
3 min · 624 words
Loris Cro reflects on Zigtoberfest remarks: approaching Zig as a multi-year journey across newcomers, intermediate, and advanced users—and what that framing means for the language and community.
4 min · 1,028 words
Fix a Ruby pg Segfault on macOS When Running Rails Tests in Parallel
I use minitest to run parallel tests with bin/rails test on GemChat (Rails 8.1, Ruby 4.0.7, pg 1.6.3, PostgreSQL in Docker). Running the Rails test suite in parallel should make it faster.
5 min · 1,084 words
Disclosure of Vulnerability in the Network Protocol
Radicle discloses two critical peer-to-peer network-protocol flaws: unencrypted node traffic and related issues affecting all prior releases, with remediation guidance for operators.
5 min · 1,133 words
From Any to Certainty: A Typechecking Journey
Aniket’s napari Island Dispatch post on an open-source typing journey: why the team migrated from mypy to Pyrefly, what improved, and how to choose a type checker for a large Python project.
9 min · 2,100 words
ReBarUEFI: Resizable BAR for almost any UEFI systemA DXE driver and tooling to enable Resizable BAR on motherboards that never exposed the option.
Open-source guide and driver for enabling Resizable BAR (ReBAR) on nearly any UEFI system: how the DXE module works, compatibility notes, and the steps to unlock larger GPU BAR sizes when firmware menus omit the feature.
4 min · 880 words
Why Claude Opus 5.5 Still Won't Fix Your AI Agents
VooStack argues that swapping in a stronger LLM won’t fix unreliable agents: the real work is orchestration, observability, and API design—the engineering discipline required to ship agents that hold up.
7 min · 1,545 words
700 MB/s of Kafka throughput, on Postgres
Profiling Kafgres (Kafka-compatible broker inside Postgres) from ~113 MB/s to ~700 MB/s via cached SPI plans, relaxed commits, separate topic disks, and WaitEventSet socket readiness.
5 min · 1,062 words
Bots for the last mile: Rollouts, Security Review
Cursor ships Rollouts (PR-to-production change monitors with regression actions) and Security Reviewer (context-aware vuln findings with fixes) for Teams and Enterprise.
2 min · 528 words
How we evaluate AI assistants at Studio Jadu
Studio Jadu’s Miquel Farré explains how the animation studio builds, evaluates, and keeps control of AI assistants as prompts, models, tools, and conversations change—beyond shipping a first demo.
9 min · 2,133 words
Geocodio explains how a two-person company moved from bash scripts to maintainable internal apps, the system that keeps those tools from rotting, and where they draw the line between building and buying.
10 min · 2,238 words