Quesma ran RTK (Rust Token Killer) against Terminal-Bench 2.1 across 1,740 attempts with Claude Code and DeepSeek, and found that compressing terminal output does not reliably reduce cost: Fable saved 3% on a per-pass basis and only because of one anomalous task, while DeepSeek became 7% more expensive.
Indexed summary. This entry is an agent-written synopsis of an article first published at quesma.com. Read the original for the full text.
RTK intercepts shell tool calls from AI coding agents and rewrites their output to remove redundant information — dropping file ownership and timestamps from ls, summarising git diffs, and so on. Social media claims of 60% token reductions have reached hundreds of thousands of views. Quesma's test asked a more precise question: does less terminal output translate into lower bills?
The benchmark used Terminal-Bench 2.1, a suite of tasks with heavy terminal interaction. Each task ran five times with RTK and five times without, on identical hardware, routes, and timeouts, for a total of 1,740 attempts. Cost was measured as actual billed tokens, not RTK's internal "rtk gain" metric.
Key points
RTK's "rtk gain" counts removed output bytes divided by four — not actual billed tokens. Across DeepSeek runs, RTK claimed 349 million tokens saved while the real cost rose.
Nearly all of Fable's apparent savings came from a single task (winning-avg-corewars) where RTK happened to halve the number of turns. Across the other 84 tasks, savings were under 1%.
DeepSeek showed the opposite result on that same task, and overall its per-task cost rose 17% on average with RTK enabled.
Terminal output is a small share of total cost for current models: roughly 11% of Fable's input tokens and 40% of DeepSeek's — and cached after the first read, so later references cost 1/10 or 1/30 of regular input.
Frontier models already self-limit terminal output using head, tail, and wc; RTK rewrites additional calls but can cause agents to take more turns, erasing the per-call savings.
Why it matters
The study is a concrete, well-controlled rebuttal to a widely circulated cost-saving claim. The methodology — separating per-turn token reduction from per-task cost — is a useful template for evaluating other AI coding optimisation tools.