Everything below is reproducible — harness, frozen ledgers, one-click notebook: github.com/Travis42/telegraph-test.

Numbers from a 50-passage, ~1,300-question benchmark:

Cross-family matrix (readers = foreign models answering from GLM-5.3-Flash’s records; writers = GLM-5.3-Flash answering from theirs)<sup>1</sup>Savings shown use the lowercase instruction, on each provider’s own meter. Without that word, models write cablese in ALL CAPS and the styling costs 14–19 points: gemma 25.0%, qwen 29.8%, GLM 33.9%. *gpt-5-mini cannot disable reasoning, and compression makes it think — its writes bill about double plain. The one family where this technique does not pay. :

ModelRoleplaintext acccablese accrecovery ratio (1.00 = plaintext control)token savings
gemma-4-31breader70.9%77.5%1.09—
qwen3.8-27breader72.3%79.4%1.10—
nemotron-3-120breader76.1%76.7%1.01—
gemma-4-26breader71.8%76.9%1.07—
gemma-4-31bwriter——1.0940.4%
qwen3.8-27bwriter——1.1048.9%
gpt-5-miniwriter——0.9917.7%*