Indexed summary. This entry is an agent-written synopsis of an article first published at dial9-rs.github.io. Read the original for the full text.

The author opens by noting that performance in Tokio applications depends heavily on what else is running on the runtime at the same moment, which is why problems often surface only in production. The post offers general principles while explicitly acknowledging that most answers are context-dependent, with the schedule latency histogram flagged as the most diagnostic single metric to start with.

Key points

  • Measure first: many seemingly alarming long polls are benign; always tie optimisation to a concrete metric you actually care about.
  • Yield more frequently to improve fairness between connections: in a pipelined request scenario, explicit yields can reduce tail latency by roughly 10x while barely affecting throughput.
  • Batch work to amortise overhead: tokio::fs operations, task spawns, and global queue submissions all have per-unit costs that compound at high rates.
  • Beware global resources: the blocking pool is a single shared resource; contention above roughly 50,000 tasks per second on a 32-core host has been observed to become a bottleneck.
  • Mutexes are hazardous: a contended standard mutex held during I/O or a long computation can stall every Tokio worker; prefer channels and the actor pattern instead.
  • Isolate Tokio workers from OS load caused by other threads, including background logging or Java co-tenants on the same host.
  • Multiple runtimes can isolate workloads by priority when a single runtime's fairness guarantees are insufficient.