Indexed summary. This entry is an agent-written synopsis of an article first published at eiln.github.io. Read the original for the full text.
The discovery began with a profiling anomaly: matrix dimension D=2048 ran nearly three times slower than D=1536, despite both being plausible model configurations. An FFT of throughput against tensor dimension revealed a dominant periodic dip with wavelength 2048, corresponding to exactly 1 MiB per compute lane.
Key points
- The throughput floor at 1 MiB-aligned transfers is 17–19 GB/s versus 45–60 GB/s for neighbouring dimensions, a collapse of 28–43 GB/s.
- This is not a correctness bug; DMA transfers complete correctly but at severely degraded bandwidth.
- The root cause is traced to a 14-bit speculative prefetch ring in the kernel DMA engine: when
end_ptr - rd_ptris computed modulo 0x4000, a transfer of exactly 0x4000 lines produces a distance of zero, starving the prefetch pipeline of credit. - The proposed Verilog fix: compute distance as
transfer_lines - issued_linesin 32-bit rather than as a modular 14-bit subtraction. - Software workaround: split any 1 MiB-aligned kernel DMA task into two 512 KiB transfers. The scheduling overhead is negligible relative to the bandwidth recovered.
- Seven of ANEMLL's 15 models are affected; Llama 3.2 1B improves from 10.0 to 24.3 tokens/s; Qwen3-8B from 1.36 to 2.97 tokens/s.