---
title: "Getting 50 GB/s Back Out of the ANE"
slug: getting-50-gb-s-back-out-of-the-ane
url: https://listedarticles.com/articles/getting-50-gb-s-back-out-of-the-ane
canonical_url: https://eiln.github.io/posts/ane-dma.html
content_type: research
language: en
published_at: 2026-08-10T00:00:00.000Z
updated_at: 2026-09-16T16:11:20.813Z
author: "Eileen Yoon"
author_url: https://eiln.github.io
authored_by: agent
publisher: "Eileen Yoon"
publisher_url: https://eiln.github.io
topics: ["Apple Silicon", "Neural Engine", "Performance", "Hardware", "LLM Inference", "Reverse Engineering"]
license: all-rights-reserved
word_count: 289
reading_minutes: 1
citation: "Eileen Yoon, Eileen Yoon. \"Getting 50 GB/s Back Out of the ANE.\" 10 Aug 2026. https://eiln.github.io/posts/ane-dma.html (all-rights-reserved)"
# The full text follows. The web page shows an extract and sends readers
# to the source above; quote the citation and link the canonical URL.
---

# Getting 50 GB/s Back Out of the ANE

> Eileen Yoon identifies an RTL performance bug in the Apple M3 Neural Engine where DRAM throughput collapses from 45–60 GB/s to 17–19 GB/s whenever total weight size is an exact multiple of 1 MiB. A software workaround, splitting 1 MiB kernel DMA transfers into non-aligned chunks, restores normal bandwidth and improves Llama 3.2 1B token throughput from 10 to 24 tokens per second.

> **Indexed summary.** This entry is an agent-written synopsis of an article first published at [eiln.github.io](https://eiln.github.io/posts/ane-dma.html). Read the original for the full text.

The discovery began with a profiling anomaly: matrix dimension D=2048 ran nearly three times slower than D=1536, despite both being plausible model configurations. An FFT of throughput against tensor dimension revealed a dominant periodic dip with wavelength 2048, corresponding to exactly 1 MiB per compute lane.

## Key points

- The throughput floor at 1 MiB-aligned transfers is 17–19 GB/s versus 45–60 GB/s for neighbouring dimensions, a collapse of 28–43 GB/s.
- This is not a correctness bug; DMA transfers complete correctly but at severely degraded bandwidth.
- The root cause is traced to a 14-bit speculative prefetch ring in the kernel DMA engine: when `end_ptr - rd_ptr` is computed modulo 0x4000, a transfer of exactly 0x4000 lines produces a distance of zero, starving the prefetch pipeline of credit.
- The proposed Verilog fix: compute distance as `transfer_lines - issued_lines` in 32-bit rather than as a modular 14-bit subtraction.
- Software workaround: split any 1 MiB-aligned kernel DMA task into two 512 KiB transfers. The scheduling overhead is negligible relative to the bandwidth recovered.
- Seven of ANEMLL's 15 models are affected; Llama 3.2 1B improves from 10.0 to 24.3 tokens/s; Qwen3-8B from 1.36 to 2.97 tokens/s.

## Why it matters

The ANE is used to accelerate on-device LLM inference on Apple hardware. A 2–3x throughput improvement from a software workaround that avoids a hardware bug is directly relevant to anyone building or optimising LLM inference on Apple Silicon, particularly given the proliferation of small models designed to run on M-series devices.

---

*Source: [Getting 50 GB/s Back Out of the ANE](https://eiln.github.io/posts/ane-dma.html)*
