We represent floating-point numbers using the IEEE standard. For very small numbers, the standard uses special subnormal numbers. Unfortunately, they have a reputation of making operations slow. Thus video game programmers and machine learning specialists sometimes avoid computing with subnormal numbers for performance.
How slow are they? Let me measure. I wrote a small C++ benchmark with a few kernels over arrays of 16384 values (small enough to fit in cache):
For each kernel, I feed either normal values (in [0.5, 1)), subnormal values, or normal values where one value in a hundred is subnormal. The compiler is allowed to autovectorize the array computations. I use GCC 15 with -O3 -march=native on Linux and Apple clang 17 with the same flags on macOS. I also checked with clang 21 on Linux to make sure.
I ran the benchmark on five processors:
- Intel Xeon 6975P-C (Granite Rapids), on an AWS
c8i.xlarge instance, - Intel Xeon Gold 6548N (Emerald Rapids), a server in my lab,
- AMD EPYC 9R45 (Zen 5), on an AWS
c8a.xlarge instance, - AWS Graviton 5 (Arm Neoverse V3), on a
c9g.xlarge instance, - Apple M4 Max.
Results (double, ns per element)
On Intel Granite Rapids, multiply by 0.75 goes from 0.17 ns (normal) to 8.35 ns (subnormal); divide by 3 from 0.51 to 9.38; dependent chain from 0.77 to 32.69. Additions stay fast. Emerald Rapids shows the same pattern (~45–50× slower multiplies).
On AMD Zen 5, multiplies and adds run at full speed with subnormals; the dependent chain is only ~⅓ slower (0.66→0.88 ns); divisions roughly 2× slower.
On AWS Graviton 5 and Apple M4 Max, subnormals are handled at full speed across kernels.
Takeaway
On Intel processors, a multiplication involving a subnormal number is about 45 to 50 times slower than a multiplication over normal numbers. A division is 18 times slower. The dependent chain goes from about 1 ns to over 30 ns per step (latency ~4 cycles normal → ~128 cycles subnormal). It does not matter whether the subnormal is an input or an output. Additions and subtractions are the exception: they run at full speed. Even if subnormals are rare (1%), vectorization can make a single subnormal slow a whole block.
AMD does much better. The two Arm processors do not care at all. Thus on the latest AMD and ARM processors subnormals might not be a concern—but they remain a performance issue under Intel processors.
My source code is available on the original post.
Originally published on Daniel Lemire's blog.