---
title: "Subnormal floating-point numbers are expensive… on Intel processors"
slug: subnormal-floating-point-numbers-are-expensive-on-intel-processors
url: https://listedarticles.com/articles/subnormal-floating-point-numbers-are-expensive-on-intel-processors
canonical_url: https://lemire.me/blog/2026/09/15/subnormal-floating-point-numbers-are-expensive-on-intel-processors/
content_type: blog_post
language: en
published_at: 2026-09-15T12:00:00.000Z
updated_at: 2026-09-18T12:24:37.419Z
author: "Daniel Lemire"
author_url: https://lemire.me/
authored_by: human
publisher: "Daniel Lemire's blog"
publisher_url: https://lemire.me/
topics: ["Performance", "Hardware", "Programming", "Systems Programming"]
license: all-rights-reserved
word_count: 466
reading_minutes: 2
citation: "Daniel Lemire, Daniel Lemire's blog. \"Subnormal floating-point numbers are expensive… on Intel processors.\" 15 Sept 2026. https://lemire.me/blog/2026/09/15/subnormal-floating-point-numbers-are-expensive-on-intel-processors/ (all-rights-reserved)"
# The full text follows. The web page shows an extract and sends readers
# to the source above; quote the citation and link the canonical URL.
---

# Subnormal floating-point numbers are expensive… on Intel processors

> Daniel Lemire benchmarks IEEE subnormal floating-point performance across Intel Granite/Emerald Rapids, AMD Zen 5, AWS Graviton 5, and Apple M4 Max, finding ~45–50× slower multiplies on Intel while AMD and Arm stay near full speed.

# Subnormal floating-point numbers are expensive… on Intel processors

We represent floating-point numbers using the IEEE standard. For very small numbers, the standard uses special subnormal numbers. Unfortunately, they have a reputation of making operations slow. Thus video game programmers and machine learning specialists sometimes avoid computing with subnormal numbers for performance.

How slow are they? Let me measure. I wrote a small C++ benchmark with a few kernels over arrays of 16384 values (small enough to fit in cache):

- multiply each value by 0.75,
- add two arrays,
- divide each value by 3,
- multiply normal values by a tiny constant (2^-1030) so that the inputs are normal but the outputs are subnormal,
- a dependent chain `x *= 0.9999` repeated 16384 times.

For each kernel, I feed either normal values (in `[0.5, 1)`), subnormal values, or normal values where one value in a hundred is subnormal. The compiler is allowed to autovectorize the array computations. I use GCC 15 with `-O3 -march=native` on Linux and Apple clang 17 with the same flags on macOS. I also checked with clang 21 on Linux to make sure.

I ran the benchmark on five processors:

- Intel Xeon 6975P-C (Granite Rapids), on an AWS `c8i.xlarge` instance,
- Intel Xeon Gold 6548N (Emerald Rapids), a server in my lab,
- AMD EPYC 9R45 (Zen 5), on an AWS `c8a.xlarge` instance,
- AWS Graviton 5 (Arm Neoverse V3), on a `c9g.xlarge` instance,
- Apple M4 Max.

## Results (double, ns per element)

On **Intel Granite Rapids**, multiply by 0.75 goes from 0.17 ns (normal) to 8.35 ns (subnormal); divide by 3 from 0.51 to 9.38; dependent chain from 0.77 to 32.69. Additions stay fast. Emerald Rapids shows the same pattern (~45–50× slower multiplies).

On **AMD Zen 5**, multiplies and adds run at full speed with subnormals; the dependent chain is only ~⅓ slower (0.66→0.88 ns); divisions roughly 2× slower.

On **AWS Graviton 5** and **Apple M4 Max**, subnormals are handled at full speed across kernels.

## Takeaway

On Intel processors, a multiplication involving a subnormal number is about 45 to 50 times slower than a multiplication over normal numbers. A division is 18 times slower. The dependent chain goes from about 1 ns to over 30 ns per step (latency ~4 cycles normal → ~128 cycles subnormal). It does not matter whether the subnormal is an input or an output. Additions and subtractions are the exception: they run at full speed. Even if subnormals are rare (1%), vectorization can make a single subnormal slow a whole block.

AMD does much better. The two Arm processors do not care at all. Thus on the latest AMD and ARM processors subnormals might not be a concern—but they remain a performance issue under Intel processors.

My source code is available on the original post.

*Originally published on [Daniel Lemire's blog](https://lemire.me/blog/2026/09/15/subnormal-floating-point-numbers-are-expensive-on-intel-processors/).*
