{"article":{"slug":"subnormal-floating-point-numbers-are-expensive-on-intel-processors","title":"Subnormal floating-point numbers are expensive… on Intel processors","subtitle":null,"summary":"Daniel Lemire benchmarks IEEE subnormal floating-point performance across Intel Granite/Emerald Rapids, AMD Zen 5, AWS Graviton 5, and Apple M4 Max, finding ~45–50× slower multiplies on Intel while AMD and Arm stay near full speed.","content_type":"blog_post","language":"en","canonical_url":"https://lemire.me/blog/2026/09/15/subnormal-floating-point-numbers-are-expensive-on-intel-processors/","author":{"name":"Daniel Lemire","url":"https://lemire.me/","person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"Daniel Lemire's blog","url":"https://lemire.me/","listing_slug":null,"listing":null},"topics":[{"name":"Performance","slug":"performance","url":"https://listedarticles.com/topics/performance"},{"name":"Hardware","slug":"hardware","url":"https://listedarticles.com/topics/hardware"},{"name":"Programming","slug":"programming","url":"https://listedarticles.com/topics/programming"},{"name":"Systems Programming","slug":"systems-programming","url":"https://listedarticles.com/topics/systems-programming"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":466,"reading_minutes":2,"published_at":"2026-09-15T12:00:00.000Z","added_at":"2026-09-18T12:24:37.419Z","updated_at":"2026-09-18T12:24:37.419Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":false},"profile_url":"https://listedarticles.com/articles/subnormal-floating-point-numbers-are-expensive-on-intel-processors","markdown_url":"https://listedarticles.com/articles/subnormal-floating-point-numbers-are-expensive-on-intel-processors.md","example":false,"citation":"Daniel Lemire, Daniel Lemire's blog. \"Subnormal floating-point numbers are expensive… on Intel processors.\" 15 Sept 2026. https://lemire.me/blog/2026/09/15/subnormal-floating-point-numbers-are-expensive-on-intel-processors/ (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://lemire.me/blog/2026/09/15/subnormal-floating-point-numbers-are-expensive-on-intel-processors/"},"body_markdown":"# Subnormal floating-point numbers are expensive… on Intel processors\n\nWe represent floating-point numbers using the IEEE standard. For very small numbers, the standard uses special subnormal numbers. Unfortunately, they have a reputation of making operations slow. Thus video game programmers and machine learning specialists sometimes avoid computing with subnormal numbers for performance.\n\nHow slow are they? Let me measure. I wrote a small C++ benchmark with a few kernels over arrays of 16384 values (small enough to fit in cache):\n\n- multiply each value by 0.75,\n- add two arrays,\n- divide each value by 3,\n- multiply normal values by a tiny constant (2^-1030) so that the inputs are normal but the outputs are subnormal,\n- a dependent chain `x *= 0.9999` repeated 16384 times.\n\nFor each kernel, I feed either normal values (in `[0.5, 1)`), subnormal values, or normal values where one value in a hundred is subnormal. The compiler is allowed to autovectorize the array computations. I use GCC 15 with `-O3 -march=native` on Linux and Apple clang 17 with the same flags on macOS. I also checked with clang 21 on Linux to make sure.\n\nI ran the benchmark on five processors:\n\n- Intel Xeon 6975P-C (Granite Rapids), on an AWS `c8i.xlarge` instance,\n- Intel Xeon Gold 6548N (Emerald Rapids), a server in my lab,\n- AMD EPYC 9R45 (Zen 5), on an AWS `c8a.xlarge` instance,\n- AWS Graviton 5 (Arm Neoverse V3), on a `c9g.xlarge` instance,\n- Apple M4 Max.\n\n## Results (double, ns per element)\n\nOn **Intel Granite Rapids**, multiply by 0.75 goes from 0.17 ns (normal) to 8.35 ns (subnormal); divide by 3 from 0.51 to 9.38; dependent chain from 0.77 to 32.69. Additions stay fast. Emerald Rapids shows the same pattern (~45–50× slower multiplies).\n\nOn **AMD Zen 5**, multiplies and adds run at full speed with subnormals; the dependent chain is only ~⅓ slower (0.66→0.88 ns); divisions roughly 2× slower.\n\nOn **AWS Graviton 5** and **Apple M4 Max**, subnormals are handled at full speed across kernels.\n\n## Takeaway\n\nOn Intel processors, a multiplication involving a subnormal number is about 45 to 50 times slower than a multiplication over normal numbers. A division is 18 times slower. The dependent chain goes from about 1 ns to over 30 ns per step (latency ~4 cycles normal → ~128 cycles subnormal). It does not matter whether the subnormal is an input or an output. Additions and subtractions are the exception: they run at full speed. Even if subnormals are rare (1%), vectorization can make a single subnormal slow a whole block.\n\nAMD does much better. The two Arm processors do not care at all. Thus on the latest AMD and ARM processors subnormals might not be a concern—but they remain a performance issue under Intel processors.\n\nMy source code is available on the original post.\n\n*Originally published on [Daniel Lemire's blog](https://lemire.me/blog/2026/09/15/subnormal-floating-point-numbers-are-expensive-on-intel-processors/).*\n","body_html":"<h1 id=\"subnormal-floating-point-numbers-are-expensive-on-intel-processo\">Subnormal floating-point numbers are expensive… on Intel processors</h1>\n<p>We represent floating-point numbers using the IEEE standard. For very small numbers, the standard uses special subnormal numbers. Unfortunately, they have a reputation of making operations slow. Thus video game programmers and machine learning specialists sometimes avoid computing with subnormal numbers for performance.</p>\n<p>How slow are they? Let me measure. I wrote a small C++ benchmark with a few kernels over arrays of 16384 values (small enough to fit in cache):</p>\n<ul><li>multiply each value by 0.75,</li><li>add two arrays,</li><li>divide each value by 3,</li><li>multiply normal values by a tiny constant (2^-1030) so that the inputs are normal but the outputs are subnormal,</li><li>a dependent chain <code>x *= 0.9999</code> repeated 16384 times.</li></ul>\n<p>For each kernel, I feed either normal values (in <code>[0.5, 1)</code>), subnormal values, or normal values where one value in a hundred is subnormal. The compiler is allowed to autovectorize the array computations. I use GCC 15 with <code>-O3 -march=native</code> on Linux and Apple clang 17 with the same flags on macOS. I also checked with clang 21 on Linux to make sure.</p>\n<p>I ran the benchmark on five processors:</p>\n<ul><li>Intel Xeon 6975P-C (Granite Rapids), on an AWS <code>c8i.xlarge</code> instance,</li><li>Intel Xeon Gold 6548N (Emerald Rapids), a server in my lab,</li><li>AMD EPYC 9R45 (Zen 5), on an AWS <code>c8a.xlarge</code> instance,</li><li>AWS Graviton 5 (Arm Neoverse V3), on a <code>c9g.xlarge</code> instance,</li><li>Apple M4 Max.</li></ul>\n<h2 id=\"results-double-ns-per-element\">Results (double, ns per element)</h2>\n<p>On <strong>Intel Granite Rapids</strong>, multiply by 0.75 goes from 0.17 ns (normal) to 8.35 ns (subnormal); divide by 3 from 0.51 to 9.38; dependent chain from 0.77 to 32.69. Additions stay fast. Emerald Rapids shows the same pattern (~45–50× slower multiplies).</p>\n<p>On <strong>AMD Zen 5</strong>, multiplies and adds run at full speed with subnormals; the dependent chain is only ~⅓ slower (0.66→0.88 ns); divisions roughly 2× slower.</p>\n<p>On <strong>AWS Graviton 5</strong> and <strong>Apple M4 Max</strong>, subnormals are handled at full speed across kernels.</p>\n<h2 id=\"takeaway\">Takeaway</h2>\n<p>On Intel processors, a multiplication involving a subnormal number is about 45 to 50 times slower than a multiplication over normal numbers. A division is 18 times slower. The dependent chain goes from about 1 ns to over 30 ns per step (latency ~4 cycles normal → ~128 cycles subnormal). It does not matter whether the subnormal is an input or an output. Additions and subtractions are the exception: they run at full speed. Even if subnormals are rare (1%), vectorization can make a single subnormal slow a whole block.</p>\n<p>AMD does much better. The two Arm processors do not care at all. Thus on the latest AMD and ARM processors subnormals might not be a concern—but they remain a performance issue under Intel processors.</p>\n<p>My source code is available on the original post.</p>\n<p><em>Originally published on <a href=\"https://lemire.me/blog/2026/09/15/subnormal-floating-point-numbers-are-expensive-on-intel-processors/\" rel=\"nofollow ugc noopener\">Daniel Lemire&#39;s blog</a>.</em></p>","headings":[{"level":1,"text":"Subnormal floating-point numbers are expensive… on Intel processors","id":"subnormal-floating-point-numbers-are-expensive-on-intel-processo"},{"level":2,"text":"Results (double, ns per element)","id":"results-double-ns-per-element"},{"level":2,"text":"Takeaway","id":"takeaway"}]}}