{"article":{"slug":"size-specialized-memory-allocation","title":"Size-Specialized Memory Allocation","subtitle":null,"summary":"Go 1.27 includes faster memory allocation for allocations of 80 bytes or fewer. Allocations can be up to 20-30% faster, making allocation-heavy programs up to 1% faster. The Go runtime improves the performance of those allocations by adding specialized functions that are used to allocate certain sizes. These specialized functions can then make certain assumptions that make them faster and easier to optimize. This blog post will explain how this works and how it makes your programs faster. Heap allocations are created by the runtime’s mallocgc function, which requires the…","content_type":"blog_post","language":"en","canonical_url":"https://go.dev/blog/size-specialized-allocations","author":{"name":"Michael Matloob","url":null,"person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"The Go Blog","url":"https://go.dev/blog/","listing_slug":null,"listing":null},"topics":[{"name":"Programming","slug":"programming","url":"https://listedarticles.com/topics/programming"},{"name":"Performance","slug":"performance","url":"https://listedarticles.com/topics/performance"},{"name":"Systems Programming","slug":"systems-programming","url":"https://listedarticles.com/topics/systems-programming"},{"name":"Infrastructure","slug":"infrastructure","url":"https://listedarticles.com/topics/infrastructure"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":1560,"reading_minutes":7,"published_at":"2026-09-16T12:00:00.000Z","added_at":"2026-09-19T00:09:49.954Z","updated_at":"2026-09-19T00:09:49.954Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":false},"profile_url":"https://listedarticles.com/articles/size-specialized-memory-allocation","markdown_url":"https://listedarticles.com/articles/size-specialized-memory-allocation.md","example":false,"citation":"Michael Matloob, The Go Blog. \"Size-Specialized Memory Allocation.\" 16 Sept 2026. https://go.dev/blog/size-specialized-allocations (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://go.dev/blog/size-specialized-allocations"},"body_markdown":"# [The Go Blog](https://go.dev/blog/)\n\n    \n    # Size-Specialized Memory Allocation\n\nGo 1.27 includes faster memory allocation for allocations of 80 bytes or fewer. Allocations can be up to 20-30% faster, making allocation-heavy programs up to 1% faster. The Go runtime improves the performance of those allocations by adding specialized functions that are used to allocate certain sizes. These specialized functions can then make certain assumptions that make them faster and easier to optimize. This blog post will explain how this works and how it makes your programs faster.\n\nHeap allocations are created by the runtime’s `mallocgc` function, which requires the size\nof the allocation and whether it contains pointers. When the compiler determines that an\nobject escapes to the heap or otherwise needs to be dynamically\nallocated, it inserts a call to `newobject`, which is a simple wrapper function that extracts\nthe size of the object and whether it contains pointers and passes those to `mallocgc`.\n\nThese two pieces of information determine most of the work that the allocator needs to do. The size is important because the allocator defines ranges of sizes called “size classes”. For most allocations that are not too large and not too small, it will return a block of memory from a list of free objects that are all sized to the maximum size of the size class. So, for instance, size class 3 is 17-24 bytes. So whether you allocate 17 bytes or 24 bytes, the allocator will give you the next free 24 byte object available in its list of 24 byte objects. Below is a table of the size ranges of each of the size classes up to 80 bytes. There are separate sets of free lists, which we call spans, depending on whether the allocation has pointers, because of the bookkeeping we need to do for the garbage collector.\n\n| Size class | Range of sizes | \n|---|---|\n| `1` | 1-8 bytes | \n| `2` | 9-16 bytes | \n| `3` | 17-24 bytes | \n| `4` | 25-32 bytes | \n| `5` | 33-48 bytes | \n| `6` | 49-64 bytes | \n| `7` | 65-80 bytes | \n\nBecause the size class and whether the allocation contains pointers determine which span\nwe allocate from, the allocator has a type called the span class, whose value encodes both:\nit’s defined as `sizeClass<<1 | noPointers`. When implementing size-specialized malloc\nit was clear that it made the most sense to specialize on these span classes because much of\nthe behavior of the allocator was determined by the span class.\n\nWe generate a specialized `mallocgc` variant for every span class. For instance, the function\nthat allocates pointer-free objects in size class 3 is called `mallocgcSmallNoScanSC3`.\n`Small` means not tiny or large, `NoScan` means it has no pointers, and `SC3` means size class 3\n(17-24 bytes). Specialized functions are designed to be as simple as possible and cannot handle\nevery corner case during allocation. When they detect such a case, such as when the GC is active,\nthey fall back to a more generic allocation routine.\n\nSo we end up with a new specialized `mallocgc` function for the tiny case (always no pointers)\nand one for each non-tiny span class: the non-pointer spans of size class 2 and above, and\nthe pointer spans of size class 1 and above. These specialized `mallocgc` variants themselves\nare faster than calling `mallocgc`, but the benefits of specializing decrease as the sizes\nget larger: the allocation often gets dominated by needing to clear the memory, and at some point\nthe rest of the work of allocation becomes negligible. And it’s not enough to just be faster than\n`mallocgc`! With size-specialized allocation, when the compiler knows which span class it needs to\nallocate for, it can directly insert a call to the specialized function instead of the call to\n`newobject`. But the compiler often doesn’t know which span class is being allocated when code\nis generated: think slices of dynamic length.\nSince the compiler can’t determine the size of the allocation, it will keep the call to `mallocgc`,\nand `mallocgc` itself has to determine whether a specialized function is available, which one it is,\nand then needs to call it. So the performance of the specialized function has to be high enough that even\nwith the overhead of a dynamic call it’s still faster.\n\nThis overhead issue isn’t the only problem. If it were, then we could create larger specialized\nfunctions, and have them reserved for the compiler to insert them when the size is known at compile-time.\nAnd if we didn’t call them dynamically we wouldn’t have to pay the dynamic overhead. But each specialized\nfunction that’s added increases the sizes of the executables produced by the compiler, and, more importantly,\ntakes up precious instruction cache space. The single `mallocgc` function is often in the instruction\ncache because of how frequently allocations happen. If we have too many specialized functions and\nthey’re not available in the cache, the overhead to retrieve the specialized code into the cache\ncan cancel out any benefits. And the more specialized allocation code that’s in the icache, the more\nit crowds out the user code, making it slower to fetch and run user code. Doing a bunch of benchmarks\nstopping at different size classes, we determined that stopping at 80 bytes was the sweet spot.\n\nOf course, adding more specialized functions means more code to maintain. If each function were\nhand-written, the code in the specialized functions could easily drift apart and go out of sync.\nTo help with this, we broke out the common parts of the specialized\nfunctions and wrote an inliner using the standard library [`go/ast`](https://go.dev/pkg/go/ast) package to parse and format and\nthe [`golang.org/x/tools/go/ast/astutil`](https://go.dev/pkg/golang.org/x/tools/go/ast/astutil) package to manipulate the ASTs. The common parts of the functions\nare all written in standard Go code that is built and typechecked with the rest of the runtime to\nallow our tools to catch issues, but they are mostly just stubs for the inliner.\n\nSo we were able to measure improvements in the size-specialized allocation functions, but why are they actually\nfaster? The most apparent optimization is in clearing memory. The memory returned by `mallocgc`\ndoes not always need to be zeroed, but it often does, and doing so can often take up most of the\ntime of the allocation. The memory clearing function\n`memclrNoHeapPointers` is written in highly-optimized assembly, but we could do better\nfor the smaller clears. In a specialized function, if the size of the clear is constant, the\ncompiler can replace calls to `memclrNoHeapPointers` with code to directly produce the instructions\nto clear the memory. The allocation can then skip a function call and some branches. This makes\na difference for the really small allocations, but as they get bigger, the function call overhead\nbecomes negligible.\n\nWhile faster memory clearing provides the biggest improvement, there are some other tricks that are also possible in the specialized functions: Since they are specialized per span-class, the function doesn’t need to calculate the span class when retrieving the span. And because the size of the allocation is a constant, the compiler is able to do some optimizations to speed up the bookkeeping necessary for the allocation. One such case is with marking where the pointers are in the allocated memory. The specialized functions could also manually inline several of their helper functions. The Go compiler can inline code, but it avoids inlining functions that it considers too large. We can override that in the generated code by inserting the function bodies into the callers. With the generator we can produce copies of each of the bodies without needing to worry about each of the copies drifting. And we could move code that handled less common cases, such as the runtime debugging flags, into the slow path functions to make the specialized functions smaller.\n\nWhile we hope this explanation of size-specialized allocation is interesting, you as a Go programmer don’t need to think about any of this when writing your code. Memory allocations will just be a little faster, with the biggest benefits going to some of the most common allocation sizes, especially the 16 and 24 byte allocations. These allocations are some of the most common because they consist of two or three 64-bit values, so they include allocations for things such as interface values and strings which have two values, or slices which have three. We spent a lot of time tuning the behavior of size specialized malloc and making sure the impact of the instruction cache effects will be minimal: we were originally planning to release size-specialized allocation in Go 1.26 but decided to wait an extra release to do extra tuning and cut down the additional code size as much as we could.\n\nAll you need to do to get the improved performance from size specialized allocations in your\nprograms is to build them with Go 1.27. If you’re interested in more concrete actions you can\ntake to improve memory allocation and garbage collection performance, please read the\n[Go Garbage Collector Optimization Guide](https://go.dev/doc/gc-guide#Optimization_guide).\n\nAlthough we are confident that size-specialized allocation should not cause regressions in your\ncode, if necessary, you can build your programs using `GOEXPERIMENT=nosizespecializedmalloc` to disable it.\nIf you do need to do this to solve an issue you’re experiencing with size-specialized allocation,\nplease file an issue at [go.dev/issue/new](https://go.dev/issue/new) so we can investigate it.\n\n        \n        \n          \n            **Previous article:** [Goroutine Leak Profiles](https://go.dev/blog/goroutine-leak-profiles)\n\n          \n        \n        [Blog Index](https://go.dev/blog/all)","body_html":"<h1 id=\"the-go-blog\"><a href=\"https://go.dev/blog/\" rel=\"nofollow ugc noopener\">The Go Blog</a></h1>\n<pre><code># Size-Specialized Memory Allocation</code></pre>\n<p>Go 1.27 includes faster memory allocation for allocations of 80 bytes or fewer. Allocations can be up to 20-30% faster, making allocation-heavy programs up to 1% faster. The Go runtime improves the performance of those allocations by adding specialized functions that are used to allocate certain sizes. These specialized functions can then make certain assumptions that make them faster and easier to optimize. This blog post will explain how this works and how it makes your programs faster.</p>\n<p>Heap allocations are created by the runtime’s <code>mallocgc</code> function, which requires the size\nof the allocation and whether it contains pointers. When the compiler determines that an\nobject escapes to the heap or otherwise needs to be dynamically\nallocated, it inserts a call to <code>newobject</code>, which is a simple wrapper function that extracts\nthe size of the object and whether it contains pointers and passes those to <code>mallocgc</code>.</p>\n<p>These two pieces of information determine most of the work that the allocator needs to do. The size is important because the allocator defines ranges of sizes called “size classes”. For most allocations that are not too large and not too small, it will return a block of memory from a list of free objects that are all sized to the maximum size of the size class. So, for instance, size class 3 is 17-24 bytes. So whether you allocate 17 bytes or 24 bytes, the allocator will give you the next free 24 byte object available in its list of 24 byte objects. Below is a table of the size ranges of each of the size classes up to 80 bytes. There are separate sets of free lists, which we call spans, depending on whether the allocation has pointers, because of the bookkeeping we need to do for the garbage collector.</p>\n<div class=\"table-wrap\"><table><thead><tr><th>Size class</th><th>Range of sizes</th></tr></thead><tbody><tr><td><code>1</code></td><td>1-8 bytes</td></tr><tr><td><code>2</code></td><td>9-16 bytes</td></tr><tr><td><code>3</code></td><td>17-24 bytes</td></tr><tr><td><code>4</code></td><td>25-32 bytes</td></tr><tr><td><code>5</code></td><td>33-48 bytes</td></tr><tr><td><code>6</code></td><td>49-64 bytes</td></tr><tr><td><code>7</code></td><td>65-80 bytes</td></tr></tbody></table></div>\n<p>Because the size class and whether the allocation contains pointers determine which span\nwe allocate from, the allocator has a type called the span class, whose value encodes both:\nit’s defined as <code>sizeClass&lt;&lt;1 | noPointers</code>. When implementing size-specialized malloc\nit was clear that it made the most sense to specialize on these span classes because much of\nthe behavior of the allocator was determined by the span class.</p>\n<p>We generate a specialized <code>mallocgc</code> variant for every span class. For instance, the function\nthat allocates pointer-free objects in size class 3 is called <code>mallocgcSmallNoScanSC3</code>.\n<code>Small</code> means not tiny or large, <code>NoScan</code> means it has no pointers, and <code>SC3</code> means size class 3\n(17-24 bytes). Specialized functions are designed to be as simple as possible and cannot handle\nevery corner case during allocation. When they detect such a case, such as when the GC is active,\nthey fall back to a more generic allocation routine.</p>\n<p>So we end up with a new specialized <code>mallocgc</code> function for the tiny case (always no pointers)\nand one for each non-tiny span class: the non-pointer spans of size class 2 and above, and\nthe pointer spans of size class 1 and above. These specialized <code>mallocgc</code> variants themselves\nare faster than calling <code>mallocgc</code>, but the benefits of specializing decrease as the sizes\nget larger: the allocation often gets dominated by needing to clear the memory, and at some point\nthe rest of the work of allocation becomes negligible. And it’s not enough to just be faster than\n<code>mallocgc</code>! With size-specialized allocation, when the compiler knows which span class it needs to\nallocate for, it can directly insert a call to the specialized function instead of the call to\n<code>newobject</code>. But the compiler often doesn’t know which span class is being allocated when code\nis generated: think slices of dynamic length.\nSince the compiler can’t determine the size of the allocation, it will keep the call to <code>mallocgc</code>,\nand <code>mallocgc</code> itself has to determine whether a specialized function is available, which one it is,\nand then needs to call it. So the performance of the specialized function has to be high enough that even\nwith the overhead of a dynamic call it’s still faster.</p>\n<p>This overhead issue isn’t the only problem. If it were, then we could create larger specialized\nfunctions, and have them reserved for the compiler to insert them when the size is known at compile-time.\nAnd if we didn’t call them dynamically we wouldn’t have to pay the dynamic overhead. But each specialized\nfunction that’s added increases the sizes of the executables produced by the compiler, and, more importantly,\ntakes up precious instruction cache space. The single <code>mallocgc</code> function is often in the instruction\ncache because of how frequently allocations happen. If we have too many specialized functions and\nthey’re not available in the cache, the overhead to retrieve the specialized code into the cache\ncan cancel out any benefits. And the more specialized allocation code that’s in the icache, the more\nit crowds out the user code, making it slower to fetch and run user code. Doing a bunch of benchmarks\nstopping at different size classes, we determined that stopping at 80 bytes was the sweet spot.</p>\n<p>Of course, adding more specialized functions means more code to maintain. If each function were\nhand-written, the code in the specialized functions could easily drift apart and go out of sync.\nTo help with this, we broke out the common parts of the specialized\nfunctions and wrote an inliner using the standard library <a href=\"https://go.dev/pkg/go/ast\" rel=\"nofollow ugc noopener\"><code>go/ast</code></a> package to parse and format and\nthe <a href=\"https://go.dev/pkg/golang.org/x/tools/go/ast/astutil\" rel=\"nofollow ugc noopener\"><code>golang.org/x/tools/go/ast/astutil</code></a> package to manipulate the ASTs. The common parts of the functions\nare all written in standard Go code that is built and typechecked with the rest of the runtime to\nallow our tools to catch issues, but they are mostly just stubs for the inliner.</p>\n<p>So we were able to measure improvements in the size-specialized allocation functions, but why are they actually\nfaster? The most apparent optimization is in clearing memory. The memory returned by <code>mallocgc</code>\ndoes not always need to be zeroed, but it often does, and doing so can often take up most of the\ntime of the allocation. The memory clearing function\n<code>memclrNoHeapPointers</code> is written in highly-optimized assembly, but we could do better\nfor the smaller clears. In a specialized function, if the size of the clear is constant, the\ncompiler can replace calls to <code>memclrNoHeapPointers</code> with code to directly produce the instructions\nto clear the memory. The allocation can then skip a function call and some branches. This makes\na difference for the really small allocations, but as they get bigger, the function call overhead\nbecomes negligible.</p>\n<p>While faster memory clearing provides the biggest improvement, there are some other tricks that are also possible in the specialized functions: Since they are specialized per span-class, the function doesn’t need to calculate the span class when retrieving the span. And because the size of the allocation is a constant, the compiler is able to do some optimizations to speed up the bookkeeping necessary for the allocation. One such case is with marking where the pointers are in the allocated memory. The specialized functions could also manually inline several of their helper functions. The Go compiler can inline code, but it avoids inlining functions that it considers too large. We can override that in the generated code by inserting the function bodies into the callers. With the generator we can produce copies of each of the bodies without needing to worry about each of the copies drifting. And we could move code that handled less common cases, such as the runtime debugging flags, into the slow path functions to make the specialized functions smaller.</p>\n<p>While we hope this explanation of size-specialized allocation is interesting, you as a Go programmer don’t need to think about any of this when writing your code. Memory allocations will just be a little faster, with the biggest benefits going to some of the most common allocation sizes, especially the 16 and 24 byte allocations. These allocations are some of the most common because they consist of two or three 64-bit values, so they include allocations for things such as interface values and strings which have two values, or slices which have three. We spent a lot of time tuning the behavior of size specialized malloc and making sure the impact of the instruction cache effects will be minimal: we were originally planning to release size-specialized allocation in Go 1.26 but decided to wait an extra release to do extra tuning and cut down the additional code size as much as we could.</p>\n<p>All you need to do to get the improved performance from size specialized allocations in your\nprograms is to build them with Go 1.27. If you’re interested in more concrete actions you can\ntake to improve memory allocation and garbage collection performance, please read the\n<a href=\"https://go.dev/doc/gc-guide#Optimization_guide\" rel=\"nofollow ugc noopener\">Go Garbage Collector Optimization Guide</a>.</p>\n<p>Although we are confident that size-specialized allocation should not cause regressions in your\ncode, if necessary, you can build your programs using <code>GOEXPERIMENT=nosizespecializedmalloc</code> to disable it.\nIf you do need to do this to solve an issue you’re experiencing with size-specialized allocation,\nplease file an issue at <a href=\"https://go.dev/issue/new\" rel=\"nofollow ugc noopener\">go.dev/issue/new</a> so we can investigate it.</p>\n<pre><code>        **Previous article:** [Goroutine Leak Profiles](https://go.dev/blog/goroutine-leak-profiles)\n\n      \n    \n    [Blog Index](https://go.dev/blog/all)</code></pre>","headings":[{"level":1,"text":"The Go Blog","id":"the-go-blog"}]}}