{"article":{"slug":"a-40ms-go-gc-pause-caused-by-swap","title":"A 40ms Go GC pause caused by swap","subtitle":null,"summary":"Running swap in production to absorb memory spikes, the author found that Go's garbage collector reads off-heap metadata during a stop-the-world pause, and when that metadata is swapped out the pause jumps from a median of about 51 microseconds to 40ms. Includes eBPF measurements and a note on Go 1.26's Green Tea GC.","content_type":"blog_post","language":"en","canonical_url":"https://frn.sh/go-gc/","author":{"name":"Fernando Simões","url":"https://frn.sh/","person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"frn.sh","url":"https://frn.sh/","listing_slug":null,"listing":null},"topics":[{"name":"Programming","slug":"programming","url":"https://listedarticles.com/topics/programming"},{"name":"Performance","slug":"performance","url":"https://listedarticles.com/topics/performance"},{"name":"Linux","slug":"linux","url":"https://listedarticles.com/topics/linux"},{"name":"Systems Programming","slug":"systems-programming","url":"https://listedarticles.com/topics/systems-programming"}],"about_listings":[],"cover_image_url":null,"license":"CC-BY-SA-4.0","word_count":785,"reading_minutes":3,"published_at":"2026-09-13T00:00:00.000Z","added_at":"2026-10-05T05:08:02.526Z","updated_at":"2026-10-05T05:08:02.526Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":true},"profile_url":"https://listedarticles.com/articles/a-40ms-go-gc-pause-caused-by-swap","markdown_url":"https://listedarticles.com/articles/a-40ms-go-gc-pause-caused-by-swap.md","example":false,"citation":"Fernando Simões, frn.sh. \"A 40ms Go GC pause caused by swap.\" 13 Sept 2026. https://frn.sh/go-gc/ (CC-BY-SA-4.0)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://frn.sh/go-gc/"},"body_markdown":"I’m writing this so you don’t slap your forehead like I *almost* did when I decided to run swap in production to absorb memory spikes.\n\nI had a cgroup with two processes: one is a Go process that calls `io.ReadAll` and then `proto.Unmarshal`, creating a blob and then a graph struct (which is marked as `scan` by Go’s allocator). The other process is an HTTP server that mostly stays quiet.\n\nWhenever the collector runs, it reads those `scan` spans, pointer by pointer, and decides what to do with them. So I thought: ok, under memory pressure, the kernel is going to evict pages to the swap device, but since the eviction is per cgroup, and not per process, both processes’ pages are going to be evicted - so there is only a small chance that this will turn into a sad dance of swap-in and swap-out between the kernel and the garbage collector.\n\nI was wrong. While experimenting with this, I found a problem that could’ve hurt me: Go’s garbage collector reads its metadata (outside the heap, in a region that is not [freed](https://cs.opensource.google/go/go/+/release-branch.go1.23:src/runtime/HACKING.md;l=214)) in a stop-the-world pause, and that metadata can be in swap.\n\nI did a mock run on a Hetzner box using kernel 6.8 with MGLRU enabled[1](https://frn.sh#fn:1)<sup>1</sup> You can find everything about these experiments: plots, the mock allocator, bpf scripts, python scripts, etc., here: [https://github.com/frnsimoes/go-gc-swap-cost](https://github.com/frnsimoes/go-gc-swap-cost). The median pause was around 51 us. With the metadata on the NVMe, the worst pause was 40ms.\n\nTo check where those 40ms went, I wrote a small [bpf script](https://github.com/frnsimoes/go-gc-swap-cost/blob/main/bpf/stw.bt) that counts page faults while the world is stopped. This was the worst one: `39902 us, faults during it 228, 39013 us in faults`. 39 of those 40ms were spent in 228 page faults. Those faults happened inside the GC’s bookkeeping:\n\n```\n// addr2line output\n0x42e5c8  runtime.(*spanSet).reset         /usr/local/go/src/internal/runtime/atomic/types.go:194\n0x4219de  runtime.finishsweep_m            /usr/local/go/src/runtime/mcentral.go:71\n0x4629cf  runtime.gcStart.func2            /usr/local/go/src/runtime/mgc.go:724\n0x46dd8a  runtime.systemstack              /usr/local/go/src/runtime/asm_amd64.s:518\n0x4169dc  runtime.gcStart                  /usr/local/go/src/runtime/mgc.go:722\n0x4276a4  runtime.nextMarkBitArenaEpoch    /usr/local/go/src/runtime/mheap.go:2481\n0x421a65  runtime.finishsweep_m            /usr/local/go/src/runtime/mgcsweep.go:268\n0x4629cf  runtime.gcStart.func2            /usr/local/go/src/runtime/mgc.go:724\n0x46dd8a  runtime.systemstack              /usr/local/go/src/runtime/asm_amd64.s:518\n0x4169dc  runtime.gcStart                  /usr/local/go/src/runtime/mgc.go:722\n```\nThat’s a potential failure mode. Go’s GC has to stop the world at two points: when it performs a [sweep termination](https://cs.opensource.google/go/go/+/release-branch.go1.23:src/runtime/mgc.go;l=24), and when it performs a [mark termination](https://cs.opensource.google/go/go/+/release-branch.go1.23:src/runtime/mgc.go;l=61). We had 312 of those pauses in 30 minutes.\n\nSo here is why this happens: the runtime allocates those pages. They are not freed, but reused. Those pages are read in GC cycles. Because the kernel evicts pages [by age](https://elixir.bootlin.com/linux/v6.8/source/include/linux/mmzone.h), it sends the least recently accessed pages to swap. The GC runs, stops the world, tries to read those pages, but now we have a major page fault. The kernel needs to [read PTEs](https://elixir.bootlin.com/linux/v6.8/source/mm/memory.c#L5154), and then [call `do_swap_page`](https://elixir.bootlin.com/linux/v6.8/source/mm/memory.c#L5167), find a [new frame](https://elixir.bootlin.com/linux/v6.8/source/mm/swap_state.c#L454), [charge it to the cgroup](https://elixir.bootlin.com/linux/v6.8/source/mm/swap_state.c#L498), [read the pages](https://elixir.bootlin.com/linux/v6.8/source/mm/memory.c#L3913), [submit a bio](https://elixir.bootlin.com/linux/v6.8/source/mm/page_io.c#L482), [wait for the disk](https://elixir.bootlin.com/linux/v6.8/source/mm/memory.c#L3946), and put them [back in memory](https://elixir.bootlin.com/linux/v6.8/source/mm/memory.c#L4105) - just to keep it short.\n\nThose 40ms seem harmless at first. But we are talking about a stop-the-world pause. Those 40ms mean everything has stopped - in Go’s terminology, [every `P` has stopped](https://cs.opensource.google/go/go/+/release-branch.go1.23:src/runtime/proc.go;l=1586), so, for example, if a goroutine was waiting for I/O, during that pause the I/O might return and there would be no one to handle it. 40ms is 800 times the median pause. It happens two or three times per memory spike during the test. *It is* a lot.\n\nAnd then I noticed another thing: building one 511 KiB message, which usually takes 3-5 ms, jumped to 105 ms on the NVMe and 903 ms on Hetzner’s network volume. Per message, this costs more than the metadata pause. But only the goroutine doing the allocation pays that price, so at least it’s localized, and not global like the metadata one.\n\nI haven’t confirmed where that time goes, but even so I wanted to mention it here, because that’s another cost you would have to pay. In any case, I agree with [Chris Down](https://chrisdown.name/2018/01/02/in-defence-of-swap.html), swap is not *evil*. But, yeah, it didn’t behave well with garbage collection, and, in production, I’m collecting a lot.\n\n**Update, September 14**.\n\nSomeone asked me if Go 1.26’s [Green Tea garbage collector](https://go.dev/blog/greenteagc) changed the way the GC reads metadata. I measured it, and the impact is negligible.\n\n1.\nYou can find everything about these experiments: plots, the mock allocator, bpf scripts, python scripts, etc., here: [https://github.com/frnsimoes/go-gc-swap-cost](https://github.com/frnsimoes/go-gc-swap-cost)[↩︎](https://frn.sh#fnref:1)\n","body_html":"<p>I’m writing this so you don’t slap your forehead like I <em>almost</em> did when I decided to run swap in production to absorb memory spikes.</p>\n<p>I had a cgroup with two processes: one is a Go process that calls <code>io.ReadAll</code> and then <code>proto.Unmarshal</code>, creating a blob and then a graph struct (which is marked as <code>scan</code> by Go’s allocator). The other process is an HTTP server that mostly stays quiet.</p>\n<p>Whenever the collector runs, it reads those <code>scan</code> spans, pointer by pointer, and decides what to do with them. So I thought: ok, under memory pressure, the kernel is going to evict pages to the swap device, but since the eviction is per cgroup, and not per process, both processes’ pages are going to be evicted - so there is only a small chance that this will turn into a sad dance of swap-in and swap-out between the kernel and the garbage collector.</p>\n<p>I was wrong. While experimenting with this, I found a problem that could’ve hurt me: Go’s garbage collector reads its metadata (outside the heap, in a region that is not <a href=\"https://cs.opensource.google/go/go/+/release-branch.go1.23:src/runtime/HACKING.md;l=214\" rel=\"nofollow ugc noopener\">freed</a>) in a stop-the-world pause, and that metadata can be in swap.</p>\n<p>I did a mock run on a Hetzner box using kernel 6.8 with MGLRU enabled<a href=\"https://frn.sh#fn:1\" rel=\"nofollow ugc noopener\">1</a>&lt;sup&gt;1&lt;/sup&gt; You can find everything about these experiments: plots, the mock allocator, bpf scripts, python scripts, etc., here: <a href=\"https://github.com/frnsimoes/go-gc-swap-cost\" rel=\"nofollow ugc noopener\"><a href=\"https://github.com/frnsimoes/go-gc-swap-cost\" rel=\"nofollow ugc noopener\">https://github.com/frnsimoes/go-gc-swap-cost</a></a>. The median pause was around 51 us. With the metadata on the NVMe, the worst pause was 40ms.</p>\n<p>To check where those 40ms went, I wrote a small <a href=\"https://github.com/frnsimoes/go-gc-swap-cost/blob/main/bpf/stw.bt\" rel=\"nofollow ugc noopener\">bpf script</a> that counts page faults while the world is stopped. This was the worst one: <code>39902 us, faults during it 228, 39013 us in faults</code>. 39 of those 40ms were spent in 228 page faults. Those faults happened inside the GC’s bookkeeping:</p>\n<pre><code>// addr2line output\n0x42e5c8  runtime.(*spanSet).reset         /usr/local/go/src/internal/runtime/atomic/types.go:194\n0x4219de  runtime.finishsweep_m            /usr/local/go/src/runtime/mcentral.go:71\n0x4629cf  runtime.gcStart.func2            /usr/local/go/src/runtime/mgc.go:724\n0x46dd8a  runtime.systemstack              /usr/local/go/src/runtime/asm_amd64.s:518\n0x4169dc  runtime.gcStart                  /usr/local/go/src/runtime/mgc.go:722\n0x4276a4  runtime.nextMarkBitArenaEpoch    /usr/local/go/src/runtime/mheap.go:2481\n0x421a65  runtime.finishsweep_m            /usr/local/go/src/runtime/mgcsweep.go:268\n0x4629cf  runtime.gcStart.func2            /usr/local/go/src/runtime/mgc.go:724\n0x46dd8a  runtime.systemstack              /usr/local/go/src/runtime/asm_amd64.s:518\n0x4169dc  runtime.gcStart                  /usr/local/go/src/runtime/mgc.go:722</code></pre>\n<p>That’s a potential failure mode. Go’s GC has to stop the world at two points: when it performs a <a href=\"https://cs.opensource.google/go/go/+/release-branch.go1.23:src/runtime/mgc.go;l=24\" rel=\"nofollow ugc noopener\">sweep termination</a>, and when it performs a <a href=\"https://cs.opensource.google/go/go/+/release-branch.go1.23:src/runtime/mgc.go;l=61\" rel=\"nofollow ugc noopener\">mark termination</a>. We had 312 of those pauses in 30 minutes.</p>\n<p>So here is why this happens: the runtime allocates those pages. They are not freed, but reused. Those pages are read in GC cycles. Because the kernel evicts pages <a href=\"https://elixir.bootlin.com/linux/v6.8/source/include/linux/mmzone.h\" rel=\"nofollow ugc noopener\">by age</a>, it sends the least recently accessed pages to swap. The GC runs, stops the world, tries to read those pages, but now we have a major page fault. The kernel needs to <a href=\"https://elixir.bootlin.com/linux/v6.8/source/mm/memory.c#L5154\" rel=\"nofollow ugc noopener\">read PTEs</a>, and then <a href=\"https://elixir.bootlin.com/linux/v6.8/source/mm/memory.c#L5167\" rel=\"nofollow ugc noopener\">call <code>do_swap_page</code></a>, find a <a href=\"https://elixir.bootlin.com/linux/v6.8/source/mm/swap_state.c#L454\" rel=\"nofollow ugc noopener\">new frame</a>, <a href=\"https://elixir.bootlin.com/linux/v6.8/source/mm/swap_state.c#L498\" rel=\"nofollow ugc noopener\">charge it to the cgroup</a>, <a href=\"https://elixir.bootlin.com/linux/v6.8/source/mm/memory.c#L3913\" rel=\"nofollow ugc noopener\">read the pages</a>, <a href=\"https://elixir.bootlin.com/linux/v6.8/source/mm/page_io.c#L482\" rel=\"nofollow ugc noopener\">submit a bio</a>, <a href=\"https://elixir.bootlin.com/linux/v6.8/source/mm/memory.c#L3946\" rel=\"nofollow ugc noopener\">wait for the disk</a>, and put them <a href=\"https://elixir.bootlin.com/linux/v6.8/source/mm/memory.c#L4105\" rel=\"nofollow ugc noopener\">back in memory</a> - just to keep it short.</p>\n<p>Those 40ms seem harmless at first. But we are talking about a stop-the-world pause. Those 40ms mean everything has stopped - in Go’s terminology, <a href=\"https://cs.opensource.google/go/go/+/release-branch.go1.23:src/runtime/proc.go;l=1586\" rel=\"nofollow ugc noopener\">every <code>P</code> has stopped</a>, so, for example, if a goroutine was waiting for I/O, during that pause the I/O might return and there would be no one to handle it. 40ms is 800 times the median pause. It happens two or three times per memory spike during the test. <em>It is</em> a lot.</p>\n<p>And then I noticed another thing: building one 511 KiB message, which usually takes 3-5 ms, jumped to 105 ms on the NVMe and 903 ms on Hetzner’s network volume. Per message, this costs more than the metadata pause. But only the goroutine doing the allocation pays that price, so at least it’s localized, and not global like the metadata one.</p>\n<p>I haven’t confirmed where that time goes, but even so I wanted to mention it here, because that’s another cost you would have to pay. In any case, I agree with <a href=\"https://chrisdown.name/2018/01/02/in-defence-of-swap.html\" rel=\"nofollow ugc noopener\">Chris Down</a>, swap is not <em>evil</em>. But, yeah, it didn’t behave well with garbage collection, and, in production, I’m collecting a lot.</p>\n<p><strong>Update, September 14</strong>.</p>\n<p>Someone asked me if Go 1.26’s <a href=\"https://go.dev/blog/greenteagc\" rel=\"nofollow ugc noopener\">Green Tea garbage collector</a> changed the way the GC reads metadata. I measured it, and the impact is negligible.</p>\n<p>1.\nYou can find everything about these experiments: plots, the mock allocator, bpf scripts, python scripts, etc., here: <a href=\"https://github.com/frnsimoes/go-gc-swap-cost\" rel=\"nofollow ugc noopener\"><a href=\"https://github.com/frnsimoes/go-gc-swap-cost\" rel=\"nofollow ugc noopener\">https://github.com/frnsimoes/go-gc-swap-cost</a></a><a href=\"https://frn.sh#fnref:1\" rel=\"nofollow ugc noopener\">↩︎</a></p>","headings":[]}}