{"article":{"slug":"page-table-memory-consumption","title":"Page table memory consumption","subtitle":null,"summary":"Fernando Simões digs into when page tables stop being a rounding error: shared mappings across many processes, Linux kernel history from Arcangeli to Qi Zheng's empty-page-table fix, and the trade-offs of huge pages.","content_type":"essay","language":"en","canonical_url":"https://frn.sh/pagetables/","author":{"name":"Fernando Simões","url":"https://frn.sh/","person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"frn.sh","url":"https://frn.sh/","listing_slug":null,"listing":null},"topics":[{"name":"Linux","slug":"linux","url":"https://listedarticles.com/topics/linux"},{"name":"Systems Programming","slug":"systems-programming","url":"https://listedarticles.com/topics/systems-programming"},{"name":"Performance","slug":"performance","url":"https://listedarticles.com/topics/performance"},{"name":"Hardware","slug":"hardware","url":"https://listedarticles.com/topics/hardware"},{"name":"History","slug":"history","url":"https://listedarticles.com/topics/history"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":1309,"reading_minutes":6,"published_at":"2026-10-01T00:00:00.000Z","added_at":"2026-10-04T23:17:11.139Z","updated_at":"2026-10-04T23:17:11.139Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":true},"profile_url":"https://listedarticles.com/articles/page-table-memory-consumption","markdown_url":"https://listedarticles.com/articles/page-table-memory-consumption.md","example":false,"citation":"Fernando Simões, frn.sh. \"Page table memory consumption.\" 1 Oct 2026. https://frn.sh/pagetables/ (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://frn.sh/pagetables/"},"body_markdown":"The other day I was reading [Linus Torvalds punch some people](https://yarchive.net/comp/powerpc_page_tables.html) over hashed page tables. His argument: in a tree, the entries of neighboring pages are adjacent, which means one cache-line fills several TLB entries at once. A hash table, on the other hand, scatters neighbors in buckets:\n\nThe fact is, you just don’t know what GOOD actually is.\n\n\nI’ll tell you: t a *good* TLB fill should pre-populate the TLB with all\nthe entries it can fit in one cache-line.  Do you realize that a\nbog-standard Intel CPU will fetch 8 TLB entries in one go? Together with\na self-mapping (or, as Andy points out, you can just cache the other\nlevels in dedicated caches), that means that with a single memory\nreference you get *eight_times* the coverage that the silly Power hash\ntables get.\n\n\nThis discussion happened in 2003, in 1997, on his master’s [thesis](https://www.cs.helsinki.fi/u/kutvonen/index_files/linus.pdf), he explains the adoption of the three multi-level page tables on the kernel. He seemed quite interested in latency, but maybe less interested in memory size:\n\nThere are also secondary concerns: the virtual machine memory mappings must be memoryefficient, so that the mapping information does not take up a lot of physical memory that could be used to better advantage for file system caching or running user programs.\n\n\nIn 2020, Shakeel Butt found the exact case where memory efficiency is not a secondary concern at all. [Shakeel](https://lore.kernel.org/linux-mm/20201126005603.1293012-1-shakeelb@google.com/#R) sent a patch adding a page table metric to `memory.stat`. [Roman Gushchin](https://lkml.rescloud.iu.edu/hypermail/linux/kernel/2011.3/06476.html) asked what was the use case for that statistic, arguing page tables are about 1/512 of mapped memory, which is under 1% of most cgroups (“for all but very large cgroups the value will be in the noise of per-cpu counters”). But Shakeel mentioned an [interesting case](https://lkml.iu.edu/hypermail/linux/kernel/2011.3/06492.html) he was seeing: a user space network driver that maps application memory for zero copy with a big `pgtables` usage.\n\nThose 1/512 make sense when each page is mapped only once. But if N processes[1](https://frn.sh#fn:1)<sup>1</sup> To be fair, N should be “address spaces”, not “processes”. Every task [`task_struct`](https://elixir.bootlin.com/linux/v6.18.6/source/include/linux/sched.h#L962) has a pointer to an `mm_struct`. map the same page, each process needs its own PTEs for it. So the page tables cost about N/512 of memory they map. In general terms, this means that if you are running 512 processes with identical PTEs, the page tables cost as much as data pages.[2](https://frn.sh#fn:2)<sup>2</sup> A data page has 4 KiB. Each PTE has 8 bytes. If 512 processes map the same page, they need 512 PTEs: 512 * 8 = 4 KiB.\n\nOther issues similar to Shakeel’s that I found while searching:\n\n**Andrea Arcangeli, 2002.** [Andrea found](https://lkml.iu.edu/hypermail/linux/kernel/0201.2/0133.html) out why users with 64 GiB x86 machines were running out of memory even though there was a bunch of free memory in the system. Their workloads had hundreds of processes each mapping the same 1 GiB of shared memory. Data pages were shared, and each process needed its own page tables to map them - about 2 MiB with PAE[3](https://frn.sh#fn:3)<sup>3</sup> x86-32 feature that widens physical addresses to 36 bits allowing 64 GiB of RAM.. page tables live in the “lowmem” area, in 32-bit x86 that corresponds to ~896 MiB of physical memory. The patch moved the page tables from lowmem to highmem. This was an edge case scenario that only happened because users were using a 32-bit machine - with all the constraints of it, including lowmem - with a feature that allowed them to widen physical addresses. Pretty nice.\n\n**Khalid Aziz, 2022.** 20 years later, [Khalid reported](https://lkml.rescloud.iu.edu/hypermail/linux/kernel/2201.2/03246.html) a similar problem: an Oracle database server with 512 GB of RAM crashing with OOM when 1500+ clients attached to a 300 GB SGA (shared memory). As in Andrea’s case, data pages were shared, but each process needed its own page tables to map them. With 8-byte PTEs, 2 thousand processes mapping the same 4 KiB page need 16 KB of PTEs for that page. In the worst case, where every process maps the whole SGA, PTEs alone would take 878 GB. He proposed `mshare` to let processes share page tables. [As far as I checked](https://lore.kernel.org/all/?q=%22Add+support+for+shared+PTEs+across+processes%22), no version of mshare has been merged yet.\n\n**Qi Zheng, 2021.** [Qi reported](https://lkml.rescloud.iu.edu/2109.0/01549.html) a process with 590 GiB RSS and 110 GiB of page tables (!), when mapping that much memory should only need around 1.2 GiB of PTEs. That workload used jemalloc and tcmalloc, which return memory to the kernel with `madvise(MADV_DONTNEED)` instead of `munmap()`[4](https://frn.sh#fn:4)<sup>4</sup> `munmap()` at some point calls [`free_pgtables`](https://elixir.bootlin.com/linux/v6.8/source/mm/memory.c#L362), which ends up calling [`free_pgd_range`](https://elixir.bootlin.com/linux/v6.8/source/mm/memory.c#L300).. `MADV_DONTNEED` frees data pages and clears the PTEs, but keeps the page tables allocated, so empty page tables piled up. Unlike Andrea’s and Khalid’s cases, there were no shared mappings: it was a single process which kept page tables for freed memory. This patch got several rewrites, and was [merged in 2025](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=6375e95f381e3dc85065b6f74263a61522736203).\n\nWhile these might read like edge cases from kernel mailing lists, the pattern is common: many processes mapping a large shared memory segment, each with its own page tables for it. I found two articles related to databases.\n\n**Percona, 2021.** [Jobin Augustine reproduced](https://www.percona.com/blog/why-linux-hugepages-are-super-important-for-database-servers-a-case-with-postgresql/) a failure he kept seeing in production incidents: Postgres getting hit by OOM kills. 192 GB machine, shared buffers at 138 GiB, only 80 connections. page tables grew from 45 MiB to over 25 GiB as backends touched more of the cache. Memory got pressured and the machine started swapping. He solved the problem with huge pages: page tables stayed at 61 MiB for the same workload.\n\n**ClickHouse, 2026.** [Kaushik Iska measured](https://clickhouse.com/blog/huge-pages-clickhouse-managed-postgres) the growth directly. A machine with 128 GiB and shared buffers at 32 GiB. Each backend scanning a 15.6 GiB cached table added 31 MiB of page tables. At 200 connections, page tables took 6.1 GiB. With huge pages, 200 connections only cost 111 MiB.\n\nHuge pages aren’t free though. Reserving them needs unfragmented memory: ClickHouse notes that on a machine under load, “the same request routinely fails”, and THP can stall on allocation. [In 2016, Mel Gorman](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=444eb2a449ef36fe115431ed7b71467c4563c7f1) proposed the kernel stop defragmenting for THP (Transparent huge pages) as default, “it’s been years and it’s time to throw in the towel”. [Nelson Elhage](https://blog.nelhage.com/post/transparent-hugepages/) wrote about possible memory leaks and CPU usage issues that THP could cause.\n\n**The problem with page tables in NUMA machines**\n\nNew problems arrived with NUMA machines related to latency and memory size. With multiple nodes, a thread running on another node has to walk page tables in remote memory on TLB misses. [Mitosis](https://arxiv.org/pdf/1910.05398) (ASPLOS 2020) showed that remote page tables can slow down an application as much as remote data. And the [Hydra paper](https://www.usenix.org/system/files/atc24-gao-bin-scalable.pdf) (USENIX ATC 2024) reproduced this on a 8-socket machine with 8 TB of RAM.\n\nMitosis’s choice was to replicate the whole page table tree on every node - which costs memory (page table size * number of nodes), and every change has to update every copy. With Hydra, a PTE is only copied to a node when one of its threads faults on that node, and each page table page keeps a list of the nodes that hold a copy.\n\nIn machines with multiple nodes, you either keep one copy of the page tables and pay the cost of reading remote tables on TLB misses, or you keep a copy on each node and pay to keep them in sync.\n\n1. \nTo be fair, N should be “address spaces”, not “processes”. Every task [`task_struct`](https://elixir.bootlin.com/linux/v6.18.6/source/include/linux/sched.h#L962) has a pointer to an`mm_struct` .[↩︎](https://frn.sh#fnref:1)\n2. \nA data page has 4 KiB. Each PTE has 8 bytes. If 512 processes map the same page, they need 512 PTEs: 512 * 8 = 4 KiB. [↩︎](https://frn.sh#fnref:2)\n3. \nx86-32 feature that widens physical addresses to 36 bits allowing 64 GiB of RAM. [↩︎](https://frn.sh#fnref:3)\n4. \n`munmap()` at some point calls[`free_pgtables`](https://elixir.bootlin.com/linux/v6.8/source/mm/memory.c#L362) , which ends up calling[`free_pgd_range`](https://elixir.bootlin.com/linux/v6.8/source/mm/memory.c#L300) .[↩︎](https://frn.sh#fnref:4)\n","body_html":"<p>The other day I was reading <a href=\"https://yarchive.net/comp/powerpc_page_tables.html\" rel=\"nofollow ugc noopener\">Linus Torvalds punch some people</a> over hashed page tables. His argument: in a tree, the entries of neighboring pages are adjacent, which means one cache-line fills several TLB entries at once. A hash table, on the other hand, scatters neighbors in buckets:</p>\n<p>The fact is, you just don’t know what GOOD actually is.</p>\n<p>I’ll tell you: t a <em>good</em> TLB fill should pre-populate the TLB with all\nthe entries it can fit in one cache-line.  Do you realize that a\nbog-standard Intel CPU will fetch 8 TLB entries in one go? Together with\na self-mapping (or, as Andy points out, you can just cache the other\nlevels in dedicated caches), that means that with a single memory\nreference you get <em>eight_times</em> the coverage that the silly Power hash\ntables get.</p>\n<p>This discussion happened in 2003, in 1997, on his master’s <a href=\"https://www.cs.helsinki.fi/u/kutvonen/index_files/linus.pdf\" rel=\"nofollow ugc noopener\">thesis</a>, he explains the adoption of the three multi-level page tables on the kernel. He seemed quite interested in latency, but maybe less interested in memory size:</p>\n<p>There are also secondary concerns: the virtual machine memory mappings must be memoryefficient, so that the mapping information does not take up a lot of physical memory that could be used to better advantage for file system caching or running user programs.</p>\n<p>In 2020, Shakeel Butt found the exact case where memory efficiency is not a secondary concern at all. <a href=\"https://lore.kernel.org/linux-mm/20201126005603.1293012-1-shakeelb@google.com/#R\" rel=\"nofollow ugc noopener\">Shakeel</a> sent a patch adding a page table metric to <code>memory.stat</code>. <a href=\"https://lkml.rescloud.iu.edu/hypermail/linux/kernel/2011.3/06476.html\" rel=\"nofollow ugc noopener\">Roman Gushchin</a> asked what was the use case for that statistic, arguing page tables are about 1/512 of mapped memory, which is under 1% of most cgroups (“for all but very large cgroups the value will be in the noise of per-cpu counters”). But Shakeel mentioned an <a href=\"https://lkml.iu.edu/hypermail/linux/kernel/2011.3/06492.html\" rel=\"nofollow ugc noopener\">interesting case</a> he was seeing: a user space network driver that maps application memory for zero copy with a big <code>pgtables</code> usage.</p>\n<p>Those 1/512 make sense when each page is mapped only once. But if N processes<a href=\"https://frn.sh#fn:1\" rel=\"nofollow ugc noopener\">1</a>&lt;sup&gt;1&lt;/sup&gt; To be fair, N should be “address spaces”, not “processes”. Every task <a href=\"https://elixir.bootlin.com/linux/v6.18.6/source/include/linux/sched.h#L962\" rel=\"nofollow ugc noopener\"><code>task_struct</code></a> has a pointer to an <code>mm_struct</code>. map the same page, each process needs its own PTEs for it. So the page tables cost about N/512 of memory they map. In general terms, this means that if you are running 512 processes with identical PTEs, the page tables cost as much as data pages.<a href=\"https://frn.sh#fn:2\" rel=\"nofollow ugc noopener\">2</a>&lt;sup&gt;2&lt;/sup&gt; A data page has 4 KiB. Each PTE has 8 bytes. If 512 processes map the same page, they need 512 PTEs: 512 * 8 = 4 KiB.</p>\n<p>Other issues similar to Shakeel’s that I found while searching:</p>\n<p><strong>Andrea Arcangeli, 2002.</strong> <a href=\"https://lkml.iu.edu/hypermail/linux/kernel/0201.2/0133.html\" rel=\"nofollow ugc noopener\">Andrea found</a> out why users with 64 GiB x86 machines were running out of memory even though there was a bunch of free memory in the system. Their workloads had hundreds of processes each mapping the same 1 GiB of shared memory. Data pages were shared, and each process needed its own page tables to map them - about 2 MiB with PAE<a href=\"https://frn.sh#fn:3\" rel=\"nofollow ugc noopener\">3</a>&lt;sup&gt;3&lt;/sup&gt; x86-32 feature that widens physical addresses to 36 bits allowing 64 GiB of RAM.. page tables live in the “lowmem” area, in 32-bit x86 that corresponds to ~896 MiB of physical memory. The patch moved the page tables from lowmem to highmem. This was an edge case scenario that only happened because users were using a 32-bit machine - with all the constraints of it, including lowmem - with a feature that allowed them to widen physical addresses. Pretty nice.</p>\n<p><strong>Khalid Aziz, 2022.</strong> 20 years later, <a href=\"https://lkml.rescloud.iu.edu/hypermail/linux/kernel/2201.2/03246.html\" rel=\"nofollow ugc noopener\">Khalid reported</a> a similar problem: an Oracle database server with 512 GB of RAM crashing with OOM when 1500+ clients attached to a 300 GB SGA (shared memory). As in Andrea’s case, data pages were shared, but each process needed its own page tables to map them. With 8-byte PTEs, 2 thousand processes mapping the same 4 KiB page need 16 KB of PTEs for that page. In the worst case, where every process maps the whole SGA, PTEs alone would take 878 GB. He proposed <code>mshare</code> to let processes share page tables. <a href=\"https://lore.kernel.org/all/?q=%22Add+support+for+shared+PTEs+across+processes%22\" rel=\"nofollow ugc noopener\">As far as I checked</a>, no version of mshare has been merged yet.</p>\n<p><strong>Qi Zheng, 2021.</strong> <a href=\"https://lkml.rescloud.iu.edu/2109.0/01549.html\" rel=\"nofollow ugc noopener\">Qi reported</a> a process with 590 GiB RSS and 110 GiB of page tables (!), when mapping that much memory should only need around 1.2 GiB of PTEs. That workload used jemalloc and tcmalloc, which return memory to the kernel with <code>madvise(MADV_DONTNEED)</code> instead of <code>munmap()</code><a href=\"https://frn.sh#fn:4\" rel=\"nofollow ugc noopener\">4</a>&lt;sup&gt;4&lt;/sup&gt; <code>munmap()</code> at some point calls <a href=\"https://elixir.bootlin.com/linux/v6.8/source/mm/memory.c#L362\" rel=\"nofollow ugc noopener\"><code>free_pgtables</code></a>, which ends up calling <a href=\"https://elixir.bootlin.com/linux/v6.8/source/mm/memory.c#L300\" rel=\"nofollow ugc noopener\"><code>free_pgd_range</code></a>.. <code>MADV_DONTNEED</code> frees data pages and clears the PTEs, but keeps the page tables allocated, so empty page tables piled up. Unlike Andrea’s and Khalid’s cases, there were no shared mappings: it was a single process which kept page tables for freed memory. This patch got several rewrites, and was <a href=\"https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=6375e95f381e3dc85065b6f74263a61522736203\" rel=\"nofollow ugc noopener\">merged in 2025</a>.</p>\n<p>While these might read like edge cases from kernel mailing lists, the pattern is common: many processes mapping a large shared memory segment, each with its own page tables for it. I found two articles related to databases.</p>\n<p><strong>Percona, 2021.</strong> <a href=\"https://www.percona.com/blog/why-linux-hugepages-are-super-important-for-database-servers-a-case-with-postgresql/\" rel=\"nofollow ugc noopener\">Jobin Augustine reproduced</a> a failure he kept seeing in production incidents: Postgres getting hit by OOM kills. 192 GB machine, shared buffers at 138 GiB, only 80 connections. page tables grew from 45 MiB to over 25 GiB as backends touched more of the cache. Memory got pressured and the machine started swapping. He solved the problem with huge pages: page tables stayed at 61 MiB for the same workload.</p>\n<p><strong>ClickHouse, 2026.</strong> <a href=\"https://clickhouse.com/blog/huge-pages-clickhouse-managed-postgres\" rel=\"nofollow ugc noopener\">Kaushik Iska measured</a> the growth directly. A machine with 128 GiB and shared buffers at 32 GiB. Each backend scanning a 15.6 GiB cached table added 31 MiB of page tables. At 200 connections, page tables took 6.1 GiB. With huge pages, 200 connections only cost 111 MiB.</p>\n<p>Huge pages aren’t free though. Reserving them needs unfragmented memory: ClickHouse notes that on a machine under load, “the same request routinely fails”, and THP can stall on allocation. <a href=\"https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=444eb2a449ef36fe115431ed7b71467c4563c7f1\" rel=\"nofollow ugc noopener\">In 2016, Mel Gorman</a> proposed the kernel stop defragmenting for THP (Transparent huge pages) as default, “it’s been years and it’s time to throw in the towel”. <a href=\"https://blog.nelhage.com/post/transparent-hugepages/\" rel=\"nofollow ugc noopener\">Nelson Elhage</a> wrote about possible memory leaks and CPU usage issues that THP could cause.</p>\n<p><strong>The problem with page tables in NUMA machines</strong></p>\n<p>New problems arrived with NUMA machines related to latency and memory size. With multiple nodes, a thread running on another node has to walk page tables in remote memory on TLB misses. <a href=\"https://arxiv.org/pdf/1910.05398\" rel=\"nofollow ugc noopener\">Mitosis</a> (ASPLOS 2020) showed that remote page tables can slow down an application as much as remote data. And the <a href=\"https://www.usenix.org/system/files/atc24-gao-bin-scalable.pdf\" rel=\"nofollow ugc noopener\">Hydra paper</a> (USENIX ATC 2024) reproduced this on a 8-socket machine with 8 TB of RAM.</p>\n<p>Mitosis’s choice was to replicate the whole page table tree on every node - which costs memory (page table size * number of nodes), and every change has to update every copy. With Hydra, a PTE is only copied to a node when one of its threads faults on that node, and each page table page keeps a list of the nodes that hold a copy.</p>\n<p>In machines with multiple nodes, you either keep one copy of the page tables and pay the cost of reading remote tables on TLB misses, or you keep a copy on each node and pay to keep them in sync.</p>\n<ol><li></li></ol>\n<p>To be fair, N should be “address spaces”, not “processes”. Every task <a href=\"https://elixir.bootlin.com/linux/v6.18.6/source/include/linux/sched.h#L962\" rel=\"nofollow ugc noopener\"><code>task_struct</code></a> has a pointer to an<code>mm_struct</code> .<a href=\"https://frn.sh#fnref:1\" rel=\"nofollow ugc noopener\">↩︎</a></p>\n<ol start=\"2\"><li></li></ol>\n<p>A data page has 4 KiB. Each PTE has 8 bytes. If 512 processes map the same page, they need 512 PTEs: 512 * 8 = 4 KiB. <a href=\"https://frn.sh#fnref:2\" rel=\"nofollow ugc noopener\">↩︎</a></p>\n<ol start=\"3\"><li></li></ol>\n<p>x86-32 feature that widens physical addresses to 36 bits allowing 64 GiB of RAM. <a href=\"https://frn.sh#fnref:3\" rel=\"nofollow ugc noopener\">↩︎</a></p>\n<ol start=\"4\"><li></li></ol>\n<p><code>munmap()</code> at some point calls<a href=\"https://elixir.bootlin.com/linux/v6.8/source/mm/memory.c#L362\" rel=\"nofollow ugc noopener\"><code>free_pgtables</code></a> , which ends up calling<a href=\"https://elixir.bootlin.com/linux/v6.8/source/mm/memory.c#L300\" rel=\"nofollow ugc noopener\"><code>free_pgd_range</code></a> .<a href=\"https://frn.sh#fnref:4\" rel=\"nofollow ugc noopener\">↩︎</a></p>","headings":[]}}