{"article":{"slug":"the-forgetful-cpu-linux-on-m4","title":"The forgetful CPU (Linux on M4)","subtitle":null,"summary":"A detailed write-up of first-booting Linux on an M4 Mac mini, covering CPU quirks, bring-up details, and the hardware surprises along the way.","content_type":"blog_post","language":"en","canonical_url":"https://yuka.dev/blog-2026-10-02-linux-m4.html","author":{"name":"Yureka Lilian","url":"https://yuka.dev/","person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"Yureka Lilian","url":"https://yuka.dev/","listing_slug":null,"listing":null},"topics":[{"name":"Linux","slug":"linux","url":"https://listedarticles.com/topics/linux"},{"name":"Hardware","slug":"hardware","url":"https://listedarticles.com/topics/hardware"},{"name":"Systems Programming","slug":"systems-programming","url":"https://listedarticles.com/topics/systems-programming"},{"name":"Engineering","slug":"engineering","url":"https://listedarticles.com/topics/engineering"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":1650,"reading_minutes":7,"published_at":"2026-10-02T00:00:00.000Z","added_at":"2026-10-02T18:14:46.029Z","updated_at":"2026-10-02T18:14:46.029Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":true},"profile_url":"https://listedarticles.com/articles/the-forgetful-cpu-linux-on-m4","markdown_url":"https://listedarticles.com/articles/the-forgetful-cpu-linux-on-m4.md","example":false,"citation":"Yureka Lilian, Yureka Lilian. \"The forgetful CPU (Linux on M4).\" 2 Oct 2026. https://yuka.dev/blog-2026-10-02-linux-m4.html (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://yuka.dev/blog-2026-10-02-linux-m4.html"},"body_markdown":"# The forgetful CPU (Linux on M4)\n\n*This blog post goes into quite some detail about how I first\nbooted Linux on my M4 Mac mini. I encourage you to look up terms and\nconcepts you don’t know, since I cannot explain all the background in\nthis post ;)*\n\nBefore saying anything further, I need to thank the entire Asahi Linux team for all their prior work and help during the journey. Please consider donating to the Asahi Open Collective if you want to see more mainline Linux on Apple Silicon work!\n\n### The Beginnings\n\nIn November 2024, I bought an M4 Mac mini, gambling that it would be similar to the M1-M3 Apple Silicon machines and could be quickly supported in Asahi Linux. While the M4 was sitting on my desk for several months, more details started surfacing about this SoC.\n\nIt turned out to be more difficult, since the M4 machines are the first\ngeneration of Apple Silicon to mandate SPTM (Secure Page Table Monitor),\nwhich provides hardening against vulnerabilities in the XNU kernel of\nmacOS. In previous generations, the Linux bringup was largely based on\nMMIO traces captured using the m1n1 hypervisor, allowing the analysis of\ninteractions between the original macOS drivers and the hardware. With\nSPTM, major changes to m1n1 are required to get macOS running under the\nhypervisor, and these changes certainly go beyond what I could come up\nwith as a newbie in this space.\n\nStill, it wasn’t entirely a lost cause. In parallel to the hypervisor\nwork, I started my first attempts to boot Linux on the M4. This meant\ndisabling strict boot security, installing m1n1 as a custom boot object\nvia macOS recovery<sup>1</sup>, and obtaining a serial console<sup>2</sup> to inspect the logs.\n\n### Locked registers\n\nInitially, m1n1 was only able to start in BRINGUP mode and\nimmediately crashed while attempting to initialize GXF<sup>3</sup>\notherwise. GXF functionality turned out to be disabled/locked in raw\nboot mode on these M4+ SoCs, so making its initialization conditional\nand skipping it on these machines was the correct thing to do. Besides\nthe disabled GXF feature, there was also the RVBAR (Reset Vector Base\nAddress Register): a place in memory (one per CPU core) that determines\nwhere the core starts executing when powered on. The m1n1 code writes\nits entrypoint address to the RVBAR for each core when booting the\nkernel or chainloading another m1n1. Writing to it also led to a crash\non M4, but long story short, it already contained the correct value so\nthis write needed to be skipped as well.\n\n`2024-12-01 17:37 <yuka> can confirm this state gives a working usb proxy, as in a \"Generic m1n1 uartproxy v1.4.17-61-ga24ff77\" turns up in lsusb, and I can open a shell :)` ### `debug_putc`\n\nAt this point, a long time went by without me touching the Mac mini. At\nthe Chaos Communication Congress at the end of 2025, I found some new\nmotivation. Finally, I hacked together a very minimal device tree\ncontaining only the CPU cores and AIC interrupt controller and loaded\nthe Linux kernel (using m1n1’s `linux.py`) with the\n`earlycon` parameter, but I did not see any output after\n“Vectoring to next stage”. Since I wasn’t getting any useful output from\nthe kernel, I could only take wild guesses at what was going wrong,\nright? I decided to try the brute-force method: good old\n`println`-debugging. I took the `debug_putc`\nassembly routine from m1n1 and adjusted it to print a single ‘a’\ncharacter <sup>4</sup>. I inserted this into the Linux\nkernel code very early in boot, and sure enough, I got an ‘a’ after\n“Vectoring to next stage”! Essentially, I bisected the Linux boot code\nand landed at the MMU init code (which is still *very* early,\nstill in the assembly code in `arch/arm64/kernel/head.S`).\nWas the MMU initialization crashing the CPU somehow? Not quite: the UART\nis accessed using memory-mapped I/O. Once the MMU is enabled, all memory\naccesses are directed to virtual addresses, which are mapped to a\ncorresponding physical addresses through the page-tables. While m1n1\ncreates mappings to expose the MMIO address space at identical virtual\naddresses, Linux does not do this, meaning we end up accessing unmapped\nspace instead of the UART once the MMU is enabled. I modified the\ninitial pagetables to add this 1:1 mapping for the MMIO space, and now\nmy `debug_putc` worked much further into the boot process.\nBisecting again, the print now worked up to somewhere in the interrupt\ncontroller initialization! I narrowed it down to a write to the\nimplementation-specific CPU register\n`SYS_IMP_APL_VM_TMR_FIQ_ENA_EL2`, which triggered the new\ncrash. After I commented out this write, the kernel booted to a shell.\nLet’s goooo! `SYS_IMP_APL_VM_TMR_FIQ_ENA_EL2`, which is\nrelated to virtualization, has since been unlocked in new iBoot\nversions, so commenting out the write is no longer necessary.\n\nNow that I knew that Linux code was actually being executed, I looked\nonce again into why I hadn’t gotten any `printk` output on\nthe serial console earlier despite the `earlycon` boot\nparameter.\n\n```\n2026-01-23 12:35 <yuka> added earlycon=s5l,0x3ad200000 to my bootargs\n2026-01-23 12:36 <yuka> and now I'm actually getting useful output before the crash is happening! \n```\nSure enough, the device tree was just missing\n`stdout-path = \"serial0\"`, and after adding it, I got full\nregister dumps and stack traces from Linux on early crashes.\n\n### Secondary cores and WFI\n\nFor now, m1n1 didn’t start the secondary cores because\n`smp_start_offset` was missing (an offset hardcoded in m1n1;\nwithout it, `smp_init` is skipped). I tried the offset used\nfor base M1 - M3 and was able to start the secondary cores. Once I\nattempted loading Linux again, I arrived at another mysterious\ncrash.\n\nPrevious Apple Silicon CPUs already had known quirks regarding the\nWFI instruction. Depending on the state of the *chicken bit* (a\nbit in a register, that allows the vendor to disable some CPU\noptimization or feature, or to *chicken out*) called\n`ARM64_REG_CYC_OVRD_ok2pwrdn_force_mask` in the XNU OSS code,\nthe WFI instruction causes the CPU registers x0-x31 to be zeroed on\nthese previous generations. XNU saves these registers to the stack\nbefore the WFI and restores them afterwards. In M1-M3 SoCs, m1n1\ndisables this behavior so the CPU behaves like other arm64 processors.\nThe Asahi kernel later specifically re-enables the behavior to allow the\nCPU cores to reach deeper sleep states and save power. This is also\nrequired to allow one core in the cluster to boost to higher clock\nspeeds when all other cores are in this deeper WFI sleep state.\n\nIt seems this chicken bit is either locked or has been removed on M4,\nand the default behavior does not comply with the ARM64 specification\n(specifically: “If the system is configured such that the WFI\ninstruction can be completed, then the WFI instruction must not cause a\nloss of architectural state.”<sup>5</sup>). In April 2026, I\nmanaged to boot Linux on the M4 with all cores enabled by replacing all\nWFI and WFIT (Wait For Interrupt with Timeout) instructions in my kernel\nwith NOP (no-op).\n\nWith this, a long journey towards an upstream solution for the WFI\nproblem began. At first, I looked into how errata (misbehaviors of the\nsilicon) are usually handled. There is a whole framework in Linux to\nallow it to patch itself in memory during early boot. However, this\napproach was discarded because it is very difficult to accurately detect\ncases in which WFI should be NOP’ed. Namely, a virtual machine running\nunder the macOS hypervisor would also trigger the erratum logic, but WFI\nis trapped by the macOS hypervisor and used to schedule different guests\nefficiently. Detecting virtualization (especially when nested\nvirtualization is enabled) is also complicated, so Will Deacon suggested\nan alternative solution: The kernel should gain support for disabling\nWFI idle using a new bootarg<sup>6</sup>, and then m1n1 can add\nthe appropriate bootargs conditionally when booting on bare-metal\nmachines known to have broken WFI. We will then add a mechanism to allow\nLinux to put the cores into sleep states. For now, the downstream\ncpuidle-apple driver can be used for this, but Sven’s PSCI EFI conduit\nwork is very promising for a future upstream solution.\n\nI’m happy to report that the mechanism for preventing Linux from\ncrashing on WFI and WFIT instructions has been merged into mainline\nLinux<sup>7</sup> <sup>8</sup> and m1n1<sup>9</sup>, so\nthe latest releases of these can boot natively with secondary cores on\nM4 Macs!\n\n### Next steps\n\nThis work has so far uncovered and addressed a fundamental issue with running Linux on the M4 and later Apple Silicon chips, allowing Linux to boot to a shell with all cores usable (the same WFI workarounds have been found to work on M4 Pro, M4 Max and M5 chips!). Reverse engineering the peripherals, on which I will not go into details here, is going at a slow but steady pace. Sven’s tireless work on making the m1n1 hypervisor able to boot and trace macOS on these devices is going to be very useful for tackling the more complex components such as the builtin camera, display controller, and GPU initialization. For the most part, I’m submitting my work directly to the respective upstream projects, allowing everyone to take advantage of it. Sometimes it would be nice if certain projects getting a lot of funding were more transparent about how they’re benefitting from the upstream projects’ progress.\n\n**If you want to support my work on Linux on the M4, I take\ndonations over at LiberaPay\nor GitHub Sponsors.\nPlease also consider donating to the Asahi Open Collective\nfor more Linux on Apple Silicon mainlining.**\n\nAs always, I hope you could learn something :)","body_html":"<h1 id=\"the-forgetful-cpu-linux-on-m4\">The forgetful CPU (Linux on M4)</h1>\n<p>*This blog post goes into quite some detail about how I first\nbooted Linux on my M4 Mac mini. I encourage you to look up terms and\nconcepts you don’t know, since I cannot explain all the background in\nthis post ;)*</p>\n<p>Before saying anything further, I need to thank the entire Asahi Linux team for all their prior work and help during the journey. Please consider donating to the Asahi Open Collective if you want to see more mainline Linux on Apple Silicon work!</p>\n<h3 id=\"the-beginnings\">The Beginnings</h3>\n<p>In November 2024, I bought an M4 Mac mini, gambling that it would be similar to the M1-M3 Apple Silicon machines and could be quickly supported in Asahi Linux. While the M4 was sitting on my desk for several months, more details started surfacing about this SoC.</p>\n<p>It turned out to be more difficult, since the M4 machines are the first\ngeneration of Apple Silicon to mandate SPTM (Secure Page Table Monitor),\nwhich provides hardening against vulnerabilities in the XNU kernel of\nmacOS. In previous generations, the Linux bringup was largely based on\nMMIO traces captured using the m1n1 hypervisor, allowing the analysis of\ninteractions between the original macOS drivers and the hardware. With\nSPTM, major changes to m1n1 are required to get macOS running under the\nhypervisor, and these changes certainly go beyond what I could come up\nwith as a newbie in this space.</p>\n<p>Still, it wasn’t entirely a lost cause. In parallel to the hypervisor\nwork, I started my first attempts to boot Linux on the M4. This meant\ndisabling strict boot security, installing m1n1 as a custom boot object\nvia macOS recovery&lt;sup&gt;1&lt;/sup&gt;, and obtaining a serial console&lt;sup&gt;2&lt;/sup&gt; to inspect the logs.</p>\n<h3 id=\"locked-registers\">Locked registers</h3>\n<p>Initially, m1n1 was only able to start in BRINGUP mode and\nimmediately crashed while attempting to initialize GXF&lt;sup&gt;3&lt;/sup&gt;\notherwise. GXF functionality turned out to be disabled/locked in raw\nboot mode on these M4+ SoCs, so making its initialization conditional\nand skipping it on these machines was the correct thing to do. Besides\nthe disabled GXF feature, there was also the RVBAR (Reset Vector Base\nAddress Register): a place in memory (one per CPU core) that determines\nwhere the core starts executing when powered on. The m1n1 code writes\nits entrypoint address to the RVBAR for each core when booting the\nkernel or chainloading another m1n1. Writing to it also led to a crash\non M4, but long story short, it already contained the correct value so\nthis write needed to be skipped as well.</p>\n<p><code>2024-12-01 17:37 &lt;yuka&gt; can confirm this state gives a working usb proxy, as in a &quot;Generic m1n1 uartproxy v1.4.17-61-ga24ff77&quot; turns up in lsusb, and I can open a shell :)</code> ### <code>debug_putc</code></p>\n<p>At this point, a long time went by without me touching the Mac mini. At\nthe Chaos Communication Congress at the end of 2025, I found some new\nmotivation. Finally, I hacked together a very minimal device tree\ncontaining only the CPU cores and AIC interrupt controller and loaded\nthe Linux kernel (using m1n1’s <code>linux.py</code>) with the\n<code>earlycon</code> parameter, but I did not see any output after\n“Vectoring to next stage”. Since I wasn’t getting any useful output from\nthe kernel, I could only take wild guesses at what was going wrong,\nright? I decided to try the brute-force method: good old\n<code>println</code>-debugging. I took the <code>debug_putc</code>\nassembly routine from m1n1 and adjusted it to print a single ‘a’\ncharacter &lt;sup&gt;4&lt;/sup&gt;. I inserted this into the Linux\nkernel code very early in boot, and sure enough, I got an ‘a’ after\n“Vectoring to next stage”! Essentially, I bisected the Linux boot code\nand landed at the MMU init code (which is still <em>very</em> early,\nstill in the assembly code in <code>arch/arm64/kernel/head.S</code>).\nWas the MMU initialization crashing the CPU somehow? Not quite: the UART\nis accessed using memory-mapped I/O. Once the MMU is enabled, all memory\naccesses are directed to virtual addresses, which are mapped to a\ncorresponding physical addresses through the page-tables. While m1n1\ncreates mappings to expose the MMIO address space at identical virtual\naddresses, Linux does not do this, meaning we end up accessing unmapped\nspace instead of the UART once the MMU is enabled. I modified the\ninitial pagetables to add this 1:1 mapping for the MMIO space, and now\nmy <code>debug_putc</code> worked much further into the boot process.\nBisecting again, the print now worked up to somewhere in the interrupt\ncontroller initialization! I narrowed it down to a write to the\nimplementation-specific CPU register\n<code>SYS_IMP_APL_VM_TMR_FIQ_ENA_EL2</code>, which triggered the new\ncrash. After I commented out this write, the kernel booted to a shell.\nLet’s goooo! <code>SYS_IMP_APL_VM_TMR_FIQ_ENA_EL2</code>, which is\nrelated to virtualization, has since been unlocked in new iBoot\nversions, so commenting out the write is no longer necessary.</p>\n<p>Now that I knew that Linux code was actually being executed, I looked\nonce again into why I hadn’t gotten any <code>printk</code> output on\nthe serial console earlier despite the <code>earlycon</code> boot\nparameter.</p>\n<pre><code>2026-01-23 12:35 &lt;yuka&gt; added earlycon=s5l,0x3ad200000 to my bootargs\n2026-01-23 12:36 &lt;yuka&gt; and now I&#39;m actually getting useful output before the crash is happening! </code></pre>\n<p>Sure enough, the device tree was just missing\n<code>stdout-path = &quot;serial0&quot;</code>, and after adding it, I got full\nregister dumps and stack traces from Linux on early crashes.</p>\n<h3 id=\"secondary-cores-and-wfi\">Secondary cores and WFI</h3>\n<p>For now, m1n1 didn’t start the secondary cores because\n<code>smp_start_offset</code> was missing (an offset hardcoded in m1n1;\nwithout it, <code>smp_init</code> is skipped). I tried the offset used\nfor base M1 - M3 and was able to start the secondary cores. Once I\nattempted loading Linux again, I arrived at another mysterious\ncrash.</p>\n<p>Previous Apple Silicon CPUs already had known quirks regarding the\nWFI instruction. Depending on the state of the <em>chicken bit</em> (a\nbit in a register, that allows the vendor to disable some CPU\noptimization or feature, or to <em>chicken out</em>) called\n<code>ARM64_REG_CYC_OVRD_ok2pwrdn_force_mask</code> in the XNU OSS code,\nthe WFI instruction causes the CPU registers x0-x31 to be zeroed on\nthese previous generations. XNU saves these registers to the stack\nbefore the WFI and restores them afterwards. In M1-M3 SoCs, m1n1\ndisables this behavior so the CPU behaves like other arm64 processors.\nThe Asahi kernel later specifically re-enables the behavior to allow the\nCPU cores to reach deeper sleep states and save power. This is also\nrequired to allow one core in the cluster to boost to higher clock\nspeeds when all other cores are in this deeper WFI sleep state.</p>\n<p>It seems this chicken bit is either locked or has been removed on M4,\nand the default behavior does not comply with the ARM64 specification\n(specifically: “If the system is configured such that the WFI\ninstruction can be completed, then the WFI instruction must not cause a\nloss of architectural state.”&lt;sup&gt;5&lt;/sup&gt;). In April 2026, I\nmanaged to boot Linux on the M4 with all cores enabled by replacing all\nWFI and WFIT (Wait For Interrupt with Timeout) instructions in my kernel\nwith NOP (no-op).</p>\n<p>With this, a long journey towards an upstream solution for the WFI\nproblem began. At first, I looked into how errata (misbehaviors of the\nsilicon) are usually handled. There is a whole framework in Linux to\nallow it to patch itself in memory during early boot. However, this\napproach was discarded because it is very difficult to accurately detect\ncases in which WFI should be NOP’ed. Namely, a virtual machine running\nunder the macOS hypervisor would also trigger the erratum logic, but WFI\nis trapped by the macOS hypervisor and used to schedule different guests\nefficiently. Detecting virtualization (especially when nested\nvirtualization is enabled) is also complicated, so Will Deacon suggested\nan alternative solution: The kernel should gain support for disabling\nWFI idle using a new bootarg&lt;sup&gt;6&lt;/sup&gt;, and then m1n1 can add\nthe appropriate bootargs conditionally when booting on bare-metal\nmachines known to have broken WFI. We will then add a mechanism to allow\nLinux to put the cores into sleep states. For now, the downstream\ncpuidle-apple driver can be used for this, but Sven’s PSCI EFI conduit\nwork is very promising for a future upstream solution.</p>\n<p>I’m happy to report that the mechanism for preventing Linux from\ncrashing on WFI and WFIT instructions has been merged into mainline\nLinux&lt;sup&gt;7&lt;/sup&gt; &lt;sup&gt;8&lt;/sup&gt; and m1n1&lt;sup&gt;9&lt;/sup&gt;, so\nthe latest releases of these can boot natively with secondary cores on\nM4 Macs!</p>\n<h3 id=\"next-steps\">Next steps</h3>\n<p>This work has so far uncovered and addressed a fundamental issue with running Linux on the M4 and later Apple Silicon chips, allowing Linux to boot to a shell with all cores usable (the same WFI workarounds have been found to work on M4 Pro, M4 Max and M5 chips!). Reverse engineering the peripherals, on which I will not go into details here, is going at a slow but steady pace. Sven’s tireless work on making the m1n1 hypervisor able to boot and trace macOS on these devices is going to be very useful for tackling the more complex components such as the builtin camera, display controller, and GPU initialization. For the most part, I’m submitting my work directly to the respective upstream projects, allowing everyone to take advantage of it. Sometimes it would be nice if certain projects getting a lot of funding were more transparent about how they’re benefitting from the upstream projects’ progress.</p>\n<p><strong>If you want to support my work on Linux on the M4, I take\ndonations over at LiberaPay\nor GitHub Sponsors.\nPlease also consider donating to the Asahi Open Collective\nfor more Linux on Apple Silicon mainlining.</strong></p>\n<p>As always, I hope you could learn something :)</p>","headings":[{"level":1,"text":"The forgetful CPU (Linux on M4)","id":"the-forgetful-cpu-linux-on-m4"},{"level":3,"text":"The Beginnings","id":"the-beginnings"},{"level":3,"text":"Locked registers","id":"locked-registers"},{"level":3,"text":"Secondary cores and WFI","id":"secondary-cores-and-wfi"},{"level":3,"text":"Next steps","id":"next-steps"}]}}