{"article":{"slug":"when-ldaxr-doesnt-work-exclusive-accesses-and-cacheability-on-aarch64","title":"When ldaxr Doesn't Work: Exclusive Accesses and Cacheability on AArch64","subtitle":null,"summary":"Joel Siks debugs why spinlocks built on ldaxr/stxr exclusive accesses hung on a real Raspberry Pi 5 while working in QEMU in his hobby AArch64 kernel floss, tracing it to memory cacheability attributes and the differences between QEMU, the Arm architecture manual and the Cortex-A76.","content_type":"blog_post","language":"en","canonical_url":"https://joelsiks.com/posts/aarch64-exclusive-access-cacheability/","author":{"name":"Joel Siks","url":null,"person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"joelsiks.com","url":"https://joelsiks.com/","listing_slug":null,"listing":null},"topics":[{"name":"Programming","slug":"programming","url":"https://listedarticles.com/topics/programming"},{"name":"Engineering","slug":"engineering","url":"https://listedarticles.com/topics/engineering"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":2658,"reading_minutes":12,"published_at":"2026-10-10T00:00:00.000Z","added_at":"2026-10-11T14:10:16.587Z","updated_at":"2026-10-11T14:10:16.587Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":true},"profile_url":"https://listedarticles.com/articles/when-ldaxr-doesnt-work-exclusive-accesses-and-cacheability-on-aarch64","markdown_url":"https://listedarticles.com/articles/when-ldaxr-doesnt-work-exclusive-accesses-and-cacheability-on-aarch64.md","example":false,"citation":"Joel Siks, joelsiks.com. \"When ldaxr Doesn't Work: Exclusive Accesses and Cacheability on AArch64.\" 10 Oct 2026. https://joelsiks.com/posts/aarch64-exclusive-access-cacheability/ (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://joelsiks.com/posts/aarch64-exclusive-access-cacheability/"},"body_markdown":"# When ldaxr Doesn't Work: Exclusive Accesses and Cacheability on AArch64\n\n# Introduction\n\nI’ve now been working on [floss](https://github.com/joelsiks/floss), my own operating system (OS) from scratch for AArch64, which is really just a kernel so far, for a couple of months since my last post about [Interrupts and the Generic Interrupt Controller on AArch64](../aarch64-interrupts). Right now, floss targets QEMU’s virt and Raspberry Pi 4 (rpi4) boards, as well as real Raspberry Pi 5 (rpi5) hardware. Testing on real hardware has been rewarding and makes the project exciting. However, as we’ll cover in much detail about in this post, real hardware reveals a new dimension of possible failures/behavior that QEMU doesn’t necessarily reveal.\n\nThis post will cover my journey of implementing a basic spin lock and figuring out why executing a load exclusive raises an exception on real hardware but not in QEMU. We will cover an exciting combination of areas, which include: spin locks, exception handling, virtual memory, memory types, memory attributes, cacheability, the differences between QEMU and real hardware, and most notably, how exclusive memory operations work.\n\n# Implementing a Spin Lock\n\nOnce I got to the point of enabling and bringing up secondary cores on the CPU, it was a good time to get some form of mutual exclusion going. As a trivial example, each core prints its own id in the setup path, e.g., “Running core 0”. Without mutual exclusion, the cores will race with each other and interleave their output, producing a garbled mess as shown in the snippet below. To get legible output, we want to guard the print with some form of mutual exclusion, so that only one core will print at a time.\n\n```\nRunning core 0\nRRRununing cournnien ng 2cni\nonrge  1co\nre 3\n```\n\nTo warm up on locks and synchronization, I’ve read and recapped Arm’s [Implementation Software Synchronization Primitives in A64](https://support.arm.com/documentation/110478/0100), as well as my favourite OS book [Operating Systems: Three Easy Pieces](https://pages.cs.wisc.edu/~remzi/OSTEP/). As a good starting point, I decided to implement a spin lock, which is extremely straightforward in its simplest form. I also decided it would be a good exercise to write the spin lock by hand in assembly, which is much more fun than writing it in C/C++.\n\nMy basic version uses the load exclusive and store exclusive instructions, which work in concert to achieve atomicity. The idea is that the `stxr` instruction will only succeed if no other writes to the same memory location were observed since the `ldx[a]r` instruction took place.\n\n```\nvoid SpinLock::lock() {\n  uint32_t lock_read;\n  uint32_t store_result;\n\n  asm volatile(\n    \"1:                    \\t\\n\\\n     ldaxr  %w0, [%2]      \\t\\n\\\n     cbnz   %w0, 1b        \\t\\n\\\n     mov    %w0, #1        \\t\\n\\\n     stxr   %w1, %w0, [%2] \\t\\n\\\n     cbnz   %w1, 1b        \\t\\n\\\n     \"\n     : \"=&r\"(lock_read), \"=&r\"(store_result) // outputs\n     : \"r\"(&_lock) // input\n     : \"memory\" // clobber\n  );\n}\n\nvoid SpinLock::unlock() {\n  asm volatile(\n    \"stlr   wzr, [%0]\"\n     : // no output\n     : \"r\"(&_lock) // input\n     : \"memory\" // clobber\n  );\n}\n```\n\nThe lock uses a 32-bit locking word (hence the use of `w` registers), where a 0 indicates unlocked and 1 indicates locked. To satisfy correctness, the load exclusive is marked with acquire semantics (the “a” in `ldaxr`), which guarantees that all memory accesses after the load are not reordered before it. Similarly, the store in the unlock path is marked with release semantics (the “l” in `stlr`), which guarantees that all memory accesses before the store are not reordered after it.\n\nLocks are usually evaluated in three categories: *correctness*, *fairness*, and *performance*. We’ll empirically validate correctness later, and we’ve already considered memory ordering. Fairness is a no-go, which is alright in this basic version, but a core might get starved and never get to take the lock if other cores manage to always get it first. Performance is not that good, since cores will spend cycles busy-waiting to get the lock.\n\nOne approach to address the performance issue is to use Arm’s `wfe` (wait for event) and `sev` (send event) pair, which puts the core in a low-power state waiting for an event, and wakes all cores up that were waiting. That is something for future work.\n\n# Validation\n\nI initially tested the spin lock in the path with the garbled output shown in #Implementing a Spin Lock, to get legible output during the bringup of multiple cores. In QEMU’s virt and rpi4 boards, everything seemed to be working fine. I reran the program a couple of thousand times for good measure and never observed garbled output.\n\n```\n// SpinLockGuard is a RAII object that locks during\n// construction and unlocks during destruction\n{\n  SpinLockGuard guard(&_print_lock);\n  kprintf(\"Running core %d\\n\", cpu_id);\n}\n```\n\nHowever, when I ran the same code on my rpi5 I got a mysterious exception. The output below is from my custom exception handler, where we see a Data Abort exception with the Exception Class (EC) value of 37 (0x25). According to Arm’s [“Types of exception”](https://support.arm.com/documentation/den0042/0100/Exceptions-and-Interrupts/Types-of-exception): *“A data abort exception happens as a result of a load or store instruction”*.\n\n```\nGot an exception (from 0x84474)\n- Syndrome is 0x96000410\n- EC is 37 (Data Abort)\nRegister dump:\n  x0:  0x2bcf88\n  ...\n  x30: 0x85554\n```\n\nTaking a closer look at the instruction the CPU was executing when the exception occurred (0x84474), we see that it is right on the `ldaxr` instruction. This aligns with the Data Abort value we also got, so something must have gone awry with the `ldaxr` memory access. As a sanity check, replacing the load exclusive and store exclusive with a “normal” load acquire (`ldar`) and a plain store (`str`) makes the problem go away and no exception is thrown. This tells me there is something blocking the exclusive operations from working.\n\n```\n0000000000084470 <_ZN13SpinLockGuardC1EP8SpinLock>:\n   84470:       f9000001        str     x1, [x0]\n   84474:       885ffc20        ldaxr   w0, [x1] <-- The CPU was here\n   84478:       35ffffe0        cbnz    w0, 84474\n   8447c:       52800020        mov     w0, #0x1\n   84480:       88027c20        stxr    w2, w0, [x1]\n   84484:       35ffff82        cbnz    w2, 84474\n   84488:       d65f03c0        ret\n```\n\n# Figuring Out Why\n\nFortunately, since I had read [Implementation Software Synchronization Primitives in A64](https://support.arm.com/documentation/110478/0100), I remember there was a specific section on [memory attributes](https://support.arm.com/documentation/110478/0100/AArch64-Atomic-Instructions/The-memory-attribute-of-address-which-is-atomic-accessed). This section highlights the memory attributes for which it is architecturally guaranteed that atomic instructions are atomic. These are:\n\n* Inner Shareable, Inner Write-Back, Outer Write-Back Normal memory with Read allocation hints and Write allocation hints and not transient\n* Outer Shareable, Inner Write-Back, Outer Write-Back Normal memory with Read allocation hints and Write allocation hints and not transient\n\nThe page also lists memory types which might not support atomic instructions. The most notable bullet point of those is:\n\n* Device, Non-cacheable memory, or memory that is treated as Non-cacheable, in an implementation that does support hardware cache coherency\n\n## Memory Types\n\nThe default memory type for all memory is Device memory. However, I set up and map virtual memory particularly early in floss, associating all regions of RAM with the “Normal” memory type, so that unaligned accesses can be performed (e.g., an 8-byte load from an address that is not 8-byte aligned). This is because Device memory requires alignment checking. I became aware of this when upgrading from QEMU v8.2.2 to v11.1.50, which included a (not so recent) [patch](https://gitlab.com/qemu-project/qemu/-/commit/59754f85ed35cbd5f4bf2663ca2136c78d5b2413) to add alignment checking to device memory, which raised a rather infuriating exception that took a while to debug as well.\n\nNormal memory is a pretty fuzzy term, but is generally all types of memory that are not Device. Normal memory can generally be cached, speculative, and reordered depending on attributes. Device memory that is accessed via memory mapped IO (MMIO) can generally not be speculative or cached like RAM.\n\nAll types of memory are defined by a set of attributes, which can be configured for individual virtual memory mappings.\n\n## Memory Attributes\n\nMemory attributes are configured via the Memory Attribute Indirection Register (`MAIR_EL1` for EL1). This is a 64-bit register that is divided into 8 different sets of attributes of 8 bits each. An entry in the translation table (page table in Linux terminology) that points to a block of physical memory (i.e., a [Block or Page descriptor](https://support.arm.com/documentation/102376/0200/Describing-memory-in-AArch64)) has three bits that reference an index in the MAIR register, indicating the attributes for that specific memory mapping. Additionally, such an entry has bits that configure execute permissions, access permission, shareability, and more.\n\nRelevant to this post, and specifically to exclusive memory operations, is the shareability configuration, which can be set to one of: *Non-Shareable*, *Inner-Shareable*, or *Outer-Shareable*. Loosely, shareability determines in which domain, and for which observers, memory operations should be observed to complete in a finite amount of time and without using explicit cache maintenance. This topic is pretty complex, and I refer you to the [Arm ARM, section B 2.7.1](https://support.arm.com/architectures/a-profile-architecture) for more information. In floss, the Normal memory type is configured to be Inner-Shareable, which is what Arm [expects](https://support.arm.com/documentation/102376/0200/Cacheability-and-shareability-attributes) operating systems to do.\n\nCurrently, floss only configures two sets of memory attributes, one for Device memory and one for Normal memory. Let’s ignore the specifics for Device and focus on Normal memory.\n\n* **Device**: `0b00000000`, Device-nGnRnE memory\n* **Normal**: `0b11111111`, Normal memory, Outer Write-Back, Inner Write-Back, with Read allocation hints and Write allocation hints and Non-transient\n\nThe Write-Back property means that writes should be written to cache, and flushed to memory only when the cache-line is evicted. In summary, this means that the memory attributes I’ve configured for Normal memory has *full* support for caching. It also fully matches the architectural requirements for atomics to be atomic.\n\nAn alternative to Write-Back is Write-Through, which writes to both cache and memory immediately, without waiting for the cache-line to be evicted.\n\n## Global and Local Cacheability Scopes\n\nAs we’ve concluded, the memory attributes for the Normal memory type do have the cacheable Inner and Outer Write-Back property. Even so, with all the details available to us so far, the only logical explanation for why an exception is raised when executing the load exclusive is cacheability.\n\nWith the help of my LLM, I was able to figure out that there is a global setting in the [System Control Register (`SCTLR`)](https://support.arm.com/documentation/100442/0100/register-descriptions/aarch32-system-registers/sctlr--system-control-register): the “C” bit, which controls cacheability for data accesses at EL1 and EL0 (floss runs at EL1). Querying this bit, I see that it’s set to its default value of 0 on all cores, which disables the data cache for all of them. Setting this bit to 1 makes the exception go away on the rpi5. Hoorah!\n\nReading up on the documentation of [L1 memory system and cache behavior](https://support.arm.com/documentation/100798/0401/L1-memory-system/Cache-behavior/Data-cache-disabled-behavior?lang=en) on the Cortex-A76 CPU (the CPU of the rpi5), there is some revealing information:\n\n* When the data cache is disabled, instructions and operations are affected as follows:\n  + All load and store instructions to cacheable memory are treated as if they were Non-cacheable and are incoherent with the caches in both this core and other cores in the cluster. Software must take this into account.\n\n# Cacheability\n\nWe’ve figured out why the exception was raised on the rpi5, because `SCTLR.C` with a value of 0 effectively treated the memory as non-cacheable. Interestingly, if I enable the global cacheability setting for each core by setting `SCTLR.C` to 1, but change the MAIR attribute to something that does not have Write-Back enabled, the same exception is raised. I did a small test of different MAIR configurations for the memory being referenced by `ldaxr`, and this is the result when executing on the rpi5.\n\n| Bit pattern | Description | Exception raised |\n| --- | --- | --- |\n| `0b00110011` | Normal memory, Outer/Inner Write-Through Transient | Yes |\n| `0b01000100` | Normal memory, Outer/Inner Non-cacheable | Yes |\n| `0b01110111` | Normal memory, Outer/Inner Write-Back Transient | **No** |\n| `0b10111011` | Normal memory, Outer/Inner Write-Through Non-transient | Yes |\n| `0b11111111` | Normal memory, Outer/Inner Write-Back Non-transient | **No** |\n\nFrom this experiment, it appears as if the load exclusive instruction doesn’t work without Write-Back. In fact, the documentation of [Memory attributes](https://support.arm.com/documentation/100798/0401/Memory-Management-Unit/Specific-behaviors-on-aborts-and-memory-attributes/Memory-attributes?lang=en) on the Cortex-A76 CPU says that: *\"… a page is cacheable only if the Inner and Outer memory attributes are Write-Back. In all other cases, all pages are downgraded to Non-cacheable Normal memory\"*. This means that even though Write-Through technically means “write to cache”, the Cortex-A76 CPU does not in fact cache that type of memory.\n\nFrom this, it appears as if the load exclusive operation only completes without an exception if cacheability is enabled, both globally with `SCTLR.C` set to 1, and locally with a Write-Back attribute for the translation table entry.\n\nIt’s a bit frustrating that I wasn’t able to catch this in QEMU. QEMU doesn’t have support for the rpi5 board yet, and there are likely a few differences between the actual rpi5 hardware and the available rpi4 QEMU board. Nonetheless, when digging through the QEMU source code, it turns out that cacheability is simply not being modelled at all (see references in commits [fa2ef212df](https://gitlab.com/qemu-project/qemu/-/commit/fa2ef212df), [b0fe242751](https://gitlab.com/qemu-project/qemu/-/commit/b0fe242751)), reinforced by the constant for `SCTLR_C` not being referenced [anywhere at all](https://gitlab.com/search?search=SCTLR_C&nav_source=navbar&project_id=11167699&group_id=3038080&search_code=true&repository_ref=v11.1.2). So, even though it runs fine in QEMU, don’t assume everything’s perfect until you test stuff out in the real world.\n\n# Exclusive Monitors\n\nTo achieve atomicity via the load exclusive and store exclusive instructions, one or more exclusive monitors are used. Each core has its own local monitor, and all cores have access to a global monitor. Each monitor has two possible states: *open* and *exclusive*. When executing a load exclusive instruction, the monitor(s) are moved to the exclusive state, and the address of the load is recorded. When executing a store exclusive instruction, the monitor(s) are checked to see if they are all in the exclusive state and if the address matches, in which case it succeeds.\n\nMore precisely, the address stored in the exclusive monitor(s) is only the upper address bits, referred to as the Exclusive Reservation Granule (ERG). Two address locations within the same ERG will appear as the same entry in the monitor(s). On a typical system, the ERG matches the size of a cache line.\n\nExclusive accesses to memory locations marked as Non-shareable will only access the local monitor, and memory locations that are marked shareable (either Inner or Outer) will access both the local and global monitors.\n\nAgain, Arm’s [Implementation Software Synchronization Primitives in A64](https://support.arm.com/documentation/110478/0100/Exclusives/Exclusive-monitors/Global-monitor) highlights a very interesting aspect in its documentation of exclusive monitors:\n\n* Although we describe the local and global monitors as being different things, they might actually share logic. For example, one implementation approach is to use a combination of the per-core local monitors and the cache coherency logic to provide global monitor functionality for cacheable locations.\n\nThis feels like the cherry on the top of this mystery. Apparently (or empirically?), the rpi5’s Cortex-A76 CPU seems to require cacheability to be enabled for a load exclusive instruction to not generate a Data Abort exception. This could very well be because it implements the global monitor via cache coherency logic, and therefore, any exclusive memory operation to a memory location marked as shareable (either Inner or Outer) that access the global monitor, will fail when cacheability is disabled.\n\n# Conclusion\n\nLet’s recap: I wanted mutual exclusion and opted for a “hand-made” spin lock using load exclusive and store exclusive. It worked fine in QEMU, but I got an exception on rpi5 with real hardware. I then ended up scouring documentation (Arm ARM, Cortex-A76, Implementation Software Synchronization Primitives in A64) to understand what’s going on.\n\nFrom this adventure, I can conclude that on this implementation and configuration of the Raspberry Pi 5 with an Arm Cortex-A76 CPU, I observe that a load exclusive operation on shareable memory (which presumably accesses the global exclusive monitor), does not work unless cacheability is enabled, both globally (`SCTLR.C` set to 1) and locally (memory attribute for translation table entry configured with Write-Back).\n\nDuring development I iterate quickly with QEMU, but from time to time also verify on real hardware with the Raspberry Pi 5. This creates several layers, which make this a bit confusing, but very interesting at the same time. QEMU says one thing, the Arm ARM another, and then finally the implementation of the Cortex-A76 a third thing.\n\nThis is by far the most interesting bug I have encountered so far while developing floss, but likely not the last. Thank you for reading!\n","body_html":"<h1 id=\"when-ldaxr-doesn-t-work-exclusive-accesses-and-cacheability-on-a\">When ldaxr Doesn&#39;t Work: Exclusive Accesses and Cacheability on AArch64</h1>\n<h1 id=\"introduction\">Introduction</h1>\n<p>I’ve now been working on <a href=\"https://github.com/joelsiks/floss\" rel=\"nofollow ugc noopener\">floss</a>, my own operating system (OS) from scratch for AArch64, which is really just a kernel so far, for a couple of months since my last post about Interrupts and the Generic Interrupt Controller on AArch64. Right now, floss targets QEMU’s virt and Raspberry Pi 4 (rpi4) boards, as well as real Raspberry Pi 5 (rpi5) hardware. Testing on real hardware has been rewarding and makes the project exciting. However, as we’ll cover in much detail about in this post, real hardware reveals a new dimension of possible failures/behavior that QEMU doesn’t necessarily reveal.</p>\n<p>This post will cover my journey of implementing a basic spin lock and figuring out why executing a load exclusive raises an exception on real hardware but not in QEMU. We will cover an exciting combination of areas, which include: spin locks, exception handling, virtual memory, memory types, memory attributes, cacheability, the differences between QEMU and real hardware, and most notably, how exclusive memory operations work.</p>\n<h1 id=\"implementing-a-spin-lock\">Implementing a Spin Lock</h1>\n<p>Once I got to the point of enabling and bringing up secondary cores on the CPU, it was a good time to get some form of mutual exclusion going. As a trivial example, each core prints its own id in the setup path, e.g., “Running core 0”. Without mutual exclusion, the cores will race with each other and interleave their output, producing a garbled mess as shown in the snippet below. To get legible output, we want to guard the print with some form of mutual exclusion, so that only one core will print at a time.</p>\n<pre><code>Running core 0\nRRRununing cournnien ng 2cni\nonrge  1co\nre 3</code></pre>\n<p>To warm up on locks and synchronization, I’ve read and recapped Arm’s <a href=\"https://support.arm.com/documentation/110478/0100\" rel=\"nofollow ugc noopener\">Implementation Software Synchronization Primitives in A64</a>, as well as my favourite OS book <a href=\"https://pages.cs.wisc.edu/~remzi/OSTEP/\" rel=\"nofollow ugc noopener\">Operating Systems: Three Easy Pieces</a>. As a good starting point, I decided to implement a spin lock, which is extremely straightforward in its simplest form. I also decided it would be a good exercise to write the spin lock by hand in assembly, which is much more fun than writing it in C/C++.</p>\n<p>My basic version uses the load exclusive and store exclusive instructions, which work in concert to achieve atomicity. The idea is that the <code>stxr</code> instruction will only succeed if no other writes to the same memory location were observed since the <code>ldx[a]r</code> instruction took place.</p>\n<pre><code>void SpinLock::lock() {\n  uint32_t lock_read;\n  uint32_t store_result;\n\n  asm volatile(\n    &quot;1:                    \\t\\n\\\n     ldaxr  %w0, [%2]      \\t\\n\\\n     cbnz   %w0, 1b        \\t\\n\\\n     mov    %w0, #1        \\t\\n\\\n     stxr   %w1, %w0, [%2] \\t\\n\\\n     cbnz   %w1, 1b        \\t\\n\\\n     &quot;\n     : &quot;=&amp;r&quot;(lock_read), &quot;=&amp;r&quot;(store_result) // outputs\n     : &quot;r&quot;(&amp;_lock) // input\n     : &quot;memory&quot; // clobber\n  );\n}\n\nvoid SpinLock::unlock() {\n  asm volatile(\n    &quot;stlr   wzr, [%0]&quot;\n     : // no output\n     : &quot;r&quot;(&amp;_lock) // input\n     : &quot;memory&quot; // clobber\n  );\n}</code></pre>\n<p>The lock uses a 32-bit locking word (hence the use of <code>w</code> registers), where a 0 indicates unlocked and 1 indicates locked. To satisfy correctness, the load exclusive is marked with acquire semantics (the “a” in <code>ldaxr</code>), which guarantees that all memory accesses after the load are not reordered before it. Similarly, the store in the unlock path is marked with release semantics (the “l” in <code>stlr</code>), which guarantees that all memory accesses before the store are not reordered after it.</p>\n<p>Locks are usually evaluated in three categories: <em>correctness</em>, <em>fairness</em>, and <em>performance</em>. We’ll empirically validate correctness later, and we’ve already considered memory ordering. Fairness is a no-go, which is alright in this basic version, but a core might get starved and never get to take the lock if other cores manage to always get it first. Performance is not that good, since cores will spend cycles busy-waiting to get the lock.</p>\n<p>One approach to address the performance issue is to use Arm’s <code>wfe</code> (wait for event) and <code>sev</code> (send event) pair, which puts the core in a low-power state waiting for an event, and wakes all cores up that were waiting. That is something for future work.</p>\n<h1 id=\"validation\">Validation</h1>\n<p>I initially tested the spin lock in the path with the garbled output shown in #Implementing a Spin Lock, to get legible output during the bringup of multiple cores. In QEMU’s virt and rpi4 boards, everything seemed to be working fine. I reran the program a couple of thousand times for good measure and never observed garbled output.</p>\n<pre><code>// SpinLockGuard is a RAII object that locks during\n// construction and unlocks during destruction\n{\n  SpinLockGuard guard(&amp;_print_lock);\n  kprintf(&quot;Running core %d\\n&quot;, cpu_id);\n}</code></pre>\n<p>However, when I ran the same code on my rpi5 I got a mysterious exception. The output below is from my custom exception handler, where we see a Data Abort exception with the Exception Class (EC) value of 37 (0x25). According to Arm’s <a href=\"https://support.arm.com/documentation/den0042/0100/Exceptions-and-Interrupts/Types-of-exception\" rel=\"nofollow ugc noopener\">“Types of exception”</a>: <em>“A data abort exception happens as a result of a load or store instruction”</em>.</p>\n<pre><code>Got an exception (from 0x84474)\n- Syndrome is 0x96000410\n- EC is 37 (Data Abort)\nRegister dump:\n  x0:  0x2bcf88\n  ...\n  x30: 0x85554</code></pre>\n<p>Taking a closer look at the instruction the CPU was executing when the exception occurred (0x84474), we see that it is right on the <code>ldaxr</code> instruction. This aligns with the Data Abort value we also got, so something must have gone awry with the <code>ldaxr</code> memory access. As a sanity check, replacing the load exclusive and store exclusive with a “normal” load acquire (<code>ldar</code>) and a plain store (<code>str</code>) makes the problem go away and no exception is thrown. This tells me there is something blocking the exclusive operations from working.</p>\n<pre><code>0000000000084470 &lt;_ZN13SpinLockGuardC1EP8SpinLock&gt;:\n   84470:       f9000001        str     x1, [x0]\n   84474:       885ffc20        ldaxr   w0, [x1] &lt;-- The CPU was here\n   84478:       35ffffe0        cbnz    w0, 84474\n   8447c:       52800020        mov     w0, #0x1\n   84480:       88027c20        stxr    w2, w0, [x1]\n   84484:       35ffff82        cbnz    w2, 84474\n   84488:       d65f03c0        ret</code></pre>\n<h1 id=\"figuring-out-why\">Figuring Out Why</h1>\n<p>Fortunately, since I had read <a href=\"https://support.arm.com/documentation/110478/0100\" rel=\"nofollow ugc noopener\">Implementation Software Synchronization Primitives in A64</a>, I remember there was a specific section on <a href=\"https://support.arm.com/documentation/110478/0100/AArch64-Atomic-Instructions/The-memory-attribute-of-address-which-is-atomic-accessed\" rel=\"nofollow ugc noopener\">memory attributes</a>. This section highlights the memory attributes for which it is architecturally guaranteed that atomic instructions are atomic. These are:</p>\n<ul><li>Inner Shareable, Inner Write-Back, Outer Write-Back Normal memory with Read allocation hints and Write allocation hints and not transient</li><li>Outer Shareable, Inner Write-Back, Outer Write-Back Normal memory with Read allocation hints and Write allocation hints and not transient</li></ul>\n<p>The page also lists memory types which might not support atomic instructions. The most notable bullet point of those is:</p>\n<ul><li>Device, Non-cacheable memory, or memory that is treated as Non-cacheable, in an implementation that does support hardware cache coherency</li></ul>\n<h2 id=\"memory-types\">Memory Types</h2>\n<p>The default memory type for all memory is Device memory. However, I set up and map virtual memory particularly early in floss, associating all regions of RAM with the “Normal” memory type, so that unaligned accesses can be performed (e.g., an 8-byte load from an address that is not 8-byte aligned). This is because Device memory requires alignment checking. I became aware of this when upgrading from QEMU v8.2.2 to v11.1.50, which included a (not so recent) <a href=\"https://gitlab.com/qemu-project/qemu/-/commit/59754f85ed35cbd5f4bf2663ca2136c78d5b2413\" rel=\"nofollow ugc noopener\">patch</a> to add alignment checking to device memory, which raised a rather infuriating exception that took a while to debug as well.</p>\n<p>Normal memory is a pretty fuzzy term, but is generally all types of memory that are not Device. Normal memory can generally be cached, speculative, and reordered depending on attributes. Device memory that is accessed via memory mapped IO (MMIO) can generally not be speculative or cached like RAM.</p>\n<p>All types of memory are defined by a set of attributes, which can be configured for individual virtual memory mappings.</p>\n<h2 id=\"memory-attributes\">Memory Attributes</h2>\n<p>Memory attributes are configured via the Memory Attribute Indirection Register (<code>MAIR_EL1</code> for EL1). This is a 64-bit register that is divided into 8 different sets of attributes of 8 bits each. An entry in the translation table (page table in Linux terminology) that points to a block of physical memory (i.e., a <a href=\"https://support.arm.com/documentation/102376/0200/Describing-memory-in-AArch64\" rel=\"nofollow ugc noopener\">Block or Page descriptor</a>) has three bits that reference an index in the MAIR register, indicating the attributes for that specific memory mapping. Additionally, such an entry has bits that configure execute permissions, access permission, shareability, and more.</p>\n<p>Relevant to this post, and specifically to exclusive memory operations, is the shareability configuration, which can be set to one of: <em>Non-Shareable</em>, <em>Inner-Shareable</em>, or <em>Outer-Shareable</em>. Loosely, shareability determines in which domain, and for which observers, memory operations should be observed to complete in a finite amount of time and without using explicit cache maintenance. This topic is pretty complex, and I refer you to the <a href=\"https://support.arm.com/architectures/a-profile-architecture\" rel=\"nofollow ugc noopener\">Arm ARM, section B 2.7.1</a> for more information. In floss, the Normal memory type is configured to be Inner-Shareable, which is what Arm <a href=\"https://support.arm.com/documentation/102376/0200/Cacheability-and-shareability-attributes\" rel=\"nofollow ugc noopener\">expects</a> operating systems to do.</p>\n<p>Currently, floss only configures two sets of memory attributes, one for Device memory and one for Normal memory. Let’s ignore the specifics for Device and focus on Normal memory.</p>\n<ul><li><strong>Device</strong>: <code>0b00000000</code>, Device-nGnRnE memory</li><li><strong>Normal</strong>: <code>0b11111111</code>, Normal memory, Outer Write-Back, Inner Write-Back, with Read allocation hints and Write allocation hints and Non-transient</li></ul>\n<p>The Write-Back property means that writes should be written to cache, and flushed to memory only when the cache-line is evicted. In summary, this means that the memory attributes I’ve configured for Normal memory has <em>full</em> support for caching. It also fully matches the architectural requirements for atomics to be atomic.</p>\n<p>An alternative to Write-Back is Write-Through, which writes to both cache and memory immediately, without waiting for the cache-line to be evicted.</p>\n<h2 id=\"global-and-local-cacheability-scopes\">Global and Local Cacheability Scopes</h2>\n<p>As we’ve concluded, the memory attributes for the Normal memory type do have the cacheable Inner and Outer Write-Back property. Even so, with all the details available to us so far, the only logical explanation for why an exception is raised when executing the load exclusive is cacheability.</p>\n<p>With the help of my LLM, I was able to figure out that there is a global setting in the <a href=\"https://support.arm.com/documentation/100442/0100/register-descriptions/aarch32-system-registers/sctlr--system-control-register\" rel=\"nofollow ugc noopener\">System Control Register (<code>SCTLR</code>)</a>: the “C” bit, which controls cacheability for data accesses at EL1 and EL0 (floss runs at EL1). Querying this bit, I see that it’s set to its default value of 0 on all cores, which disables the data cache for all of them. Setting this bit to 1 makes the exception go away on the rpi5. Hoorah!</p>\n<p>Reading up on the documentation of <a href=\"https://support.arm.com/documentation/100798/0401/L1-memory-system/Cache-behavior/Data-cache-disabled-behavior?lang=en\" rel=\"nofollow ugc noopener\">L1 memory system and cache behavior</a> on the Cortex-A76 CPU (the CPU of the rpi5), there is some revealing information:</p>\n<ul><li>When the data cache is disabled, instructions and operations are affected as follows:<ul><li>All load and store instructions to cacheable memory are treated as if they were Non-cacheable and are incoherent with the caches in both this core and other cores in the cluster. Software must take this into account.</li></ul></li></ul>\n<h1 id=\"cacheability\">Cacheability</h1>\n<p>We’ve figured out why the exception was raised on the rpi5, because <code>SCTLR.C</code> with a value of 0 effectively treated the memory as non-cacheable. Interestingly, if I enable the global cacheability setting for each core by setting <code>SCTLR.C</code> to 1, but change the MAIR attribute to something that does not have Write-Back enabled, the same exception is raised. I did a small test of different MAIR configurations for the memory being referenced by <code>ldaxr</code>, and this is the result when executing on the rpi5.</p>\n<div class=\"table-wrap\"><table><thead><tr><th>Bit pattern</th><th>Description</th><th>Exception raised</th></tr></thead><tbody><tr><td><code>0b00110011</code></td><td>Normal memory, Outer/Inner Write-Through Transient</td><td>Yes</td></tr><tr><td><code>0b01000100</code></td><td>Normal memory, Outer/Inner Non-cacheable</td><td>Yes</td></tr><tr><td><code>0b01110111</code></td><td>Normal memory, Outer/Inner Write-Back Transient</td><td><strong>No</strong></td></tr><tr><td><code>0b10111011</code></td><td>Normal memory, Outer/Inner Write-Through Non-transient</td><td>Yes</td></tr><tr><td><code>0b11111111</code></td><td>Normal memory, Outer/Inner Write-Back Non-transient</td><td><strong>No</strong></td></tr></tbody></table></div>\n<p>From this experiment, it appears as if the load exclusive instruction doesn’t work without Write-Back. In fact, the documentation of <a href=\"https://support.arm.com/documentation/100798/0401/Memory-Management-Unit/Specific-behaviors-on-aborts-and-memory-attributes/Memory-attributes?lang=en\" rel=\"nofollow ugc noopener\">Memory attributes</a> on the Cortex-A76 CPU says that: <em>&quot;… a page is cacheable only if the Inner and Outer memory attributes are Write-Back. In all other cases, all pages are downgraded to Non-cacheable Normal memory&quot;</em>. This means that even though Write-Through technically means “write to cache”, the Cortex-A76 CPU does not in fact cache that type of memory.</p>\n<p>From this, it appears as if the load exclusive operation only completes without an exception if cacheability is enabled, both globally with <code>SCTLR.C</code> set to 1, and locally with a Write-Back attribute for the translation table entry.</p>\n<p>It’s a bit frustrating that I wasn’t able to catch this in QEMU. QEMU doesn’t have support for the rpi5 board yet, and there are likely a few differences between the actual rpi5 hardware and the available rpi4 QEMU board. Nonetheless, when digging through the QEMU source code, it turns out that cacheability is simply not being modelled at all (see references in commits <a href=\"https://gitlab.com/qemu-project/qemu/-/commit/fa2ef212df\" rel=\"nofollow ugc noopener\">fa2ef212df</a>, <a href=\"https://gitlab.com/qemu-project/qemu/-/commit/b0fe242751\" rel=\"nofollow ugc noopener\">b0fe242751</a>), reinforced by the constant for <code>SCTLR_C</code> not being referenced <a href=\"https://gitlab.com/search?search=SCTLR_C&amp;nav_source=navbar&amp;project_id=11167699&amp;group_id=3038080&amp;search_code=true&amp;repository_ref=v11.1.2\" rel=\"nofollow ugc noopener\">anywhere at all</a>. So, even though it runs fine in QEMU, don’t assume everything’s perfect until you test stuff out in the real world.</p>\n<h1 id=\"exclusive-monitors\">Exclusive Monitors</h1>\n<p>To achieve atomicity via the load exclusive and store exclusive instructions, one or more exclusive monitors are used. Each core has its own local monitor, and all cores have access to a global monitor. Each monitor has two possible states: <em>open</em> and <em>exclusive</em>. When executing a load exclusive instruction, the monitor(s) are moved to the exclusive state, and the address of the load is recorded. When executing a store exclusive instruction, the monitor(s) are checked to see if they are all in the exclusive state and if the address matches, in which case it succeeds.</p>\n<p>More precisely, the address stored in the exclusive monitor(s) is only the upper address bits, referred to as the Exclusive Reservation Granule (ERG). Two address locations within the same ERG will appear as the same entry in the monitor(s). On a typical system, the ERG matches the size of a cache line.</p>\n<p>Exclusive accesses to memory locations marked as Non-shareable will only access the local monitor, and memory locations that are marked shareable (either Inner or Outer) will access both the local and global monitors.</p>\n<p>Again, Arm’s <a href=\"https://support.arm.com/documentation/110478/0100/Exclusives/Exclusive-monitors/Global-monitor\" rel=\"nofollow ugc noopener\">Implementation Software Synchronization Primitives in A64</a> highlights a very interesting aspect in its documentation of exclusive monitors:</p>\n<ul><li>Although we describe the local and global monitors as being different things, they might actually share logic. For example, one implementation approach is to use a combination of the per-core local monitors and the cache coherency logic to provide global monitor functionality for cacheable locations.</li></ul>\n<p>This feels like the cherry on the top of this mystery. Apparently (or empirically?), the rpi5’s Cortex-A76 CPU seems to require cacheability to be enabled for a load exclusive instruction to not generate a Data Abort exception. This could very well be because it implements the global monitor via cache coherency logic, and therefore, any exclusive memory operation to a memory location marked as shareable (either Inner or Outer) that access the global monitor, will fail when cacheability is disabled.</p>\n<h1 id=\"conclusion\">Conclusion</h1>\n<p>Let’s recap: I wanted mutual exclusion and opted for a “hand-made” spin lock using load exclusive and store exclusive. It worked fine in QEMU, but I got an exception on rpi5 with real hardware. I then ended up scouring documentation (Arm ARM, Cortex-A76, Implementation Software Synchronization Primitives in A64) to understand what’s going on.</p>\n<p>From this adventure, I can conclude that on this implementation and configuration of the Raspberry Pi 5 with an Arm Cortex-A76 CPU, I observe that a load exclusive operation on shareable memory (which presumably accesses the global exclusive monitor), does not work unless cacheability is enabled, both globally (<code>SCTLR.C</code> set to 1) and locally (memory attribute for translation table entry configured with Write-Back).</p>\n<p>During development I iterate quickly with QEMU, but from time to time also verify on real hardware with the Raspberry Pi 5. This creates several layers, which make this a bit confusing, but very interesting at the same time. QEMU says one thing, the Arm ARM another, and then finally the implementation of the Cortex-A76 a third thing.</p>\n<p>This is by far the most interesting bug I have encountered so far while developing floss, but likely not the last. Thank you for reading!</p>","headings":[{"level":1,"text":"When ldaxr Doesn't Work: Exclusive Accesses and Cacheability on AArch64","id":"when-ldaxr-doesn-t-work-exclusive-accesses-and-cacheability-on-a"},{"level":1,"text":"Introduction","id":"introduction"},{"level":1,"text":"Implementing a Spin Lock","id":"implementing-a-spin-lock"},{"level":1,"text":"Validation","id":"validation"},{"level":1,"text":"Figuring Out Why","id":"figuring-out-why"},{"level":2,"text":"Memory Types","id":"memory-types"},{"level":2,"text":"Memory Attributes","id":"memory-attributes"},{"level":2,"text":"Global and Local Cacheability Scopes","id":"global-and-local-cacheability-scopes"},{"level":1,"text":"Cacheability","id":"cacheability"},{"level":1,"text":"Exclusive Monitors","id":"exclusive-monitors"},{"level":1,"text":"Conclusion","id":"conclusion"}]}}