{"article":{"slug":"the-scourge-of-x86-emulation","title":"The scourge of x86 emulation","subtitle":null,"summary":"Welcome to the first feature article on our site. We’re going to cover an ongoing problem with x86 emulation that affects every application that we\nemulate. This comes down to a single over-arching term that has wide-reaching ramifications; Emulating the x86 Total Store Ordering memory model\n(x86-TSO).","content_type":"essay","language":"en","canonical_url":"https://fex-emu.com/Scourge-of-emulation/","author":{"name":null,"url":null,"person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"FEX-Emu","url":"https://fex-emu.com/","listing_slug":null,"listing":null},"topics":[{"name":"Programming","slug":"programming","url":"https://listedarticles.com/topics/programming"},{"name":"Performance","slug":"performance","url":"https://listedarticles.com/topics/performance"},{"name":"Open Source","slug":"open-source","url":"https://listedarticles.com/topics/open-source"},{"name":"Hardware","slug":"hardware","url":"https://listedarticles.com/topics/hardware"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":7172,"reading_minutes":31,"published_at":"2026-09-17T00:00:00.000Z","added_at":"2026-09-18T06:16:46.876Z","updated_at":"2026-09-18T06:16:46.876Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":true},"profile_url":"https://listedarticles.com/articles/the-scourge-of-x86-emulation","markdown_url":"https://listedarticles.com/articles/the-scourge-of-x86-emulation.md","example":false,"citation":"FEX-Emu. \"The scourge of x86 emulation.\" 17 Sept 2026. https://fex-emu.com/Scourge-of-emulation/ (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://fex-emu.com/Scourge-of-emulation/"},"body_markdown":"Welcome to the first feature article on our site. We’re going to cover an ongoing problem with x86 emulation that affects every application that we\nemulate. This comes down to a single over-arching term that has wide-reaching ramifications; Emulating the [x86 Total Store Ordering memory model\n(x86-TSO)](https://en.wikipedia.org/wiki/Processor_consistency#Similarity_to_SPARC_V8_TSO,_IBM-370,_and_x86-TSO_memory_models).\n\nThe problems with emulating this memory model on the [weak ordering memory model](https://en.wikipedia.org/wiki/Consistency_model#Weak_ordering)\nthat ARM defines is multi-faceted and covers multiple issues. We’re going to go over all the problems that we can encounter and the ways we solve\n(or in some cases can’t solve) in this article. Get yourself a snack and a warm drink to enjoy, this is going to be a long one.\n\n# What exactly is x86-TSO?\n\nBefore diving in to how we work around the x86 memory model problem, we need to first discuss exactly what it is. A memory model is a set of rules for how memory accesses in a system behave in relation to each other. The rules will dictate how loads and stores interact in a single-threaded or a multi-threaded environment. There’s a handful of popular memory memory models implemented in various forms of hardware, but the two we care about today is ARM’s relaxed (or weak) consistency model, and the x86 variant of Total-Store-Ordering consistency model. These two models are basically the two extremes of the spectrum; where ARM is the most relaxed, allowing significant hardware optimizations; and x86 is the most strict, enforcing a very strong coherency model that doesn’t allow a lot of room for optimization. One thing to be careful about when discussing memory models is the difference between consistency and atomicity. While these are related, they are not the same nor guaranteed in all cases.\n\nThe best way to explain how the differences in memory models work is to start with how x86 handles this. With TSO being very strict in how it operates, the programmer can assume that when a memory store occurs, that this will be coherently visible to all other processors in the system. This additionally means that when a memory load occurs, all stores before it “logically” will have been completed, or at least visible. This matches programmer expectations, you write to memory, it becomes visible as at the point of writing, as this is intuitive to think about when programming. The stores are effectively ordering the visibility of the loads, thus the name of the model. There’s a bit of nuance with how this operates but isn’t strictly necessary to understand.\n\nThe weak memory model that ARM has is a bit less intuitive about how it operates. By default the regular memory loads and stores that ARM uses aren’t strictly coherent across processors in your system, allowing the CPU to operate more efficiently most of the time. When a store instruction executes, that piece of memory (the cacheline) isn’t immediately visible to other processors in the system. Saving on precious power and efficiency because it’s expensive in hardware to invalidate other core’s cachelines, or allow them to snoop another processor’s caches. Relatedly if a processor is loading data from memory that another processor has written to, it’s not guaranteed that this load will even see this updated memory. This sounds like it would cause some significant problems in a multi-threaded application right? Older versions of ARM (ARMv7 and older) used a memory barrier instruction to ensure ordering, which had significant performance implications.\n\nTo get around this limitation of consistency, ARM also introduced load-acquire, and store-release memory instructions. In C++ parlance this maps to\n**std::atomic’s** *memory_order_acquire* and *memory_order_release* definitions respectively. In ARM’s terminology, these instructions also aren’t\n***technically*** considered to be atomic operations, but programmers conflate the two. FEX has used the terms atomic-load and atomic-store to mean\nthe same thing! The distinction *usually* doesn’t matter, but when discussing these topics it may be better to be pedantic about it.\n\nThe primary use case for these instructions is to force memory ordering between these class of instructions. ARM calls this the “Release Consistency sequentially consistent (RCsc)” model. Without getting too far in to the weeds about how this model operates, the basic gist is that the load-acquire instructions must be observed sequentially without reordering, and the store-release instructions must as well while fulfilling “barrier-ordered-before” semantics. Removing the costly memory barrier instruction required in older ARM architecture versions.\n\n# The humble beginnings of ARMv8.0-a\n\nThis is the premise of where we start in ARMv8.0-a when we’re emulating the x86-TSO memory model. We make all x86 memory loads turn in to ARM’s\n**load-acquire** instructions, and x86 memory stores turn in to **store-release** instructions. This gives FEX effectively the same memory semantics\nas x86, although we are actually being *more* strict than what is necessary. This is because we had no middle-ground which exactly matches behaviour.\nAs one might think, it is *exceedingly* costly to emulate TSO wth this instructions and we have microbenchmarks that can show this.\nAs ARM CPUs weren’t designed to have these relatively rare acquire/release instructions suddenly become the vast majority of instructions executed.\n\nFirst let’s start with something easy and use a microbenchmark that is fairly nice to the hardware. No tricky edge-cases, just accessing memory in in the common case. This gives us some baseline numbers for what the best-case situation should be.\n\nLet’s break down this graph as it tells us a few interesting stories. The Load and Store columns of each machine is representing our baseline performance number that our hardware should be attempting to achieve. These aren’t trying to max out the memory bandwidth of each system, but do the same amount of work for each type of operation. If we turn our attention to the acquire-load results, we can see that out of the five CPUs tests, three of them have their performance hindered quite a bit by using acquire-loads! Additionally we can see that the AmpereOne CPU has release-store instructions that are strikingly low compared to the other results, and the M1 Acquire/LRCPC load instructions are quite a bit lower than the baseline as well.\n\nThe AmpereOne results in particular showcase how bad this legacy path can get. These instructions were never designed to be used this way. Using acquire-release semantics for every load for x86 emulation actually imposes some really strict limitations on ARM CPUs in that the load instructions can no longer be ordered around each other at all. So when you have millions of them in flight per second, the performance isn’t really expected to be good. But because these are the only instructions we had with ARMv8.0-a, it’s what we had to use. While Cortex-X4 and Cortex-X925 have amazing performance for these, you can see how the Oryon-3 has deprioritized their importance.\n\n## Where do we go from here?\n\nLet’s take a closer look at the LRCPC-load instructions, which is mandatory since ARMv8.3. This extension adds a bunch of new load instructions to the ARM ISA and adds a new memory model on top of\nARM’s *RCsc* model from before. This new “Release Consistency processor consistent (RCpc)” memory model is what we’ve been wanting! This extension is\ndesigned around the requirements that x86 emulation requires, and is expected to get utilized heavily on hardware that implements it. As you can see from the\ngraph, almost all of the platforms have their LRCPC-loads matching their regular loads in performance.\n\nWith this new extension that is mandated by newer ARM versions, we basically get **solved**\nmemory performance. At least according to this microbenchmark that seems to be the case. Once FEX detects this extension we stop using Acquire-Load instructions\nentirely and switch over to LRCPC-Load instead. But what’s going on with that Apple M1 result..?\n\nThis is where we need to commend Apple’s path towards solving this problem. With their Apple Silicon processors they directly added support for the\nx86-TSO memory model. When the CPU feature is toggled, their ***regular*** load/store ARM instructions change behaviour to match what x86 requires. They went\nthis route knowing that they will need a high performance solution for their hardware when switching to the ARM ecosystem exclusively.\nThis is why on their hardware the LRCPC-load instructions are actually aliases of their acquire-load instructions, because their x86\nemulator doesn’t even use these instructions! Because they implement the x86-memory model, they just use regular load/store instructions, which can be seen in our\nmicrobench results as indiscernable performance overhead. To be fair to the other platforms, this thread-wide TSO mode toggle does have some\nperformance impact, we just don’t see it here. When FEX detects this CPU feature from [Asahi Linux](https://asahilinux.org/) we will also enable this\nand get the “free” performance improvement. A potential concern is that when jumping between x86 emulation and ARM code, that the ARM code\nwill pay unnecessary overhead due to all its accesses being TSO now. While this is a reasonable concern, the amount of ARM native code executing under\nemulation approaches 0%. As a developer, you don’t care about 1% of memory accesses becoming 10% slower, you care about 99% of accesses becoming 15% of\nthe “ideal” (As shown in AmpereOne results).\n\nAs a note, we think a TSO mode is the best path forward for ensuring high performance x86 emulation on the platform. Because this ensures that every memory\naccess instruction behaves how we want or expect. This is shown with the official **FEAT_LRCPC** extension actually having three versions that\napply bandages to the implementation each time.\n\n- **FEAT_LRCPC** - Adds basic GPR TSO load instructions\n- **FEAT_LRCPC2** - Adds small offset immediate to TSO load instructions\n- **FEAT_LRCPC3** - Adds basic vector and stack-based TSO load & store instructions\n\nEven with these three extensions, there is edge-case behaviour that can’t be emulated as nicely as if we had a TSO hardware toggle. We are expecting there to be additional extensions versions as time goes on, trying to fix some of the additional problems we’ll discuss later in the article.\n\n# I thought accessing memory was the easy bit?\n\nIn the previous section, we were being nice to the ARM hardware and playing along with the underlying hardware’s alignment requirements to get a\nbaseline for what the performance should look like. When emulating x86 although, we run face first in to a glaring problem right from the start. Your\nfavourite x86 applications don’t care about alignment! They’ll access memory however they please, crossing cacheline granularities, doing atomics that\naren’t aligned. You think of the alignment problems, these games are doing it. This problem is so bad that we have a term associated with it, called **split-locks.**\nThese are such a big deal that even the Linux kernel will capture when these occur and slow down games when they do it! Causing many gamers to tinker\nwith kernel options to avoid the slowdown!\n\nBut we aren’t going to talk about full on split-locks yet, let’s get started with just load-store instructions in an environment that doesn’t care\nabout alignment. x86 makes certain guarantees to the programmer; if you do a load-store and it is inside of a cacheline then that load-store will be both atomic and still match the coherency model as described before. However, to be a\nlittle bit nice to the hardware developers, if the load-store *does* cross a cacheline, the data isn’t atomic and other threads can and will see it\ntear. So the programmer needs to be careful as a basic load-store is not a split-lock.\n\nThe problem with emulating these basic accesses with load-acquire/store-release is that ARMv8.0 requires what is known as **natural alignment**.\nThis means that for whatever size of data being accessed, the offset in memory must match the size. So for an 8-byte access, it must be at offsets; 0,\n8, 16, 24, etc. This works well for native ARM applications, but what happens when we don’t obey natural alignment requirements? For ARM, this means\nthe instruction with raise an **alignment fault.**\nThe hardware validates that the alignment requirements are fulfilled and if they are not then the CPU will fault. This usually results in a crash but\nFEX does special handling.\n\nInside of FEX’s JIT mechanism we keep track of memory load-store instructions that are emulating the\nx86 load-stores. When we know that a load-store can cause an alignment fault we have what is known as a **patchpoint** in the code. For load-store\ninstructions, this shows up as a **NOP** instruction either before or after the load-store. When a alignment fault occurs as one of these patchpoints, FEX\nwill capture the fault, patch the code from a load-acquire/store-release instruction to a **basic** equivalent load-store, and wraps the instruction\nin a data memory barrier. Then it continues executing!\n\nBefore then after patching\n\nThat entire discussion from before about how ARMv8.0-a added these new fancy load-acquire, store-release instructions? We immediately fall back to the classic memory barrier instruction instead when alignment behaviour doesn’t match. Our previous chart didn’t show this bad case, so let’s bring in some fresh data.\n\nOh, that’s a lot of data to sift through. While again good to see how far away the hardware is from the “optimal” path while emulating TSO, it’s not what we care about here. It is interesting to note that this microbench doesn’t showcase much of a difference between aligned and unaligned for regular load/stores so we just calculated an average between the two. We’ll be removing the x86 CPU and the regular load-store data from the ARM columns, as these aren’t the common FEX paths. This way we’ll have a more targeted view about how badly unaligned memory accesses hurt under emulation.\n\nNow that we have a much more reasonable graph of data, let’s walk from left to right on this and discuss what is going on.\n\n## AmpereOne\n\nThis one is pretty interesting, both the aligned and unaligned load instructions are roughly equivalent and fall within noise. This means that even though the unaligned loads are getting hit with a data memory barrier penalty, the CPU just handles it. This might be the case that the benchmark is bottlenecked by other things, considering how much lower the performance is compared to other platforms.\n\nMeanwhile the store side is not looking to be in a good shape even without unaligned. It nearly isn’t visible on the chart! When hitting unaligned stores we’re looking at ~8.5% of a performance hit, but because we are already starting so low it is hard to notice. This is also in stark contrast to regular store instructions getting ~28GB/s in this bench.\n\nThe only conclusion we can come to here is that Ampere is optimizing for some server class workload and doesn’t really match consumer hardware behaviour. It’s an interesting datapoint, but our users aren’t typically running games on this class of hardware.\n\n## Cortex-X4\n\nThis is a highly popular CPU core that is living inside the [Qualcomm Snapdragon 8 Gen 3](https://en.wikipedia.org/wiki/Steam_Frame). We only tested\nthis one core from the SoC to not overwhelm the chart with data. Quite a large number of handhelds ship with this so it’s an interesting\ntarget. This CPU actually does **surprisingly** well considering it’s the only cellphone SoC on this list. Overall this core kind of falls in line\nwith what we would expect from it and the graph trends follow with the next-generation Cortex in that chart.\n\nThe main topics for this CPU are that its aligned loads and stores are reasonably powerful, getting around 11.5GB/s and\n6.7GB/s respectively. What’s interesting is the performance falloff when it needs to deal with unaligned loadstores, hitting the **DMB** instructions\npenalizes the core roughly evenly between loads and stores at around 50% in this benchmark.\n\nThis seems to imply that the CPU can keep a decent number of LRCPC-release loadstores in flight so the **DMB** instructions hurt more when they are\nencountered, but it isn’t causing world-ending performance. Just that a 50% performance hit due to alignment isn’t an amazing result.\n\n## Cortex-X925\n\nFollowing up the X4, let’s stop by the [DGX Spark](https://www.nvidia.com/en-us/products/workstations/dgx-spark/) and its X925 cores. Not only is this\na newer CPU core from ARM, it’s running on a system with dramatically more memory bandwidth. 273GB/s in the platform versus the previous 76.8GB/s. This\nmeans that we get fairly similar results to the X4 even, just the graph scales a little higher. Interestingly enough, the performance penalty for\nunaligned accesses roughly match the X4 even. Although it looks like the stores can recover a little faster, likely due to the faster memory helping\nout. No surprises here, just consistently matching performance across the generations.\n\n## Oryon-3\n\nThis CPU core design is hot off the presses from Qualcomm. Linux support is still in the process of coming up but it already has a strong showing. The most interesting result from this actually comes from the fact that aligned LRCPC-load instructions are matching the performance of regular loads! That means in the case of a well-behaved application we can typically expect full performance. This continues onward to the release-store instructions being quite capable, although it doesn’t quite match regular stores with only 68% of the bandwidth. Not a bad showing in the slightest.\n\nThis CPU also can’t escape from the penalty of unaligned LRCPC-release loadstores. The load side is roughly matching the ~70% performance penalty of the Cortex-X925, likely because the Snapdragon X2 Elite also has tons of bandwidth. But the store side actually gets off a little worse at ~43% of the performance. Even with these performance hits of unaligned accesses, this platform is actually faster than the aligned accesses from the Cortex offerings.\n\nOne of the weird things about this platform is that it was advertised to have “Fully coherent 96KB 6-way L1 cache with 64B coherency granules.” Which to our reading implied that unaligned accesses should have dramatically less of a performance impact. Interesting… keep that in mind.\n\n## Apple M1\n\nThis is the big one we need to talk about. This is the one that was a game changer, it was the “Apple moment.” It showed everyone that ARM was\nnot only feasible, it could be faster. These numbers on this chart are amazing and it’s the result of Apple sticking the TSO memory\nmodel directly in to their hardware. Instead of using LRCPC-release accesses for this one, we just enabled their TSO feature and the aligned\nversions basically match the unaligned version. Maybe a 5% performance hit on the stores? Compared to every other device on that chart, it’s\neffectively nothing. This primarily comes down to unaligned accesses no longer requiring **DMB** instructions to be backpatched in to the code, as the\nhardware just handles it directly.\n\nFor us, this is what it means to take x86 emulation seriously on ARM and it really shows that Apple cared that their customers would have a good\nexperience running software both natively and emulated. They saw the problem and just solved it, making it go away.\nThat said, when the TSO mode *is* enabled, you do get a performance hit. Comparing to the previous graph it’s only getting 76% of the regular\nstore performance, and the load performance basically matches; that’s much more tolerable to bear when everything is so much faster.\n\n## Wrapping up unaligned LRCPC/release accesses\n\nWrapping up this section, we need to talk about one of the performance improvements that all of these vendors actually support. This is an\nextension that ARM whipped up called **FEAT_LSE2** which all of these tested platforms implement. We previously talked about how acquire/LRCPC/release\nmemory accesses require natural alignment in order to not incur the wrath of the CPU raising alignment faults. ARM actually thought about\nthis problem and implemented this extension which helps x86 emulation (and probably other workloads). This extension loosens the alignment\nrequirements of not only acquire/LRCPC/release load store instructions, it *also* loosens the requirement for read-modify-write atomics!\n\nThat sounds all well and good, but here’s the kick to the teeth: that means it only provides marginal performance gains for x86 emulation. This\nextension only loosens the alignment requirements to allow unaligned memory accesses inside of a 16-byte granule. Any access that crosses that 16-byte\ngranule still receives an alignment fault. x86 applications don’t really care about the alignment of their memory accesses, so we get\nunaligned accesses across the entire cacheline. It’s only read-modify-write atomics that **try** to avoid crossing a cacheline on x86!\n\nSo thanks for the attempt, it’s nice to see, but it doesn’t really move the needle. Since we’re already talking about it, let’s dive in to those RMW atomics shall we?\n\n# Oh no, what are these atomic instructions?\n\nLike most modern instruction sets, x86 supports atomic memory operations. These are instructions that execute an ALU operation on data in memory\natomically, allowing no intermediate state to be visible. In x86 terms this operates on memory that is both atomic and coherent, while ARM lets you\nchoose to be only atomic *or* both atomic and coherent. We touched on this briefly before but there is actually a difference between operating on data\natomically, and coherency of that data. What difference does it make?\n\nFor all of the previous x86 memory model discussion we have been talking about the coherency implications of loads and stores being visible to other\nprocessors in the system. What we entirely glossed over is the atomicity requirements of these memory accesses. In the world of x86 a load or store\n*usually* completes atomically even when unaligned. This means that if you’re storing 8-bytes of data, and another thread is loading those 8-bytes in\na race condition it will never suddenly see a mix of the data from before the store and after the store. In ARM these atomicity guarantees are\n**significantly** weaker, meaning if you do an unaligned store instruction the specification of the ISA has zero guarantees about reading a **tear** in\nthe data. Thankfully for **naturally** aligned load-store instructions, ARM has a specification called **“single-copy atomicity”** which guarantees\nyou don’t get a tear for these accesses. Also good news; that **FEAT_LSE2** extension from before? It actually extends the\nsingle-copy atomicity guarantees to *any* unaligned access inside of a 16-byte granule! The downside is that x86 has single-copy atomicity\nguarantees across a full cacheline, so once again the extension still didn’t solve anything completely, just reduced the number of occurences.\n\nEnough about the differences in atomicity and coherency. Where’s the actual atomic instructions? What do they do? Starting in ARMv8.1-a, our ISA has gained instructions that mostly matches x86 atomic instructions in behaviour. Let’s just give the full list to show how they map directly in our JIT.\n\n| **x86** | **ARMv8.1-a** | \n|---|---|\n| LOCK DEC | ldaddal | \n| LOCK INC | ldaddal | \n| LOCK NEG | [???](https://github.com/FEX-Emu/FEX/blob/e02953dc174531551219712df20355dcf4afc089/unittests/InstructionCountCI/FlagM/Atomics.json#L1368-L1381) | \n| LOCK NOT | ldeoral | \n| LOCK ADC | ldaddal | \n| LOCK ADD | ldaddal | \n| LOCK AND | ldclral | \n| LOCK OR | ldsetal | \n| LOCK SBB | ldaddal | \n| LOCK SUB | ldaddal | \n| LOCK XADD | ldaddal | \n| LOCK XOR | ldeoral | \n| LOCK BTC | ldclralb | \n| LOCK BTR | ldeoralb | \n| LOCK BTS | ldsetalb | \n| LOCK CMPXCHG | casal | \n| CMPXCHG8B | caspal | \n| CMPXCHG16B | caspal | \n\nWell would you look at that, we have a full list of the 18 atomic RMW operations and they basically map directly to some ARM instructions. Ignore the questionable one as it’s not used in real workloads and we would get far too in to the weeds talking about it. We have a pretty clear 1:1 mapping between the architectures, job’s done right? That’s the funny thing about x86 emulation, just because we have these instructions doesn’t mean we get to wire them up without problems. We spent all this time talking about how unaligned accesses can really hurt performance of regular loads and stores, this same problem also applies to RMW atomics!\n\nWith this graph, we are looking at a single atomic instruction with its memory address landing somewhere within a cacheline. If we included all of the\ndata for all 18 atomic operations then this data would be even more overwhelming than it already is. All these atomic operations behave *roughly*\nequivalent so it would be redundant and wouldn’t matter for what we’re discussing here anyway. This is also the first graph in this post that is\nactually using logarithmic scaling, so when reading it make sure to understand that the performance difference from the fastest to slowest result is\non the scale of around 1000x.\n\nStarting with the x86 Zen processor on this graph; these are the results that our emulation should be striving to achieve. As we can see, if the access is\nfully contained within a cacheline then the latency of the instruction is the same at 1.44ns. This can be explained by x86 having “atomic cachelines”\nor “coherent cachelines”, where as long as an unaligned atomic operation stays within a cacheline then it roughly costs the same. This is a really\npowerful feature of x86 that has been supported for decades at this point so games end up relying on this heavily without even realizing it. The\nstand-out result for x86 is the final result that is crossing a 64-byte granule and taking ~660ns! That’s an amazingly slow result at ~458x slower\ncompared to the other results because this is finally the hardware using **split-locks**.\n\nWe need to take a moment here to shout out an article that [Chips and Cheese](https://chipsandcheese.com/p/investigating-split-locks-on-x86) wrote\nwhile we were preparing to write our article. They do a great deep dive in to why these **split-locks** are so dramatically slower and is worth the\nread if you’re unaware of how they work. Specifically we need to mention that x86 split-locks maintain the atomicity and coherency requirements of\nx86-TSO and will ***never*** tear the data even when crossing a cacheline. This is kind of nuts and we’ll explain this more later.\n\nNow for our ARM processors, let’s start with the natural alignment latency numbers. As we can see, all of our platforms perform fairly well but even\nthe latest cores don’t get anywhere near x86. Even our fastest ARM platform is ~3x the latency compared to x86; This directly impacts performance of\ngames but usually isn’t the direct bottleneck so it’s hard to measure exactly how much. Continuing onward to the next data point, we can actually\ncombine the results for 16-byte granule and 64-byte granule crossing with most of our ARM platforms. Due to how the ARM specification defines how\nunaligned atomics work, both of these results are roughly equivalent and FEX treats them the same as the x86 **split-lock** problem.\n\nWe keep bringing up this split-lock problem but how exactly does FEX emulate them and what makes it so slow? “I thought Apple M1 added x86-TSO support\nin the hardware, why is it still slow?” If you recall how we brought up before that **FEAT_LSE2** introduced support for unaligned memory accesses within\na 16-byte granule; these split-lock operations end up hitting the same alignment problems as before but are dramatically slower. FEX\ncan’t backpatch any of these instructions to just do a **DMB** operation, so we cause an **alignment-fault** every time one gets executed. This means that\nwe do a kernel -> userspace signal handler -> kernel -> original code dance. *every—single—time* one of this split-lock operations execute.\nJumping between kernel-space and userspace is slow on every platform and when you’re executing thousands of these per second it adds up very quickly.\nThis is why the emulation of these feature is so terribly slow on ARM.\n\nOne ARM platform today actually partially resolved this problem although. The Oryon-3 CPU cores introduced what they advertised as “coherent\ncachelines” and we can see this in our microbenchmark results here. Just like with x86, if the atomic memory access in anywhere inside of the 64-byte\ncacheline, the performance matches the natural alignment version! This is a tremendous improvement that means the CPU is on par with x86 in\nfeature support until the point it tries to cross a cacheline. We need to applaud Qualcomm on implementing this feature, as it resolves a major\nperformance and correctness problem around split-locks for x86 emulation. The hardware still doesn’t support 64-byte **split-locks** so we still fall\ndown the FEX emulated path in that instance although.\n\nContinuing on to the Apple result; even though they added x86-TSO memory accesses to their hardware for some reason they neglected to implement full cacheline unaligned atomics like Oryon did. It seems like they should have expected this edge case to surface and implement it but that’s just speculation. This is why you can see the cross 16-byte granule behaving the same as other platforms even with the TSO hardware toggle enabled.\n\nYou might have also noticed another little data quirk in the graph. We have an asterisk on the Cortex-X4 result in this benchmark and the performance\nof the unaligned atomics are dramatically faster than significantly newer CPUs. It is somehow managing to have only\n~209ns latency, while the X925 is latency is 1060ns; that’s a 5x perf improvement! How can this possibly be the case? This is actually some fun\n“special sauce” that is shipping on the platform we’re testing on, which is of course the **[Valve Steam\nFrame](https://store.steampowered.com/sale/steamframe)**. Because Valve cares about the performance of their existing gaming catalogue, they are shipping a\n[kernel patch](https://github.com/bylaws/linux/commit/7ae989a43ae7e3cb8007ac21c28dacc24c9d8320) that one of the FEX developers whipped up. This allows\nthe Linux kernel itself to handle the unaligned atomic without that slow dance with FEX and userspace, allowing it to be dramatically faster. If other\nplatforms want to ship this patch in the kernel then we recommend picking it up as and FEX will automatically start using it.\n\nSpeaking of kernel intervention, we need to talk about how split-lock emulation is not actually quite correct under FEX due to limitations in the hardware. In order to implement this mandatory feature of x86 correctly, any time we do a 16-byte or 64-byte split-lock, the only way to handle it is to have the kernel implement the feature. Right now FEX implements this as a “best-effort” attempt that can actually tear the data in some cases. You’ll recall that before we said split-locks on x86 will never tear right? Not even the Oryon-3 with its “coherent cachelines” have resolved this problem yet.\n\n# What do you mean split-lock is mandatory?\n\nImplementing split-lock emulation with today’s ARM hardware in a performant matter is actually really difficult to do. A naive implementation is to\nuse a global mutex and whenever a split-lock occurs we will ensure to acquire the mutex before doing the operation. This means that any\n*participating* split-lock operation will funnel through this mutex. This is correct except for the issue that any aligned atomic operation\nisn’t a split-lock and won’t participate. Due to the split-lock emulation code needed to be implemented as two 64-bit compare-exchange\noperations with each half straddling the granularity boundary, we can get a tear with a non-participating atomic still. A trivial example is one\nthread constantly modifying an atomic in the middle of the cacheline, and then another thread modifying *only* the integer on one half. This might sound\nlike a contrived example initially, but there are lock-less [linked-list](https://en.wikipedia.org/wiki/Linked_list) implementations that behave exactly like this!\nDepending on which half the aligned thread is modifying, either the first or second CAS in the split-lock code will fail. If the first CAS fails, then\nthat’s safe and the code can retry, if the *second* CAS fails that means the data has torn and we can do nothing but hope it doesn’t corrupt data and\ncrash. This will entirely depend on the algorithm that the guest application is using so we don’t control it.\n\nAn alternative approach that is completely untenable is to have the kernel track all processes and threads that are sharing memory with each other,\nthen when a thread needs to emulate a split-lock the kernel can halt ***every*** process that is sharing memory with that process, do the split-lock\nin isolation, and then restart the world. The performance implications of this approach aren’t viable. Applications and games can end up doing thousands or more\nsplit-locks per second and halting the world will have an intractable performance hit that is dramatically worse than even x86 native.\n\nIf we want to ensure correctness in the emulation of split-locks FEX needs to have hardware support in some form to support these. Although we’re not\nsaying that all atomic operations should now support split-locks like x86, that would also not be viable. The good news is that ARM actually has an\nextension for this that does exactly what we want. ARM has an extension call **[Transactional Memory\nExtension](https://en.wikipedia.org/wiki/Transactional_memory)** that could solve our problem. This extension allows our code to do some number of\noperations inside of a transactional region, then commit that work atomically; if the commit operation fails, then we can simply retry. The downside\nof this extension? ARM has officially deprecated the extension and no one ever shipped it. This is likely for the best as the x86 version of the\nextension has had an abundance of problems that caused it to be disabled on many platforms.\n\nSo we need something else to emulate split-locks correctly. For a solution that we believe works for both FEX needs and ARM vendor needs, we have come\nup with the idea that a 128-bit **CASP** instruction can be given the ability to have each half of the CASP perfectly straddle\nthe atomic granule boundary, 64-bits on the lower half, and 64-bits on the upper half. Then *only* in that case does the instruction not raise an\nalignment-fault and tries to do the CAS operation. This works because x86 only has up to 64-bit unaligned atomic operations, so both halves of the\noperation can always be fully enclosed by our single operation.\n\nBut you may be asking yourself, “how is this any better than the hardware just supporting split-locks?” That’s a good thought and we need to be\ncareful with the how exactly we describe this operation. For x86 their atomic operations must *always* succeed without tear. For our emulated\napproach, we can have this ARM **CASP** instruction fail safely and then we can try again. This is one of the benefits of CAS is that\nthe operation can fail for *any* reason and it must be tried again. The instruction then also returns the data that it loaded from memory in that time\nso the program has the latest up to date memory. This is an important distinction since that means FEX can retry the **CAS** operations infinite times\nuntil it inevitably succeeds! This is a benefit of ARM LL/SC architecture that basically allows this to work. A tricky thing is that the hardware does\nneed to guarantee forward progress at *some* point but it already has support for that for other reasons so it’s completely viable! The only newly\nadded failure mode to the *CAS* instruction is purely if one of the two cachelines got acquired by another core before it could do the full operation.\nEven if the hardware still requires up to a couple thousand cycles to guarantee forward progress, that basically matches x86 behaviour.\n\nWe think this would be the best way forward for x86 emulation of split-locks on ARM platforms, but we’re not hardware architects so all we can do is complain and hope someone solves it for us. We’ll leave the split-lock discussion there for now so we can move on to another interesting problem.\n\n# Wait, uncached memory needs to work?\n\nBefore we get in to this topic we need to talk about the term “uncached” because it can mean a couple of things depending on your view of the\nworld. For the purposes of this article, we are using Vulkan terminology because we care about games primarily. In Vulkan terms we have\n*VK_MEMORY_HOST_CACHED_BIT* which means that the host CPU caches this memory. The lack of this bit is what we care about here, and what we refer to as\n“uncached.” As for what this means to the memory subsystem, it gets a little more complicated than you would think. In particular when the memory is\nliving on a GPU, potentially over PCIe, when the memory is “uncached” it will also typically (but not always!) also gain the flag\n*VK_MEMORY_HOST_COHERENT*. This means that because of the uncacheable property of the memory, the CPU and GPU always have a coherent world memory view\nwith each other.\n\nFor the CPU this typically means the memory can be mapped up to three ways. When asking for “cached” memory, this typically has a memory type of\n[Write-back](https://en.wikipedia.org/wiki/Cache_%28computing%29#WRITE-BACK) which is also what regular memory mapping types are. “uncached” mapping\ncan be either [Write-Combine](https://en.wikipedia.org/wiki/Uncacheable_speculative_write_combining) or “Strong Uncacheable”. The “Strong Uncacheable”\nimplementation is basically non-existant for userspace applications so we can ignore that for today’s discussion. This limits us to effectively **WB**\n(cached) and **WC** (uncached) memory types. Cached is what games typically use for staging buffers, and then uncached is what we use when passing data directly\nto the GPU.\n\nThis is code-ified in many game engines that if you don’t expose support for uncached buffer types then some don’t work. This comes down to a\nbehaviour detail around the differences of [UMA](https://en.wikipedia.org/wiki/Unified_memory_architecture) systems like APUs and PCIe GPUs. UMA\nsystems will typically expose the ability to allocate memory that is cached, coherent, and GPU visible. Where PCIe GPUs can’t guarantee that behaviour\nso game developers need to either use a staging buffer and an async copy of the data over to the GPU, or use “uncached” memory to very carefully\nshuffle the data over to the GPU through PCIe. Because of how ubiquitous PCIe is with PC gaming, some engines won’t even do UMA specific code\npaths and will do the uncached approach regardless!\n\nWith that little introduction out of the way for what uncached means for us. Let’s bring up a benchmark for how fast cached memory is on some UMA Snapdragon systems. This will let us get a baseline for how the performance should be regularly.\n\nFor both the **Steam Frame** and **Snapdragon X2 Elite** these are some really good results. As we would expect, the Oryon-3 platform has more memory\nbandwidth so it is able to scale higher in the chart, but both are hitting dozens of gigabytes per second in their results. This graph sets a good\nbaseline for what “normal” **write-back** memory can achieve. Let’s now show uncached results to see the performance differences.\n\nThere’s some strange things happening here so we had to use logarithmic again on this graph. Let’s talk about the good first that has shown up.\nDue to uncached memory buffers being write-combine, we can see that the regular stores for our ARM platforms match the cached benchmark\nresults. This comes down to write-combine memory using what is coined as [write combine buffers](https://en.wikipedia.org/wiki/Write_combine_buffer)\nthat actually *very* temporarily keep around a cacheline of data so that write-combine can burst a cacheline of memory at a time. Interestingly enough\nit looks like the Zen 4’s WCB can’t quite keep up with cached, but considering this is expected to be going over a PCIe bus it’s probably fine.\n\nNow let’s get in to the really ugly results that we have here. Starting off with the easier to explain is the load bandwidth from write-combined\nmemory is abysmal on all platforms tested. If we’re using Zen as our baseline for performance, then our regular load instructions are ARM are winning,\nbut the LRCPC loads are worse. What’s going on here? This is a quirk of how write-combined memory operates, because it is uncached our load\ninstructions are required to go out to system memory for every single access to maintain semantics. Then when we add LRCPC-loads on top of that, it\njust compounds the problem even further. But the worst case out of all of this is just how badly the store performance is, compared\nto the performance that Zen gets on the stores, this is basically a showstopper. Up to **816x worse** bandwidth! We had games like [Hollow\nKnight: Silksong](https://store.steampowered.com/app/1030300/Hollow_Knight_Silksong/) and [Subnautica\n2](https://store.steampowered.com/app/1962700/Subnautica_2/) run at less than 1FPS because of this performance cliff.\n\nAs we were saying above, when there are PCIe GPUs in the mix then games will need to use uncached memory to pass data to the GPU. When emulating x86\ngames on platforms with a dedicated PCIe GPU then we are in an unwinnable situation and we are guaranteed to run dramatically slower. Remember how ARM\nhas added the family of **FEAT_LRCPC1/2/3** extensions from before to improve x86 memory model emulation? This is what happens when we hit an\nedge-case that isn’t supported. All of these extensions add new instructions to handle loading memory using x86-TSO memory model semantics but none of\nthem solve storing to write-combine memory with x86-TSO semantics. All the way from ARMv8.0-a our store instructions use the regular `store-release`\ninstructions regardless of the backing memory type. The only way for FEX to work around this problem is to selectively disable TSO-emulation when it\nbecomes an issue, so x86 emulation platforms with PCIe GPUs will always be a worse experience than UMA. At least until we get another **FEAT_LRCPC4**\nor similar to resolve the issue.\n\nFor users on UMA systems then rejoice, there’s a workaround for gaming that we use to improve performance. Because we know when a platform supports\ncache-coherent CPU and GPU combinations, we can have the video driver ***always*** use cached buffers and never encounter this problem.\nNVIDIA already does this on their Tegra platforms, Snapdragon has been supporting this since at least Adreno 600 class GPUs, and there are many\nMali platforms where this is also the case. We have a [Adreno Turnip](https://gitlab.freedesktop.org/mesa/mesa/-/merge_requests/41323) patch that\nensures when FEX is running, we never hit uncached memory for platforms that support it. A funny thing is that since Asahi users have a hardware TSO\nbit, they just naturally don’t encounter this problem in the wild, but getting a PCIe GPU on to that platform is a different story altogether. There’s\nalso a fun quirk where Radeon GPUs on ARM platforms hide all write-combine memory to instead be write-back but we’ll talk about that another time.\n\n# Looking towards a brighter future\n\nAfter that marathon of an article we hope you have a better understanding of some of the challenges that emulating the x86-TSO memory model brings. Where we started with ARMv8.0 as a minimum spec and where the hardware has provided dramatic improvements over the years in nothing short of astounding. While not all of the edge-cases are yet resolved at the architecture level, it looks like there is a genuine commitment across the ecosystem for trying to improve the worst cases. We have various vendors solving some parts of the problem and moving the needle forward for better compatibility. Maybe in another decade as we look back at this time we’ll laugh about the problems we were encountering now, while enjoying some quality x86 games that will never see a port to ARM hardware. Keeping the legacy of the PC gaming ecosystem alive, regardless of where we might end up playing it.","body_html":"<p>Welcome to the first feature article on our site. We’re going to cover an ongoing problem with x86 emulation that affects every application that we\nemulate. This comes down to a single over-arching term that has wide-reaching ramifications; Emulating the <a href=\"https://en.wikipedia.org/wiki/Processor_consistency#Similarity_to_SPARC_V8_TSO,_IBM-370,_and_x86-TSO_memory_models\" rel=\"nofollow ugc noopener\">x86 Total Store Ordering memory model\n(x86-TSO)</a>.</p>\n<p>The problems with emulating this memory model on the <a href=\"https://en.wikipedia.org/wiki/Consistency_model#Weak_ordering\" rel=\"nofollow ugc noopener\">weak ordering memory model</a>\nthat ARM defines is multi-faceted and covers multiple issues. We’re going to go over all the problems that we can encounter and the ways we solve\n(or in some cases can’t solve) in this article. Get yourself a snack and a warm drink to enjoy, this is going to be a long one.</p>\n<h1 id=\"what-exactly-is-x86-tso\">What exactly is x86-TSO?</h1>\n<p>Before diving in to how we work around the x86 memory model problem, we need to first discuss exactly what it is. A memory model is a set of rules for how memory accesses in a system behave in relation to each other. The rules will dictate how loads and stores interact in a single-threaded or a multi-threaded environment. There’s a handful of popular memory memory models implemented in various forms of hardware, but the two we care about today is ARM’s relaxed (or weak) consistency model, and the x86 variant of Total-Store-Ordering consistency model. These two models are basically the two extremes of the spectrum; where ARM is the most relaxed, allowing significant hardware optimizations; and x86 is the most strict, enforcing a very strong coherency model that doesn’t allow a lot of room for optimization. One thing to be careful about when discussing memory models is the difference between consistency and atomicity. While these are related, they are not the same nor guaranteed in all cases.</p>\n<p>The best way to explain how the differences in memory models work is to start with how x86 handles this. With TSO being very strict in how it operates, the programmer can assume that when a memory store occurs, that this will be coherently visible to all other processors in the system. This additionally means that when a memory load occurs, all stores before it “logically” will have been completed, or at least visible. This matches programmer expectations, you write to memory, it becomes visible as at the point of writing, as this is intuitive to think about when programming. The stores are effectively ordering the visibility of the loads, thus the name of the model. There’s a bit of nuance with how this operates but isn’t strictly necessary to understand.</p>\n<p>The weak memory model that ARM has is a bit less intuitive about how it operates. By default the regular memory loads and stores that ARM uses aren’t strictly coherent across processors in your system, allowing the CPU to operate more efficiently most of the time. When a store instruction executes, that piece of memory (the cacheline) isn’t immediately visible to other processors in the system. Saving on precious power and efficiency because it’s expensive in hardware to invalidate other core’s cachelines, or allow them to snoop another processor’s caches. Relatedly if a processor is loading data from memory that another processor has written to, it’s not guaranteed that this load will even see this updated memory. This sounds like it would cause some significant problems in a multi-threaded application right? Older versions of ARM (ARMv7 and older) used a memory barrier instruction to ensure ordering, which had significant performance implications.</p>\n<p>To get around this limitation of consistency, ARM also introduced load-acquire, and store-release memory instructions. In C++ parlance this maps to\n<strong>std::atomic’s</strong> <em>memory_order_acquire</em> and <em>memory_order_release</em> definitions respectively. In ARM’s terminology, these instructions also aren’t\n<strong><em>technically</em></strong> considered to be atomic operations, but programmers conflate the two. FEX has used the terms atomic-load and atomic-store to mean\nthe same thing! The distinction <em>usually</em> doesn’t matter, but when discussing these topics it may be better to be pedantic about it.</p>\n<p>The primary use case for these instructions is to force memory ordering between these class of instructions. ARM calls this the “Release Consistency sequentially consistent (RCsc)” model. Without getting too far in to the weeds about how this model operates, the basic gist is that the load-acquire instructions must be observed sequentially without reordering, and the store-release instructions must as well while fulfilling “barrier-ordered-before” semantics. Removing the costly memory barrier instruction required in older ARM architecture versions.</p>\n<h1 id=\"the-humble-beginnings-of-armv8-0-a\">The humble beginnings of ARMv8.0-a</h1>\n<p>This is the premise of where we start in ARMv8.0-a when we’re emulating the x86-TSO memory model. We make all x86 memory loads turn in to ARM’s\n<strong>load-acquire</strong> instructions, and x86 memory stores turn in to <strong>store-release</strong> instructions. This gives FEX effectively the same memory semantics\nas x86, although we are actually being <em>more</em> strict than what is necessary. This is because we had no middle-ground which exactly matches behaviour.\nAs one might think, it is <em>exceedingly</em> costly to emulate TSO wth this instructions and we have microbenchmarks that can show this.\nAs ARM CPUs weren’t designed to have these relatively rare acquire/release instructions suddenly become the vast majority of instructions executed.</p>\n<p>First let’s start with something easy and use a microbenchmark that is fairly nice to the hardware. No tricky edge-cases, just accessing memory in in the common case. This gives us some baseline numbers for what the best-case situation should be.</p>\n<p>Let’s break down this graph as it tells us a few interesting stories. The Load and Store columns of each machine is representing our baseline performance number that our hardware should be attempting to achieve. These aren’t trying to max out the memory bandwidth of each system, but do the same amount of work for each type of operation. If we turn our attention to the acquire-load results, we can see that out of the five CPUs tests, three of them have their performance hindered quite a bit by using acquire-loads! Additionally we can see that the AmpereOne CPU has release-store instructions that are strikingly low compared to the other results, and the M1 Acquire/LRCPC load instructions are quite a bit lower than the baseline as well.</p>\n<p>The AmpereOne results in particular showcase how bad this legacy path can get. These instructions were never designed to be used this way. Using acquire-release semantics for every load for x86 emulation actually imposes some really strict limitations on ARM CPUs in that the load instructions can no longer be ordered around each other at all. So when you have millions of them in flight per second, the performance isn’t really expected to be good. But because these are the only instructions we had with ARMv8.0-a, it’s what we had to use. While Cortex-X4 and Cortex-X925 have amazing performance for these, you can see how the Oryon-3 has deprioritized their importance.</p>\n<h2 id=\"where-do-we-go-from-here\">Where do we go from here?</h2>\n<p>Let’s take a closer look at the LRCPC-load instructions, which is mandatory since ARMv8.3. This extension adds a bunch of new load instructions to the ARM ISA and adds a new memory model on top of\nARM’s <em>RCsc</em> model from before. This new “Release Consistency processor consistent (RCpc)” memory model is what we’ve been wanting! This extension is\ndesigned around the requirements that x86 emulation requires, and is expected to get utilized heavily on hardware that implements it. As you can see from the\ngraph, almost all of the platforms have their LRCPC-loads matching their regular loads in performance.</p>\n<p>With this new extension that is mandated by newer ARM versions, we basically get <strong>solved</strong>\nmemory performance. At least according to this microbenchmark that seems to be the case. Once FEX detects this extension we stop using Acquire-Load instructions\nentirely and switch over to LRCPC-Load instead. But what’s going on with that Apple M1 result..?</p>\n<p>This is where we need to commend Apple’s path towards solving this problem. With their Apple Silicon processors they directly added support for the\nx86-TSO memory model. When the CPU feature is toggled, their <strong><em>regular</em></strong> load/store ARM instructions change behaviour to match what x86 requires. They went\nthis route knowing that they will need a high performance solution for their hardware when switching to the ARM ecosystem exclusively.\nThis is why on their hardware the LRCPC-load instructions are actually aliases of their acquire-load instructions, because their x86\nemulator doesn’t even use these instructions! Because they implement the x86-memory model, they just use regular load/store instructions, which can be seen in our\nmicrobench results as indiscernable performance overhead. To be fair to the other platforms, this thread-wide TSO mode toggle does have some\nperformance impact, we just don’t see it here. When FEX detects this CPU feature from <a href=\"https://asahilinux.org/\" rel=\"nofollow ugc noopener\">Asahi Linux</a> we will also enable this\nand get the “free” performance improvement. A potential concern is that when jumping between x86 emulation and ARM code, that the ARM code\nwill pay unnecessary overhead due to all its accesses being TSO now. While this is a reasonable concern, the amount of ARM native code executing under\nemulation approaches 0%. As a developer, you don’t care about 1% of memory accesses becoming 10% slower, you care about 99% of accesses becoming 15% of\nthe “ideal” (As shown in AmpereOne results).</p>\n<p>As a note, we think a TSO mode is the best path forward for ensuring high performance x86 emulation on the platform. Because this ensures that every memory\naccess instruction behaves how we want or expect. This is shown with the official <strong>FEAT_LRCPC</strong> extension actually having three versions that\napply bandages to the implementation each time.</p>\n<ul><li><strong>FEAT_LRCPC</strong> - Adds basic GPR TSO load instructions</li><li><strong>FEAT_LRCPC2</strong> - Adds small offset immediate to TSO load instructions</li><li><strong>FEAT_LRCPC3</strong> - Adds basic vector and stack-based TSO load &amp; store instructions</li></ul>\n<p>Even with these three extensions, there is edge-case behaviour that can’t be emulated as nicely as if we had a TSO hardware toggle. We are expecting there to be additional extensions versions as time goes on, trying to fix some of the additional problems we’ll discuss later in the article.</p>\n<h1 id=\"i-thought-accessing-memory-was-the-easy-bit\">I thought accessing memory was the easy bit?</h1>\n<p>In the previous section, we were being nice to the ARM hardware and playing along with the underlying hardware’s alignment requirements to get a\nbaseline for what the performance should look like. When emulating x86 although, we run face first in to a glaring problem right from the start. Your\nfavourite x86 applications don’t care about alignment! They’ll access memory however they please, crossing cacheline granularities, doing atomics that\naren’t aligned. You think of the alignment problems, these games are doing it. This problem is so bad that we have a term associated with it, called <strong>split-locks.</strong>\nThese are such a big deal that even the Linux kernel will capture when these occur and slow down games when they do it! Causing many gamers to tinker\nwith kernel options to avoid the slowdown!</p>\n<p>But we aren’t going to talk about full on split-locks yet, let’s get started with just load-store instructions in an environment that doesn’t care\nabout alignment. x86 makes certain guarantees to the programmer; if you do a load-store and it is inside of a cacheline then that load-store will be both atomic and still match the coherency model as described before. However, to be a\nlittle bit nice to the hardware developers, if the load-store <em>does</em> cross a cacheline, the data isn’t atomic and other threads can and will see it\ntear. So the programmer needs to be careful as a basic load-store is not a split-lock.</p>\n<p>The problem with emulating these basic accesses with load-acquire/store-release is that ARMv8.0 requires what is known as <strong>natural alignment</strong>.\nThis means that for whatever size of data being accessed, the offset in memory must match the size. So for an 8-byte access, it must be at offsets; 0,\n8, 16, 24, etc. This works well for native ARM applications, but what happens when we don’t obey natural alignment requirements? For ARM, this means\nthe instruction with raise an <strong>alignment fault.</strong>\nThe hardware validates that the alignment requirements are fulfilled and if they are not then the CPU will fault. This usually results in a crash but\nFEX does special handling.</p>\n<p>Inside of FEX’s JIT mechanism we keep track of memory load-store instructions that are emulating the\nx86 load-stores. When we know that a load-store can cause an alignment fault we have what is known as a <strong>patchpoint</strong> in the code. For load-store\ninstructions, this shows up as a <strong>NOP</strong> instruction either before or after the load-store. When a alignment fault occurs as one of these patchpoints, FEX\nwill capture the fault, patch the code from a load-acquire/store-release instruction to a <strong>basic</strong> equivalent load-store, and wraps the instruction\nin a data memory barrier. Then it continues executing!</p>\n<p>Before then after patching</p>\n<p>That entire discussion from before about how ARMv8.0-a added these new fancy load-acquire, store-release instructions? We immediately fall back to the classic memory barrier instruction instead when alignment behaviour doesn’t match. Our previous chart didn’t show this bad case, so let’s bring in some fresh data.</p>\n<p>Oh, that’s a lot of data to sift through. While again good to see how far away the hardware is from the “optimal” path while emulating TSO, it’s not what we care about here. It is interesting to note that this microbench doesn’t showcase much of a difference between aligned and unaligned for regular load/stores so we just calculated an average between the two. We’ll be removing the x86 CPU and the regular load-store data from the ARM columns, as these aren’t the common FEX paths. This way we’ll have a more targeted view about how badly unaligned memory accesses hurt under emulation.</p>\n<p>Now that we have a much more reasonable graph of data, let’s walk from left to right on this and discuss what is going on.</p>\n<h2 id=\"ampereone\">AmpereOne</h2>\n<p>This one is pretty interesting, both the aligned and unaligned load instructions are roughly equivalent and fall within noise. This means that even though the unaligned loads are getting hit with a data memory barrier penalty, the CPU just handles it. This might be the case that the benchmark is bottlenecked by other things, considering how much lower the performance is compared to other platforms.</p>\n<p>Meanwhile the store side is not looking to be in a good shape even without unaligned. It nearly isn’t visible on the chart! When hitting unaligned stores we’re looking at ~8.5% of a performance hit, but because we are already starting so low it is hard to notice. This is also in stark contrast to regular store instructions getting ~28GB/s in this bench.</p>\n<p>The only conclusion we can come to here is that Ampere is optimizing for some server class workload and doesn’t really match consumer hardware behaviour. It’s an interesting datapoint, but our users aren’t typically running games on this class of hardware.</p>\n<h2 id=\"cortex-x4\">Cortex-X4</h2>\n<p>This is a highly popular CPU core that is living inside the <a href=\"https://en.wikipedia.org/wiki/Steam_Frame\" rel=\"nofollow ugc noopener\">Qualcomm Snapdragon 8 Gen 3</a>. We only tested\nthis one core from the SoC to not overwhelm the chart with data. Quite a large number of handhelds ship with this so it’s an interesting\ntarget. This CPU actually does <strong>surprisingly</strong> well considering it’s the only cellphone SoC on this list. Overall this core kind of falls in line\nwith what we would expect from it and the graph trends follow with the next-generation Cortex in that chart.</p>\n<p>The main topics for this CPU are that its aligned loads and stores are reasonably powerful, getting around 11.5GB/s and\n6.7GB/s respectively. What’s interesting is the performance falloff when it needs to deal with unaligned loadstores, hitting the <strong>DMB</strong> instructions\npenalizes the core roughly evenly between loads and stores at around 50% in this benchmark.</p>\n<p>This seems to imply that the CPU can keep a decent number of LRCPC-release loadstores in flight so the <strong>DMB</strong> instructions hurt more when they are\nencountered, but it isn’t causing world-ending performance. Just that a 50% performance hit due to alignment isn’t an amazing result.</p>\n<h2 id=\"cortex-x925\">Cortex-X925</h2>\n<p>Following up the X4, let’s stop by the <a href=\"https://www.nvidia.com/en-us/products/workstations/dgx-spark/\" rel=\"nofollow ugc noopener\">DGX Spark</a> and its X925 cores. Not only is this\na newer CPU core from ARM, it’s running on a system with dramatically more memory bandwidth. 273GB/s in the platform versus the previous 76.8GB/s. This\nmeans that we get fairly similar results to the X4 even, just the graph scales a little higher. Interestingly enough, the performance penalty for\nunaligned accesses roughly match the X4 even. Although it looks like the stores can recover a little faster, likely due to the faster memory helping\nout. No surprises here, just consistently matching performance across the generations.</p>\n<h2 id=\"oryon-3\">Oryon-3</h2>\n<p>This CPU core design is hot off the presses from Qualcomm. Linux support is still in the process of coming up but it already has a strong showing. The most interesting result from this actually comes from the fact that aligned LRCPC-load instructions are matching the performance of regular loads! That means in the case of a well-behaved application we can typically expect full performance. This continues onward to the release-store instructions being quite capable, although it doesn’t quite match regular stores with only 68% of the bandwidth. Not a bad showing in the slightest.</p>\n<p>This CPU also can’t escape from the penalty of unaligned LRCPC-release loadstores. The load side is roughly matching the ~70% performance penalty of the Cortex-X925, likely because the Snapdragon X2 Elite also has tons of bandwidth. But the store side actually gets off a little worse at ~43% of the performance. Even with these performance hits of unaligned accesses, this platform is actually faster than the aligned accesses from the Cortex offerings.</p>\n<p>One of the weird things about this platform is that it was advertised to have “Fully coherent 96KB 6-way L1 cache with 64B coherency granules.” Which to our reading implied that unaligned accesses should have dramatically less of a performance impact. Interesting… keep that in mind.</p>\n<h2 id=\"apple-m1\">Apple M1</h2>\n<p>This is the big one we need to talk about. This is the one that was a game changer, it was the “Apple moment.” It showed everyone that ARM was\nnot only feasible, it could be faster. These numbers on this chart are amazing and it’s the result of Apple sticking the TSO memory\nmodel directly in to their hardware. Instead of using LRCPC-release accesses for this one, we just enabled their TSO feature and the aligned\nversions basically match the unaligned version. Maybe a 5% performance hit on the stores? Compared to every other device on that chart, it’s\neffectively nothing. This primarily comes down to unaligned accesses no longer requiring <strong>DMB</strong> instructions to be backpatched in to the code, as the\nhardware just handles it directly.</p>\n<p>For us, this is what it means to take x86 emulation seriously on ARM and it really shows that Apple cared that their customers would have a good\nexperience running software both natively and emulated. They saw the problem and just solved it, making it go away.\nThat said, when the TSO mode <em>is</em> enabled, you do get a performance hit. Comparing to the previous graph it’s only getting 76% of the regular\nstore performance, and the load performance basically matches; that’s much more tolerable to bear when everything is so much faster.</p>\n<h2 id=\"wrapping-up-unaligned-lrcpc-release-accesses\">Wrapping up unaligned LRCPC/release accesses</h2>\n<p>Wrapping up this section, we need to talk about one of the performance improvements that all of these vendors actually support. This is an\nextension that ARM whipped up called <strong>FEAT_LSE2</strong> which all of these tested platforms implement. We previously talked about how acquire/LRCPC/release\nmemory accesses require natural alignment in order to not incur the wrath of the CPU raising alignment faults. ARM actually thought about\nthis problem and implemented this extension which helps x86 emulation (and probably other workloads). This extension loosens the alignment\nrequirements of not only acquire/LRCPC/release load store instructions, it <em>also</em> loosens the requirement for read-modify-write atomics!</p>\n<p>That sounds all well and good, but here’s the kick to the teeth: that means it only provides marginal performance gains for x86 emulation. This\nextension only loosens the alignment requirements to allow unaligned memory accesses inside of a 16-byte granule. Any access that crosses that 16-byte\ngranule still receives an alignment fault. x86 applications don’t really care about the alignment of their memory accesses, so we get\nunaligned accesses across the entire cacheline. It’s only read-modify-write atomics that <strong>try</strong> to avoid crossing a cacheline on x86!</p>\n<p>So thanks for the attempt, it’s nice to see, but it doesn’t really move the needle. Since we’re already talking about it, let’s dive in to those RMW atomics shall we?</p>\n<h1 id=\"oh-no-what-are-these-atomic-instructions\">Oh no, what are these atomic instructions?</h1>\n<p>Like most modern instruction sets, x86 supports atomic memory operations. These are instructions that execute an ALU operation on data in memory\natomically, allowing no intermediate state to be visible. In x86 terms this operates on memory that is both atomic and coherent, while ARM lets you\nchoose to be only atomic <em>or</em> both atomic and coherent. We touched on this briefly before but there is actually a difference between operating on data\natomically, and coherency of that data. What difference does it make?</p>\n<p>For all of the previous x86 memory model discussion we have been talking about the coherency implications of loads and stores being visible to other\nprocessors in the system. What we entirely glossed over is the atomicity requirements of these memory accesses. In the world of x86 a load or store\n<em>usually</em> completes atomically even when unaligned. This means that if you’re storing 8-bytes of data, and another thread is loading those 8-bytes in\na race condition it will never suddenly see a mix of the data from before the store and after the store. In ARM these atomicity guarantees are\n<strong>significantly</strong> weaker, meaning if you do an unaligned store instruction the specification of the ISA has zero guarantees about reading a <strong>tear</strong> in\nthe data. Thankfully for <strong>naturally</strong> aligned load-store instructions, ARM has a specification called <strong>“single-copy atomicity”</strong> which guarantees\nyou don’t get a tear for these accesses. Also good news; that <strong>FEAT_LSE2</strong> extension from before? It actually extends the\nsingle-copy atomicity guarantees to <em>any</em> unaligned access inside of a 16-byte granule! The downside is that x86 has single-copy atomicity\nguarantees across a full cacheline, so once again the extension still didn’t solve anything completely, just reduced the number of occurences.</p>\n<p>Enough about the differences in atomicity and coherency. Where’s the actual atomic instructions? What do they do? Starting in ARMv8.1-a, our ISA has gained instructions that mostly matches x86 atomic instructions in behaviour. Let’s just give the full list to show how they map directly in our JIT.</p>\n<div class=\"table-wrap\"><table><thead><tr><th><strong>x86</strong></th><th><strong>ARMv8.1-a</strong></th></tr></thead><tbody><tr><td>LOCK DEC</td><td>ldaddal</td></tr><tr><td>LOCK INC</td><td>ldaddal</td></tr><tr><td>LOCK NEG</td><td><a href=\"https://github.com/FEX-Emu/FEX/blob/e02953dc174531551219712df20355dcf4afc089/unittests/InstructionCountCI/FlagM/Atomics.json#L1368-L1381\" rel=\"nofollow ugc noopener\">???</a></td></tr><tr><td>LOCK NOT</td><td>ldeoral</td></tr><tr><td>LOCK ADC</td><td>ldaddal</td></tr><tr><td>LOCK ADD</td><td>ldaddal</td></tr><tr><td>LOCK AND</td><td>ldclral</td></tr><tr><td>LOCK OR</td><td>ldsetal</td></tr><tr><td>LOCK SBB</td><td>ldaddal</td></tr><tr><td>LOCK SUB</td><td>ldaddal</td></tr><tr><td>LOCK XADD</td><td>ldaddal</td></tr><tr><td>LOCK XOR</td><td>ldeoral</td></tr><tr><td>LOCK BTC</td><td>ldclralb</td></tr><tr><td>LOCK BTR</td><td>ldeoralb</td></tr><tr><td>LOCK BTS</td><td>ldsetalb</td></tr><tr><td>LOCK CMPXCHG</td><td>casal</td></tr><tr><td>CMPXCHG8B</td><td>caspal</td></tr><tr><td>CMPXCHG16B</td><td>caspal</td></tr></tbody></table></div>\n<p>Well would you look at that, we have a full list of the 18 atomic RMW operations and they basically map directly to some ARM instructions. Ignore the questionable one as it’s not used in real workloads and we would get far too in to the weeds talking about it. We have a pretty clear 1:1 mapping between the architectures, job’s done right? That’s the funny thing about x86 emulation, just because we have these instructions doesn’t mean we get to wire them up without problems. We spent all this time talking about how unaligned accesses can really hurt performance of regular loads and stores, this same problem also applies to RMW atomics!</p>\n<p>With this graph, we are looking at a single atomic instruction with its memory address landing somewhere within a cacheline. If we included all of the\ndata for all 18 atomic operations then this data would be even more overwhelming than it already is. All these atomic operations behave <em>roughly</em>\nequivalent so it would be redundant and wouldn’t matter for what we’re discussing here anyway. This is also the first graph in this post that is\nactually using logarithmic scaling, so when reading it make sure to understand that the performance difference from the fastest to slowest result is\non the scale of around 1000x.</p>\n<p>Starting with the x86 Zen processor on this graph; these are the results that our emulation should be striving to achieve. As we can see, if the access is\nfully contained within a cacheline then the latency of the instruction is the same at 1.44ns. This can be explained by x86 having “atomic cachelines”\nor “coherent cachelines”, where as long as an unaligned atomic operation stays within a cacheline then it roughly costs the same. This is a really\npowerful feature of x86 that has been supported for decades at this point so games end up relying on this heavily without even realizing it. The\nstand-out result for x86 is the final result that is crossing a 64-byte granule and taking ~660ns! That’s an amazingly slow result at ~458x slower\ncompared to the other results because this is finally the hardware using <strong>split-locks</strong>.</p>\n<p>We need to take a moment here to shout out an article that <a href=\"https://chipsandcheese.com/p/investigating-split-locks-on-x86\" rel=\"nofollow ugc noopener\">Chips and Cheese</a> wrote\nwhile we were preparing to write our article. They do a great deep dive in to why these <strong>split-locks</strong> are so dramatically slower and is worth the\nread if you’re unaware of how they work. Specifically we need to mention that x86 split-locks maintain the atomicity and coherency requirements of\nx86-TSO and will <strong><em>never</em></strong> tear the data even when crossing a cacheline. This is kind of nuts and we’ll explain this more later.</p>\n<p>Now for our ARM processors, let’s start with the natural alignment latency numbers. As we can see, all of our platforms perform fairly well but even\nthe latest cores don’t get anywhere near x86. Even our fastest ARM platform is ~3x the latency compared to x86; This directly impacts performance of\ngames but usually isn’t the direct bottleneck so it’s hard to measure exactly how much. Continuing onward to the next data point, we can actually\ncombine the results for 16-byte granule and 64-byte granule crossing with most of our ARM platforms. Due to how the ARM specification defines how\nunaligned atomics work, both of these results are roughly equivalent and FEX treats them the same as the x86 <strong>split-lock</strong> problem.</p>\n<p>We keep bringing up this split-lock problem but how exactly does FEX emulate them and what makes it so slow? “I thought Apple M1 added x86-TSO support\nin the hardware, why is it still slow?” If you recall how we brought up before that <strong>FEAT_LSE2</strong> introduced support for unaligned memory accesses within\na 16-byte granule; these split-lock operations end up hitting the same alignment problems as before but are dramatically slower. FEX\ncan’t backpatch any of these instructions to just do a <strong>DMB</strong> operation, so we cause an <strong>alignment-fault</strong> every time one gets executed. This means that\nwe do a kernel -&gt; userspace signal handler -&gt; kernel -&gt; original code dance. <em>every—single—time</em> one of this split-lock operations execute.\nJumping between kernel-space and userspace is slow on every platform and when you’re executing thousands of these per second it adds up very quickly.\nThis is why the emulation of these feature is so terribly slow on ARM.</p>\n<p>One ARM platform today actually partially resolved this problem although. The Oryon-3 CPU cores introduced what they advertised as “coherent\ncachelines” and we can see this in our microbenchmark results here. Just like with x86, if the atomic memory access in anywhere inside of the 64-byte\ncacheline, the performance matches the natural alignment version! This is a tremendous improvement that means the CPU is on par with x86 in\nfeature support until the point it tries to cross a cacheline. We need to applaud Qualcomm on implementing this feature, as it resolves a major\nperformance and correctness problem around split-locks for x86 emulation. The hardware still doesn’t support 64-byte <strong>split-locks</strong> so we still fall\ndown the FEX emulated path in that instance although.</p>\n<p>Continuing on to the Apple result; even though they added x86-TSO memory accesses to their hardware for some reason they neglected to implement full cacheline unaligned atomics like Oryon did. It seems like they should have expected this edge case to surface and implement it but that’s just speculation. This is why you can see the cross 16-byte granule behaving the same as other platforms even with the TSO hardware toggle enabled.</p>\n<p>You might have also noticed another little data quirk in the graph. We have an asterisk on the Cortex-X4 result in this benchmark and the performance\nof the unaligned atomics are dramatically faster than significantly newer CPUs. It is somehow managing to have only\n~209ns latency, while the X925 is latency is 1060ns; that’s a 5x perf improvement! How can this possibly be the case? This is actually some fun\n“special sauce” that is shipping on the platform we’re testing on, which is of course the <strong><a href=\"https://store.steampowered.com/sale/steamframe\" rel=\"nofollow ugc noopener\">Valve Steam\nFrame</a></strong>. Because Valve cares about the performance of their existing gaming catalogue, they are shipping a\n<a href=\"https://github.com/bylaws/linux/commit/7ae989a43ae7e3cb8007ac21c28dacc24c9d8320\" rel=\"nofollow ugc noopener\">kernel patch</a> that one of the FEX developers whipped up. This allows\nthe Linux kernel itself to handle the unaligned atomic without that slow dance with FEX and userspace, allowing it to be dramatically faster. If other\nplatforms want to ship this patch in the kernel then we recommend picking it up as and FEX will automatically start using it.</p>\n<p>Speaking of kernel intervention, we need to talk about how split-lock emulation is not actually quite correct under FEX due to limitations in the hardware. In order to implement this mandatory feature of x86 correctly, any time we do a 16-byte or 64-byte split-lock, the only way to handle it is to have the kernel implement the feature. Right now FEX implements this as a “best-effort” attempt that can actually tear the data in some cases. You’ll recall that before we said split-locks on x86 will never tear right? Not even the Oryon-3 with its “coherent cachelines” have resolved this problem yet.</p>\n<h1 id=\"what-do-you-mean-split-lock-is-mandatory\">What do you mean split-lock is mandatory?</h1>\n<p>Implementing split-lock emulation with today’s ARM hardware in a performant matter is actually really difficult to do. A naive implementation is to\nuse a global mutex and whenever a split-lock occurs we will ensure to acquire the mutex before doing the operation. This means that any\n<em>participating</em> split-lock operation will funnel through this mutex. This is correct except for the issue that any aligned atomic operation\nisn’t a split-lock and won’t participate. Due to the split-lock emulation code needed to be implemented as two 64-bit compare-exchange\noperations with each half straddling the granularity boundary, we can get a tear with a non-participating atomic still. A trivial example is one\nthread constantly modifying an atomic in the middle of the cacheline, and then another thread modifying <em>only</em> the integer on one half. This might sound\nlike a contrived example initially, but there are lock-less <a href=\"https://en.wikipedia.org/wiki/Linked_list\" rel=\"nofollow ugc noopener\">linked-list</a> implementations that behave exactly like this!\nDepending on which half the aligned thread is modifying, either the first or second CAS in the split-lock code will fail. If the first CAS fails, then\nthat’s safe and the code can retry, if the <em>second</em> CAS fails that means the data has torn and we can do nothing but hope it doesn’t corrupt data and\ncrash. This will entirely depend on the algorithm that the guest application is using so we don’t control it.</p>\n<p>An alternative approach that is completely untenable is to have the kernel track all processes and threads that are sharing memory with each other,\nthen when a thread needs to emulate a split-lock the kernel can halt <strong><em>every</em></strong> process that is sharing memory with that process, do the split-lock\nin isolation, and then restart the world. The performance implications of this approach aren’t viable. Applications and games can end up doing thousands or more\nsplit-locks per second and halting the world will have an intractable performance hit that is dramatically worse than even x86 native.</p>\n<p>If we want to ensure correctness in the emulation of split-locks FEX needs to have hardware support in some form to support these. Although we’re not\nsaying that all atomic operations should now support split-locks like x86, that would also not be viable. The good news is that ARM actually has an\nextension for this that does exactly what we want. ARM has an extension call <strong><a href=\"https://en.wikipedia.org/wiki/Transactional_memory\" rel=\"nofollow ugc noopener\">Transactional Memory\nExtension</a></strong> that could solve our problem. This extension allows our code to do some number of\noperations inside of a transactional region, then commit that work atomically; if the commit operation fails, then we can simply retry. The downside\nof this extension? ARM has officially deprecated the extension and no one ever shipped it. This is likely for the best as the x86 version of the\nextension has had an abundance of problems that caused it to be disabled on many platforms.</p>\n<p>So we need something else to emulate split-locks correctly. For a solution that we believe works for both FEX needs and ARM vendor needs, we have come\nup with the idea that a 128-bit <strong>CASP</strong> instruction can be given the ability to have each half of the CASP perfectly straddle\nthe atomic granule boundary, 64-bits on the lower half, and 64-bits on the upper half. Then <em>only</em> in that case does the instruction not raise an\nalignment-fault and tries to do the CAS operation. This works because x86 only has up to 64-bit unaligned atomic operations, so both halves of the\noperation can always be fully enclosed by our single operation.</p>\n<p>But you may be asking yourself, “how is this any better than the hardware just supporting split-locks?” That’s a good thought and we need to be\ncareful with the how exactly we describe this operation. For x86 their atomic operations must <em>always</em> succeed without tear. For our emulated\napproach, we can have this ARM <strong>CASP</strong> instruction fail safely and then we can try again. This is one of the benefits of CAS is that\nthe operation can fail for <em>any</em> reason and it must be tried again. The instruction then also returns the data that it loaded from memory in that time\nso the program has the latest up to date memory. This is an important distinction since that means FEX can retry the <strong>CAS</strong> operations infinite times\nuntil it inevitably succeeds! This is a benefit of ARM LL/SC architecture that basically allows this to work. A tricky thing is that the hardware does\nneed to guarantee forward progress at <em>some</em> point but it already has support for that for other reasons so it’s completely viable! The only newly\nadded failure mode to the <em>CAS</em> instruction is purely if one of the two cachelines got acquired by another core before it could do the full operation.\nEven if the hardware still requires up to a couple thousand cycles to guarantee forward progress, that basically matches x86 behaviour.</p>\n<p>We think this would be the best way forward for x86 emulation of split-locks on ARM platforms, but we’re not hardware architects so all we can do is complain and hope someone solves it for us. We’ll leave the split-lock discussion there for now so we can move on to another interesting problem.</p>\n<h1 id=\"wait-uncached-memory-needs-to-work\">Wait, uncached memory needs to work?</h1>\n<p>Before we get in to this topic we need to talk about the term “uncached” because it can mean a couple of things depending on your view of the\nworld. For the purposes of this article, we are using Vulkan terminology because we care about games primarily. In Vulkan terms we have\n<em>VK_MEMORY_HOST_CACHED_BIT</em> which means that the host CPU caches this memory. The lack of this bit is what we care about here, and what we refer to as\n“uncached.” As for what this means to the memory subsystem, it gets a little more complicated than you would think. In particular when the memory is\nliving on a GPU, potentially over PCIe, when the memory is “uncached” it will also typically (but not always!) also gain the flag\n<em>VK_MEMORY_HOST_COHERENT</em>. This means that because of the uncacheable property of the memory, the CPU and GPU always have a coherent world memory view\nwith each other.</p>\n<p>For the CPU this typically means the memory can be mapped up to three ways. When asking for “cached” memory, this typically has a memory type of\n<a href=\"https://en.wikipedia.org/wiki/Cache_%28computing%29#WRITE-BACK\" rel=\"nofollow ugc noopener\">Write-back</a> which is also what regular memory mapping types are. “uncached” mapping\ncan be either <a href=\"https://en.wikipedia.org/wiki/Uncacheable_speculative_write_combining\" rel=\"nofollow ugc noopener\">Write-Combine</a> or “Strong Uncacheable”. The “Strong Uncacheable”\nimplementation is basically non-existant for userspace applications so we can ignore that for today’s discussion. This limits us to effectively <strong>WB</strong>\n(cached) and <strong>WC</strong> (uncached) memory types. Cached is what games typically use for staging buffers, and then uncached is what we use when passing data directly\nto the GPU.</p>\n<p>This is code-ified in many game engines that if you don’t expose support for uncached buffer types then some don’t work. This comes down to a\nbehaviour detail around the differences of <a href=\"https://en.wikipedia.org/wiki/Unified_memory_architecture\" rel=\"nofollow ugc noopener\">UMA</a> systems like APUs and PCIe GPUs. UMA\nsystems will typically expose the ability to allocate memory that is cached, coherent, and GPU visible. Where PCIe GPUs can’t guarantee that behaviour\nso game developers need to either use a staging buffer and an async copy of the data over to the GPU, or use “uncached” memory to very carefully\nshuffle the data over to the GPU through PCIe. Because of how ubiquitous PCIe is with PC gaming, some engines won’t even do UMA specific code\npaths and will do the uncached approach regardless!</p>\n<p>With that little introduction out of the way for what uncached means for us. Let’s bring up a benchmark for how fast cached memory is on some UMA Snapdragon systems. This will let us get a baseline for how the performance should be regularly.</p>\n<p>For both the <strong>Steam Frame</strong> and <strong>Snapdragon X2 Elite</strong> these are some really good results. As we would expect, the Oryon-3 platform has more memory\nbandwidth so it is able to scale higher in the chart, but both are hitting dozens of gigabytes per second in their results. This graph sets a good\nbaseline for what “normal” <strong>write-back</strong> memory can achieve. Let’s now show uncached results to see the performance differences.</p>\n<p>There’s some strange things happening here so we had to use logarithmic again on this graph. Let’s talk about the good first that has shown up.\nDue to uncached memory buffers being write-combine, we can see that the regular stores for our ARM platforms match the cached benchmark\nresults. This comes down to write-combine memory using what is coined as <a href=\"https://en.wikipedia.org/wiki/Write_combine_buffer\" rel=\"nofollow ugc noopener\">write combine buffers</a>\nthat actually <em>very</em> temporarily keep around a cacheline of data so that write-combine can burst a cacheline of memory at a time. Interestingly enough\nit looks like the Zen 4’s WCB can’t quite keep up with cached, but considering this is expected to be going over a PCIe bus it’s probably fine.</p>\n<p>Now let’s get in to the really ugly results that we have here. Starting off with the easier to explain is the load bandwidth from write-combined\nmemory is abysmal on all platforms tested. If we’re using Zen as our baseline for performance, then our regular load instructions are ARM are winning,\nbut the LRCPC loads are worse. What’s going on here? This is a quirk of how write-combined memory operates, because it is uncached our load\ninstructions are required to go out to system memory for every single access to maintain semantics. Then when we add LRCPC-loads on top of that, it\njust compounds the problem even further. But the worst case out of all of this is just how badly the store performance is, compared\nto the performance that Zen gets on the stores, this is basically a showstopper. Up to <strong>816x worse</strong> bandwidth! We had games like <a href=\"https://store.steampowered.com/app/1030300/Hollow_Knight_Silksong/\" rel=\"nofollow ugc noopener\">Hollow\nKnight: Silksong</a> and <a href=\"https://store.steampowered.com/app/1962700/Subnautica_2/\" rel=\"nofollow ugc noopener\">Subnautica\n2</a> run at less than 1FPS because of this performance cliff.</p>\n<p>As we were saying above, when there are PCIe GPUs in the mix then games will need to use uncached memory to pass data to the GPU. When emulating x86\ngames on platforms with a dedicated PCIe GPU then we are in an unwinnable situation and we are guaranteed to run dramatically slower. Remember how ARM\nhas added the family of <strong>FEAT_LRCPC1/2/3</strong> extensions from before to improve x86 memory model emulation? This is what happens when we hit an\nedge-case that isn’t supported. All of these extensions add new instructions to handle loading memory using x86-TSO memory model semantics but none of\nthem solve storing to write-combine memory with x86-TSO semantics. All the way from ARMv8.0-a our store instructions use the regular <code>store-release</code>\ninstructions regardless of the backing memory type. The only way for FEX to work around this problem is to selectively disable TSO-emulation when it\nbecomes an issue, so x86 emulation platforms with PCIe GPUs will always be a worse experience than UMA. At least until we get another <strong>FEAT_LRCPC4</strong>\nor similar to resolve the issue.</p>\n<p>For users on UMA systems then rejoice, there’s a workaround for gaming that we use to improve performance. Because we know when a platform supports\ncache-coherent CPU and GPU combinations, we can have the video driver <strong><em>always</em></strong> use cached buffers and never encounter this problem.\nNVIDIA already does this on their Tegra platforms, Snapdragon has been supporting this since at least Adreno 600 class GPUs, and there are many\nMali platforms where this is also the case. We have a <a href=\"https://gitlab.freedesktop.org/mesa/mesa/-/merge_requests/41323\" rel=\"nofollow ugc noopener\">Adreno Turnip</a> patch that\nensures when FEX is running, we never hit uncached memory for platforms that support it. A funny thing is that since Asahi users have a hardware TSO\nbit, they just naturally don’t encounter this problem in the wild, but getting a PCIe GPU on to that platform is a different story altogether. There’s\nalso a fun quirk where Radeon GPUs on ARM platforms hide all write-combine memory to instead be write-back but we’ll talk about that another time.</p>\n<h1 id=\"looking-towards-a-brighter-future\">Looking towards a brighter future</h1>\n<p>After that marathon of an article we hope you have a better understanding of some of the challenges that emulating the x86-TSO memory model brings. Where we started with ARMv8.0 as a minimum spec and where the hardware has provided dramatic improvements over the years in nothing short of astounding. While not all of the edge-cases are yet resolved at the architecture level, it looks like there is a genuine commitment across the ecosystem for trying to improve the worst cases. We have various vendors solving some parts of the problem and moving the needle forward for better compatibility. Maybe in another decade as we look back at this time we’ll laugh about the problems we were encountering now, while enjoying some quality x86 games that will never see a port to ARM hardware. Keeping the legacy of the PC gaming ecosystem alive, regardless of where we might end up playing it.</p>","headings":[{"level":1,"text":"What exactly is x86-TSO?","id":"what-exactly-is-x86-tso"},{"level":1,"text":"The humble beginnings of ARMv8.0-a","id":"the-humble-beginnings-of-armv8-0-a"},{"level":2,"text":"Where do we go from here?","id":"where-do-we-go-from-here"},{"level":1,"text":"I thought accessing memory was the easy bit?","id":"i-thought-accessing-memory-was-the-easy-bit"},{"level":2,"text":"AmpereOne","id":"ampereone"},{"level":2,"text":"Cortex-X4","id":"cortex-x4"},{"level":2,"text":"Cortex-X925","id":"cortex-x925"},{"level":2,"text":"Oryon-3","id":"oryon-3"},{"level":2,"text":"Apple M1","id":"apple-m1"},{"level":2,"text":"Wrapping up unaligned LRCPC/release accesses","id":"wrapping-up-unaligned-lrcpc-release-accesses"},{"level":1,"text":"Oh no, what are these atomic instructions?","id":"oh-no-what-are-these-atomic-instructions"},{"level":1,"text":"What do you mean split-lock is mandatory?","id":"what-do-you-mean-split-lock-is-mandatory"},{"level":1,"text":"Wait, uncached memory needs to work?","id":"wait-uncached-memory-needs-to-work"},{"level":1,"text":"Looking towards a brighter future","id":"looking-towards-a-brighter-future"}]}}