{"article":{"slug":"performance-improvements-in-net-11","title":"Performance Improvements in .NET 11","subtitle":null,"summary":"Take a tour through hundreds of performance improvements in .NET 11.","content_type":"blog_post","language":"en","canonical_url":"https://devblogs.microsoft.com/dotnet/performance-improvements-in-net-11/","author":{"name":"Stephen Toub","url":null,"person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"Microsoft","url":"https://devblogs.microsoft.com/dotnet/","listing_slug":null,"listing":null},"topics":[{"name":"Programming","slug":"programming","url":"https://listedarticles.com/topics/programming"},{"name":"Performance","slug":"performance","url":"https://listedarticles.com/topics/performance"}],"about_listings":[],"cover_image_url":"https://devblogs.microsoft.com/dotnet/wp-content/uploads/sites/10/2026/09/net11perf.webp","license":"all-rights-reserved","word_count":37742,"reading_minutes":164,"published_at":"2026-09-15T12:00:00.000Z","added_at":"2026-09-17T12:14:16.290Z","updated_at":"2026-09-17T12:14:16.290Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":false},"profile_url":"https://listedarticles.com/articles/performance-improvements-in-net-11","markdown_url":"https://listedarticles.com/articles/performance-improvements-in-net-11.md","example":false,"citation":"Stephen Toub, Microsoft. \"Performance Improvements in .NET 11.\" 15 Sept 2026. https://devblogs.microsoft.com/dotnet/performance-improvements-in-net-11/ (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://devblogs.microsoft.com/dotnet/performance-improvements-in-net-11/"},"body_markdown":"> **Syndication note:** This copy is truncated to fit the index size limit. Read the full article at the [canonical URL](https://devblogs.microsoft.com/dotnet/performance-improvements-in-net-11/).\n\nBefore television shows like *The Office* and *Parks and Recreation* cemented the mockumentary in the minds of millions, there was Christopher Guest. He didn’t invent the genre, but he’s widely recognized as one of its most influential practitioners, and for my money, there’s none better. I’ve watched *Waiting for Guffman* and *Best in Show* more times than I can count. But the one that has stuck with me the most, the one I quote at the slightest provocation, is *This Is Spinal Tap*.\n\nIf you’ve seen it you already know where this is going (and if you haven’t, you now have weekend plans). The film is a fictional documentary about an aging English rock band named Spinal Tap, whose members are everything we picture when we picture over-the-top rock stars. In one of its more memorable scenes, the guitarist (Nigel) gives the filmmaker (Marty) a tour of his most prized gear, in particular showing off an amplifier unlike any other: its dials don’t stop at ten. That leads to what might be the single most quoted exchange in the entire movie:\n\n**Nigel:** “You see, most blokes, you know, will be playing at ten. You’re on ten here, all the way up, all the way up, all the way up, you’re on ten on your guitar. Where can you go from there? Where?”\n\n**Marty:** “I don’t know.”\n\n**Nigel:** “Nowhere. Exactly. What we do is, if we need that extra push over the cliff, you know what we do?”\n\n**Marty:** “Put it up to eleven?”\n\n**Nigel:** “Eleven. Exactly. One louder.”\n\n\nThis is .NET 11. It’s one louder, with another year’s worth of performance work having gone into making the runtime and libraries that much faster. Of course, the premise of Nigel’s special amplifier is ludicrous, as is exemplified in the subsequent few lines of dialog:\n\n**Marty:** “Why don’t you just make ten louder and make ten be the top number and make that a little louder?”\n\n**Nigel:** (pauses) “…these go to eleven.”\n\n\nIn contrast, .NET 11 is actually one higher, one louder. The sections that follow are full of real improvements. A bounds check removed, an allocation that no longer happens, a lock that isn’t taken, a loop that runs in fewer cycles than it did a year ago, a comparison folded to a constant here, a redundant check hoisted out of a loop there, a couple of instructions fused into one, a syscall sidestepped, an array copy handed off to SIMD, and on and on. That’s how real performance work goes, accumulating gain after gain, each compounding on the last, until the whole thing is measurably, provably louder. And so, in this post, as I’ve done in past years with .NET 10, .NET 9, .NET 8, .NET 7, .NET 6, .NET 5, .NET Core 3.0, .NET Core 2.1, and .NET Core 2.0 before it, we’ll take an unhurried tour through hundreds of them.\n\nThis is a long one. It’s meant to be. Grab your hot beverage of choice, settle in, and let’s turn it up.\n\n## Benchmarking Setup\n\nAs in previous years, the post is chock full of micro-benchmarks that demonstrate the individual improvements. Almost all of them use BenchmarkDotNet, and each is written to be self-contained so you can try it out yourself.\n\nStart by ensuring you have both .NET 10 and .NET 11 installed (most of the benchmarks compare the same code running on both versions) and create a new console project in a fresh `benchmarks` directory:\n\n```\ndotnet new console -o benchmarks\ncd benchmarks\n```\nReplace the contents of the generated `benchmarks.csproj` with the following, which multi-targets both versions so that BenchmarkDotNet can build for each:\n\n```\n<Project Sdk=\"Microsoft.NET.Sdk\">\n  <PropertyGroup>\n    <OutputType>Exe</OutputType>\n    <TargetFrameworks>net11.0;net10.0</TargetFrameworks>\n    <LangVersion>preview</LangVersion>\n    <ImplicitUsings>enable</ImplicitUsings>\n    <Nullable>enable</Nullable>\n    <AllowUnsafeBlocks>true</AllowUnsafeBlocks>\n    <ServerGarbageCollection>true</ServerGarbageCollection>\n    <SystemPackageVersion Condition=\"'$(TargetFramework)' == 'net10.0'\">10.0.12</SystemPackageVersion>\n    <SystemPackageVersion Condition=\"'$(TargetFramework)' == 'net11.0'\">11.0.0-rc.1.26425.128</SystemPackageVersion>\n  </PropertyGroup>\n  <ItemGroup>\n    <PackageReference Include=\"BenchmarkDotNet\" Version=\"0.16.0-preview.1\" />\n    <PackageReference Include=\"System.IO.Hashing\" Version=\"$(SystemPackageVersion)\" />\n    <PackageReference Include=\"System.Runtime.Caching\" Version=\"$(SystemPackageVersion)\" />\n    <PackageReference Include=\"System.Numerics.Tensors\" Version=\"$(SystemPackageVersion)\" />\n  </ItemGroup>\n</Project>\n```\nFor a given benchmark to test, copy its complete contents over everything in `Program.cs` and then run it. Each benchmark includes as a comment at the top the exact command to use. In most cases, it’s:\n\n`dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0`\nwhich builds in Release and runs the benchmark against both .NET 10 and .NET 11, emitting a side-by-side comparison. The other common form, used when a benchmark is comparing two coding approaches on a single runtime (rather than the same code across two runtimes) is:\n\n`dotnet run -c Release -f net11.0 --filter \"*\"`\nThe usual disclaimer applies: these are micro-benchmarks, many measuring operations so short that a blink would miss them. Your results will vary with your hardware, OS, runtime configuration, what else your machine happens to be doing at that exact moment, and whether Mercury is in retrograde.\n\nEvery line of managed code ultimately ends up at the just-in-time compiler, so let’s start there.\n\n## JIT\n\nOf all the places to improve .NET’s performance, few have as broad an impact as the just-in-time (JIT) compiler. C#, F#, and Visual Basic are typically compiled first to intermediate language (IL), and the JIT ultimately turns that IL into the native instructions the CPU executes. A JIT improvement can therefore benefit application and library code wherever the optimized pattern occurs, often with no source changes or recompilation of the application itself. Even removing a single instruction or proving one check unnecessary can add up when the code is on a very hot path.\n\n### Deabstraction\n\nWe as developers love our abstractions. They let us write clean, reusable, object-oriented code, but we don’t want to pay for every abstraction at run time. The runtime can often undo an abstraction when it proves the effects aren’t observable. It can look at a virtual call and determine which concrete method it’ll invoke, look at a heap allocation and recognize that the object never leaves the current stack frame, or look at an interface cast and reuse a type fact already established earlier in the method. This process is called “deabstraction.” .NET has improved steadily in this area for years, and that continues in .NET 11.\n\nEvery time you write `interface` in C#, you’re creating a contract, a promise that any type implementing that interface can be substituted for any other. That flexibility is enormously valuable because, for example, it’s what lets us write `IEnumerable<T>` and have it work equally well over arrays, lists, other collections, LINQ, custom iterators, and so on. But the CPU doesn’t know anything about these contracts; it just knows how to execute instructions. Turning “call whatever method this interface reference points to” into actual machine instructions requires special machinery. Consider this example:\n\n```\n// dotnet run -c Release -f net11.0 --filter \"*\"\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Runtime.CompilerServices;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private Animal _animal = Environment.TickCount >= 0 ? new Dog() : new Cat();\n    [Benchmark]\n    public int Speak() => _animal.Speak();\n    public abstract class Animal\n    {\n        public abstract int Speak();\n    }\n    private sealed class Dog : Animal\n    {\n        [MethodImpl(MethodImplOptions.NoInlining)]\n        public override int Speak() => 1;\n    }\n    private sealed class Cat : Animal\n    {\n        [MethodImpl(MethodImplOptions.NoInlining)]\n        public override int Speak() => 2;\n    }\n}\n```\nAt compile time, all else equal, the JIT doesn’t know whether `_animal` is a `Dog` or a `Cat`. It generates code that loads the instance’s “method table pointer” (its object type handle), sometimes called a “vtable pointer”, stored at the beginning of every .NET object, indexes into the method table at the known slot for `Speak`, and calls the function pointer found there:\n\n```\n; x64\nmov     rcx, [rcx+8]   ; load _animal\nmov     rax, [rcx]     ; load method table\nmov     rax, [rax+40]  ; load vtable chunk\ncall    qword ptr [rax+20]\n```\nFor this one call to `Speak`, we pay three dependent memory dereferences and an indirect call because the processor doesn’t know for certain in advance where the call is going (it might guess, or “speculatively execute”, but it has to be prepared for the possibility it was wrong), and because the call target is indirect, the JIT can’t inline the callee. Whatever `Speak` does, its code can’t be folded into the calling method.\n\nThat’s a performance problem. Those indirections have overhead, but the bigger cost is the lost opportunity to inline. Inlining not only saves function call overhead, more importantly it opens the callee’s code up to the same optimizations that are operating on the caller, such as constant propagation, dead code elimination, bounds check elimination, further devirtualization, etc. That means a series of small virtual calls that each look innocent can, when devirtualized and inlined, collapse into a handful of instructions that would be unrecognizable and way cheaper when compared to the original source code. Without inlining, each callee is an opaque box; with it, the JIT can see through the layers.\n\nWe as .NET developers constantly rely on the JIT’s sophisticated heuristics for inlining that weigh the IL size of the callee, the exact work the callee is performing, the call frequency of the method, the expected benefit from constant arguments, and dozens of other factors. For virtual calls, the JIT needs to know what the actual target of the call will be; it needs to “devirtualize”. In some cases, it can determine that statically, where it has exact-type knowledge. For example, if the JIT can prove that `animal` is always a `Dog`, whether because it was just allocated with `new Dog()`:\n\n```\nAnimal animal = GetSomeAnimal();\nanimal.Speak();\n...\nstatic Animal GetSomeAnimal() => new Dog(); // inlineable\n```\nor because the variable’s type is a sealed class:\n\n```\nDog animal = GetSomeAnimal();\nanimal.Speak();\n...\nsealed class Dog { ... } // impossible for `animal` to be anything other than a `Dog`\n```\nor with NativeAOT and whole-program compilation, if it sees that `Animal` is abstract and the only type in the whole application that derives from `Animal` is `Dog`:\n\n```\nAnimal animal = GetSomeAnimal();\nanimal.Speak();\n...\nabstract class Animal { ... }\nclass Dog : Animal { ... } // no other such derived type\n```\nor other such validation, it can emit a call to `Dog.Speak()` directly, and the inliner can take its shot.\n\nBut for other cases where it can’t prove this with static analysis, the JIT turns to profile-guided optimization (PGO). PGO sounds fancy, but it’s conceptually simple. With “tiered compilation”, when a method is first invoked, it can be compiled “just in time” with few-to-no optimizations (this is referred to as Tier 0). The JIT can include in this compilation additional probes (think “printf debugging”) that let it track a bunch of interesting information about the nature of the code, recording what actually happens when it runs: which branches are taken, what are the concrete types that show up at virtual call sites or cast attempts, and so on. If the method is invoked enough or loops enough times, the runtime can ask the JIT to produce a new optimized version (referred to as Tier 1). That compilation can then factor in all of the learnings gathered as part of that profiling.\n\nThe JIT, of course, still needs to generate code that’s always correct. Even if a dynamic profile says `animal` was `Dog` 100% of the time, that doesn’t guarantee it’ll always be `Dog` in the future; it could be that the first 1000 calls passed in a `Dog` but the 1001st call is going to pass in `Dolphin`. How can the JIT incorporate this learning then? By emitting a run-time check. The `Dog` path can get a direct call, which may then be inlinable, and the other path keeps the original virtual call as the fallback. The speed comes from making the common case tiny, while correctness comes from leaving the uncommon case intact.\n\n```\n// Approximately what the JIT generates\nif (animal?.GetType() == typeof(Dog))\n{\n    ((Dog)animal).Speak();  // devirtualized, inlinable\n}\nelse\n{\n    animal.Speak(); // original virtual call, hopefully rare\n}\n```\nThis “guess and verify” pattern, called “guarded devirtualization” (GDV), accounts for many of the biggest throughput wins in real workloads. It’s applicable not only to virtual dispatch but also to interface dispatch, which also happens to be a bit more expensive than virtual dispatch because a type can implement any number of interfaces and that means the interface slots don’t simply map to fixed vtable positions.\n\nDeabstraction can also make object creation more efficient when it reveals what kind of object is involved. In general, objects in .NET are allocated on the garbage collected heap, tracked by the garbage collector (GC), and collected when no longer reachable. Heap allocation is typically fast, often effectively just bumping a pointer. However, when there’s not enough space available to bump the pointer, it can get much more expensive, including needing to incur a garbage collection. Every allocated object also effectively incurs the amortized cost of all collections, as every allocated object eventually needs to be cleaned up.\n\n“Escape analysis” is the compiler technique that lets us ask whether this object ever “escapes” the current method. If an object reference to a newly allocated object provably doesn’t escape, then the JIT can more efficiently allocate it. It needn’t store it on the GC heap, because nothing could possibly need to reference that object again, so it can instead allocate the object on the stack, making both allocation and cleanup essentially free. Stack allocation is even faster than heap bump-pointer allocation; it’s just decrementing the stack pointer, which is typically already in a register. And more importantly it means zero GC impact, because the stack frame is freed atomically on function return.\n\nThe JIT’s been progressively expanding escape analysis over the past several .NET releases, with .NET 9 and 10 seeing significant investments in stack-allocating delegates and closures, `Nullable<T>` temporaries, and small helper objects. The key theme is that every false positive escape, every time the JIT incorrectly concludes an object may escape when it really doesn’t, represents a heap allocation that could have been avoided, and we want to whittle away at that false positive list. In .NET 11, the JIT trims that list in several ways.\n\nWe’ll start with nullable boxing. Consider this benchmark:\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[MemoryDiagnoser(false), HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private int? _nullableNull;\n    private int? _nullableValue = 42;\n    [Benchmark]\n    public object? BoxNullableNull() => (object?)_nullableNull;\n    [Benchmark]\n    public object? BoxNullableValue() => (object?)_nullableValue;\n    [Benchmark]\n    public string? FormatNullableInt() => Format(_nullableValue);\n    private static string? Format<T>(T value)\n    {\n        if (value is IFormattable formattable)\n            return formattable.ToString(null, null);\n        return null;\n    }\n}\n```\n| Method | Runtime | Mean | Ratio | Allocated | Alloc Ratio | \n|---|---|---|---|---|---|\n| BoxNullableNull | .NET 10.0 | 2.095 ns | 1.00 | – | – | \n| BoxNullableNull | .NET 11.0 | 1.764 ns | 0.84 | – | – | \n| BoxNullableValue | .NET 10.0 | 9.213 ns | 1.00 | 24 B | 1.00 | \n| BoxNullableValue | .NET 11.0 | 4.126 ns | 0.45 | 24 B | 1.00 | \n| FormatNullableInt | .NET 10.0 | 9.583 ns | 1.00 | 24 B | 1.00 | \n| FormatNullableInt | .NET 11.0 | 1.987 ns | 0.21 | – | 0 | \n\ndotnet/runtime#122167 expands nullable boxing inside the JIT, exposing the temporary box to escape analysis; previously, a runtime helper hid it. For a `null` input, there’s no allocation on either version, because nothing gets boxed. And on both versions, `BoxNullableValue` returns the boxed object, meaning the object escapes, so the 24-byte allocation remains. However, for `FormatNullableInt`, the JIT in .NET 11 can now see that the temporary 24-byte box doesn’t escape and eliminates that heap allocation entirely.\n\nEscape analysis improved further for enumerators, through a mechanism called Conditional Escape Analysis (CEA). Support for CEA was introduced in .NET 10, but .NET 11 extends the set of patterns that this analysis can safely recognize. The existing escape analysis asks whether a reference created by an allocation can flow somewhere the JIT can no longer track, such as an unknown call. If it can, the object must remain on the heap. That analysis is necessarily conservative and largely flow-insensitive: if an object might be passed to an interface call on any path, it doesn’t try to prove that the path containing that call is mutually exclusive with the path containing the allocation.\n\nUnfortunately, that’s exactly what GDV produces when it optimizes a `foreach` over an `IEnumerable<T>`. As noted earlier, GDV turns an interface call into a type check with two branches: a fast branch for the likely collection type and a fallback branch containing the original interface call. Devirtualization and inlining along the fast branch will often reveal an enumerator allocation for the collection type, while later enumerator guards retain fallback calls such as `IEnumerator<T>.MoveNext`. The existing analysis sees those calls and concludes that the locally allocated enumerator might escape. CEA instead records the relationship between the fast-path allocation and the enumerator local tested by the later guards. If every apparent escape occurs only behind a failed type check, the JIT can clone the region into a hot version where those checks are known to succeed. In that clone, the object can’t reach the fallback calls, so it can be stack-allocated and often promoted into separate scalar locals. The original region remains as the general slow path.\n\nOne case .NET 10 didn’t handle, though, was a `GetEnumerator()` implementation that returns the result of another `GetEnumerator()` call. A collection expression converted to `IEnumerable<int>`, for example, uses a compiler-generated read-only-array wrapper with exactly this structure: the wrapper’s `GetEnumerator()` delegates to the underlying array’s `GetEnumerator`. With dotnet/runtime#122946, the JIT in .NET 11 handles this “chaining”:\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[MemoryDiagnoser(false), HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private static readonly IEnumerable<int> s_readOnlyStatic = [1, 2, 3, 4, 5];\n    private readonly IEnumerable<int> _readOnlyInstance = [1, 2, 3, 4, 5];\n    [Benchmark]\n    public int ReadOnlyStatic()\n    {\n        int sum = 0;\n        foreach (int item in s_readOnlyStatic) sum += item;\n        return sum;\n    }\n    [Benchmark]\n    public int ReadOnlyInstance()\n    {\n        int sum = 0;\n        foreach (int item in _readOnlyInstance) sum += item;\n        return sum;\n    }\n}\n```\n| Method | Runtime | Mean | Ratio | Allocated | Alloc Ratio | \n|---|---|---|---|---|---|\n| ReadOnlyStatic | .NET 10.0 | 2.665 ns | 1.00 | – | – | \n| ReadOnlyStatic | .NET 11.0 | 2.666 ns | 1.00 | – | – | \n| ReadOnlyInstance | .NET 10.0 | 13.874 ns | 1.00 | 32 B | 1.00 | \n| ReadOnlyInstance | .NET 11.0 | 2.674 ns | 0.19 | – | 0 | \n\n`ReadOnlyStatic`, whose `static readonly` field the JIT can effectively treat as a constant, was already optimized in .NET 10. In .NET 11, the instance-field case also loses its 32-byte enumerator allocation and converges on the same throughput.\n\ndotnet/runtime#121918 from @MichalPetryka fixes another way an address could unnecessarily make an object appear to escape. The IL `constrained.` prefix lets one generic `callvirt` sequence work for both value types and reference types: it can avoid boxing a value type, while for a reference type it dereferences the receiver and performs normal virtual dispatch. `ObjectEqualityComparer<T>.Equals`, used in the following benchmark by `EqualityComparer<T>.Default`, contains such a call to `value.Equals(other)`. The receiver was represented as an indirect read through the address of a local. Merely taking that address marked the local as exposed, preventing the newly allocated `Value` from being considered for stack allocation. The receiver is now represented as a direct value load instead, and the 24-byte heap allocation disappears.\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Collections.Generic;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[MemoryDiagnoser(false), HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private static readonly Value s_other = new(42);\n    [Benchmark]\n    public bool Equals() => EqualityComparer<Value>.Default.Equals(new Value(42), s_other);\n    private sealed class Value(int value)\n    {\n        private readonly int _value = value;\n        public override bool Equals(object? obj) => obj is Value other && _value == other._value;\n        public override int GetHashCode() => _value;\n    }\n}\n```\n| Method | Runtime | Mean | Ratio | Allocated | Alloc Ratio | \n|---|---|---|---|---|---|\n| Equals | .NET 10.0 | 3.874 ns | 1.00 | 24 B | 1.00 | \n| Equals | .NET 11.0 | 1.786 ns | 0.46 | – | 0 | \n\nWhile CEA can move a non-escaping object off the GC heap, sometimes the JIT can go further and prove an allocation need not exist at all. Generic code provides a common source of such opportunities through boxing. For example, the `ArgumentNullException.ThrowIfNull` method accepts an `object value`. That means when you have a method like this:\n\n```\nstatic void Test<T>(T value)\n{\n    ArgumentNullException.ThrowIfNull(value);\n    ...\n}\n```\nwhen `T` is constrained to a non-nullable struct, boxing is incurred, in order to pass `value` as `object`. `ThrowIfNull` here is a nop if `value` is non-`null` (since the method is simply `if (value is null) Throw();`), and previous releases successfully optimized away that boxing in optimized code. However, in Tier 0, that optimization wasn’t applied, and `ThrowIfNull` would end up allocating. While this wouldn’t negatively impact steady-state throughput, it would lead to annoying noise in profiling, as well as additional overhead during startup, where such use wasn’t yet promoted out of Tier 0. In .NET 11, dotnet/runtime#129392 adds support for this in Tier 0 as well.\n\nOn the virtual-dispatch side, multiple PRs contribute to improving generic virtual methods (GVMs). dotnet/runtime#120866 from @hez2010 stops eagerly spilling `ldvirtftn` call targets into a temporary, and lets generic virtual target resolution move ahead of argument setup when legal. dotnet/runtime#122023 from @hez2010 then enables the JIT to devirtualize non-shared GVMs, carrying the generic context needed to turn the indirect dispatch into a direct, and potentially inlineable, call. And dotnet/runtime#128702 from @hez2010 extends that support to shared GVMs and default interface implementations that require an instantiating stub. These optimizations can increase total code size when the newly direct calls are inlined, but that’s generally the desired trade: more of the actual work becomes visible to the optimizer. Consider the following benchmark:\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Runtime.CompilerServices;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[MemoryDiagnoser(false), HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    [Benchmark]\n    public int NonShared() => ((IProcessor)new Processor()).SizeOf(42);\n    [Benchmark]\n    public int Shared() => ((IProcessor)new Processor()).SizeOf(\"hello\");\n    private interface IProcessor\n    {\n        int SizeOf<T>(T value);\n    }\n    private sealed class Processor : IProcessor\n    {\n        public int SizeOf<T>(T value) => Unsafe.SizeOf<T>();\n    }\n}\n```\nCasting a freshly allocated `Processor` to `IProcessor` incurs an interface generic virtual call in the IL, but the JIT is now able to see the receiver’s exact type, even in the shared `string` case, such that .NET 11 devirtualizes and inlines both calls. That in turn exposes `Unsafe.SizeOf<T>()` as a constant and proves that the short-lived `Processor` doesn’t need to be allocated at all.\n\n| Method | Runtime | Mean | Ratio | Allocated | Alloc Ratio | \n|---|---|---|---|---|---|\n| NonShared | .NET 10.0 | 6.678 ns | 1.00 | 24 B | 1.00 | \n| NonShared | .NET 11.0 | 1.764 ns | 0.26 | – | 0 | \n| Shared | .NET 10.0 | 7.166 ns | 1.00 | 24 B | 1.00 | \n| Shared | .NET 11.0 | 1.764 ns | 0.25 | – | 0 | \n\nBuilding on that, dotnet/runtime#123183 from @hez2010 enables ReadyToRun compilation to resolve and devirtualize more non-shared generic virtual calls that would otherwise remain indirect, and dotnet/runtime#130202 from @hez2010 extends that support to NativeAOT. NativeAOT represents some generic virtual targets as “fat pointers” (pointers that are more than just an address, typically an address and associated metadata, and that in this case carry both a code address and generic context); by deferring that transformation until after exact-type devirtualization has had a chance to run, the JIT can turn an interface call site with a single known target to a non-shared GVM into a direct call that may then be inlined.\n\nType information also needs to survive the transformations the JIT performs internally. If the JIT spills a reference expression into a temporary while restructuring a tree, losing the expression’s exact class information can turn a call that was devirtualizable back into an opaque virtual call. That’s what happens here in .NET 10: `Value` gets boxed and `SetValue` is invoked through `IValue`. dotnet/runtime#128485 from @hez2010 preserves the class handle and exactness on the temporary. With that information still available, .NET 11 devirtualizes and inlines the call, eliminating the box and its 24-byte allocation.\n\nSeparately, dotnet/runtime#127433 relaxes the inliner’s budget heuristics for callees on `[Intrinsic]` types like `Span` and `Vector`. These types intentionally expose many small, composable methods that serve as gateways to JIT-recognized operations. If a wrapper remains as a call, the caller pays the call overhead and optimizations around it see an opaque boundary. If it inlines, the importer can replace its body with an intrinsic node and optimize that node together with the surrounding indexing, bounds checks, and vector operations. Giving such wrappers more favorable budgeting therefore keeps more of them inlineable and exposes more of the actual operation to the rest of the optimizer.\n\nOne of the core abstraction-enabling mechanisms in .NET is delegates: they let us pass around objects representing functions to be invoked, carrying with them associated required state. Deabstraction enables avoiding paying for the overheads associated with delegates in some cases. For the rest, we still want those delegates to be as cheap as possible. dotnet/runtime#99200 from @MichalPetryka simplifies CoreCLR’s delegate representation, removing one pointer-sized field from every delegate object. That saves 8 bytes per delegate in a 64-bit CoreCLR process. dotnet/runtime#129304 from @MichalPetryka improves Native AOT’s delegate layout separately by reordering its existing four fields so related values are adjacent. The updated layouts also give equality and hash-code operations more direct access to the method identity they need.\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[MemoryDiagnoser(false), HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private static readonly Target s_target = new();\n    private static readonly Func<int> s_first = s_target.GetValue;\n    private static readonly Func<int> s_second = s_target.GetValue;\n    [Benchmark]\n    public Func<int> ClosedInstance() => s_target.GetValue;\n    [Benchmark]\n    public bool DelegateEquals() => s_first.Equals(s_second);\n    [Benchmark]\n    public int DelegateGetHashCode() => s_first.GetHashCode();\n    private sealed class Target\n    {\n        public int GetValue() => 42;\n    }\n}\n```\n| Method | Runtime | Mean | Ratio | Allocated | Alloc Ratio | \n|---|---|---|---|---|---|\n| ClosedInstance | .NET 10.0 | 7.395 ns | 1.00 | 64 B | 1.00 | \n| ClosedInstance | .NET 11.0 | 6.844 ns | 0.93 | 56 B | 0.88 | \n| DelegateEquals | .NET 10.0 | 3.254 ns | 1.00 | – | – | \n| DelegateEquals | .NET 11.0 | 2.215 ns | 0.68 | – | – | \n| DelegateGetHashCode | .NET 10.0 | 5.623 ns | 1.00 | – | – | \n| DelegateGetHashCode | .NET 11.0 | 3.741 ns | 0.67 | – | – | \n\ndotnet/runtime#129410 from @MichalPetryka follows up on the CoreCLR layout by placing the target object and method pointer next to each other. Those are commonly consumed together during invocation, and the adjacency enables paired loads on architectures such as Arm64.\n\n### Runtime Async\n\nFor more than a decade, `async` and `await` have let us write asynchronous code that looks remarkably similar to synchronous code: we can put a `try`/`catch` around an `await`, use local variables on either side of it, return a value and generally reason about the method in source order. When execution reaches an `await` for something that isn’t yet complete, however, the method can’t simply leave its current stack frame in place and wait for the operation to finish. The thread needs to be freed up to do other work, while the work after the `await`, including whatever local state it will need later, must survive somewhere. In C#, the compiler has traditionally been responsible for transforming the method into a representation that enables that continuation.\n\nI went into the history and mechanics of that transformation in How async/await really works. The very short version is that the compiler traditionally replaces an `async` method with a small entry method and a generated state machine whose `MoveNext` method contains the transformed user code. Parameters, locals that need to survive an incomplete await, spilled expression values, awaiters, the current state number, and a method builder all become fields on a heap-allocated object. The generated `MoveNext` method runs the user’s code until an awaiter reports that it isn’t yet complete. It stores enough information to know where and with what values to resume, registers `MoveNext` as the continuation, and returns. When the operation completes, `MoveNext` is invoked again, jumps to the right location based on the saved state number (think `goto` and a label), retrieves the result from a value-producing awaiter, and continues. If every awaiter is already complete, `MoveNext` can run all the way through synchronously. When the method completes or throws, the builder publishes the result, cancellation, or exception through the returned `Task`, `Task<T>`, `ValueTask`, or `ValueTask<T>` (or, in the rare case, a custom task-like type).\n\nFor example, consider this tiny method:\n\n```\nstatic async Task<int> ReadLengthAsync(Stream stream, CancellationToken cancellationToken)\n{\n    var buffer = new byte[4096];\n    int bytesRead = await stream.ReadAsync(buffer, 0, buffer.Length, cancellationToken);\n    return bytesRead;\n}\n```\nWhile the code that gets generated for this changes over time and differs between debug and release builds, the lowering by the C# compiler has looked something like this:\n\n```\n[AsyncStateMachine(typeof(<ReadLengthAsync>d__0))]\nstatic Task<int> ReadLengthAsync(Stream stream, CancellationToken cancellationToken)\n{\n    <ReadLengthAsync>d__0 stateMachine = default;\n    stateMachine.builder = AsyncTaskMethodBuilder<int>.Create();\n    stateMachine.state = -1;\n    stateMachine.stream = stream;\n    stateMachine.cancellationToken = cancellationToken;\n    stateMachine.builder.Start(ref stateMachine);\n    return stateMachine.builder.Task;\n}\nstruct <ReadLengthAsync>d__0 : IAsyncStateMachine\n{\n    public int state;\n    public AsyncTaskMethodBuilder<int> builder;\n    public Stream stream;\n    public CancellationToken cancellationToken;\n    private TaskAwaiter<int> awaiter;\n    public void MoveNext()\n    {\n        int result;\n        try\n        {\n            TaskAwaiter<int> localAwaiter;\n            if (state != 0)\n            {\n                byte[] buffer = new byte[4096];\n                localAwaiter = stream.ReadAsync(buffer, 0, buffer.Length, cancellationToken).GetAwaiter();\n                if (!localAwaiter.IsCompleted)\n                {\n                    state = 0;\n                    awaiter = localAwaiter;\n                    builder.AwaitUnsafeOnCompleted(ref localAwaiter, ref this);\n                    return;\n                }\n            }\n            else\n            {\n                localAwaiter = awaiter;\n                awaiter = default;\n                state = -1;\n            }\n            result = localAwaiter.GetResult();\n        }\n        catch (Exception e)\n        {\n            state = -2;\n            builder.SetException(e);\n            return;\n        }\n        state = -2;\n        builder.SetResult(result);\n    }\n}\n```\nThat’s quite a lot of generated code for three lines of C#. The compiler has to make decisions before the program runs about the state-machine layout, which values might need to survive, how many awaiter fields are required, and how all the suspension points fit into one `MoveNext` dispatch. The runtime and JIT have optimized the resulting pattern heavily over the years, including combining the task, state machine, continuation, and `ExecutionContext` into a single allocation, but by the time the JIT sees the IL, the transformation has already happened, leaving it with a very complicated system to try to optimize.\n\n.NET 11 introduces a new way to split that responsibility, a reimplementation of the `async`/`await` infrastructure referred to as “runtime async”. Rather than the C# compiler being responsible for the transformation, the JIT is. The C# compiler emits a much smaller suspension-aware IL contract for each eligible `async` method and marks the method as `async` in metadata. The runtime and JIT then do the work that depends on runtime knowledge: creating the externally visible `Task` or `ValueTask`, recognizing direct async calls, deciding which values are actually alive at each suspension point, laying out continuation objects, and generating the control flow that suspends and resumes the method. Effectively, the transformation moves from C# to the runtime, where more information is available to optimize it.\n\nThe programming model hasn’t changed. This is still C# `async`/`await`; `await` still obeys the awaiter pattern, exceptions and cancellation still surface through the returned task-like object, `ConfigureAwait` still has its usual meaning, synchronous completion is still synchronous completion, and on and on. An explicit goal for the feature has been 100% behavioral compatibility: whether an `async` method is lowered by the language compiler or by the runtime is an implementation detail, and any observable semantic difference is a bug.\n\nIn .NET 11, application code opts in with a compiler feature switch:\n\n```\n<Project Sdk=\"Microsoft.NET.Sdk\">\n  <PropertyGroup>\n    <TargetFramework>net11.0</TargetFramework>\n    <Features>$(Features);runtime-async=on</Features>\n  </PropertyGroup>\n</Project>\n```\nNote that there’s no new C# syntax involved, so `LangVersion=preview` isn’t required, nor is `EnablePreviewFeatures`. While this is opt-in at the application layer, most of the in-box shared framework is already built this way for .NET 11. The `async`/`await` performance goal for .NET 11 is parity with .NET 10, and in general runtime async is already as good as or better than the older implementation in many important paths. It isn’t yet fully optimized, though, and there are known cases where it still produces less efficient code. I’d encourage you to experiment in .NET 11 with opting-in your applications and services; just make sure to measure. My hope is that it’ll be on by default starting in .NET 12.\n\nMoving the transformation from the C# compiler to the runtime has the added benefit of reducing binary size. As noted, the traditional lowering emits an entry method, a generated state-machine type, fields for captured state, and a `MoveNext` body, for every async method. Runtime async leaves a much smaller method body for the runtime to transform. The following tiny app contains ten `Task<int>`-returning async methods, each awaiting the next, and compiles the same source once with compiler lowering and once with runtime async:\n\n```\n<Project Sdk=\"Microsoft.NET.Sdk\">\n  <PropertyGroup>\n    <OutputType>Exe</OutputType>\n    <TargetFramework>net11.0</TargetFramework>\n    <AssemblyName>SizeProbe</AssemblyName>\n    <ImplicitUsings>enable</ImplicitUsings>\n    <Nullable>enable</Nullable>\n    <Features Condition=\"'$(RuntimeAsync)' == 'true'\">$(Features);runtime-async=on</Features>\n  </PropertyGroup>\n</Project>\n```\n```\n// dotnet build -c Release -p:RuntimeAsync=false -o classic --no-incremental; dotnet build -c Release -p:RuntimeAsync=true -o runtime --no-incremental; Get-Item .\\classic\\SizeProbe.dll, .\\runtime\\SizeProbe.dll | Select-Object Directory, Length\nConsole.WriteLine(await Benchmarks.Layer0());\npublic class Benchmarks\n{\n    public static async Task<int> Layer0() => await Layer1();\n    private static async Task<int> Layer1() => await Layer2();\n    private static async Task<int> Layer2() => await Layer3();\n    private static async Task<int> Layer3() => await Layer4();\n    private static async Task<int> Layer4() => await Layer5();\n    private static async Task<int> Layer5() => await Layer6();\n    private static async Task<int> Layer6() => await Layer7();\n    private static async Task<int> Layer7() => await Layer8();\n    private static async Task<int> Layer8() => await Layer9();\n    private static async Task<int> Layer9()\n    {\n        await Task.Yield();\n        return 42;\n    }\n}\n```\n| Lowering | SizeProbe.dll | Ratio | \n|---|---|---|\n| Compiler | 10,752 bytes | 1.00 | \n| Runtime async | 5,632 bytes | 0.52 | \n\nFor a method such as:\n\n`static async Task<int> CallerAsync() => await CalleeAsync();`\nwith runtime async enabled, the C# compiler generates IL like the following:\n\n```\n; MSIL\n.method private hidebysig static\n    class System.Threading.Tasks.Task`1<int32> CallerAsync() cil managed async\n{\n    call class System.Threading.Tasks.Task`1<int32> CalleeAsync()\n    call int32 System.Runtime.CompilerServices.AsyncHelpers::Await<int32>(\n        class System.Threading.Tasks.Task`1<int32>)\n    ret\n}\n```\nThere is no generated `<CallerAsync>d__0` type, no `IAsyncStateMachine`, no `MoveNext`, no `AsyncTaskMethodBuilder<int>`, and no `AsyncStateMachineAttribute`. Previously, `async` on a C# method evaporated at compile time. Now, the method has a new `MethodImpl` `async` bit, represented in IL assembly syntax by that `async` modifier, and the body calls helpers in `System.Runtime.CompilerServices.AsyncHelpers`.\n\nAt first glance the `ret` looks impossible because the declared signature returns `Task<int>` while the value on the IL evaluation stack is an `int`. This clearly isn’t a normal calling convention. The VM can give a `Task`-returning method two related identities, or MethodDescs, where one has the normal signature the rest of managed code sees, `Task<int> CallerAsync()`. The other is the AsyncCall variant, which effectively returns `int` and has an implicit channel for a continuation. Both refer to the same logical method and metadata token, but they have different calling conventions and different jobs. If regular managed code invokes `CallerAsync`, the VM-generated outer thunk preserves the public contract and returns a `Task<int>`. If another runtime async method directly awaits it, the JIT can instead call the AsyncCall variant and receive the result directly when the call completes synchronously, or a continuation when it suspends. In other words, it can hand back the `T` directly and avoid allocating a `Task<T>`.\n\nThat pairing works in both directions. For a method compiled with runtime async, the AsyncCall variant owns the generated (newly compact) IL while the public `Task`-returning entry point is an adapter thunk; for a traditionally compiled method, the public method owns its usual IL while the VM can create an AsyncCall adapter around it. That means runtime async code remains able to await existing libraries and code compiled by older compilers, a critical capability for our goal of 100% compat. The largest wins naturally appear as more of an async call chain is compiled with runtime async.\n\nThis is where the JIT gets an opportunity that simply didn’t exist when every boundary was already expressed as a task and a generated state machine. Suppose `A` awaits `B`, which awaits `C`:\n\n```\nstatic async Task<int> A(bool yield) => await B(yield);\nstatic async Task<int> B(bool yield) => await C(yield);\nstatic async Task<int> C(bool yield)\n{\n    if (yield)\n        await Task.Yield();\n    return 42;\n}\n```\nTraditionally, each method has its own compiler-generated state machine and its own task-like result. `C` suspends and eventually completes its task, which wakes `B`‘s state machine; `B` then completes its task, which wakes `A`‘s state machine; and `A` completes the root task observed by the caller. There has been an enormous amount of work done over the years to reduce the costs of those objects and transitions.\n\nWith runtime async, the importer recognizes the adjacent pattern of “call a Task-returning method, then await that task.” In the simple case it can call the callee’s AsyncCall variant instead. When `yield` is false and `C` completes synchronously, the `int` flows back through `B` and `A` as a plain value, and only the outermost boundary needs to turn it into the `Task<int>` promised to the original caller. When `yield` is true and `C` suspends, the runtime links continuation state for the chain and eventually resumes it without requiring an intermediate `Task<int>` at every directly fused edge. The `Task` contract hasn’t vanished, it just moved to the place where a `Task` is actually needed.\n\nRuntime async doesn’t make every asynchronous operation allocation-free, though. Rather, it gives the JIT enough information to avoid materializing some task objects that existed only to carry a result from one async method directly into the next. If a consumer stores the task in a collection, manually hooks up a continuation, or otherwise observes the task as an object, that object is still needed. The optimization is about not paying for boundaries that aren’t observably boundaries.\n\nThe impact is already visible with just two layers:\n\n```\n// dotnet run -c Release -f net11.0 --filter \"*\"\n// The project also needs the `runtime-async=on` feature switch set.\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Configs;\nusing BenchmarkDotNet.Running;\nusing System.Runtime.CompilerServices;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[MemoryDiagnoser(false)]\n[GroupBenchmarksBy(BenchmarkLogicalGroupRule.ByCategory)]\n[HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private static readonly Task<int> s_completed = Task.FromResult(42);\n    [Benchmark(Baseline = true), BenchmarkCategory(\"Completed\")]\n    public Task<int> ClassicCompleted() => ClassicCompletedOuter();\n    [Benchmark, BenchmarkCategory(\"Completed\")]\n    public Task<int> RuntimeCompleted() => RuntimeCompletedOuter();\n    [Benchmark(Baseline = true), BenchmarkCategory(\"Yielding\")]\n    public Task<int> ClassicYielding() => ClassicYieldingOuter();\n    [Benchmark, BenchmarkCategory(\"Yielding\")]\n    public Task<int> RuntimeYielding() => RuntimeYieldingOuter();\n    [RuntimeAsyncMethodGeneration(false)]\n    private static async Task<int> ClassicCompletedOuter() => await ClassicCompletedInner();\n    [RuntimeAsyncMethodGeneration(false)]\n    private static async Task<int> ClassicCompletedInner() => await s_completed;\n    private static async Task<int> RuntimeCompletedOuter() => await RuntimeCompletedInner();\n    private static async Task<int> RuntimeCompletedInner() => await s_completed;\n    [RuntimeAsyncMethodGeneration(false)]\n    private static async Task<int> ClassicYieldingOuter() => await ClassicYieldingInner();\n    [RuntimeAsyncMethodGeneration(false)]\n    private static async Task<int> ClassicYieldingInner()\n    {\n        await Task.Yield();\n        return 42;\n    }\n    private static async Task<int> RuntimeYieldingOuter() => await RuntimeYieldingInner();\n    private static async Task<int> RuntimeYieldingInner()\n    {\n        await Task.Yield();\n        return 42;\n    }\n}\nnamespace System.Runtime.CompilerServices\n{\n    [AttributeUsage(AttributeTargets.Method)]\n    internal sealed class RuntimeAsyncMethodGenerationAttribute(bool runtimeAsync) : Attribute\n    {\n        public bool RuntimeAsync => runtimeAsync;\n    }\n}\n```\n| Method | Mean | Ratio | Allocated | Alloc Ratio | \n|---|---|---|---|---|\n| ClassicCompleted | 21.221 ns | 1.00 | 144 B | 1.00 | \n| RuntimeCompleted | 6.151 ns | 0.29 | 0 B | 0.00 | \n| ClassicYielding | 254.139 ns | 1.00 | 248 B | 1.00 | \n| RuntimeYielding | 116.927 ns | 0.46 | 168 B | 0.68 | \n\nThe synchronously completing chain is more than 3x faster and avoids both intermediate task allocations. Even after a real suspension, the same two-layer chain takes less than half the time and allocates 80 fewer bytes.\n\nException handling amplifies the difference. Again consider an async method `A` calling an async method `B` calling an async method `C`. The transformation generated by the C# compiler of each method results in a `try`/`catch` block around the whole body of the `MoveNext` method so that any unhandled exception can be stored into the returned `Task`. Let’s say code in `C` throws an unhandled exception. That’s then caught by this manufactured `catch` block and stored into the `Task` returned to `B`. The awaiter in `B` then retrieves that exception from the `Task` object and throws it. It’s then caught by `B`‘s generated catch and stored into its `Task`. And so on. An exception crossing ten such async helpers can therefore be thrown, caught, and stored ten times even though none of the source methods has an explicit handler. That is super expensive. But runtime async doesn’t need to re-enter a pass-through frame with no handler. On the synchronous path the exception unwinds through the fused calls normally, and after a real suspension, one dispatch-loop catch walks past continuation records that have no handler and faults the observable root task once.\n\nThe following benchmark measures both a fully synchronous throw and an exception after one real `Task.Yield` suspension. It uses a compiler-recognized per-method escape hatch (`RuntimeAsyncMethodGeneration`) so that the classic and runtime async methods run in the same process on the same .NET 11 runtime and differ only in how the compiler lowers them. (Note that this attribute is experimental and isn’t a public API exposed from the core libraries; as with other attributes known to the C# compiler, it recognizes them by name and signature.)\n\n```\n// dotnet run -c Release -f net11.0 --filter \"*\"\n// The project also needs the `runtime-async=on` feature switch set.\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Runtime.CompilerServices;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[MemoryDiagnoser(false), HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    [Params(1, 10, 30)]\n    public int Depth;\n    [Params(false, true)]\n    public bool Yield;\n    [Benchmark(Baseline = true)]\n    public int Classic() => Invoke(ClassicThrowAsync(Depth));\n    [Benchmark]\n    public int Runtime() => Invoke(RuntimeThrowAsync(Depth));\n    private static int Invoke(Task<int> task)\n    {\n        try\n        {\n            return task.GetAwaiter().GetResult();\n        }\n        catch (InvalidOperationException)\n        {\n            return -1;\n        }\n    }\n    [RuntimeAsyncMethodGeneration(false)]\n    private async Task<int> ClassicThrowAsync(int depth)\n    {\n        if (depth == 0)\n        {\n            if (Yield) await Task.Yield();\n            throw new InvalidOperationException(\"uh oh\");\n        }\n        return await ClassicThrowAsync(depth - 1);\n    }\n    private async Task<int> RuntimeThrowAsync(int depth)\n    {\n        if (depth == 0)\n        {\n            if (Yield) await Task.Yield();\n            throw new InvalidOperationException(\"uh oh\");\n        }\n        return await RuntimeThrowAsync(depth - 1);\n    }\n}\nnamespace System.Runtime.CompilerServices\n{\n    [AttributeUsage(AttributeTargets.Method)]\n    internal sealed class RuntimeAsyncMethodGenerationAttribute(bool runtimeAsync) : Attribute\n    {\n        public bool RuntimeAsync => runtimeAsync;\n    }\n}\n```\n| Depth | Yield | Method | Mean | Ratio | Allocated | Alloc Ratio | \n|---|---|---|---|---|---|---|\n| 1 | False | Classic | 4.727 μs | 1.00 | 1.6 KB | 1.00 | \n| 1 | False | Runtime | 3.558 μs | 0.75 | 1.16 KB | 0.72 | \n| 1 | True | Classic | 6.308 μs | 1.00 | 1.68 KB | 1.00 | \n| 1 | True | Runtime | 8.211 μs | 1.30 | 1.42 KB | 0.85 | \n| 10 | False | Classic | 19.469 μs | 1.00 | 15.13 KB | 1.00 | \n| 10 | False | Runtime | 5.885 μs | 0.30 | 2.13 KB | 0.14 | \n| 10 | True | Classic | 24.923 μs | 1.00 | 15.63 KB | 1.00 | \n| 10 | True | Runtime | 6.122 μs | 0.25 | 2.88 KB | 0.18 | \n| 30 | False | Classic | 51.302 μs | 1.00 | 84.2 KB | 1.00 | \n| 30 | False | Runtime | 10.721 μs | 0.21 | 5.71 KB | 0.07 | \n| 30 | True | Classic | 65.974 μs | 1.00 | 85.53 KB | 1.00 | \n| 30 | True | Runtime | 11.254 μs | 0.17 | 7.72 KB | 0.09 | \n\nRuntime async supports `Task`, `Task<T>`, `ValueTask`, and `ValueTask<T>` as method return types, but as of today it doesn’t support `async void`, async iterators, or arbitrary custom task-like return types with custom builders; those continue to use the traditional compiler transformation. For `ValueTask<T>`, the existing reasons to use the type still apply. A `ValueTask<T>` can carry a result directly, wrap a `Task<T>`, or refer to an `IValueTaskSource<T>`. That’s made it useful for APIs where synchronous completion is common enough that avoiding a `Task` allocation outweighs the larger return value and the more restrictive consumption rules, or where asynchronous completion can have its costs amortized via a reusable backing object. Runtime async then addresses some of the scenarios that would have led developers to use `ValueTask<T>`. Does that mean everyone should stop using `ValueTask<T>`? No. Choosing `Task` versus `ValueTask` remains an API design decision based on completion patterns, allocation sensitivity, call frequency, and how consumers need to use the result. Write the return type that makes sense for the API, then let the compiler, VM, and JIT optimize it as best they can.\n\nWorkloads with many layers of small async methods can benefit the most from runtime async, because those layers are exactly where intermediate tasks and state machines often accumulate. Shared framework code, for example, is full of this pattern: a public method validates arguments and awaits a private helper, which awaits a transport helper, which awaits an operating-system operation. Application services similarly compose authentication, retry, logging, serialization, and I/O helpers. Runtime async can make the source-level decomposition cheaper without asking the developer to flatten the code into one giant method in order to avoid “implementation detail” costs.\n\nThe work required to reach this point has been extensive. A GitHub search of the runtime async tracking label on September 14, 2026 returned 235 pull requests, far too many for me to enumerate one by one. So I won’t try; you can peruse that label in your spare time. The work is also not only about direct performance improvements but also about\nimprovements to diagnostics and performance tooling that help you to make better\nuse of async in your own code. When an async method\nsuspends, its physical thread stack unwinds. That method’s continuation might later run\non a different thread whose physical stack begins in the thread pool, with the\nmethods that led to the original `await` nowhere to be found. A sampling\nCPU profiler can see where the processor is spending time, but without additional\ninformation, it can’t reliably connect those traces back through the logical async\ncall chain, making it hard to answer questions about what async call paths were actually costing.\nProfiling tools like the async profiler in Visual Studio have traditionally reconstructed those chains from\nevents emitted by `Task`‘s infrastructure, but async-heavy applications can generate enormous volumes of\nthose very chatty events. The resulting overhead easily perturbs the workload being measured, making\nit all but unusable in production. dotnet/runtime#127238 added a\nnew lightweight async-profiler event stream for .NET 11 and runtime async. Rather than sending every small\ntransition through the eventing system as its own full event, the runtime\nwrites compact records into per-thread buffers, delta-encoding timestamps and\ninstruction pointers and flushing the data in batches. It also puts a small\nidentifiable wrapper frame into the physical stack when invoking a\ncontinuation. A profiler can use that frame as an anchor, joining ordinary CPU\nsamples to the logical async call stack represented by the event stream. In some measurements,\nthis new approach added less than 1% overhead and shrank the traced data by an order of magnitude.\ndotnet/runtime#129043 and a few follow-up PRs extended\nthe same approach to the compiler-generated state machines used by existing\nasync code. Thus this\nisn’t useful only to applications that opt into runtime async; tooling gets one\nconsistent representation across both implementations.\n\nWhat should you as a developer do differently with runtime async in the picture? Mostly nothing. Keep writing asynchronous code the way you want it to read, and break a large operation into helpers when that makes the code clearer. Use `Task` by default and choose `ValueTask` where its API and usage tradeoffs genuinely fit. And don’t contort source code to remove a clean `await` just because today’s implementation might allocate an intermediate `Task`. The lowering strategy should “just work” as an implementation detail, preserve behavior, and make existing source get better as the runtime improves.\n\n### Bounds Checks\n\nC# is a memory-safe language. Accesses to arrays, strings, and spans are guaranteed by the runtime to be in-bounds; if you try to access `someArray[i]`, `someString[i]`, or `someSpan[i]` with an index less than 0 or greater than or equal to the length of the array/string/span, you’ll get an exception, not silently corrupted memory or a process crash. The runtime guarantees that all permitted accesses are within bounds, and that means it needs to be able to prove the access is in bounds. The main method the JIT has for achieving that is by injecting code that performs a bounds check, as if instead of:\n\n```\nint[] array = ...;\nint value = array[i];\n```\nyou’d written:\n\n```\nint[] array = ...;\nif ((uint)i >= array.Length) throw new IndexOutOfRangeException();\nint value = array[i];\n```\nAt the assembly level, a bounds check looks something like:\n\n```\n; x64\ncmp ecx, dword ptr [rax+8]        ; compare index with array length\njae THROW                         ; unsigned index >= length\nmov edx, dword ptr [rax+rcx*4+16] ; load the element\n```\nThe JIT could just inject such code on every access and call it a day, but such code adds overhead, so the JIT works to elide those checks and that overhead wherever it can prove the index is valid. Proving an index is valid means the JIT needs to be able to see from other evidence that it couldn’t possibly be out of bounds.\n\nThe quintessential example of that is a `for` loop over the full contents of an array or span:\n\n```\nfor (int i = 0; i < array.Length; i++)\n{\n    Use(array[i]);\n}\n```\nThe JIT recognizes from this idiom that, within the loop body, `i` is guaranteed to be in the range `[0, array.Length)`, and avoids emitting the bounds check for the `array[i]` access. The JIT has long handled this particular case. Other cases, not so much. Bounds-check elimination has improved in virtually every .NET release; more recent releases added range propagation for derived expressions (`.NET 7` and `.NET 8` saw significant improvements here), SSA-based reasoning (`.NET 9`), and better handling of `Span<T>`, whose length sits in a field rather than an object header, complicating tracking. Each year, the developers contributing to the JIT find new patterns that were being missed, that show up in the wild, and that are fixable. .NET 11 improves several such patterns.\n\nRange analysis in the JIT tracks intervals for each variable, an upper bound and a lower bound. For example, taking the true branch of `x < 5` gives the range for `x` in that branch an upper bound of 4 while taking the true branch of `x > 2` makes the lower bound 3. What about `x != 5`? On the true edge, we know `x` isn’t 5, and if the current range for `x` is `[5, 10]`, then we know the range must actually be `[6, 10]`… the lower bound can be tightened because the only value at the lower end is excluded. Similarly, if the range is `[0, 5]`, an `x != 5` assertion tells us the range is actually the narrower `[0, 4]`. Or, at least, that’s what you’d hope it would do. The JIT had this relevant comment:\n\n```\n// We have a != assertion, but it doesn't tell us much about the interval. So just skip it.\ncontinue;\n```\nIn .NET 11, dotnet/runtime#121273 replaces that logic with productive reasoning. It checks whether the excluded constant is at either edge of the currently tracked range, adding in the new insights if so. C# list patterns, introduced in C# 11, generate just such comparison sequences. For example, the pattern `name is [] or [':'] or [':', not ':', ..]` lowers to something like this:\n\n```\nif (name != null)\n{\n    int num = name.Length;\n    if (num == 0) return true;\n    if (num == 1)\n    {\n        if (name[0] == ':') return true;\n    }\n    else if (name[0] == ':' && name[1] != ':')\n    {\n        return true;\n    }\n    return false;\n}\n```\nRange analysis then proceeds with something like this:\n\n1. We know that `Array.Length` is never negative, so it has a range of`[0, Array.MaxLength]` .\n2. On the false edge of `num == 0` , we know that`num != 0` , so the range is narrowed now to`[1, Array.MaxLength]` .\n3. Similarly, on the false edge of `num == 1` , we know that`num != 1` , so the range is narrowed now to`[2, Array.MaxLength]` .\n4. We then access `name[0]` and`name[1]` , both of which are guaranteed in bounds based on the lower bound of 2 that was established.\n\nWithout the `!= constant` tightening, that narrowing wouldn’t happen, and the bounds checks in step 4 couldn’t be elided. Thankfully, they now can be in .NET 11. Consider this example:\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Runtime.CompilerServices;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private string[] _inputs = [\"\", \":\", \":x\", \"abc\", \":ab\", \"x\", \"ab:cd\"];\n    [Benchmark]\n    public int ClassifyAll()\n    {\n        int total = 0;\n        foreach (string s in _inputs) total += Classify(s);\n        return total;\n    }\n    [MethodImpl(MethodImplOptions.NoInlining)]\n    private static int Classify(ReadOnlySpan<char> name) =>\n        name switch\n        {\n            [] => 0,\n            [':'] => 1,\n            [':', not ':', ..] => 10 + name[0] + name[1],\n            _ => 3\n        };\n}\n```\nIn .NET 10, we can see the call to `CORINFO_HELP_RNGCHKFAIL` at the bottom of the method. That’s the tell-tale sign there was at least one bounds check in the method. With .NET 11, that sign is removed.\n\n```\n; Arm64\n--- .NET 10\n+++ .NET 11\n@@ -10,17 +10,15 @@\n             beq     G_M000_IG08\n G_M000_IG04:\n-            ldrh    w2, [x0]\n-            cmp     w2, #58\n+            ldrh    w1, [x0]\n+            cmp     w1, #58\n             bne     G_M000_IG06\n G_M000_IG05:\n-            cmp     w1, #1\n-            bls     G_M000_IG11\n             ldrh    w0, [x0, #0x02]\n             cmp     w0, #58\n             beq     G_M000_IG06\n-            add     w0, w2, w0\n+            add     w0, w1, w0\n             add     w0, w0, #10\n             b       G_M000_IG07\n@@ -44,8 +42,4 @@\n             mov     w0, wzr\n             b       G_M000_IG07\n-G_M000_IG11:\n-            bl      CORINFO_HELP_RNGCHKFAIL\n-            brk     #0\n-\n-; Total bytes of code 112\n+; Total bytes of code 96\n```\n“Assertion” machinery in the JIT propagates learned facts (like the aforementioned range information) between “basic blocks” (a sequence of instructions with one entry point, one exit point, and no branches into or out of the middle of it), so information established in block A flows to block B if A “dominates” B (meaning the only way to get to B is through A). But what about facts established earlier within the same block? That’s the gap that dotnet/runtime#121527 addresses. Consider this code:\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Runtime.CompilerServices;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private int[] _arr = new int[512];\n    [Benchmark]\n    public int RunMany()\n    {\n        int touched = 0;\n        for (int i = 0; i < _arr.Length - 2; i++)\n        {\n            Test(_arr, i);\n            touched++;\n        }\n        return touched;\n    }\n    [MethodImpl(MethodImplOptions.NoInlining)]\n    private static void Test(int[] arr, int i)\n    {\n        arr[i] = 0;  // 1: establishes 'i >= 0 && i < arr.Length'\n        i++;         // 2: same block\n        if (i < arr.Length) arr[i] = 0;  // 3: proven safe from 1's assertion\n    }\n}\n```\nStatements 1, 2, and 3 are all in the same basic block, up to the conditional; after statement 1 executes, if we reach statement 2, the bounds check on statement 1 passed, we know `i >= 0` and `i < arr.Length`, and after statement 2, `i` becomes `i + 1`. After the `if` guard `i < arr.Length` we know the incremented `i` is still within bounds. But when the range check pass in the .NET 10 JIT examined statement 3’s bounds check, it saw the assertions propagated from predecessor blocks. Since the assertion from statement 1 is generated within the current block, the range check couldn’t see it. The PR fixed it to walk the current block’s tree in execution order, accumulating assertions as it went. When we reach statement 3’s bounds check, we’ve already walked past statement 1 and picked up its `i >= 0 && i < arr.Length` assertion.\n\n```\n; Arm64\n--- .NET 10\n+++ .NET 11\n@@ -13,8 +13,6 @@\n             ble     G_M000_IG04\n G_M000_IG03:\n-            cmp     w1, w2\n-            bhs     G_M000_IG05\n             str     wzr, [x0, w1, UXTW #2]\n G_M000_IG04:\n@@ -25,4 +23,4 @@\n             bl      CORINFO_HELP_RNGCHKFAIL\n             brk     #0\n-; Total bytes of code 68\n+; Total bytes of code 60\n```\nThere are almost an infinite number of things the JIT could look for and special-case. But every special case requires code, maintenance, and, most importantly, compilation time. A “just-in-time” compiler typically runs while the application is running, so the JIT itself must be optimized and spend its limited budget only where there’s a likely payoff. That pushes the developers building it toward patterns that occur in real workloads. One such pattern, often seen in libraries like format decoders, builds a table\nindex with bitwise operations on a byte, for example\n`((b & 0x03) << 4) | ((b & 0xf0) >> 4)`. Each masked piece has a tiny upper\nbound, so the OR of those pieces is always in `[0..63]`, safely in range for\ne.g. a Base64 alphabet table. Until\ndotnet/runtime#122263, the JIT\noften failed to prove that combined bound and left a bounds check on the\nindex. Existing range-check code understood the upper bounds produced by\nbitwise AND and shifts, but not OR; the change lets the JIT combine the known\nbounds of both OR operands and remove the remaining array check.\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Runtime.CompilerServices;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private readonly byte[] _input = new byte[4096];\n    [GlobalSetup]\n    public void Setup() => new Random(42).NextBytes(_input);\n    [Benchmark]\n    public int Base64LikeIndex() => Sum(_input);\n    [MethodImpl(MethodImplOptions.NoInlining)]\n    private static int Sum(ReadOnlySpan<byte> input)\n    {\n        int sum = 0;\n        foreach (byte b in input)\n        {\n            int index = ((b & 0x03) << 4) | ((b & 0xF0) >> 4);\n            sum += \"ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789+/=\"u8[index];\n        }\n        return sum;\n    }\n}\n```\nThe .NET 10 assembly checks the computed index against the 65-byte lookup table on every iteration. In .NET 11, range analysis proves the index is at most 63, so both the comparison and the branch to the range-check failure helper disappear:\n\n```\n; x64\n M01_L00:\n        movzx    r9d, byte ptr [rdx+r8]\n        mov      r11d, r9d\n        and      r11d, 3\n        shl      r11d, 4\n        and      r9d, 0F0\n        sar      r9d, 4\n        or       r9d, r11d\n-       cmp      r9d, 41\n-       jae      short M01_L02\n        movzx    r9d, byte ptr [r10+r9]\n        add      eax, r9d\n        inc      r8d\n        cmp      r8d, ecx\n        jl       short M01_L00\n-M01_L02:\n-       call     CORINFO_HELP_RNGCHKFAIL\n-       int      3\n-\n-; Total bytes of code 95\n+; Total bytes of code 79\n```\nAs another example, dotnet/runtime#125056 improves the handling of guards like `(uint)i < span.Length` that are pervasive in performance-sensitive code. Consider this benchmark:\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Runtime.CompilerServices;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private int[] _data = Enumerable.Range(0, 512).ToArray();\n    [Benchmark]\n    public int RunMany()\n    {\n        int sum = 0;\n        for (int i = 0; i < _data.Length; i++)\n            sum += Test(_data, i);\n        return sum;\n    }\n    [MethodImpl(MethodImplOptions.NoInlining)]\n    private static int Test(Span<int> span, int i)\n    {\n        if ((uint)i < (uint)span.Length)\n        {\n            if (i != 0)\n                return span[i - 1] + span[i];\n            return span[i];\n        }\n        return 0;\n    }\n}\n```\nBecause the comparison is unsigned, `(uint)i` would be a large positive number if `i` were negative, making it impossible for `(uint)i < (uint)span.Length` to be true (since a span’s length is never negative, `(uint)span.Length` is at most `int.MaxValue`). Inside the true branch, `i` is therefore in `[0, span.Length - 1]`. Previously, the JIT wasn’t always recording the lower bound `i >= 0` when it processed the `(uint)i < span.Length` assertion, and that could leave bounds checks on expressions like `i - 1` in place. The fix adds the `[0, int.MaxValue - 1]` lower bound deduction for the index variable upon entering the true arm of a `(uint)i < span.Length` check. Combined with the existing range tracking for the upper bound, this gives the JIT a complete picture of `i`‘s range inside the guarded block.\n\n```\n; Arm64\n--- .NET 10\n+++ .NET 11\n@@ -8,10 +8,8 @@\n             cbz     w2, G_M000_IG05\n G_M000_IG03:\n-            sub     w3, w2, #1\n-            cmp     w3, w1\n-            bhs     G_M000_IG09\n-            ldr     w1, [x0, w3, UXTW #2]\n+            sub     w1, w2, #1\n+            ldr     w1, [x0, w1, UXTW #2]\n             ldr     w0, [x0, w2, UXTW #2]\n             add     w0, w1, w0\n@@ -33,8 +31,4 @@\n             ldp     fp, lr, [sp], #0x10\n             ret     lr\n-G_M000_IG09:\n-            bl      CORINFO_HELP_RNGCHKFAIL\n-            brk     #0\n-\n-; Total bytes of code 84\n+; Total bytes of code 68\n```\nBounds check elision is generally based on forms of range analysis, where the JIT needs to prove that a given index is guaranteed to be within the range of the data structure. But the same range analysis-based facts can prove that other checks are unnecessary. For example, once the JIT knows that an integer is in `[0..100]`, it can prove both that converting it to `byte` can’t lose data and that multiplying it by 10 can’t overflow. dotnet/runtime#124147 enables the JIT to use such facts to avoid unnecessary branches as part of `checked` operations. When range analysis proves that the operands are in ranges whose result can’t overflow, making `checked` a nop, the backend can now emit plain add/multiply/subtract instructions, without the jump to failure, as in the following example:\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Runtime.CompilerServices;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private readonly int[] _array = new int[99];\n    [Benchmark]\n    public int ArrayLengthPlusConstant() => AddToLength(_array);\n    [MethodImpl(MethodImplOptions.NoInlining)]\n    private static int AddToLength(int[] array) => checked(array.Length + 10);\n    [Benchmark]\n    public int GuardedLengthTimesConstant() => Multiply(_array);\n    [MethodImpl(MethodImplOptions.NoInlining)]\n    private static int Multiply(Span<int> span)\n    {\n        if (span.Length >= 100) return 0;\n        return checked(span.Length * 10);\n    }\n}\n```\n```\n; Arm64\n--- .NET 10\n+++ .NET 11\n G_M000_IG02:\n             cmp     w1, #100\n             bge     G_M000_IG05\n G_M000_IG03:\n             mov     w0, #10\n-            smull   x0, w1, w0\n-            lsr     x2, x0, #32\n-            cmp     w2, w0, ASR #31\n+            mul     w0, w1, w0\n-            bne     G_M000_IG07\n G_M000_IG04:\n             ldp     fp, lr, [sp], #0x10\n             ret     lr\n-G_M000_IG07:\n-            bl      CORINFO_HELP_OVERFLOW\n-            brk     #0\n-\n-; Total bytes of code 64\n+; Total bytes of code 44\n```\nThat makes the change broadly applicable: any time you write `checked` arithmetic on quantities that are inherently bounded, such as collection counts, lengths, or indices constrained by prior comparisons, the JIT now has a chance to prove at compile time that the overflow can’t happen and thus eliminate the run-time check entirely. Building on that range-check work, dotnet/runtime#124184 teaches the JIT to eliminate “narrowing casts” under the same kinds of guards:\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Runtime.CompilerServices;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private uint _value = 100;\n    [Benchmark]\n    public byte GuardedNarrowingCast() => Narrow(_value);\n    [MethodImpl(MethodImplOptions.NoInlining)]\n    private static byte Narrow(uint value)\n    {\n        if (value > 100) return 0;\n        return checked((byte)value);\n    }\n}\n```\n```\n; Arm64\n--- .NET 10\n+++ .NET 11\n@@ -4,23 +4,10 @@\n G_M000_IG02:\n             cmp     w0, #100\n-            bhi     G_M000_IG04\n-            cmp     w0, #255\n-            bhi     G_M000_IG06\n+            csel    w0, w0, wzr, ls\n G_M000_IG03:\n             ldp     fp, lr, [sp], #0x10\n             ret     lr\n-G_M000_IG04:\n-            mov     w0, wzr\n-\n-G_M000_IG05:\n-            ldp     fp, lr, [sp], #0x10\n-            ret     lr\n-\n-G_M000_IG06:\n-            bl      CORINFO_HELP_OVERFLOW\n-            brk     #0\n-\n-; Total bytes of code 52\n+; Total bytes of code 24\n```\nSuch use of `checked` is common in serialization and protocol code where you validate a value’s range prior to truncating it. In this benchmark I’ve used `checked` explicitly, but the more common form is with the whole project compiled with `<CheckForOverflowUnderflow>true</CheckForOverflowUnderflow>` in the .csproj, such that this `checked` becomes implicit. After the change, the range analysis sees that `value` is in the range `[0, 100]`, knows `byte` fits values up to 255, and elides the check.\n\ndotnet/runtime#128620 further teaches range analysis the possible results of leading-zero count, trailing-zero count, and population count instructions. Those results are often used to index small lookup tables… knowing their bounds lets the JIT remove the bounds check.\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Numerics;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private static readonly int[] s_lookup =\n        Enumerable.Range(0, 33).Select(i => i * i).ToArray();\n    private uint[] _values;\n    [GlobalSetup]\n    public void Setup()\n    {\n        Random rng = new(42);\n        _values = Enumerable.Range(0, 1024).Select(i => (uint)rng.Next(1, int.MaxValue)).ToArray();\n    }\n    [Benchmark]\n    public int SumLookupByLeadingZeroCount()\n    {\n        int sum = 0;\n        foreach (var v in _values)\n            sum += s_lookup[BitOperations.LeadingZeroCount(v)];\n        return sum;\n    }\n}\n```\nThe lookup improves because the JIT now knows `LeadingZeroCount(uint)` is between 0 and 32 and can remove the bounds check.\n\n| Method | Runtime | Mean | Ratio | \n|---|---|---|---|\n| SumLookupByLeadingZeroCount | .NET 10.0 | 516.1 ns | 1.00 | \n| SumLookupByLeadingZeroCount | .NET 11.0 | 438.5 ns | 0.85 | \n\nThe JIT is also able to conditionally apply range check-based elision via “cloning”. Cloning is a mechanism where the JIT takes one piece of code and duplicates it. One of the copies it leaves as it was originally, and the other copy it special cases. So, for example, if you had code like:\n\n`int value = array[i];`\nthe JIT could theoretically clone that in order to avoid the implicit bounds check, e.g.\n\n```\nint value;\nif ((uint)i < array.Length)\n{\n    // no bounds check emitted by JIT, e.g.\n    value = Unsafe.Add(ref MemoryMarshal.GetArrayDataReference(array), i);\n}\nelse\n{\n    // bounds check emitted\n    value = array[i];\n}\n```\nThat particular code looks silly, as we’re just trading an implicit bounds check for an explicit one. It becomes less silly when the JIT is able to elide multiple bounds checks with a single branch, e.g.\n\n```\nint sum;\nif (4 < array.Length)\n{\n    // zero bounds checks\n    ref int startRef = ref MemoryMarshal.GetArrayDataReference(array);\n    sum =\n        startRef +\n        Unsafe.Add(ref startRef, 1) +\n        Unsafe.Add(ref startRef, 2) +\n        Unsafe.Add(ref startRef, 3);\n}\nelse\n{\n    // potentially four bounds checks\n    sum =\n        array[0] +\n        array[1] +\n        array[2] +\n        array[3];\n}\n```\nSuch optimizations are already handled in the JIT, via its `optRangeCheckCloning` phase. It groups bounds checks from a basic block, emits one guard for the largest required range, and duplicates the affected code into a fast path where the individual checks can be removed and a fallback path where they remain. However, one long-standing limitation of range-check cloning is that it refused to process the last statement of any “terminator” block, a block that ends with a jump or return instruction. For a method like:\n\n`static int ArrayAccess(int[] abcd) => abcd[0] + abcd[1] + abcd[2] + abcd[3];`\nall four array accesses live in the return statement, the last statement of a return block, so nothing got cloned and the hot path retained four separate bounds checks. In .NET 11, dotnet/runtime#124705 removes that restriction, making the return statement eligible for range-check cloning and allowing a single fast-path guard to cover all four accesses.\n\nBut even without range-check cloning, there’s really no reason such accesses should require four bounds checks: the JIT should be able to see that the array or span needs to have a length of at least 4 and guard all accesses by that single check. If there were intervening operations that had side effects, the JIT would need to maintain order of operations, at least enough to maintain the observable behavior of those effects, but that’s not the case here. With dotnet/runtime#127439 in .NET 11, the JIT will now coalesce those checks within a basic block, strengthening the first check to the largest constant index and removing the rest.\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private readonly int[] _values = Enumerable.Range(0, 16).ToArray();\n    [Benchmark]\n    public int Sum16()\n    {\n        int[] values = _values;\n        return\n            values[0] + values[1] + values[2] + values[3] +\n            values[4] + values[5] + values[6] + values[7] +\n            values[8] + values[9] + values[10] + values[11] +\n            values[12] + values[13] + values[14] + values[15];\n    }\n}\n```\nIn previous releases, you’d sometimes see a proactive developer doing a similar optimization manually, e.g. reordering the accesses in an example like that to put the largest read first. That’s no longer necessary.\n\n| Method | Runtime | Mean | Ratio | \n|---|---|---|---|\n| Sum16 | .NET 10.0 | 2.958 ns | 1.00 | \n| Sum16 | .NET 11.0 | 1.828 ns | 0.62 | \n\nAnother bounds checking improvement comes in dotnet/runtime#127488, which actually targets explicitly-implemented bounds checks (rather than the implicit ones we’ve been discussing) and targets code that reads a fixed-size value from the end of a span, such as `BinaryPrimitives.ReadInt32BigEndian(span.Slice(span.Length - 4))` behind a `span.Length >= 4` guard, e.g.\n\n```\nif (span.Length >= 4)\n{\n    // Parse an int from the end of the span\n    ... = ReadInt32BigEndian(span.Slice(span.Length - 4));\n    ...\n}\n```\nThere shouldn’t be any additional bounds checking required here. However, `Span.Slice` begins with:\n\n```\nif ((uint)start > (uint)_length)\n    ThrowHelper.ThrowArgumentOutOfRangeException();\n```\nand `ReadInt32BigEndian` begins with:\n\n```\nif (sizeof(T) > source.Length)\n    ThrowHelper.ThrowArgumentOutOfRangeException();\n```\nso even though our `span.Length >= 4` check should have been sufficient, we’re still ending up with two additional checks. To address that, the JIT needed two things.\n\nFirst, it needed to be able to identify that `x - (x + a)` is the same as `-a`. Without this identity, `length - (length - 4)` is just an opaque subtraction of two expressions with no obvious constant result. With the identity, the JIT can recognize the inner expression `(length - 4)` as `length + (-4)`, apply `x - (x + a) == -a` with `x == length` and `a == -4`, and end up with `-(-4) == 4`. Now `ReadInt32BigEndian`‘s check against 4 becomes `4 >= 4`, which the JIT can trivially see is true.\n\nSecond, `Slice(start)` must establish that `start` is between zero and the span’s length. When `start` is `length - 4`, the existing `length >= 4` guard proves the result is non-negative, while subtracting a positive constant means the result can’t exceed `length`. The improved range analysis connects that guard to the subtraction and removes `Slice`‘s check.\n\nBoth fixes together mean the above example now elides both extra bounds checks. That’s useful in particular for libraries like parsers, network protocol implementations, and cryptographic code, all of which frequently on hot paths do things like “read the last N bytes of a buffer.”\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Buffers.Binary;\nusing System.Runtime.CompilerServices;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser]\n[HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private readonly byte[] _buffer = new byte[64];\n    [Benchmark]\n    public int ReadLastInt32() => ReadLastInt32(_buffer);\n    [MethodImpl(MethodImplOptions.NoInlining)]\n    private static int ReadLastInt32(ReadOnlySpan<byte> span)\n    {\n        if (span.Length >= sizeof(int))\n        {\n            return BinaryPrimitives.ReadInt32BigEndian(span.Slice(span.Length - sizeof(int)));\n        }\n        return -1;\n    }\n}\n```\nIn .NET 10, the helper is 73 bytes and includes both additional checks and their throw paths:\n\n```\n; x64\ncmp       ecx,4\njl        RETURN_MINUS_ONE\nlea       edx,[rcx-4]\ncmp       edx,ecx\nja        THROW_SLICE\nmov       r8d,edx\nadd       rax,r8\nsub       ecx,edx\ncmp       ecx,4\njl        THROW_READ\nmovbe     eax,[rax]\n```\nIn .NET 11, the helper is 28 bytes, and only the original length guard remains:\n\n```\n; x64\ncmp       ecx,4\njl        RETURN_MINUS_ONE\nadd       ecx,-4\nadd       rax,rcx\nmovbe     eax,[rax]\n```\ndotnet/runtime#122040 and dotnet/runtime#127117 similarly help to remove bounds checks involving `span.Slice`. Vectorized loops often work through a span a chunk at a time, slicing off the elements they’ve already processed. The JIT hasn’t always been able to keep track of how those progressively smaller slices relate to the original span, so it could end up checking the same limits again on each iteration. These changes improve that tracking, enabling more of those repeated checks to be removed.\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing System.Runtime.CompilerServices;\nusing System.Runtime.Intrinsics;\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private readonly int[] _data = Enumerable.Repeat(1, 1_024).ToArray();\n    [Benchmark]\n    public Vector256<int> CreateFromSlice() => CreateFromSlice(_data);\n    [Benchmark]\n    public int SumSliced() => SumSliced(_data);\n    [MethodImpl(MethodImplOptions.NoInlining)]\n    private static Vector256<int> CreateFromSlice(Span<int> values)\n    {\n        if (values.Length < 16)\n            return default;\n        return Vector256.Create(values.Slice(8));\n    }\n    [MethodImpl(MethodImplOptions.NoInlining)]\n    private static int SumSliced(ReadOnlySpan<int> data)\n    {\n        Vector128<int> sum = default;\n        while (data.Length >= Vector128<int>.Count)\n        {\n            sum += Vector128.Create(data);\n            data = data.Slice(Vector128<int>.Count);\n        }\n        int result = Vector128.Sum(sum);\n        foreach (int value in data)\n            result += value;\n        return result;\n    }\n}\n```\nIn .NET 10, the loop condition proves that at least one vector remains, but the construction of the vector from the current span performs the same check again. .NET 11 retains the length relationship, so the loop body begins directly with the vector addition:\n\n```\n; x64, vector loop\n-cmp       esi, 4\n-jl        THROW_ARGUMENT_OUT_OF_RANGE\n-vpaddd    xmm6, xmm6, [rbx]\n-add       rbx, 10\n-add       esi, 0FFFFFFFC\n-cmp       esi, 4\n+vpaddd    xmm0, xmm0, [rax]\n+add       rax, 10\n+add       ecx, 0FFFFFFFC\n+cmp       ecx, 4\n jge       LOOP\n```\nWe saw earlier how range-check cloning enables duplicating a sequence of instructions in order to eliminate bounds checks. “Loop cloning” extends that to a whole loop. Consider a loop that processes the first `count` elements of an array:\n\n```\nfor (int i = 0; i < count; i++)\n    sum += values[i];\n```\nThe test `i < count` doesn’t by itself prove that `i < values.Length`, so by default the compilation would need a bounds check in the body, which would mean a bounds check for every `values[i]` access. Loop cloning gives the JIT another option. Instead of generating the equivalent of:\n\n```\nfor (int i = 0; i < count; i++)\n    sum += values[i]; // bounds check!\n```\nit can generate the equivalent of:\n\n```\nif ((uint)count <= (uint)values.Length)\n{\n    // no bounds checks\n    ref int startRef = ref MemoryMarshal.GetArrayDataReference(values);\n    for (int i = 0; i < count; i++)\n    {\n        sum += Unsafe.Add(ref startRef, i);\n    }\n}\nelse\n{\n    // bounds check per iteration\n    for (int i = 0; i < count; i++)\n    {\n        sum += values[i];\n    }\n}\n```\nFor the common case where the iteration is in bounds, execution proceeds through a cloned loop with no per-iteration bounds checks, whereas the original checked loop remains as the fallback that preserves exceptional behavior for invalid inputs. The normal path pays for one guard and avoids a check on every iteration, but that comes at the expense of duplicating code. The JIT therefore needs to apply the optimization selectively.\n\nThe JIT has long employed loop cloning, but it didn’t always kick in even in cases it seemed applicable. The previous example showed loop cloning with `<` in the iteration condition. For whatever reason, however, some developers used `!=`, and loop cloning didn’t apply (I’m guessing they used `!=` because they thought it was more efficient, and they actually end up deoptimizing). Thanks to dotnet/runtime#129268, in .NET 11 `!=` is now also handled, as long as specific conditions are met, such as the stride being exactly 1 or -1, e.g. `i++` qualifies, while `i += 2` doesn’t. dotnet/runtime#129303 also improves loops that terminate with `i != bound`, giving the JIT a tighter understanding of the values `i` can take and allowing it to remove some bounds checks even when it can’t clone the whole loop.\n\nLookahead in arrays and spans is another recurring pattern, especially in parsers. dotnet/runtime#124242 and dotnet/runtime#125235 recognize conditions such as `(uint)(i + 2) < (uint)span.Length` and use that relation to remove the follow-on checks for `span[i + 1]` and `span[i + 2]`. Consider this benchmark:\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System;\nusing System.Linq;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private readonly string _text = string.Concat(Enumerable.Repeat(\"%FE\", 128));\n    [Benchmark]\n    public bool ContainsPercentFF()\n    {\n        ReadOnlySpan<char> span = _text;\n        for (int i = 0; i < span.Length; i++)\n        {\n            if (span[i] == '%' &&\n                (uint)(i + 2) < (uint)span.Length &&\n                span[i + 1] == 'F' &&\n                span[i + 2] == 'F')\n            {\n                return true;\n            }\n        }\n        return false;\n    }\n}\n```\n| Method | Runtime | Mean | Ratio | \n|---|---|---|---|\n| ContainsPercentFF | .NET 10.0 | 232.1 ns | 1.00 | \n| ContainsPercentFF | .NET 11.0 | 194.9 ns | 0.84 | \n\nSeveral smaller changes broaden the range of code from which the JIT can remove bounds checks:\n\n- dotnet/runtime#121640 helps in a situation where once an access using a chosen index has been checked, a later access to the same array at that index need not be checked again.\n- dotnet/runtime#121683 enables the JIT to trace an array’s length through calculations performed earlier in the method, exposing more redundant checks, including some involving index-from-end expressions.\n- dotnet/runtime#124387 and dotnet/runtime#130326 teach the optimizer to rely on a span’s length always being non-negative.\n- dotnet/runtime#124571 improves sequences of index-from-end accesses: once an access like `arr[^4]` establishes that the array has at least four elements, the JIT reuses that information for nearby accesses such as`arr[^3]` .\n- dotnet/runtime#129101 improves how the JIT combines and carries forward the possible ranges of arithmetic expressions, including expressions involving bitwise OR and unsigned division. Those tighter ranges can show that more values are non-negative or within bounds.\n\nBounds-check elimination is only one payoff from understanding a loop’s structure. The JIT analyzes induction variables (values like loop counters that change predictably each iteration) and puts loops into standard forms so that later optimizations can reason about them. .NET 11 broadens the range of loops for which that works:\n\n- dotnet/runtime#122184 recognizes another representation of a 32-to-64-bit zero extension. That lets pointer loops using expressions such as `data[(uint)i]` replace the repeated index extension and address calculation with a pointer increment.\n- dotnet/runtime#119537 follows simple control-flow predecessors when finding an induction variable’s initialization and zero-trip test, while dotnet/runtime#128303 gives loops with multiple backedges a single canonical latch block.\n- dotnet/runtime#128532 makes loop cloning tolerate more statements around the update and test.\n- dotnet/runtime#129309 extends cloning to more span loops with non-unit strides and offset limits.\n- dotnet/runtime#129349 handles large strides in array loops with an explicit safety guard rather than rejecting them outright.\n- dotnet/runtime#129472 allows loop inversion to spend more of its budget on likely cloning candidates.\n- dotnet/runtime#130205 removes comparisons that are redundant given the induction variable’s known range.\n- dotnet/runtime#131362 corrects profile weights after inversion changes a loop’s exit.\n\nMuch of this work wasn’t motivated by contrived benchmarks containing nothing\nbut array indexing as I’m prone to use in these posts. Rather, many of the improvements\nstemmed from an ongoing audit of unsafe code throughout the\n.NET libraries, part of a broader\neffort to improve memory safety in .NET. .NET and C# are memory safe, but as with other memory safe languages like Rust,\nit provides escape hatches that enable turning off the guardrails provided by the compiler and runtime.\nThis effort is about reducing where and when developers feel compelled to use those escape hatches, since every\noccurrence is an opportunity for increased risk. Unsafe code was often introduced years earlier to manually avoid bounds\nchecks, typically by walking a buffer with pointers, byrefs, or `Unsafe.Add`.\nSometimes the audit found that the unsafe code was no longer needed and could\nsimply be removed. Sometimes a “safe” rewrite (meaning not using `unsafe` and friends) was already just as fast or even faster.\nAnd sometimes the rewrite exposed an optimization the JIT was missing, in which\ncase the answer was to improve the JIT and then rewrite the library code to use\nnormal, bounds-checked C#. Several of the optimizations discussed in this\nsection are the result of exactly that feedback loop. dotnet/runtime#127429 is a\nparticularly nice example. The vectorized implementation of `Enumerable.Sum`\nused `MemoryMarshal.GetReference`, `Vector.LoadUnsafe`, and `Unsafe.Add` to walk\nits input without bounds checks. With the span-slicing improvements described\nearlier, it could instead use `Vector.Create(span)`, `span.Slice(...)`, and a\n`foreach` for the tail. That’s easier to reason about, removes the unchecked\nindexing, and ended up being faster.\ndotnet/runtime#114757 similarly\nreplaced an unsafe pointer-based header-name accessor with a generic\n`ReadOnlySpan<T>` implementation without loss of performance.\nSimilarly, dotnet/runtime#121270 removed\nmore unsafe code from `Uri` parsing and actually improved performance of the cited code measurably.\n\nThere’s a useful “go do” here for libraries outside of dotnet/runtime, as well.\nUnsafe code written to work around the JIT is a snapshot of what the JIT could\ndo at the time that code was written. If you own code that has hand-written pointer or\n`Unsafe`-based loops whose purpose is to avoid bounds checks, it’s worth\nrewriting them with safe, bounds-checked C# and measuring again on .NET 11.\nChances are, you’ll find the gap at this point is either non-existent or small\nenough that it’s not worth the increased maintenance and risk for managing the safety yourself.\nAnd if the revised version is still slower, that’s a great opportunity for you to share a repro\nin the dotnet/runtime repo, hopefully serving as inspiration for one of the first performance improvements\nto go into the JIT for .NET 12. `unsafe` code is still necessary for scenarios like interop,\nbut performance alone shouldn’t be a permanent reason to eschew all the valuable guardrails .NET provides.\n\nThe C# 15 memory-safety preview\npushes in the same direction and is part and parcel of this effort. Historically, C# has largely equated pointers\nwith unsafe code: simply declaring or manipulating a pointer generally required\nan `unsafe` context, even if the code never accessed the memory to which it\npoints. In the preview, pointer plumbing such as declaring a pointer, taking an\naddress with `&`, using `fixed`, converting `stackalloc` to a pointer, and\napplying `sizeof` to an unmanaged type no longer requires an `unsafe` context.\nOperations that actually access the pointed-to memory, including `*p`,\n`p->member`, and `p[i]`, still do. C# 15 also adds an `unsafe(expression)` form,\nanalogous to `checked(expression)`, so an unsafe context can cover one precise\nexpression rather than a larger statement block. Those changes are the first preview slice of a larger, multi-release\nunsafe evolution.\nThe end goal is to make unsafe regions smaller, make their assumptions visible through the call graph,\nand make them easier for reviewers and tools to find. Pairing that with a JIT\nthat makes idiomatic safe code fast removes a lot of the historical pressure to\nuse unsafe code in the first place.\n\n### Assertion Propagation\n\nAs discussed earlier, the JIT continually learns facts while compiling a method: a value equals a constant, a reference isn’t null, an integer falls within a particular range, and so on. “Assertion propagation” carries those facts forward so they can simplify later code. “Value numbering” complements it by letting the JIT recognize when two expressions compute the same value, even if they appear in different places or use different variables. Together, these mechanisms enable optimizations such as removing redundant null and bounds checks, folding conditions to constants, and reusing repeated computations. .NET 11 improves assertion propagation primarily by fixing places where useful facts were either never recorded or weren’t recognized later.\n\nFor example, reading an array’s length normally carries an implicit null-check: if the array reference is `null`, the read must throw. Once global assertion propagation already knows the reference is non-null, however, we should be able to avoid the implicit null check. In .NET 11, dotnet/runtime#124291 takes care of that for `Array.Length`:\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Runtime.CompilerServices;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private readonly int[] _values = new int[1024];\n    [Benchmark]\n    public void DeadLength() => Test(_values);\n    [MethodImpl(MethodImplOptions.NoInlining)]\n    private static void Test(int[]? values)\n    {\n        if (values is not null)\n            _ = values.Length;\n    }\n}\n```\nThe .NET 10 code still tests the reference and reads the length. In .NET 11, the guard proves the read can’t throw, and since its result isn’t used, the access disappears:\n\n```\n; Arm64\n--- .NET 10\n+++ .NET 11\n G_M000_IG01:\n             stp     fp, lr, [sp, #-0x10]!\n             mov     fp, sp\n G_M000_IG02:\n-            cbz     x0, G_M000_IG04\n-\n-G_M000_IG03:\n-            ldr     wzr, [x0, #0x08]\n-\n-G_M000_IG04:\n             ldp     fp, lr, [sp], #0x10\n             ret     lr\n-; Total bytes of code 24\n+; Total bytes of code 16\n```\ndotnet/runtime#119474 improves\nthe starting point for integer range analysis. The JIT now uses facts inherent\nin a value itself, e.g. a constant has one exact value, while a value converted\nto `byte`, for example, must be between 0 and 255. That can eliminate bounds\nchecks and conditions even when no preceding `if` explicitly established the\nrange. dotnet/runtime#124415\nfurther refines this handling of casts, combining what is known about both the\nsource value and the destination type to derive the tightest useful range.\n\nThose improvements derive ranges from facts inherent in a value, but ranges\ncan also come from control flow. After `if ((uint)x < 10)`, for example, the\nJIT knows that `x` is between 0 and 9 on the true path, which may be enough to\nremove a later comparison or array bounds check.\ndotnet/runtime#123624\nderives tighter ranges from assertions and casts, including proving that some\ncomparisons are always true or false.\ndotnet/runtime#129390\npreserves range information more accurately when control-flow paths merge.\n\nOther changes make better use of the ranges once known. dotnet/runtime#129354 traces values back through their definitions to fold more span- and slice-related comparisons, and dotnet/runtime#126917 uses narrowed ranges to remove more relational branches.\n\ndotnet/runtime#124711 teaches the JIT to learn implicit facts from operations that have already completed successfully. For example:\n\n- Creating an array proves its requested length wasn’t negative.\n- A reference-array store may need a runtime covariance check, because a value typed as `object[]` can actually refer to a`string[]` ; the helper that performs that type check also validates the index, so if it returns successfully, the index was in range.\n- Integer division or modulo proves the divisor wasn’t zero.\n\nAnd so on. Those facts can then remove redundant checks and conditions later in the method.\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private readonly object?[] _objArr = new object?[8];\n    private readonly object _value = new();\n    [Benchmark]\n    public object? CovariantArrayStore()\n    {\n        object?[] objArr = _objArr;\n        objArr[3] = _value;\n        return objArr[2];\n    }\n}\n```\nA successful store to element 3 proves that particular array has at least four elements; since an array’s length can’t change, the subsequent read of element 2 doesn’t need another bounds check.\n\n| Method | Runtime | Mean | Ratio | \n|---|---|---|---|\n| CovariantArrayStore | .NET 10.0 | 3.565 ns | 1.00 | \n| CovariantArrayStore | .NET 11.0 | 2.985 ns | 0.84 | \n\ndotnet/runtime#128522 simplifies how the global assertion pass identifies values, making it less likely to miss a fact learned earlier. One practical impact of this is better propagation of a static `string`‘s known length, which can turn a general string comparison into a fixed-size vectorized comparison.\n\ndotnet/runtime#127810 improves null-check elimination where control flow merges. With `??=`, which is a very common operator used for lazy initialization, the resulting value is non-null whether it came from the existing field or from the newly allocated object. The JIT now combines the facts from both paths and recognizes that the subsequent call doesn’t need another null check.\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Runtime.CompilerServices;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    Inner? _inner;\n    [Benchmark]\n    [Arguments(42)]\n    public int Invoke(int n) => (_inner ??= new()).Increment(n);\n    private sealed class Inner\n    {\n        [MethodImpl(MethodImplOptions.NoInlining)]\n        public int Increment(int n) => n + 1;\n    }\n}\n```\nThe generated code consequently loses the null check on the merged value:\n\n```\n; x64\n M00_L00:\n        mov      edx, esi\n-       cmp      [rcx], ecx\n        call     qword ptr [...] ; Inner.Increment(Int32)\n-; Total bytes of code 75\n+; Total bytes of code 73\n```\ndotnet/runtime#128701 removes similarly redundant null checks from copies of structs that contain object references. Such copies use a runtime helper so the garbage collector is correctly notified about the reference writes, but lowering had been adding probes for both source and destination without preserving whether either address could actually fault. It now emits only the probes that are needed.\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private FourRefs _src = new()\n    {\n        A = new(),\n        B = new(),\n        C = new(),\n        D = new()\n    };\n    private FourRefs _dst;\n    [Benchmark]\n    public void BulkStructCopy() => _dst = _src;\n    private struct FourRefs\n    {\n        public object? A;\n        public object? B;\n        public object? C;\n        public object? D;\n    }\n}\n```\nBoth `_src` and `_dst` are fields of the same object, so after probing the source address has established that the object isn’t null, probing the destination address can’t provide any additional information. .NET 11 removes that second probe:\n\n```\n; Arm64\n G_M000_IG02:\n             add     x1, x0, #8\n             ldrsb   wzr, [x1]\n             add     x0, x0, #40\n-            ldrsb   wzr, [x0]\n             movz    x2, ...\n             ldr     x3, [x2]\n             mov     x2, #32\n             blr     x3      // CORINFO_HELP_BULK_WRITEBARRIER\n-; Total bytes of code 56\n+; Total bytes of code 52\n```\nAdditionally, dotnet/runtime#125215 lets the JIT retain and efficiently find more assertions in larger methods, increasing the opportunities for the same kinds of simplification. And dotnet/runtime#129312 removes unnecessary temporary variables when the same simple field address is used multiple times, enabling more efficient loads and stores.\n\n### Simplification\n\nAssertion propagation is largely about proving things to help the generated code. Once the JIT knows enough about an operation’s inputs, it can often replace the operation with something simpler and cheaper.\n\n“Constant folding” is a fancy way of saying the compiler does work once so it doesn’t need to be repeated at run time. If the compiler has everything it needs to compute an answer when building, it can bake that answer in to the generated code and avoid needing the code to re-compute it. That answer can then be further used by other computations at build time, potentially folding further. The C# compiler handles constant folding expressions composed entirely of language constants, while the JIT compiler can go further after inlining and after learning things about values and control flow. The JIT already does a ton of folding, and as with every release, it goes further in .NET 11.\n\nOne straightforward example is the offset of a field within a struct. dotnet/runtime#122297 recognizes more cases where two addresses refer to the same struct and replaces their difference with the known field offset. Here, the second `int` field begins four bytes into the struct:\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Runtime.CompilerServices;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic unsafe class Benchmarks\n{\n    private struct MyStruct\n    {\n        public int A;\n        public int Field;\n    }\n    [MethodImpl(MethodImplOptions.AggressiveInlining)]\n    private static nint OffsetOfFieldInline()\n    {\n        MyStruct dummy;\n        return (nint)((byte*)&dummy.Field - (byte*)&dummy);\n    }\n    [Benchmark]\n    [Arguments(1_000)]\n    public nint OffsetOfFieldLoop(int n)\n    {\n        nint sum = 0;\n        for (int i = 0; i < n; i++)\n            sum += OffsetOfFieldInline();\n        return sum;\n    }\n}\n```\nWithout the fold, the loop repeatedly computes the field offset. With the fold, each iteration simply adds the constant `4`.\n\n```\n; Arm64\n--- .NET 10\n+++ .NET 11\n@@ -1,7 +1,6 @@\n G_M000_IG01:\n-            stp     fp, lr, [sp, #-0x20]!\n+            stp     fp, lr, [sp, #-0x10]!\n             mov     fp, sp\n-            str     xzr, [fp, #0x18]\n G_M000_IG02:\n             mov     x0, xzr\n@@ -9,22 +8,18 @@\n             ble     G_M000_IG05\n G_M000_IG03:\n-            add     x2, fp, #0x1C\n-            add     x3, fp, #24\n-            sub     x2, x2, x3\n             align   [0 bytes for IG04]\n             align   [0 bytes]\n             align   [0 bytes]\n             align   [0 bytes]\n G_M000_IG04:\n-            str     xzr, [fp, #0x18]\n-            add     x0, x2, x0\n+            add     x0, x0, #4\n             sub     w1, w1, #1\n             cbnz    w1, G_M000_IG04\n G_M000_IG05:\n-            ldp     fp, lr, [sp], #0x20\n+            ldp     fp, lr, [sp], #0x10\n             ret     lr\n-; Total bytes of code 60\n+; Total bytes of code 40\n```\ndotnet/runtime#121985 from @hez2010 enables the JIT to evaluate `SequenceEqual` at compile time when both inputs are known. `SequenceEqual` normally walks two sequences element by element, stopping at the first mismatch. But if inlining exposes both sequences as constants, there’s nothing useful left to do at run time: the JIT can compare them while compiling and replace the whole operation with a constant `true` or `false`. This intrinsic underpins APIs including `MemoryExtensions.SequenceEqual`, `ReadOnlySpan<T>.SequenceEqual`, and `string.Equals`.\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private static string AlphaLower => \"abcdefghijklmnopqrstuvwxyz\";\n    private static string AlphaUpper => \"ABCDEFGHIJKLMNOPQRSTUVWXYZ\";\n    [Benchmark]\n    public bool CompareEqual() => AlphaLower.Equals(AlphaLower);\n    [Benchmark]\n    public bool CompareDistinct() => AlphaLower.Equals(AlphaUpper);\n}\n```\nBecause these properties aren’t `const`, the C# compiler can’t evaluate the comparisons. The JIT, however, can see the string literals after inlining. It now folds comparisons of the same input whose contents are available at the time of compilation. `CompareDistinct` therefore becomes a constant `false`.\n\n```\n; x64\n--- .NET 10\n+++ .NET 11\n-mov       rax,LOWER_STRING\n-mov       rcx,UPPER_STRING\n-add       rax,0C\n-vmovups   ymm0,[rax]\n-vmovups   ymm1,[rax+14]\n-vmovups   ymm2,[rcx]\n-vpxor     ymm0,ymm2,ymm0\n-vpxor     ymm1,ymm1,[rcx+14]\n-vpor      ymm0,ymm1,ymm0\n-vptest    ymm0,ymm0\n-sete      al\n-movzx     eax,al\n-vzeroupper\n+xor       eax,eax\n ret\n-; Total bytes of code 65\n+; Total bytes of code 3\n```\nFolding an operation is only the first step, though. The result can then simplify later code, even when it’s a vector. dotnet/runtime#127124 extends assertion propagation to 128-bit integer vector constants. If a branch establishes that a vector is zero, uses of that vector within the branch can now be replaced with zero and simplified just like scalar values.\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Runtime.CompilerServices;\nusing System.Runtime.Intrinsics;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private int _selector;\n    [MethodImpl(MethodImplOptions.NoInlining)]\n    private Vector128<int> Compute() => _selector == 0 ? Vector128<int>.Zero : Vector128.Create(7);\n    [Benchmark]\n    public int AndNotIfZero()\n    {\n        Vector128<int> v = Compute();\n        if (v == Vector128<int>.Zero)\n        {\n            Vector128<int> masked = Vector128.AndNot(v, Vector128.Create(0x00FF00FF));\n            return masked[0];\n        }\n        return -1;\n    }\n}\n```\nIn the benchmark’s zero branch, the JIT can now fold away the mask creation, `AndNot`, and lane extraction, reducing the Arm64 method from 68 bytes to 56 bytes. This currently applies to integer vectors up to 128 bits (floating-point equality has additional NaN and signed-zero semantics that prevent the same reasoning at present).\n\n```\n; Arm64\n--- .NET 10\n+++ .NET 11\n@@ -11,14 +11,11 @@\n             umaxp   v16.4s, v0.4s, v0.4s\n             umov    x0, v16.d[0]\n             movn    w1, #0\n-            movi    v16.8h, #0xFF,  LSL #8\n-            and     v16.4s, v0.4s, v16.4s\n-            smov    x2, v16.s[0]\n             cmp     x0, #0\n-            csel    w0, w1, w2, ne\n+            cinc    w0, w1, eq\nG_M000_IG03:\n             ldp     fp, lr, [sp], #0x10\n             ret     lr\n-; Total bytes of code 68\n+; Total bytes of code 56\n```\nTwo backend cleanups take advantage of simpler expressions. dotnet/runtime#124332 from @jonathandavies-arm removes an unnecessary negation when Arm64 code compares a negated value with zero. And dotnet/runtime#124642 from @yykkibbb lets short-circuit Boolean returns fold even when inlining has left unused writes in the same block; those stores previously obscured the simple Boolean expression from the optimizer.\n\nBranches offer another opportunity for simplification. Modern processors work on several instructions at different stages at the same time. When a processor encounters a conditional branch, it predicts which path will be taken so that it can continue fetching and executing instructions speculatively. A correct prediction hides much of the branch’s cost. A misprediction throws away that speculative work, redirects instruction fetch to the correct path, and refills the processor’s execution pipeline. That can make the predictability of a branch as important as the work in either branch. The JIT can sometimes avoid that variability, particularly inside small hot loops, by replacing a branch with a conditional move instruction or by recognizing that several branches describe one simpler condition. This isn’t always profitable: branchless code may evaluate work that a predictable branch would skip, making the branching code less expensive in the majority case. But it can be valuable for small, data-dependent choices.\n\ndotnet/runtime#124567\nrecognizes zero-based equality chains, e.g.\n`value == 0 || value == 1 || value == 2`. Such chains can be replaced with an\nunsigned range check, e.g. `(uint)value <= 2`, producing a branchless result.\nThe unsigned comparison also handles negative inputs: when interpreted as\nunsigned, any negative `int` is larger than the upper bound.\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser]\n[HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private int _value = 2;\n    [Benchmark]\n    public bool IsLetterCategory() =>\n        _value == 0 ||\n        _value == 1 ||\n        _value == 2 ||\n        _value == 3 ||\n        _value == 4;\n}\n```\nThe .NET 10 JIT already combines the first four comparisons, but still needs\na branch and a separate comparison for `4`:\n\n```\n; x64\nmov       ecx,[rcx+8]\ncmp       ecx,3\nja        CHECK_FOUR\nmov       eax,1\nret\nCHECK_FOUR:\ncmp       ecx,4\nsete      al\nmovzx     eax,al\nret\n```\n.NET 11 recognizes the whole chain as one unsigned range check, reducing the method from 24 bytes to 13:\n\n```\n; x64\nmov       eax,[rcx+8]\ncmp       eax,5\nsetb      al\nmovzx     eax,al\nret\n```\ndotnet/runtime#128524 from @BoyBaykiller extends the same optimization to contiguous ranges that don’t start at zero. For example, `x == 3 || x == 4 || x == 5` can become `(uint)(x - 3) <= 2`.\n\nCasts can obscure an equally simple comparison. dotnet/runtime#128091 from @BoyBaykiller broadens cast-comparison optimization to equality and inequality. In this benchmark, converting a `uint` to `ulong` adds no information needed to compare it with `uint.MaxValue`, so the JIT can keep the comparison at 32 bits:\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Linq;\nusing System.Runtime.CompilerServices;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private readonly int[] _values = Enumerable.Range(0, 128).ToArray();\n    [Benchmark]\n    public int CastEquality()\n    {\n        int matches = 0;\n        foreach (int value in _values)\n            if ((ulong)(uint)value == uint.MaxValue)\n                matches++;\n        return matches;\n    }\n}\n```\nThe widening cast disappears, reducing the Arm64 method from 80 bytes to 76 bytes.\n\n```\n; Arm64\n--- .NET 10\n+++ .NET 11\n@@ -18,8 +18,7 @@\n G_M000_IG04:\n             ldr     w3, [x0]\n-            mov     x4, #0xFFFFFFFF\n-            cmp     x3, x4\n+            cmn     w3, #1\n             beq     G_M000_IG08\n G_M000_IG05:\n@@ -38,4 +37,4 @@\n             add     w1, w1, #1\n             b       G_M000_IG05\n-; Total bytes of code 80\n+; Total bytes of code 76\n```\nThe examples thus far simplify individual comparisons. dotnet/runtime#127181 also combines multiple comparisons in the same expression. For example, `(x >= c) && (x <= c)` can only be true when `x == c`; corresponding OR forms can be simplified similarly.\n\nOnce the JIT can reason about one comparison in terms of another, it can apply the same idea across branches. dotnet/runtime#126587 removes an earlier test when a later, stronger test subsumes it. For example, `if (x > 0) if (x > 1)` needs only the `x > 1` test, as reaching the nested body with `x > 1` necessarily also means `x > 0`.\n\nRather than simply removing a test, the JIT can sometimes use the outcome of an earlier branch to choose the destination of a later one. This is known as “jump threading”: the JIT threads a control-flow path through the intervening jumps directly to its eventual destination. For example, consider:\n\n```\nint value = condition ? 1 : 2;\nif (value == 1)\n{\n    One();\n}\nelse\n{\n    Two();\n}\n```\nThe path where `condition` is true can go directly to `One`, while the false path can go directly to `Two`, eliminating the second test, effectively:\n\n```\nint value;\nif (condition)\n{\n    value = 1;\n    One();\n}\nelse\n{\n    value = 2;\n    Two();\n}\n```\ndotnet/runtime#126812 lets this continue through more places where paths rejoin, and dotnet/runtime#127103 ensures the rewritten values remain correct in more of those cases. dotnet/runtime#127950 carries relationships between values further, so facts like `a > 10` and `b > a` can simplify later branches or bounds. The same reasoning can apply to type information. dotnet/runtime#128500 combines the known types of instances arriving from multiple paths; if every value derives from the tested base type, the JIT can remove the `is` test after the paths merge. And dotnet/runtime#127434 from @hez2010 lets redundant-branch elimination look through empty jump blocks. Such a block contains no work of its own and exists only to redirect control elsewhere, but it could still hide the relationship between two conditions from the optimizer. Consider this benchmark:\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Runtime.CompilerServices;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private static object? s_sink;\n    private int _x = 20;\n    private int _y = 30;\n    private bool _flag = true;\n    private int _count = 5;\n    [Benchmark]\n    public bool TransitiveComparison() => TransitiveComparison(_x, _y);\n    [Benchmark]\n    public bool MergedTypeCheck() => MergedTypeCheck(_flag);\n    [Benchmark]\n    public int NestedThresholds() => NestedThresholds(_count);\n    [MethodImpl(MethodImplOptions.NoInlining)]\n    private static bool TransitiveComparison(int x, int y)\n    {\n        if (x > 10 && x < 100 && y > x)\n            return y > 0;\n        return false;\n    }\n    [MethodImpl(MethodImplOptions.NoInlining)]\n    private static bool MergedTypeCheck(bool flag)\n    {\n        object shape = flag ? new Circle() : new Rectangle();\n        s_sink = shape;\n        return shape is Shape;\n    }\n    [MethodImpl(MethodImplOptions.NoInlining)]\n    private static int NestedThresholds(int count)\n    {\n        if (count > 1)\n            if (count > 2)\n                if (count > 3)\n                    if (count > 4)\n                        return 1;\n        return 3;\n    }\n    private abstract class Shape;\n    private sealed class Circle : Shape;\n    private sealed class Rectangle : Shape;\n}\n```\nIn `TransitiveComparison`, reaching `y > 0` means the JIT already knows that `x > 10` and `y > x`, which together prove that `y` is positive. The final comparison disappears, reducing the Arm64 method from 40 bytes to 36 bytes:\n\n```\n; Arm64\n--- .NET 10\n+++ .NET 11\n             cmp     w1, w0\n             ccmp    w2, w3, c, gt\n-            ccmp    w1, #0, nzc, ls\n-            cset    x0, gt\n+            cset    x0, ls\n-; Total bytes of code 40\n+; Total bytes of code 36\n```\nIn `MergedTypeCheck`, each path creates a different concrete type, but both derive from `Shape`. .NET 11 keeps the allocations and the store that make the example observable, but replaces the `is` helper call and its result test with the constant `true`, reducing the method from 112 bytes to 88 bytes:\n\n```\n; Arm64\n--- .NET 10\n+++ .NET 11\n             bl      CORINFO_HELP_ASSIGN_REF\n-            movz    x0, #0xEA30\n-            movk    x0, #0x4EB LSL #16\n-            movk    x0, #0x7FFF LSL #32\n-            bl      CORINFO_HELP_ISINSTANCEOFCLASS\n-            cmp     x0, #0\n-            cset    x0, ne\n+            mov     w0, #1\n-; Total bytes of code 112\n+; Total bytes of code 88\n```\nFor `NestedThresholds`, reaching the `return 1` requires `count` to be greater than all four constants, which is equivalent to just `count > 4`. Once redundant-branch elimination can see through the empty jump blocks left behind while simplifying the nested conditions, the other three comparisons disappear:\n\n```\n; Arm64\n--- .NET 10\n+++ .NET 11\n             mov     w1, #3\n             mov     w2, #1\n-            cmp     w0, #1\n-            ccmp    w0, #2, nzc, gt\n-            ccmp    w0, #3, nzc, gt\n-            ccmp    w0, #4, nzc, gt\n+            cmp     w0, #4\n             csel    w0, w1, w2, le\n-; Total bytes of code 44\n+; Total bytes of code 32\n```\nRemoving a redundant branch is ideal; why do work when it’s provably unnecessary? Often, however, the branch is necessary, as both outcomes are possible (or at least not provably impossible). In such cases, the JIT may still be able to avoid branching via specialized instructions that bake the choice into the instruction. “If-conversion” replaces a small `if`/`else` with a conditional-move instruction or another branchless form when both alternatives are cheap. The JIT has been able to do this for several releases, and improves in .NET 11. dotnet/runtime#124738 from @BoyBaykiller recognizes an earlier default assignment as the implicit `else`, so `bool x = false; if (cond) x = true;` can become the same branchless form as an explicit `else`. dotnet/runtime#127915 from @BoyBaykiller handles the opposite cleanup, removing a conditional selection when both outcomes are the same constant while preserving any side effects from evaluating the condition. dotnet/runtime#128533 from @BoyBaykiller also helps these Boolean optimizations meet in the middle by normalizing power-of-two bit tests. A power of two has exactly one bit set, in which case `(A & bit) == bit` is equivalent to `(A & bit) != 0`; putting both forms into the same canonical representation makes them easier to combine with surrounding conditions. All three improvements are visible in the following benchmarks:\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Runtime.CompilerServices;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private int _left = 1;\n    private int _right = 2;\n    private double _double = 0.0;\n    private int _bits = 4;\n    [Benchmark]\n    public bool ImplicitElse() => ImplicitElse(_left, _right);\n    [Benchmark]\n    public bool IsDefaultValue() => IsDefaultValue(_double);\n    [Benchmark]\n    public bool HasEitherBit() => HasEitherBit(_bits);\n    [MethodImpl(MethodImplOptions.NoInlining)]\n    private static bool ImplicitElse(int left, int right)\n    {\n        bool leftIsSmaller = false;\n        if (left < right)\n            leftIsSmaller = true;\n        return leftIsSmaller;\n    }\n    [MethodImpl(MethodImplOptions.NoInlining)]\n    private static bool IsDefaultValue(double value) => 0.0.Equals(value);\n    [MethodImpl(MethodImplOptions.NoInlining)]\n    private static bool HasEitherBit(int value) =>\n        ((value & 4) == 4) || ((value & 8) == 8);\n}\n```\nFor `ImplicitElse`, .NET 10 already avoids a branch, but it still materializes both Boolean values and selects between them. In .NET 11, the method becomes just the comparison and a `cset`, shrinking from 36 bytes to 24 bytes:\n\n```\n; Arm64\n--- .NET 10\n+++ .NET 11\n-            mov     w2, wzr\n-            mov     w3, #1\n             cmp     w0, w1\n-            csel    w2, w2, w3, ge\n-            mov     w0, w2\n+            cset    x0, lt\n-; Total bytes of code 36\n+; Total bytes of code 24\n```\n`0.0.Equals(value)` needs to account for `NaN`, but because the left operand is zero, the case where both operands are `NaN` can never apply. Removing the conditional selection for that case leaves one floating-point comparison and one `cset`, reducing `IsDefaultValue` from 40 bytes to 24 bytes:\n\n```\n; Arm64\n--- .NET 10\n+++ .NET 11\n             fcmp    d0, #0.0\n-            beq     G_M000_IG04\n-\n-G_M000_IG03:\n-            fcmp    d0, d0\n-            csel    w0, wzr, wzr, eq\n-            b       G_M000_IG05\n-\n-G_M000_IG04:\n-            mov     w0, #1\n-\n-G_M000_IG05:\n+            cset    x0, eq\n+\n+G_M000_IG03:\n             ldp     fp, lr, [sp], #0x10\n             ret     lr\n-; Total bytes of code 40\n+; Total bytes of code 24\n```\nFinally, normalizing both power-of-two comparisons lets the JIT combine their results. The short-circuit branch in `HasEitherBit` is replaced by two masks and an `or`, reducing the method from 40 bytes to 36 bytes:\n\n```\n; Arm64\n--- .NET 10\n+++ .NET 11\n-            tbz     w0, #2, G_M000_IG05\n-\n-G_M000_IG03:\n-            mov     w0, #1\n-\n-G_M000_IG04:\n-            ldp     fp, lr, [sp], #0x10\n-            ret     lr\n-\n-G_M000_IG05:\n-            tst     w0, #8\n+            and     w1, w0, #4\n+            and     w0, w0, #8\n+            orr     w0, w1, w0\n+            cmp     w0, #0\n             cset    x0, ne\n-G_M000_IG06:\n+G_M000_IG03:\n             ldp     fp, lr, [sp], #0x10\n             ret     lr\n-; Total bytes of code 40\n+; Total bytes of code 36\n```\nNot every simplification depends on broader control-flow reasoning.\n“Peephole optimizations” instead replace a short, recognizable pattern with an\nequivalent cheaper one. Each may save only an instruction or expose a form\nthat another optimization understands, but these patterns can occur very\nfrequently on hot paths throughout generated code. For example, dotnet/runtime#126529 from @BoyBaykiller recognizes that `255 - x` for a `byte` is equivalent to `x ^ 255`: both simply flip all eight bits, but the latter can remove an instruction if it’s able to replace a negation and add with an xor. Similarly, `-1 - x` can turn into the equivalent of `~x`.\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private readonly byte[] _data = new byte[4096];\n    [GlobalSetup]\n    public void Setup() => new Random(42).NextBytes(_data);\n    [Benchmark]\n    public int InvertBytes()\n    {\n        int sum = 0;\n        foreach (byte b in _data) sum += 255 - b;\n        return sum;\n    }\n}\n```\nIn .NET 11, the loop loses a separate negate and add:\n\n```\n; Arm64\n--- .NET 10\n+++ .NET 11\n@@ -19,9 +19,8 @@\n G_M000_IG04:\n             ldrb    w4, [x0, w2, UXTW]\n-            neg     w4, w4\n+            eor     w4, w4, #255\n             add     w1, w4, w1\n-            add     w1, w1, #255\n             add     w2, w2, #1\n             cmp     w3, w2\n             bgt     G_M000_IG04\n@@ -33,4 +32,4 @@\n             ldp     fp, lr, [sp], #0x10\n             ret     lr\n-; Total bytes of code 76\n+; Total bytes of code 72\n```\ndotnet/runtime#129361 removes another unnecessary instruction when comparing an `sbyte` with a constant that fits in eight bits. The JIT can compare the byte directly, with no sign extension:\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Runtime.CompilerServices;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private sbyte _value = -65;\n    [Benchmark]\n    public bool IsLow() => IsLow(_value);\n    [MethodImpl(MethodImplOptions.NoInlining)]\n    private static bool IsLow(sbyte value) => value < -64;\n}\n```\nThe optimized codegen then compares the byte directly, removing the `movsx` sign-extension instruction (though the JIT still retains it in the few comparison forms that require a full-width sign bit for correctness).\n\n```\n; x64\n--- .NET 10\n+++ .NET 11\n-movsx  rax,cl\n-cmp    eax,0FFFFFFC0\n+cmp    cl,0C0\n setl   al\n movzx  eax,al\n ret\n-; Total bytes of code 14\n+; Total bytes of code 10\n```\ndotnet/runtime#125180 from @saucecontrol improves non-overflowing `float` and `double` conversions to `long` and `ulong` on x86 machines with AVX-512 or AVX10.2. These casts have defined behavior for NaN and out-of-range values, so older code used a helper to preserve those semantics. The newer instruction set lets the JIT keep the normal path inline and register-based, avoiding the helper call; machines that don’t support these instructions retain the existing fallback.\n\n```\n// Run with 32-bit x86 dotnet on a machine with AVX-512 or AVX10.2:\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private float _single = 123_456.75f;\n    private double _double = 123_456.75;\n    [Benchmark] public long SingleToInt64() => (long)_single;\n    [Benchmark] public ulong SingleToUInt64() => (ulong)_single;\n    [Benchmark] public long DoubleToInt64() => (long)_double;\n    [Benchmark] public ulong DoubleToUInt64() => (ulong)_double;\n}\n```\n| Method | Runtime | Mean | Ratio | Code Size | \n|---|---|---|---|---|\n| SingleToInt64 | .NET 10.0 | 5.103 ns | 1.00 | 31 B | \n| SingleToInt64 | .NET 11.0 | 2.093 ns | 0.41 | 54 B | \n| SingleToUInt64 | .NET 10.0 | 4.810 ns | 1.00 | 31 B | \n| SingleToUInt64 | .NET 11.0 | 1.366 ns | 0.28 | 34 B | \n| DoubleToInt64 | .NET 10.0 | 4.834 ns | 1.00 | 31 B | \n| DoubleToInt64 | .NET 11.0 | 2.101 ns | 0.43 | 54 B | \n| DoubleToUInt64 | .NET 10.0 | 4.663 ns | 1.00 | 31 B | \n| DoubleToUInt64 | .NET 11.0 | 1.363 ns | 0.29 | 34 B | \n\n### Vectorization\n\nSIMD, or “single instruction, multiple data”, is the concept of one instruction applying the same operation to several values at once. A “scalar” `add`, for example, might combine one pair of 32-bit integers, while a 128-bit SIMD `add` can combine “vectors” of four pairs in the same instruction; 256- and 512-bit variants can handle vectors of eight and sixteen pairs, respectively. When the iterations of an operation are independent, “vectorizing” a loop can therefore replace several scalar iterations with one, improving the throughput of the loop significantly.\n\n.NET exposes portable (they work on any machine) variable-width vector type `Vector<T>` (which can represent different counts of `T` depending on the current hardware), fixed-width `Vector64<T>` through `Vector512<T>` types (which always represent the same count of `T`), and architecture-specific intrinsics (performing operations on such vector types which the JIT then maps to the right underlying hardware instructions). Each element in a vector is often referred to as a “lane”. Because the JIT recognizes these operations directly, it can fold constants, select instructions, and remove unsupported paths without treating them as normal method calls.\n\nA variety of PRs in .NET 11 improve AVX-512 broadcasting and masking. Embedded broadcasting lets an instruction load a single scalar value and replicate it across all vector lanes, avoiding the need to materialize a full-width vector constant in memory to feed into the instruction. For example, this bitwise AND instruction:\n\n```\n; x64\nvpandd  zmm0, zmm1, dword ptr [reloc @RWD00] {1to16}\n```\ncan replace this one:\n\n```\n; x64\nvpandd  zmm0, zmm1, zmmword ptr [reloc @RWD00]\n```\nstoring only 4 bytes in the read-only data section rather than 64. Because the broadcast is handled as part of the load, there’s no additional instruction-level latency; the primary benefit is reduced data size and cache footprint.\n\nEmbedded masking similarly lets an instruction update only a subset of the lanes. A mask is one bit per vector lane, where each bit indicates whether and how the operation should affect the corresponding lane. Without embedded masking, code often needs to compute every lane and then blend that result with the old value, so folding the mask into the operation can remove both the separate blend and a zero-vector setup. dotnet/runtime#117700 from @saucecontrol improves broadcast selection when an intrinsic’s natural element size differs from its managed vector type. VNNI, the Vector Neural Network Instructions used for small-integer multiply-accumulate operations, and bitwise operations can now use the smallest valid repeated constant, avoiding a full-vector load.\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\n// Requires AVX-VNNI and AVX-512F.\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Runtime.Intrinsics;\nusing System.Runtime.Intrinsics.X86;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private readonly Vector128<byte> _bytes = Vector128.Create((byte)1);\n    private readonly Vector128<ulong> _u64 = Vector128.Create(1UL);\n    private readonly Vector128<uint> _u32 = Vector128.Create(1U);\n    private readonly Vector512<int> _v512 = Vector512.Create(1);\n    private int _n = 42;\n    [Benchmark]\n    public Vector128<int> VnniBroadcast() =>\n        AvxVnni.MultiplyWideningAndAdd(\n            Vector128<int>.Zero, _bytes, Vector128<sbyte>.One);\n    [Benchmark]\n    public Vector128<uint> MaskAnd() =>\n        Vector128.ConditionalSelect(\n            Vector128.GreaterThan(_u32, Vector128<uint>.Zero),\n            (_u64 & Vector128<uint>.One.AsUInt64()).AsUInt32(),\n            Vector128<uint>.Zero);\n    [Benchmark]\n    public Vector512<int> BlendMaskAllOnes() =>\n        Avx512F.BlendVariable(\n            Vector512.Create(_n),\n            _v512,\n            Vector512.Create(-1));\n    [Benchmark]\n    public Vector512<int> MultiInsert() =>\n        Vector512.ConditionalSelect(\n            Vector512.Create(0, -1, 0, 0, 0, -1, 0, 0, 0, -1, 0, 0, 0, -1, 0, 0),\n            _v512,\n            Vector512.Create(_n));\n    [Benchmark]\n    public Vector512<int> MultiInsertZero() =>\n        Avx512F.BlendVariable(\n            _v512,\n            Vector512<int>.Zero,\n            Vector512.Create(0, -1, 0, 0, 0, -1, 0, 0, 0, -1, 0, 0, 0, -1, 0, 0));\n}\n```\nIn .NET 11, this results in 12 fewer bytes in the read-only data section, and 12 fewer bytes of constant-pool cache footprint.\n\n```\n; x64\n-C4E279503500000000   vpdpbusd xmm6, xmm0, xmmword ptr [reloc @RWD00]\n+62F27D18503500000000 vpdpbusd xmm6, xmm0, dword ptr [reloc @RWD00] {1to4}\n-RWD00  dq 0101010101010101h, 0101010101010101h\n+RWD00  dd 01010101h\n```\nThe fix also impacts embedded masking. For example, with `MaskAnd` previously, the AND used a qword broadcast, `{1to2}`, and a separate blend then moved the masked result, meaning two instructions. Now that `Vector128<uint>.One` can be broadcast at dword granularity, the mask’s element size and the AND’s element size agree, unlocking using the single merged-masked form. This pattern shows up throughout vectorized algorithms that do lots of bitwise manipulation and hashing, including implementations in System.Numerics.Tensors, System.IO.Hashing, and System.Private.CoreLib.\n\n```\n; x64\n-       vpandq   xmm0, xmm0, qword ptr [reloc @RWD00] {1to2}\n-       vpblendmd xmm0 {k1}{z}, xmm0, xmm0\n+       vpandd   xmm0 {k1}{z}, xmm0, dword ptr [reloc @RWD00] {1to4}\n; Code: 45 → 39 bytes; data: 8 bytes → 4 bytes\n```\nThat `MaskAnd` example starts as an AND followed by a blend, an operation that\nchooses independently for each vector lane whether to take its value from one\ninput or the other, with the JIT able to fold those two operations together.\nSimilar opportunities arise with blends more generally. Sometimes the mask or\none of the inputs makes the choice trivial, e.g. an all-ones mask always selects the\nsame input, so the blend is just a move. If one input is zero, it can often\nbecome an `AND` or `ANDN`. AVX-512 provides more options still, as constant\nmasks and zeroing can be encoded directly in the instruction.\ndotnet/runtime#123146 from\n@saucecontrol makes these simplifications\nconsistently across the portable and hardware-specific APIs. A blend with an all-ones mask provides a particularly clear example:\n\n```\n// Run on x64 with AVX-512:\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing System.Runtime.Intrinsics;\nusing System.Runtime.Intrinsics.X86;\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser]\n[HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private readonly Vector512<int> _values = Vector512.Create(1);\n    private int _n = 42;\n    [GlobalSetup]\n    public void Setup()\n    {\n        if (!Avx512F.IsSupported)\n            throw new PlatformNotSupportedException();\n    }\n    [Benchmark]\n    public Vector512<int> BlendMaskAllOnes() =>\n        Avx512F.BlendVariable(Vector512.Create(_n), _values, Vector512.Create(-1));\n}\n```\nThe generated code no longer needs to create the first input, load the mask, or perform the blend. It simply loads the input the all-ones mask would always select:\n\n```\n; x64\n-vpbroadcastd zmm0, dword ptr [rcx+8]\n-kmovq       k1, qword ptr [RWD00]\n-vpblendmd   zmm0 {k1}, zmm0, [rcx+48]\n+vmovups     zmm0, [rcx+48]\n vmovups     [rdx], zmm0\n mov         rax, rdx\n vzeroupper\n ret\n; 39 bytes → 23 bytes\n```\n### Intrinsics\n\nAn intrinsic is a managed API that the JIT recognizes and special-cases. Often that special-casing involves actually replacing calls to the method with custom code that’s behaviorally equivalent but better in some way (faster, smaller, etc.)\n\nAs an example, dotnet/runtime#128678 improves recognition of generic-math calls to `IBinaryNumber<T>.Log2`. The method computes the base-2 logarithm of an integer, equivalent to the index of the number’s highest set bit; for example, `Log2(16)` is `4`. Previously, the JIT’s normalized integer type lost the signedness needed to import the operation directly as an intrinsic. Inlining the managed implementation could still produce the same optimized code, but when inlining didn’t happen, the managed call remained. In .NET 11, the JIT consults the precise type and imports the operation directly: unsigned and non-negative signed inputs can become leading-zero-count or bit-scan arithmetic, while a negative signed value retains the managed fallback and its exact exception behavior.\n\nSometimes the JIT has a perfectly good intrinsic lowering but doesn’t recognize a call that should use it. `Enum.Equals` from a generic `T : Enum` context was a good example. Even though both arguments to the generic helper are strongly typed as `T`, an enum doesn’t provide an `Equals(T)` method; it inherits the virtual `Enum.Equals(object)` implementation. The second argument therefore needs to be boxed to pass it as `object`. The receiver is invoked with a constrained virtual call, but because the concrete enum doesn’t override the method itself, it too needs to be boxed to invoke the implementation on `System.Enum`. Thus, what looks like a strongly-typed comparison can end up allocating two boxes and making a virtual call. In .NET 11, dotnet/runtime#122779 eliminates this overhead by teaching the JIT to recognize the call and fold it to a direct comparison of the enum’s underlying integer values. For example:\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Runtime.CompilerServices;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[MemoryDiagnoser(false), HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private static readonly StringComparison[] s_values =\n    {\n        StringComparison.Ordinal, StringComparison.OrdinalIgnoreCase, \n        StringComparison.CurrentCulture, StringComparison.CurrentCultureIgnoreCase,\n        StringComparison.InvariantCulture, StringComparison.Ordinal,\n    };\n    [Benchmark]\n    public int CountOrdinal_Generic()\n    {\n        int count = 0;\n        foreach (var v in s_values)\n            if (EqualsGeneric(v, StringComparison.Ordinal))\n                count++;\n        return count;\n    }\n    [MethodImpl(MethodImplOptions.NoInlining)]\n    private static bool EqualsGeneric<T>(T a, T b) where T : Enum => a.Equals(b);\n}\n```\nOnce the JIT knows the callee is `Enum.Equals` and knows the exact enum type, it asks the runtime for the underlying integer type and replaces the virtual call with a direct comparison. That in turn makes both box/unbox pairs redundant, and the generated code contains neither allocation. For the six comparisons performed here, .NET 10 creates twelve boxes, totaling 288 bytes. In .NET 11, the helper becomes just the integer comparison, eliminating both the allocations and the virtual dispatch.\n\n| Method | Runtime | Mean | Ratio | Allocated | \n|---|---|---|---|---|\n| CountOrdinal_Generic | .NET 10.0 | 59.24 ns | 1.00 | 288 B | \n| CountOrdinal_Generic | .NET 11.0 | 10.01 ns | 0.17 | – | \n\nNativeAOT had been carrying an equivalent optimization for years, implemented as IL rewriting in ILCompiler that patches `Enum.Equals` to use typed comparisons. With the JIT now handling it, including in NativeAOT’s own use of the JIT (NativeAOT uses the JIT ahead of time rather than just in time), dotnet/runtime#123086 deletes that rewriting and its supporting machinery.\n\ndotnet/runtime#127329 improves the `Vector256.Sum` and `Vector512.Sum` intrinsics. The JIT now performs most of the reduction at full width and combines the per-lane results at the end, avoiding the extracts and duplicate shuffle sequences needed when splitting wide vectors into 128-bit pieces. And dotnet/runtime#127402 extends vector-constant propagation from 128-bit vectors to `Vector256` and `Vector512`. Code that compares a wide vector with a known sentinel can now simplify subsequent uses just as narrower vectors already could. The following benchmark exemplifies both:\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\n// Requires AVX2 for the assembly shown below.\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Runtime.CompilerServices;\nusing System.Runtime.Intrinsics;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private readonly Vector256<float> _floats =\n        Vector256.Create(1.0f, 2.0f, 3.0f, 4.0f, 5.0f, 6.0f, 7.0f, 8.0f);\n    private int _selector;\n    [Benchmark]\n    public float Sum() => Vector256.Sum(_floats);\n    [Benchmark]\n    public int TransformWhenKnown()\n    {\n        Vector256<int> value = GetVector();\n        if (value == Vector256.Create(0, 1, 2, 3, 4, 5, 6, 7))\n            return (value + Vector256.Create(10)).GetElement(6);\n        return -1;\n    }\n    [MethodImpl(MethodImplOptions.NoInlining)]\n    private Vector256<int> GetVector() =>\n        _selector == 0 ?\n            Vector256.Create(0, 1, 2, 3, 4, 5, 6, 7) :\n            Vector256.Create(7);\n}\n```\nFor `Sum`, .NET 10 separately reduces each 128-bit half and then adds the two scalar results. In .NET 11, the permutes and adds operate on both halves in parallel as 256-bit instructions, after which only the two already-reduced halves need to be combined:\n\n```\n; x64\n vmovups   ymm0, [rcx+28]\n-vmovaps   ymm1, ymm0\n-vpermilps xmm2, xmm1, 0B1\n-vaddps    xmm1, xmm2, xmm1\n-vpermilps xmm2, xmm1, 4E\n-vaddps    xmm1, xmm2, xmm1\n-vextractf128 xmm0, ymm0, 1\n-vpermilps xmm2, xmm0, 0B1\n-vaddps    xmm0, xmm2, xmm0\n-vpermilps xmm2, xmm0, 4E\n-vaddps    xmm0, xmm2, xmm0\n-vaddss    xmm0, xmm1, xmm0\n+vpermilps ymm1, ymm0, 0B1\n+vaddps    ymm0, ymm1, ymm0\n+vpermilps ymm1, ymm0, 4E\n+vaddps    ymm0, ymm1, ymm0\n+vextractf128 xmm1, ymm0, 1\n+vaddps    xmm0, xmm1, xmm0\n; 63 bytes → 39 bytes\n```\n`TransformWhenKnown` uses a deliberately non-repeating constant across its eight lanes. On the branch where the comparison succeeds, .NET 11 can replace `value` with that constant, fold the vector addition, and determine that element 6 is `16`. The `vpaddd`, extraction, second 32-byte constant, and associated control flow all disappear:\n\n```\n; x64\n-cmp      eax, 0FFFFFFFF\n-jne      M00_L00\n-vmovups  ymm0, [rsp+20]\n-vpaddd   ymm0, ymm0, [RWD32]\n-vextracti128 xmm0, ymm0, 1\n-vpextrd  eax, xmm0, 2\n-vzeroupper\n-add      rsp, 58\n-ret\n-\n-M00_L00:\n-mov      eax, 0FFFFFFFF\n+mov      ecx, 0FFFFFFFF\n+mov      edx, 10\n+cmp      eax, 0FFFFFFFF\n+mov      eax, edx\n+cmovne   eax, ecx\n vzeroupper\n add      rsp, 58\n ret\n; 85 bytes → 59 bytes\n```\nOne of the goals of .NET is that you can write code once and have it run anywhere, optimized for whatever that “anywhere” has to offer. For vectorization, that means providing portable operations whenever the intent is common across instruction sets, while retaining architecture-specific APIs for algorithms that really do need to target a particular machine.\n\nWhenever possible, we want to enable developers to express their algorithms using the portable APIs, and each release of .NET fills additional gaps there. Including .NET 11. dotnet/runtime#129627 from @hez2010 adds portable APIs for constructing common lane sequences (e.g. `[1, 2, 4, 8]` or `[a, b, a, b]`), concatenating half-vectors (the lower halves of `[a, b, c, d]` and `[w, x, y, z]` producing `[a, b, w, x]`), interleaving (`[a, b]` and `[x, y]` producing `[a, x, b, y]`), de-interleaving (`[a, x, b, y]` producing `[a, b]` and `[x, y]`), and reversal (`[a, b, c, d]` producing `[d, c, b, a]`), along with their JIT intrinsification. These operations were already expressible, but only verbosely and only if you knew which hardware instruction to reach for, e.g. writing `Zip` by hand meant targeting a platform-specific API like `AdvSimd.Arm64.ZipLow`. The new APIs let the code state the transformation and leave instruction selection to the JIT.\n\nOnce the intrinsic operation has been recognized, the backend still needs to keep it in a useful vector form while assigning registers and selecting instructions. Vector values are structs, and the JIT will often apply “struct promotion,” tracking a struct’s fields as independent locals so that each can be optimized separately. That’s useful for ordinary structs, but counterproductive when a value is meant to remain in a vector or mask register: splitting it can introduce extra moves and obscure what should be a single whole-value store, particularly after inlining introduces more local stores. dotnet/runtime#128013 consistently marks SIMD and mask stores as intrinsic-related across platforms, including 32-bit x86 and x64 mask stores, so those locals remain intact. dotnet/runtime#129563 extends that principle to user-defined structs that are bitcast to SIMD types. This trades away struct promotion for those locals, but enables the JIT to preserve their vector representation.\n\nThis matters for user-defined numerical types that store the same data as a\nhardware vector but expose named fields or domain-specific operations. The\nfollowing `Vector2Double` is laid out as two adjacent `double` values, so it can\nbe bitcast to `Vector128<double>`, operated on with SIMD, and bitcast back:\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing System.Runtime.CompilerServices;\nusing System.Runtime.InteropServices;\nusing System.Runtime.Intrinsics;\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\npublic struct Vector2Double(double x, double y)\n{\n    public double X = x;\n    public double Y = y;\n    public static Vector2Double operator +(Vector2Double left, Vector2Double right)\n    {\n        Vector128<double> simdLeft = Unsafe.BitCast<Vector2Double, Vector128<double>>(left);\n        Vector128<double> simdRight = Unsafe.BitCast<Vector2Double, Vector128<double>>(right);\n        return Unsafe.BitCast<Vector128<double>, Vector2Double>(simdLeft + simdRight);\n    }\n}\n[DisassemblyDiagnoser, HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private readonly Vector2Double _a = new(1.0, 2.0);\n    private readonly Vector2Double _b = new(3.0, 4.0);\n    private readonly Vector2Double _c = new(5.0, 6.0);\n    [Benchmark]\n    public Vector2Double Add() => Add(_a, _b, _c);\n    [MethodImpl(MethodImplOptions.NoInlining)]\n    private static Vector2Double Add(Vector2Double a, Vector2Double b, Vector2Double c) =>\n        a + b + c;\n}\n```\nIn .NET 10, promotion of the intermediate struct sends the first SIMD result\nthrough two stack locations before the second addition. .NET 11 keeps that\nvalue in `xmm0`, reducing the helper from 52 bytes to 22 bytes:\n\n```\n; x64\n-sub       rsp, 28\n vmovups   xmm0, [rdx]\n vaddpd    xmm0, xmm0, [r8]\n-vmovaps   [rsp], xmm0\n-vmovups   xmm0, [rsp]\n-vmovups   [rsp+18], xmm0\n-vmovups   xmm0, [rsp+18]\n vaddpd    xmm0, xmm0, [r9]\n vmovups   [rcx], xmm0\n mov       rax, rcx\n-add       rsp, 28\n ret\n; 52 bytes → 22 bytes\n```\ndotnet/runtime#128350 gives the xarch register allocator more freedom around fused multiply-add (FMA) and AVX-512 ternary-logic operations. These instructions can read and overwrite operands in several equivalent arrangements; choosing the arrangement that already matches the surrounding registers avoids otherwise necessary moves.\n\nGeneric vector code introduces another wrinkle. Operators like\n`Vector128<T>.operator ==` return `bool`, so the return type doesn’t reveal the\nvector’s element type. The JIT instead needs to obtain that type from the\noperands in order to select the right comparison instruction. In some generic\ncontexts, including helpers built on the internal `ISimdVector` abstraction,\nthe JIT was consulting the wrong type information and failed to import the\noperator as an intrinsic. It then executed the managed fallback, which compares\nthe lanes individually. dotnet/runtime#130086 marks these operators so their element type is taken from the first argument. As an example, the generic helpers used internally by ordinal-ignore-case string comparer benefit from this.\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private readonly string _lower = new('a', 256);\n    private readonly string _upper = new('A', 256);\n    [Benchmark]\n    public bool OrdinalIgnoreCase() => string.Equals(_lower, _upper, StringComparison.OrdinalIgnoreCase);\n}\n```\n| Method | Runtime | Mean | Ratio | \n|---|---|---|---|\n| OrdinalIgnoreCase | .NET 10.0 | 27.794 ns | 1.00 | \n| OrdinalIgnoreCase | .NET 11.0 | 21.861 ns | 0.79 | \n\n.NET 11 adds support for newer x86 capabilities while also\nimproving code generated for existing hardware. These changes benefit both\ndirect users of hardware intrinsics and portable vector code selected by the\nJIT. For example, dotnet/runtime#124114 from @saucecontrol improves 32-bit x86 without AVX-512, where converting `uint` to `float` or `double` previously required a runtime helper. Older x86 conversion instructions accept signed integers, and half of the `uint` range doesn’t fit in a signed 32-bit value, which is why the helper existed. The JIT now emits an inline vector-instruction sequence that handles the high bit explicitly, avoiding the call and its register and stack overhead.\n\n```\n// Run with 32-bit x86 dotnet and AVX-512 disabled (DOTNET_EnableAVX512=0)\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private uint _value = 0xF123_4567;\n    [Benchmark] public float UInt32ToSingle() => _value;\n    [Benchmark] public double UInt32ToDouble() => _value;\n}\n```\n| Method | Runtime | Mean | Ratio | Code Size | \n|---|---|---|---|---|\n| UInt32ToSingle | .NET 10.0 | 4.752 ns | 1.00 | 37 B | \n| UInt32ToSingle | .NET 11.0 | 2.403 ns | 0.51 | 43 B | \n| UInt32ToDouble | .NET 10.0 | 4.727 ns | 1.00 | 39 B | \n| UInt32ToDouble | .NET 11.0 | 2.402 ns | 0.51 | 41 B | \n\ndotnet/runtime#124804 from @alexcovington adds the AVX-512 Bit Matrix Multiply APIs. A binary matrix treats each bit as an element and combines rows and columns with bitwise operations, not integer multiplication. The instructions are useful in areas such as error correction and CRC computation. Each replaces a much longer sequence of shifts, masks, and exclusive-ORs. And dotnet/runtime#128365 from @jamesburton adds `AvxVnni.V512`, extending the AVX-VNNI APIs from 256-bit to 512-bit operands so the small-integer dot products used by quantized machine-learning models can process 64 bytes per operation instead of 32.\n\ndotnet/runtime#126062 from @saucecontrol also avoids converting a vector selector into an AVX-512 mask register when the eventual operation still needs the vector form. In such cases, the older-looking vector blend is actually shorter and uses fewer resources:\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Runtime.Intrinsics;\nusing System.Runtime.Intrinsics.X86;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private readonly Vector128<float> _v1 = Vector128.Create(-1.0f, 2.0f, -3.0f, 4.0f);\n    private readonly Vector128<float> _v2 = Vector128.Create(10.0f);\n    [GlobalSetup]\n    public void Setup()\n    {\n        if (!Sse41.IsSupported)\n            throw new PlatformNotSupportedException();\n    }\n    [Benchmark]\n    public Vector128<float> AddToNegative() =>\n        Sse41.BlendVariable(_v1, _v1 + _v2, _v1);\n}\n```\nIn .NET 11, you get the simpler `vblendvps` form that avoids an unnecessary k-register operation.\n\n```\n; x64\n  vmovups   xmm0, [rcx+8]\n- vpmovd2m  k1, xmm0\n- vaddps    xmm0 {k1}, xmm0, [rcx+18]\n+ vaddps    xmm1, xmm0, [rcx+18]\n+ vblendvps xmm0, xmm0, xmm1, xmm0\n  vmovups   [rdx], xmm0\n; 29 bytes → 24 bytes\n```\nThe masked EVEX form looks more modern, but when the mask originates from a vector anyway, the vector-blend sequence is five bytes shorter and avoids writing a mask register. There are only 8 k-registers, and some microarchitectures have port contention for instructions that write them.\n\nA compiler’s cost model assigns estimates to operations and instructions, such as their execution cost or throughput and their impact on code size, and uses those estimates to choose between otherwise legal transformations or instruction sequences. Wrong estimates can still produce semantically correct code, just slower or larger code. With dotnet/runtime#127048, which updates the JIT’s xarch floating-point and SIMD cost model, the JIT’s cost model reflects modern instruction throughput and encoded size, replacing old x87 assumptions and a flat cost for every intrinsic. That leads to better decisions about common-subexpression elimination and loop unrolling, particularly for 512-bit operations.\n\ndotnet/runtime#130422 folds a vector lane extraction followed by `WithElement` into one `insertps` that reads the source lane directly. Code such as `destination.WithElement(0, source.GetElement(2))` conceptually extracts a scalar and then inserts it elsewhere. `insertps`, however, has an immediate operand whose bits select both the source lane and destination lane. The JIT can therefore pass the original source vector to the instruction and encode lane 2 in that immediate, instead of first shuffling lane 2 into the scalar position and then inserting it.\n\nThree more xarch changes tighten public SIMD operations on the hardware where they apply. dotnet/runtime#125666 from @alexcovington replaces the dedicated AVX dot-product instruction with a multiply, add, and permute reduction that has better throughput on contemporary cores:\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing System.Numerics;\nusing System.Runtime.Intrinsics;\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private readonly Plane _plane = new(new Vector3(1.0f, 2.0f, 3.0f), 4.0f);\n    private readonly Vector4 _vector4 = new(5.0f, 6.0f, 7.0f, 8.0f);\n    private readonly Quaternion _quaternion1 = new(1.0f, 2.0f, 3.0f, 4.0f);\n    private readonly Quaternion _quaternion2 = new(5.0f, 6.0f, 7.0f, 8.0f);\n    private readonly Vector128<float> _vector1 = Vector128.Create(1.0f, 2.0f, 3.0f, 4.0f);\n    private readonly Vector128<float> _vector2 = Vector128.Create(5.0f, 6.0f, 7.0f, 8.0f);\n    [Benchmark]\n    public float PlaneDot() => Plane.Dot(_plane, _vector4);\n    [Benchmark]\n    public float QuaternionDot() => Quaternion.Dot(_quaternion1, _quaternion2);\n    [Benchmark]\n    public float Vector128Dot() => Vector128.Dot(_vector1, _vector2);\n}\n```\n| Method | Runtime | Mean | Ratio | Code Size | \n|---|---|---|---|---|\n| PlaneDot | .NET 10.0 | 2.616 ns | 1.00 | 13 B | \n| PlaneDot | .NET 11.0 | 1.365 ns | 0.52 | 31 B | \n| QuaternionDot | .NET 10.0 | 2.640 ns | 1.00 | 13 B | \n| QuaternionDot | .NET 11.0 | 1.326 ns | 0.50 | 31 B | \n| Vector128Dot | .NET 10.0 | 2.597 ns | 1.00 | 13 B | \n| Vector128Dot | .NET 11.0 | 1.366 ns | 0.53 | 31 B | \n\nMultiplying vectors of bytes is more involved than multiplying vectors of larger integer types because x86 doesn’t provide a packed byte-multiply instruction. The implementation needs to combine wider 16-bit multiplications while retaining only the low byte of each product. When it couldn’t widen the whole operation to the next vector size, .NET 10 split the input into two halves, widened and multiplied each half, narrowed both results, and joined them again. dotnet/runtime#126348 from @saucecontrol instead separates the even and odd bytes with masks and shifts, performs two 16-bit multiplications over the full vector width, and recombines the low bytes:\n\n```\n// Run on x64 with AVX-512:\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing System.Runtime.Intrinsics;\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private readonly Vector512<byte> _left = Vector512.Create((byte)17);\n    private readonly Vector512<byte> _right = Vector512.Create((byte)19);\n    [Benchmark]\n    public Vector512<byte> Multiply() => _left * _right;\n}\n```\n| Method | Runtime | Mean | Ratio | Code Size | \n|---|---|---|---|---|\n| Multiply | .NET 10.0 | 3.752 ns | 1.00 | 114 B | \n| Multiply | .NET 11.0 | 2.174 ns | 0.58 | 73 B | \n\nThe .NET 11 sequence no longer extracts, widens, narrows, and reinserts both 256-bit halves:\n\n```\n; x64\n vmovups     zmm0, [rcx+8]\n-vmovaps     zmm1, zmm0\n-vpmovzxbw   zmm1, ymm1\n-vmovups     zmm2, [rcx+48]\n-vmovaps     zmm3, zmm2\n-vpmovzxbw   zmm3, ymm3\n-vpmullw     zmm1, zmm3, zmm1\n-vpmovwb     ymm1, zmm1\n-vextracti32x8 ymm0, zmm0, 1\n-vpmovzxbw   zmm0, ymm0\n-vextracti32x8 ymm2, zmm2, 1\n-vpmovzxbw   zmm2, ymm2\n-vpmullw     zmm0, zmm2, zmm0\n-vpmovwb     ymm0, zmm0\n-vinserti32x8 zmm0, zmm1, ymm0, 1\n+vmovups     zmm1, [rcx+48]\n+vpmullw     zmm2, zmm0, zmm1\n+vpsrlw      zmm0, zmm0, 8\n+vpandd      zmm1, zmm1, dword bcst [RWD00]\n+vpmullw     zmm0, zmm1, zmm0\n+vpternlogd  zmm0, zmm2, dword bcst [RWD04], 0F8\n vmovups     [rdx], zmm0\n; 114 bytes → 73 bytes\n```\ndotnet/runtime#127094 lets scalar conversions between `Half` and `float` use F16C’s `vcvtps2ph` and `vcvtph2ps` instructions when AVX2 is enabled:\n\n```\n// Run on x64 with AVX2 enabled:\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private Half _half = (Half)123.5f;\n    private float _single = 123.5f;\n    [Benchmark] public float HalfToSingle() => (float)_half;\n    [Benchmark] public Half SingleToHalf() => (Half)_single;\n}\n```\n| Method | Runtime | Mean | Ratio | Code Size | \n|---|---|---|---|---|\n| HalfToSingle | .NET 10.0 | 2.506 ns | 1.00 | 104 B | \n| HalfToSingle | .NET 11.0 | 1.380 ns | 0.55 | 14 B | \n| SingleToHalf | .NET 10.0 | 2.598 ns | 1.00 | 134 B | \n| SingleToHalf | .NET 11.0 | 1.351 ns | 0.52 | 19 B | \n\nFinally, dotnet/runtime#127536 from @Ruihan-Yin completes support for APX, Intel’s Advanced Performance Extensions. In addition to expanding the general-purpose register set, APX adds forms of many instructions that don’t overwrite the processor’s condition flags. That gives the register allocator and instruction scheduler more freedom to keep values and pending conditions alive at the same time. Its `CTEST` and `CFCMOV` instructions can also represent chained conditions without branches and replace some compare-with-zero forms with shorter encodings. Applications don’t need to call APX-specific APIs to benefit; when the hardware and operating system expose APX, the JIT is able to utilize the additional instructions automatically.\n\nOn Arm64, the work in .NET 11 spans both conventional code generation and continued support for SVE (Scalable Vector Extension). Unlike 128-bit AdvSimd vectors, an SVE vector doesn’t have one width fixed by the instruction set; each processor chooses a supported width, and the same compiled loop uses predicate masks to operate on however many elements fit. That makes SVE well suited to loops whose trip counts are not exact multiples of a particular vector size.\n\ndotnet/runtime#121986 improves zeroing for larger stack allocations on Arm64. The JIT can store two zeroed 128-bit vector registers at a time, doubling the amount cleared by each instruction:\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Runtime.CompilerServices;\nusing System.Runtime.Intrinsics.Arm;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    [Benchmark] public void Stackalloc512() => Consume(stackalloc byte[512]);\n    [Benchmark] public void Stackalloc1024() => Consume(stackalloc byte[1024]);\n    [Benchmark] public void Stackalloc16384() => Consume(stackalloc byte[16384]);\n    [MethodImpl(MethodImplOptions.NoInlining)]\n    private static void Consume(Span<byte> x) { }\n}\n```\n| Method | Runtime | Mean | Ratio | \n|---|---|---|---|\n| Stackalloc512 | .NET 10.0 | 13.65 ns | 1.00 | \n| Stackalloc512 | .NET 11.0 | 9.557 ns | 0.70 | \n| Stackalloc1024 | .NET 10.0 | 25.35 ns | 1.00 | \n| Stackalloc1024 | .NET 11.0 | 14.332 ns | 0.57 | \n| Stackalloc16384 | .NET 10.0 | 312.97 ns | 1.00 | \n| Stackalloc16384 | .NET 11.0 | 162.656 ns | 0.52 | \n\nA wave of smaller Arm64 changes improves instruction selection. In .NET 11, dotnet/runtime#119758 from @jonathandavies-arm lets a comparison with zero consume condition flags set as a side effect of the preceding arithmetic or logical instruction, avoiding a separate `cmp`. dotnet/runtime#123138 from @jonathandavies-arm recognizes bit-extraction idioms such as `(value >> 6) & 0x3F` and maps them to the dedicated `ubfx` instruction. And dotnet/runtime#123546 from @jonathandavies-arm removes a non-overflowing `int`-to-`long` widening cast when the result is immediately truncated to a smaller integer type.\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\", \"left\", \"right\", \"value\")]\npublic class Benchmarks\n{\n    [Benchmark]\n    [Arguments(-1, 2)]\n    public bool CompareWithZero(int left, int right) => (left & right) <= 0;\n    [Benchmark]\n    [Arguments(0x7F65_4321)]\n    public int ExtractBits(int value) => (value >> 6) & 0x3F;\n    [Benchmark]\n    [Arguments(0x1122_3344)]\n    public sbyte TruncateAfterWidening(int value) => (sbyte)(long)value;\n}\n```\nEach example removes one instruction. `CompareWithZero` changes `and` to its flag-setting `ands` form and drops the subsequent `cmp`; `ExtractBits` replaces a shift and mask with `ubfx`; and `TruncateAfterWidening` drops the `sxtw` that widened the value to 64 bits only for `sxtb` to immediately truncate it again:\n\n```\n; Arm64\n; CompareWithZero: 28 bytes → 24 bytes\n-            and     w0, w1, w2\n-            cmp     w0, #0\n+            ands    w0, w1, w2\n             cset    x0, le\n; ExtractBits: 24 bytes → 20 bytes\n-            asr     w0, w1, #6\n-            and     w0, w0, #63\n+            ubfx    w0, w1, #6, #6\n; TruncateAfterWidening: 24 bytes → 20 bytes\n-            sxtw    x0, w1\n-            sxtb    w0, w0\n+            sxtb    w0, w1\n```\nInstruction selection also improves where values move between registers and memory. In .NET 11, dotnet/runtime#126803 changes `ToScalar` on a vector of 64-bit integers to use `fmov Xd, Dn` rather than the lane-extract instruction `umov`; in both cases lane zero moves to a general-purpose register, but `fmov` is the more direct form. For ReadyToRun code, dotnet/runtime#129589 folds relocatable indirection-cell loads from `adrp + add + ldr` into `adrp + ldr #:lo12:`, removing the separate address addition. And dotnet/runtime#129932 re-enables `ldp`/`stp` formation for negative unscaled offsets, letting two adjacent loads or stores become one paired instruction.\n\nThe first and third changes are easy to see with small methods:\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Runtime.CompilerServices;\nusing System.Runtime.Intrinsics;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private Vector128<long> _vector = Vector128.Create(42L, 84L);\n    private nint[] _storage = new nint[8];\n    [Benchmark]\n    public long ToScalar() => ToScalarCore(_vector);\n    [Benchmark]\n    public void ClearPrevious() => ClearPreviousCore(ref _storage[4]);\n    [MethodImpl(MethodImplOptions.NoInlining)]\n    private static long ToScalarCore(Vector128<long> value) => value.ToScalar();\n    [MethodImpl(MethodImplOptions.NoInlining)]\n    private static void ClearPreviousCore(ref nint value)\n    {\n        Unsafe.Add(ref value, -1) = 0;\n        Unsafe.Add(ref value, -2) = 0;\n        Unsafe.Add(ref value, -3) = 0;\n        Unsafe.Add(ref value, -4) = 0;\n    }\n}\n```\nThe `ToScalar` change is a direct instruction substitution, while the negative-offset stores collapse from four instructions to two, reducing the helper from 32 bytes to 24 bytes:\n\n```\n; Arm64\n; ToScalarCore\n-            umov    x0, v0.d[0]\n+            fmov    x0, d0\n; ClearPreviousCore\n-            str     xzr, [x0, #-0x08]\n-            str     xzr, [x0, #-0x10]\n-            str     xzr, [x0, #-0x18]\n-            str     xzr, [x0, #-0x20]\n+            stp     xzr, xzr, [x0, #-0x10]\n+            stp     xzr, xzr, [x0, #-0x20]\n```\nBit-counting operations benefit as well. `PopCount` counts the one bits in a value, while `TrailingZeroCount` counts the zero bits below its least-significant one bit. dotnet/runtime#128677 imports both as dedicated Arm64 intrinsics, making their intent visible to later optimization. On processors with the FEAT_CSSC extension, dotnet/runtime#130332 can then lower them directly to the scalar `cnt` and `ctz` instructions.\n\nComparison masks are another place where spelling out the intent enables much better code. Portable SIMD code often compares vectors, calls `ExtractMostSignificantBits`, and then asks whether any lane matched, counts matching lanes, or finds the first or last match. dotnet/runtime#129688 from @jonathandavies-arm recognizes those consumers on Arm64 and avoids materializing the full scalar mask: it can horizontally reduce the vector mask directly.\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Numerics;\nusing System.Runtime.CompilerServices;\nusing System.Runtime.Intrinsics;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private Vector128<int> _value = Vector128.Create(1, -2, 3, -4);\n    [Benchmark]\n    public bool AnyLessThan() => AnyLessThanCore(_value, 0);\n    [Benchmark]\n    public int CountLessThan() => CountLessThanCore(_value, 0);\n    [MethodImpl(MethodImplOptions.NoInlining)]\n    private static bool AnyLessThanCore(Vector128<int> value, int limit) =>\n        Vector128.LessThan(value, Vector128.Create(limit))\n            .ExtractMostSignificantBits() != 0;\n    [MethodImpl(MethodImplOptions.NoInlining)]\n    private static int CountLessThanCore(Vector128<int> value, int limit) =>\n        BitOperations.PopCount(\n            Vector128.LessThan(value, Vector128.Create(limit))\n                .ExtractMostSignificantBits());\n}\n```\nIn .NET 10, both helpers first pack the most-significant bit from every comparison lane into a scalar. .NET 11 instead keeps the comparison as a vector.\n\n```\n; Arm64\n; AnyLessThanCore\n             cmgt    v16.4s, v16.4s, v0.4s\n-            movi    v17.4s, #0x80, LSL #24\n-            and     v16.4s, v16.4s, v17.4s\n-            ldr     q17, [@RWD00]\n-            ushl    v16.4s, v16.4s, v17.4s\n-            addv    s16, v16.4s\n-            smov    x0, v16.s[0]\n+            umaxv   s16, v16.4s\n+            umov    w0, v16.s[0]\n             cmp     w0, #0\n             cset    x0, ne\n; CountLessThanCore\n             cmgt    v16.4s, v16.4s, v0.4s\n-            movi    v17.4s, #0x80, LSL #24\n-            and     v16.4s, v16.4s, v17.4s\n-            ldr     q17, [@RWD00]\n-            ushl    v16.4s, v16.4s, v17.4s\n-            addv    s16, v16.4s\n-            movi    v17.2s, #0\n-            smov    x0, v16.s[0]\n-            ins     v17.s[0], w0\n-            cnt     v16.8b, v17.8b\n-            addv    b16, v16.8b\n-            umov    w0, v16.b[0]\n+            ushr    v16.4s, v16.4s, #31\n+            addv    s16, v16.4s\n+            umov    w0, v16.s[0]\n```\nOn the SVE and SVE2 side, dotnet/runtime#129852 from @snickolls-arm removes the old 128-bit size ceiling for `Vector<T>` on Arm64 and lets the runtime size the type from the process’s actual SVE vector length. (Scalable `Vector<T>` remains experimental and disabled by default in .NET 11, so this expands what the experimental mode can do; it doesn’t speed up the default `Vector<T>` configuration.)\n\nThe public intrinsic surface also grows. In .NET 11, dotnet/runtime#118957 from @SwapnilGaikwad exposes odd-lane floating-point conversions; “odd lane” here means converting elements 1, 3, 5, and so on, which is useful when widening or narrowing interleaved data. dotnet/runtime#123890 from @ylpoonlg and dotnet/runtime#123892 from @ylpoonlg add non-temporal gather loads and scatter stores, which read from or write to multiple non-contiguous addresses (the “gather” part) while hinting that the data need not remain in cache (the “non-temporal” part).\n\nOther changes improve the predicates that make scalable loops work. dotnet/runtime#127538 adds hardware-generated predicate masks for more loop and memory-access patterns, while dotnet/runtime#126398 from @ylpoonlg reduces setup moves for masked operations. And dotnet/runtime#128326 from @snickolls-arm improves how SVE masks flow through the JIT, allowing zeroing forms of instructions to replace separate constant setup. dotnet/runtime#127520 from @a74nh enables scalable vector and mask constants, and dotnet/runtime#128148 from @snickolls-arm uses vector stores to initialize scalable vector locals, replacing scalar loops.\n\n### Register Allocation\n\nGenerated code constantly moves values between the CPU’s limited set of fast registers and temporary stack slots. Register allocation in a compiler decides which values stay in registers and which are “spilled” to the stack; avoiding one spill can remove both the store and the later reload.\n\nSome small structs are passed with multiple fields packed into one register. In .NET 11, dotnet/runtime#112740 lets the JIT extract those fields directly, avoiding a “spill” to a temporary stack slot followed by a reload of each field:\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Drawing;\nusing System.Runtime.CompilerServices;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private readonly Memory<int>[] _memories = CreateMemories();\n    private static Memory<int>[] CreateMemories()\n    {\n        Random rng = new(42);\n        var memories = new Memory<int>[4096];\n        for (int i = 0; i < memories.Length; i++)\n            memories[i] = new int[rng.Next(0, 20)];\n        return memories;\n    }\n    [MethodImpl(MethodImplOptions.NoInlining)]\n    private static bool Test(Memory<int> mem) => mem.Length > 10;\n    [Benchmark]\n    public int MemoryLengthExtract_Loop()\n    {\n        int count = 0;\n        for (int i = 0; i < _memories.Length; i++)\n            if (Test(_memories[i]))\n                count++;\n        return count;\n    }\n}\n```\nThe measured row uses `Memory<int>` because its length arrives packed into part of an argument register on Arm64. The new extraction avoids a stack round-trip on every call.\n\n| Method | Runtime | Mean | Ratio | \n|---|---|---|---|\n| MemoryLengthExtract_Loop | .NET 10.0 | 27.15 μs | 1.00 | \n| MemoryLengthExtract_Loop | .NET 11.0 | 23.95 μs | 0.88 | \n\nTwo broader register-allocation changes reduce unnecessary copies and spills: dotnet/runtime#125214 handles more conflicts directly, while dotnet/runtime#125219 steers short-lived values away from registers an upcoming operation will overwrite. dotnet/runtime#126552 from @SingleAccretion removes an old restriction on method prologs, eliminating jumps that existed only to satisfy that encoding rule.\n\n### Write Barriers and Garbage Collection\n\nThe .NET garbage collector is generational: new objects start in gen0, while objects that survive collections are promoted to gen1 and gen2. That enables the GC to collect younger generations without having to scan the whole heap. Of course, a reference to a younger object could get written to a field of an older one, in which case only scanning the younger generation would lead to problems. To ensure such references aren’t missed, whenever a write could create one, the JIT emits a small piece of code to update the GC’s bookkeeping; that code is known as a GC write barrier. Reference writes happen a lot, so it’s really important for performance that those barriers be as cheap as possible, and elided if they’re provably not needed at all.\n\nManaged reference stores may require both an array covariance check and a GC write barrier. Arrays in .NET are covariant, meaning a `TDerived[]` can be used as a `TBase[]`, e.g. a `string[]` can be used as an `object[]`; consequently, storing an instance into an `object[]` must validate that the instance is actually of the right type (otherwise, you could have a `TDerived1[]` masquerading as a `TBase[]` and try to store a `TDerived2` into it, which would cause badness if it were to store successfully). dotnet/runtime#126547 expands calls to the runtime’s array-store helper into the individual operations it performs, exposing both the covariance check and write barrier to the JIT. When the JIT knows the array’s exact type, it can then eliminate the covariance check and optimize the barrier:\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Runtime.CompilerServices;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private readonly object[] _array = new object[4096];\n    private object _value = new();\n    [MethodImpl(MethodImplOptions.NoInlining)]\n    private static void StoreAll(object[] arr, object value)\n    {\n        for (int i = 0; i < arr.Length; i++)\n            arr[i] = value;\n    }\n    [Benchmark]\n    public object[] CovariantStore_Loop()\n    {\n        StoreAll(_array, _value);\n        return _array;\n    }\n}\n```\n| Method | Runtime | Mean | Ratio | \n|---|---|---|---|\n| CovariantStore_Loop | .NET 10.0 | 10.85 μs | 1.00 | \n| CovariantStore_Loop | .NET 11.0 | 6.042 μs | 0.56 | \n\nSometimes writes are done one at a time, but sometimes they can be batched, as happens when copying structs. dotnet/runtime#128238 extends the JIT’s heap-destination analysis from individual stores to whole-struct copies. dotnet/runtime#128542 then replaces a specialized helper that copied one reference field at a time with reference stores and vector stores for the non-reference data. Together, they let the JIT choose more efficient write barriers and copy the rest of a mixed struct with SIMD.\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Runtime.CompilerServices;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    [InlineArray(4)]\n    public struct InlineArray4Long\n    {\n        private long _element0;\n    }\n    public struct MyStruct\n    {\n        public string A;\n        public InlineArray4Long G;\n        public string B;\n    }\n    private MyStruct _src;\n    private MyStruct _dst;\n    [GlobalSetup]\n    public void Setup()\n    {\n        _src = new MyStruct { A = \"hello\", B = \"world\" };\n        _src.G[0] = 1;\n        _src.G[1] = 2;\n        _src.G[2] = 3;\n        _src.G[3] = 4;\n    }\n    [Benchmark]\n    public void HeapStructCopy() => _dst = _src;\n}\n```\n| Method | Runtime | Mean | Ratio | \n|---|---|---|---|\n| HeapStructCopy | .NET 10.0 | 4.132 ns | 1.00 | \n| HeapStructCopy | .NET 11.0 | 3.071 ns | 0.74 | \n\ndotnet/runtime#130535 handles the equivalent case for small structs that don’t contain object references. Once the JIT has turned the copy into several writes to adjacent fields, it can combine them into fewer, wider writes.\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Runtime.CompilerServices;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private Int128 _value;\n    [Benchmark]\n    public void StoreInt128() => _value = 123456789;\n}\n```\n.NET 10 stores the low and high halves separately. .NET 11 loads the value into a vector register and writes all 16 bytes at once.\n\n```\n; x64\n; StoreInt128\n-       mov      qword ptr [rcx+8], 75BCD15\n-       xor      eax, eax\n-       mov      [rcx+10], rax\n+       vmovss   xmm0, dword ptr [RWD00]\n+       vmovups  [rcx+8], xmm0\n-; Total bytes of code 15\n+; Total bytes of code 14\n```\nThe same idea applies when the source code assigns neighboring fields individually. dotnet/runtime#126562 enables this for promoted struct locals, while dotnet/runtime#130107 extends it to adjacent fields at constant static addresses:\n\n```\n// dotnet run -c Release -f net11.0 --filter \"*\"\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Runtime.CompilerServices;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private static Point s_point;\n    [Benchmark]\n    public void SetPoint() => Set();\n    [MethodImpl(MethodImplOptions.NoInlining | MethodImplOptions.AggressiveOptimization)]\n    private static void Set()\n    {\n        s_point.X = 1;\n        s_point.Y = 2;\n    }\n    private struct Point\n    {\n        public int X;\n        public int Y;\n    }\n}\n```\nThe referenced .NET 11 x64 build combines the two 32-bit constants and writes both fields with one 64-bit store:\n\n```\n; x64\nmov     rax, 200000001\nmov     rcx, <address of s_point>\nmov     [rcx], rax\n```\ndotnet/runtime#127487 applies a related improvement when stack protection requires a struct parameter to be copied. It uses consistently sized writes so a subsequent wider read doesn’t need to wait for the processor to reconcile overlapping stores.\n\nWrite barriers are only one part of the interaction between generated code\nand the garbage collector. During a compacting collection, the GC needs to\nplan where surviving objects will move and then update references to them. To\ndo that efficiently, it records their addresses, sorts those addresses, and\ngroups adjacent survivors into regions called “plugs.” With enough live\nobjects, sorting these mark lists becomes a meaningful part of the collection.\nRecent x86/x64 runtimes use a vectorized `vxsort` implementation for\nsufficiently large lists. In .NET 11,\ndotnet/runtime#110692 from\n@a74nh extends that support to Arm64.\n\nThe generation assigned to GC metadata matters just as much as the speed of\none collection. .NET’s generational GC is based on the observation that most\nobjects die young: generation 0 and generation 1 collections, collectively\ncalled ephemeral collections, run frequently and should avoid revisiting\nstate that has already survived into generation 2. A dependent handle\nassociates a primary object with a secondary object, keeping the secondary\nalive while the primary remains reachable; `ConditionalWeakTable<TKey, TValue>` is built on this mechanism. Previously, the handle itself didn’t age\nwith its referents, so every ephemeral collection continued scanning it even\nafter both objects had become long-lived.\ndotnet/runtime#78746 ages\ndependent handles accordingly and moves a handle back to a younger generation\nwhen necessary. Old handles can therefore be skipped by young collections\nwithout compromising reachability.\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing System.Runtime.CompilerServices;\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private ConditionalWeakTable<object, object> _table = new();\n    private object[] _keys = [];\n    [Params(100_000, 1_000_000)]\n    public int Handles { get; set; }\n    [GlobalSetup]\n    public void Setup()\n    {\n        _table = new();\n        _keys = new object[Handles];\n        for (int i = 0; i < _keys.Length; i++)\n        {\n            object key = new();\n            _keys[i] = key;\n            _table.Add(key, new object());\n        }\n        GC.Collect(2, GCCollectionMode.Forced, blocking: true, compacting: true);\n    }\n    [Benchmark]\n    public void CollectGen0() =>\n        GC.Collect(0, GCCollectionMode.Forced, blocking: true, compacting: false);\n}\n```\n| Method | Runtime | Handles | Mean | Ratio | \n|---|---|---|---|---|\n| CollectGen0 | .NET 10.0 | 100000 | 1.522 ms | 1.00 | \n| CollectGen0 | .NET 11.0 | 100000 | 255.5 μs | 0.17 | \n| CollectGen0 | .NET 10.0 | 1000000 | 10.737 ms | 1.00 | \n| CollectGen0 | .NET 11.0 | 1000000 | 310.7 μs | 0.029 | \n\n### Runtime Knowledge and Frozen Data\n\nThe JIT can optimize only the facts it knows. Some facts come from its own analysis; others are contracts supplied by the runtime, such as which helpers have side effects, the length of a newly allocated string, or whether a data object will ever move.\n\nA generic virtual call such as `baseReference.Foo<string>()` may need help from the runtime to find the implementation for both the object’s actual type and the generic argument. If that lookup appears to have arbitrary side effects, the JIT has to perform it exactly where it occurs, rather than possibly resulting on a cached answer from a previous lookup. In .NET 11, dotnet/runtime#122017 teaches the JIT more precisely which exceptions these runtime helpers can throw and whether they otherwise have side effects. The JIT can then share repeated lookups or move an unchanging lookup out of a loop:\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Runtime.CompilerServices;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    public abstract class Base\n    {\n        public abstract void Foo<T>();\n    }\n    public class Derived : Base\n    {\n        public override void Foo<T>() { }\n    }\n    private Base _b = new Derived();\n    [Benchmark]\n    public void GvmCseHoist()\n    {\n        Base b = _b;\n        b.Foo<string>();\n        b.Foo<int>();\n        b.Foo<string>();\n        b.Foo<int>();\n        for (int i = 0; i < 10; i++)\n            b.Foo<double>();\n    }\n}\n```\nIn .NET 11, the repeated lookups outside the loop are shared and the loop’s lookup is performed once, not ten times.\n\n| Method | Runtime | Mean | Ratio | \n|---|---|---|---|\n| GvmCseHoist | .NET 10.0 | 45.75 ns | 1.00 | \n| GvmCseHoist | .NET 11.0 | 24.06 ns | 0.53 | \n\nProfile data is another way the JIT learns what matters. Inlining could previously hide important work from the instrumentation used to gather that data. dotnet/runtime#119658 allows the inlined code to be instrumented as well, giving later PGO-driven compilation a more complete picture of the hot paths.\n\n### JIT Throughput and Cleanup\n\nThe quality of the generated code isn’t the only concern; the time spent producing it matters too. Every analysis the JIT performs has a cost. dotnet/runtime#123856 removes checks and maps from Global Assertion Propagation whose bookkeeping wasn’t paying for itself. This is the recurring balancing act in the development of the JIT: retaining the information that enables meaningful optimizations while avoiding analysis overhead whose code-quality benefit is negligible.\n\ndotnet/runtime#127363 makes profile-guided optimization more resilient with OSR (on-stack replacement), which replaces a method while one of its loops is already running. Because that execution begins in the middle of the method rather than at its normal entry, reconstructed profile data doesn’t always line up perfectly with the paths actually available. The JIT now estimates the likelihood of those paths rather than asserting or abandoning the profile.\n\nOptimizations can leave behind code that’s no longer reachable, so the JIT also needs to be good at dead code removal. dotnet/runtime#126223 runs another sweep whenever the method’s branching structure changes, catching blocks made obsolete by earlier transformations.\n\nAnd dotnet/runtime#128515 from @BoyBaykiller repeatedly combines equivalent return and throw endings, removing duplicate exit paths and sometimes exposing more code that can be shared.\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    [Benchmark]\n    [Arguments((byte)9)]\n    public bool IsLinearWhiteSpace(byte value) =>\n        value <= 32 &&\n        (value == 32 || value == 10 || value == 13 || value == 9);\n}\n```\nIn .NET 10, tail merging combines the paths that return `false`, but not both paths that return `true`. As a result, the JIT’s bit test covers three of the four values, with a separate comparison for `9`. In .NET 11, the true returns are merged as well, enabling all four values to be handled by the same bit test:\n\n```\n; x64\n-       movzx    ecx, dl\n-       cmp      ecx, 20\n-       jg       M00_L02\n-       cmp      ecx, 20\n-       ja       M00_L01\n-       mov      eax, 0FFFFDBFF\n-       bt       rax, rcx\n-       jae      M00_L00\n-       mov      eax, 1\n-       ret\n-M00_L00:\n-       cmp      ecx, 9\n-       sete     al\n-       movzx    eax, al\n-       ret\n-M00_L01:\n+       movzx    eax, dl\n+       cmp      eax, 20\n+       jg       M00_L00\n+       cmp      eax, 20\n+       ja       M00_L00\n+       mov      ecx, 0FFFFD9FF\n+       bt       rcx, rax\n+       jb       M00_L00\n+       mov      eax, 1\n+       ret\n+M00_L00:\n        xor      eax, eax\n        ret\n-; Total bytes of code 43\n+; Total bytes of code 33\n```\nAlso related to dead code, a call that never returns, such as one that always throws, makes everything after it unreachable. In .NET 11, after inlining, dotnet/runtime#128513 removes the remaining statements and outgoing paths from such a block and marks it as ending in a throw, exposing the dead code early enough for the cleanup passes above to remove it.\n\n## Startup and Deployment\n\nBefore managed `Main` can run, the native host needs to locate the application’s dependencies, CoreCLR needs to load enough types and code to begin execution, and various pieces of framework infrastructure need to initialize themselves. Work removed from any of those stages helps the application get going sooner, improving startup time.\n\nThe host starts by reading the application’s `.deps.json`, turning its entries into paths, and building the trusted platform assembly (TPA) list. That list tells CoreCLR which framework and application assemblies it can resolve by simple name. Several costs in this process scaled with the number of assets rather than with the amount of useful work. dotnet/runtime#123568 in .NET 11 avoids checking every asset against a servicing directory unless the resolver is actually probing that directory. dotnet/runtime#123919 avoids repeatedly comparing the servicing-directory name and copying every dependency asset while constructing the TPA list, avoiding a lot of allocation. dotnet/runtime#125251 removes more allocation by normalizing each asset’s directory separators once when parsing the `.deps.json`, rather than normalizing the path again every time it is used.\n\nOnce the host hands off to CoreCLR, ReadyToRun (R2R) code helps avoid compiling methods before they can execute. However, initializing `Comparer<T>.Default` and `EqualityComparer<T>.Default` called a reflection-based helper whose resulting concrete comparer type wasn’t known when the R2R image was built. The comparer constructor and operations could consequently fall back to being interpreted. In .NET 11, dotnet/runtime#126204 uses specialized helpers that R2R can compile ahead of time and ensures the required comparer types are included in the image.\n\nEven better than making initialization faster is avoiding it altogether. An `EventSource` normally discovers its event metadata and computes its provider GUID when it is initialized. dotnet/runtime#121180 adds an internal source generator that performs this work when the framework is built and emits the result for its `EventSource` implementations, including the ones for core runtime tracing. Applications then don’t need to pay the reflection and setup costs when those event sources are first used.\n\nStartup also has a memory footprint outside the managed heap. Native AOT’s `AllocHeap` typically holds only small amounts of runtime metadata. On Windows, however, its virtual-memory allocator reserved a 64 KB region for each block even when it initially needed only 4 KB. In .NET 11, dotnet/runtime#122822 instead uses ordinary `new` and `delete` for these small blocks, matching the allocation strategy to the amount of memory normally involved.\n\nNote that the aforementioned R2R work wasn’t motivated only by desktop and server startup. It was also part of the substantial effort to make CoreCLR the runtime for .NET on mobile. Starting with .NET 11, .NET MAUI moved to CoreCLR for Android, iOS, and Mac Catalyst, the last .NET MAUI platforms that had still been using Mono. This is much more than swapping one execution engine for another. Those apps now use the same runtime as ASP.NET Core, cloud services, and desktop .NET, with the same JIT, garbage collector, diagnostics infrastructure, performance improvements, and bug fixes. It also brings CoreCLR’s tiered compilation, ReadyToRun, and profile-guided optimization to mobile, while providing a common foundation for NativeAOT. That combination is important: R2R and packaged profiles can precompile the code most important to startup, while the optimizing JIT can produce higher-quality code for hot methods on platforms where dynamic compilation is available. Improvements like the comparer specialization mentioned earlier keep more code on the compiled path instead of falling back to interpretation.\n\n## Threading\n\nThreading is a cross-cutting concern that impacts almost every application and service. Whether code is protecting shared state, queueing work, or coordinating asynchronous operations, small costs in the underlying machinery can quickly add up. As such, it’s something that’s revisited in every release of .NET.\n\n`Monitor` is the synchronization primitive historically used to implement `lock`, providing the most pervasively used support for mutual exclusion. It also supports sending signals, such that one thread can wait on a `Monitor` with `Monitor.Wait` for another thread to `Pulse` it. The internal object that tracks these waiters is a “condition variable.” dotnet/runtime#129083 stores that condition directly on the lock, removing a separate `ConditionalWeakTable` lookup from this already synchronization-heavy path.\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private const int RoundTripsPerInvoke = 2_000;\n    private readonly object _gate = new();\n    private int _ping;\n    private int _pong;\n    private bool _stop;\n    private Thread _responder = null!;\n    [GlobalSetup]\n    public void Setup()\n    {\n        _responder = new Thread(ResponderLoop) { IsBackground = true };\n        _responder.Start();\n    }\n    [GlobalCleanup]\n    public void Cleanup()\n    {\n        lock (_gate)\n        {\n            _stop = true;\n            Monitor.PulseAll(_gate);\n        }\n        _responder.Join();\n    }\n    private void ResponderLoop()\n    {\n        lock (_gate)\n        {\n            int seen = 0;\n            while (true)\n            {\n                while (_ping == seen && !_stop)\n                    Monitor.Wait(_gate);\n                if (_stop)\n                    return;\n                seen = _ping;\n                _pong = seen;\n                Monitor.PulseAll(_gate);\n            }\n        }\n    }\n    [Benchmark(OperationsPerInvoke = RoundTripsPerInvoke)]\n    public int PingPong_MonitorWaitPulse()\n    {\n        lock (_gate)\n        {\n            for (int i = 0; i < RoundTripsPerInvoke; i++)\n            {\n                _ping++;\n                int expected = _ping;\n                Monitor.PulseAll(_gate);\n                while (_pong != expected)\n                    Monitor.Wait(_gate);\n            }\n            return _pong;\n        }\n    }\n}\n```\n| Method | Runtime | Mean | Ratio | \n|---|---|---|---|\n| PingPong_MonitorWaitPulse | .NET 10.0 | 4.194 μs | 1.00 | \n| PingPong_MonitorWaitPulse | .NET 11.0 | 3.517 μs | 0.84 | \n\nIn the case of `Monitor`, that improvement targeted the specific shared implementation. In other cases, the costs are spread out in a more peanut butter manner across lots of code. dotnet/runtime#125274 removes some of that peanut butter by removing unnecessary `volatile` annotations from a wide range of library fields whose correctness already comes from locks, `Interlocked`, or one-time initialization. On x86/x64 hardware, which already provides a strong memory model, those annotations generally don’t result in extra instructions, though they can still constrain compiler optimizations. Arm, however, permits more reordering, so the JIT often needs to emit memory fences to provide `volatile`‘s guarantees. Removing the annotations where they’re redundant therefore can end up removing unnecessary fences from Arm’s generated code.\n\nSimilar considerations apply to code in the runtime. dotnet/runtime#125259 replaces\nfull memory barriers in the runtime’s `HashMap` with the narrower acquire and\nrelease operations actually required. On top of that, many VM\nhash tables, including its `EEHashTable`, are read constantly but updated only\noccasionally. dotnet/runtime#124822\nadds epoch-based reclamation, enabling readers to avoid entering cooperative\nGC mode simply to keep an old set of buckets alive. And dotnet/runtime#129640 replaces\nthe previous byte-at-a-time hash used by these tables with an xxHash\nimplementation that consumes four bytes at a time.\n\nAlong the same lines, in .NET 11 dotnet/runtime#122726 reduces the scheduling overhead around small thread-pool work items. It removes unnecessary memory fences and shared-state updates, checks in with the thread-pool controller once per batch rather than once per item, spends less time spinning on a semaphore, and requests another worker only when the queued work shows one is needed. The result is less coordination overhead and fewer workers woken just as the queue becomes empty.\n\nEarlier in this post, we talked about runtime async, which can have a significant impact on the performance of `async`/`await` code, how they produce `Task`s, and so on. They’re not the only improvements in .NET 11 related to `Task`s, though.\n\nOne fun one is a new analyzer, CA2027, introduced in dotnet/sdk#51452. With that, the SDK can point out problematic usage of `Task.Delay` that I’ve seen on multiple occasions to lead to non-trivial performance issues in large scale services. Consider this code:\n\n```\nTask someTask = ...;\nif (await Task.WhenAny(someTask, Task.Delay(timeout)) != someTask) // oops!\n{\n    throw new TimeoutException();\n}\n```\nThe developer that wrote this is obviously trying to implement a timeout. The problem, however, is that this leaks. In the hopefully common case where `someTask` completes really quickly, the `Task.Delay` will still be pending. That `Delay` has associated with it a `System.Threading.Timer` that’s consuming valuable resources, as well as other data in memory, and if this `timeout` is long and this code is on a hotter path, we could accumulate thousands upon thousands of those timers. That in turn can increase memory use and slow down other calls that interact with timers.\n\nThe fix is to instead use the `Task.WaitAsync` method, introduced all the way back in .NET 6. It provides a much more efficient mechanism for doing this same kind of timed waiting, and it correctly handles all the relevant cleanup. CA2027 will detect common forms of this issue and recommend the replacement.\n\n## Numerics\n\n`BigInteger` is one of those types that many applications may never need, but\nfor those that do, there’s often no practical substitute. It powers workloads\nranging from cryptography and number theory to compilers and applications that\nneed to parse, format, or compute with integers larger than the fixed-width\nprimitives can hold. Despite that need, however, `BigInteger` hasn’t received the\nsame steady stream of performance investment as many of .NET’s other core\ntypes. Thankfully, in .NET 11 it gets a makeover.\n\ndotnet/runtime#125799 rewrote significant portions of `BigInteger`‘s implementation, changing its limbs (the fixed-size pieces stored in its backing array) from `uint` to `nuint` (`UIntPtr`). That makes no effective difference on a 32-bit machine. On a 64-bit machine, however, each limb grows from 32 to 64 bits; since most arithmetic on a 64-bit value on a 64-bit platform costs no more than the corresponding 32-bit operation, each step can therefore process twice as many bits in the same number of cycles. The implementation also improves the algorithms around those wider limbs, including Montgomery multiplication and sliding-window exponentiation in `ModPow`, fused bitwise steps, additional hardware intrinsics, loop unrolling, and caching. That all builds on top of other optimizations that were done previously in the release, such as faster conversion of huge values to decimal text in dotnet/runtime#112178 from @kzrnm, dotnet/runtime#112876 from @kzrnm using Toom-Cook multiplication for sufficiently large operands, and improved shifts and rotations thanks to dotnet/runtime#113005 from @kzrnm. Toom-Cook splits each operand into several chunks and combines smaller products, doing less work than the straightforward every-limb-by-every-limb algorithm once the operands are large enough.\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Numerics;\nusing System.Globalization;\nusing System.Text;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    [Params(64, 512)]\n    public int Limbs;\n    private BigInteger _a;\n    private BigInteger _b;\n    private BigInteger _shiftSubject;\n    private BigInteger _hugeValueForToString;\n    private string _decimalDigits100000 = \"\";\n    private byte[] _utf8Digits1000 = [];\n    private byte[] _utf8FormatBuffer = new byte[120_000];\n    private BigInteger _divideDividendBelowThreshold;\n    private BigInteger _divideDivisorBelowThreshold;\n    private BigInteger _divideDividendAboveThreshold;\n    private BigInteger _divideDivisorAboveThreshold;\n    [GlobalSetup]\n    public void Setup()\n    {\n        _a = MakeDeterministicBigInteger(Limbs, seed: 1);\n        _b = MakeDeterministicBigInteger(Limbs, seed: 2);\n        _shiftSubject = MakeDeterministicBigInteger(Limbs, seed: 3);\n        _decimalDigits100000 = MakeDeterministicDecimalDigits(100_000);\n        _hugeValueForToString = BigInteger.Parse(_decimalDigits100000, CultureInfo.InvariantCulture);\n        string decimalDigits1000 = MakeDeterministicDecimalDigits(1_000);\n        _utf8Digits1000 = Encoding.UTF8.GetBytes(decimalDigits1000);\n        _divideDivisorBelowThreshold = MakeDeterministicBigInteger(16, seed: 4);\n        _divideDividendBelowThreshold = MakeDeterministicBigInteger(16 + 96, seed: 5);\n        _divideDivisorAboveThreshold = MakeDeterministicBigInteger(128, seed: 6);\n        _divideDividendAboveThreshold = MakeDeterministicBigInteger(128 + 96, seed: 7);\n    }\n    private static BigInteger MakeDeterministicBigInteger(int limbCount, int seed)\n    {\n        Random rng = new(seed);\n        byte[] bytes = new byte[(limbCount * 4) + 1]; // trailing 0 byte keeps the value positive\n        rng.NextBytes(bytes);\n        bytes[^1] = 0;\n        return new BigInteger(bytes);\n    }\n    private static string MakeDeterministicDecimalDigits(int digitCount)\n    {\n        StringBuilder sb = new(digitCount);\n        sb.Append('9'); // avoid a leading zero, which would shorten the effective digit count\n        Random rng = new(42);\n        for (int i = 1; i < digitCount; i++)\n            sb.Append((char)('0' + rng.Next(0, 10)));\n        return sb.ToString();\n    }\n    [Benchmark]\n    public BigInteger Divide_BelowBurnikelZieglerThreshold() => _divideDividendBelowThreshold / _divideDivisorBelowThreshold;\n    [Benchmark]\n    public BigInteger Divide_AboveBurnikelZieglerThreshold() => _divideDividendAboveThreshold / _divideDivisorAboveThreshold;\n    [Benchmark]\n    public BigInteger Multiply() => _a * _b;\n    [Benchmark]\n    public BigInteger ShiftLeft() => _shiftSubject << 12345;\n    [Benchmark]\n    public BigInteger ParseLargeDecimal() => BigInteger.Parse(_decimalDigits100000, CultureInfo.InvariantCulture);\n    [Benchmark]\n    public string ToStringLargeDecimal() => _hugeValueForToString.ToString(CultureInfo.InvariantCulture);\n}\n```\n| Method | Runtime | Limbs | Mean | Ratio | \n|---|---|---|---|---|\n| Divide_BelowBurnikelZieglerThreshold | .NET 10.0 | 64 | 2,954.3 ns | 1.00 | \n| Divide_BelowBurnikelZieglerThreshold | .NET 11.0 | 64 | 1,493.8 ns | 0.51 | \n| Divide_AboveBurnikelZieglerThreshold | .NET 10.0 | 64 | 10,766.7 ns | 1.00 | \n| Divide_AboveBurnikelZieglerThreshold | .NET 11.0 | 64 | 6,303.4 ns | 0.59 | \n| Multiply | .NET 10.0 | 64 | 2,359.5 ns | 1.00 | \n| Multiply | .NET 11.0 | 64 | 1,328.2 ns | 0.56 | \n| ShiftLeft | .NET 10.0 | 64 | 217.1 ns | 1.00 | \n| ShiftLeft | .NET 11.0 | 64 | 121.6 ns | 0.56 | \n| ParseLargeDecimal | .NET 10.0 | 64 | 9,111,926.1 ns | 1.00 | \n| ParseLargeDecimal | .NET 11.0 | 64 | 3,899,478.0 ns | 0.43 | \n| ToStringLargeDecimal | .NET 10.0 | 64 | 135,894,135.4 ns | 1.00 | \n| ToStringLargeDecimal | .NET 11.0 | 64 | 7,868,359.3 ns | 0.058 | \n| Divide_BelowBurnikelZieglerThreshold | .NET 10.0 | 512 | 2,894.6 ns | 1.00 | \n| Divide_BelowBurnikelZieglerThreshold | .NET 11.0 | 512 | 1,482.0 ns | 0.51 | \n| Divide_AboveBurnikelZieglerThreshold | .NET 10.0 | 512 | 10,751.7 ns | 1.00 | \n| Divide_AboveBurnikelZieglerThreshold | .NET 11.0 | 512 | 6,317.8 ns | 0.59 | \n| Multiply | .NET 10.0 | 512 | 68,063.6 ns | 1.00 | \n| Multiply | .NET 11.0 | 512 | 35,443.5 ns | 0.52 | \n| ShiftLeft | .NET 10.0 | 512 | 673.0 ns | 1.00 | \n| ShiftLeft | .NET 11.0 | 512 | 309.4 ns | 0.46 | \n| ParseLargeDecimal | .NET 10.0 | 512 | 9,149,793.0 ns | 1.00 | \n| ParseLargeDecimal | .NET 11.0 | 512 | 3,891,313.0 ns | 0.43 | \n| ToStringLargeDecimal | .NET 10.0 | 512 | 135,829,594.6 ns | 1.00 | \n| ToStringLargeDecimal | .NET 11.0 | 512 | 7,880,631.1 ns | 0.058 | \n\nIn addition to internal changes, `BigInteger` also gained new public APIs that avoid transcoding. Protocols and storage formats increasingly expose text as UTF-8 bytes, but the previous parsing and formatting APIs required UTF-16 characters. Callers therefore had to decode the input into a temporary string before parsing, or format into characters and encode the result back to bytes. dotnet/runtime#117745 adds direct UTF-8 parsing and formatting to both `BigInteger` and `Complex`, sharing the generic numeric machinery used for UTF-16 and letting those consumers operate on their original representation.\n\ndotnet/runtime#130721 improves a different `BigInteger` boundary: casting to `double` and `float`. The general conversion needs to inspect the arbitrary-width magnitude, locate its highest set bits, and perform the rounding required by the target floating-point format. But many `BigInteger` instances are much smaller than that machinery is designed for… the implementation now recognizes values that fit in 64 bits and routes them through the hardware’s native integer conversion support.\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Numerics;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private readonly BigInteger _small = (BigInteger.One << 63) + 123;\n    private readonly BigInteger _large = (BigInteger.One << 1023) + (BigInteger.One << 511) + 123;\n    [Benchmark] public double SmallToDouble() => (double)_small;\n    [Benchmark] public float SmallToSingle() => (float)_small;\n    [Benchmark] public double LargeToDouble() => (double)_large;\n    [Benchmark] public float LargeToSingle() => (float)_large;\n}\n```\n| Method | Runtime | Mean | Ratio | \n|---|---|---|---|\n| SmallToDouble | .NET 10.0 | 2.873 ns | 1.00 | \n| SmallToDouble | .NET 11.0 | 1.764 ns | 0.61 | \n| SmallToSingle | .NET 10.0 | 3.548 ns | 1.00 | \n| SmallToSingle | .NET 11.0 | 1.764 ns | 0.50 | \n| LargeToDouble | .NET 10.0 | 2.863 ns | 1.00 | \n| LargeToDouble | .NET 11.0 | 2.797 ns | 0.98 | \n| LargeToSingle | .NET 10.0 | 3.559 ns | 1.00 | \n| LargeToSingle | .NET 11.0 | 2.849 ns | 0.80 | \n\nThe same limb-widening advantages given to `BigInteger` in .NET 11 were also extended to the core floating-point types. Parsing a very long decimal input and formatting a floating-point value with many requested digits both need temporary arbitrary-precision arithmetic once the value no longer fits in the normal mantissa. .NET uses a separate internal `Number.BigInteger` for that work. dotnet/runtime#132577 applies the same native-width limb representation to that type, reducing the amount of per-limb work in floating-point parsing, formatting, and rounding.\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing System.Globalization;\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private readonly string _longFraction = \"0.\" + new string('1', 768);\n    [Benchmark]\n    public double ParseLongFraction() => double.Parse(_longFraction, CultureInfo.InvariantCulture);\n    [Benchmark]\n    public string FormatSubnormal() => double.Epsilon.ToString(\"G99\", CultureInfo.InvariantCulture);\n}\n```\n| Method | Runtime | Mean | Ratio | \n|---|---|---|---|\n| ParseLongFraction | .NET 10.0 | 8.592 μs | 1.00 | \n| ParseLongFraction | .NET 11.0 | 3.569 μs | 0.42 | \n| FormatSubnormal | .NET 10.0 | 6.884 μs | 1.00 | \n| FormatSubnormal | .NET 11.0 | 1.237 μs | 0.18 | \n\nThe .NET 11 improvements aren’t limited to the scalar representations underlying\n`BigInteger` and floating-point parsing and formatting. Other numerical types improve as well. Consider `Matrix4x4`. A 4×4 matrix\ndeterminant combines products of many independent matrix elements, making it a\nnatural fit for SIMD. dotnet/runtime#123954\nfrom @alexcovington adds an SSE\nimplementation of `Matrix4x4.GetDeterminant`, evaluating several of those\nproducts in parallel:\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Numerics;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private readonly Matrix4x4 _matrix =\n        Matrix4x4.CreateFromYawPitchRoll(0.4f, 0.8f, 1.1f) *\n        Matrix4x4.CreateTranslation(1.5f, -2.5f, 3.25f) *\n        Matrix4x4.CreateScale(1.1f, 0.9f, 1.05f);\n    [Benchmark]\n    public float GetDeterminant() => _matrix.GetDeterminant();\n}\n```\n| Method | Runtime | Mean | Ratio | \n|---|---|---|---|\n| GetDeterminant | .NET 10.0 | 3.836 ns | 1.00 | \n| GetDeterminant | .NET 11.0 | 2.645 ns | 0.69 | \n\nThe `System.Numerics.Tensors` APIs are designed to perform the same numerical operation over many values, making them a natural fit for SIMD. dotnet/runtime#126052 adds vector implementations of inverse sine to the portable vector types and uses them in `TensorPrimitives.Asin`. The tensor loop now evaluates a polynomial approximation for several inputs together, with special handling near the ends of the function’s `[-1, 1]` domain, rather than calling `MathF.Asin` or `Math.Asin` separately for every element:\n\n```\n// Run separately so each target uses its matching System.Numerics.Tensors package:\n// dotnet run -c Release -f net10.0 --filter \"*\"\n// dotnet run -c Release -f net11.0 --filter \"*\"\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Numerics.Tensors;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private const int Length = 4096;\n    private float[] _floatsIn = new float[Length];\n    private float[] _floatsOut = new float[Length];\n    private double[] _doublesIn = new double[Length];\n    private double[] _doublesOut = new double[Length];\n    [GlobalSetup]\n    public void Setup()\n    {\n        Random rng = new(42);\n        for (int i = 0; i < Length; i++)\n        {\n            float v = (float)((rng.NextDouble() * 2.0) - 1.0); // Asin's domain is [-1, 1]\n            _floatsIn[i] = v;\n            _doublesIn[i] = v;\n        }\n    }\n    [Benchmark]\n    public float AsinFloat()\n    {\n        TensorPrimitives.Asin(_floatsIn, _floatsOut);\n        return _floatsOut[0];\n    }\n    [Benchmark]\n    public double AsinDouble()\n    {\n        TensorPrimitives.Asin(_doublesIn, _doublesOut);\n        return _doublesOut[0];\n    }\n}\n```\n| Method | Runtime | Mean | Ratio | \n|---|---|---|---|\n| AsinFloat | .NET 10.0 | 33.72 μs | 1.00 | \n| AsinFloat | .NET 11.0 | 8.240 μs | 0.24 | \n| AsinDouble | .NET 10.0 | 35.80 μs | 1.00 | \n| AsinDouble | .NET 11.0 | 10.975 μs | 0.31 | \n\n`TensorPrimitives` also picked up a few more targeted SIMD improvements. For floating-point values, `BitIncrement` and `BitDecrement` move to the immediately adjacent representable value; despite their names, they can’t simply add or subtract one, as they also need to handle signed zero, infinities, and NaNs correctly. dotnet/runtime#123610 and dotnet/runtime#123754 process multiple `float`/`double` and `Half` values at once, respectively. The `Half` path works directly with the raw `ushort` bit patterns, avoiding conversion to `float` and back, and both paths use vector masks and conditional selection rather than calling a scalar helper for every element.\n\n```\n// Run separately so each target uses its matching System.Numerics.Tensors package:\n// dotnet run -c Release -f net10.0 --filter \"*\"\n// dotnet run -c Release -f net11.0 --filter \"*\"\nusing System.Numerics.Tensors;\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private const int Length = 4096;\n    private readonly float[] _floats = new float[Length];\n    private readonly float[] _floatDestination = new float[Length];\n    private readonly double[] _doubles = new double[Length];\n    private readonly double[] _doubleDestination = new double[Length];\n    private readonly Half[] _halves = new Half[Length];\n    private readonly Half[] _halfDestination = new Half[Length];\n    [GlobalSetup]\n    public void Setup()\n    {\n        for (int i = 0; i < Length; i++)\n        {\n            float value = (i & 7) switch\n            {\n                0 => 0,\n                1 => -0.0f,\n                2 => float.PositiveInfinity,\n                3 => float.NegativeInfinity,\n                4 => float.NaN,\n                _ => i / 7.0f,\n            };\n            _floats[i] = value;\n            _doubles[i] = value;\n            _halves[i] = (Half)value;\n        }\n    }\n    [Benchmark]\n    public void BitIncrementFloat() => TensorPrimitives.BitIncrement(_floats, _floatDestination);\n    [Benchmark]\n    public void BitIncrementDouble() => TensorPrimitives.BitIncrement(_doubles, _doubleDestination);\n    [Benchmark]\n    public void BitIncrementHalf() => TensorPrimitives.BitIncrement(_halves, _halfDestination);\n}\n```\n| Method | Runtime | Mean | Ratio | \n|---|---|---|---|\n| BitIncrementFloat | .NET 10.0 | 4.223 μs | 1.00 | \n| BitIncrementFloat | .NET 11.0 | 1,071.1 ns | 0.25 | \n| BitIncrementDouble | .NET 10.0 | 4.223 μs | 1.00 | \n| BitIncrementDouble | .NET 11.0 | 2,140.6 ns | 0.51 | \n| BitIncrementHalf | .NET 10.0 | 3.918 μs | 1.00 | \n| BitIncrementHalf | .NET 11.0 | 573.6 ns | 0.15 | \n\ndotnet/runtime#124280 removes a more mechanical cost from `TensorPrimitives.Round`: for `digits == 0`, the old code invoked a full-span rounding kernel and then continued through another full-span pass. Returning immediately removes that redundant traversal and overwrite of the destination.\n\n```\n// Run separately so each target uses its matching System.Numerics.Tensors package:\n// dotnet run -c Release -f net10.0 --filter \"*\"\n// dotnet run -c Release -f net11.0 --filter \"*\"\nusing System.Numerics.Tensors;\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private const int Length = 4096;\n    private readonly float[] _source = new float[Length];\n    private readonly float[] _destination = new float[Length];\n    [Benchmark]\n    public void RoundZero() => TensorPrimitives.Round(_source, 0, MidpointRounding.ToEven, _destination);\n}\n```\n| Method | Runtime | Mean | Ratio | \n|---|---|---|---|\n| RoundZero | .NET 10.0 | 882.7 ns | 1.00 | \n| RoundZero | .NET 11.0 | 205.5 ns | 0.23 | \n\n`Half` comparisons are faster as well. Previously, `Half.CompareTo` separately\nasked whether one value was less than, greater than, or equal to the other,\nrepeating the special handling required for NaN and signed zero each time.\ndotnet/runtime#131297 performs\nthat work once and then arranges the underlying bits into a form that can be\ncompared directly, while still treating `+0` and `-0` as equal. On x64 with\nAVX2, it also makes `CompareTo`, `<`, and `<=` faster by converting the operands\nto `float`, which the hardware can do very efficiently. Equality remains\nbit-based, as that’s already the cheaper approach.\n\n```\n// Run on x64 with AVX2:\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private readonly Half[] _left = Enumerable.Range(0, 4096).Select(i => (Half)(i - 2048)).ToArray();\n    private readonly Half[] _right = Enumerable.Range(0, 4096).Select(i => (Half)(2048 - i)).ToArray();\n    [Benchmark]\n    public int CompareTo()\n    {\n        int sum = 0;\n        for (int i = 0; i < _left.Length; i++)\n            sum += _left[i].CompareTo(_right[i]);\n        return sum;\n    }\n    [Benchmark]\n    public int LessThan()\n    {\n        int count = 0;\n        for (int i = 0; i < _left.Length; i++)\n            count += _left[i] < _right[i] ? 1 : 0;\n        return count;\n    }\n}\n```\n| Method | Runtime | Mean | Ratio | \n|---|---|---|---|\n| CompareTo | .NET 10.0 | 9.340 μs | 1.00 | \n| CompareTo | .NET 11.0 | 5.745 μs | 0.62 | \n| LessThan | .NET 10.0 | 7.514 μs | 1.00 | \n| LessThan | .NET 11.0 | 5.672 μs | 0.75 | \n\nMultiplying two 64-bit integers produces a 128-bit result, and x64 has instructions that provide both 64-bit halves directly. dotnet/runtime#117261 from @Daniel-Svensson exposes those signed and unsigned forms through an `X86Base.X64.BigMul` intrinsic. `Math.BigMul` can then map directly to `imul` or `mul` and return both halves in registers, avoiding the extra instructions and register shuffling required by the previous paths.\n\n```\n// Run on x64:\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private readonly long _signedLeft = 0x1234_5678_9ABC_DEF;\n    private readonly long _signedRight = 0x0FED_CBA9_8765_432;\n    private readonly ulong _unsignedLeft = 0xFEDC_BA98_7654_3210;\n    private readonly ulong _unsignedRight = 0x1234_5678_9ABC_DEF0;\n    [Benchmark]\n    public long Signed()\n    {\n        long high = Math.BigMul(_signedLeft, _signedRight, out long low);\n        return high ^ low;\n    }\n    [Benchmark]\n    public ulong Unsigned()\n    {\n        ulong high = Math.BigMul(_unsignedLeft, _unsignedRight, out ulong low);\n        return high ^ low;\n    }\n}\n```\n| Method | Runtime | Mean | Ratio | Code Size | \n|---|---|---|---|---|\n| Signed | .NET 10.0 | 2.210 ns | 1.00 | 65 B | \n| Signed | .NET 11.0 | 1.344 ns | 0.61 | 12 B | \n| Unsigned | .NET 10.0 | 1.446 ns | 1.00 | 39 B | \n| Unsigned | .NET 11.0 | 1.323 ns | 0.92 | 12 B | \n\nFixed-format numeric and identifier helpers benefit from a much simpler technique: establish the exact span length once, then let the JIT reuse that fact. dotnet/runtime#119254 from @xtqqczze applies that pattern in `Decimal`, `Guid`, and `IPAddress`, removing repeated bounds checks.\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private const string Value = \"a8098c1a-f86e-11da-bd1a-00112444be1e\";\n    [Benchmark]\n    public bool TryParseExactD() => Guid.TryParseExact(Value, \"D\", out _);\n}\n```\n| Method | Runtime | Mean | Ratio | \n|---|---|---|---|\n| TryParseExactD | .NET 10.0 | 15.94 ns | 1.00 | \n| TryParseExactD | .NET 11.0 | 12.60 ns | 0.79 | \n\n`Guid` has been improving every .NET release, and sees several improvements in .NET 11. Whenever possible, .NET tries to maintain similar performance and behaviors across operating systems, but low-level functionality often simply delegates to the operating system, exposing that OS’ characteristics. When it comes to random number generation, historically cryptographically-secure random number generation, as is used in `Guid.NewGuid`, has been a bit slower on Linux than on Windows due to using `/dev/urandom` as the source of entropy. dotnet/runtime#123540 from @reedz moves `Guid.NewGuid()` off of that file-descriptor path to the `getrandom()` syscall, avoiding descriptor setup and reads through the file abstraction.\n\nAnd on the subject of randomness, dotnet/runtime#119890 from @hamarb123 removes two pieces of work from `Random.Shuffle`: an unnecessary copy of the span length and a branch that skipped swapping an element with itself. A self-swap is harmless and uncommon, while testing for it adds an unpredictable branch to every iteration. The difference is most visible for short arrays and small value types, where the swap itself is cheap:\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    [Params(16, 4096)]\n    public int Length;\n    private readonly Random _random = new(42);\n    private int[] _values = [];\n    [GlobalSetup]\n    public void Setup() => _values = Enumerable.Range(0, Length).ToArray();\n    [Benchmark]\n    public int ShuffleSmallValueType()\n    {\n        _random.Shuffle(_values);\n        return _values[0] + _values[^1];\n    }\n}\n```\n| Method | Runtime | Length | Mean | Ratio | \n|---|---|---|---|---|\n| ShuffleSmallValueType | .NET 10.0 | 16 | 138.9 ns | 1.00 | \n| ShuffleSmallValueType | .NET 11.0 | 16 | 89.34 ns | 0.64 | \n| ShuffleSmallValueType | .NET 10.0 | 4096 | 25,925.1 ns | 1.00 | \n| ShuffleSmallValueType | .NET 11.0 | 4096 | 14,151.08 ns | 0.55 | \n\n`Random` itself picked up a small but pointed code-generation fix. `Random.InternalSample` contains a condition that’s inherently hard for the processor to predict, so it’s better implemented with conditional instructions than with a branch. The JIT’s if-conversion support we previously discussed would have been able to do that transformation, except it doesn’t currently support if-conversion inside of loops, which is a pretty common place to find an inlined `Random.Next` call. dotnet/runtime#131714 marks the helper as `[MethodImpl(MethodImplOptions.NoInlining)]` to preserve the branch-free form; once the JIT can perform if-conversion inside loops, that annotation can be reconsidered.\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private readonly Random _random = new(42);\n    [Benchmark]\n    public int Next()\n    {\n        int sum = 0;\n        for (int i = 0; i < 1024; i++)\n            sum += _random.Next();\n        return sum;\n    }\n}\n```\n| Method | Runtime | Mean | Ratio | \n|---|---|---|---|\n| Next | .NET 10.0 | 5.954 μs | 1.00 | \n| Next | .NET 11.0 | 3.275 μs | 0.55 | \n\n## Globalization\n\nMany globalization-related APIs sit atop data that can be expensive to locate\nand interpret. `DateTime.Now`, for example, depends on time-zone transition\ndata, while casing and parsing depend on native globalization services and\nculture-specific tables.\n\ndotnet/runtime#119662 substantially reworks `TimeZoneInfo` around that observation. Determining an offset isn’t always a fixed arithmetic operation: daylight-saving rules can vary by year, and historical rules can contain multiple transitions and exceptional cases. Once the transitions for a zone and year have been interpreted, however, other conversions in that year can reuse them. Similarly, the local offset used by `DateTime.Now` can’t change between transition instants. Conversions now reuse cached per-year transition data rather than repeatedly walking adjustment rules, while `DateTime.Now` caches the active UTC offset together with the instant at which it next needs to be recomputed.\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private readonly DateTime _utc = new(2026, 7, 15, 12, 0, 0, DateTimeKind.Utc);\n    private readonly DateTime _local = new(2026, 7, 15, 5, 0, 0, DateTimeKind.Unspecified);\n    private readonly TimeZoneInfo _zone = TimeZoneInfo.FindSystemTimeZoneById(\n        OperatingSystem.IsWindows() ? \"Pacific Standard Time\" : \"America/Los_Angeles\");\n    [Benchmark]\n    public DateTime ConvertTimeFromUtc() => TimeZoneInfo.ConvertTimeFromUtc(_utc, _zone);\n    [Benchmark]\n    public DateTime ConvertTimeToUtc() => TimeZoneInfo.ConvertTimeToUtc(_local, _zone);\n    [Benchmark]\n    public DateTime GetLocalNow() => DateTime.Now;\n}\n```\n| Method | Runtime | Mean | Ratio | \n|---|---|---|---|\n| ConvertTimeFromUtc | .NET 10.0 | 45.13 ns | 1.00 | \n| ConvertTimeFromUtc | .NET 11.0 | 19.44 ns | 0.43 | \n| ConvertTimeToUtc | .NET 10.0 | 51.97 ns | 1.00 | \n| ConvertTimeToUtc | .NET 11.0 | 20.23 ns | 0.39 | \n| GetLocalNow | .NET 10.0 | 76.41 ns | 1.00 | \n| GetLocalNow | .NET 11.0 | 34.39 ns | 0.45 | \n\ndotnet/runtime#120685 separates two costs in invariant casing. With the normal globalization configuration, `ToUpperInvariant` and `ToLowerInvariant` now try a managed ASCII path first, so casing ASCII text can avoid or delay initialization of ICU, the native library .NET uses for culture-aware globalization. In invariant-globalization mode, where ICU isn’t loaded at all, that managed path also improves ASCII casing throughput. Non-ASCII input still needs the appropriate globalization path.\n\n```\n// DOTNET_SYSTEM_GLOBALIZATION_INVARIANT=1 dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private readonly string _short = \"runtime\";\n    private readonly string _long = new('a', 139);\n    [Benchmark]\n    public string ShortAscii() => _short.ToUpperInvariant();\n    [Benchmark]\n    public string LongAscii() => _long.ToUpperInvariant();\n}\n```\n| Method | Runtime | Mean | Ratio | Allocated | \n|---|---|---|---|---|\n| ShortAscii | .NET 10.0 | 18.12 ns | 1.00 | 40 B | \n| ShortAscii | .NET 11.0 | 14.62 ns | 0.81 | 40 B | \n| LongAscii | .NET 10.0 | 248.38 ns | 1.00 | 304 B | \n| LongAscii | .NET 11.0 | 39.22 ns | 0.16 | 304 B | \n\nSeveral smaller changes remove setup around date and culture data. dotnet/runtime#123886 allocates the `DateTimeFormatInfo` date-word table only for cultures that actually contain such words. And dotnet/runtime#122918 replaces synchronized, boxing `Hashtable` caches used by time-zone and encoding tables with typed `ConcurrentDictionary` instances.\n\nThe round-trip `\"O\"` date format always contains exactly seven fractional-second digits, matching the 10,000,000 ticks in a second. dotnet/runtime#129005 parses those digits directly as ticks, avoiding a conversion through `double` followed by division, multiplication, and rounding. Formatting benefits from specialization as well. dotnet/runtime#129374 routes invariant `DateTime.ToString(\"G\")` through the existing fixed-format fast path, bypassing the general culture-aware formatter. `DateTimeOffset` retains the general path because its offset changes the output:\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Running;\nusing System.Globalization;\nusing BenchmarkDotNet.Attributes;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private readonly DateTime _dateTime = new(2024, 3, 15, 13, 45, 30, DateTimeKind.Utc);\n    [Benchmark]\n    public string DateTime_ToString_G() => _dateTime.ToString(\"G\", CultureInfo.InvariantCulture);\n}\n```\n| Method | Runtime | Mean | Ratio | \n|---|---|---|---|\n| DateTime_ToString_G | .NET 10.0 | 65.54 ns | 1.00 | \n| DateTime_ToString_G | .NET 11.0 | 29.76 ns | 0.45 | \n\n## Strings and Spans\n\nUTF-8 is everywhere, from web protocols and JSON payloads to files on disk. Since .NET strings use UTF-16, applications frequently need to convert between the two, making it especially important for those conversions to be fast. UTF-8 encoding must validate UTF-16 surrogate pairs as it counts and converts them. On Arm64, the vectorized implementation in .NET 10 still examined individual elements when counting the resulting UTF-8 bytes and checking that surrogates were correctly paired. That gets expensive for text containing many supplementary characters, as every surrogate-heavy vector falls back to this element-by-element work. dotnet/runtime#121981 from @ylpoonlg instead performs the counting and surrogate checks with vector-wide operations. As part of that work, it also unifies most of the x86 and Arm64 implementations, retaining small platform-specific helpers where the instruction sets differ:\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Text;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private const int Length = 4096;\n    private string _validWithSurrogatePairs = string.Empty;\n    [GlobalSetup]\n    public void Setup()\n    {\n        Random rng = new(42);\n        StringBuilder sb = new(Length);\n        while (sb.Length < Length - 2)\n        {\n            sb.Append((char)('A' + rng.Next(0, 26)));\n            sb.Append(\"\\U0001F600\"); // emoji -> surrogate pair\n        }\n        _validWithSurrogatePairs = sb.ToString();\n    }\n    [Benchmark]\n    public int ValidWithSurrogatePairs() => Encoding.UTF8.GetByteCount(_validWithSurrogatePairs);\n}\n```\nThis input deliberately contains a surrogate pair for every ASCII character, making the removed per-element work especially visible.\n\n| Method | Runtime | Mean | Ratio | \n|---|---|---|---|\n| ValidWithSurrogatePairs | .NET 10.0 | 2.994 μs | 1.00 | \n| ValidWithSurrogatePairs | .NET 11.0 | 856.8 ns | 0.29 | \n\nThe byte-to-char direction was also improved on Arm. UTF-8 decoding can copy ASCII bytes directly to UTF-16 characters, but as soon as we find the first non-ASCII byte, we need the full multi-byte decoder. The vector loop therefore needs both a fast test for whether any lane is non-ASCII and, only when one is found, its exact position. Calculating that position for every all-ASCII vector wastes work on the overwhelmingly common fast path. dotnet/runtime#121382 from @ylpoonlg first performs the cheap vector-wide test and then computes the lane index only after that test succeeds.\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Text;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private readonly byte[] _ascii = Enumerable.Repeat((byte)'a', 16_384).ToArray();\n    [Benchmark]\n    public int Utf8GetCharCount() => Encoding.UTF8.GetCharCount(_ascii);\n}\n```\nWith all-ASCII input, every vector can stay on the cheap path:\n\n| Method | Runtime | Mean | Ratio | \n|---|---|---|---|\n| Utf8GetCharCount | .NET 10.0 | 443.5 ns | 1.00 | \n| Utf8GetCharCount | .NET 11.0 | 210.6 ns | 0.47 | \n\nBase64 is commonly used when binary data needs to travel through\ntext-oriented formats and protocols. Its encoder naturally works in groups of\nthree input bytes and four output characters, but the line-breaking option\nalso needs to stop at the MIME-style 76-character boundary and insert `\\r\\n`.\nThe older implementation handled that formatting through a separate scalar\npath. dotnet/runtime#123403 brings the optimized span-based Base64 encoder to `Convert.ToBase64String` with `Base64FormattingOptions.InsertLineBreaks`, processing each line with the same vectorized core and handling the separators around it:\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Running;\nusing BenchmarkDotNet.Attributes;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    [Params(57, 570)]\n    public int ByteLength { get; set; }\n    private byte[] _bytes = [];\n    [GlobalSetup]\n    public void Setup()\n    {\n        _bytes = new byte[ByteLength];\n        new Random(42).NextBytes(_bytes);\n    }\n    [Benchmark]\n    public string ToBase64String_InsertLineBreaks() => Convert.ToBase64String(_bytes, Base64FormattingOptions.InsertLineBreaks);\n}\n```\n| Method | Runtime | ByteLength | Mean | Ratio | \n|---|---|---|---|---|\n| ToBase64String_InsertLineBreaks | .NET 10.0 | 57 | 60.95 ns | 1.00 | \n| ToBase64String_InsertLineBreaks | .NET 11.0 | 57 | 23.91 ns | 0.39 | \n| ToBase64String_InsertLineBreaks | .NET 10.0 | 570 | 560.66 ns | 1.00 | \n| ToBase64String_InsertLineBreaks | .NET 11.0 | 570 | 194.99 ns | 0.35 | \n\nBase64 decoding got the same treatment from the other direction. `Base64.DecodeFromUtf8InPlace` decodes in place, overwriting the encoded input with the decoded bytes. In .NET 10, it still employed a scalar loop, long after the out-of-place `DecodeFromUtf8` had acquired AVX-512, AVX2, AdvSimd, and SSSE3 paths. In-place decoding turns out to be safe to vectorize precisely because of Base64’s ratio: 4 bytes read produce 3 bytes written, so the write cursor always trails the read cursor, and each vector store, including its zero-padded overshoot, ends at or before the next vector load and never clobbers source that hasn’t been read yet. dotnet/runtime#131333 therefore reuses the existing decode helpers for the in-place path.\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing System.Buffers;\nusing System.Buffers.Text;\nusing System.Text;\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private readonly byte[] _encoded = Encoding.ASCII.GetBytes(Convert.ToBase64String(new byte[16_384]));\n    private byte[] _buffer = [];\n    [IterationSetup]\n    public void Setup() => _buffer = (byte[])_encoded.Clone();\n    [Benchmark]\n    public OperationStatus Decode() => Base64.DecodeFromUtf8InPlace(_buffer, out _);\n}\n```\n| Method | Runtime | Mean | Ratio | \n|---|---|---|---|\n| Decode | .NET 10.0 | 10.70 μs | 1.00 | \n| Decode | .NET 11.0 | 2.256 μs | 0.21 | \n\n`MemoryExtensions.CommonPrefixLength` compares two spans and returns how many\nelements they share at the beginning (“hello” and “help”, for example, have a\ncommon prefix length of 3). Internally, it utilizes a helper that slices whichever input was longer to\nthe length of the shorter one. dotnet/runtime#121104\nfrom @xtqqczze simplifies that helper: after\nshortening the second span if necessary, it always slices the first span to\nthe second’s length. That gives the JIT the same explicit relationship between\nthe two lengths regardless of which input started out longer.\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System;\nusing System.Linq;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private readonly string[] _shorter = Enumerable.Repeat(\"value\", 64).ToArray();\n    private readonly string[] _longer = Enumerable.Repeat(\"value\", 128).ToArray();\n    [Benchmark]\n    public int ShorterFirst() => _shorter.AsSpan().CommonPrefixLength(_longer);\n    [Benchmark]\n    public int LongerFirst() => _longer.AsSpan().CommonPrefixLength(_shorter);\n}\n```\nThe longer-first case was already efficient. The change improves the shorter-first case, bringing the two orderings to essentially the same throughput:\n\n| Method | Runtime | Mean | Ratio | \n|---|---|---|---|\n| ShorterFirst | .NET 10.0 | 45.16 ns | 1.00 | \n| ShorterFirst | .NET 11.0 | 25.43 ns | 0.56 | \n| LongerFirst | .NET 10.0 | 26.40 ns | 1.00 | \n| LongerFirst | .NET 11.0 | 26.31 ns | 1.00 | \n\nText processing often starts by obtaining an `Encoding`. Properties such as\n`Encoding.UTF8` provide fast access to popular encodings, while legacy code\npages can be made available by registering `CodePagesEncodingProvider`. In\n.NET 10, that provider’s tables, including the name lookup used by\n`Encoding.GetEncoding(string)` once the provider is registered, used\nreader-writer locks. dotnet/runtime#125001\nreplaces those caches with `ConcurrentDictionary` instances, allowing\nwarmed-up provider lookups to proceed without acquiring the reader lock.\n\nOn `string` itself, dotnet/runtime#130361 from @prozolic recognizes when `string.Concat(IEnumerable<string?>)` receives a `string[]` or `List<string?>` and passes its contiguous storage directly to the span-based implementation:\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing System.Collections.Generic;\nusing System.Linq;\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[MemoryDiagnoser(false), HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private readonly IEnumerable<string?> _array = Enumerable.Range(0, 1_000).Select(i => i.ToString()).ToArray();\n    private readonly IEnumerable<string?> _list = Enumerable.Range(0, 1_000).Select(i => i.ToString()).ToList();\n    [Benchmark]\n    public string Array() => string.Concat(_array);\n    [Benchmark]\n    public string List() => string.Concat(_list);\n}\n```\n| Method | Runtime | Mean | Ratio | Allocated | \n|---|---|---|---|---|\n| Array | .NET 10.0 | 4.230 μs | 1.00 | 5.7 KB | \n| Array | .NET 11.0 | 3.285 μs | 0.78 | 5.67 KB | \n| List | .NET 10.0 | 6.703 μs | 1.00 | 5.71 KB | \n| List | .NET 11.0 | 3.253 μs | 0.49 | 5.67 KB | \n\nSome of my favorite improvements in .NET are the tiny ones that show up everywhere. A good example of that is in dotnet/roslyn#82729. Previously, when you wrote `span[start..]`, the compiler would lower that to the equivalent of `span.Slice(start, span.Length - start)`. The JIT has made strides towards compiling this exactly how it would `span.Slice(start)`, but everyone is better off if the C# compiler just emits that in the first place. And it now does. The difference is clear in the IL for a method that returns `span[start..]`:\n\n```\n; Platform-independent IL\n-// Before: 21 bytes\n+// After: 9 bytes\n-.locals init ([0] System.Span<char>&, [1] int32)\n ldarga.s span\n-stloc.0\n ldarg.1\n-stloc.1\n-ldloc.0\n-ldloc.1\n-ldloc.0\n-call instance int32 System.Span<char>::get_Length()\n-ldloc.1\n-sub\n-call instance System.Span<char> System.Span<char>::Slice(int32, int32)\n+call instance System.Span<char> System.Span<char>::Slice(int32)\n ret\n```\n### Searching and Comparing\n\nSearching in one way, shape, or form is one of the most common things programs do. And when it comes to searching text, regular expressions are an extremely common and helpful way to specify and perform said search. .NET’s regex support has improved by leaps and bounds over the years, with significant investments in .NET 5 and .NET 7 and then every release since, including .NET 11.\n\nWhen a `Regex` instance is created, it needs to parse the incoming regular expression pattern and turn it into a form it can utilize for performing the actual searches. The regex language is very expressive and enables multiple ways of specifying the same pattern, some more efficient to process than others, so as part of parsing, `Regex` applies a variety of simplifications and optimizations over the parsed tree in order to put it into an ideal form, as well as to learn facts about the pattern to further optimize later processing (such as discovering a minimum and maximum length of any possible match). Each of these transformations can in turn expose more opportunity for other transformations, but based on the order the transformations are applied, sometimes those opportunities can be missed. In .NET 11, dotnet/runtime#125289 gives compiled and source-generated regexes one final cleanup pass after the whole-pattern optimizations have reshaped the pattern. Consider the pattern `[ab]+c[ab]+|[ab]+`. On input containing a long run of `a`s with no `c`, the .NET 10 source-generated matcher first scans the whole run for the first alternative, fails when it doesn’t find the `c`, and then scans the same run again for the second alternative. The final cleanup pass factors out the common `[ab]+`, leaving `c[ab]+` as an optional suffix:\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\"\n// dotnet run -c Release -f net11.0 --filter \"*\"\nusing System.Text.RegularExpressions;\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic partial class Benchmarks\n{\n    private readonly string _input = new('a', 4096);\n    [Benchmark]\n    public bool SharedPrefix() => SharedPrefixRegex().IsMatch(_input);\n    [GeneratedRegex(\"[ab]+c[ab]+|[ab]+\")]\n    private static partial Regex SharedPrefixRegex();\n}\n```\n| Method | Runtime | Mean | Ratio | \n|---|---|---|---|\n| SharedPrefix | .NET 10.0 | 550.9 ns | 1.00 | \n| SharedPrefix | .NET 11.0 | 282.8 ns | 0.51 | \n\nBeyond doing additional passes, several other changes improve what those\nanalysis passes can see. For example, for a pattern like `(http|https)` with\nordinal ignore-case matching, for uninteresting reasons previously the engine\nwould extract a prefix of `\"htt\"`, even though it could have extracted\n`\"http\"`. dotnet/runtime#124881\nimproves that, enabling the engine to skip far more false candidates. The\ninput here contains 25,000 `\"htt\"` prefixes that aren’t followed by a `p`\nbefore the final match:\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\"\n// dotnet run -c Release -f net11.0 --filter \"*\"\nusing System.Text.RegularExpressions;\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic partial class Benchmarks\n{\n    private readonly string _input = string.Concat(Enumerable.Repeat(\"httx\", 25_000)) + \"https\";\n    [Benchmark]\n    public bool IgnoreCaseAlternation() => Http.IsMatch(_input);\n    [GeneratedRegex(\"(http|https)\", RegexOptions.IgnoreCase)]\n    private static partial Regex Http { get; }\n}\n```\n| Method | Runtime | Mean | Ratio | \n|---|---|---|---|\n| IgnoreCaseAlternation | .NET 10.0 | 415.0 μs | 1.00 | \n| IgnoreCaseAlternation | .NET 11.0 | 7.012 μs | 0.017 | \n\nWhen those transformation passes are looking for various patterns, sometimes small\nthings obscure what they’re trying to see, and they miss optimizations.\ndotnet/runtime#124842\nimproves a case where captures were getting in the way of identifying a\nsearchable prefix. For a pattern like `\\b(in)\\b` with\n`RegexOptions.IgnoreCase`, it will now discover it can search for\nordinal-ignore-case `\"in\"`.\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\"\n// dotnet run -c Release -f net11.0 --filter \"*\"\nusing System.Text.RegularExpressions;\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic partial class Benchmarks\n{\n    private readonly string _input = string.Concat(Enumerable.Repeat(\"xn \", 33_333)) + \"in\";\n    [Benchmark]\n    public bool IgnoreCaseCapturedPrefix() => CapturedPrefix.IsMatch(_input);\n    [GeneratedRegex(@\"\\b(in)\\b\", RegexOptions.IgnoreCase)]\n    private static partial Regex CapturedPrefix { get; }\n}\n```\n| Method | Runtime | Mean | Ratio | \n|---|---|---|---|\n| IgnoreCaseCapturedPrefix | .NET 10.0 | 277.7 μs | 1.00 | \n| IgnoreCaseCapturedPrefix | .NET 11.0 | 7.286 μs | 0.026 | \n\nAs these cases highlight, one of the most impactful things we can do for regular expression processing is improve the engine’s ability to find things to search for as the next possible place a match could apply, and to optimize that search. dotnet/runtime#124736 does that. For compiled, source-generated, and `NonBacktracking` regexes, it improves how the engine is able to search for one of several literal prefixes. For `agggtaaa|tttaccct`, for example, the .NET 10 source generator first searched for `[ag]` at offset 3 and then checked nearby characters for `[gt]`. That’s a weak filter for an input full of `a` characters, where almost every position becomes a candidate. The .NET 11 generator instead searches for the complete `agggtaaa` and `tttaccct` strings with `SearchValues<string>`. A frequency heuristic selects this approach only for case-sensitive alternatives where whole-string searching is expected to reject more false candidates than the available character-set filter.\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\"\n// dotnet run -c Release -f net11.0 --filter \"*\"\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Text.RegularExpressions;\nBenchmarkSwitcher.FromAssembly(typeof(RegexPrefixBenchmarks).Assembly).Run(args);\n[HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic partial class RegexPrefixBenchmarks\n{\n    private const string Pattern = \"agggtaaa|tttaccct\";\n    private readonly string _match = new string('a', 100_000) + \"tttaccct\";\n    private readonly string _miss = new('a', 100_000);\n    [Benchmark]\n    public bool Match() => Generated.IsMatch(_match);\n    [Benchmark]\n    public bool Miss() => Generated.IsMatch(_miss);\n    [GeneratedRegex(Pattern)]\n    private static partial Regex Generated { get; }\n}\n```\n| Method | Runtime | Mean | Ratio | \n|---|---|---|---|\n| Match | .NET 10.0 | 861.9 μs | 1.00 | \n| Match | .NET 11.0 | 9.251 μs | 0.011 | \n| Miss | .NET 10.0 | 861.4 μs | 1.00 | \n| Miss | .NET 11.0 | 9.647 μs | 0.011 | \n\nOf course, searching for the next place to match isn’t the only opportunity for improvement. Once you’ve found that place, you need to try to match, and we want to optimize that further, too.\n\nConsider the pattern `\\b\\w+n\\b`. The `\\w+` can match `n`, which means we can’t automatically treat this loop as being atomic. Normally, after matching the loop greedily and failing to match `n`, we’d need to backtrack looking for the next viable place to match the `n`. But if what comes after the `n` (in this case, a boundary) can’t possibly match the loop, we can avoid doing that search. dotnet/runtime#125636 teaches the compiled and source-generated engines to prove that and test the final position directly rather than searching backward through the loop’s existing match. The same idea applies to other loops followed by a literal when the engine can prove that trying earlier positions can’t change the result.\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\"\n// dotnet run -c Release -f net11.0 --filter \"*\"\nusing BenchmarkDotNet.Running;\nusing System.Text.RegularExpressions;\nusing BenchmarkDotNet.Attributes;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic partial class Benchmarks\n{\n    private const int WordLength = 5000;\n    private readonly string _matchingWord = new string('a', WordLength - 1) + \"n\";\n    private readonly string _nonMatchingWord = new string('a', WordLength - 1) + \"b\";\n    [GeneratedRegex(@\"\\b\\w+n\\b\")]\n    private static partial Regex Generated { get; }\n    [Benchmark]\n    public bool Matching() => Generated.IsMatch(_matchingWord);\n    [Benchmark]\n    public bool NonMatching() => Generated.IsMatch(_nonMatchingWord);\n}\n```\n| Method | Runtime | Mean | Ratio | \n|---|---|---|---|\n| Matching | .NET 10.0 | 3.441 μs | 1.00 | \n| Matching | .NET 11.0 | 3.118 μs | 0.91 | \n| NonMatching | .NET 10.0 | 83.656 μs | 1.00 | \n| NonMatching | .NET 11.0 | 69.984 μs | 0.84 | \n\nA match can sometimes be ruled out before examining any of the input’s\ncharacters. When matching starts at position zero, a fixed-length pattern with\na leading `\\A` or non-multiline `^` and a trailing `\\z` can match only when the\nwhole input has exactly that length.\ndotnet/runtime#120916 emits\nthat length check up front for the compiled and source-generated engines when\nthe computed maximum length equals the minimum required length. Here, the\npattern requires exactly 512 characters while the input contains 513:\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\"\n// dotnet run -c Release -f net11.0 --filter \"*\"\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Text.RegularExpressions;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic partial class Benchmarks\n{\n    private readonly string _tooLong = new('a', 513);\n    [GeneratedRegex(@\"\\A[a-z]{512}\\z\")]\n    private static partial Regex Generated { get; }\n    [Benchmark]\n    public bool AnchoredReject() => Generated.IsMatch(_tooLong);\n}\n```\n| Method | Runtime | Mean | Ratio | \n|---|---|---|---|\n| AnchoredReject | .NET 10.0 | 41.06 ns | 1.00 | \n| AnchoredReject | .NET 11.0 | 16.05 ns | 0.39 | \n\nIn general, we’ve tried to keep the compilers behind `RegexOptions.Compiled` (which emits IL) and the source generator (which emits C#) as close to 1:1 as possible. There are a few cases, however, where they have diverged from each other, generally where one was able to easily utilize some feature of the target language the other didn’t have. A good example is with alternations. If several left-to-right atomic branches each begin with a different literal character, the engine can read that character and jump straight to the matching branch rather than testing each branch in order. With C#, we emitted a `switch`, which the C# compiler could then lower to IL using various strategies. For IL, in .NET 10 and earlier, without the C# compiler to provide those optimizations, we just skipped the optimization. Now in .NET 11, dotnet/runtime#122959 emits a similar implementation to what the C# compiler would have, bringing this optimization to `RegexOptions.Compiled`.\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing System.Text.RegularExpressions;\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private readonly string _input = string.Concat(Enumerable.Repeat(\"p15\", 10_000));\n    private readonly Regex _regex = new(@\"(?>a0|b1|c2|d3|e4|f5|g6|h7|i8|j9|k10|l11|m12|n13|o14|p15)\", RegexOptions.Compiled);\n    [Benchmark]\n    public int DispatchToFinalBranch() => _regex.Count(_input);\n}\n```\n| Method | Runtime | Mean | Ratio | \n|---|---|---|---|\n| DispatchToFinalBranch | .NET 10.0 | 233.9 μs | 1.00 | \n| DispatchToFinalBranch | .NET 11.0 | 159.5 μs | 0.68 | \n\nAnother of the few differences between compiled and source-generated regexes had to do with backreferences. A case-sensitive backreference, such as the `\\1` in `([a-z]+)-\\1`, asks whether the next input equals text that was previously captured in the match. Source-generated regexes were using the optimized `SequenceEqual` to do that comparison, whereas `RegexOptions.Compiled` wasn’t. With dotnet/runtime#123914 in .NET 11, now it does.\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Text.RegularExpressions;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private readonly string _input = new string('a', 256) + \"-\" + new string('a', 256);\n    private readonly Regex _regex = new(@\"^([a-z]{256})-\\1$\", RegexOptions.Compiled);\n    [Benchmark]\n    public bool Backreference() => _regex.IsMatch(_input);\n}\n```\n| Method | Runtime | Mean | Ratio | \n|---|---|---|---|\n| Backreference | .NET 10.0 | 168.4 ns | 1.00 | \n| Backreference | .NET 11.0 | 49.59 ns | 0.29 | \n\nSearching isn’t limited to `Regex`, of course. Many other methods in .NET help finding things and comparing things, some of which get notable bumps in .NET 11.\n\nThe `Ascii` class provides optimized helpers for validating and manipulating ASCII text. Members like `Equals` are already vectorized in .NET 10, but in .NET 11, dotnet/runtime#123115 improves that implementation by ensuring that inputs of length 8 through 15 can be vectorized, as well.\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Running;\nusing BenchmarkDotNet.Attributes;\nusing System.Text;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    [Params(8, 15)]\n    public int Length { get; set; }\n    private byte[] _bytes = [];\n    private char[] _charsMatching = [];\n    [GlobalSetup]\n    public void Setup()\n    {\n        _bytes = new byte[Length];\n        _charsMatching = new char[Length];\n        for (int i = 0; i < Length; i++)\n        {\n            byte b = (byte)('a' + (i % 26));\n            _bytes[i] = b;\n            _charsMatching[i] = (char)b;\n        }\n    }\n    [Benchmark]\n    public bool Equals_Matching() => Ascii.Equals(_bytes, _charsMatching);\n}\n```\n| Method | Runtime | Length | Mean | Ratio | \n|---|---|---|---|---|\n| Equals_Matching | .NET 10.0 | 8 | 3.834 ns | 1.00 | \n| Equals_Matching | .NET 11.0 | 8 | 1.966 ns | 0.51 | \n| Equals_Matching | .NET 10.0 | 15 | 6.177 ns | 1.00 | \n| Equals_Matching | .NET 11.0 | 15 | 2.398 ns | 0.39 | \n\ndotnet/runtime#130644\nalso improves equality performance, in this case with `SequenceEqual` over a\nspan of `Guid` or `Int128`. Previously, `SequenceEqual` treated these as\narbitrary structures and compared them one element at a time. The PR teaches\nthe runtime that their fixed bitwise representations are suitable for\ncomparison as raw bytes. That enables the same optimized memory-comparison\npath used for primitive types, including JIT unrolling and vectorization:\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private readonly Guid[] _guids1 = new Guid[2];\n    private readonly Guid[] _guids2 = new Guid[2];\n    private readonly Int128[] _int128s1 = new Int128[2];\n    private readonly Int128[] _int128s2 = new Int128[2];\n    [Benchmark]\n    public bool GuidEqual() => _guids1.AsSpan().SequenceEqual(_guids2);\n    [Benchmark]\n    public bool Int128Equal() => _int128s1.AsSpan().SequenceEqual(_int128s2);\n}\n```\n| Method | Runtime | Mean | Ratio | \n|---|---|---|---|\n| GuidEqual | .NET 10.0 | 2.960 ns | 1.00 | \n| GuidEqual | .NET 11.0 | 2.077 ns | 0.70 | \n| Int128Equal | .NET 10.0 | 3.547 ns | 1.00 | \n| Int128Equal | .NET 11.0 | 2.077 ns | 0.59 | \n\nAnother improvement in .NET 11 is to `string.Split`.\nBefore `string.Split` can produce the resulting strings, it first needs to\nfind the characters that separate them and record their positions. In .NET 10, that search is\nalready vectorized: rather than examine one UTF-16 character at a time, it\nloads a vector’s worth, compares all of its lanes against the separator in\nparallel, and turns the comparison result into a mask identifying any matches.\nIt then advances to the next vector, or uses the mask to record the matching\npositions. In .NET 11, on x86/x64, dotnet/runtime#125379\nfrom @hamarb123 makes the no-match path cheaper\nfor ASCII separators. It loads two vectors of UTF-16 characters, packs their\n16-bit elements into one vector of bytes, and checks that combined vector for\nthe separator. If there isn’t a match, it has skipped twice as much input with\none packed comparison; only a possible match requires the full 16-bit\ncomparisons needed to determine its exact position. (This same packing technique\nis already employed elsewhere, such as in various `SearchValues<T>` implementations.)\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private readonly string _input = new('a', 16_384);\n    [Benchmark]\n    public int SplitNoSeparators() => _input.Split(',').Length;\n}\n```\n| Method | Runtime | Mean | Ratio | \n|---|---|---|---|\n| SplitNoSeparators | .NET 10.0 | 786.0 ns | 1.00 | \n| SplitNoSeparators | .NET 11.0 | 404.7 ns | 0.51 | \n\nA related Arm64 text-search improvement comes from\ndotnet/runtime#126678.\nA vector comparison produces a vector whose elements are all zero for\nnon-matches and all one bits for matches. Finding the first or last match then\nrequires condensing those bits into a scalar value and counting its leading or\ntrailing zeros. On x86, the runtime can use a movemask instruction for that\ncondensing step. Arm64 has no direct equivalent, and the old implementation\nneeded a sequence of shifts, widening operations, and a horizontal add to achieve it.\nThe .NET 11 implementation now uses `shrn`, Arm64’s shift-right-and-narrow\ninstruction, to pack the relevant bits directly. `SearchValues<char>` uses these helpers, so the following benchmark reaches\nthe affected code while searching for a match at the end of the input.\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Buffers;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser(maxDepth: 3)]\n[HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private const int Length = 8_192;\n    private static readonly SearchValues<char> s_vowels = SearchValues.Create(\"aeiouAEIOU\");\n    private static readonly string s_input = new string('x', Length - 1) + 'e';\n    [Benchmark]\n    public int IndexOfAny() => s_input.AsSpan().IndexOfAny(s_vowels);\n}\n```\nThe .NET 10 match-index path requires this sequence:\n\n```\n; Arm64\n; .NET 10\ncmeq    v16.16b, v16.16b, #0\nmovi    v17.16b, #0x80\nand     v16.16b, v16.16b, v17.16b\nldr     q17, [MASK]\nushl    v16.16b, v16.16b, v17.16b\nuxtl2   v17.8h, v16.16b\nshl     v17.8h, v17.8h, #8\nuaddw   v16.8h, v17.8h, v16.8b\naddv    h16, v16.8h\numov    w2, v16.h[0]\nmvn     w2, w2\nrbit    w2, w2\nclz     w2, w2\n```\nIn .NET 11, the equivalent work is simpler:\n\n```\n; Arm64\n; .NET 11\ncmeq    v16.16b, v16.16b, #0\nmvn     v16.16b, v16.16b\nshrn    v16.8b, v16.8h, #4\nfmov    x2, d16\nrbit    x2, x2\nclz     x2, x2\nlsr     w2, w2, #2\n```\n`MemoryExtensions` already provides span-based searches for one or more values\nwith `IndexOfAny`, and for contiguous ranges with `IndexOfAnyInRange`, along\nwith `Except`, `Contains`, and last-index variants of these operations. For\nexample, `span.IndexOfAnyInRange('0', '9')` finds the next ASCII digit.\nWhitespace is also common to search for, but the characters recognized by\n`char.IsWhiteSpace` are spread across multiple parts of Unicode rather than\nforming one contiguous range. To avoid requiring every caller to construct\nthe same `SearchValues<char>`,\ndotnet/runtime#111439 from\n@AlexRadch adds\n`ContainsAnyWhiteSpace`, `IndexOfAnyWhiteSpace`,\n`IndexOfAnyExceptWhiteSpace`, `LastIndexOfAnyWhiteSpace`, and\n`LastIndexOfAnyExceptWhiteSpace` for `ReadOnlySpan<char>`. Their shared\n`SearchValues<char>`-based implementation vectorizes these searches for\nparsers, validators, trimming code, and other text-processing code.\n\nThis is, however, a good example of how vectorization isn’t always a win. Take trimming. To trim leading whitespace, code needs to find the first character that isn’t whitespace. That character could be deep into the string, but in the most common case, there’s little or nothing to trim. A scalar loop can then return after inspecting just one or two characters, whereas the vectorized helper has fixed setup cost. It’s still worth vectorizing, because that overhead is small and the benefits when there is a lot to scan can be significant. Something to keep in mind.\n\n```\n// dotnet run -c Release -f net11.0 --filter \"*\"\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private readonly string _input = new(' ', 256);\n    [Benchmark(Baseline = true)]\n    public int Scalar()\n    {\n        ReadOnlySpan<char> input = _input;\n        for (int i = 0; i < input.Length; i++)\n        {\n            if (!char.IsWhiteSpace(input[i]))\n                return i;\n        }\n        return -1;\n    }\n    [Benchmark]\n    public int Vectorized() => _input.AsSpan().IndexOfAnyExceptWhiteSpace();\n}\n```\n| Method | Mean | Ratio | \n|---|---|---|\n| Scalar | 127.87 ns | 1.00 | \n| Vectorized | 13.14 ns | 0.10 | \n\nThis method is a particularly good fit when needing to validate that input does not contain any whitespace; that requires searching the entirety of input, which is where the vectorization in these methods shines. As an example of this, dotnet/runtime#127123 uses it to accelerate the parsing of the `\"X\"` GUID format:\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private static readonly string s_noWhitespace =\n        Guid.Parse(\"a8098c1a-f86e-11da-bd1a-00112444be1e\").ToString(\"X\");\n    [Benchmark]\n    public Guid ParseExactX() => Guid.ParseExact(s_noWhitespace, \"X\");\n}\n```\n| Method | Runtime | Mean | Ratio | \n|---|---|---|---|\n| ParseExactX | .NET 10.0 | 120.6 ns | 1.00 | \n| ParseExactX | .NET 11.0 | 87.10 ns | 0.72 | \n\nClosely related to searching is sorting. Years ago, sorting methods for `Span<T>` were added to `MemoryExtensions`. Interestingly, the method wasn’t added as `Sort<T>` but rather as `Sort<T, TComparer>` where `TComparer : IComparer<T>`. That signature enables a caller to provide a struct comparer without allocating a delegate or class-based comparer. Because the comparer is a constrained value type, the JIT should also be able to inline the comparison into the hot sorting loop. In practice, the implementation boxed the struct into an `IComparer<T>`, both allocating and turning every comparison back into an interface call. This was known at the time, but avoiding the box used generic implementation techniques that then carried too much runtime and code-size cost. Those supporting costs have since been addressed, so dotnet/runtime#116109 from @2A5F now carries a value-type comparer through `Span<T>.Sort` without boxing it. The JIT can specialize the sorting routine for that comparer and inline the comparison.\n\nThe generic specialization does increase generated code and very large comparer structs can be more expensive to copy; this optimization is aimed at the small stateless or lightly stateful structs for which the API was designed.\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[MemoryDiagnoser(false), HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private readonly int[] _source = Enumerable.Range(0, 512).Select(i => (i * 257) % 512).ToArray();\n    private int[] _values = [];\n    [IterationSetup]\n    public void Setup() => _values = (int[])_source.Clone();\n    [Benchmark]\n    public void Sort() => _values.AsSpan().Sort(new DescendingComparer());\n    private readonly struct DescendingComparer : IComparer<int>\n    {\n        public int Compare(int x, int y) => y.CompareTo(x);\n    }\n}\n```\n| Method | Runtime | Mean | Ratio | Allocated | \n|---|---|---|---|---|\n| Sort | .NET 10.0 | 11.62 μs | 1.00 | 88 B | \n| Sort | .NET 11.0 | 3.533 μs | 0.30 | – | \n\n## Collections and LINQ\n\nMuch of the collection and LINQ work in .NET 11 comes from taking better\nadvantage of information that’s already available. A collection often knows\nmuch more than an `IEnumerable<T>` can express: its count, its contiguous\nstorage, its comparer, or the layout of its hash table. Similarly, a LINQ\niterator can know how many elements it represents or how its operations were\ncomposed. Preserving that information can avoid enumeration, temporary\nstorage, repeated hashing, and other work a general-purpose implementation\nwould otherwise need to perform.\n\ndotnet/runtime#119896 from @prozolic changes `ImmutableArray.Create` to use `Array.Copy` rather than a hand-written element loop. A general element-by-element copy repeatedly performs indexing and assignment, while the runtime can specialize `Array.Copy` for the element type and size. For blittable data, it can use optimized bulk memory copies, and for reference types, which need GC write barriers, it performs the required write barriers in the runtime’s tuned copy helpers. The change therefore both simplifies the managed code and gives `ImmutableArray` access to those optimized implementations.\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Running;\nusing System.Collections.Immutable;\nusing BenchmarkDotNet.Attributes;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(\"Job\", \"Error\", \"StdDev\", \"RatioSD\", \"Median\")]\npublic class Benchmarks\n{\n    private readonly int[] _source = Enumerable.Range(0, 1_000).ToArray();\n    [Benchmark]\n    public ImmutableArray<int> CreateSlice() => ImmutableArray.Create(_source, 0, _source.Length);\n}\n```\nThis in particular makes larger copies much faster.\n\n| Method | Runtime | Mean | Ratio | \n|---|---|---|---|\n| CreateSlice | .NET 10.0 | 552.5 ns | 1.00 | \n| CreateSlice | .NET 11.0 | 277.7 ns | 0.50 | \n\ndotnet/runtime#118932 from @prozolic similarly keeps `ImmutableArrayExtensions.SequenceEqual` on optimized paths when the other sequence is an array, list, or another `ICollection<T>`.\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Running;\nusing System.Collections.Immutable;\nusing BenchmarkDotNet.Attributes;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(\"Job\", \"Error\", \"StdDev\", \"RatioSD\", \"Median\")]\npublic class Benchmarks\n{\n    private ImmutableArray<int> _immutable;\n    private List<int> _list = [];\n    [GlobalSetup]\n    public void Setup()\n    {\n        int[] values = Enumerable.Range(0, 1_000).ToArray();\n        _immutable = ImmutableArray.Create(values);\n        _list = [.. values];\n    }\n    [Benchmark]\n    public bool SequenceEqual() => _immutable.SequenceEqual(_list);\n}\n```\n| Method | Runtime | Mean | Ratio | \n|---|---|---|---|\n| SequenceEqual | .NET 10.0 | 925.8 ns | 1.00 | \n| SequenceEqual | .NET 11.0 | 122.2 ns | 0.13 | \n\n`Array.FindAll` has the opposite job: it produces a new collection. For a small result, its temporary storage used to cost more than the result itself.\ndotnet/runtime#120336 from @Henr1k80 has `Array.FindAll` collect its first four matches in an inline stack buffer rather than an intermediate `List<T>`:\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Running;\nusing BenchmarkDotNet.Attributes;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[MemoryDiagnoser(false), HideColumns(\"Job\", \"Error\", \"StdDev\", \"RatioSD\", \"Median\")]\npublic class Benchmarks\n{\n    private int[] _data = [];\n    [Params(4, 5)]\n    public int Size { get; set; }\n    [GlobalSetup]\n    public void Setup() => _data = Enumerable.Range(0, Size).ToArray();\n    [Benchmark]\n    public int[] FindAllMatch() => Array.FindAll(_data, static _ => true);\n}\n```\n| Method | Runtime | Size | Mean | Ratio | Allocated | Alloc Ratio | \n|---|---|---|---|---|---|---|\n| FindAllMatch | .NET 10.0 | 4 | 27.61 ns | 1.00 | 112 B | 1.00 | \n| FindAllMatch | .NET 11.0 | 4 | 9.212 ns | 0.33 | 40 B | 0.36 | \n| FindAllMatch | .NET 10.0 | 5 | 36.93 ns | 1.00 | 176 B | 1.00 | \n| FindAllMatch | .NET 11.0 | 5 | 11.102 ns | 0.30 | 48 B | 0.27 | \n\n`Dictionary<TKey, TValue>.Remove` had also missed an optimization already used by lookup and insertion. dotnet/runtime#125884 gives value-type keys a streamlined loop for the common default-comparer case. Because that path doesn’t need a virtual comparer call, the JIT can keep more of the operation’s state in registers; reference-type keys and custom comparers continue to use the general path.\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private readonly Guid[] _keys = Enumerable.Range(0, 512).Select(i => new Guid(i, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0)).ToArray();\n    private Dictionary<Guid, int> _dictionary = [];\n    [IterationSetup]\n    public void Setup() => _dictionary = _keys.ToDictionary(key => key, key => key.GetHashCode());\n    [Benchmark(OperationsPerInvoke = 512)]\n    public void Remove()\n    {\n        foreach (Guid key in _keys)\n            _dictionary.Remove(key);\n    }\n}\n```\n| Method | Runtime | Mean | Ratio | \n|---|---|---|---|\n| Remove | .NET 10.0 | 5.285 ns | 1.00 | \n| Remove | .NET 11.0 | 4.321 ns | 0.82 | \n\ndotnet/runtime#125893 changes `HashSet<T>`‘s internal chain walks to test the entry index against the array length with an unsigned comparison. That proves the subsequent array access is in range, allowing the JIT to remove its bounds check.\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private readonly HashSet<int> _set = Enumerable.Range(0, 4096).ToHashSet();\n    private readonly int[] _probes = Enumerable.Range(0, 4096).ToArray();\n    [Benchmark]\n    public int ContainsHits()\n    {\n        int count = 0;\n        foreach (int value in _probes)\n            count += _set.Contains(value) ? 1 : 0;\n        return count;\n    }\n}\n```\n| Method | Runtime | Mean | Ratio | \n|---|---|---|---|\n| ContainsHits | .NET 10.0 | 7.870 μs | 1.00 | \n| ContainsHits | .NET 11.0 | 7.388 μs | 0.94 | \n\ndotnet/runtime#128988 from @prozolic removes a second hash-table lookup when removing a matching key-value pair from `OrderedDictionary<TKey, TValue>` through `ICollection<KeyValuePair<TKey, TValue>>`. That interface operation must first find the key and verify that its stored value equals the supplied value. Once both checks have succeeded, the implementation already has the entry index needed for removal. Looking up the key again unnecessarily repeats its hash computation and collision-chain walk, so the updated path removes the known entry directly.\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Running;\nusing System.Collections.Generic;\nusing BenchmarkDotNet.Attributes;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(\"Job\", \"Error\", \"StdDev\", \"RatioSD\", \"Median\")]\npublic class Benchmarks\n{\n    private const int N = 10_000;\n    private OrderedDictionary<string, int> _dict = [];\n    private KeyValuePair<string, int>[] _pairs = Enumerable.Range(0, N)\n        .Select(i => new KeyValuePair<string, int>($\"key{i}\", i))\n        .ToArray();\n    [IterationSetup]\n    public void IterationSetup() => _dict = new OrderedDictionary<string, int>(_pairs);\n    [Benchmark]\n    public int Remove_ExplicitInterface()\n    {\n        ICollection<KeyValuePair<string, int>> col = _dict;\n        int removed = 0;\n        foreach (var pair in _pairs)\n            if (col.Remove(pair))\n                removed++;\n        return removed;\n    }\n}\n```\nFor 10,000 entries:\n\n| Method | Runtime | Mean | Ratio | \n|---|---|---|---|\n| Remove_ExplicitInterface | .NET 10.0 | 220.7 ms | 1.00 | \n| Remove_ExplicitInterface | .NET 11.0 | 179.0 ms | 0.81 | \n\ndotnet/runtime#122952 goes further when two hash tables have compatible layouts. Normally, `UnionWith` enumerates the source and inserts every element independently, recomputing hashes, checking for duplicates, and potentially resizing the destination along the way. If the destination is empty and both sets use compatible comparers, every source entry is already unique under exactly the equality rules the destination needs. `UnionWith` can therefore use the existing `HashSet<T>` copy-constructor fast path to clone the populated storage rather than rebuilding the same table entry by entry.\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Running;\nusing System.Collections.Generic;\nusing System.Linq;\nusing BenchmarkDotNet.Attributes;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[MemoryDiagnoser(false), HideColumns(\"Job\", \"Error\", \"StdDev\", \"RatioSD\", \"Median\")]\npublic class Benchmarks\n{\n    private readonly HashSet<int> _source = new(Enumerable.Range(0, 4_096));\n    [Benchmark]\n    public HashSet<int> FreshDestinationUnionWith()\n    {\n        HashSet<int> destination = [];\n        destination.UnionWith(_source);\n        return destination;\n    }\n}\n```\n| Method | Runtime | Mean | Ratio | Allocated | Alloc Ratio | \n|---|---|---|---|---|---|\n| FreshDestinationUnionWith | .NET 10.0 | 46.207 μs | 1.00 | 252.27 KB | 1.00 | \n| FreshDestinationUnionWith | .NET 11.0 | 2.433 μs | 0.05 | 76.07 KB | 0.30 | \n\ndotnet/runtime#128300\nfrom @AndrewP-GH also helps with collection construction. Building a `FrozenDictionary<TKey, TValue>` first requires collecting the input elements into a regular `Dictionary<TKey, TValue>` if they’re not already in one. That temporary dictionary resolves duplicate keys before the final frozen representation is chosen, but in .NET 10 it was growing incrementally even when the source’s count was readily available. This PR uses that count as the\ndictionary’s initial capacity, avoiding repeated allocation, copying, and\nrehashing as it’s populated.\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Collections.Concurrent;\nusing System.Collections.Frozen;\nusing System.Collections.Generic;\nusing System.Collections.Immutable;\nusing System.Linq;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[MemoryDiagnoser(false), HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private readonly KeyValuePair<int, int>[] _array =\n        Enumerable.Range(0, 4096).Select(i => new KeyValuePair<int, int>(i, i)).ToArray();\n    [Benchmark]\n    public FrozenDictionary<int, int> FromArray() => _array.ToFrozenDictionary();\n}\n```\n| Method | Runtime | Allocated | Alloc Ratio | \n|---|---|---|---|\n| FromArray | .NET 10.0 | 347.17 KB | 1.00 | \n| FromArray | .NET 11.0 | 127.16 KB | 0.37 | \n\n`SetEquals` asks whether two sets contain the same values, regardless of insertion order. The general implementation needs a temporary mutable set so it can account for duplicates and arbitrary enumeration order. When the other input is already a hash set with a compatible comparer, though, that reconstruction is unnecessary. dotnet/runtime#126309 from @aw0lid adds to `ImmutableHashSet<T>.SetEquals` direct zero-allocation paths for compatible `ImmutableHashSet<T>` and `HashSet<T>` inputs; with an identical comparer, the sets can be considered equal if they have the same count and if every element from one is found in the other.\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing System.Collections.Immutable;\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[MemoryDiagnoser(false)]\n[HideColumns(\"Job\", \"Error\", \"StdDev\", \"Median\", \"RatioSD\")]\npublic class Benchmarks\n{\n    private ImmutableHashSet<int> _set = ImmutableHashSet<int>.Empty;\n    private ImmutableHashSet<int> _immutable = ImmutableHashSet<int>.Empty;\n    private HashSet<int> _mutable = [];\n    [GlobalSetup]\n    public void Setup()\n    {\n        int[] items = Enumerable.Range(0, 10_000).ToArray();\n        _set = ImmutableHashSet.CreateRange(items);\n        _immutable = ImmutableHashSet.CreateRange(items);\n        _mutable = new(items);\n    }\n    [Benchmark]\n    public bool EqualImmutableHashSet() => _set.SetEquals(_immutable);\n    [Benchmark]\n    public bool EqualHashSet() => _set.SetEquals(_mutable);\n}\n```\n| Method | Runtime | Mean | Ratio | Allocated | Alloc Ratio | \n|---|---|---|---|---|---|\n| EqualImmutableHashSet | .NET 10.0 | 775.6 μs | 1.00 | 158.16 KB | 1.00 | \n| EqualImmutableHashSet | .NET 11.0 | 559.7 μs | 0.72 | – | 0 | \n| EqualHashSet | .NET 10.0 | 478.6 μs | 1.00 | 157.99 KB | 1.00 | \n| EqualHashSet | .NET 11.0 | 230.9 μs | 0.48 | – | 0 | \n\nSorted sets have a related case. `SetEquals` can be passed any `IEnumerable<T>`. That sequence might be unordered and might contain duplicate values, so `ImmutableSortedSet<T>` previously copied it into a temporary `SortedSet<T>` before performing the comparison. However, when the input is another sorted set using the same ordering comparer, both sets contain unique values and enumerate those values\nin the same order. Equality can then be determined by first comparing their\ncounts and, if those match, advancing both enumerators together. The first\nunequal pair proves the sets are different, and reaching the end without finding\na difference proves they’re equal. dotnet/runtime#126549 from\n@aw0lid recognizes this case for\n`ImmutableSortedSet<T>`, avoiding the temporary `SortedSet<T>` and comparing the two sorted sequences directly in one linear pass:\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Running;\nusing System.Collections.Generic;\nusing System.Collections.Immutable;\nusing BenchmarkDotNet.Attributes;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[MemoryDiagnoser(false), HideColumns(\"Job\", \"Error\", \"StdDev\", \"RatioSD\", \"Median\")]\npublic class Benchmarks\n{\n    private const int N = 10_000;\n    private ImmutableSortedSet<int> _set = ImmutableSortedSet<int>.Empty;\n    private ImmutableSortedSet<int> _equalSet = ImmutableSortedSet<int>.Empty;\n    [GlobalSetup]\n    public void Setup()\n    {\n        var items = new int[N];\n        for (int i = 0; i < N; i++) items[i] = i;\n        _set = ImmutableSortedSet.CreateRange(items);\n        _equalSet = ImmutableSortedSet.CreateRange(items);\n    }\n    [Benchmark]\n    public bool SetEquals_EqualImmutableSortedSet() => _set.SetEquals(_equalSet);\n}\n```\nFor 10,000 elements:\n\n| Method | Runtime | Mean | Ratio | Allocated | Alloc Ratio | \n|---|---|---|---|---|---|\n| SetEquals_EqualImmutableSortedSet | .NET 10.0 | 767.7 μs | 1.00 | 430.02 KB | 1.00 | \n| SetEquals_EqualImmutableSortedSet | .NET 11.0 | 117.6 μs | 0.15 | – | 0 | \n\n`SortedSet<T>` already enjoyed an optimization for that case in .NET 10, but it’s not left out of .NET 11 improvements. `SortedSet<T>.GetViewBetween` returns a `SortedSet<T>` view, effectively a slice of another `SortedSet<T>`, a live window onto a range of another set: changes through the view affect the original set. Clearing a view therefore can’t replace the view with an empty collection; it must find and remove every original node in that range. dotnet/runtime#126410 from @prozolic reduces the temporary storage used for that operation. The implementation pre-sizes the list of elements to remove and walks it by index rather than repeatedly removing from and shrinking the temporary list.\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Running;\nusing BenchmarkDotNet.Attributes;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[MemoryDiagnoser(false), HideColumns(\"Job\", \"Error\", \"StdDev\", \"RatioSD\", \"Median\")]\npublic class Benchmarks\n{\n    private const int N = 10_000;\n    private SortedSet<int> _fullSet = [];\n    [IterationSetup]\n    public void Setup() => _fullSet = new SortedSet<int>(Enumerable.Range(0, N));\n    [Benchmark]\n    public int GetViewBetweenThenClear()\n    {\n        SortedSet<int> view = _fullSet.GetViewBetween(0, N - 1);\n        view.Clear();\n        return _fullSet.Count;\n    }\n}\n```\n| Method | Runtime | Allocated | Alloc Ratio | \n|---|---|---|---|\n| GetViewBetweenThenClear | .NET 10.0 | 193.15 KB | 1.00 | \n| GetViewBetweenThenClear | .NET 11.0 | 103.93 KB | 0.54 | \n\nCollections are frequently consumed through LINQ. Although its operators work\nin terms of the general `IEnumerable<T>` abstraction, LINQ’s internal\niterators can preserve useful facts about their sources and the operations\nalready applied. Those facts can sometimes answer a query without enumerating\nthe source at all.\n\nFor example, consider `source.Append(x).Skip(10).LastOrDefault()`. LINQ queries are lazy, so the actual search begins only when `LastOrDefault` asks the `Skip` iterator for its last element. If `source.Append(x)` contains ten or fewer elements,\n`Skip(10)` necessarily removes all of them, leaving an empty sequence from\nwhich `LastOrDefault` must return the default value. `Append`, `Prepend`, and\n`Concat` iterators can cheaply report their total count when their underlying\nsources can do so. dotnet/runtime#123306 from @prozolic teaches the last-element path for `Skip` to compare that count with the number being skipped and immediately\nreport that there is no element, rather than searching a sequence it already\nknows is empty.\n\n```\n// dotnet run -c Release -f net10.0 --filter \"*\" --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Running;\nusing System.Linq;\nusing BenchmarkDotNet.Attributes;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[MemoryDiagnoser(false), HideColumns(\"Job\", \"Error\", \"StdDev\", \"RatioSD\", \"Median\")]\npublic class Benchmarks\n{\n    private readonly int[] _source = [1, 2, 3, 4, 5];\n    [Benchmark]\n    public int AppendSkipLastOrDefault() => _source.Append(6).Skip(10).LastOrDefault();\n}\n```\n| Method | Runtime | Mean | Ratio | Allocated | Alloc Ratio | \n|---|---|---|---|---|---|\n| AppendSkipLastOrDefault | .NET 10.0 | 44.89 ns | 1.00 | 144 B | 1.00 | \n| AppendSkipLastOrDefault | .NET 11.0 | 16.32 ns | 0.36 | 112 B | 0.78 | \n\n.NET 11 also improve’s LINQ’s `Enumerable.Sum`. `Sum` already uses SIMD. The\nmain loop processes four vectors at a time, alternating between two\naccumulators so that the additions don’t require extra moves. However, `Sum`\nalso promises to throw if the result overflows. Alongside each vector addition,\nthe implementation uses the signs of the two inputs and the result to update\nanother vector that tracks whether any lane overflowed. After every\ngroup of four vectors, the loop tests that tracking vector and branches to the\nthrowing path if needed. Overflow is rare, though, so on the common path that test and branch almost always just\nconfirm that nothing happened. dotnet/runtime#127429\nremoves that repeated work in .NET 11. It accumulates the overflow information\nacross all of the vector processing and tests it once after the vector loops\nhave completed. The checked-overflow behavior remains the same, but the normal\npath no longer needs to stop and check after every four vectors. The PR also\nsimplifies how the method walks the input, replacing unsafe reference and index\narithmetic with span-based vector loads, progressively slicing off the elements\nalready processed, and using a `foreach` for the final scalar elements.","body_html":"<blockquote><p><strong>Syndication note:</strong> This copy is truncated to fit the index size limit. Read the full article at the <a href=\"https://devblogs.microsoft.com/dotnet/performance-improvements-in-net-11/\" rel=\"nofollow ugc noopener\">canonical URL</a>.</p></blockquote>\n<p>Before television shows like <em>The Office</em> and <em>Parks and Recreation</em> cemented the mockumentary in the minds of millions, there was Christopher Guest. He didn’t invent the genre, but he’s widely recognized as one of its most influential practitioners, and for my money, there’s none better. I’ve watched <em>Waiting for Guffman</em> and <em>Best in Show</em> more times than I can count. But the one that has stuck with me the most, the one I quote at the slightest provocation, is <em>This Is Spinal Tap</em>.</p>\n<p>If you’ve seen it you already know where this is going (and if you haven’t, you now have weekend plans). The film is a fictional documentary about an aging English rock band named Spinal Tap, whose members are everything we picture when we picture over-the-top rock stars. In one of its more memorable scenes, the guitarist (Nigel) gives the filmmaker (Marty) a tour of his most prized gear, in particular showing off an amplifier unlike any other: its dials don’t stop at ten. That leads to what might be the single most quoted exchange in the entire movie:</p>\n<p><strong>Nigel:</strong> “You see, most blokes, you know, will be playing at ten. You’re on ten here, all the way up, all the way up, all the way up, you’re on ten on your guitar. Where can you go from there? Where?”</p>\n<p><strong>Marty:</strong> “I don’t know.”</p>\n<p><strong>Nigel:</strong> “Nowhere. Exactly. What we do is, if we need that extra push over the cliff, you know what we do?”</p>\n<p><strong>Marty:</strong> “Put it up to eleven?”</p>\n<p><strong>Nigel:</strong> “Eleven. Exactly. One louder.”</p>\n<p>This is .NET 11. It’s one louder, with another year’s worth of performance work having gone into making the runtime and libraries that much faster. Of course, the premise of Nigel’s special amplifier is ludicrous, as is exemplified in the subsequent few lines of dialog:</p>\n<p><strong>Marty:</strong> “Why don’t you just make ten louder and make ten be the top number and make that a little louder?”</p>\n<p><strong>Nigel:</strong> (pauses) “…these go to eleven.”</p>\n<p>In contrast, .NET 11 is actually one higher, one louder. The sections that follow are full of real improvements. A bounds check removed, an allocation that no longer happens, a lock that isn’t taken, a loop that runs in fewer cycles than it did a year ago, a comparison folded to a constant here, a redundant check hoisted out of a loop there, a couple of instructions fused into one, a syscall sidestepped, an array copy handed off to SIMD, and on and on. That’s how real performance work goes, accumulating gain after gain, each compounding on the last, until the whole thing is measurably, provably louder. And so, in this post, as I’ve done in past years with .NET 10, .NET 9, .NET 8, .NET 7, .NET 6, .NET 5, .NET Core 3.0, .NET Core 2.1, and .NET Core 2.0 before it, we’ll take an unhurried tour through hundreds of them.</p>\n<p>This is a long one. It’s meant to be. Grab your hot beverage of choice, settle in, and let’s turn it up.</p>\n<h2 id=\"benchmarking-setup\">Benchmarking Setup</h2>\n<p>As in previous years, the post is chock full of micro-benchmarks that demonstrate the individual improvements. Almost all of them use BenchmarkDotNet, and each is written to be self-contained so you can try it out yourself.</p>\n<p>Start by ensuring you have both .NET 10 and .NET 11 installed (most of the benchmarks compare the same code running on both versions) and create a new console project in a fresh <code>benchmarks</code> directory:</p>\n<pre><code>dotnet new console -o benchmarks\ncd benchmarks</code></pre>\n<p>Replace the contents of the generated <code>benchmarks.csproj</code> with the following, which multi-targets both versions so that BenchmarkDotNet can build for each:</p>\n<pre><code>&lt;Project Sdk=&quot;Microsoft.NET.Sdk&quot;&gt;\n  &lt;PropertyGroup&gt;\n    &lt;OutputType&gt;Exe&lt;/OutputType&gt;\n    &lt;TargetFrameworks&gt;net11.0;net10.0&lt;/TargetFrameworks&gt;\n    &lt;LangVersion&gt;preview&lt;/LangVersion&gt;\n    &lt;ImplicitUsings&gt;enable&lt;/ImplicitUsings&gt;\n    &lt;Nullable&gt;enable&lt;/Nullable&gt;\n    &lt;AllowUnsafeBlocks&gt;true&lt;/AllowUnsafeBlocks&gt;\n    &lt;ServerGarbageCollection&gt;true&lt;/ServerGarbageCollection&gt;\n    &lt;SystemPackageVersion Condition=&quot;&#39;$(TargetFramework)&#39; == &#39;net10.0&#39;&quot;&gt;10.0.12&lt;/SystemPackageVersion&gt;\n    &lt;SystemPackageVersion Condition=&quot;&#39;$(TargetFramework)&#39; == &#39;net11.0&#39;&quot;&gt;11.0.0-rc.1.26425.128&lt;/SystemPackageVersion&gt;\n  &lt;/PropertyGroup&gt;\n  &lt;ItemGroup&gt;\n    &lt;PackageReference Include=&quot;BenchmarkDotNet&quot; Version=&quot;0.16.0-preview.1&quot; /&gt;\n    &lt;PackageReference Include=&quot;System.IO.Hashing&quot; Version=&quot;$(SystemPackageVersion)&quot; /&gt;\n    &lt;PackageReference Include=&quot;System.Runtime.Caching&quot; Version=&quot;$(SystemPackageVersion)&quot; /&gt;\n    &lt;PackageReference Include=&quot;System.Numerics.Tensors&quot; Version=&quot;$(SystemPackageVersion)&quot; /&gt;\n  &lt;/ItemGroup&gt;\n&lt;/Project&gt;</code></pre>\n<p>For a given benchmark to test, copy its complete contents over everything in <code>Program.cs</code> and then run it. Each benchmark includes as a comment at the top the exact command to use. In most cases, it’s:</p>\n<p><code>dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0</code>\nwhich builds in Release and runs the benchmark against both .NET 10 and .NET 11, emitting a side-by-side comparison. The other common form, used when a benchmark is comparing two coding approaches on a single runtime (rather than the same code across two runtimes) is:</p>\n<p><code>dotnet run -c Release -f net11.0 --filter &quot;*&quot;</code>\nThe usual disclaimer applies: these are micro-benchmarks, many measuring operations so short that a blink would miss them. Your results will vary with your hardware, OS, runtime configuration, what else your machine happens to be doing at that exact moment, and whether Mercury is in retrograde.</p>\n<p>Every line of managed code ultimately ends up at the just-in-time compiler, so let’s start there.</p>\n<h2 id=\"jit\">JIT</h2>\n<p>Of all the places to improve .NET’s performance, few have as broad an impact as the just-in-time (JIT) compiler. C#, F#, and Visual Basic are typically compiled first to intermediate language (IL), and the JIT ultimately turns that IL into the native instructions the CPU executes. A JIT improvement can therefore benefit application and library code wherever the optimized pattern occurs, often with no source changes or recompilation of the application itself. Even removing a single instruction or proving one check unnecessary can add up when the code is on a very hot path.</p>\n<h3 id=\"deabstraction\">Deabstraction</h3>\n<p>We as developers love our abstractions. They let us write clean, reusable, object-oriented code, but we don’t want to pay for every abstraction at run time. The runtime can often undo an abstraction when it proves the effects aren’t observable. It can look at a virtual call and determine which concrete method it’ll invoke, look at a heap allocation and recognize that the object never leaves the current stack frame, or look at an interface cast and reuse a type fact already established earlier in the method. This process is called “deabstraction.” .NET has improved steadily in this area for years, and that continues in .NET 11.</p>\n<p>Every time you write <code>interface</code> in C#, you’re creating a contract, a promise that any type implementing that interface can be substituted for any other. That flexibility is enormously valuable because, for example, it’s what lets us write <code>IEnumerable&lt;T&gt;</code> and have it work equally well over arrays, lists, other collections, LINQ, custom iterators, and so on. But the CPU doesn’t know anything about these contracts; it just knows how to execute instructions. Turning “call whatever method this interface reference points to” into actual machine instructions requires special machinery. Consider this example:</p>\n<pre><code>// dotnet run -c Release -f net11.0 --filter &quot;*&quot;\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Runtime.CompilerServices;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private Animal _animal = Environment.TickCount &gt;= 0 ? new Dog() : new Cat();\n    [Benchmark]\n    public int Speak() =&gt; _animal.Speak();\n    public abstract class Animal\n    {\n        public abstract int Speak();\n    }\n    private sealed class Dog : Animal\n    {\n        [MethodImpl(MethodImplOptions.NoInlining)]\n        public override int Speak() =&gt; 1;\n    }\n    private sealed class Cat : Animal\n    {\n        [MethodImpl(MethodImplOptions.NoInlining)]\n        public override int Speak() =&gt; 2;\n    }\n}</code></pre>\n<p>At compile time, all else equal, the JIT doesn’t know whether <code>_animal</code> is a <code>Dog</code> or a <code>Cat</code>. It generates code that loads the instance’s “method table pointer” (its object type handle), sometimes called a “vtable pointer”, stored at the beginning of every .NET object, indexes into the method table at the known slot for <code>Speak</code>, and calls the function pointer found there:</p>\n<pre><code>; x64\nmov     rcx, [rcx+8]   ; load _animal\nmov     rax, [rcx]     ; load method table\nmov     rax, [rax+40]  ; load vtable chunk\ncall    qword ptr [rax+20]</code></pre>\n<p>For this one call to <code>Speak</code>, we pay three dependent memory dereferences and an indirect call because the processor doesn’t know for certain in advance where the call is going (it might guess, or “speculatively execute”, but it has to be prepared for the possibility it was wrong), and because the call target is indirect, the JIT can’t inline the callee. Whatever <code>Speak</code> does, its code can’t be folded into the calling method.</p>\n<p>That’s a performance problem. Those indirections have overhead, but the bigger cost is the lost opportunity to inline. Inlining not only saves function call overhead, more importantly it opens the callee’s code up to the same optimizations that are operating on the caller, such as constant propagation, dead code elimination, bounds check elimination, further devirtualization, etc. That means a series of small virtual calls that each look innocent can, when devirtualized and inlined, collapse into a handful of instructions that would be unrecognizable and way cheaper when compared to the original source code. Without inlining, each callee is an opaque box; with it, the JIT can see through the layers.</p>\n<p>We as .NET developers constantly rely on the JIT’s sophisticated heuristics for inlining that weigh the IL size of the callee, the exact work the callee is performing, the call frequency of the method, the expected benefit from constant arguments, and dozens of other factors. For virtual calls, the JIT needs to know what the actual target of the call will be; it needs to “devirtualize”. In some cases, it can determine that statically, where it has exact-type knowledge. For example, if the JIT can prove that <code>animal</code> is always a <code>Dog</code>, whether because it was just allocated with <code>new Dog()</code>:</p>\n<pre><code>Animal animal = GetSomeAnimal();\nanimal.Speak();\n...\nstatic Animal GetSomeAnimal() =&gt; new Dog(); // inlineable</code></pre>\n<p>or because the variable’s type is a sealed class:</p>\n<pre><code>Dog animal = GetSomeAnimal();\nanimal.Speak();\n...\nsealed class Dog { ... } // impossible for `animal` to be anything other than a `Dog`</code></pre>\n<p>or with NativeAOT and whole-program compilation, if it sees that <code>Animal</code> is abstract and the only type in the whole application that derives from <code>Animal</code> is <code>Dog</code>:</p>\n<pre><code>Animal animal = GetSomeAnimal();\nanimal.Speak();\n...\nabstract class Animal { ... }\nclass Dog : Animal { ... } // no other such derived type</code></pre>\n<p>or other such validation, it can emit a call to <code>Dog.Speak()</code> directly, and the inliner can take its shot.</p>\n<p>But for other cases where it can’t prove this with static analysis, the JIT turns to profile-guided optimization (PGO). PGO sounds fancy, but it’s conceptually simple. With “tiered compilation”, when a method is first invoked, it can be compiled “just in time” with few-to-no optimizations (this is referred to as Tier 0). The JIT can include in this compilation additional probes (think “printf debugging”) that let it track a bunch of interesting information about the nature of the code, recording what actually happens when it runs: which branches are taken, what are the concrete types that show up at virtual call sites or cast attempts, and so on. If the method is invoked enough or loops enough times, the runtime can ask the JIT to produce a new optimized version (referred to as Tier 1). That compilation can then factor in all of the learnings gathered as part of that profiling.</p>\n<p>The JIT, of course, still needs to generate code that’s always correct. Even if a dynamic profile says <code>animal</code> was <code>Dog</code> 100% of the time, that doesn’t guarantee it’ll always be <code>Dog</code> in the future; it could be that the first 1000 calls passed in a <code>Dog</code> but the 1001st call is going to pass in <code>Dolphin</code>. How can the JIT incorporate this learning then? By emitting a run-time check. The <code>Dog</code> path can get a direct call, which may then be inlinable, and the other path keeps the original virtual call as the fallback. The speed comes from making the common case tiny, while correctness comes from leaving the uncommon case intact.</p>\n<pre><code>// Approximately what the JIT generates\nif (animal?.GetType() == typeof(Dog))\n{\n    ((Dog)animal).Speak();  // devirtualized, inlinable\n}\nelse\n{\n    animal.Speak(); // original virtual call, hopefully rare\n}</code></pre>\n<p>This “guess and verify” pattern, called “guarded devirtualization” (GDV), accounts for many of the biggest throughput wins in real workloads. It’s applicable not only to virtual dispatch but also to interface dispatch, which also happens to be a bit more expensive than virtual dispatch because a type can implement any number of interfaces and that means the interface slots don’t simply map to fixed vtable positions.</p>\n<p>Deabstraction can also make object creation more efficient when it reveals what kind of object is involved. In general, objects in .NET are allocated on the garbage collected heap, tracked by the garbage collector (GC), and collected when no longer reachable. Heap allocation is typically fast, often effectively just bumping a pointer. However, when there’s not enough space available to bump the pointer, it can get much more expensive, including needing to incur a garbage collection. Every allocated object also effectively incurs the amortized cost of all collections, as every allocated object eventually needs to be cleaned up.</p>\n<p>“Escape analysis” is the compiler technique that lets us ask whether this object ever “escapes” the current method. If an object reference to a newly allocated object provably doesn’t escape, then the JIT can more efficiently allocate it. It needn’t store it on the GC heap, because nothing could possibly need to reference that object again, so it can instead allocate the object on the stack, making both allocation and cleanup essentially free. Stack allocation is even faster than heap bump-pointer allocation; it’s just decrementing the stack pointer, which is typically already in a register. And more importantly it means zero GC impact, because the stack frame is freed atomically on function return.</p>\n<p>The JIT’s been progressively expanding escape analysis over the past several .NET releases, with .NET 9 and 10 seeing significant investments in stack-allocating delegates and closures, <code>Nullable&lt;T&gt;</code> temporaries, and small helper objects. The key theme is that every false positive escape, every time the JIT incorrectly concludes an object may escape when it really doesn’t, represents a heap allocation that could have been avoided, and we want to whittle away at that false positive list. In .NET 11, the JIT trims that list in several ways.</p>\n<p>We’ll start with nullable boxing. Consider this benchmark:</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[MemoryDiagnoser(false), HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private int? _nullableNull;\n    private int? _nullableValue = 42;\n    [Benchmark]\n    public object? BoxNullableNull() =&gt; (object?)_nullableNull;\n    [Benchmark]\n    public object? BoxNullableValue() =&gt; (object?)_nullableValue;\n    [Benchmark]\n    public string? FormatNullableInt() =&gt; Format(_nullableValue);\n    private static string? Format&lt;T&gt;(T value)\n    {\n        if (value is IFormattable formattable)\n            return formattable.ToString(null, null);\n        return null;\n    }\n}</code></pre>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Mean</th><th>Ratio</th><th>Allocated</th><th>Alloc Ratio</th></tr></thead><tbody><tr><td>BoxNullableNull</td><td>.NET 10.0</td><td>2.095 ns</td><td>1.00</td><td>–</td><td>–</td></tr><tr><td>BoxNullableNull</td><td>.NET 11.0</td><td>1.764 ns</td><td>0.84</td><td>–</td><td>–</td></tr><tr><td>BoxNullableValue</td><td>.NET 10.0</td><td>9.213 ns</td><td>1.00</td><td>24 B</td><td>1.00</td></tr><tr><td>BoxNullableValue</td><td>.NET 11.0</td><td>4.126 ns</td><td>0.45</td><td>24 B</td><td>1.00</td></tr><tr><td>FormatNullableInt</td><td>.NET 10.0</td><td>9.583 ns</td><td>1.00</td><td>24 B</td><td>1.00</td></tr><tr><td>FormatNullableInt</td><td>.NET 11.0</td><td>1.987 ns</td><td>0.21</td><td>–</td><td>0</td></tr></tbody></table></div>\n<p>dotnet/runtime#122167 expands nullable boxing inside the JIT, exposing the temporary box to escape analysis; previously, a runtime helper hid it. For a <code>null</code> input, there’s no allocation on either version, because nothing gets boxed. And on both versions, <code>BoxNullableValue</code> returns the boxed object, meaning the object escapes, so the 24-byte allocation remains. However, for <code>FormatNullableInt</code>, the JIT in .NET 11 can now see that the temporary 24-byte box doesn’t escape and eliminates that heap allocation entirely.</p>\n<p>Escape analysis improved further for enumerators, through a mechanism called Conditional Escape Analysis (CEA). Support for CEA was introduced in .NET 10, but .NET 11 extends the set of patterns that this analysis can safely recognize. The existing escape analysis asks whether a reference created by an allocation can flow somewhere the JIT can no longer track, such as an unknown call. If it can, the object must remain on the heap. That analysis is necessarily conservative and largely flow-insensitive: if an object might be passed to an interface call on any path, it doesn’t try to prove that the path containing that call is mutually exclusive with the path containing the allocation.</p>\n<p>Unfortunately, that’s exactly what GDV produces when it optimizes a <code>foreach</code> over an <code>IEnumerable&lt;T&gt;</code>. As noted earlier, GDV turns an interface call into a type check with two branches: a fast branch for the likely collection type and a fallback branch containing the original interface call. Devirtualization and inlining along the fast branch will often reveal an enumerator allocation for the collection type, while later enumerator guards retain fallback calls such as <code>IEnumerator&lt;T&gt;.MoveNext</code>. The existing analysis sees those calls and concludes that the locally allocated enumerator might escape. CEA instead records the relationship between the fast-path allocation and the enumerator local tested by the later guards. If every apparent escape occurs only behind a failed type check, the JIT can clone the region into a hot version where those checks are known to succeed. In that clone, the object can’t reach the fallback calls, so it can be stack-allocated and often promoted into separate scalar locals. The original region remains as the general slow path.</p>\n<p>One case .NET 10 didn’t handle, though, was a <code>GetEnumerator()</code> implementation that returns the result of another <code>GetEnumerator()</code> call. A collection expression converted to <code>IEnumerable&lt;int&gt;</code>, for example, uses a compiler-generated read-only-array wrapper with exactly this structure: the wrapper’s <code>GetEnumerator()</code> delegates to the underlying array’s <code>GetEnumerator</code>. With dotnet/runtime#122946, the JIT in .NET 11 handles this “chaining”:</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[MemoryDiagnoser(false), HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private static readonly IEnumerable&lt;int&gt; s_readOnlyStatic = [1, 2, 3, 4, 5];\n    private readonly IEnumerable&lt;int&gt; _readOnlyInstance = [1, 2, 3, 4, 5];\n    [Benchmark]\n    public int ReadOnlyStatic()\n    {\n        int sum = 0;\n        foreach (int item in s_readOnlyStatic) sum += item;\n        return sum;\n    }\n    [Benchmark]\n    public int ReadOnlyInstance()\n    {\n        int sum = 0;\n        foreach (int item in _readOnlyInstance) sum += item;\n        return sum;\n    }\n}</code></pre>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Mean</th><th>Ratio</th><th>Allocated</th><th>Alloc Ratio</th></tr></thead><tbody><tr><td>ReadOnlyStatic</td><td>.NET 10.0</td><td>2.665 ns</td><td>1.00</td><td>–</td><td>–</td></tr><tr><td>ReadOnlyStatic</td><td>.NET 11.0</td><td>2.666 ns</td><td>1.00</td><td>–</td><td>–</td></tr><tr><td>ReadOnlyInstance</td><td>.NET 10.0</td><td>13.874 ns</td><td>1.00</td><td>32 B</td><td>1.00</td></tr><tr><td>ReadOnlyInstance</td><td>.NET 11.0</td><td>2.674 ns</td><td>0.19</td><td>–</td><td>0</td></tr></tbody></table></div>\n<p><code>ReadOnlyStatic</code>, whose <code>static readonly</code> field the JIT can effectively treat as a constant, was already optimized in .NET 10. In .NET 11, the instance-field case also loses its 32-byte enumerator allocation and converges on the same throughput.</p>\n<p>dotnet/runtime#121918 from @MichalPetryka fixes another way an address could unnecessarily make an object appear to escape. The IL <code>constrained.</code> prefix lets one generic <code>callvirt</code> sequence work for both value types and reference types: it can avoid boxing a value type, while for a reference type it dereferences the receiver and performs normal virtual dispatch. <code>ObjectEqualityComparer&lt;T&gt;.Equals</code>, used in the following benchmark by <code>EqualityComparer&lt;T&gt;.Default</code>, contains such a call to <code>value.Equals(other)</code>. The receiver was represented as an indirect read through the address of a local. Merely taking that address marked the local as exposed, preventing the newly allocated <code>Value</code> from being considered for stack allocation. The receiver is now represented as a direct value load instead, and the 24-byte heap allocation disappears.</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Collections.Generic;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[MemoryDiagnoser(false), HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private static readonly Value s_other = new(42);\n    [Benchmark]\n    public bool Equals() =&gt; EqualityComparer&lt;Value&gt;.Default.Equals(new Value(42), s_other);\n    private sealed class Value(int value)\n    {\n        private readonly int _value = value;\n        public override bool Equals(object? obj) =&gt; obj is Value other &amp;&amp; _value == other._value;\n        public override int GetHashCode() =&gt; _value;\n    }\n}</code></pre>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Mean</th><th>Ratio</th><th>Allocated</th><th>Alloc Ratio</th></tr></thead><tbody><tr><td>Equals</td><td>.NET 10.0</td><td>3.874 ns</td><td>1.00</td><td>24 B</td><td>1.00</td></tr><tr><td>Equals</td><td>.NET 11.0</td><td>1.786 ns</td><td>0.46</td><td>–</td><td>0</td></tr></tbody></table></div>\n<p>While CEA can move a non-escaping object off the GC heap, sometimes the JIT can go further and prove an allocation need not exist at all. Generic code provides a common source of such opportunities through boxing. For example, the <code>ArgumentNullException.ThrowIfNull</code> method accepts an <code>object value</code>. That means when you have a method like this:</p>\n<pre><code>static void Test&lt;T&gt;(T value)\n{\n    ArgumentNullException.ThrowIfNull(value);\n    ...\n}</code></pre>\n<p>when <code>T</code> is constrained to a non-nullable struct, boxing is incurred, in order to pass <code>value</code> as <code>object</code>. <code>ThrowIfNull</code> here is a nop if <code>value</code> is non-<code>null</code> (since the method is simply <code>if (value is null) Throw();</code>), and previous releases successfully optimized away that boxing in optimized code. However, in Tier 0, that optimization wasn’t applied, and <code>ThrowIfNull</code> would end up allocating. While this wouldn’t negatively impact steady-state throughput, it would lead to annoying noise in profiling, as well as additional overhead during startup, where such use wasn’t yet promoted out of Tier 0. In .NET 11, dotnet/runtime#129392 adds support for this in Tier 0 as well.</p>\n<p>On the virtual-dispatch side, multiple PRs contribute to improving generic virtual methods (GVMs). dotnet/runtime#120866 from @hez2010 stops eagerly spilling <code>ldvirtftn</code> call targets into a temporary, and lets generic virtual target resolution move ahead of argument setup when legal. dotnet/runtime#122023 from @hez2010 then enables the JIT to devirtualize non-shared GVMs, carrying the generic context needed to turn the indirect dispatch into a direct, and potentially inlineable, call. And dotnet/runtime#128702 from @hez2010 extends that support to shared GVMs and default interface implementations that require an instantiating stub. These optimizations can increase total code size when the newly direct calls are inlined, but that’s generally the desired trade: more of the actual work becomes visible to the optimizer. Consider the following benchmark:</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Runtime.CompilerServices;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[MemoryDiagnoser(false), HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    [Benchmark]\n    public int NonShared() =&gt; ((IProcessor)new Processor()).SizeOf(42);\n    [Benchmark]\n    public int Shared() =&gt; ((IProcessor)new Processor()).SizeOf(&quot;hello&quot;);\n    private interface IProcessor\n    {\n        int SizeOf&lt;T&gt;(T value);\n    }\n    private sealed class Processor : IProcessor\n    {\n        public int SizeOf&lt;T&gt;(T value) =&gt; Unsafe.SizeOf&lt;T&gt;();\n    }\n}</code></pre>\n<p>Casting a freshly allocated <code>Processor</code> to <code>IProcessor</code> incurs an interface generic virtual call in the IL, but the JIT is now able to see the receiver’s exact type, even in the shared <code>string</code> case, such that .NET 11 devirtualizes and inlines both calls. That in turn exposes <code>Unsafe.SizeOf&lt;T&gt;()</code> as a constant and proves that the short-lived <code>Processor</code> doesn’t need to be allocated at all.</p>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Mean</th><th>Ratio</th><th>Allocated</th><th>Alloc Ratio</th></tr></thead><tbody><tr><td>NonShared</td><td>.NET 10.0</td><td>6.678 ns</td><td>1.00</td><td>24 B</td><td>1.00</td></tr><tr><td>NonShared</td><td>.NET 11.0</td><td>1.764 ns</td><td>0.26</td><td>–</td><td>0</td></tr><tr><td>Shared</td><td>.NET 10.0</td><td>7.166 ns</td><td>1.00</td><td>24 B</td><td>1.00</td></tr><tr><td>Shared</td><td>.NET 11.0</td><td>1.764 ns</td><td>0.25</td><td>–</td><td>0</td></tr></tbody></table></div>\n<p>Building on that, dotnet/runtime#123183 from @hez2010 enables ReadyToRun compilation to resolve and devirtualize more non-shared generic virtual calls that would otherwise remain indirect, and dotnet/runtime#130202 from @hez2010 extends that support to NativeAOT. NativeAOT represents some generic virtual targets as “fat pointers” (pointers that are more than just an address, typically an address and associated metadata, and that in this case carry both a code address and generic context); by deferring that transformation until after exact-type devirtualization has had a chance to run, the JIT can turn an interface call site with a single known target to a non-shared GVM into a direct call that may then be inlined.</p>\n<p>Type information also needs to survive the transformations the JIT performs internally. If the JIT spills a reference expression into a temporary while restructuring a tree, losing the expression’s exact class information can turn a call that was devirtualizable back into an opaque virtual call. That’s what happens here in .NET 10: <code>Value</code> gets boxed and <code>SetValue</code> is invoked through <code>IValue</code>. dotnet/runtime#128485 from @hez2010 preserves the class handle and exactness on the temporary. With that information still available, .NET 11 devirtualizes and inlines the call, eliminating the box and its 24-byte allocation.</p>\n<p>Separately, dotnet/runtime#127433 relaxes the inliner’s budget heuristics for callees on <code>[Intrinsic]</code> types like <code>Span</code> and <code>Vector</code>. These types intentionally expose many small, composable methods that serve as gateways to JIT-recognized operations. If a wrapper remains as a call, the caller pays the call overhead and optimizations around it see an opaque boundary. If it inlines, the importer can replace its body with an intrinsic node and optimize that node together with the surrounding indexing, bounds checks, and vector operations. Giving such wrappers more favorable budgeting therefore keeps more of them inlineable and exposes more of the actual operation to the rest of the optimizer.</p>\n<p>One of the core abstraction-enabling mechanisms in .NET is delegates: they let us pass around objects representing functions to be invoked, carrying with them associated required state. Deabstraction enables avoiding paying for the overheads associated with delegates in some cases. For the rest, we still want those delegates to be as cheap as possible. dotnet/runtime#99200 from @MichalPetryka simplifies CoreCLR’s delegate representation, removing one pointer-sized field from every delegate object. That saves 8 bytes per delegate in a 64-bit CoreCLR process. dotnet/runtime#129304 from @MichalPetryka improves Native AOT’s delegate layout separately by reordering its existing four fields so related values are adjacent. The updated layouts also give equality and hash-code operations more direct access to the method identity they need.</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[MemoryDiagnoser(false), HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private static readonly Target s_target = new();\n    private static readonly Func&lt;int&gt; s_first = s_target.GetValue;\n    private static readonly Func&lt;int&gt; s_second = s_target.GetValue;\n    [Benchmark]\n    public Func&lt;int&gt; ClosedInstance() =&gt; s_target.GetValue;\n    [Benchmark]\n    public bool DelegateEquals() =&gt; s_first.Equals(s_second);\n    [Benchmark]\n    public int DelegateGetHashCode() =&gt; s_first.GetHashCode();\n    private sealed class Target\n    {\n        public int GetValue() =&gt; 42;\n    }\n}</code></pre>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Mean</th><th>Ratio</th><th>Allocated</th><th>Alloc Ratio</th></tr></thead><tbody><tr><td>ClosedInstance</td><td>.NET 10.0</td><td>7.395 ns</td><td>1.00</td><td>64 B</td><td>1.00</td></tr><tr><td>ClosedInstance</td><td>.NET 11.0</td><td>6.844 ns</td><td>0.93</td><td>56 B</td><td>0.88</td></tr><tr><td>DelegateEquals</td><td>.NET 10.0</td><td>3.254 ns</td><td>1.00</td><td>–</td><td>–</td></tr><tr><td>DelegateEquals</td><td>.NET 11.0</td><td>2.215 ns</td><td>0.68</td><td>–</td><td>–</td></tr><tr><td>DelegateGetHashCode</td><td>.NET 10.0</td><td>5.623 ns</td><td>1.00</td><td>–</td><td>–</td></tr><tr><td>DelegateGetHashCode</td><td>.NET 11.0</td><td>3.741 ns</td><td>0.67</td><td>–</td><td>–</td></tr></tbody></table></div>\n<p>dotnet/runtime#129410 from @MichalPetryka follows up on the CoreCLR layout by placing the target object and method pointer next to each other. Those are commonly consumed together during invocation, and the adjacency enables paired loads on architectures such as Arm64.</p>\n<h3 id=\"runtime-async\">Runtime Async</h3>\n<p>For more than a decade, <code>async</code> and <code>await</code> have let us write asynchronous code that looks remarkably similar to synchronous code: we can put a <code>try</code>/<code>catch</code> around an <code>await</code>, use local variables on either side of it, return a value and generally reason about the method in source order. When execution reaches an <code>await</code> for something that isn’t yet complete, however, the method can’t simply leave its current stack frame in place and wait for the operation to finish. The thread needs to be freed up to do other work, while the work after the <code>await</code>, including whatever local state it will need later, must survive somewhere. In C#, the compiler has traditionally been responsible for transforming the method into a representation that enables that continuation.</p>\n<p>I went into the history and mechanics of that transformation in How async/await really works. The very short version is that the compiler traditionally replaces an <code>async</code> method with a small entry method and a generated state machine whose <code>MoveNext</code> method contains the transformed user code. Parameters, locals that need to survive an incomplete await, spilled expression values, awaiters, the current state number, and a method builder all become fields on a heap-allocated object. The generated <code>MoveNext</code> method runs the user’s code until an awaiter reports that it isn’t yet complete. It stores enough information to know where and with what values to resume, registers <code>MoveNext</code> as the continuation, and returns. When the operation completes, <code>MoveNext</code> is invoked again, jumps to the right location based on the saved state number (think <code>goto</code> and a label), retrieves the result from a value-producing awaiter, and continues. If every awaiter is already complete, <code>MoveNext</code> can run all the way through synchronously. When the method completes or throws, the builder publishes the result, cancellation, or exception through the returned <code>Task</code>, <code>Task&lt;T&gt;</code>, <code>ValueTask</code>, or <code>ValueTask&lt;T&gt;</code> (or, in the rare case, a custom task-like type).</p>\n<p>For example, consider this tiny method:</p>\n<pre><code>static async Task&lt;int&gt; ReadLengthAsync(Stream stream, CancellationToken cancellationToken)\n{\n    var buffer = new byte[4096];\n    int bytesRead = await stream.ReadAsync(buffer, 0, buffer.Length, cancellationToken);\n    return bytesRead;\n}</code></pre>\n<p>While the code that gets generated for this changes over time and differs between debug and release builds, the lowering by the C# compiler has looked something like this:</p>\n<pre><code>[AsyncStateMachine(typeof(&lt;ReadLengthAsync&gt;d__0))]\nstatic Task&lt;int&gt; ReadLengthAsync(Stream stream, CancellationToken cancellationToken)\n{\n    &lt;ReadLengthAsync&gt;d__0 stateMachine = default;\n    stateMachine.builder = AsyncTaskMethodBuilder&lt;int&gt;.Create();\n    stateMachine.state = -1;\n    stateMachine.stream = stream;\n    stateMachine.cancellationToken = cancellationToken;\n    stateMachine.builder.Start(ref stateMachine);\n    return stateMachine.builder.Task;\n}\nstruct &lt;ReadLengthAsync&gt;d__0 : IAsyncStateMachine\n{\n    public int state;\n    public AsyncTaskMethodBuilder&lt;int&gt; builder;\n    public Stream stream;\n    public CancellationToken cancellationToken;\n    private TaskAwaiter&lt;int&gt; awaiter;\n    public void MoveNext()\n    {\n        int result;\n        try\n        {\n            TaskAwaiter&lt;int&gt; localAwaiter;\n            if (state != 0)\n            {\n                byte[] buffer = new byte[4096];\n                localAwaiter = stream.ReadAsync(buffer, 0, buffer.Length, cancellationToken).GetAwaiter();\n                if (!localAwaiter.IsCompleted)\n                {\n                    state = 0;\n                    awaiter = localAwaiter;\n                    builder.AwaitUnsafeOnCompleted(ref localAwaiter, ref this);\n                    return;\n                }\n            }\n            else\n            {\n                localAwaiter = awaiter;\n                awaiter = default;\n                state = -1;\n            }\n            result = localAwaiter.GetResult();\n        }\n        catch (Exception e)\n        {\n            state = -2;\n            builder.SetException(e);\n            return;\n        }\n        state = -2;\n        builder.SetResult(result);\n    }\n}</code></pre>\n<p>That’s quite a lot of generated code for three lines of C#. The compiler has to make decisions before the program runs about the state-machine layout, which values might need to survive, how many awaiter fields are required, and how all the suspension points fit into one <code>MoveNext</code> dispatch. The runtime and JIT have optimized the resulting pattern heavily over the years, including combining the task, state machine, continuation, and <code>ExecutionContext</code> into a single allocation, but by the time the JIT sees the IL, the transformation has already happened, leaving it with a very complicated system to try to optimize.</p>\n<p>.NET 11 introduces a new way to split that responsibility, a reimplementation of the <code>async</code>/<code>await</code> infrastructure referred to as “runtime async”. Rather than the C# compiler being responsible for the transformation, the JIT is. The C# compiler emits a much smaller suspension-aware IL contract for each eligible <code>async</code> method and marks the method as <code>async</code> in metadata. The runtime and JIT then do the work that depends on runtime knowledge: creating the externally visible <code>Task</code> or <code>ValueTask</code>, recognizing direct async calls, deciding which values are actually alive at each suspension point, laying out continuation objects, and generating the control flow that suspends and resumes the method. Effectively, the transformation moves from C# to the runtime, where more information is available to optimize it.</p>\n<p>The programming model hasn’t changed. This is still C# <code>async</code>/<code>await</code>; <code>await</code> still obeys the awaiter pattern, exceptions and cancellation still surface through the returned task-like object, <code>ConfigureAwait</code> still has its usual meaning, synchronous completion is still synchronous completion, and on and on. An explicit goal for the feature has been 100% behavioral compatibility: whether an <code>async</code> method is lowered by the language compiler or by the runtime is an implementation detail, and any observable semantic difference is a bug.</p>\n<p>In .NET 11, application code opts in with a compiler feature switch:</p>\n<pre><code>&lt;Project Sdk=&quot;Microsoft.NET.Sdk&quot;&gt;\n  &lt;PropertyGroup&gt;\n    &lt;TargetFramework&gt;net11.0&lt;/TargetFramework&gt;\n    &lt;Features&gt;$(Features);runtime-async=on&lt;/Features&gt;\n  &lt;/PropertyGroup&gt;\n&lt;/Project&gt;</code></pre>\n<p>Note that there’s no new C# syntax involved, so <code>LangVersion=preview</code> isn’t required, nor is <code>EnablePreviewFeatures</code>. While this is opt-in at the application layer, most of the in-box shared framework is already built this way for .NET 11. The <code>async</code>/<code>await</code> performance goal for .NET 11 is parity with .NET 10, and in general runtime async is already as good as or better than the older implementation in many important paths. It isn’t yet fully optimized, though, and there are known cases where it still produces less efficient code. I’d encourage you to experiment in .NET 11 with opting-in your applications and services; just make sure to measure. My hope is that it’ll be on by default starting in .NET 12.</p>\n<p>Moving the transformation from the C# compiler to the runtime has the added benefit of reducing binary size. As noted, the traditional lowering emits an entry method, a generated state-machine type, fields for captured state, and a <code>MoveNext</code> body, for every async method. Runtime async leaves a much smaller method body for the runtime to transform. The following tiny app contains ten <code>Task&lt;int&gt;</code>-returning async methods, each awaiting the next, and compiles the same source once with compiler lowering and once with runtime async:</p>\n<pre><code>&lt;Project Sdk=&quot;Microsoft.NET.Sdk&quot;&gt;\n  &lt;PropertyGroup&gt;\n    &lt;OutputType&gt;Exe&lt;/OutputType&gt;\n    &lt;TargetFramework&gt;net11.0&lt;/TargetFramework&gt;\n    &lt;AssemblyName&gt;SizeProbe&lt;/AssemblyName&gt;\n    &lt;ImplicitUsings&gt;enable&lt;/ImplicitUsings&gt;\n    &lt;Nullable&gt;enable&lt;/Nullable&gt;\n    &lt;Features Condition=&quot;&#39;$(RuntimeAsync)&#39; == &#39;true&#39;&quot;&gt;$(Features);runtime-async=on&lt;/Features&gt;\n  &lt;/PropertyGroup&gt;\n&lt;/Project&gt;</code></pre>\n<pre><code>// dotnet build -c Release -p:RuntimeAsync=false -o classic --no-incremental; dotnet build -c Release -p:RuntimeAsync=true -o runtime --no-incremental; Get-Item .\\classic\\SizeProbe.dll, .\\runtime\\SizeProbe.dll | Select-Object Directory, Length\nConsole.WriteLine(await Benchmarks.Layer0());\npublic class Benchmarks\n{\n    public static async Task&lt;int&gt; Layer0() =&gt; await Layer1();\n    private static async Task&lt;int&gt; Layer1() =&gt; await Layer2();\n    private static async Task&lt;int&gt; Layer2() =&gt; await Layer3();\n    private static async Task&lt;int&gt; Layer3() =&gt; await Layer4();\n    private static async Task&lt;int&gt; Layer4() =&gt; await Layer5();\n    private static async Task&lt;int&gt; Layer5() =&gt; await Layer6();\n    private static async Task&lt;int&gt; Layer6() =&gt; await Layer7();\n    private static async Task&lt;int&gt; Layer7() =&gt; await Layer8();\n    private static async Task&lt;int&gt; Layer8() =&gt; await Layer9();\n    private static async Task&lt;int&gt; Layer9()\n    {\n        await Task.Yield();\n        return 42;\n    }\n}</code></pre>\n<div class=\"table-wrap\"><table><thead><tr><th>Lowering</th><th>SizeProbe.dll</th><th>Ratio</th></tr></thead><tbody><tr><td>Compiler</td><td>10,752 bytes</td><td>1.00</td></tr><tr><td>Runtime async</td><td>5,632 bytes</td><td>0.52</td></tr></tbody></table></div>\n<p>For a method such as:</p>\n<p><code>static async Task&lt;int&gt; CallerAsync() =&gt; await CalleeAsync();</code>\nwith runtime async enabled, the C# compiler generates IL like the following:</p>\n<pre><code>; MSIL\n.method private hidebysig static\n    class System.Threading.Tasks.Task`1&lt;int32&gt; CallerAsync() cil managed async\n{\n    call class System.Threading.Tasks.Task`1&lt;int32&gt; CalleeAsync()\n    call int32 System.Runtime.CompilerServices.AsyncHelpers::Await&lt;int32&gt;(\n        class System.Threading.Tasks.Task`1&lt;int32&gt;)\n    ret\n}</code></pre>\n<p>There is no generated <code>&lt;CallerAsync&gt;d__0</code> type, no <code>IAsyncStateMachine</code>, no <code>MoveNext</code>, no <code>AsyncTaskMethodBuilder&lt;int&gt;</code>, and no <code>AsyncStateMachineAttribute</code>. Previously, <code>async</code> on a C# method evaporated at compile time. Now, the method has a new <code>MethodImpl</code> <code>async</code> bit, represented in IL assembly syntax by that <code>async</code> modifier, and the body calls helpers in <code>System.Runtime.CompilerServices.AsyncHelpers</code>.</p>\n<p>At first glance the <code>ret</code> looks impossible because the declared signature returns <code>Task&lt;int&gt;</code> while the value on the IL evaluation stack is an <code>int</code>. This clearly isn’t a normal calling convention. The VM can give a <code>Task</code>-returning method two related identities, or MethodDescs, where one has the normal signature the rest of managed code sees, <code>Task&lt;int&gt; CallerAsync()</code>. The other is the AsyncCall variant, which effectively returns <code>int</code> and has an implicit channel for a continuation. Both refer to the same logical method and metadata token, but they have different calling conventions and different jobs. If regular managed code invokes <code>CallerAsync</code>, the VM-generated outer thunk preserves the public contract and returns a <code>Task&lt;int&gt;</code>. If another runtime async method directly awaits it, the JIT can instead call the AsyncCall variant and receive the result directly when the call completes synchronously, or a continuation when it suspends. In other words, it can hand back the <code>T</code> directly and avoid allocating a <code>Task&lt;T&gt;</code>.</p>\n<p>That pairing works in both directions. For a method compiled with runtime async, the AsyncCall variant owns the generated (newly compact) IL while the public <code>Task</code>-returning entry point is an adapter thunk; for a traditionally compiled method, the public method owns its usual IL while the VM can create an AsyncCall adapter around it. That means runtime async code remains able to await existing libraries and code compiled by older compilers, a critical capability for our goal of 100% compat. The largest wins naturally appear as more of an async call chain is compiled with runtime async.</p>\n<p>This is where the JIT gets an opportunity that simply didn’t exist when every boundary was already expressed as a task and a generated state machine. Suppose <code>A</code> awaits <code>B</code>, which awaits <code>C</code>:</p>\n<pre><code>static async Task&lt;int&gt; A(bool yield) =&gt; await B(yield);\nstatic async Task&lt;int&gt; B(bool yield) =&gt; await C(yield);\nstatic async Task&lt;int&gt; C(bool yield)\n{\n    if (yield)\n        await Task.Yield();\n    return 42;\n}</code></pre>\n<p>Traditionally, each method has its own compiler-generated state machine and its own task-like result. <code>C</code> suspends and eventually completes its task, which wakes <code>B</code>‘s state machine; <code>B</code> then completes its task, which wakes <code>A</code>‘s state machine; and <code>A</code> completes the root task observed by the caller. There has been an enormous amount of work done over the years to reduce the costs of those objects and transitions.</p>\n<p>With runtime async, the importer recognizes the adjacent pattern of “call a Task-returning method, then await that task.” In the simple case it can call the callee’s AsyncCall variant instead. When <code>yield</code> is false and <code>C</code> completes synchronously, the <code>int</code> flows back through <code>B</code> and <code>A</code> as a plain value, and only the outermost boundary needs to turn it into the <code>Task&lt;int&gt;</code> promised to the original caller. When <code>yield</code> is true and <code>C</code> suspends, the runtime links continuation state for the chain and eventually resumes it without requiring an intermediate <code>Task&lt;int&gt;</code> at every directly fused edge. The <code>Task</code> contract hasn’t vanished, it just moved to the place where a <code>Task</code> is actually needed.</p>\n<p>Runtime async doesn’t make every asynchronous operation allocation-free, though. Rather, it gives the JIT enough information to avoid materializing some task objects that existed only to carry a result from one async method directly into the next. If a consumer stores the task in a collection, manually hooks up a continuation, or otherwise observes the task as an object, that object is still needed. The optimization is about not paying for boundaries that aren’t observably boundaries.</p>\n<p>The impact is already visible with just two layers:</p>\n<pre><code>// dotnet run -c Release -f net11.0 --filter &quot;*&quot;\n// The project also needs the `runtime-async=on` feature switch set.\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Configs;\nusing BenchmarkDotNet.Running;\nusing System.Runtime.CompilerServices;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[MemoryDiagnoser(false)]\n[GroupBenchmarksBy(BenchmarkLogicalGroupRule.ByCategory)]\n[HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private static readonly Task&lt;int&gt; s_completed = Task.FromResult(42);\n    [Benchmark(Baseline = true), BenchmarkCategory(&quot;Completed&quot;)]\n    public Task&lt;int&gt; ClassicCompleted() =&gt; ClassicCompletedOuter();\n    [Benchmark, BenchmarkCategory(&quot;Completed&quot;)]\n    public Task&lt;int&gt; RuntimeCompleted() =&gt; RuntimeCompletedOuter();\n    [Benchmark(Baseline = true), BenchmarkCategory(&quot;Yielding&quot;)]\n    public Task&lt;int&gt; ClassicYielding() =&gt; ClassicYieldingOuter();\n    [Benchmark, BenchmarkCategory(&quot;Yielding&quot;)]\n    public Task&lt;int&gt; RuntimeYielding() =&gt; RuntimeYieldingOuter();\n    [RuntimeAsyncMethodGeneration(false)]\n    private static async Task&lt;int&gt; ClassicCompletedOuter() =&gt; await ClassicCompletedInner();\n    [RuntimeAsyncMethodGeneration(false)]\n    private static async Task&lt;int&gt; ClassicCompletedInner() =&gt; await s_completed;\n    private static async Task&lt;int&gt; RuntimeCompletedOuter() =&gt; await RuntimeCompletedInner();\n    private static async Task&lt;int&gt; RuntimeCompletedInner() =&gt; await s_completed;\n    [RuntimeAsyncMethodGeneration(false)]\n    private static async Task&lt;int&gt; ClassicYieldingOuter() =&gt; await ClassicYieldingInner();\n    [RuntimeAsyncMethodGeneration(false)]\n    private static async Task&lt;int&gt; ClassicYieldingInner()\n    {\n        await Task.Yield();\n        return 42;\n    }\n    private static async Task&lt;int&gt; RuntimeYieldingOuter() =&gt; await RuntimeYieldingInner();\n    private static async Task&lt;int&gt; RuntimeYieldingInner()\n    {\n        await Task.Yield();\n        return 42;\n    }\n}\nnamespace System.Runtime.CompilerServices\n{\n    [AttributeUsage(AttributeTargets.Method)]\n    internal sealed class RuntimeAsyncMethodGenerationAttribute(bool runtimeAsync) : Attribute\n    {\n        public bool RuntimeAsync =&gt; runtimeAsync;\n    }\n}</code></pre>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Mean</th><th>Ratio</th><th>Allocated</th><th>Alloc Ratio</th></tr></thead><tbody><tr><td>ClassicCompleted</td><td>21.221 ns</td><td>1.00</td><td>144 B</td><td>1.00</td></tr><tr><td>RuntimeCompleted</td><td>6.151 ns</td><td>0.29</td><td>0 B</td><td>0.00</td></tr><tr><td>ClassicYielding</td><td>254.139 ns</td><td>1.00</td><td>248 B</td><td>1.00</td></tr><tr><td>RuntimeYielding</td><td>116.927 ns</td><td>0.46</td><td>168 B</td><td>0.68</td></tr></tbody></table></div>\n<p>The synchronously completing chain is more than 3x faster and avoids both intermediate task allocations. Even after a real suspension, the same two-layer chain takes less than half the time and allocates 80 fewer bytes.</p>\n<p>Exception handling amplifies the difference. Again consider an async method <code>A</code> calling an async method <code>B</code> calling an async method <code>C</code>. The transformation generated by the C# compiler of each method results in a <code>try</code>/<code>catch</code> block around the whole body of the <code>MoveNext</code> method so that any unhandled exception can be stored into the returned <code>Task</code>. Let’s say code in <code>C</code> throws an unhandled exception. That’s then caught by this manufactured <code>catch</code> block and stored into the <code>Task</code> returned to <code>B</code>. The awaiter in <code>B</code> then retrieves that exception from the <code>Task</code> object and throws it. It’s then caught by <code>B</code>‘s generated catch and stored into its <code>Task</code>. And so on. An exception crossing ten such async helpers can therefore be thrown, caught, and stored ten times even though none of the source methods has an explicit handler. That is super expensive. But runtime async doesn’t need to re-enter a pass-through frame with no handler. On the synchronous path the exception unwinds through the fused calls normally, and after a real suspension, one dispatch-loop catch walks past continuation records that have no handler and faults the observable root task once.</p>\n<p>The following benchmark measures both a fully synchronous throw and an exception after one real <code>Task.Yield</code> suspension. It uses a compiler-recognized per-method escape hatch (<code>RuntimeAsyncMethodGeneration</code>) so that the classic and runtime async methods run in the same process on the same .NET 11 runtime and differ only in how the compiler lowers them. (Note that this attribute is experimental and isn’t a public API exposed from the core libraries; as with other attributes known to the C# compiler, it recognizes them by name and signature.)</p>\n<pre><code>// dotnet run -c Release -f net11.0 --filter &quot;*&quot;\n// The project also needs the `runtime-async=on` feature switch set.\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Runtime.CompilerServices;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[MemoryDiagnoser(false), HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    [Params(1, 10, 30)]\n    public int Depth;\n    [Params(false, true)]\n    public bool Yield;\n    [Benchmark(Baseline = true)]\n    public int Classic() =&gt; Invoke(ClassicThrowAsync(Depth));\n    [Benchmark]\n    public int Runtime() =&gt; Invoke(RuntimeThrowAsync(Depth));\n    private static int Invoke(Task&lt;int&gt; task)\n    {\n        try\n        {\n            return task.GetAwaiter().GetResult();\n        }\n        catch (InvalidOperationException)\n        {\n            return -1;\n        }\n    }\n    [RuntimeAsyncMethodGeneration(false)]\n    private async Task&lt;int&gt; ClassicThrowAsync(int depth)\n    {\n        if (depth == 0)\n        {\n            if (Yield) await Task.Yield();\n            throw new InvalidOperationException(&quot;uh oh&quot;);\n        }\n        return await ClassicThrowAsync(depth - 1);\n    }\n    private async Task&lt;int&gt; RuntimeThrowAsync(int depth)\n    {\n        if (depth == 0)\n        {\n            if (Yield) await Task.Yield();\n            throw new InvalidOperationException(&quot;uh oh&quot;);\n        }\n        return await RuntimeThrowAsync(depth - 1);\n    }\n}\nnamespace System.Runtime.CompilerServices\n{\n    [AttributeUsage(AttributeTargets.Method)]\n    internal sealed class RuntimeAsyncMethodGenerationAttribute(bool runtimeAsync) : Attribute\n    {\n        public bool RuntimeAsync =&gt; runtimeAsync;\n    }\n}</code></pre>\n<div class=\"table-wrap\"><table><thead><tr><th>Depth</th><th>Yield</th><th>Method</th><th>Mean</th><th>Ratio</th><th>Allocated</th><th>Alloc Ratio</th></tr></thead><tbody><tr><td>1</td><td>False</td><td>Classic</td><td>4.727 μs</td><td>1.00</td><td>1.6 KB</td><td>1.00</td></tr><tr><td>1</td><td>False</td><td>Runtime</td><td>3.558 μs</td><td>0.75</td><td>1.16 KB</td><td>0.72</td></tr><tr><td>1</td><td>True</td><td>Classic</td><td>6.308 μs</td><td>1.00</td><td>1.68 KB</td><td>1.00</td></tr><tr><td>1</td><td>True</td><td>Runtime</td><td>8.211 μs</td><td>1.30</td><td>1.42 KB</td><td>0.85</td></tr><tr><td>10</td><td>False</td><td>Classic</td><td>19.469 μs</td><td>1.00</td><td>15.13 KB</td><td>1.00</td></tr><tr><td>10</td><td>False</td><td>Runtime</td><td>5.885 μs</td><td>0.30</td><td>2.13 KB</td><td>0.14</td></tr><tr><td>10</td><td>True</td><td>Classic</td><td>24.923 μs</td><td>1.00</td><td>15.63 KB</td><td>1.00</td></tr><tr><td>10</td><td>True</td><td>Runtime</td><td>6.122 μs</td><td>0.25</td><td>2.88 KB</td><td>0.18</td></tr><tr><td>30</td><td>False</td><td>Classic</td><td>51.302 μs</td><td>1.00</td><td>84.2 KB</td><td>1.00</td></tr><tr><td>30</td><td>False</td><td>Runtime</td><td>10.721 μs</td><td>0.21</td><td>5.71 KB</td><td>0.07</td></tr><tr><td>30</td><td>True</td><td>Classic</td><td>65.974 μs</td><td>1.00</td><td>85.53 KB</td><td>1.00</td></tr><tr><td>30</td><td>True</td><td>Runtime</td><td>11.254 μs</td><td>0.17</td><td>7.72 KB</td><td>0.09</td></tr></tbody></table></div>\n<p>Runtime async supports <code>Task</code>, <code>Task&lt;T&gt;</code>, <code>ValueTask</code>, and <code>ValueTask&lt;T&gt;</code> as method return types, but as of today it doesn’t support <code>async void</code>, async iterators, or arbitrary custom task-like return types with custom builders; those continue to use the traditional compiler transformation. For <code>ValueTask&lt;T&gt;</code>, the existing reasons to use the type still apply. A <code>ValueTask&lt;T&gt;</code> can carry a result directly, wrap a <code>Task&lt;T&gt;</code>, or refer to an <code>IValueTaskSource&lt;T&gt;</code>. That’s made it useful for APIs where synchronous completion is common enough that avoiding a <code>Task</code> allocation outweighs the larger return value and the more restrictive consumption rules, or where asynchronous completion can have its costs amortized via a reusable backing object. Runtime async then addresses some of the scenarios that would have led developers to use <code>ValueTask&lt;T&gt;</code>. Does that mean everyone should stop using <code>ValueTask&lt;T&gt;</code>? No. Choosing <code>Task</code> versus <code>ValueTask</code> remains an API design decision based on completion patterns, allocation sensitivity, call frequency, and how consumers need to use the result. Write the return type that makes sense for the API, then let the compiler, VM, and JIT optimize it as best they can.</p>\n<p>Workloads with many layers of small async methods can benefit the most from runtime async, because those layers are exactly where intermediate tasks and state machines often accumulate. Shared framework code, for example, is full of this pattern: a public method validates arguments and awaits a private helper, which awaits a transport helper, which awaits an operating-system operation. Application services similarly compose authentication, retry, logging, serialization, and I/O helpers. Runtime async can make the source-level decomposition cheaper without asking the developer to flatten the code into one giant method in order to avoid “implementation detail” costs.</p>\n<p>The work required to reach this point has been extensive. A GitHub search of the runtime async tracking label on September 14, 2026 returned 235 pull requests, far too many for me to enumerate one by one. So I won’t try; you can peruse that label in your spare time. The work is also not only about direct performance improvements but also about\nimprovements to diagnostics and performance tooling that help you to make better\nuse of async in your own code. When an async method\nsuspends, its physical thread stack unwinds. That method’s continuation might later run\non a different thread whose physical stack begins in the thread pool, with the\nmethods that led to the original <code>await</code> nowhere to be found. A sampling\nCPU profiler can see where the processor is spending time, but without additional\ninformation, it can’t reliably connect those traces back through the logical async\ncall chain, making it hard to answer questions about what async call paths were actually costing.\nProfiling tools like the async profiler in Visual Studio have traditionally reconstructed those chains from\nevents emitted by <code>Task</code>‘s infrastructure, but async-heavy applications can generate enormous volumes of\nthose very chatty events. The resulting overhead easily perturbs the workload being measured, making\nit all but unusable in production. dotnet/runtime#127238 added a\nnew lightweight async-profiler event stream for .NET 11 and runtime async. Rather than sending every small\ntransition through the eventing system as its own full event, the runtime\nwrites compact records into per-thread buffers, delta-encoding timestamps and\ninstruction pointers and flushing the data in batches. It also puts a small\nidentifiable wrapper frame into the physical stack when invoking a\ncontinuation. A profiler can use that frame as an anchor, joining ordinary CPU\nsamples to the logical async call stack represented by the event stream. In some measurements,\nthis new approach added less than 1% overhead and shrank the traced data by an order of magnitude.\ndotnet/runtime#129043 and a few follow-up PRs extended\nthe same approach to the compiler-generated state machines used by existing\nasync code. Thus this\nisn’t useful only to applications that opt into runtime async; tooling gets one\nconsistent representation across both implementations.</p>\n<p>What should you as a developer do differently with runtime async in the picture? Mostly nothing. Keep writing asynchronous code the way you want it to read, and break a large operation into helpers when that makes the code clearer. Use <code>Task</code> by default and choose <code>ValueTask</code> where its API and usage tradeoffs genuinely fit. And don’t contort source code to remove a clean <code>await</code> just because today’s implementation might allocate an intermediate <code>Task</code>. The lowering strategy should “just work” as an implementation detail, preserve behavior, and make existing source get better as the runtime improves.</p>\n<h3 id=\"bounds-checks\">Bounds Checks</h3>\n<p>C# is a memory-safe language. Accesses to arrays, strings, and spans are guaranteed by the runtime to be in-bounds; if you try to access <code>someArray[i]</code>, <code>someString[i]</code>, or <code>someSpan[i]</code> with an index less than 0 or greater than or equal to the length of the array/string/span, you’ll get an exception, not silently corrupted memory or a process crash. The runtime guarantees that all permitted accesses are within bounds, and that means it needs to be able to prove the access is in bounds. The main method the JIT has for achieving that is by injecting code that performs a bounds check, as if instead of:</p>\n<pre><code>int[] array = ...;\nint value = array[i];</code></pre>\n<p>you’d written:</p>\n<pre><code>int[] array = ...;\nif ((uint)i &gt;= array.Length) throw new IndexOutOfRangeException();\nint value = array[i];</code></pre>\n<p>At the assembly level, a bounds check looks something like:</p>\n<pre><code>; x64\ncmp ecx, dword ptr [rax+8]        ; compare index with array length\njae THROW                         ; unsigned index &gt;= length\nmov edx, dword ptr [rax+rcx*4+16] ; load the element</code></pre>\n<p>The JIT could just inject such code on every access and call it a day, but such code adds overhead, so the JIT works to elide those checks and that overhead wherever it can prove the index is valid. Proving an index is valid means the JIT needs to be able to see from other evidence that it couldn’t possibly be out of bounds.</p>\n<p>The quintessential example of that is a <code>for</code> loop over the full contents of an array or span:</p>\n<pre><code>for (int i = 0; i &lt; array.Length; i++)\n{\n    Use(array[i]);\n}</code></pre>\n<p>The JIT recognizes from this idiom that, within the loop body, <code>i</code> is guaranteed to be in the range <code>[0, array.Length)</code>, and avoids emitting the bounds check for the <code>array[i]</code> access. The JIT has long handled this particular case. Other cases, not so much. Bounds-check elimination has improved in virtually every .NET release; more recent releases added range propagation for derived expressions (<code>.NET 7</code> and <code>.NET 8</code> saw significant improvements here), SSA-based reasoning (<code>.NET 9</code>), and better handling of <code>Span&lt;T&gt;</code>, whose length sits in a field rather than an object header, complicating tracking. Each year, the developers contributing to the JIT find new patterns that were being missed, that show up in the wild, and that are fixable. .NET 11 improves several such patterns.</p>\n<p>Range analysis in the JIT tracks intervals for each variable, an upper bound and a lower bound. For example, taking the true branch of <code>x &lt; 5</code> gives the range for <code>x</code> in that branch an upper bound of 4 while taking the true branch of <code>x &gt; 2</code> makes the lower bound 3. What about <code>x != 5</code>? On the true edge, we know <code>x</code> isn’t 5, and if the current range for <code>x</code> is <code>[5, 10]</code>, then we know the range must actually be <code>[6, 10]</code>… the lower bound can be tightened because the only value at the lower end is excluded. Similarly, if the range is <code>[0, 5]</code>, an <code>x != 5</code> assertion tells us the range is actually the narrower <code>[0, 4]</code>. Or, at least, that’s what you’d hope it would do. The JIT had this relevant comment:</p>\n<pre><code>// We have a != assertion, but it doesn&#39;t tell us much about the interval. So just skip it.\ncontinue;</code></pre>\n<p>In .NET 11, dotnet/runtime#121273 replaces that logic with productive reasoning. It checks whether the excluded constant is at either edge of the currently tracked range, adding in the new insights if so. C# list patterns, introduced in C# 11, generate just such comparison sequences. For example, the pattern <code>name is [] or [&#39;:&#39;] or [&#39;:&#39;, not &#39;:&#39;, ..]</code> lowers to something like this:</p>\n<pre><code>if (name != null)\n{\n    int num = name.Length;\n    if (num == 0) return true;\n    if (num == 1)\n    {\n        if (name[0] == &#39;:&#39;) return true;\n    }\n    else if (name[0] == &#39;:&#39; &amp;&amp; name[1] != &#39;:&#39;)\n    {\n        return true;\n    }\n    return false;\n}</code></pre>\n<p>Range analysis then proceeds with something like this:</p>\n<ol><li>We know that <code>Array.Length</code> is never negative, so it has a range of<code>[0, Array.MaxLength]</code> .</li><li>On the false edge of <code>num == 0</code> , we know that<code>num != 0</code> , so the range is narrowed now to<code>[1, Array.MaxLength]</code> .</li><li>Similarly, on the false edge of <code>num == 1</code> , we know that<code>num != 1</code> , so the range is narrowed now to<code>[2, Array.MaxLength]</code> .</li><li>We then access <code>name[0]</code> and<code>name[1]</code> , both of which are guaranteed in bounds based on the lower bound of 2 that was established.</li></ol>\n<p>Without the <code>!= constant</code> tightening, that narrowing wouldn’t happen, and the bounds checks in step 4 couldn’t be elided. Thankfully, they now can be in .NET 11. Consider this example:</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Runtime.CompilerServices;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private string[] _inputs = [&quot;&quot;, &quot;:&quot;, &quot;:x&quot;, &quot;abc&quot;, &quot;:ab&quot;, &quot;x&quot;, &quot;ab:cd&quot;];\n    [Benchmark]\n    public int ClassifyAll()\n    {\n        int total = 0;\n        foreach (string s in _inputs) total += Classify(s);\n        return total;\n    }\n    [MethodImpl(MethodImplOptions.NoInlining)]\n    private static int Classify(ReadOnlySpan&lt;char&gt; name) =&gt;\n        name switch\n        {\n            [] =&gt; 0,\n            [&#39;:&#39;] =&gt; 1,\n            [&#39;:&#39;, not &#39;:&#39;, ..] =&gt; 10 + name[0] + name[1],\n            _ =&gt; 3\n        };\n}</code></pre>\n<p>In .NET 10, we can see the call to <code>CORINFO_HELP_RNGCHKFAIL</code> at the bottom of the method. That’s the tell-tale sign there was at least one bounds check in the method. With .NET 11, that sign is removed.</p>\n<pre><code>; Arm64\n--- .NET 10\n+++ .NET 11\n@@ -10,17 +10,15 @@\n             beq     G_M000_IG08\n G_M000_IG04:\n-            ldrh    w2, [x0]\n-            cmp     w2, #58\n+            ldrh    w1, [x0]\n+            cmp     w1, #58\n             bne     G_M000_IG06\n G_M000_IG05:\n-            cmp     w1, #1\n-            bls     G_M000_IG11\n             ldrh    w0, [x0, #0x02]\n             cmp     w0, #58\n             beq     G_M000_IG06\n-            add     w0, w2, w0\n+            add     w0, w1, w0\n             add     w0, w0, #10\n             b       G_M000_IG07\n@@ -44,8 +42,4 @@\n             mov     w0, wzr\n             b       G_M000_IG07\n-G_M000_IG11:\n-            bl      CORINFO_HELP_RNGCHKFAIL\n-            brk     #0\n-\n-; Total bytes of code 112\n+; Total bytes of code 96</code></pre>\n<p>“Assertion” machinery in the JIT propagates learned facts (like the aforementioned range information) between “basic blocks” (a sequence of instructions with one entry point, one exit point, and no branches into or out of the middle of it), so information established in block A flows to block B if A “dominates” B (meaning the only way to get to B is through A). But what about facts established earlier within the same block? That’s the gap that dotnet/runtime#121527 addresses. Consider this code:</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Runtime.CompilerServices;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private int[] _arr = new int[512];\n    [Benchmark]\n    public int RunMany()\n    {\n        int touched = 0;\n        for (int i = 0; i &lt; _arr.Length - 2; i++)\n        {\n            Test(_arr, i);\n            touched++;\n        }\n        return touched;\n    }\n    [MethodImpl(MethodImplOptions.NoInlining)]\n    private static void Test(int[] arr, int i)\n    {\n        arr[i] = 0;  // 1: establishes &#39;i &gt;= 0 &amp;&amp; i &lt; arr.Length&#39;\n        i++;         // 2: same block\n        if (i &lt; arr.Length) arr[i] = 0;  // 3: proven safe from 1&#39;s assertion\n    }\n}</code></pre>\n<p>Statements 1, 2, and 3 are all in the same basic block, up to the conditional; after statement 1 executes, if we reach statement 2, the bounds check on statement 1 passed, we know <code>i &gt;= 0</code> and <code>i &lt; arr.Length</code>, and after statement 2, <code>i</code> becomes <code>i + 1</code>. After the <code>if</code> guard <code>i &lt; arr.Length</code> we know the incremented <code>i</code> is still within bounds. But when the range check pass in the .NET 10 JIT examined statement 3’s bounds check, it saw the assertions propagated from predecessor blocks. Since the assertion from statement 1 is generated within the current block, the range check couldn’t see it. The PR fixed it to walk the current block’s tree in execution order, accumulating assertions as it went. When we reach statement 3’s bounds check, we’ve already walked past statement 1 and picked up its <code>i &gt;= 0 &amp;&amp; i &lt; arr.Length</code> assertion.</p>\n<pre><code>; Arm64\n--- .NET 10\n+++ .NET 11\n@@ -13,8 +13,6 @@\n             ble     G_M000_IG04\n G_M000_IG03:\n-            cmp     w1, w2\n-            bhs     G_M000_IG05\n             str     wzr, [x0, w1, UXTW #2]\n G_M000_IG04:\n@@ -25,4 +23,4 @@\n             bl      CORINFO_HELP_RNGCHKFAIL\n             brk     #0\n-; Total bytes of code 68\n+; Total bytes of code 60</code></pre>\n<p>There are almost an infinite number of things the JIT could look for and special-case. But every special case requires code, maintenance, and, most importantly, compilation time. A “just-in-time” compiler typically runs while the application is running, so the JIT itself must be optimized and spend its limited budget only where there’s a likely payoff. That pushes the developers building it toward patterns that occur in real workloads. One such pattern, often seen in libraries like format decoders, builds a table\nindex with bitwise operations on a byte, for example\n<code>((b &amp; 0x03) &lt;&lt; 4) | ((b &amp; 0xf0) &gt;&gt; 4)</code>. Each masked piece has a tiny upper\nbound, so the OR of those pieces is always in <code>[0..63]</code>, safely in range for\ne.g. a Base64 alphabet table. Until\ndotnet/runtime#122263, the JIT\noften failed to prove that combined bound and left a bounds check on the\nindex. Existing range-check code understood the upper bounds produced by\nbitwise AND and shifts, but not OR; the change lets the JIT combine the known\nbounds of both OR operands and remove the remaining array check.</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Runtime.CompilerServices;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private readonly byte[] _input = new byte[4096];\n    [GlobalSetup]\n    public void Setup() =&gt; new Random(42).NextBytes(_input);\n    [Benchmark]\n    public int Base64LikeIndex() =&gt; Sum(_input);\n    [MethodImpl(MethodImplOptions.NoInlining)]\n    private static int Sum(ReadOnlySpan&lt;byte&gt; input)\n    {\n        int sum = 0;\n        foreach (byte b in input)\n        {\n            int index = ((b &amp; 0x03) &lt;&lt; 4) | ((b &amp; 0xF0) &gt;&gt; 4);\n            sum += &quot;ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789+/=&quot;u8[index];\n        }\n        return sum;\n    }\n}</code></pre>\n<p>The .NET 10 assembly checks the computed index against the 65-byte lookup table on every iteration. In .NET 11, range analysis proves the index is at most 63, so both the comparison and the branch to the range-check failure helper disappear:</p>\n<pre><code>; x64\n M01_L00:\n        movzx    r9d, byte ptr [rdx+r8]\n        mov      r11d, r9d\n        and      r11d, 3\n        shl      r11d, 4\n        and      r9d, 0F0\n        sar      r9d, 4\n        or       r9d, r11d\n-       cmp      r9d, 41\n-       jae      short M01_L02\n        movzx    r9d, byte ptr [r10+r9]\n        add      eax, r9d\n        inc      r8d\n        cmp      r8d, ecx\n        jl       short M01_L00\n-M01_L02:\n-       call     CORINFO_HELP_RNGCHKFAIL\n-       int      3\n-\n-; Total bytes of code 95\n+; Total bytes of code 79</code></pre>\n<p>As another example, dotnet/runtime#125056 improves the handling of guards like <code>(uint)i &lt; span.Length</code> that are pervasive in performance-sensitive code. Consider this benchmark:</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Runtime.CompilerServices;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private int[] _data = Enumerable.Range(0, 512).ToArray();\n    [Benchmark]\n    public int RunMany()\n    {\n        int sum = 0;\n        for (int i = 0; i &lt; _data.Length; i++)\n            sum += Test(_data, i);\n        return sum;\n    }\n    [MethodImpl(MethodImplOptions.NoInlining)]\n    private static int Test(Span&lt;int&gt; span, int i)\n    {\n        if ((uint)i &lt; (uint)span.Length)\n        {\n            if (i != 0)\n                return span[i - 1] + span[i];\n            return span[i];\n        }\n        return 0;\n    }\n}</code></pre>\n<p>Because the comparison is unsigned, <code>(uint)i</code> would be a large positive number if <code>i</code> were negative, making it impossible for <code>(uint)i &lt; (uint)span.Length</code> to be true (since a span’s length is never negative, <code>(uint)span.Length</code> is at most <code>int.MaxValue</code>). Inside the true branch, <code>i</code> is therefore in <code>[0, span.Length - 1]</code>. Previously, the JIT wasn’t always recording the lower bound <code>i &gt;= 0</code> when it processed the <code>(uint)i &lt; span.Length</code> assertion, and that could leave bounds checks on expressions like <code>i - 1</code> in place. The fix adds the <code>[0, int.MaxValue - 1]</code> lower bound deduction for the index variable upon entering the true arm of a <code>(uint)i &lt; span.Length</code> check. Combined with the existing range tracking for the upper bound, this gives the JIT a complete picture of <code>i</code>‘s range inside the guarded block.</p>\n<pre><code>; Arm64\n--- .NET 10\n+++ .NET 11\n@@ -8,10 +8,8 @@\n             cbz     w2, G_M000_IG05\n G_M000_IG03:\n-            sub     w3, w2, #1\n-            cmp     w3, w1\n-            bhs     G_M000_IG09\n-            ldr     w1, [x0, w3, UXTW #2]\n+            sub     w1, w2, #1\n+            ldr     w1, [x0, w1, UXTW #2]\n             ldr     w0, [x0, w2, UXTW #2]\n             add     w0, w1, w0\n@@ -33,8 +31,4 @@\n             ldp     fp, lr, [sp], #0x10\n             ret     lr\n-G_M000_IG09:\n-            bl      CORINFO_HELP_RNGCHKFAIL\n-            brk     #0\n-\n-; Total bytes of code 84\n+; Total bytes of code 68</code></pre>\n<p>Bounds check elision is generally based on forms of range analysis, where the JIT needs to prove that a given index is guaranteed to be within the range of the data structure. But the same range analysis-based facts can prove that other checks are unnecessary. For example, once the JIT knows that an integer is in <code>[0..100]</code>, it can prove both that converting it to <code>byte</code> can’t lose data and that multiplying it by 10 can’t overflow. dotnet/runtime#124147 enables the JIT to use such facts to avoid unnecessary branches as part of <code>checked</code> operations. When range analysis proves that the operands are in ranges whose result can’t overflow, making <code>checked</code> a nop, the backend can now emit plain add/multiply/subtract instructions, without the jump to failure, as in the following example:</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Runtime.CompilerServices;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private readonly int[] _array = new int[99];\n    [Benchmark]\n    public int ArrayLengthPlusConstant() =&gt; AddToLength(_array);\n    [MethodImpl(MethodImplOptions.NoInlining)]\n    private static int AddToLength(int[] array) =&gt; checked(array.Length + 10);\n    [Benchmark]\n    public int GuardedLengthTimesConstant() =&gt; Multiply(_array);\n    [MethodImpl(MethodImplOptions.NoInlining)]\n    private static int Multiply(Span&lt;int&gt; span)\n    {\n        if (span.Length &gt;= 100) return 0;\n        return checked(span.Length * 10);\n    }\n}</code></pre>\n<pre><code>; Arm64\n--- .NET 10\n+++ .NET 11\n G_M000_IG02:\n             cmp     w1, #100\n             bge     G_M000_IG05\n G_M000_IG03:\n             mov     w0, #10\n-            smull   x0, w1, w0\n-            lsr     x2, x0, #32\n-            cmp     w2, w0, ASR #31\n+            mul     w0, w1, w0\n-            bne     G_M000_IG07\n G_M000_IG04:\n             ldp     fp, lr, [sp], #0x10\n             ret     lr\n-G_M000_IG07:\n-            bl      CORINFO_HELP_OVERFLOW\n-            brk     #0\n-\n-; Total bytes of code 64\n+; Total bytes of code 44</code></pre>\n<p>That makes the change broadly applicable: any time you write <code>checked</code> arithmetic on quantities that are inherently bounded, such as collection counts, lengths, or indices constrained by prior comparisons, the JIT now has a chance to prove at compile time that the overflow can’t happen and thus eliminate the run-time check entirely. Building on that range-check work, dotnet/runtime#124184 teaches the JIT to eliminate “narrowing casts” under the same kinds of guards:</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Runtime.CompilerServices;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private uint _value = 100;\n    [Benchmark]\n    public byte GuardedNarrowingCast() =&gt; Narrow(_value);\n    [MethodImpl(MethodImplOptions.NoInlining)]\n    private static byte Narrow(uint value)\n    {\n        if (value &gt; 100) return 0;\n        return checked((byte)value);\n    }\n}</code></pre>\n<pre><code>; Arm64\n--- .NET 10\n+++ .NET 11\n@@ -4,23 +4,10 @@\n G_M000_IG02:\n             cmp     w0, #100\n-            bhi     G_M000_IG04\n-            cmp     w0, #255\n-            bhi     G_M000_IG06\n+            csel    w0, w0, wzr, ls\n G_M000_IG03:\n             ldp     fp, lr, [sp], #0x10\n             ret     lr\n-G_M000_IG04:\n-            mov     w0, wzr\n-\n-G_M000_IG05:\n-            ldp     fp, lr, [sp], #0x10\n-            ret     lr\n-\n-G_M000_IG06:\n-            bl      CORINFO_HELP_OVERFLOW\n-            brk     #0\n-\n-; Total bytes of code 52\n+; Total bytes of code 24</code></pre>\n<p>Such use of <code>checked</code> is common in serialization and protocol code where you validate a value’s range prior to truncating it. In this benchmark I’ve used <code>checked</code> explicitly, but the more common form is with the whole project compiled with <code>&lt;CheckForOverflowUnderflow&gt;true&lt;/CheckForOverflowUnderflow&gt;</code> in the .csproj, such that this <code>checked</code> becomes implicit. After the change, the range analysis sees that <code>value</code> is in the range <code>[0, 100]</code>, knows <code>byte</code> fits values up to 255, and elides the check.</p>\n<p>dotnet/runtime#128620 further teaches range analysis the possible results of leading-zero count, trailing-zero count, and population count instructions. Those results are often used to index small lookup tables… knowing their bounds lets the JIT remove the bounds check.</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Numerics;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private static readonly int[] s_lookup =\n        Enumerable.Range(0, 33).Select(i =&gt; i * i).ToArray();\n    private uint[] _values;\n    [GlobalSetup]\n    public void Setup()\n    {\n        Random rng = new(42);\n        _values = Enumerable.Range(0, 1024).Select(i =&gt; (uint)rng.Next(1, int.MaxValue)).ToArray();\n    }\n    [Benchmark]\n    public int SumLookupByLeadingZeroCount()\n    {\n        int sum = 0;\n        foreach (var v in _values)\n            sum += s_lookup[BitOperations.LeadingZeroCount(v)];\n        return sum;\n    }\n}</code></pre>\n<p>The lookup improves because the JIT now knows <code>LeadingZeroCount(uint)</code> is between 0 and 32 and can remove the bounds check.</p>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Mean</th><th>Ratio</th></tr></thead><tbody><tr><td>SumLookupByLeadingZeroCount</td><td>.NET 10.0</td><td>516.1 ns</td><td>1.00</td></tr><tr><td>SumLookupByLeadingZeroCount</td><td>.NET 11.0</td><td>438.5 ns</td><td>0.85</td></tr></tbody></table></div>\n<p>The JIT is also able to conditionally apply range check-based elision via “cloning”. Cloning is a mechanism where the JIT takes one piece of code and duplicates it. One of the copies it leaves as it was originally, and the other copy it special cases. So, for example, if you had code like:</p>\n<p><code>int value = array[i];</code>\nthe JIT could theoretically clone that in order to avoid the implicit bounds check, e.g.</p>\n<pre><code>int value;\nif ((uint)i &lt; array.Length)\n{\n    // no bounds check emitted by JIT, e.g.\n    value = Unsafe.Add(ref MemoryMarshal.GetArrayDataReference(array), i);\n}\nelse\n{\n    // bounds check emitted\n    value = array[i];\n}</code></pre>\n<p>That particular code looks silly, as we’re just trading an implicit bounds check for an explicit one. It becomes less silly when the JIT is able to elide multiple bounds checks with a single branch, e.g.</p>\n<pre><code>int sum;\nif (4 &lt; array.Length)\n{\n    // zero bounds checks\n    ref int startRef = ref MemoryMarshal.GetArrayDataReference(array);\n    sum =\n        startRef +\n        Unsafe.Add(ref startRef, 1) +\n        Unsafe.Add(ref startRef, 2) +\n        Unsafe.Add(ref startRef, 3);\n}\nelse\n{\n    // potentially four bounds checks\n    sum =\n        array[0] +\n        array[1] +\n        array[2] +\n        array[3];\n}</code></pre>\n<p>Such optimizations are already handled in the JIT, via its <code>optRangeCheckCloning</code> phase. It groups bounds checks from a basic block, emits one guard for the largest required range, and duplicates the affected code into a fast path where the individual checks can be removed and a fallback path where they remain. However, one long-standing limitation of range-check cloning is that it refused to process the last statement of any “terminator” block, a block that ends with a jump or return instruction. For a method like:</p>\n<p><code>static int ArrayAccess(int[] abcd) =&gt; abcd[0] + abcd[1] + abcd[2] + abcd[3];</code>\nall four array accesses live in the return statement, the last statement of a return block, so nothing got cloned and the hot path retained four separate bounds checks. In .NET 11, dotnet/runtime#124705 removes that restriction, making the return statement eligible for range-check cloning and allowing a single fast-path guard to cover all four accesses.</p>\n<p>But even without range-check cloning, there’s really no reason such accesses should require four bounds checks: the JIT should be able to see that the array or span needs to have a length of at least 4 and guard all accesses by that single check. If there were intervening operations that had side effects, the JIT would need to maintain order of operations, at least enough to maintain the observable behavior of those effects, but that’s not the case here. With dotnet/runtime#127439 in .NET 11, the JIT will now coalesce those checks within a basic block, strengthening the first check to the largest constant index and removing the rest.</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private readonly int[] _values = Enumerable.Range(0, 16).ToArray();\n    [Benchmark]\n    public int Sum16()\n    {\n        int[] values = _values;\n        return\n            values[0] + values[1] + values[2] + values[3] +\n            values[4] + values[5] + values[6] + values[7] +\n            values[8] + values[9] + values[10] + values[11] +\n            values[12] + values[13] + values[14] + values[15];\n    }\n}</code></pre>\n<p>In previous releases, you’d sometimes see a proactive developer doing a similar optimization manually, e.g. reordering the accesses in an example like that to put the largest read first. That’s no longer necessary.</p>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Mean</th><th>Ratio</th></tr></thead><tbody><tr><td>Sum16</td><td>.NET 10.0</td><td>2.958 ns</td><td>1.00</td></tr><tr><td>Sum16</td><td>.NET 11.0</td><td>1.828 ns</td><td>0.62</td></tr></tbody></table></div>\n<p>Another bounds checking improvement comes in dotnet/runtime#127488, which actually targets explicitly-implemented bounds checks (rather than the implicit ones we’ve been discussing) and targets code that reads a fixed-size value from the end of a span, such as <code>BinaryPrimitives.ReadInt32BigEndian(span.Slice(span.Length - 4))</code> behind a <code>span.Length &gt;= 4</code> guard, e.g.</p>\n<pre><code>if (span.Length &gt;= 4)\n{\n    // Parse an int from the end of the span\n    ... = ReadInt32BigEndian(span.Slice(span.Length - 4));\n    ...\n}</code></pre>\n<p>There shouldn’t be any additional bounds checking required here. However, <code>Span.Slice</code> begins with:</p>\n<pre><code>if ((uint)start &gt; (uint)_length)\n    ThrowHelper.ThrowArgumentOutOfRangeException();</code></pre>\n<p>and <code>ReadInt32BigEndian</code> begins with:</p>\n<pre><code>if (sizeof(T) &gt; source.Length)\n    ThrowHelper.ThrowArgumentOutOfRangeException();</code></pre>\n<p>so even though our <code>span.Length &gt;= 4</code> check should have been sufficient, we’re still ending up with two additional checks. To address that, the JIT needed two things.</p>\n<p>First, it needed to be able to identify that <code>x - (x + a)</code> is the same as <code>-a</code>. Without this identity, <code>length - (length - 4)</code> is just an opaque subtraction of two expressions with no obvious constant result. With the identity, the JIT can recognize the inner expression <code>(length - 4)</code> as <code>length + (-4)</code>, apply <code>x - (x + a) == -a</code> with <code>x == length</code> and <code>a == -4</code>, and end up with <code>-(-4) == 4</code>. Now <code>ReadInt32BigEndian</code>‘s check against 4 becomes <code>4 &gt;= 4</code>, which the JIT can trivially see is true.</p>\n<p>Second, <code>Slice(start)</code> must establish that <code>start</code> is between zero and the span’s length. When <code>start</code> is <code>length - 4</code>, the existing <code>length &gt;= 4</code> guard proves the result is non-negative, while subtracting a positive constant means the result can’t exceed <code>length</code>. The improved range analysis connects that guard to the subtraction and removes <code>Slice</code>‘s check.</p>\n<p>Both fixes together mean the above example now elides both extra bounds checks. That’s useful in particular for libraries like parsers, network protocol implementations, and cryptographic code, all of which frequently on hot paths do things like “read the last N bytes of a buffer.”</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Buffers.Binary;\nusing System.Runtime.CompilerServices;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser]\n[HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private readonly byte[] _buffer = new byte[64];\n    [Benchmark]\n    public int ReadLastInt32() =&gt; ReadLastInt32(_buffer);\n    [MethodImpl(MethodImplOptions.NoInlining)]\n    private static int ReadLastInt32(ReadOnlySpan&lt;byte&gt; span)\n    {\n        if (span.Length &gt;= sizeof(int))\n        {\n            return BinaryPrimitives.ReadInt32BigEndian(span.Slice(span.Length - sizeof(int)));\n        }\n        return -1;\n    }\n}</code></pre>\n<p>In .NET 10, the helper is 73 bytes and includes both additional checks and their throw paths:</p>\n<pre><code>; x64\ncmp       ecx,4\njl        RETURN_MINUS_ONE\nlea       edx,[rcx-4]\ncmp       edx,ecx\nja        THROW_SLICE\nmov       r8d,edx\nadd       rax,r8\nsub       ecx,edx\ncmp       ecx,4\njl        THROW_READ\nmovbe     eax,[rax]</code></pre>\n<p>In .NET 11, the helper is 28 bytes, and only the original length guard remains:</p>\n<pre><code>; x64\ncmp       ecx,4\njl        RETURN_MINUS_ONE\nadd       ecx,-4\nadd       rax,rcx\nmovbe     eax,[rax]</code></pre>\n<p>dotnet/runtime#122040 and dotnet/runtime#127117 similarly help to remove bounds checks involving <code>span.Slice</code>. Vectorized loops often work through a span a chunk at a time, slicing off the elements they’ve already processed. The JIT hasn’t always been able to keep track of how those progressively smaller slices relate to the original span, so it could end up checking the same limits again on each iteration. These changes improve that tracking, enabling more of those repeated checks to be removed.</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing System.Runtime.CompilerServices;\nusing System.Runtime.Intrinsics;\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private readonly int[] _data = Enumerable.Repeat(1, 1_024).ToArray();\n    [Benchmark]\n    public Vector256&lt;int&gt; CreateFromSlice() =&gt; CreateFromSlice(_data);\n    [Benchmark]\n    public int SumSliced() =&gt; SumSliced(_data);\n    [MethodImpl(MethodImplOptions.NoInlining)]\n    private static Vector256&lt;int&gt; CreateFromSlice(Span&lt;int&gt; values)\n    {\n        if (values.Length &lt; 16)\n            return default;\n        return Vector256.Create(values.Slice(8));\n    }\n    [MethodImpl(MethodImplOptions.NoInlining)]\n    private static int SumSliced(ReadOnlySpan&lt;int&gt; data)\n    {\n        Vector128&lt;int&gt; sum = default;\n        while (data.Length &gt;= Vector128&lt;int&gt;.Count)\n        {\n            sum += Vector128.Create(data);\n            data = data.Slice(Vector128&lt;int&gt;.Count);\n        }\n        int result = Vector128.Sum(sum);\n        foreach (int value in data)\n            result += value;\n        return result;\n    }\n}</code></pre>\n<p>In .NET 10, the loop condition proves that at least one vector remains, but the construction of the vector from the current span performs the same check again. .NET 11 retains the length relationship, so the loop body begins directly with the vector addition:</p>\n<pre><code>; x64, vector loop\n-cmp       esi, 4\n-jl        THROW_ARGUMENT_OUT_OF_RANGE\n-vpaddd    xmm6, xmm6, [rbx]\n-add       rbx, 10\n-add       esi, 0FFFFFFFC\n-cmp       esi, 4\n+vpaddd    xmm0, xmm0, [rax]\n+add       rax, 10\n+add       ecx, 0FFFFFFFC\n+cmp       ecx, 4\n jge       LOOP</code></pre>\n<p>We saw earlier how range-check cloning enables duplicating a sequence of instructions in order to eliminate bounds checks. “Loop cloning” extends that to a whole loop. Consider a loop that processes the first <code>count</code> elements of an array:</p>\n<pre><code>for (int i = 0; i &lt; count; i++)\n    sum += values[i];</code></pre>\n<p>The test <code>i &lt; count</code> doesn’t by itself prove that <code>i &lt; values.Length</code>, so by default the compilation would need a bounds check in the body, which would mean a bounds check for every <code>values[i]</code> access. Loop cloning gives the JIT another option. Instead of generating the equivalent of:</p>\n<pre><code>for (int i = 0; i &lt; count; i++)\n    sum += values[i]; // bounds check!</code></pre>\n<p>it can generate the equivalent of:</p>\n<pre><code>if ((uint)count &lt;= (uint)values.Length)\n{\n    // no bounds checks\n    ref int startRef = ref MemoryMarshal.GetArrayDataReference(values);\n    for (int i = 0; i &lt; count; i++)\n    {\n        sum += Unsafe.Add(ref startRef, i);\n    }\n}\nelse\n{\n    // bounds check per iteration\n    for (int i = 0; i &lt; count; i++)\n    {\n        sum += values[i];\n    }\n}</code></pre>\n<p>For the common case where the iteration is in bounds, execution proceeds through a cloned loop with no per-iteration bounds checks, whereas the original checked loop remains as the fallback that preserves exceptional behavior for invalid inputs. The normal path pays for one guard and avoids a check on every iteration, but that comes at the expense of duplicating code. The JIT therefore needs to apply the optimization selectively.</p>\n<p>The JIT has long employed loop cloning, but it didn’t always kick in even in cases it seemed applicable. The previous example showed loop cloning with <code>&lt;</code> in the iteration condition. For whatever reason, however, some developers used <code>!=</code>, and loop cloning didn’t apply (I’m guessing they used <code>!=</code> because they thought it was more efficient, and they actually end up deoptimizing). Thanks to dotnet/runtime#129268, in .NET 11 <code>!=</code> is now also handled, as long as specific conditions are met, such as the stride being exactly 1 or -1, e.g. <code>i++</code> qualifies, while <code>i += 2</code> doesn’t. dotnet/runtime#129303 also improves loops that terminate with <code>i != bound</code>, giving the JIT a tighter understanding of the values <code>i</code> can take and allowing it to remove some bounds checks even when it can’t clone the whole loop.</p>\n<p>Lookahead in arrays and spans is another recurring pattern, especially in parsers. dotnet/runtime#124242 and dotnet/runtime#125235 recognize conditions such as <code>(uint)(i + 2) &lt; (uint)span.Length</code> and use that relation to remove the follow-on checks for <code>span[i + 1]</code> and <code>span[i + 2]</code>. Consider this benchmark:</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System;\nusing System.Linq;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private readonly string _text = string.Concat(Enumerable.Repeat(&quot;%FE&quot;, 128));\n    [Benchmark]\n    public bool ContainsPercentFF()\n    {\n        ReadOnlySpan&lt;char&gt; span = _text;\n        for (int i = 0; i &lt; span.Length; i++)\n        {\n            if (span[i] == &#39;%&#39; &amp;&amp;\n                (uint)(i + 2) &lt; (uint)span.Length &amp;&amp;\n                span[i + 1] == &#39;F&#39; &amp;&amp;\n                span[i + 2] == &#39;F&#39;)\n            {\n                return true;\n            }\n        }\n        return false;\n    }\n}</code></pre>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Mean</th><th>Ratio</th></tr></thead><tbody><tr><td>ContainsPercentFF</td><td>.NET 10.0</td><td>232.1 ns</td><td>1.00</td></tr><tr><td>ContainsPercentFF</td><td>.NET 11.0</td><td>194.9 ns</td><td>0.84</td></tr></tbody></table></div>\n<p>Several smaller changes broaden the range of code from which the JIT can remove bounds checks:</p>\n<ul><li>dotnet/runtime#121640 helps in a situation where once an access using a chosen index has been checked, a later access to the same array at that index need not be checked again.</li><li>dotnet/runtime#121683 enables the JIT to trace an array’s length through calculations performed earlier in the method, exposing more redundant checks, including some involving index-from-end expressions.</li><li>dotnet/runtime#124387 and dotnet/runtime#130326 teach the optimizer to rely on a span’s length always being non-negative.</li><li>dotnet/runtime#124571 improves sequences of index-from-end accesses: once an access like <code>arr[^4]</code> establishes that the array has at least four elements, the JIT reuses that information for nearby accesses such as<code>arr[^3]</code> .</li><li>dotnet/runtime#129101 improves how the JIT combines and carries forward the possible ranges of arithmetic expressions, including expressions involving bitwise OR and unsigned division. Those tighter ranges can show that more values are non-negative or within bounds.</li></ul>\n<p>Bounds-check elimination is only one payoff from understanding a loop’s structure. The JIT analyzes induction variables (values like loop counters that change predictably each iteration) and puts loops into standard forms so that later optimizations can reason about them. .NET 11 broadens the range of loops for which that works:</p>\n<ul><li>dotnet/runtime#122184 recognizes another representation of a 32-to-64-bit zero extension. That lets pointer loops using expressions such as <code>data[(uint)i]</code> replace the repeated index extension and address calculation with a pointer increment.</li><li>dotnet/runtime#119537 follows simple control-flow predecessors when finding an induction variable’s initialization and zero-trip test, while dotnet/runtime#128303 gives loops with multiple backedges a single canonical latch block.</li><li>dotnet/runtime#128532 makes loop cloning tolerate more statements around the update and test.</li><li>dotnet/runtime#129309 extends cloning to more span loops with non-unit strides and offset limits.</li><li>dotnet/runtime#129349 handles large strides in array loops with an explicit safety guard rather than rejecting them outright.</li><li>dotnet/runtime#129472 allows loop inversion to spend more of its budget on likely cloning candidates.</li><li>dotnet/runtime#130205 removes comparisons that are redundant given the induction variable’s known range.</li><li>dotnet/runtime#131362 corrects profile weights after inversion changes a loop’s exit.</li></ul>\n<p>Much of this work wasn’t motivated by contrived benchmarks containing nothing\nbut array indexing as I’m prone to use in these posts. Rather, many of the improvements\nstemmed from an ongoing audit of unsafe code throughout the\n.NET libraries, part of a broader\neffort to improve memory safety in .NET. .NET and C# are memory safe, but as with other memory safe languages like Rust,\nit provides escape hatches that enable turning off the guardrails provided by the compiler and runtime.\nThis effort is about reducing where and when developers feel compelled to use those escape hatches, since every\noccurrence is an opportunity for increased risk. Unsafe code was often introduced years earlier to manually avoid bounds\nchecks, typically by walking a buffer with pointers, byrefs, or <code>Unsafe.Add</code>.\nSometimes the audit found that the unsafe code was no longer needed and could\nsimply be removed. Sometimes a “safe” rewrite (meaning not using <code>unsafe</code> and friends) was already just as fast or even faster.\nAnd sometimes the rewrite exposed an optimization the JIT was missing, in which\ncase the answer was to improve the JIT and then rewrite the library code to use\nnormal, bounds-checked C#. Several of the optimizations discussed in this\nsection are the result of exactly that feedback loop. dotnet/runtime#127429 is a\nparticularly nice example. The vectorized implementation of <code>Enumerable.Sum</code>\nused <code>MemoryMarshal.GetReference</code>, <code>Vector.LoadUnsafe</code>, and <code>Unsafe.Add</code> to walk\nits input without bounds checks. With the span-slicing improvements described\nearlier, it could instead use <code>Vector.Create(span)</code>, <code>span.Slice(...)</code>, and a\n<code>foreach</code> for the tail. That’s easier to reason about, removes the unchecked\nindexing, and ended up being faster.\ndotnet/runtime#114757 similarly\nreplaced an unsafe pointer-based header-name accessor with a generic\n<code>ReadOnlySpan&lt;T&gt;</code> implementation without loss of performance.\nSimilarly, dotnet/runtime#121270 removed\nmore unsafe code from <code>Uri</code> parsing and actually improved performance of the cited code measurably.</p>\n<p>There’s a useful “go do” here for libraries outside of dotnet/runtime, as well.\nUnsafe code written to work around the JIT is a snapshot of what the JIT could\ndo at the time that code was written. If you own code that has hand-written pointer or\n<code>Unsafe</code>-based loops whose purpose is to avoid bounds checks, it’s worth\nrewriting them with safe, bounds-checked C# and measuring again on .NET 11.\nChances are, you’ll find the gap at this point is either non-existent or small\nenough that it’s not worth the increased maintenance and risk for managing the safety yourself.\nAnd if the revised version is still slower, that’s a great opportunity for you to share a repro\nin the dotnet/runtime repo, hopefully serving as inspiration for one of the first performance improvements\nto go into the JIT for .NET 12. <code>unsafe</code> code is still necessary for scenarios like interop,\nbut performance alone shouldn’t be a permanent reason to eschew all the valuable guardrails .NET provides.</p>\n<p>The C# 15 memory-safety preview\npushes in the same direction and is part and parcel of this effort. Historically, C# has largely equated pointers\nwith unsafe code: simply declaring or manipulating a pointer generally required\nan <code>unsafe</code> context, even if the code never accessed the memory to which it\npoints. In the preview, pointer plumbing such as declaring a pointer, taking an\naddress with <code>&amp;</code>, using <code>fixed</code>, converting <code>stackalloc</code> to a pointer, and\napplying <code>sizeof</code> to an unmanaged type no longer requires an <code>unsafe</code> context.\nOperations that actually access the pointed-to memory, including <code>*p</code>,\n<code>p-&gt;member</code>, and <code>p[i]</code>, still do. C# 15 also adds an <code>unsafe(expression)</code> form,\nanalogous to <code>checked(expression)</code>, so an unsafe context can cover one precise\nexpression rather than a larger statement block. Those changes are the first preview slice of a larger, multi-release\nunsafe evolution.\nThe end goal is to make unsafe regions smaller, make their assumptions visible through the call graph,\nand make them easier for reviewers and tools to find. Pairing that with a JIT\nthat makes idiomatic safe code fast removes a lot of the historical pressure to\nuse unsafe code in the first place.</p>\n<h3 id=\"assertion-propagation\">Assertion Propagation</h3>\n<p>As discussed earlier, the JIT continually learns facts while compiling a method: a value equals a constant, a reference isn’t null, an integer falls within a particular range, and so on. “Assertion propagation” carries those facts forward so they can simplify later code. “Value numbering” complements it by letting the JIT recognize when two expressions compute the same value, even if they appear in different places or use different variables. Together, these mechanisms enable optimizations such as removing redundant null and bounds checks, folding conditions to constants, and reusing repeated computations. .NET 11 improves assertion propagation primarily by fixing places where useful facts were either never recorded or weren’t recognized later.</p>\n<p>For example, reading an array’s length normally carries an implicit null-check: if the array reference is <code>null</code>, the read must throw. Once global assertion propagation already knows the reference is non-null, however, we should be able to avoid the implicit null check. In .NET 11, dotnet/runtime#124291 takes care of that for <code>Array.Length</code>:</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Runtime.CompilerServices;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private readonly int[] _values = new int[1024];\n    [Benchmark]\n    public void DeadLength() =&gt; Test(_values);\n    [MethodImpl(MethodImplOptions.NoInlining)]\n    private static void Test(int[]? values)\n    {\n        if (values is not null)\n            _ = values.Length;\n    }\n}</code></pre>\n<p>The .NET 10 code still tests the reference and reads the length. In .NET 11, the guard proves the read can’t throw, and since its result isn’t used, the access disappears:</p>\n<pre><code>; Arm64\n--- .NET 10\n+++ .NET 11\n G_M000_IG01:\n             stp     fp, lr, [sp, #-0x10]!\n             mov     fp, sp\n G_M000_IG02:\n-            cbz     x0, G_M000_IG04\n-\n-G_M000_IG03:\n-            ldr     wzr, [x0, #0x08]\n-\n-G_M000_IG04:\n             ldp     fp, lr, [sp], #0x10\n             ret     lr\n-; Total bytes of code 24\n+; Total bytes of code 16</code></pre>\n<p>dotnet/runtime#119474 improves\nthe starting point for integer range analysis. The JIT now uses facts inherent\nin a value itself, e.g. a constant has one exact value, while a value converted\nto <code>byte</code>, for example, must be between 0 and 255. That can eliminate bounds\nchecks and conditions even when no preceding <code>if</code> explicitly established the\nrange. dotnet/runtime#124415\nfurther refines this handling of casts, combining what is known about both the\nsource value and the destination type to derive the tightest useful range.</p>\n<p>Those improvements derive ranges from facts inherent in a value, but ranges\ncan also come from control flow. After <code>if ((uint)x &lt; 10)</code>, for example, the\nJIT knows that <code>x</code> is between 0 and 9 on the true path, which may be enough to\nremove a later comparison or array bounds check.\ndotnet/runtime#123624\nderives tighter ranges from assertions and casts, including proving that some\ncomparisons are always true or false.\ndotnet/runtime#129390\npreserves range information more accurately when control-flow paths merge.</p>\n<p>Other changes make better use of the ranges once known. dotnet/runtime#129354 traces values back through their definitions to fold more span- and slice-related comparisons, and dotnet/runtime#126917 uses narrowed ranges to remove more relational branches.</p>\n<p>dotnet/runtime#124711 teaches the JIT to learn implicit facts from operations that have already completed successfully. For example:</p>\n<ul><li>Creating an array proves its requested length wasn’t negative.</li><li>A reference-array store may need a runtime covariance check, because a value typed as <code>object[]</code> can actually refer to a<code>string[]</code> ; the helper that performs that type check also validates the index, so if it returns successfully, the index was in range.</li><li>Integer division or modulo proves the divisor wasn’t zero.</li></ul>\n<p>And so on. Those facts can then remove redundant checks and conditions later in the method.</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private readonly object?[] _objArr = new object?[8];\n    private readonly object _value = new();\n    [Benchmark]\n    public object? CovariantArrayStore()\n    {\n        object?[] objArr = _objArr;\n        objArr[3] = _value;\n        return objArr[2];\n    }\n}</code></pre>\n<p>A successful store to element 3 proves that particular array has at least four elements; since an array’s length can’t change, the subsequent read of element 2 doesn’t need another bounds check.</p>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Mean</th><th>Ratio</th></tr></thead><tbody><tr><td>CovariantArrayStore</td><td>.NET 10.0</td><td>3.565 ns</td><td>1.00</td></tr><tr><td>CovariantArrayStore</td><td>.NET 11.0</td><td>2.985 ns</td><td>0.84</td></tr></tbody></table></div>\n<p>dotnet/runtime#128522 simplifies how the global assertion pass identifies values, making it less likely to miss a fact learned earlier. One practical impact of this is better propagation of a static <code>string</code>‘s known length, which can turn a general string comparison into a fixed-size vectorized comparison.</p>\n<p>dotnet/runtime#127810 improves null-check elimination where control flow merges. With <code>??=</code>, which is a very common operator used for lazy initialization, the resulting value is non-null whether it came from the existing field or from the newly allocated object. The JIT now combines the facts from both paths and recognizes that the subsequent call doesn’t need another null check.</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Runtime.CompilerServices;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    Inner? _inner;\n    [Benchmark]\n    [Arguments(42)]\n    public int Invoke(int n) =&gt; (_inner ??= new()).Increment(n);\n    private sealed class Inner\n    {\n        [MethodImpl(MethodImplOptions.NoInlining)]\n        public int Increment(int n) =&gt; n + 1;\n    }\n}</code></pre>\n<p>The generated code consequently loses the null check on the merged value:</p>\n<pre><code>; x64\n M00_L00:\n        mov      edx, esi\n-       cmp      [rcx], ecx\n        call     qword ptr [...] ; Inner.Increment(Int32)\n-; Total bytes of code 75\n+; Total bytes of code 73</code></pre>\n<p>dotnet/runtime#128701 removes similarly redundant null checks from copies of structs that contain object references. Such copies use a runtime helper so the garbage collector is correctly notified about the reference writes, but lowering had been adding probes for both source and destination without preserving whether either address could actually fault. It now emits only the probes that are needed.</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private FourRefs _src = new()\n    {\n        A = new(),\n        B = new(),\n        C = new(),\n        D = new()\n    };\n    private FourRefs _dst;\n    [Benchmark]\n    public void BulkStructCopy() =&gt; _dst = _src;\n    private struct FourRefs\n    {\n        public object? A;\n        public object? B;\n        public object? C;\n        public object? D;\n    }\n}</code></pre>\n<p>Both <code>_src</code> and <code>_dst</code> are fields of the same object, so after probing the source address has established that the object isn’t null, probing the destination address can’t provide any additional information. .NET 11 removes that second probe:</p>\n<pre><code>; Arm64\n G_M000_IG02:\n             add     x1, x0, #8\n             ldrsb   wzr, [x1]\n             add     x0, x0, #40\n-            ldrsb   wzr, [x0]\n             movz    x2, ...\n             ldr     x3, [x2]\n             mov     x2, #32\n             blr     x3      // CORINFO_HELP_BULK_WRITEBARRIER\n-; Total bytes of code 56\n+; Total bytes of code 52</code></pre>\n<p>Additionally, dotnet/runtime#125215 lets the JIT retain and efficiently find more assertions in larger methods, increasing the opportunities for the same kinds of simplification. And dotnet/runtime#129312 removes unnecessary temporary variables when the same simple field address is used multiple times, enabling more efficient loads and stores.</p>\n<h3 id=\"simplification\">Simplification</h3>\n<p>Assertion propagation is largely about proving things to help the generated code. Once the JIT knows enough about an operation’s inputs, it can often replace the operation with something simpler and cheaper.</p>\n<p>“Constant folding” is a fancy way of saying the compiler does work once so it doesn’t need to be repeated at run time. If the compiler has everything it needs to compute an answer when building, it can bake that answer in to the generated code and avoid needing the code to re-compute it. That answer can then be further used by other computations at build time, potentially folding further. The C# compiler handles constant folding expressions composed entirely of language constants, while the JIT compiler can go further after inlining and after learning things about values and control flow. The JIT already does a ton of folding, and as with every release, it goes further in .NET 11.</p>\n<p>One straightforward example is the offset of a field within a struct. dotnet/runtime#122297 recognizes more cases where two addresses refer to the same struct and replaces their difference with the known field offset. Here, the second <code>int</code> field begins four bytes into the struct:</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Runtime.CompilerServices;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic unsafe class Benchmarks\n{\n    private struct MyStruct\n    {\n        public int A;\n        public int Field;\n    }\n    [MethodImpl(MethodImplOptions.AggressiveInlining)]\n    private static nint OffsetOfFieldInline()\n    {\n        MyStruct dummy;\n        return (nint)((byte*)&amp;dummy.Field - (byte*)&amp;dummy);\n    }\n    [Benchmark]\n    [Arguments(1_000)]\n    public nint OffsetOfFieldLoop(int n)\n    {\n        nint sum = 0;\n        for (int i = 0; i &lt; n; i++)\n            sum += OffsetOfFieldInline();\n        return sum;\n    }\n}</code></pre>\n<p>Without the fold, the loop repeatedly computes the field offset. With the fold, each iteration simply adds the constant <code>4</code>.</p>\n<pre><code>; Arm64\n--- .NET 10\n+++ .NET 11\n@@ -1,7 +1,6 @@\n G_M000_IG01:\n-            stp     fp, lr, [sp, #-0x20]!\n+            stp     fp, lr, [sp, #-0x10]!\n             mov     fp, sp\n-            str     xzr, [fp, #0x18]\n G_M000_IG02:\n             mov     x0, xzr\n@@ -9,22 +8,18 @@\n             ble     G_M000_IG05\n G_M000_IG03:\n-            add     x2, fp, #0x1C\n-            add     x3, fp, #24\n-            sub     x2, x2, x3\n             align   [0 bytes for IG04]\n             align   [0 bytes]\n             align   [0 bytes]\n             align   [0 bytes]\n G_M000_IG04:\n-            str     xzr, [fp, #0x18]\n-            add     x0, x2, x0\n+            add     x0, x0, #4\n             sub     w1, w1, #1\n             cbnz    w1, G_M000_IG04\n G_M000_IG05:\n-            ldp     fp, lr, [sp], #0x20\n+            ldp     fp, lr, [sp], #0x10\n             ret     lr\n-; Total bytes of code 60\n+; Total bytes of code 40</code></pre>\n<p>dotnet/runtime#121985 from @hez2010 enables the JIT to evaluate <code>SequenceEqual</code> at compile time when both inputs are known. <code>SequenceEqual</code> normally walks two sequences element by element, stopping at the first mismatch. But if inlining exposes both sequences as constants, there’s nothing useful left to do at run time: the JIT can compare them while compiling and replace the whole operation with a constant <code>true</code> or <code>false</code>. This intrinsic underpins APIs including <code>MemoryExtensions.SequenceEqual</code>, <code>ReadOnlySpan&lt;T&gt;.SequenceEqual</code>, and <code>string.Equals</code>.</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private static string AlphaLower =&gt; &quot;abcdefghijklmnopqrstuvwxyz&quot;;\n    private static string AlphaUpper =&gt; &quot;ABCDEFGHIJKLMNOPQRSTUVWXYZ&quot;;\n    [Benchmark]\n    public bool CompareEqual() =&gt; AlphaLower.Equals(AlphaLower);\n    [Benchmark]\n    public bool CompareDistinct() =&gt; AlphaLower.Equals(AlphaUpper);\n}</code></pre>\n<p>Because these properties aren’t <code>const</code>, the C# compiler can’t evaluate the comparisons. The JIT, however, can see the string literals after inlining. It now folds comparisons of the same input whose contents are available at the time of compilation. <code>CompareDistinct</code> therefore becomes a constant <code>false</code>.</p>\n<pre><code>; x64\n--- .NET 10\n+++ .NET 11\n-mov       rax,LOWER_STRING\n-mov       rcx,UPPER_STRING\n-add       rax,0C\n-vmovups   ymm0,[rax]\n-vmovups   ymm1,[rax+14]\n-vmovups   ymm2,[rcx]\n-vpxor     ymm0,ymm2,ymm0\n-vpxor     ymm1,ymm1,[rcx+14]\n-vpor      ymm0,ymm1,ymm0\n-vptest    ymm0,ymm0\n-sete      al\n-movzx     eax,al\n-vzeroupper\n+xor       eax,eax\n ret\n-; Total bytes of code 65\n+; Total bytes of code 3</code></pre>\n<p>Folding an operation is only the first step, though. The result can then simplify later code, even when it’s a vector. dotnet/runtime#127124 extends assertion propagation to 128-bit integer vector constants. If a branch establishes that a vector is zero, uses of that vector within the branch can now be replaced with zero and simplified just like scalar values.</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Runtime.CompilerServices;\nusing System.Runtime.Intrinsics;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private int _selector;\n    [MethodImpl(MethodImplOptions.NoInlining)]\n    private Vector128&lt;int&gt; Compute() =&gt; _selector == 0 ? Vector128&lt;int&gt;.Zero : Vector128.Create(7);\n    [Benchmark]\n    public int AndNotIfZero()\n    {\n        Vector128&lt;int&gt; v = Compute();\n        if (v == Vector128&lt;int&gt;.Zero)\n        {\n            Vector128&lt;int&gt; masked = Vector128.AndNot(v, Vector128.Create(0x00FF00FF));\n            return masked[0];\n        }\n        return -1;\n    }\n}</code></pre>\n<p>In the benchmark’s zero branch, the JIT can now fold away the mask creation, <code>AndNot</code>, and lane extraction, reducing the Arm64 method from 68 bytes to 56 bytes. This currently applies to integer vectors up to 128 bits (floating-point equality has additional NaN and signed-zero semantics that prevent the same reasoning at present).</p>\n<pre><code>; Arm64\n--- .NET 10\n+++ .NET 11\n@@ -11,14 +11,11 @@\n             umaxp   v16.4s, v0.4s, v0.4s\n             umov    x0, v16.d[0]\n             movn    w1, #0\n-            movi    v16.8h, #0xFF,  LSL #8\n-            and     v16.4s, v0.4s, v16.4s\n-            smov    x2, v16.s[0]\n             cmp     x0, #0\n-            csel    w0, w1, w2, ne\n+            cinc    w0, w1, eq\nG_M000_IG03:\n             ldp     fp, lr, [sp], #0x10\n             ret     lr\n-; Total bytes of code 68\n+; Total bytes of code 56</code></pre>\n<p>Two backend cleanups take advantage of simpler expressions. dotnet/runtime#124332 from @jonathandavies-arm removes an unnecessary negation when Arm64 code compares a negated value with zero. And dotnet/runtime#124642 from @yykkibbb lets short-circuit Boolean returns fold even when inlining has left unused writes in the same block; those stores previously obscured the simple Boolean expression from the optimizer.</p>\n<p>Branches offer another opportunity for simplification. Modern processors work on several instructions at different stages at the same time. When a processor encounters a conditional branch, it predicts which path will be taken so that it can continue fetching and executing instructions speculatively. A correct prediction hides much of the branch’s cost. A misprediction throws away that speculative work, redirects instruction fetch to the correct path, and refills the processor’s execution pipeline. That can make the predictability of a branch as important as the work in either branch. The JIT can sometimes avoid that variability, particularly inside small hot loops, by replacing a branch with a conditional move instruction or by recognizing that several branches describe one simpler condition. This isn’t always profitable: branchless code may evaluate work that a predictable branch would skip, making the branching code less expensive in the majority case. But it can be valuable for small, data-dependent choices.</p>\n<p>dotnet/runtime#124567\nrecognizes zero-based equality chains, e.g.\n<code>value == 0 || value == 1 || value == 2</code>. Such chains can be replaced with an\nunsigned range check, e.g. <code>(uint)value &lt;= 2</code>, producing a branchless result.\nThe unsigned comparison also handles negative inputs: when interpreted as\nunsigned, any negative <code>int</code> is larger than the upper bound.</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser]\n[HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private int _value = 2;\n    [Benchmark]\n    public bool IsLetterCategory() =&gt;\n        _value == 0 ||\n        _value == 1 ||\n        _value == 2 ||\n        _value == 3 ||\n        _value == 4;\n}</code></pre>\n<p>The .NET 10 JIT already combines the first four comparisons, but still needs\na branch and a separate comparison for <code>4</code>:</p>\n<pre><code>; x64\nmov       ecx,[rcx+8]\ncmp       ecx,3\nja        CHECK_FOUR\nmov       eax,1\nret\nCHECK_FOUR:\ncmp       ecx,4\nsete      al\nmovzx     eax,al\nret</code></pre>\n<p>.NET 11 recognizes the whole chain as one unsigned range check, reducing the method from 24 bytes to 13:</p>\n<pre><code>; x64\nmov       eax,[rcx+8]\ncmp       eax,5\nsetb      al\nmovzx     eax,al\nret</code></pre>\n<p>dotnet/runtime#128524 from @BoyBaykiller extends the same optimization to contiguous ranges that don’t start at zero. For example, <code>x == 3 || x == 4 || x == 5</code> can become <code>(uint)(x - 3) &lt;= 2</code>.</p>\n<p>Casts can obscure an equally simple comparison. dotnet/runtime#128091 from @BoyBaykiller broadens cast-comparison optimization to equality and inequality. In this benchmark, converting a <code>uint</code> to <code>ulong</code> adds no information needed to compare it with <code>uint.MaxValue</code>, so the JIT can keep the comparison at 32 bits:</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Linq;\nusing System.Runtime.CompilerServices;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private readonly int[] _values = Enumerable.Range(0, 128).ToArray();\n    [Benchmark]\n    public int CastEquality()\n    {\n        int matches = 0;\n        foreach (int value in _values)\n            if ((ulong)(uint)value == uint.MaxValue)\n                matches++;\n        return matches;\n    }\n}</code></pre>\n<p>The widening cast disappears, reducing the Arm64 method from 80 bytes to 76 bytes.</p>\n<pre><code>; Arm64\n--- .NET 10\n+++ .NET 11\n@@ -18,8 +18,7 @@\n G_M000_IG04:\n             ldr     w3, [x0]\n-            mov     x4, #0xFFFFFFFF\n-            cmp     x3, x4\n+            cmn     w3, #1\n             beq     G_M000_IG08\n G_M000_IG05:\n@@ -38,4 +37,4 @@\n             add     w1, w1, #1\n             b       G_M000_IG05\n-; Total bytes of code 80\n+; Total bytes of code 76</code></pre>\n<p>The examples thus far simplify individual comparisons. dotnet/runtime#127181 also combines multiple comparisons in the same expression. For example, <code>(x &gt;= c) &amp;&amp; (x &lt;= c)</code> can only be true when <code>x == c</code>; corresponding OR forms can be simplified similarly.</p>\n<p>Once the JIT can reason about one comparison in terms of another, it can apply the same idea across branches. dotnet/runtime#126587 removes an earlier test when a later, stronger test subsumes it. For example, <code>if (x &gt; 0) if (x &gt; 1)</code> needs only the <code>x &gt; 1</code> test, as reaching the nested body with <code>x &gt; 1</code> necessarily also means <code>x &gt; 0</code>.</p>\n<p>Rather than simply removing a test, the JIT can sometimes use the outcome of an earlier branch to choose the destination of a later one. This is known as “jump threading”: the JIT threads a control-flow path through the intervening jumps directly to its eventual destination. For example, consider:</p>\n<pre><code>int value = condition ? 1 : 2;\nif (value == 1)\n{\n    One();\n}\nelse\n{\n    Two();\n}</code></pre>\n<p>The path where <code>condition</code> is true can go directly to <code>One</code>, while the false path can go directly to <code>Two</code>, eliminating the second test, effectively:</p>\n<pre><code>int value;\nif (condition)\n{\n    value = 1;\n    One();\n}\nelse\n{\n    value = 2;\n    Two();\n}</code></pre>\n<p>dotnet/runtime#126812 lets this continue through more places where paths rejoin, and dotnet/runtime#127103 ensures the rewritten values remain correct in more of those cases. dotnet/runtime#127950 carries relationships between values further, so facts like <code>a &gt; 10</code> and <code>b &gt; a</code> can simplify later branches or bounds. The same reasoning can apply to type information. dotnet/runtime#128500 combines the known types of instances arriving from multiple paths; if every value derives from the tested base type, the JIT can remove the <code>is</code> test after the paths merge. And dotnet/runtime#127434 from @hez2010 lets redundant-branch elimination look through empty jump blocks. Such a block contains no work of its own and exists only to redirect control elsewhere, but it could still hide the relationship between two conditions from the optimizer. Consider this benchmark:</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Runtime.CompilerServices;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private static object? s_sink;\n    private int _x = 20;\n    private int _y = 30;\n    private bool _flag = true;\n    private int _count = 5;\n    [Benchmark]\n    public bool TransitiveComparison() =&gt; TransitiveComparison(_x, _y);\n    [Benchmark]\n    public bool MergedTypeCheck() =&gt; MergedTypeCheck(_flag);\n    [Benchmark]\n    public int NestedThresholds() =&gt; NestedThresholds(_count);\n    [MethodImpl(MethodImplOptions.NoInlining)]\n    private static bool TransitiveComparison(int x, int y)\n    {\n        if (x &gt; 10 &amp;&amp; x &lt; 100 &amp;&amp; y &gt; x)\n            return y &gt; 0;\n        return false;\n    }\n    [MethodImpl(MethodImplOptions.NoInlining)]\n    private static bool MergedTypeCheck(bool flag)\n    {\n        object shape = flag ? new Circle() : new Rectangle();\n        s_sink = shape;\n        return shape is Shape;\n    }\n    [MethodImpl(MethodImplOptions.NoInlining)]\n    private static int NestedThresholds(int count)\n    {\n        if (count &gt; 1)\n            if (count &gt; 2)\n                if (count &gt; 3)\n                    if (count &gt; 4)\n                        return 1;\n        return 3;\n    }\n    private abstract class Shape;\n    private sealed class Circle : Shape;\n    private sealed class Rectangle : Shape;\n}</code></pre>\n<p>In <code>TransitiveComparison</code>, reaching <code>y &gt; 0</code> means the JIT already knows that <code>x &gt; 10</code> and <code>y &gt; x</code>, which together prove that <code>y</code> is positive. The final comparison disappears, reducing the Arm64 method from 40 bytes to 36 bytes:</p>\n<pre><code>; Arm64\n--- .NET 10\n+++ .NET 11\n             cmp     w1, w0\n             ccmp    w2, w3, c, gt\n-            ccmp    w1, #0, nzc, ls\n-            cset    x0, gt\n+            cset    x0, ls\n-; Total bytes of code 40\n+; Total bytes of code 36</code></pre>\n<p>In <code>MergedTypeCheck</code>, each path creates a different concrete type, but both derive from <code>Shape</code>. .NET 11 keeps the allocations and the store that make the example observable, but replaces the <code>is</code> helper call and its result test with the constant <code>true</code>, reducing the method from 112 bytes to 88 bytes:</p>\n<pre><code>; Arm64\n--- .NET 10\n+++ .NET 11\n             bl      CORINFO_HELP_ASSIGN_REF\n-            movz    x0, #0xEA30\n-            movk    x0, #0x4EB LSL #16\n-            movk    x0, #0x7FFF LSL #32\n-            bl      CORINFO_HELP_ISINSTANCEOFCLASS\n-            cmp     x0, #0\n-            cset    x0, ne\n+            mov     w0, #1\n-; Total bytes of code 112\n+; Total bytes of code 88</code></pre>\n<p>For <code>NestedThresholds</code>, reaching the <code>return 1</code> requires <code>count</code> to be greater than all four constants, which is equivalent to just <code>count &gt; 4</code>. Once redundant-branch elimination can see through the empty jump blocks left behind while simplifying the nested conditions, the other three comparisons disappear:</p>\n<pre><code>; Arm64\n--- .NET 10\n+++ .NET 11\n             mov     w1, #3\n             mov     w2, #1\n-            cmp     w0, #1\n-            ccmp    w0, #2, nzc, gt\n-            ccmp    w0, #3, nzc, gt\n-            ccmp    w0, #4, nzc, gt\n+            cmp     w0, #4\n             csel    w0, w1, w2, le\n-; Total bytes of code 44\n+; Total bytes of code 32</code></pre>\n<p>Removing a redundant branch is ideal; why do work when it’s provably unnecessary? Often, however, the branch is necessary, as both outcomes are possible (or at least not provably impossible). In such cases, the JIT may still be able to avoid branching via specialized instructions that bake the choice into the instruction. “If-conversion” replaces a small <code>if</code>/<code>else</code> with a conditional-move instruction or another branchless form when both alternatives are cheap. The JIT has been able to do this for several releases, and improves in .NET 11. dotnet/runtime#124738 from @BoyBaykiller recognizes an earlier default assignment as the implicit <code>else</code>, so <code>bool x = false; if (cond) x = true;</code> can become the same branchless form as an explicit <code>else</code>. dotnet/runtime#127915 from @BoyBaykiller handles the opposite cleanup, removing a conditional selection when both outcomes are the same constant while preserving any side effects from evaluating the condition. dotnet/runtime#128533 from @BoyBaykiller also helps these Boolean optimizations meet in the middle by normalizing power-of-two bit tests. A power of two has exactly one bit set, in which case <code>(A &amp; bit) == bit</code> is equivalent to <code>(A &amp; bit) != 0</code>; putting both forms into the same canonical representation makes them easier to combine with surrounding conditions. All three improvements are visible in the following benchmarks:</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Runtime.CompilerServices;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private int _left = 1;\n    private int _right = 2;\n    private double _double = 0.0;\n    private int _bits = 4;\n    [Benchmark]\n    public bool ImplicitElse() =&gt; ImplicitElse(_left, _right);\n    [Benchmark]\n    public bool IsDefaultValue() =&gt; IsDefaultValue(_double);\n    [Benchmark]\n    public bool HasEitherBit() =&gt; HasEitherBit(_bits);\n    [MethodImpl(MethodImplOptions.NoInlining)]\n    private static bool ImplicitElse(int left, int right)\n    {\n        bool leftIsSmaller = false;\n        if (left &lt; right)\n            leftIsSmaller = true;\n        return leftIsSmaller;\n    }\n    [MethodImpl(MethodImplOptions.NoInlining)]\n    private static bool IsDefaultValue(double value) =&gt; 0.0.Equals(value);\n    [MethodImpl(MethodImplOptions.NoInlining)]\n    private static bool HasEitherBit(int value) =&gt;\n        ((value &amp; 4) == 4) || ((value &amp; 8) == 8);\n}</code></pre>\n<p>For <code>ImplicitElse</code>, .NET 10 already avoids a branch, but it still materializes both Boolean values and selects between them. In .NET 11, the method becomes just the comparison and a <code>cset</code>, shrinking from 36 bytes to 24 bytes:</p>\n<pre><code>; Arm64\n--- .NET 10\n+++ .NET 11\n-            mov     w2, wzr\n-            mov     w3, #1\n             cmp     w0, w1\n-            csel    w2, w2, w3, ge\n-            mov     w0, w2\n+            cset    x0, lt\n-; Total bytes of code 36\n+; Total bytes of code 24</code></pre>\n<p><code>0.0.Equals(value)</code> needs to account for <code>NaN</code>, but because the left operand is zero, the case where both operands are <code>NaN</code> can never apply. Removing the conditional selection for that case leaves one floating-point comparison and one <code>cset</code>, reducing <code>IsDefaultValue</code> from 40 bytes to 24 bytes:</p>\n<pre><code>; Arm64\n--- .NET 10\n+++ .NET 11\n             fcmp    d0, #0.0\n-            beq     G_M000_IG04\n-\n-G_M000_IG03:\n-            fcmp    d0, d0\n-            csel    w0, wzr, wzr, eq\n-            b       G_M000_IG05\n-\n-G_M000_IG04:\n-            mov     w0, #1\n-\n-G_M000_IG05:\n+            cset    x0, eq\n+\n+G_M000_IG03:\n             ldp     fp, lr, [sp], #0x10\n             ret     lr\n-; Total bytes of code 40\n+; Total bytes of code 24</code></pre>\n<p>Finally, normalizing both power-of-two comparisons lets the JIT combine their results. The short-circuit branch in <code>HasEitherBit</code> is replaced by two masks and an <code>or</code>, reducing the method from 40 bytes to 36 bytes:</p>\n<pre><code>; Arm64\n--- .NET 10\n+++ .NET 11\n-            tbz     w0, #2, G_M000_IG05\n-\n-G_M000_IG03:\n-            mov     w0, #1\n-\n-G_M000_IG04:\n-            ldp     fp, lr, [sp], #0x10\n-            ret     lr\n-\n-G_M000_IG05:\n-            tst     w0, #8\n+            and     w1, w0, #4\n+            and     w0, w0, #8\n+            orr     w0, w1, w0\n+            cmp     w0, #0\n             cset    x0, ne\n-G_M000_IG06:\n+G_M000_IG03:\n             ldp     fp, lr, [sp], #0x10\n             ret     lr\n-; Total bytes of code 40\n+; Total bytes of code 36</code></pre>\n<p>Not every simplification depends on broader control-flow reasoning.\n“Peephole optimizations” instead replace a short, recognizable pattern with an\nequivalent cheaper one. Each may save only an instruction or expose a form\nthat another optimization understands, but these patterns can occur very\nfrequently on hot paths throughout generated code. For example, dotnet/runtime#126529 from @BoyBaykiller recognizes that <code>255 - x</code> for a <code>byte</code> is equivalent to <code>x ^ 255</code>: both simply flip all eight bits, but the latter can remove an instruction if it’s able to replace a negation and add with an xor. Similarly, <code>-1 - x</code> can turn into the equivalent of <code>~x</code>.</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private readonly byte[] _data = new byte[4096];\n    [GlobalSetup]\n    public void Setup() =&gt; new Random(42).NextBytes(_data);\n    [Benchmark]\n    public int InvertBytes()\n    {\n        int sum = 0;\n        foreach (byte b in _data) sum += 255 - b;\n        return sum;\n    }\n}</code></pre>\n<p>In .NET 11, the loop loses a separate negate and add:</p>\n<pre><code>; Arm64\n--- .NET 10\n+++ .NET 11\n@@ -19,9 +19,8 @@\n G_M000_IG04:\n             ldrb    w4, [x0, w2, UXTW]\n-            neg     w4, w4\n+            eor     w4, w4, #255\n             add     w1, w4, w1\n-            add     w1, w1, #255\n             add     w2, w2, #1\n             cmp     w3, w2\n             bgt     G_M000_IG04\n@@ -33,4 +32,4 @@\n             ldp     fp, lr, [sp], #0x10\n             ret     lr\n-; Total bytes of code 76\n+; Total bytes of code 72</code></pre>\n<p>dotnet/runtime#129361 removes another unnecessary instruction when comparing an <code>sbyte</code> with a constant that fits in eight bits. The JIT can compare the byte directly, with no sign extension:</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Runtime.CompilerServices;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private sbyte _value = -65;\n    [Benchmark]\n    public bool IsLow() =&gt; IsLow(_value);\n    [MethodImpl(MethodImplOptions.NoInlining)]\n    private static bool IsLow(sbyte value) =&gt; value &lt; -64;\n}</code></pre>\n<p>The optimized codegen then compares the byte directly, removing the <code>movsx</code> sign-extension instruction (though the JIT still retains it in the few comparison forms that require a full-width sign bit for correctness).</p>\n<pre><code>; x64\n--- .NET 10\n+++ .NET 11\n-movsx  rax,cl\n-cmp    eax,0FFFFFFC0\n+cmp    cl,0C0\n setl   al\n movzx  eax,al\n ret\n-; Total bytes of code 14\n+; Total bytes of code 10</code></pre>\n<p>dotnet/runtime#125180 from @saucecontrol improves non-overflowing <code>float</code> and <code>double</code> conversions to <code>long</code> and <code>ulong</code> on x86 machines with AVX-512 or AVX10.2. These casts have defined behavior for NaN and out-of-range values, so older code used a helper to preserve those semantics. The newer instruction set lets the JIT keep the normal path inline and register-based, avoiding the helper call; machines that don’t support these instructions retain the existing fallback.</p>\n<pre><code>// Run with 32-bit x86 dotnet on a machine with AVX-512 or AVX10.2:\n// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private float _single = 123_456.75f;\n    private double _double = 123_456.75;\n    [Benchmark] public long SingleToInt64() =&gt; (long)_single;\n    [Benchmark] public ulong SingleToUInt64() =&gt; (ulong)_single;\n    [Benchmark] public long DoubleToInt64() =&gt; (long)_double;\n    [Benchmark] public ulong DoubleToUInt64() =&gt; (ulong)_double;\n}</code></pre>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Mean</th><th>Ratio</th><th>Code Size</th></tr></thead><tbody><tr><td>SingleToInt64</td><td>.NET 10.0</td><td>5.103 ns</td><td>1.00</td><td>31 B</td></tr><tr><td>SingleToInt64</td><td>.NET 11.0</td><td>2.093 ns</td><td>0.41</td><td>54 B</td></tr><tr><td>SingleToUInt64</td><td>.NET 10.0</td><td>4.810 ns</td><td>1.00</td><td>31 B</td></tr><tr><td>SingleToUInt64</td><td>.NET 11.0</td><td>1.366 ns</td><td>0.28</td><td>34 B</td></tr><tr><td>DoubleToInt64</td><td>.NET 10.0</td><td>4.834 ns</td><td>1.00</td><td>31 B</td></tr><tr><td>DoubleToInt64</td><td>.NET 11.0</td><td>2.101 ns</td><td>0.43</td><td>54 B</td></tr><tr><td>DoubleToUInt64</td><td>.NET 10.0</td><td>4.663 ns</td><td>1.00</td><td>31 B</td></tr><tr><td>DoubleToUInt64</td><td>.NET 11.0</td><td>1.363 ns</td><td>0.29</td><td>34 B</td></tr></tbody></table></div>\n<h3 id=\"vectorization\">Vectorization</h3>\n<p>SIMD, or “single instruction, multiple data”, is the concept of one instruction applying the same operation to several values at once. A “scalar” <code>add</code>, for example, might combine one pair of 32-bit integers, while a 128-bit SIMD <code>add</code> can combine “vectors” of four pairs in the same instruction; 256- and 512-bit variants can handle vectors of eight and sixteen pairs, respectively. When the iterations of an operation are independent, “vectorizing” a loop can therefore replace several scalar iterations with one, improving the throughput of the loop significantly.</p>\n<p>.NET exposes portable (they work on any machine) variable-width vector type <code>Vector&lt;T&gt;</code> (which can represent different counts of <code>T</code> depending on the current hardware), fixed-width <code>Vector64&lt;T&gt;</code> through <code>Vector512&lt;T&gt;</code> types (which always represent the same count of <code>T</code>), and architecture-specific intrinsics (performing operations on such vector types which the JIT then maps to the right underlying hardware instructions). Each element in a vector is often referred to as a “lane”. Because the JIT recognizes these operations directly, it can fold constants, select instructions, and remove unsupported paths without treating them as normal method calls.</p>\n<p>A variety of PRs in .NET 11 improve AVX-512 broadcasting and masking. Embedded broadcasting lets an instruction load a single scalar value and replicate it across all vector lanes, avoiding the need to materialize a full-width vector constant in memory to feed into the instruction. For example, this bitwise AND instruction:</p>\n<pre><code>; x64\nvpandd  zmm0, zmm1, dword ptr [reloc @RWD00] {1to16}</code></pre>\n<p>can replace this one:</p>\n<pre><code>; x64\nvpandd  zmm0, zmm1, zmmword ptr [reloc @RWD00]</code></pre>\n<p>storing only 4 bytes in the read-only data section rather than 64. Because the broadcast is handled as part of the load, there’s no additional instruction-level latency; the primary benefit is reduced data size and cache footprint.</p>\n<p>Embedded masking similarly lets an instruction update only a subset of the lanes. A mask is one bit per vector lane, where each bit indicates whether and how the operation should affect the corresponding lane. Without embedded masking, code often needs to compute every lane and then blend that result with the old value, so folding the mask into the operation can remove both the separate blend and a zero-vector setup. dotnet/runtime#117700 from @saucecontrol improves broadcast selection when an intrinsic’s natural element size differs from its managed vector type. VNNI, the Vector Neural Network Instructions used for small-integer multiply-accumulate operations, and bitwise operations can now use the smallest valid repeated constant, avoiding a full-vector load.</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\n// Requires AVX-VNNI and AVX-512F.\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Runtime.Intrinsics;\nusing System.Runtime.Intrinsics.X86;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private readonly Vector128&lt;byte&gt; _bytes = Vector128.Create((byte)1);\n    private readonly Vector128&lt;ulong&gt; _u64 = Vector128.Create(1UL);\n    private readonly Vector128&lt;uint&gt; _u32 = Vector128.Create(1U);\n    private readonly Vector512&lt;int&gt; _v512 = Vector512.Create(1);\n    private int _n = 42;\n    [Benchmark]\n    public Vector128&lt;int&gt; VnniBroadcast() =&gt;\n        AvxVnni.MultiplyWideningAndAdd(\n            Vector128&lt;int&gt;.Zero, _bytes, Vector128&lt;sbyte&gt;.One);\n    [Benchmark]\n    public Vector128&lt;uint&gt; MaskAnd() =&gt;\n        Vector128.ConditionalSelect(\n            Vector128.GreaterThan(_u32, Vector128&lt;uint&gt;.Zero),\n            (_u64 &amp; Vector128&lt;uint&gt;.One.AsUInt64()).AsUInt32(),\n            Vector128&lt;uint&gt;.Zero);\n    [Benchmark]\n    public Vector512&lt;int&gt; BlendMaskAllOnes() =&gt;\n        Avx512F.BlendVariable(\n            Vector512.Create(_n),\n            _v512,\n            Vector512.Create(-1));\n    [Benchmark]\n    public Vector512&lt;int&gt; MultiInsert() =&gt;\n        Vector512.ConditionalSelect(\n            Vector512.Create(0, -1, 0, 0, 0, -1, 0, 0, 0, -1, 0, 0, 0, -1, 0, 0),\n            _v512,\n            Vector512.Create(_n));\n    [Benchmark]\n    public Vector512&lt;int&gt; MultiInsertZero() =&gt;\n        Avx512F.BlendVariable(\n            _v512,\n            Vector512&lt;int&gt;.Zero,\n            Vector512.Create(0, -1, 0, 0, 0, -1, 0, 0, 0, -1, 0, 0, 0, -1, 0, 0));\n}</code></pre>\n<p>In .NET 11, this results in 12 fewer bytes in the read-only data section, and 12 fewer bytes of constant-pool cache footprint.</p>\n<pre><code>; x64\n-C4E279503500000000   vpdpbusd xmm6, xmm0, xmmword ptr [reloc @RWD00]\n+62F27D18503500000000 vpdpbusd xmm6, xmm0, dword ptr [reloc @RWD00] {1to4}\n-RWD00  dq 0101010101010101h, 0101010101010101h\n+RWD00  dd 01010101h</code></pre>\n<p>The fix also impacts embedded masking. For example, with <code>MaskAnd</code> previously, the AND used a qword broadcast, <code>{1to2}</code>, and a separate blend then moved the masked result, meaning two instructions. Now that <code>Vector128&lt;uint&gt;.One</code> can be broadcast at dword granularity, the mask’s element size and the AND’s element size agree, unlocking using the single merged-masked form. This pattern shows up throughout vectorized algorithms that do lots of bitwise manipulation and hashing, including implementations in System.Numerics.Tensors, System.IO.Hashing, and System.Private.CoreLib.</p>\n<pre><code>; x64\n-       vpandq   xmm0, xmm0, qword ptr [reloc @RWD00] {1to2}\n-       vpblendmd xmm0 {k1}{z}, xmm0, xmm0\n+       vpandd   xmm0 {k1}{z}, xmm0, dword ptr [reloc @RWD00] {1to4}\n; Code: 45 → 39 bytes; data: 8 bytes → 4 bytes</code></pre>\n<p>That <code>MaskAnd</code> example starts as an AND followed by a blend, an operation that\nchooses independently for each vector lane whether to take its value from one\ninput or the other, with the JIT able to fold those two operations together.\nSimilar opportunities arise with blends more generally. Sometimes the mask or\none of the inputs makes the choice trivial, e.g. an all-ones mask always selects the\nsame input, so the blend is just a move. If one input is zero, it can often\nbecome an <code>AND</code> or <code>ANDN</code>. AVX-512 provides more options still, as constant\nmasks and zeroing can be encoded directly in the instruction.\ndotnet/runtime#123146 from\n@saucecontrol makes these simplifications\nconsistently across the portable and hardware-specific APIs. A blend with an all-ones mask provides a particularly clear example:</p>\n<pre><code>// Run on x64 with AVX-512:\n// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing System.Runtime.Intrinsics;\nusing System.Runtime.Intrinsics.X86;\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser]\n[HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private readonly Vector512&lt;int&gt; _values = Vector512.Create(1);\n    private int _n = 42;\n    [GlobalSetup]\n    public void Setup()\n    {\n        if (!Avx512F.IsSupported)\n            throw new PlatformNotSupportedException();\n    }\n    [Benchmark]\n    public Vector512&lt;int&gt; BlendMaskAllOnes() =&gt;\n        Avx512F.BlendVariable(Vector512.Create(_n), _values, Vector512.Create(-1));\n}</code></pre>\n<p>The generated code no longer needs to create the first input, load the mask, or perform the blend. It simply loads the input the all-ones mask would always select:</p>\n<pre><code>; x64\n-vpbroadcastd zmm0, dword ptr [rcx+8]\n-kmovq       k1, qword ptr [RWD00]\n-vpblendmd   zmm0 {k1}, zmm0, [rcx+48]\n+vmovups     zmm0, [rcx+48]\n vmovups     [rdx], zmm0\n mov         rax, rdx\n vzeroupper\n ret\n; 39 bytes → 23 bytes</code></pre>\n<h3 id=\"intrinsics\">Intrinsics</h3>\n<p>An intrinsic is a managed API that the JIT recognizes and special-cases. Often that special-casing involves actually replacing calls to the method with custom code that’s behaviorally equivalent but better in some way (faster, smaller, etc.)</p>\n<p>As an example, dotnet/runtime#128678 improves recognition of generic-math calls to <code>IBinaryNumber&lt;T&gt;.Log2</code>. The method computes the base-2 logarithm of an integer, equivalent to the index of the number’s highest set bit; for example, <code>Log2(16)</code> is <code>4</code>. Previously, the JIT’s normalized integer type lost the signedness needed to import the operation directly as an intrinsic. Inlining the managed implementation could still produce the same optimized code, but when inlining didn’t happen, the managed call remained. In .NET 11, the JIT consults the precise type and imports the operation directly: unsigned and non-negative signed inputs can become leading-zero-count or bit-scan arithmetic, while a negative signed value retains the managed fallback and its exact exception behavior.</p>\n<p>Sometimes the JIT has a perfectly good intrinsic lowering but doesn’t recognize a call that should use it. <code>Enum.Equals</code> from a generic <code>T : Enum</code> context was a good example. Even though both arguments to the generic helper are strongly typed as <code>T</code>, an enum doesn’t provide an <code>Equals(T)</code> method; it inherits the virtual <code>Enum.Equals(object)</code> implementation. The second argument therefore needs to be boxed to pass it as <code>object</code>. The receiver is invoked with a constrained virtual call, but because the concrete enum doesn’t override the method itself, it too needs to be boxed to invoke the implementation on <code>System.Enum</code>. Thus, what looks like a strongly-typed comparison can end up allocating two boxes and making a virtual call. In .NET 11, dotnet/runtime#122779 eliminates this overhead by teaching the JIT to recognize the call and fold it to a direct comparison of the enum’s underlying integer values. For example:</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Runtime.CompilerServices;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[MemoryDiagnoser(false), HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private static readonly StringComparison[] s_values =\n    {\n        StringComparison.Ordinal, StringComparison.OrdinalIgnoreCase, \n        StringComparison.CurrentCulture, StringComparison.CurrentCultureIgnoreCase,\n        StringComparison.InvariantCulture, StringComparison.Ordinal,\n    };\n    [Benchmark]\n    public int CountOrdinal_Generic()\n    {\n        int count = 0;\n        foreach (var v in s_values)\n            if (EqualsGeneric(v, StringComparison.Ordinal))\n                count++;\n        return count;\n    }\n    [MethodImpl(MethodImplOptions.NoInlining)]\n    private static bool EqualsGeneric&lt;T&gt;(T a, T b) where T : Enum =&gt; a.Equals(b);\n}</code></pre>\n<p>Once the JIT knows the callee is <code>Enum.Equals</code> and knows the exact enum type, it asks the runtime for the underlying integer type and replaces the virtual call with a direct comparison. That in turn makes both box/unbox pairs redundant, and the generated code contains neither allocation. For the six comparisons performed here, .NET 10 creates twelve boxes, totaling 288 bytes. In .NET 11, the helper becomes just the integer comparison, eliminating both the allocations and the virtual dispatch.</p>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Mean</th><th>Ratio</th><th>Allocated</th></tr></thead><tbody><tr><td>CountOrdinal_Generic</td><td>.NET 10.0</td><td>59.24 ns</td><td>1.00</td><td>288 B</td></tr><tr><td>CountOrdinal_Generic</td><td>.NET 11.0</td><td>10.01 ns</td><td>0.17</td><td>–</td></tr></tbody></table></div>\n<p>NativeAOT had been carrying an equivalent optimization for years, implemented as IL rewriting in ILCompiler that patches <code>Enum.Equals</code> to use typed comparisons. With the JIT now handling it, including in NativeAOT’s own use of the JIT (NativeAOT uses the JIT ahead of time rather than just in time), dotnet/runtime#123086 deletes that rewriting and its supporting machinery.</p>\n<p>dotnet/runtime#127329 improves the <code>Vector256.Sum</code> and <code>Vector512.Sum</code> intrinsics. The JIT now performs most of the reduction at full width and combines the per-lane results at the end, avoiding the extracts and duplicate shuffle sequences needed when splitting wide vectors into 128-bit pieces. And dotnet/runtime#127402 extends vector-constant propagation from 128-bit vectors to <code>Vector256</code> and <code>Vector512</code>. Code that compares a wide vector with a known sentinel can now simplify subsequent uses just as narrower vectors already could. The following benchmark exemplifies both:</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\n// Requires AVX2 for the assembly shown below.\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Runtime.CompilerServices;\nusing System.Runtime.Intrinsics;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private readonly Vector256&lt;float&gt; _floats =\n        Vector256.Create(1.0f, 2.0f, 3.0f, 4.0f, 5.0f, 6.0f, 7.0f, 8.0f);\n    private int _selector;\n    [Benchmark]\n    public float Sum() =&gt; Vector256.Sum(_floats);\n    [Benchmark]\n    public int TransformWhenKnown()\n    {\n        Vector256&lt;int&gt; value = GetVector();\n        if (value == Vector256.Create(0, 1, 2, 3, 4, 5, 6, 7))\n            return (value + Vector256.Create(10)).GetElement(6);\n        return -1;\n    }\n    [MethodImpl(MethodImplOptions.NoInlining)]\n    private Vector256&lt;int&gt; GetVector() =&gt;\n        _selector == 0 ?\n            Vector256.Create(0, 1, 2, 3, 4, 5, 6, 7) :\n            Vector256.Create(7);\n}</code></pre>\n<p>For <code>Sum</code>, .NET 10 separately reduces each 128-bit half and then adds the two scalar results. In .NET 11, the permutes and adds operate on both halves in parallel as 256-bit instructions, after which only the two already-reduced halves need to be combined:</p>\n<pre><code>; x64\n vmovups   ymm0, [rcx+28]\n-vmovaps   ymm1, ymm0\n-vpermilps xmm2, xmm1, 0B1\n-vaddps    xmm1, xmm2, xmm1\n-vpermilps xmm2, xmm1, 4E\n-vaddps    xmm1, xmm2, xmm1\n-vextractf128 xmm0, ymm0, 1\n-vpermilps xmm2, xmm0, 0B1\n-vaddps    xmm0, xmm2, xmm0\n-vpermilps xmm2, xmm0, 4E\n-vaddps    xmm0, xmm2, xmm0\n-vaddss    xmm0, xmm1, xmm0\n+vpermilps ymm1, ymm0, 0B1\n+vaddps    ymm0, ymm1, ymm0\n+vpermilps ymm1, ymm0, 4E\n+vaddps    ymm0, ymm1, ymm0\n+vextractf128 xmm1, ymm0, 1\n+vaddps    xmm0, xmm1, xmm0\n; 63 bytes → 39 bytes</code></pre>\n<p><code>TransformWhenKnown</code> uses a deliberately non-repeating constant across its eight lanes. On the branch where the comparison succeeds, .NET 11 can replace <code>value</code> with that constant, fold the vector addition, and determine that element 6 is <code>16</code>. The <code>vpaddd</code>, extraction, second 32-byte constant, and associated control flow all disappear:</p>\n<pre><code>; x64\n-cmp      eax, 0FFFFFFFF\n-jne      M00_L00\n-vmovups  ymm0, [rsp+20]\n-vpaddd   ymm0, ymm0, [RWD32]\n-vextracti128 xmm0, ymm0, 1\n-vpextrd  eax, xmm0, 2\n-vzeroupper\n-add      rsp, 58\n-ret\n-\n-M00_L00:\n-mov      eax, 0FFFFFFFF\n+mov      ecx, 0FFFFFFFF\n+mov      edx, 10\n+cmp      eax, 0FFFFFFFF\n+mov      eax, edx\n+cmovne   eax, ecx\n vzeroupper\n add      rsp, 58\n ret\n; 85 bytes → 59 bytes</code></pre>\n<p>One of the goals of .NET is that you can write code once and have it run anywhere, optimized for whatever that “anywhere” has to offer. For vectorization, that means providing portable operations whenever the intent is common across instruction sets, while retaining architecture-specific APIs for algorithms that really do need to target a particular machine.</p>\n<p>Whenever possible, we want to enable developers to express their algorithms using the portable APIs, and each release of .NET fills additional gaps there. Including .NET 11. dotnet/runtime#129627 from @hez2010 adds portable APIs for constructing common lane sequences (e.g. <code>[1, 2, 4, 8]</code> or <code>[a, b, a, b]</code>), concatenating half-vectors (the lower halves of <code>[a, b, c, d]</code> and <code>[w, x, y, z]</code> producing <code>[a, b, w, x]</code>), interleaving (<code>[a, b]</code> and <code>[x, y]</code> producing <code>[a, x, b, y]</code>), de-interleaving (<code>[a, x, b, y]</code> producing <code>[a, b]</code> and <code>[x, y]</code>), and reversal (<code>[a, b, c, d]</code> producing <code>[d, c, b, a]</code>), along with their JIT intrinsification. These operations were already expressible, but only verbosely and only if you knew which hardware instruction to reach for, e.g. writing <code>Zip</code> by hand meant targeting a platform-specific API like <code>AdvSimd.Arm64.ZipLow</code>. The new APIs let the code state the transformation and leave instruction selection to the JIT.</p>\n<p>Once the intrinsic operation has been recognized, the backend still needs to keep it in a useful vector form while assigning registers and selecting instructions. Vector values are structs, and the JIT will often apply “struct promotion,” tracking a struct’s fields as independent locals so that each can be optimized separately. That’s useful for ordinary structs, but counterproductive when a value is meant to remain in a vector or mask register: splitting it can introduce extra moves and obscure what should be a single whole-value store, particularly after inlining introduces more local stores. dotnet/runtime#128013 consistently marks SIMD and mask stores as intrinsic-related across platforms, including 32-bit x86 and x64 mask stores, so those locals remain intact. dotnet/runtime#129563 extends that principle to user-defined structs that are bitcast to SIMD types. This trades away struct promotion for those locals, but enables the JIT to preserve their vector representation.</p>\n<p>This matters for user-defined numerical types that store the same data as a\nhardware vector but expose named fields or domain-specific operations. The\nfollowing <code>Vector2Double</code> is laid out as two adjacent <code>double</code> values, so it can\nbe bitcast to <code>Vector128&lt;double&gt;</code>, operated on with SIMD, and bitcast back:</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing System.Runtime.CompilerServices;\nusing System.Runtime.InteropServices;\nusing System.Runtime.Intrinsics;\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\npublic struct Vector2Double(double x, double y)\n{\n    public double X = x;\n    public double Y = y;\n    public static Vector2Double operator +(Vector2Double left, Vector2Double right)\n    {\n        Vector128&lt;double&gt; simdLeft = Unsafe.BitCast&lt;Vector2Double, Vector128&lt;double&gt;&gt;(left);\n        Vector128&lt;double&gt; simdRight = Unsafe.BitCast&lt;Vector2Double, Vector128&lt;double&gt;&gt;(right);\n        return Unsafe.BitCast&lt;Vector128&lt;double&gt;, Vector2Double&gt;(simdLeft + simdRight);\n    }\n}\n[DisassemblyDiagnoser, HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private readonly Vector2Double _a = new(1.0, 2.0);\n    private readonly Vector2Double _b = new(3.0, 4.0);\n    private readonly Vector2Double _c = new(5.0, 6.0);\n    [Benchmark]\n    public Vector2Double Add() =&gt; Add(_a, _b, _c);\n    [MethodImpl(MethodImplOptions.NoInlining)]\n    private static Vector2Double Add(Vector2Double a, Vector2Double b, Vector2Double c) =&gt;\n        a + b + c;\n}</code></pre>\n<p>In .NET 10, promotion of the intermediate struct sends the first SIMD result\nthrough two stack locations before the second addition. .NET 11 keeps that\nvalue in <code>xmm0</code>, reducing the helper from 52 bytes to 22 bytes:</p>\n<pre><code>; x64\n-sub       rsp, 28\n vmovups   xmm0, [rdx]\n vaddpd    xmm0, xmm0, [r8]\n-vmovaps   [rsp], xmm0\n-vmovups   xmm0, [rsp]\n-vmovups   [rsp+18], xmm0\n-vmovups   xmm0, [rsp+18]\n vaddpd    xmm0, xmm0, [r9]\n vmovups   [rcx], xmm0\n mov       rax, rcx\n-add       rsp, 28\n ret\n; 52 bytes → 22 bytes</code></pre>\n<p>dotnet/runtime#128350 gives the xarch register allocator more freedom around fused multiply-add (FMA) and AVX-512 ternary-logic operations. These instructions can read and overwrite operands in several equivalent arrangements; choosing the arrangement that already matches the surrounding registers avoids otherwise necessary moves.</p>\n<p>Generic vector code introduces another wrinkle. Operators like\n<code>Vector128&lt;T&gt;.operator ==</code> return <code>bool</code>, so the return type doesn’t reveal the\nvector’s element type. The JIT instead needs to obtain that type from the\noperands in order to select the right comparison instruction. In some generic\ncontexts, including helpers built on the internal <code>ISimdVector</code> abstraction,\nthe JIT was consulting the wrong type information and failed to import the\noperator as an intrinsic. It then executed the managed fallback, which compares\nthe lanes individually. dotnet/runtime#130086 marks these operators so their element type is taken from the first argument. As an example, the generic helpers used internally by ordinal-ignore-case string comparer benefit from this.</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private readonly string _lower = new(&#39;a&#39;, 256);\n    private readonly string _upper = new(&#39;A&#39;, 256);\n    [Benchmark]\n    public bool OrdinalIgnoreCase() =&gt; string.Equals(_lower, _upper, StringComparison.OrdinalIgnoreCase);\n}</code></pre>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Mean</th><th>Ratio</th></tr></thead><tbody><tr><td>OrdinalIgnoreCase</td><td>.NET 10.0</td><td>27.794 ns</td><td>1.00</td></tr><tr><td>OrdinalIgnoreCase</td><td>.NET 11.0</td><td>21.861 ns</td><td>0.79</td></tr></tbody></table></div>\n<p>.NET 11 adds support for newer x86 capabilities while also\nimproving code generated for existing hardware. These changes benefit both\ndirect users of hardware intrinsics and portable vector code selected by the\nJIT. For example, dotnet/runtime#124114 from @saucecontrol improves 32-bit x86 without AVX-512, where converting <code>uint</code> to <code>float</code> or <code>double</code> previously required a runtime helper. Older x86 conversion instructions accept signed integers, and half of the <code>uint</code> range doesn’t fit in a signed 32-bit value, which is why the helper existed. The JIT now emits an inline vector-instruction sequence that handles the high bit explicitly, avoiding the call and its register and stack overhead.</p>\n<pre><code>// Run with 32-bit x86 dotnet and AVX-512 disabled (DOTNET_EnableAVX512=0)\n// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private uint _value = 0xF123_4567;\n    [Benchmark] public float UInt32ToSingle() =&gt; _value;\n    [Benchmark] public double UInt32ToDouble() =&gt; _value;\n}</code></pre>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Mean</th><th>Ratio</th><th>Code Size</th></tr></thead><tbody><tr><td>UInt32ToSingle</td><td>.NET 10.0</td><td>4.752 ns</td><td>1.00</td><td>37 B</td></tr><tr><td>UInt32ToSingle</td><td>.NET 11.0</td><td>2.403 ns</td><td>0.51</td><td>43 B</td></tr><tr><td>UInt32ToDouble</td><td>.NET 10.0</td><td>4.727 ns</td><td>1.00</td><td>39 B</td></tr><tr><td>UInt32ToDouble</td><td>.NET 11.0</td><td>2.402 ns</td><td>0.51</td><td>41 B</td></tr></tbody></table></div>\n<p>dotnet/runtime#124804 from @alexcovington adds the AVX-512 Bit Matrix Multiply APIs. A binary matrix treats each bit as an element and combines rows and columns with bitwise operations, not integer multiplication. The instructions are useful in areas such as error correction and CRC computation. Each replaces a much longer sequence of shifts, masks, and exclusive-ORs. And dotnet/runtime#128365 from @jamesburton adds <code>AvxVnni.V512</code>, extending the AVX-VNNI APIs from 256-bit to 512-bit operands so the small-integer dot products used by quantized machine-learning models can process 64 bytes per operation instead of 32.</p>\n<p>dotnet/runtime#126062 from @saucecontrol also avoids converting a vector selector into an AVX-512 mask register when the eventual operation still needs the vector form. In such cases, the older-looking vector blend is actually shorter and uses fewer resources:</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Runtime.Intrinsics;\nusing System.Runtime.Intrinsics.X86;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private readonly Vector128&lt;float&gt; _v1 = Vector128.Create(-1.0f, 2.0f, -3.0f, 4.0f);\n    private readonly Vector128&lt;float&gt; _v2 = Vector128.Create(10.0f);\n    [GlobalSetup]\n    public void Setup()\n    {\n        if (!Sse41.IsSupported)\n            throw new PlatformNotSupportedException();\n    }\n    [Benchmark]\n    public Vector128&lt;float&gt; AddToNegative() =&gt;\n        Sse41.BlendVariable(_v1, _v1 + _v2, _v1);\n}</code></pre>\n<p>In .NET 11, you get the simpler <code>vblendvps</code> form that avoids an unnecessary k-register operation.</p>\n<pre><code>; x64\n  vmovups   xmm0, [rcx+8]\n- vpmovd2m  k1, xmm0\n- vaddps    xmm0 {k1}, xmm0, [rcx+18]\n+ vaddps    xmm1, xmm0, [rcx+18]\n+ vblendvps xmm0, xmm0, xmm1, xmm0\n  vmovups   [rdx], xmm0\n; 29 bytes → 24 bytes</code></pre>\n<p>The masked EVEX form looks more modern, but when the mask originates from a vector anyway, the vector-blend sequence is five bytes shorter and avoids writing a mask register. There are only 8 k-registers, and some microarchitectures have port contention for instructions that write them.</p>\n<p>A compiler’s cost model assigns estimates to operations and instructions, such as their execution cost or throughput and their impact on code size, and uses those estimates to choose between otherwise legal transformations or instruction sequences. Wrong estimates can still produce semantically correct code, just slower or larger code. With dotnet/runtime#127048, which updates the JIT’s xarch floating-point and SIMD cost model, the JIT’s cost model reflects modern instruction throughput and encoded size, replacing old x87 assumptions and a flat cost for every intrinsic. That leads to better decisions about common-subexpression elimination and loop unrolling, particularly for 512-bit operations.</p>\n<p>dotnet/runtime#130422 folds a vector lane extraction followed by <code>WithElement</code> into one <code>insertps</code> that reads the source lane directly. Code such as <code>destination.WithElement(0, source.GetElement(2))</code> conceptually extracts a scalar and then inserts it elsewhere. <code>insertps</code>, however, has an immediate operand whose bits select both the source lane and destination lane. The JIT can therefore pass the original source vector to the instruction and encode lane 2 in that immediate, instead of first shuffling lane 2 into the scalar position and then inserting it.</p>\n<p>Three more xarch changes tighten public SIMD operations on the hardware where they apply. dotnet/runtime#125666 from @alexcovington replaces the dedicated AVX dot-product instruction with a multiply, add, and permute reduction that has better throughput on contemporary cores:</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing System.Numerics;\nusing System.Runtime.Intrinsics;\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private readonly Plane _plane = new(new Vector3(1.0f, 2.0f, 3.0f), 4.0f);\n    private readonly Vector4 _vector4 = new(5.0f, 6.0f, 7.0f, 8.0f);\n    private readonly Quaternion _quaternion1 = new(1.0f, 2.0f, 3.0f, 4.0f);\n    private readonly Quaternion _quaternion2 = new(5.0f, 6.0f, 7.0f, 8.0f);\n    private readonly Vector128&lt;float&gt; _vector1 = Vector128.Create(1.0f, 2.0f, 3.0f, 4.0f);\n    private readonly Vector128&lt;float&gt; _vector2 = Vector128.Create(5.0f, 6.0f, 7.0f, 8.0f);\n    [Benchmark]\n    public float PlaneDot() =&gt; Plane.Dot(_plane, _vector4);\n    [Benchmark]\n    public float QuaternionDot() =&gt; Quaternion.Dot(_quaternion1, _quaternion2);\n    [Benchmark]\n    public float Vector128Dot() =&gt; Vector128.Dot(_vector1, _vector2);\n}</code></pre>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Mean</th><th>Ratio</th><th>Code Size</th></tr></thead><tbody><tr><td>PlaneDot</td><td>.NET 10.0</td><td>2.616 ns</td><td>1.00</td><td>13 B</td></tr><tr><td>PlaneDot</td><td>.NET 11.0</td><td>1.365 ns</td><td>0.52</td><td>31 B</td></tr><tr><td>QuaternionDot</td><td>.NET 10.0</td><td>2.640 ns</td><td>1.00</td><td>13 B</td></tr><tr><td>QuaternionDot</td><td>.NET 11.0</td><td>1.326 ns</td><td>0.50</td><td>31 B</td></tr><tr><td>Vector128Dot</td><td>.NET 10.0</td><td>2.597 ns</td><td>1.00</td><td>13 B</td></tr><tr><td>Vector128Dot</td><td>.NET 11.0</td><td>1.366 ns</td><td>0.53</td><td>31 B</td></tr></tbody></table></div>\n<p>Multiplying vectors of bytes is more involved than multiplying vectors of larger integer types because x86 doesn’t provide a packed byte-multiply instruction. The implementation needs to combine wider 16-bit multiplications while retaining only the low byte of each product. When it couldn’t widen the whole operation to the next vector size, .NET 10 split the input into two halves, widened and multiplied each half, narrowed both results, and joined them again. dotnet/runtime#126348 from @saucecontrol instead separates the even and odd bytes with masks and shifts, performs two 16-bit multiplications over the full vector width, and recombines the low bytes:</p>\n<pre><code>// Run on x64 with AVX-512:\n// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing System.Runtime.Intrinsics;\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private readonly Vector512&lt;byte&gt; _left = Vector512.Create((byte)17);\n    private readonly Vector512&lt;byte&gt; _right = Vector512.Create((byte)19);\n    [Benchmark]\n    public Vector512&lt;byte&gt; Multiply() =&gt; _left * _right;\n}</code></pre>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Mean</th><th>Ratio</th><th>Code Size</th></tr></thead><tbody><tr><td>Multiply</td><td>.NET 10.0</td><td>3.752 ns</td><td>1.00</td><td>114 B</td></tr><tr><td>Multiply</td><td>.NET 11.0</td><td>2.174 ns</td><td>0.58</td><td>73 B</td></tr></tbody></table></div>\n<p>The .NET 11 sequence no longer extracts, widens, narrows, and reinserts both 256-bit halves:</p>\n<pre><code>; x64\n vmovups     zmm0, [rcx+8]\n-vmovaps     zmm1, zmm0\n-vpmovzxbw   zmm1, ymm1\n-vmovups     zmm2, [rcx+48]\n-vmovaps     zmm3, zmm2\n-vpmovzxbw   zmm3, ymm3\n-vpmullw     zmm1, zmm3, zmm1\n-vpmovwb     ymm1, zmm1\n-vextracti32x8 ymm0, zmm0, 1\n-vpmovzxbw   zmm0, ymm0\n-vextracti32x8 ymm2, zmm2, 1\n-vpmovzxbw   zmm2, ymm2\n-vpmullw     zmm0, zmm2, zmm0\n-vpmovwb     ymm0, zmm0\n-vinserti32x8 zmm0, zmm1, ymm0, 1\n+vmovups     zmm1, [rcx+48]\n+vpmullw     zmm2, zmm0, zmm1\n+vpsrlw      zmm0, zmm0, 8\n+vpandd      zmm1, zmm1, dword bcst [RWD00]\n+vpmullw     zmm0, zmm1, zmm0\n+vpternlogd  zmm0, zmm2, dword bcst [RWD04], 0F8\n vmovups     [rdx], zmm0\n; 114 bytes → 73 bytes</code></pre>\n<p>dotnet/runtime#127094 lets scalar conversions between <code>Half</code> and <code>float</code> use F16C’s <code>vcvtps2ph</code> and <code>vcvtph2ps</code> instructions when AVX2 is enabled:</p>\n<pre><code>// Run on x64 with AVX2 enabled:\n// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private Half _half = (Half)123.5f;\n    private float _single = 123.5f;\n    [Benchmark] public float HalfToSingle() =&gt; (float)_half;\n    [Benchmark] public Half SingleToHalf() =&gt; (Half)_single;\n}</code></pre>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Mean</th><th>Ratio</th><th>Code Size</th></tr></thead><tbody><tr><td>HalfToSingle</td><td>.NET 10.0</td><td>2.506 ns</td><td>1.00</td><td>104 B</td></tr><tr><td>HalfToSingle</td><td>.NET 11.0</td><td>1.380 ns</td><td>0.55</td><td>14 B</td></tr><tr><td>SingleToHalf</td><td>.NET 10.0</td><td>2.598 ns</td><td>1.00</td><td>134 B</td></tr><tr><td>SingleToHalf</td><td>.NET 11.0</td><td>1.351 ns</td><td>0.52</td><td>19 B</td></tr></tbody></table></div>\n<p>Finally, dotnet/runtime#127536 from @Ruihan-Yin completes support for APX, Intel’s Advanced Performance Extensions. In addition to expanding the general-purpose register set, APX adds forms of many instructions that don’t overwrite the processor’s condition flags. That gives the register allocator and instruction scheduler more freedom to keep values and pending conditions alive at the same time. Its <code>CTEST</code> and <code>CFCMOV</code> instructions can also represent chained conditions without branches and replace some compare-with-zero forms with shorter encodings. Applications don’t need to call APX-specific APIs to benefit; when the hardware and operating system expose APX, the JIT is able to utilize the additional instructions automatically.</p>\n<p>On Arm64, the work in .NET 11 spans both conventional code generation and continued support for SVE (Scalable Vector Extension). Unlike 128-bit AdvSimd vectors, an SVE vector doesn’t have one width fixed by the instruction set; each processor chooses a supported width, and the same compiled loop uses predicate masks to operate on however many elements fit. That makes SVE well suited to loops whose trip counts are not exact multiples of a particular vector size.</p>\n<p>dotnet/runtime#121986 improves zeroing for larger stack allocations on Arm64. The JIT can store two zeroed 128-bit vector registers at a time, doubling the amount cleared by each instruction:</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Runtime.CompilerServices;\nusing System.Runtime.Intrinsics.Arm;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    [Benchmark] public void Stackalloc512() =&gt; Consume(stackalloc byte[512]);\n    [Benchmark] public void Stackalloc1024() =&gt; Consume(stackalloc byte[1024]);\n    [Benchmark] public void Stackalloc16384() =&gt; Consume(stackalloc byte[16384]);\n    [MethodImpl(MethodImplOptions.NoInlining)]\n    private static void Consume(Span&lt;byte&gt; x) { }\n}</code></pre>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Mean</th><th>Ratio</th></tr></thead><tbody><tr><td>Stackalloc512</td><td>.NET 10.0</td><td>13.65 ns</td><td>1.00</td></tr><tr><td>Stackalloc512</td><td>.NET 11.0</td><td>9.557 ns</td><td>0.70</td></tr><tr><td>Stackalloc1024</td><td>.NET 10.0</td><td>25.35 ns</td><td>1.00</td></tr><tr><td>Stackalloc1024</td><td>.NET 11.0</td><td>14.332 ns</td><td>0.57</td></tr><tr><td>Stackalloc16384</td><td>.NET 10.0</td><td>312.97 ns</td><td>1.00</td></tr><tr><td>Stackalloc16384</td><td>.NET 11.0</td><td>162.656 ns</td><td>0.52</td></tr></tbody></table></div>\n<p>A wave of smaller Arm64 changes improves instruction selection. In .NET 11, dotnet/runtime#119758 from @jonathandavies-arm lets a comparison with zero consume condition flags set as a side effect of the preceding arithmetic or logical instruction, avoiding a separate <code>cmp</code>. dotnet/runtime#123138 from @jonathandavies-arm recognizes bit-extraction idioms such as <code>(value &gt;&gt; 6) &amp; 0x3F</code> and maps them to the dedicated <code>ubfx</code> instruction. And dotnet/runtime#123546 from @jonathandavies-arm removes a non-overflowing <code>int</code>-to-<code>long</code> widening cast when the result is immediately truncated to a smaller integer type.</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;, &quot;left&quot;, &quot;right&quot;, &quot;value&quot;)]\npublic class Benchmarks\n{\n    [Benchmark]\n    [Arguments(-1, 2)]\n    public bool CompareWithZero(int left, int right) =&gt; (left &amp; right) &lt;= 0;\n    [Benchmark]\n    [Arguments(0x7F65_4321)]\n    public int ExtractBits(int value) =&gt; (value &gt;&gt; 6) &amp; 0x3F;\n    [Benchmark]\n    [Arguments(0x1122_3344)]\n    public sbyte TruncateAfterWidening(int value) =&gt; (sbyte)(long)value;\n}</code></pre>\n<p>Each example removes one instruction. <code>CompareWithZero</code> changes <code>and</code> to its flag-setting <code>ands</code> form and drops the subsequent <code>cmp</code>; <code>ExtractBits</code> replaces a shift and mask with <code>ubfx</code>; and <code>TruncateAfterWidening</code> drops the <code>sxtw</code> that widened the value to 64 bits only for <code>sxtb</code> to immediately truncate it again:</p>\n<pre><code>; Arm64\n; CompareWithZero: 28 bytes → 24 bytes\n-            and     w0, w1, w2\n-            cmp     w0, #0\n+            ands    w0, w1, w2\n             cset    x0, le\n; ExtractBits: 24 bytes → 20 bytes\n-            asr     w0, w1, #6\n-            and     w0, w0, #63\n+            ubfx    w0, w1, #6, #6\n; TruncateAfterWidening: 24 bytes → 20 bytes\n-            sxtw    x0, w1\n-            sxtb    w0, w0\n+            sxtb    w0, w1</code></pre>\n<p>Instruction selection also improves where values move between registers and memory. In .NET 11, dotnet/runtime#126803 changes <code>ToScalar</code> on a vector of 64-bit integers to use <code>fmov Xd, Dn</code> rather than the lane-extract instruction <code>umov</code>; in both cases lane zero moves to a general-purpose register, but <code>fmov</code> is the more direct form. For ReadyToRun code, dotnet/runtime#129589 folds relocatable indirection-cell loads from <code>adrp + add + ldr</code> into <code>adrp + ldr #:lo12:</code>, removing the separate address addition. And dotnet/runtime#129932 re-enables <code>ldp</code>/<code>stp</code> formation for negative unscaled offsets, letting two adjacent loads or stores become one paired instruction.</p>\n<p>The first and third changes are easy to see with small methods:</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Runtime.CompilerServices;\nusing System.Runtime.Intrinsics;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private Vector128&lt;long&gt; _vector = Vector128.Create(42L, 84L);\n    private nint[] _storage = new nint[8];\n    [Benchmark]\n    public long ToScalar() =&gt; ToScalarCore(_vector);\n    [Benchmark]\n    public void ClearPrevious() =&gt; ClearPreviousCore(ref _storage[4]);\n    [MethodImpl(MethodImplOptions.NoInlining)]\n    private static long ToScalarCore(Vector128&lt;long&gt; value) =&gt; value.ToScalar();\n    [MethodImpl(MethodImplOptions.NoInlining)]\n    private static void ClearPreviousCore(ref nint value)\n    {\n        Unsafe.Add(ref value, -1) = 0;\n        Unsafe.Add(ref value, -2) = 0;\n        Unsafe.Add(ref value, -3) = 0;\n        Unsafe.Add(ref value, -4) = 0;\n    }\n}</code></pre>\n<p>The <code>ToScalar</code> change is a direct instruction substitution, while the negative-offset stores collapse from four instructions to two, reducing the helper from 32 bytes to 24 bytes:</p>\n<pre><code>; Arm64\n; ToScalarCore\n-            umov    x0, v0.d[0]\n+            fmov    x0, d0\n; ClearPreviousCore\n-            str     xzr, [x0, #-0x08]\n-            str     xzr, [x0, #-0x10]\n-            str     xzr, [x0, #-0x18]\n-            str     xzr, [x0, #-0x20]\n+            stp     xzr, xzr, [x0, #-0x10]\n+            stp     xzr, xzr, [x0, #-0x20]</code></pre>\n<p>Bit-counting operations benefit as well. <code>PopCount</code> counts the one bits in a value, while <code>TrailingZeroCount</code> counts the zero bits below its least-significant one bit. dotnet/runtime#128677 imports both as dedicated Arm64 intrinsics, making their intent visible to later optimization. On processors with the FEAT_CSSC extension, dotnet/runtime#130332 can then lower them directly to the scalar <code>cnt</code> and <code>ctz</code> instructions.</p>\n<p>Comparison masks are another place where spelling out the intent enables much better code. Portable SIMD code often compares vectors, calls <code>ExtractMostSignificantBits</code>, and then asks whether any lane matched, counts matching lanes, or finds the first or last match. dotnet/runtime#129688 from @jonathandavies-arm recognizes those consumers on Arm64 and avoids materializing the full scalar mask: it can horizontally reduce the vector mask directly.</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Numerics;\nusing System.Runtime.CompilerServices;\nusing System.Runtime.Intrinsics;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private Vector128&lt;int&gt; _value = Vector128.Create(1, -2, 3, -4);\n    [Benchmark]\n    public bool AnyLessThan() =&gt; AnyLessThanCore(_value, 0);\n    [Benchmark]\n    public int CountLessThan() =&gt; CountLessThanCore(_value, 0);\n    [MethodImpl(MethodImplOptions.NoInlining)]\n    private static bool AnyLessThanCore(Vector128&lt;int&gt; value, int limit) =&gt;\n        Vector128.LessThan(value, Vector128.Create(limit))\n            .ExtractMostSignificantBits() != 0;\n    [MethodImpl(MethodImplOptions.NoInlining)]\n    private static int CountLessThanCore(Vector128&lt;int&gt; value, int limit) =&gt;\n        BitOperations.PopCount(\n            Vector128.LessThan(value, Vector128.Create(limit))\n                .ExtractMostSignificantBits());\n}</code></pre>\n<p>In .NET 10, both helpers first pack the most-significant bit from every comparison lane into a scalar. .NET 11 instead keeps the comparison as a vector.</p>\n<pre><code>; Arm64\n; AnyLessThanCore\n             cmgt    v16.4s, v16.4s, v0.4s\n-            movi    v17.4s, #0x80, LSL #24\n-            and     v16.4s, v16.4s, v17.4s\n-            ldr     q17, [@RWD00]\n-            ushl    v16.4s, v16.4s, v17.4s\n-            addv    s16, v16.4s\n-            smov    x0, v16.s[0]\n+            umaxv   s16, v16.4s\n+            umov    w0, v16.s[0]\n             cmp     w0, #0\n             cset    x0, ne\n; CountLessThanCore\n             cmgt    v16.4s, v16.4s, v0.4s\n-            movi    v17.4s, #0x80, LSL #24\n-            and     v16.4s, v16.4s, v17.4s\n-            ldr     q17, [@RWD00]\n-            ushl    v16.4s, v16.4s, v17.4s\n-            addv    s16, v16.4s\n-            movi    v17.2s, #0\n-            smov    x0, v16.s[0]\n-            ins     v17.s[0], w0\n-            cnt     v16.8b, v17.8b\n-            addv    b16, v16.8b\n-            umov    w0, v16.b[0]\n+            ushr    v16.4s, v16.4s, #31\n+            addv    s16, v16.4s\n+            umov    w0, v16.s[0]</code></pre>\n<p>On the SVE and SVE2 side, dotnet/runtime#129852 from @snickolls-arm removes the old 128-bit size ceiling for <code>Vector&lt;T&gt;</code> on Arm64 and lets the runtime size the type from the process’s actual SVE vector length. (Scalable <code>Vector&lt;T&gt;</code> remains experimental and disabled by default in .NET 11, so this expands what the experimental mode can do; it doesn’t speed up the default <code>Vector&lt;T&gt;</code> configuration.)</p>\n<p>The public intrinsic surface also grows. In .NET 11, dotnet/runtime#118957 from @SwapnilGaikwad exposes odd-lane floating-point conversions; “odd lane” here means converting elements 1, 3, 5, and so on, which is useful when widening or narrowing interleaved data. dotnet/runtime#123890 from @ylpoonlg and dotnet/runtime#123892 from @ylpoonlg add non-temporal gather loads and scatter stores, which read from or write to multiple non-contiguous addresses (the “gather” part) while hinting that the data need not remain in cache (the “non-temporal” part).</p>\n<p>Other changes improve the predicates that make scalable loops work. dotnet/runtime#127538 adds hardware-generated predicate masks for more loop and memory-access patterns, while dotnet/runtime#126398 from @ylpoonlg reduces setup moves for masked operations. And dotnet/runtime#128326 from @snickolls-arm improves how SVE masks flow through the JIT, allowing zeroing forms of instructions to replace separate constant setup. dotnet/runtime#127520 from @a74nh enables scalable vector and mask constants, and dotnet/runtime#128148 from @snickolls-arm uses vector stores to initialize scalable vector locals, replacing scalar loops.</p>\n<h3 id=\"register-allocation\">Register Allocation</h3>\n<p>Generated code constantly moves values between the CPU’s limited set of fast registers and temporary stack slots. Register allocation in a compiler decides which values stay in registers and which are “spilled” to the stack; avoiding one spill can remove both the store and the later reload.</p>\n<p>Some small structs are passed with multiple fields packed into one register. In .NET 11, dotnet/runtime#112740 lets the JIT extract those fields directly, avoiding a “spill” to a temporary stack slot followed by a reload of each field:</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Drawing;\nusing System.Runtime.CompilerServices;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private readonly Memory&lt;int&gt;[] _memories = CreateMemories();\n    private static Memory&lt;int&gt;[] CreateMemories()\n    {\n        Random rng = new(42);\n        var memories = new Memory&lt;int&gt;[4096];\n        for (int i = 0; i &lt; memories.Length; i++)\n            memories[i] = new int[rng.Next(0, 20)];\n        return memories;\n    }\n    [MethodImpl(MethodImplOptions.NoInlining)]\n    private static bool Test(Memory&lt;int&gt; mem) =&gt; mem.Length &gt; 10;\n    [Benchmark]\n    public int MemoryLengthExtract_Loop()\n    {\n        int count = 0;\n        for (int i = 0; i &lt; _memories.Length; i++)\n            if (Test(_memories[i]))\n                count++;\n        return count;\n    }\n}</code></pre>\n<p>The measured row uses <code>Memory&lt;int&gt;</code> because its length arrives packed into part of an argument register on Arm64. The new extraction avoids a stack round-trip on every call.</p>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Mean</th><th>Ratio</th></tr></thead><tbody><tr><td>MemoryLengthExtract_Loop</td><td>.NET 10.0</td><td>27.15 μs</td><td>1.00</td></tr><tr><td>MemoryLengthExtract_Loop</td><td>.NET 11.0</td><td>23.95 μs</td><td>0.88</td></tr></tbody></table></div>\n<p>Two broader register-allocation changes reduce unnecessary copies and spills: dotnet/runtime#125214 handles more conflicts directly, while dotnet/runtime#125219 steers short-lived values away from registers an upcoming operation will overwrite. dotnet/runtime#126552 from @SingleAccretion removes an old restriction on method prologs, eliminating jumps that existed only to satisfy that encoding rule.</p>\n<h3 id=\"write-barriers-and-garbage-collection\">Write Barriers and Garbage Collection</h3>\n<p>The .NET garbage collector is generational: new objects start in gen0, while objects that survive collections are promoted to gen1 and gen2. That enables the GC to collect younger generations without having to scan the whole heap. Of course, a reference to a younger object could get written to a field of an older one, in which case only scanning the younger generation would lead to problems. To ensure such references aren’t missed, whenever a write could create one, the JIT emits a small piece of code to update the GC’s bookkeeping; that code is known as a GC write barrier. Reference writes happen a lot, so it’s really important for performance that those barriers be as cheap as possible, and elided if they’re provably not needed at all.</p>\n<p>Managed reference stores may require both an array covariance check and a GC write barrier. Arrays in .NET are covariant, meaning a <code>TDerived[]</code> can be used as a <code>TBase[]</code>, e.g. a <code>string[]</code> can be used as an <code>object[]</code>; consequently, storing an instance into an <code>object[]</code> must validate that the instance is actually of the right type (otherwise, you could have a <code>TDerived1[]</code> masquerading as a <code>TBase[]</code> and try to store a <code>TDerived2</code> into it, which would cause badness if it were to store successfully). dotnet/runtime#126547 expands calls to the runtime’s array-store helper into the individual operations it performs, exposing both the covariance check and write barrier to the JIT. When the JIT knows the array’s exact type, it can then eliminate the covariance check and optimize the barrier:</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Runtime.CompilerServices;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private readonly object[] _array = new object[4096];\n    private object _value = new();\n    [MethodImpl(MethodImplOptions.NoInlining)]\n    private static void StoreAll(object[] arr, object value)\n    {\n        for (int i = 0; i &lt; arr.Length; i++)\n            arr[i] = value;\n    }\n    [Benchmark]\n    public object[] CovariantStore_Loop()\n    {\n        StoreAll(_array, _value);\n        return _array;\n    }\n}</code></pre>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Mean</th><th>Ratio</th></tr></thead><tbody><tr><td>CovariantStore_Loop</td><td>.NET 10.0</td><td>10.85 μs</td><td>1.00</td></tr><tr><td>CovariantStore_Loop</td><td>.NET 11.0</td><td>6.042 μs</td><td>0.56</td></tr></tbody></table></div>\n<p>Sometimes writes are done one at a time, but sometimes they can be batched, as happens when copying structs. dotnet/runtime#128238 extends the JIT’s heap-destination analysis from individual stores to whole-struct copies. dotnet/runtime#128542 then replaces a specialized helper that copied one reference field at a time with reference stores and vector stores for the non-reference data. Together, they let the JIT choose more efficient write barriers and copy the rest of a mixed struct with SIMD.</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Runtime.CompilerServices;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    [InlineArray(4)]\n    public struct InlineArray4Long\n    {\n        private long _element0;\n    }\n    public struct MyStruct\n    {\n        public string A;\n        public InlineArray4Long G;\n        public string B;\n    }\n    private MyStruct _src;\n    private MyStruct _dst;\n    [GlobalSetup]\n    public void Setup()\n    {\n        _src = new MyStruct { A = &quot;hello&quot;, B = &quot;world&quot; };\n        _src.G[0] = 1;\n        _src.G[1] = 2;\n        _src.G[2] = 3;\n        _src.G[3] = 4;\n    }\n    [Benchmark]\n    public void HeapStructCopy() =&gt; _dst = _src;\n}</code></pre>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Mean</th><th>Ratio</th></tr></thead><tbody><tr><td>HeapStructCopy</td><td>.NET 10.0</td><td>4.132 ns</td><td>1.00</td></tr><tr><td>HeapStructCopy</td><td>.NET 11.0</td><td>3.071 ns</td><td>0.74</td></tr></tbody></table></div>\n<p>dotnet/runtime#130535 handles the equivalent case for small structs that don’t contain object references. Once the JIT has turned the copy into several writes to adjacent fields, it can combine them into fewer, wider writes.</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Runtime.CompilerServices;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private Int128 _value;\n    [Benchmark]\n    public void StoreInt128() =&gt; _value = 123456789;\n}</code></pre>\n<p>.NET 10 stores the low and high halves separately. .NET 11 loads the value into a vector register and writes all 16 bytes at once.</p>\n<pre><code>; x64\n; StoreInt128\n-       mov      qword ptr [rcx+8], 75BCD15\n-       xor      eax, eax\n-       mov      [rcx+10], rax\n+       vmovss   xmm0, dword ptr [RWD00]\n+       vmovups  [rcx+8], xmm0\n-; Total bytes of code 15\n+; Total bytes of code 14</code></pre>\n<p>The same idea applies when the source code assigns neighboring fields individually. dotnet/runtime#126562 enables this for promoted struct locals, while dotnet/runtime#130107 extends it to adjacent fields at constant static addresses:</p>\n<pre><code>// dotnet run -c Release -f net11.0 --filter &quot;*&quot;\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Runtime.CompilerServices;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private static Point s_point;\n    [Benchmark]\n    public void SetPoint() =&gt; Set();\n    [MethodImpl(MethodImplOptions.NoInlining | MethodImplOptions.AggressiveOptimization)]\n    private static void Set()\n    {\n        s_point.X = 1;\n        s_point.Y = 2;\n    }\n    private struct Point\n    {\n        public int X;\n        public int Y;\n    }\n}</code></pre>\n<p>The referenced .NET 11 x64 build combines the two 32-bit constants and writes both fields with one 64-bit store:</p>\n<pre><code>; x64\nmov     rax, 200000001\nmov     rcx, &lt;address of s_point&gt;\nmov     [rcx], rax</code></pre>\n<p>dotnet/runtime#127487 applies a related improvement when stack protection requires a struct parameter to be copied. It uses consistently sized writes so a subsequent wider read doesn’t need to wait for the processor to reconcile overlapping stores.</p>\n<p>Write barriers are only one part of the interaction between generated code\nand the garbage collector. During a compacting collection, the GC needs to\nplan where surviving objects will move and then update references to them. To\ndo that efficiently, it records their addresses, sorts those addresses, and\ngroups adjacent survivors into regions called “plugs.” With enough live\nobjects, sorting these mark lists becomes a meaningful part of the collection.\nRecent x86/x64 runtimes use a vectorized <code>vxsort</code> implementation for\nsufficiently large lists. In .NET 11,\ndotnet/runtime#110692 from\n@a74nh extends that support to Arm64.</p>\n<p>The generation assigned to GC metadata matters just as much as the speed of\none collection. .NET’s generational GC is based on the observation that most\nobjects die young: generation 0 and generation 1 collections, collectively\ncalled ephemeral collections, run frequently and should avoid revisiting\nstate that has already survived into generation 2. A dependent handle\nassociates a primary object with a secondary object, keeping the secondary\nalive while the primary remains reachable; <code>ConditionalWeakTable&lt;TKey, TValue&gt;</code> is built on this mechanism. Previously, the handle itself didn’t age\nwith its referents, so every ephemeral collection continued scanning it even\nafter both objects had become long-lived.\ndotnet/runtime#78746 ages\ndependent handles accordingly and moves a handle back to a younger generation\nwhen necessary. Old handles can therefore be skipped by young collections\nwithout compromising reachability.</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing System.Runtime.CompilerServices;\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private ConditionalWeakTable&lt;object, object&gt; _table = new();\n    private object[] _keys = [];\n    [Params(100_000, 1_000_000)]\n    public int Handles { get; set; }\n    [GlobalSetup]\n    public void Setup()\n    {\n        _table = new();\n        _keys = new object[Handles];\n        for (int i = 0; i &lt; _keys.Length; i++)\n        {\n            object key = new();\n            _keys[i] = key;\n            _table.Add(key, new object());\n        }\n        GC.Collect(2, GCCollectionMode.Forced, blocking: true, compacting: true);\n    }\n    [Benchmark]\n    public void CollectGen0() =&gt;\n        GC.Collect(0, GCCollectionMode.Forced, blocking: true, compacting: false);\n}</code></pre>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Handles</th><th>Mean</th><th>Ratio</th></tr></thead><tbody><tr><td>CollectGen0</td><td>.NET 10.0</td><td>100000</td><td>1.522 ms</td><td>1.00</td></tr><tr><td>CollectGen0</td><td>.NET 11.0</td><td>100000</td><td>255.5 μs</td><td>0.17</td></tr><tr><td>CollectGen0</td><td>.NET 10.0</td><td>1000000</td><td>10.737 ms</td><td>1.00</td></tr><tr><td>CollectGen0</td><td>.NET 11.0</td><td>1000000</td><td>310.7 μs</td><td>0.029</td></tr></tbody></table></div>\n<h3 id=\"runtime-knowledge-and-frozen-data\">Runtime Knowledge and Frozen Data</h3>\n<p>The JIT can optimize only the facts it knows. Some facts come from its own analysis; others are contracts supplied by the runtime, such as which helpers have side effects, the length of a newly allocated string, or whether a data object will ever move.</p>\n<p>A generic virtual call such as <code>baseReference.Foo&lt;string&gt;()</code> may need help from the runtime to find the implementation for both the object’s actual type and the generic argument. If that lookup appears to have arbitrary side effects, the JIT has to perform it exactly where it occurs, rather than possibly resulting on a cached answer from a previous lookup. In .NET 11, dotnet/runtime#122017 teaches the JIT more precisely which exceptions these runtime helpers can throw and whether they otherwise have side effects. The JIT can then share repeated lookups or move an unchanging lookup out of a loop:</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Runtime.CompilerServices;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    public abstract class Base\n    {\n        public abstract void Foo&lt;T&gt;();\n    }\n    public class Derived : Base\n    {\n        public override void Foo&lt;T&gt;() { }\n    }\n    private Base _b = new Derived();\n    [Benchmark]\n    public void GvmCseHoist()\n    {\n        Base b = _b;\n        b.Foo&lt;string&gt;();\n        b.Foo&lt;int&gt;();\n        b.Foo&lt;string&gt;();\n        b.Foo&lt;int&gt;();\n        for (int i = 0; i &lt; 10; i++)\n            b.Foo&lt;double&gt;();\n    }\n}</code></pre>\n<p>In .NET 11, the repeated lookups outside the loop are shared and the loop’s lookup is performed once, not ten times.</p>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Mean</th><th>Ratio</th></tr></thead><tbody><tr><td>GvmCseHoist</td><td>.NET 10.0</td><td>45.75 ns</td><td>1.00</td></tr><tr><td>GvmCseHoist</td><td>.NET 11.0</td><td>24.06 ns</td><td>0.53</td></tr></tbody></table></div>\n<p>Profile data is another way the JIT learns what matters. Inlining could previously hide important work from the instrumentation used to gather that data. dotnet/runtime#119658 allows the inlined code to be instrumented as well, giving later PGO-driven compilation a more complete picture of the hot paths.</p>\n<h3 id=\"jit-throughput-and-cleanup\">JIT Throughput and Cleanup</h3>\n<p>The quality of the generated code isn’t the only concern; the time spent producing it matters too. Every analysis the JIT performs has a cost. dotnet/runtime#123856 removes checks and maps from Global Assertion Propagation whose bookkeeping wasn’t paying for itself. This is the recurring balancing act in the development of the JIT: retaining the information that enables meaningful optimizations while avoiding analysis overhead whose code-quality benefit is negligible.</p>\n<p>dotnet/runtime#127363 makes profile-guided optimization more resilient with OSR (on-stack replacement), which replaces a method while one of its loops is already running. Because that execution begins in the middle of the method rather than at its normal entry, reconstructed profile data doesn’t always line up perfectly with the paths actually available. The JIT now estimates the likelihood of those paths rather than asserting or abandoning the profile.</p>\n<p>Optimizations can leave behind code that’s no longer reachable, so the JIT also needs to be good at dead code removal. dotnet/runtime#126223 runs another sweep whenever the method’s branching structure changes, catching blocks made obsolete by earlier transformations.</p>\n<p>And dotnet/runtime#128515 from @BoyBaykiller repeatedly combines equivalent return and throw endings, removing duplicate exit paths and sometimes exposing more code that can be shared.</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    [Benchmark]\n    [Arguments((byte)9)]\n    public bool IsLinearWhiteSpace(byte value) =&gt;\n        value &lt;= 32 &amp;&amp;\n        (value == 32 || value == 10 || value == 13 || value == 9);\n}</code></pre>\n<p>In .NET 10, tail merging combines the paths that return <code>false</code>, but not both paths that return <code>true</code>. As a result, the JIT’s bit test covers three of the four values, with a separate comparison for <code>9</code>. In .NET 11, the true returns are merged as well, enabling all four values to be handled by the same bit test:</p>\n<pre><code>; x64\n-       movzx    ecx, dl\n-       cmp      ecx, 20\n-       jg       M00_L02\n-       cmp      ecx, 20\n-       ja       M00_L01\n-       mov      eax, 0FFFFDBFF\n-       bt       rax, rcx\n-       jae      M00_L00\n-       mov      eax, 1\n-       ret\n-M00_L00:\n-       cmp      ecx, 9\n-       sete     al\n-       movzx    eax, al\n-       ret\n-M00_L01:\n+       movzx    eax, dl\n+       cmp      eax, 20\n+       jg       M00_L00\n+       cmp      eax, 20\n+       ja       M00_L00\n+       mov      ecx, 0FFFFD9FF\n+       bt       rcx, rax\n+       jb       M00_L00\n+       mov      eax, 1\n+       ret\n+M00_L00:\n        xor      eax, eax\n        ret\n-; Total bytes of code 43\n+; Total bytes of code 33</code></pre>\n<p>Also related to dead code, a call that never returns, such as one that always throws, makes everything after it unreachable. In .NET 11, after inlining, dotnet/runtime#128513 removes the remaining statements and outgoing paths from such a block and marks it as ending in a throw, exposing the dead code early enough for the cleanup passes above to remove it.</p>\n<h2 id=\"startup-and-deployment\">Startup and Deployment</h2>\n<p>Before managed <code>Main</code> can run, the native host needs to locate the application’s dependencies, CoreCLR needs to load enough types and code to begin execution, and various pieces of framework infrastructure need to initialize themselves. Work removed from any of those stages helps the application get going sooner, improving startup time.</p>\n<p>The host starts by reading the application’s <code>.deps.json</code>, turning its entries into paths, and building the trusted platform assembly (TPA) list. That list tells CoreCLR which framework and application assemblies it can resolve by simple name. Several costs in this process scaled with the number of assets rather than with the amount of useful work. dotnet/runtime#123568 in .NET 11 avoids checking every asset against a servicing directory unless the resolver is actually probing that directory. dotnet/runtime#123919 avoids repeatedly comparing the servicing-directory name and copying every dependency asset while constructing the TPA list, avoiding a lot of allocation. dotnet/runtime#125251 removes more allocation by normalizing each asset’s directory separators once when parsing the <code>.deps.json</code>, rather than normalizing the path again every time it is used.</p>\n<p>Once the host hands off to CoreCLR, ReadyToRun (R2R) code helps avoid compiling methods before they can execute. However, initializing <code>Comparer&lt;T&gt;.Default</code> and <code>EqualityComparer&lt;T&gt;.Default</code> called a reflection-based helper whose resulting concrete comparer type wasn’t known when the R2R image was built. The comparer constructor and operations could consequently fall back to being interpreted. In .NET 11, dotnet/runtime#126204 uses specialized helpers that R2R can compile ahead of time and ensures the required comparer types are included in the image.</p>\n<p>Even better than making initialization faster is avoiding it altogether. An <code>EventSource</code> normally discovers its event metadata and computes its provider GUID when it is initialized. dotnet/runtime#121180 adds an internal source generator that performs this work when the framework is built and emits the result for its <code>EventSource</code> implementations, including the ones for core runtime tracing. Applications then don’t need to pay the reflection and setup costs when those event sources are first used.</p>\n<p>Startup also has a memory footprint outside the managed heap. Native AOT’s <code>AllocHeap</code> typically holds only small amounts of runtime metadata. On Windows, however, its virtual-memory allocator reserved a 64 KB region for each block even when it initially needed only 4 KB. In .NET 11, dotnet/runtime#122822 instead uses ordinary <code>new</code> and <code>delete</code> for these small blocks, matching the allocation strategy to the amount of memory normally involved.</p>\n<p>Note that the aforementioned R2R work wasn’t motivated only by desktop and server startup. It was also part of the substantial effort to make CoreCLR the runtime for .NET on mobile. Starting with .NET 11, .NET MAUI moved to CoreCLR for Android, iOS, and Mac Catalyst, the last .NET MAUI platforms that had still been using Mono. This is much more than swapping one execution engine for another. Those apps now use the same runtime as ASP.NET Core, cloud services, and desktop .NET, with the same JIT, garbage collector, diagnostics infrastructure, performance improvements, and bug fixes. It also brings CoreCLR’s tiered compilation, ReadyToRun, and profile-guided optimization to mobile, while providing a common foundation for NativeAOT. That combination is important: R2R and packaged profiles can precompile the code most important to startup, while the optimizing JIT can produce higher-quality code for hot methods on platforms where dynamic compilation is available. Improvements like the comparer specialization mentioned earlier keep more code on the compiled path instead of falling back to interpretation.</p>\n<h2 id=\"threading\">Threading</h2>\n<p>Threading is a cross-cutting concern that impacts almost every application and service. Whether code is protecting shared state, queueing work, or coordinating asynchronous operations, small costs in the underlying machinery can quickly add up. As such, it’s something that’s revisited in every release of .NET.</p>\n<p><code>Monitor</code> is the synchronization primitive historically used to implement <code>lock</code>, providing the most pervasively used support for mutual exclusion. It also supports sending signals, such that one thread can wait on a <code>Monitor</code> with <code>Monitor.Wait</code> for another thread to <code>Pulse</code> it. The internal object that tracks these waiters is a “condition variable.” dotnet/runtime#129083 stores that condition directly on the lock, removing a separate <code>ConditionalWeakTable</code> lookup from this already synchronization-heavy path.</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private const int RoundTripsPerInvoke = 2_000;\n    private readonly object _gate = new();\n    private int _ping;\n    private int _pong;\n    private bool _stop;\n    private Thread _responder = null!;\n    [GlobalSetup]\n    public void Setup()\n    {\n        _responder = new Thread(ResponderLoop) { IsBackground = true };\n        _responder.Start();\n    }\n    [GlobalCleanup]\n    public void Cleanup()\n    {\n        lock (_gate)\n        {\n            _stop = true;\n            Monitor.PulseAll(_gate);\n        }\n        _responder.Join();\n    }\n    private void ResponderLoop()\n    {\n        lock (_gate)\n        {\n            int seen = 0;\n            while (true)\n            {\n                while (_ping == seen &amp;&amp; !_stop)\n                    Monitor.Wait(_gate);\n                if (_stop)\n                    return;\n                seen = _ping;\n                _pong = seen;\n                Monitor.PulseAll(_gate);\n            }\n        }\n    }\n    [Benchmark(OperationsPerInvoke = RoundTripsPerInvoke)]\n    public int PingPong_MonitorWaitPulse()\n    {\n        lock (_gate)\n        {\n            for (int i = 0; i &lt; RoundTripsPerInvoke; i++)\n            {\n                _ping++;\n                int expected = _ping;\n                Monitor.PulseAll(_gate);\n                while (_pong != expected)\n                    Monitor.Wait(_gate);\n            }\n            return _pong;\n        }\n    }\n}</code></pre>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Mean</th><th>Ratio</th></tr></thead><tbody><tr><td>PingPong_MonitorWaitPulse</td><td>.NET 10.0</td><td>4.194 μs</td><td>1.00</td></tr><tr><td>PingPong_MonitorWaitPulse</td><td>.NET 11.0</td><td>3.517 μs</td><td>0.84</td></tr></tbody></table></div>\n<p>In the case of <code>Monitor</code>, that improvement targeted the specific shared implementation. In other cases, the costs are spread out in a more peanut butter manner across lots of code. dotnet/runtime#125274 removes some of that peanut butter by removing unnecessary <code>volatile</code> annotations from a wide range of library fields whose correctness already comes from locks, <code>Interlocked</code>, or one-time initialization. On x86/x64 hardware, which already provides a strong memory model, those annotations generally don’t result in extra instructions, though they can still constrain compiler optimizations. Arm, however, permits more reordering, so the JIT often needs to emit memory fences to provide <code>volatile</code>‘s guarantees. Removing the annotations where they’re redundant therefore can end up removing unnecessary fences from Arm’s generated code.</p>\n<p>Similar considerations apply to code in the runtime. dotnet/runtime#125259 replaces\nfull memory barriers in the runtime’s <code>HashMap</code> with the narrower acquire and\nrelease operations actually required. On top of that, many VM\nhash tables, including its <code>EEHashTable</code>, are read constantly but updated only\noccasionally. dotnet/runtime#124822\nadds epoch-based reclamation, enabling readers to avoid entering cooperative\nGC mode simply to keep an old set of buckets alive. And dotnet/runtime#129640 replaces\nthe previous byte-at-a-time hash used by these tables with an xxHash\nimplementation that consumes four bytes at a time.</p>\n<p>Along the same lines, in .NET 11 dotnet/runtime#122726 reduces the scheduling overhead around small thread-pool work items. It removes unnecessary memory fences and shared-state updates, checks in with the thread-pool controller once per batch rather than once per item, spends less time spinning on a semaphore, and requests another worker only when the queued work shows one is needed. The result is less coordination overhead and fewer workers woken just as the queue becomes empty.</p>\n<p>Earlier in this post, we talked about runtime async, which can have a significant impact on the performance of <code>async</code>/<code>await</code> code, how they produce <code>Task</code>s, and so on. They’re not the only improvements in .NET 11 related to <code>Task</code>s, though.</p>\n<p>One fun one is a new analyzer, CA2027, introduced in dotnet/sdk#51452. With that, the SDK can point out problematic usage of <code>Task.Delay</code> that I’ve seen on multiple occasions to lead to non-trivial performance issues in large scale services. Consider this code:</p>\n<pre><code>Task someTask = ...;\nif (await Task.WhenAny(someTask, Task.Delay(timeout)) != someTask) // oops!\n{\n    throw new TimeoutException();\n}</code></pre>\n<p>The developer that wrote this is obviously trying to implement a timeout. The problem, however, is that this leaks. In the hopefully common case where <code>someTask</code> completes really quickly, the <code>Task.Delay</code> will still be pending. That <code>Delay</code> has associated with it a <code>System.Threading.Timer</code> that’s consuming valuable resources, as well as other data in memory, and if this <code>timeout</code> is long and this code is on a hotter path, we could accumulate thousands upon thousands of those timers. That in turn can increase memory use and slow down other calls that interact with timers.</p>\n<p>The fix is to instead use the <code>Task.WaitAsync</code> method, introduced all the way back in .NET 6. It provides a much more efficient mechanism for doing this same kind of timed waiting, and it correctly handles all the relevant cleanup. CA2027 will detect common forms of this issue and recommend the replacement.</p>\n<h2 id=\"numerics\">Numerics</h2>\n<p><code>BigInteger</code> is one of those types that many applications may never need, but\nfor those that do, there’s often no practical substitute. It powers workloads\nranging from cryptography and number theory to compilers and applications that\nneed to parse, format, or compute with integers larger than the fixed-width\nprimitives can hold. Despite that need, however, <code>BigInteger</code> hasn’t received the\nsame steady stream of performance investment as many of .NET’s other core\ntypes. Thankfully, in .NET 11 it gets a makeover.</p>\n<p>dotnet/runtime#125799 rewrote significant portions of <code>BigInteger</code>‘s implementation, changing its limbs (the fixed-size pieces stored in its backing array) from <code>uint</code> to <code>nuint</code> (<code>UIntPtr</code>). That makes no effective difference on a 32-bit machine. On a 64-bit machine, however, each limb grows from 32 to 64 bits; since most arithmetic on a 64-bit value on a 64-bit platform costs no more than the corresponding 32-bit operation, each step can therefore process twice as many bits in the same number of cycles. The implementation also improves the algorithms around those wider limbs, including Montgomery multiplication and sliding-window exponentiation in <code>ModPow</code>, fused bitwise steps, additional hardware intrinsics, loop unrolling, and caching. That all builds on top of other optimizations that were done previously in the release, such as faster conversion of huge values to decimal text in dotnet/runtime#112178 from @kzrnm, dotnet/runtime#112876 from @kzrnm using Toom-Cook multiplication for sufficiently large operands, and improved shifts and rotations thanks to dotnet/runtime#113005 from @kzrnm. Toom-Cook splits each operand into several chunks and combines smaller products, doing less work than the straightforward every-limb-by-every-limb algorithm once the operands are large enough.</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Numerics;\nusing System.Globalization;\nusing System.Text;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    [Params(64, 512)]\n    public int Limbs;\n    private BigInteger _a;\n    private BigInteger _b;\n    private BigInteger _shiftSubject;\n    private BigInteger _hugeValueForToString;\n    private string _decimalDigits100000 = &quot;&quot;;\n    private byte[] _utf8Digits1000 = [];\n    private byte[] _utf8FormatBuffer = new byte[120_000];\n    private BigInteger _divideDividendBelowThreshold;\n    private BigInteger _divideDivisorBelowThreshold;\n    private BigInteger _divideDividendAboveThreshold;\n    private BigInteger _divideDivisorAboveThreshold;\n    [GlobalSetup]\n    public void Setup()\n    {\n        _a = MakeDeterministicBigInteger(Limbs, seed: 1);\n        _b = MakeDeterministicBigInteger(Limbs, seed: 2);\n        _shiftSubject = MakeDeterministicBigInteger(Limbs, seed: 3);\n        _decimalDigits100000 = MakeDeterministicDecimalDigits(100_000);\n        _hugeValueForToString = BigInteger.Parse(_decimalDigits100000, CultureInfo.InvariantCulture);\n        string decimalDigits1000 = MakeDeterministicDecimalDigits(1_000);\n        _utf8Digits1000 = Encoding.UTF8.GetBytes(decimalDigits1000);\n        _divideDivisorBelowThreshold = MakeDeterministicBigInteger(16, seed: 4);\n        _divideDividendBelowThreshold = MakeDeterministicBigInteger(16 + 96, seed: 5);\n        _divideDivisorAboveThreshold = MakeDeterministicBigInteger(128, seed: 6);\n        _divideDividendAboveThreshold = MakeDeterministicBigInteger(128 + 96, seed: 7);\n    }\n    private static BigInteger MakeDeterministicBigInteger(int limbCount, int seed)\n    {\n        Random rng = new(seed);\n        byte[] bytes = new byte[(limbCount * 4) + 1]; // trailing 0 byte keeps the value positive\n        rng.NextBytes(bytes);\n        bytes[^1] = 0;\n        return new BigInteger(bytes);\n    }\n    private static string MakeDeterministicDecimalDigits(int digitCount)\n    {\n        StringBuilder sb = new(digitCount);\n        sb.Append(&#39;9&#39;); // avoid a leading zero, which would shorten the effective digit count\n        Random rng = new(42);\n        for (int i = 1; i &lt; digitCount; i++)\n            sb.Append((char)(&#39;0&#39; + rng.Next(0, 10)));\n        return sb.ToString();\n    }\n    [Benchmark]\n    public BigInteger Divide_BelowBurnikelZieglerThreshold() =&gt; _divideDividendBelowThreshold / _divideDivisorBelowThreshold;\n    [Benchmark]\n    public BigInteger Divide_AboveBurnikelZieglerThreshold() =&gt; _divideDividendAboveThreshold / _divideDivisorAboveThreshold;\n    [Benchmark]\n    public BigInteger Multiply() =&gt; _a * _b;\n    [Benchmark]\n    public BigInteger ShiftLeft() =&gt; _shiftSubject &lt;&lt; 12345;\n    [Benchmark]\n    public BigInteger ParseLargeDecimal() =&gt; BigInteger.Parse(_decimalDigits100000, CultureInfo.InvariantCulture);\n    [Benchmark]\n    public string ToStringLargeDecimal() =&gt; _hugeValueForToString.ToString(CultureInfo.InvariantCulture);\n}</code></pre>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Limbs</th><th>Mean</th><th>Ratio</th></tr></thead><tbody><tr><td>Divide_BelowBurnikelZieglerThreshold</td><td>.NET 10.0</td><td>64</td><td>2,954.3 ns</td><td>1.00</td></tr><tr><td>Divide_BelowBurnikelZieglerThreshold</td><td>.NET 11.0</td><td>64</td><td>1,493.8 ns</td><td>0.51</td></tr><tr><td>Divide_AboveBurnikelZieglerThreshold</td><td>.NET 10.0</td><td>64</td><td>10,766.7 ns</td><td>1.00</td></tr><tr><td>Divide_AboveBurnikelZieglerThreshold</td><td>.NET 11.0</td><td>64</td><td>6,303.4 ns</td><td>0.59</td></tr><tr><td>Multiply</td><td>.NET 10.0</td><td>64</td><td>2,359.5 ns</td><td>1.00</td></tr><tr><td>Multiply</td><td>.NET 11.0</td><td>64</td><td>1,328.2 ns</td><td>0.56</td></tr><tr><td>ShiftLeft</td><td>.NET 10.0</td><td>64</td><td>217.1 ns</td><td>1.00</td></tr><tr><td>ShiftLeft</td><td>.NET 11.0</td><td>64</td><td>121.6 ns</td><td>0.56</td></tr><tr><td>ParseLargeDecimal</td><td>.NET 10.0</td><td>64</td><td>9,111,926.1 ns</td><td>1.00</td></tr><tr><td>ParseLargeDecimal</td><td>.NET 11.0</td><td>64</td><td>3,899,478.0 ns</td><td>0.43</td></tr><tr><td>ToStringLargeDecimal</td><td>.NET 10.0</td><td>64</td><td>135,894,135.4 ns</td><td>1.00</td></tr><tr><td>ToStringLargeDecimal</td><td>.NET 11.0</td><td>64</td><td>7,868,359.3 ns</td><td>0.058</td></tr><tr><td>Divide_BelowBurnikelZieglerThreshold</td><td>.NET 10.0</td><td>512</td><td>2,894.6 ns</td><td>1.00</td></tr><tr><td>Divide_BelowBurnikelZieglerThreshold</td><td>.NET 11.0</td><td>512</td><td>1,482.0 ns</td><td>0.51</td></tr><tr><td>Divide_AboveBurnikelZieglerThreshold</td><td>.NET 10.0</td><td>512</td><td>10,751.7 ns</td><td>1.00</td></tr><tr><td>Divide_AboveBurnikelZieglerThreshold</td><td>.NET 11.0</td><td>512</td><td>6,317.8 ns</td><td>0.59</td></tr><tr><td>Multiply</td><td>.NET 10.0</td><td>512</td><td>68,063.6 ns</td><td>1.00</td></tr><tr><td>Multiply</td><td>.NET 11.0</td><td>512</td><td>35,443.5 ns</td><td>0.52</td></tr><tr><td>ShiftLeft</td><td>.NET 10.0</td><td>512</td><td>673.0 ns</td><td>1.00</td></tr><tr><td>ShiftLeft</td><td>.NET 11.0</td><td>512</td><td>309.4 ns</td><td>0.46</td></tr><tr><td>ParseLargeDecimal</td><td>.NET 10.0</td><td>512</td><td>9,149,793.0 ns</td><td>1.00</td></tr><tr><td>ParseLargeDecimal</td><td>.NET 11.0</td><td>512</td><td>3,891,313.0 ns</td><td>0.43</td></tr><tr><td>ToStringLargeDecimal</td><td>.NET 10.0</td><td>512</td><td>135,829,594.6 ns</td><td>1.00</td></tr><tr><td>ToStringLargeDecimal</td><td>.NET 11.0</td><td>512</td><td>7,880,631.1 ns</td><td>0.058</td></tr></tbody></table></div>\n<p>In addition to internal changes, <code>BigInteger</code> also gained new public APIs that avoid transcoding. Protocols and storage formats increasingly expose text as UTF-8 bytes, but the previous parsing and formatting APIs required UTF-16 characters. Callers therefore had to decode the input into a temporary string before parsing, or format into characters and encode the result back to bytes. dotnet/runtime#117745 adds direct UTF-8 parsing and formatting to both <code>BigInteger</code> and <code>Complex</code>, sharing the generic numeric machinery used for UTF-16 and letting those consumers operate on their original representation.</p>\n<p>dotnet/runtime#130721 improves a different <code>BigInteger</code> boundary: casting to <code>double</code> and <code>float</code>. The general conversion needs to inspect the arbitrary-width magnitude, locate its highest set bits, and perform the rounding required by the target floating-point format. But many <code>BigInteger</code> instances are much smaller than that machinery is designed for… the implementation now recognizes values that fit in 64 bits and routes them through the hardware’s native integer conversion support.</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Numerics;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private readonly BigInteger _small = (BigInteger.One &lt;&lt; 63) + 123;\n    private readonly BigInteger _large = (BigInteger.One &lt;&lt; 1023) + (BigInteger.One &lt;&lt; 511) + 123;\n    [Benchmark] public double SmallToDouble() =&gt; (double)_small;\n    [Benchmark] public float SmallToSingle() =&gt; (float)_small;\n    [Benchmark] public double LargeToDouble() =&gt; (double)_large;\n    [Benchmark] public float LargeToSingle() =&gt; (float)_large;\n}</code></pre>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Mean</th><th>Ratio</th></tr></thead><tbody><tr><td>SmallToDouble</td><td>.NET 10.0</td><td>2.873 ns</td><td>1.00</td></tr><tr><td>SmallToDouble</td><td>.NET 11.0</td><td>1.764 ns</td><td>0.61</td></tr><tr><td>SmallToSingle</td><td>.NET 10.0</td><td>3.548 ns</td><td>1.00</td></tr><tr><td>SmallToSingle</td><td>.NET 11.0</td><td>1.764 ns</td><td>0.50</td></tr><tr><td>LargeToDouble</td><td>.NET 10.0</td><td>2.863 ns</td><td>1.00</td></tr><tr><td>LargeToDouble</td><td>.NET 11.0</td><td>2.797 ns</td><td>0.98</td></tr><tr><td>LargeToSingle</td><td>.NET 10.0</td><td>3.559 ns</td><td>1.00</td></tr><tr><td>LargeToSingle</td><td>.NET 11.0</td><td>2.849 ns</td><td>0.80</td></tr></tbody></table></div>\n<p>The same limb-widening advantages given to <code>BigInteger</code> in .NET 11 were also extended to the core floating-point types. Parsing a very long decimal input and formatting a floating-point value with many requested digits both need temporary arbitrary-precision arithmetic once the value no longer fits in the normal mantissa. .NET uses a separate internal <code>Number.BigInteger</code> for that work. dotnet/runtime#132577 applies the same native-width limb representation to that type, reducing the amount of per-limb work in floating-point parsing, formatting, and rounding.</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing System.Globalization;\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private readonly string _longFraction = &quot;0.&quot; + new string(&#39;1&#39;, 768);\n    [Benchmark]\n    public double ParseLongFraction() =&gt; double.Parse(_longFraction, CultureInfo.InvariantCulture);\n    [Benchmark]\n    public string FormatSubnormal() =&gt; double.Epsilon.ToString(&quot;G99&quot;, CultureInfo.InvariantCulture);\n}</code></pre>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Mean</th><th>Ratio</th></tr></thead><tbody><tr><td>ParseLongFraction</td><td>.NET 10.0</td><td>8.592 μs</td><td>1.00</td></tr><tr><td>ParseLongFraction</td><td>.NET 11.0</td><td>3.569 μs</td><td>0.42</td></tr><tr><td>FormatSubnormal</td><td>.NET 10.0</td><td>6.884 μs</td><td>1.00</td></tr><tr><td>FormatSubnormal</td><td>.NET 11.0</td><td>1.237 μs</td><td>0.18</td></tr></tbody></table></div>\n<p>The .NET 11 improvements aren’t limited to the scalar representations underlying\n<code>BigInteger</code> and floating-point parsing and formatting. Other numerical types improve as well. Consider <code>Matrix4x4</code>. A 4×4 matrix\ndeterminant combines products of many independent matrix elements, making it a\nnatural fit for SIMD. dotnet/runtime#123954\nfrom @alexcovington adds an SSE\nimplementation of <code>Matrix4x4.GetDeterminant</code>, evaluating several of those\nproducts in parallel:</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Numerics;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private readonly Matrix4x4 _matrix =\n        Matrix4x4.CreateFromYawPitchRoll(0.4f, 0.8f, 1.1f) *\n        Matrix4x4.CreateTranslation(1.5f, -2.5f, 3.25f) *\n        Matrix4x4.CreateScale(1.1f, 0.9f, 1.05f);\n    [Benchmark]\n    public float GetDeterminant() =&gt; _matrix.GetDeterminant();\n}</code></pre>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Mean</th><th>Ratio</th></tr></thead><tbody><tr><td>GetDeterminant</td><td>.NET 10.0</td><td>3.836 ns</td><td>1.00</td></tr><tr><td>GetDeterminant</td><td>.NET 11.0</td><td>2.645 ns</td><td>0.69</td></tr></tbody></table></div>\n<p>The <code>System.Numerics.Tensors</code> APIs are designed to perform the same numerical operation over many values, making them a natural fit for SIMD. dotnet/runtime#126052 adds vector implementations of inverse sine to the portable vector types and uses them in <code>TensorPrimitives.Asin</code>. The tensor loop now evaluates a polynomial approximation for several inputs together, with special handling near the ends of the function’s <code>[-1, 1]</code> domain, rather than calling <code>MathF.Asin</code> or <code>Math.Asin</code> separately for every element:</p>\n<pre><code>// Run separately so each target uses its matching System.Numerics.Tensors package:\n// dotnet run -c Release -f net10.0 --filter &quot;*&quot;\n// dotnet run -c Release -f net11.0 --filter &quot;*&quot;\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Numerics.Tensors;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private const int Length = 4096;\n    private float[] _floatsIn = new float[Length];\n    private float[] _floatsOut = new float[Length];\n    private double[] _doublesIn = new double[Length];\n    private double[] _doublesOut = new double[Length];\n    [GlobalSetup]\n    public void Setup()\n    {\n        Random rng = new(42);\n        for (int i = 0; i &lt; Length; i++)\n        {\n            float v = (float)((rng.NextDouble() * 2.0) - 1.0); // Asin&#39;s domain is [-1, 1]\n            _floatsIn[i] = v;\n            _doublesIn[i] = v;\n        }\n    }\n    [Benchmark]\n    public float AsinFloat()\n    {\n        TensorPrimitives.Asin(_floatsIn, _floatsOut);\n        return _floatsOut[0];\n    }\n    [Benchmark]\n    public double AsinDouble()\n    {\n        TensorPrimitives.Asin(_doublesIn, _doublesOut);\n        return _doublesOut[0];\n    }\n}</code></pre>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Mean</th><th>Ratio</th></tr></thead><tbody><tr><td>AsinFloat</td><td>.NET 10.0</td><td>33.72 μs</td><td>1.00</td></tr><tr><td>AsinFloat</td><td>.NET 11.0</td><td>8.240 μs</td><td>0.24</td></tr><tr><td>AsinDouble</td><td>.NET 10.0</td><td>35.80 μs</td><td>1.00</td></tr><tr><td>AsinDouble</td><td>.NET 11.0</td><td>10.975 μs</td><td>0.31</td></tr></tbody></table></div>\n<p><code>TensorPrimitives</code> also picked up a few more targeted SIMD improvements. For floating-point values, <code>BitIncrement</code> and <code>BitDecrement</code> move to the immediately adjacent representable value; despite their names, they can’t simply add or subtract one, as they also need to handle signed zero, infinities, and NaNs correctly. dotnet/runtime#123610 and dotnet/runtime#123754 process multiple <code>float</code>/<code>double</code> and <code>Half</code> values at once, respectively. The <code>Half</code> path works directly with the raw <code>ushort</code> bit patterns, avoiding conversion to <code>float</code> and back, and both paths use vector masks and conditional selection rather than calling a scalar helper for every element.</p>\n<pre><code>// Run separately so each target uses its matching System.Numerics.Tensors package:\n// dotnet run -c Release -f net10.0 --filter &quot;*&quot;\n// dotnet run -c Release -f net11.0 --filter &quot;*&quot;\nusing System.Numerics.Tensors;\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private const int Length = 4096;\n    private readonly float[] _floats = new float[Length];\n    private readonly float[] _floatDestination = new float[Length];\n    private readonly double[] _doubles = new double[Length];\n    private readonly double[] _doubleDestination = new double[Length];\n    private readonly Half[] _halves = new Half[Length];\n    private readonly Half[] _halfDestination = new Half[Length];\n    [GlobalSetup]\n    public void Setup()\n    {\n        for (int i = 0; i &lt; Length; i++)\n        {\n            float value = (i &amp; 7) switch\n            {\n                0 =&gt; 0,\n                1 =&gt; -0.0f,\n                2 =&gt; float.PositiveInfinity,\n                3 =&gt; float.NegativeInfinity,\n                4 =&gt; float.NaN,\n                _ =&gt; i / 7.0f,\n            };\n            _floats[i] = value;\n            _doubles[i] = value;\n            _halves[i] = (Half)value;\n        }\n    }\n    [Benchmark]\n    public void BitIncrementFloat() =&gt; TensorPrimitives.BitIncrement(_floats, _floatDestination);\n    [Benchmark]\n    public void BitIncrementDouble() =&gt; TensorPrimitives.BitIncrement(_doubles, _doubleDestination);\n    [Benchmark]\n    public void BitIncrementHalf() =&gt; TensorPrimitives.BitIncrement(_halves, _halfDestination);\n}</code></pre>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Mean</th><th>Ratio</th></tr></thead><tbody><tr><td>BitIncrementFloat</td><td>.NET 10.0</td><td>4.223 μs</td><td>1.00</td></tr><tr><td>BitIncrementFloat</td><td>.NET 11.0</td><td>1,071.1 ns</td><td>0.25</td></tr><tr><td>BitIncrementDouble</td><td>.NET 10.0</td><td>4.223 μs</td><td>1.00</td></tr><tr><td>BitIncrementDouble</td><td>.NET 11.0</td><td>2,140.6 ns</td><td>0.51</td></tr><tr><td>BitIncrementHalf</td><td>.NET 10.0</td><td>3.918 μs</td><td>1.00</td></tr><tr><td>BitIncrementHalf</td><td>.NET 11.0</td><td>573.6 ns</td><td>0.15</td></tr></tbody></table></div>\n<p>dotnet/runtime#124280 removes a more mechanical cost from <code>TensorPrimitives.Round</code>: for <code>digits == 0</code>, the old code invoked a full-span rounding kernel and then continued through another full-span pass. Returning immediately removes that redundant traversal and overwrite of the destination.</p>\n<pre><code>// Run separately so each target uses its matching System.Numerics.Tensors package:\n// dotnet run -c Release -f net10.0 --filter &quot;*&quot;\n// dotnet run -c Release -f net11.0 --filter &quot;*&quot;\nusing System.Numerics.Tensors;\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private const int Length = 4096;\n    private readonly float[] _source = new float[Length];\n    private readonly float[] _destination = new float[Length];\n    [Benchmark]\n    public void RoundZero() =&gt; TensorPrimitives.Round(_source, 0, MidpointRounding.ToEven, _destination);\n}</code></pre>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Mean</th><th>Ratio</th></tr></thead><tbody><tr><td>RoundZero</td><td>.NET 10.0</td><td>882.7 ns</td><td>1.00</td></tr><tr><td>RoundZero</td><td>.NET 11.0</td><td>205.5 ns</td><td>0.23</td></tr></tbody></table></div>\n<p><code>Half</code> comparisons are faster as well. Previously, <code>Half.CompareTo</code> separately\nasked whether one value was less than, greater than, or equal to the other,\nrepeating the special handling required for NaN and signed zero each time.\ndotnet/runtime#131297 performs\nthat work once and then arranges the underlying bits into a form that can be\ncompared directly, while still treating <code>+0</code> and <code>-0</code> as equal. On x64 with\nAVX2, it also makes <code>CompareTo</code>, <code>&lt;</code>, and <code>&lt;=</code> faster by converting the operands\nto <code>float</code>, which the hardware can do very efficiently. Equality remains\nbit-based, as that’s already the cheaper approach.</p>\n<pre><code>// Run on x64 with AVX2:\n// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private readonly Half[] _left = Enumerable.Range(0, 4096).Select(i =&gt; (Half)(i - 2048)).ToArray();\n    private readonly Half[] _right = Enumerable.Range(0, 4096).Select(i =&gt; (Half)(2048 - i)).ToArray();\n    [Benchmark]\n    public int CompareTo()\n    {\n        int sum = 0;\n        for (int i = 0; i &lt; _left.Length; i++)\n            sum += _left[i].CompareTo(_right[i]);\n        return sum;\n    }\n    [Benchmark]\n    public int LessThan()\n    {\n        int count = 0;\n        for (int i = 0; i &lt; _left.Length; i++)\n            count += _left[i] &lt; _right[i] ? 1 : 0;\n        return count;\n    }\n}</code></pre>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Mean</th><th>Ratio</th></tr></thead><tbody><tr><td>CompareTo</td><td>.NET 10.0</td><td>9.340 μs</td><td>1.00</td></tr><tr><td>CompareTo</td><td>.NET 11.0</td><td>5.745 μs</td><td>0.62</td></tr><tr><td>LessThan</td><td>.NET 10.0</td><td>7.514 μs</td><td>1.00</td></tr><tr><td>LessThan</td><td>.NET 11.0</td><td>5.672 μs</td><td>0.75</td></tr></tbody></table></div>\n<p>Multiplying two 64-bit integers produces a 128-bit result, and x64 has instructions that provide both 64-bit halves directly. dotnet/runtime#117261 from @Daniel-Svensson exposes those signed and unsigned forms through an <code>X86Base.X64.BigMul</code> intrinsic. <code>Math.BigMul</code> can then map directly to <code>imul</code> or <code>mul</code> and return both halves in registers, avoiding the extra instructions and register shuffling required by the previous paths.</p>\n<pre><code>// Run on x64:\n// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser, HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private readonly long _signedLeft = 0x1234_5678_9ABC_DEF;\n    private readonly long _signedRight = 0x0FED_CBA9_8765_432;\n    private readonly ulong _unsignedLeft = 0xFEDC_BA98_7654_3210;\n    private readonly ulong _unsignedRight = 0x1234_5678_9ABC_DEF0;\n    [Benchmark]\n    public long Signed()\n    {\n        long high = Math.BigMul(_signedLeft, _signedRight, out long low);\n        return high ^ low;\n    }\n    [Benchmark]\n    public ulong Unsigned()\n    {\n        ulong high = Math.BigMul(_unsignedLeft, _unsignedRight, out ulong low);\n        return high ^ low;\n    }\n}</code></pre>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Mean</th><th>Ratio</th><th>Code Size</th></tr></thead><tbody><tr><td>Signed</td><td>.NET 10.0</td><td>2.210 ns</td><td>1.00</td><td>65 B</td></tr><tr><td>Signed</td><td>.NET 11.0</td><td>1.344 ns</td><td>0.61</td><td>12 B</td></tr><tr><td>Unsigned</td><td>.NET 10.0</td><td>1.446 ns</td><td>1.00</td><td>39 B</td></tr><tr><td>Unsigned</td><td>.NET 11.0</td><td>1.323 ns</td><td>0.92</td><td>12 B</td></tr></tbody></table></div>\n<p>Fixed-format numeric and identifier helpers benefit from a much simpler technique: establish the exact span length once, then let the JIT reuse that fact. dotnet/runtime#119254 from @xtqqczze applies that pattern in <code>Decimal</code>, <code>Guid</code>, and <code>IPAddress</code>, removing repeated bounds checks.</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private const string Value = &quot;a8098c1a-f86e-11da-bd1a-00112444be1e&quot;;\n    [Benchmark]\n    public bool TryParseExactD() =&gt; Guid.TryParseExact(Value, &quot;D&quot;, out _);\n}</code></pre>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Mean</th><th>Ratio</th></tr></thead><tbody><tr><td>TryParseExactD</td><td>.NET 10.0</td><td>15.94 ns</td><td>1.00</td></tr><tr><td>TryParseExactD</td><td>.NET 11.0</td><td>12.60 ns</td><td>0.79</td></tr></tbody></table></div>\n<p><code>Guid</code> has been improving every .NET release, and sees several improvements in .NET 11. Whenever possible, .NET tries to maintain similar performance and behaviors across operating systems, but low-level functionality often simply delegates to the operating system, exposing that OS’ characteristics. When it comes to random number generation, historically cryptographically-secure random number generation, as is used in <code>Guid.NewGuid</code>, has been a bit slower on Linux than on Windows due to using <code>/dev/urandom</code> as the source of entropy. dotnet/runtime#123540 from @reedz moves <code>Guid.NewGuid()</code> off of that file-descriptor path to the <code>getrandom()</code> syscall, avoiding descriptor setup and reads through the file abstraction.</p>\n<p>And on the subject of randomness, dotnet/runtime#119890 from @hamarb123 removes two pieces of work from <code>Random.Shuffle</code>: an unnecessary copy of the span length and a branch that skipped swapping an element with itself. A self-swap is harmless and uncommon, while testing for it adds an unpredictable branch to every iteration. The difference is most visible for short arrays and small value types, where the swap itself is cheap:</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    [Params(16, 4096)]\n    public int Length;\n    private readonly Random _random = new(42);\n    private int[] _values = [];\n    [GlobalSetup]\n    public void Setup() =&gt; _values = Enumerable.Range(0, Length).ToArray();\n    [Benchmark]\n    public int ShuffleSmallValueType()\n    {\n        _random.Shuffle(_values);\n        return _values[0] + _values[^1];\n    }\n}</code></pre>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Length</th><th>Mean</th><th>Ratio</th></tr></thead><tbody><tr><td>ShuffleSmallValueType</td><td>.NET 10.0</td><td>16</td><td>138.9 ns</td><td>1.00</td></tr><tr><td>ShuffleSmallValueType</td><td>.NET 11.0</td><td>16</td><td>89.34 ns</td><td>0.64</td></tr><tr><td>ShuffleSmallValueType</td><td>.NET 10.0</td><td>4096</td><td>25,925.1 ns</td><td>1.00</td></tr><tr><td>ShuffleSmallValueType</td><td>.NET 11.0</td><td>4096</td><td>14,151.08 ns</td><td>0.55</td></tr></tbody></table></div>\n<p><code>Random</code> itself picked up a small but pointed code-generation fix. <code>Random.InternalSample</code> contains a condition that’s inherently hard for the processor to predict, so it’s better implemented with conditional instructions than with a branch. The JIT’s if-conversion support we previously discussed would have been able to do that transformation, except it doesn’t currently support if-conversion inside of loops, which is a pretty common place to find an inlined <code>Random.Next</code> call. dotnet/runtime#131714 marks the helper as <code>[MethodImpl(MethodImplOptions.NoInlining)]</code> to preserve the branch-free form; once the JIT can perform if-conversion inside loops, that annotation can be reconsidered.</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private readonly Random _random = new(42);\n    [Benchmark]\n    public int Next()\n    {\n        int sum = 0;\n        for (int i = 0; i &lt; 1024; i++)\n            sum += _random.Next();\n        return sum;\n    }\n}</code></pre>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Mean</th><th>Ratio</th></tr></thead><tbody><tr><td>Next</td><td>.NET 10.0</td><td>5.954 μs</td><td>1.00</td></tr><tr><td>Next</td><td>.NET 11.0</td><td>3.275 μs</td><td>0.55</td></tr></tbody></table></div>\n<h2 id=\"globalization\">Globalization</h2>\n<p>Many globalization-related APIs sit atop data that can be expensive to locate\nand interpret. <code>DateTime.Now</code>, for example, depends on time-zone transition\ndata, while casing and parsing depend on native globalization services and\nculture-specific tables.</p>\n<p>dotnet/runtime#119662 substantially reworks <code>TimeZoneInfo</code> around that observation. Determining an offset isn’t always a fixed arithmetic operation: daylight-saving rules can vary by year, and historical rules can contain multiple transitions and exceptional cases. Once the transitions for a zone and year have been interpreted, however, other conversions in that year can reuse them. Similarly, the local offset used by <code>DateTime.Now</code> can’t change between transition instants. Conversions now reuse cached per-year transition data rather than repeatedly walking adjustment rules, while <code>DateTime.Now</code> caches the active UTC offset together with the instant at which it next needs to be recomputed.</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private readonly DateTime _utc = new(2026, 7, 15, 12, 0, 0, DateTimeKind.Utc);\n    private readonly DateTime _local = new(2026, 7, 15, 5, 0, 0, DateTimeKind.Unspecified);\n    private readonly TimeZoneInfo _zone = TimeZoneInfo.FindSystemTimeZoneById(\n        OperatingSystem.IsWindows() ? &quot;Pacific Standard Time&quot; : &quot;America/Los_Angeles&quot;);\n    [Benchmark]\n    public DateTime ConvertTimeFromUtc() =&gt; TimeZoneInfo.ConvertTimeFromUtc(_utc, _zone);\n    [Benchmark]\n    public DateTime ConvertTimeToUtc() =&gt; TimeZoneInfo.ConvertTimeToUtc(_local, _zone);\n    [Benchmark]\n    public DateTime GetLocalNow() =&gt; DateTime.Now;\n}</code></pre>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Mean</th><th>Ratio</th></tr></thead><tbody><tr><td>ConvertTimeFromUtc</td><td>.NET 10.0</td><td>45.13 ns</td><td>1.00</td></tr><tr><td>ConvertTimeFromUtc</td><td>.NET 11.0</td><td>19.44 ns</td><td>0.43</td></tr><tr><td>ConvertTimeToUtc</td><td>.NET 10.0</td><td>51.97 ns</td><td>1.00</td></tr><tr><td>ConvertTimeToUtc</td><td>.NET 11.0</td><td>20.23 ns</td><td>0.39</td></tr><tr><td>GetLocalNow</td><td>.NET 10.0</td><td>76.41 ns</td><td>1.00</td></tr><tr><td>GetLocalNow</td><td>.NET 11.0</td><td>34.39 ns</td><td>0.45</td></tr></tbody></table></div>\n<p>dotnet/runtime#120685 separates two costs in invariant casing. With the normal globalization configuration, <code>ToUpperInvariant</code> and <code>ToLowerInvariant</code> now try a managed ASCII path first, so casing ASCII text can avoid or delay initialization of ICU, the native library .NET uses for culture-aware globalization. In invariant-globalization mode, where ICU isn’t loaded at all, that managed path also improves ASCII casing throughput. Non-ASCII input still needs the appropriate globalization path.</p>\n<pre><code>// DOTNET_SYSTEM_GLOBALIZATION_INVARIANT=1 dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private readonly string _short = &quot;runtime&quot;;\n    private readonly string _long = new(&#39;a&#39;, 139);\n    [Benchmark]\n    public string ShortAscii() =&gt; _short.ToUpperInvariant();\n    [Benchmark]\n    public string LongAscii() =&gt; _long.ToUpperInvariant();\n}</code></pre>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Mean</th><th>Ratio</th><th>Allocated</th></tr></thead><tbody><tr><td>ShortAscii</td><td>.NET 10.0</td><td>18.12 ns</td><td>1.00</td><td>40 B</td></tr><tr><td>ShortAscii</td><td>.NET 11.0</td><td>14.62 ns</td><td>0.81</td><td>40 B</td></tr><tr><td>LongAscii</td><td>.NET 10.0</td><td>248.38 ns</td><td>1.00</td><td>304 B</td></tr><tr><td>LongAscii</td><td>.NET 11.0</td><td>39.22 ns</td><td>0.16</td><td>304 B</td></tr></tbody></table></div>\n<p>Several smaller changes remove setup around date and culture data. dotnet/runtime#123886 allocates the <code>DateTimeFormatInfo</code> date-word table only for cultures that actually contain such words. And dotnet/runtime#122918 replaces synchronized, boxing <code>Hashtable</code> caches used by time-zone and encoding tables with typed <code>ConcurrentDictionary</code> instances.</p>\n<p>The round-trip <code>&quot;O&quot;</code> date format always contains exactly seven fractional-second digits, matching the 10,000,000 ticks in a second. dotnet/runtime#129005 parses those digits directly as ticks, avoiding a conversion through <code>double</code> followed by division, multiplication, and rounding. Formatting benefits from specialization as well. dotnet/runtime#129374 routes invariant <code>DateTime.ToString(&quot;G&quot;)</code> through the existing fixed-format fast path, bypassing the general culture-aware formatter. <code>DateTimeOffset</code> retains the general path because its offset changes the output:</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Running;\nusing System.Globalization;\nusing BenchmarkDotNet.Attributes;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private readonly DateTime _dateTime = new(2024, 3, 15, 13, 45, 30, DateTimeKind.Utc);\n    [Benchmark]\n    public string DateTime_ToString_G() =&gt; _dateTime.ToString(&quot;G&quot;, CultureInfo.InvariantCulture);\n}</code></pre>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Mean</th><th>Ratio</th></tr></thead><tbody><tr><td>DateTime_ToString_G</td><td>.NET 10.0</td><td>65.54 ns</td><td>1.00</td></tr><tr><td>DateTime_ToString_G</td><td>.NET 11.0</td><td>29.76 ns</td><td>0.45</td></tr></tbody></table></div>\n<h2 id=\"strings-and-spans\">Strings and Spans</h2>\n<p>UTF-8 is everywhere, from web protocols and JSON payloads to files on disk. Since .NET strings use UTF-16, applications frequently need to convert between the two, making it especially important for those conversions to be fast. UTF-8 encoding must validate UTF-16 surrogate pairs as it counts and converts them. On Arm64, the vectorized implementation in .NET 10 still examined individual elements when counting the resulting UTF-8 bytes and checking that surrogates were correctly paired. That gets expensive for text containing many supplementary characters, as every surrogate-heavy vector falls back to this element-by-element work. dotnet/runtime#121981 from @ylpoonlg instead performs the counting and surrogate checks with vector-wide operations. As part of that work, it also unifies most of the x86 and Arm64 implementations, retaining small platform-specific helpers where the instruction sets differ:</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Text;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private const int Length = 4096;\n    private string _validWithSurrogatePairs = string.Empty;\n    [GlobalSetup]\n    public void Setup()\n    {\n        Random rng = new(42);\n        StringBuilder sb = new(Length);\n        while (sb.Length &lt; Length - 2)\n        {\n            sb.Append((char)(&#39;A&#39; + rng.Next(0, 26)));\n            sb.Append(&quot;\\U0001F600&quot;); // emoji -&gt; surrogate pair\n        }\n        _validWithSurrogatePairs = sb.ToString();\n    }\n    [Benchmark]\n    public int ValidWithSurrogatePairs() =&gt; Encoding.UTF8.GetByteCount(_validWithSurrogatePairs);\n}</code></pre>\n<p>This input deliberately contains a surrogate pair for every ASCII character, making the removed per-element work especially visible.</p>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Mean</th><th>Ratio</th></tr></thead><tbody><tr><td>ValidWithSurrogatePairs</td><td>.NET 10.0</td><td>2.994 μs</td><td>1.00</td></tr><tr><td>ValidWithSurrogatePairs</td><td>.NET 11.0</td><td>856.8 ns</td><td>0.29</td></tr></tbody></table></div>\n<p>The byte-to-char direction was also improved on Arm. UTF-8 decoding can copy ASCII bytes directly to UTF-16 characters, but as soon as we find the first non-ASCII byte, we need the full multi-byte decoder. The vector loop therefore needs both a fast test for whether any lane is non-ASCII and, only when one is found, its exact position. Calculating that position for every all-ASCII vector wastes work on the overwhelmingly common fast path. dotnet/runtime#121382 from @ylpoonlg first performs the cheap vector-wide test and then computes the lane index only after that test succeeds.</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Text;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private readonly byte[] _ascii = Enumerable.Repeat((byte)&#39;a&#39;, 16_384).ToArray();\n    [Benchmark]\n    public int Utf8GetCharCount() =&gt; Encoding.UTF8.GetCharCount(_ascii);\n}</code></pre>\n<p>With all-ASCII input, every vector can stay on the cheap path:</p>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Mean</th><th>Ratio</th></tr></thead><tbody><tr><td>Utf8GetCharCount</td><td>.NET 10.0</td><td>443.5 ns</td><td>1.00</td></tr><tr><td>Utf8GetCharCount</td><td>.NET 11.0</td><td>210.6 ns</td><td>0.47</td></tr></tbody></table></div>\n<p>Base64 is commonly used when binary data needs to travel through\ntext-oriented formats and protocols. Its encoder naturally works in groups of\nthree input bytes and four output characters, but the line-breaking option\nalso needs to stop at the MIME-style 76-character boundary and insert <code>\\r\\n</code>.\nThe older implementation handled that formatting through a separate scalar\npath. dotnet/runtime#123403 brings the optimized span-based Base64 encoder to <code>Convert.ToBase64String</code> with <code>Base64FormattingOptions.InsertLineBreaks</code>, processing each line with the same vectorized core and handling the separators around it:</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Running;\nusing BenchmarkDotNet.Attributes;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    [Params(57, 570)]\n    public int ByteLength { get; set; }\n    private byte[] _bytes = [];\n    [GlobalSetup]\n    public void Setup()\n    {\n        _bytes = new byte[ByteLength];\n        new Random(42).NextBytes(_bytes);\n    }\n    [Benchmark]\n    public string ToBase64String_InsertLineBreaks() =&gt; Convert.ToBase64String(_bytes, Base64FormattingOptions.InsertLineBreaks);\n}</code></pre>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>ByteLength</th><th>Mean</th><th>Ratio</th></tr></thead><tbody><tr><td>ToBase64String_InsertLineBreaks</td><td>.NET 10.0</td><td>57</td><td>60.95 ns</td><td>1.00</td></tr><tr><td>ToBase64String_InsertLineBreaks</td><td>.NET 11.0</td><td>57</td><td>23.91 ns</td><td>0.39</td></tr><tr><td>ToBase64String_InsertLineBreaks</td><td>.NET 10.0</td><td>570</td><td>560.66 ns</td><td>1.00</td></tr><tr><td>ToBase64String_InsertLineBreaks</td><td>.NET 11.0</td><td>570</td><td>194.99 ns</td><td>0.35</td></tr></tbody></table></div>\n<p>Base64 decoding got the same treatment from the other direction. <code>Base64.DecodeFromUtf8InPlace</code> decodes in place, overwriting the encoded input with the decoded bytes. In .NET 10, it still employed a scalar loop, long after the out-of-place <code>DecodeFromUtf8</code> had acquired AVX-512, AVX2, AdvSimd, and SSSE3 paths. In-place decoding turns out to be safe to vectorize precisely because of Base64’s ratio: 4 bytes read produce 3 bytes written, so the write cursor always trails the read cursor, and each vector store, including its zero-padded overshoot, ends at or before the next vector load and never clobbers source that hasn’t been read yet. dotnet/runtime#131333 therefore reuses the existing decode helpers for the in-place path.</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing System.Buffers;\nusing System.Buffers.Text;\nusing System.Text;\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private readonly byte[] _encoded = Encoding.ASCII.GetBytes(Convert.ToBase64String(new byte[16_384]));\n    private byte[] _buffer = [];\n    [IterationSetup]\n    public void Setup() =&gt; _buffer = (byte[])_encoded.Clone();\n    [Benchmark]\n    public OperationStatus Decode() =&gt; Base64.DecodeFromUtf8InPlace(_buffer, out _);\n}</code></pre>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Mean</th><th>Ratio</th></tr></thead><tbody><tr><td>Decode</td><td>.NET 10.0</td><td>10.70 μs</td><td>1.00</td></tr><tr><td>Decode</td><td>.NET 11.0</td><td>2.256 μs</td><td>0.21</td></tr></tbody></table></div>\n<p><code>MemoryExtensions.CommonPrefixLength</code> compares two spans and returns how many\nelements they share at the beginning (“hello” and “help”, for example, have a\ncommon prefix length of 3). Internally, it utilizes a helper that slices whichever input was longer to\nthe length of the shorter one. dotnet/runtime#121104\nfrom @xtqqczze simplifies that helper: after\nshortening the second span if necessary, it always slices the first span to\nthe second’s length. That gives the JIT the same explicit relationship between\nthe two lengths regardless of which input started out longer.</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System;\nusing System.Linq;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private readonly string[] _shorter = Enumerable.Repeat(&quot;value&quot;, 64).ToArray();\n    private readonly string[] _longer = Enumerable.Repeat(&quot;value&quot;, 128).ToArray();\n    [Benchmark]\n    public int ShorterFirst() =&gt; _shorter.AsSpan().CommonPrefixLength(_longer);\n    [Benchmark]\n    public int LongerFirst() =&gt; _longer.AsSpan().CommonPrefixLength(_shorter);\n}</code></pre>\n<p>The longer-first case was already efficient. The change improves the shorter-first case, bringing the two orderings to essentially the same throughput:</p>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Mean</th><th>Ratio</th></tr></thead><tbody><tr><td>ShorterFirst</td><td>.NET 10.0</td><td>45.16 ns</td><td>1.00</td></tr><tr><td>ShorterFirst</td><td>.NET 11.0</td><td>25.43 ns</td><td>0.56</td></tr><tr><td>LongerFirst</td><td>.NET 10.0</td><td>26.40 ns</td><td>1.00</td></tr><tr><td>LongerFirst</td><td>.NET 11.0</td><td>26.31 ns</td><td>1.00</td></tr></tbody></table></div>\n<p>Text processing often starts by obtaining an <code>Encoding</code>. Properties such as\n<code>Encoding.UTF8</code> provide fast access to popular encodings, while legacy code\npages can be made available by registering <code>CodePagesEncodingProvider</code>. In\n.NET 10, that provider’s tables, including the name lookup used by\n<code>Encoding.GetEncoding(string)</code> once the provider is registered, used\nreader-writer locks. dotnet/runtime#125001\nreplaces those caches with <code>ConcurrentDictionary</code> instances, allowing\nwarmed-up provider lookups to proceed without acquiring the reader lock.</p>\n<p>On <code>string</code> itself, dotnet/runtime#130361 from @prozolic recognizes when <code>string.Concat(IEnumerable&lt;string?&gt;)</code> receives a <code>string[]</code> or <code>List&lt;string?&gt;</code> and passes its contiguous storage directly to the span-based implementation:</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing System.Collections.Generic;\nusing System.Linq;\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[MemoryDiagnoser(false), HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private readonly IEnumerable&lt;string?&gt; _array = Enumerable.Range(0, 1_000).Select(i =&gt; i.ToString()).ToArray();\n    private readonly IEnumerable&lt;string?&gt; _list = Enumerable.Range(0, 1_000).Select(i =&gt; i.ToString()).ToList();\n    [Benchmark]\n    public string Array() =&gt; string.Concat(_array);\n    [Benchmark]\n    public string List() =&gt; string.Concat(_list);\n}</code></pre>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Mean</th><th>Ratio</th><th>Allocated</th></tr></thead><tbody><tr><td>Array</td><td>.NET 10.0</td><td>4.230 μs</td><td>1.00</td><td>5.7 KB</td></tr><tr><td>Array</td><td>.NET 11.0</td><td>3.285 μs</td><td>0.78</td><td>5.67 KB</td></tr><tr><td>List</td><td>.NET 10.0</td><td>6.703 μs</td><td>1.00</td><td>5.71 KB</td></tr><tr><td>List</td><td>.NET 11.0</td><td>3.253 μs</td><td>0.49</td><td>5.67 KB</td></tr></tbody></table></div>\n<p>Some of my favorite improvements in .NET are the tiny ones that show up everywhere. A good example of that is in dotnet/roslyn#82729. Previously, when you wrote <code>span[start..]</code>, the compiler would lower that to the equivalent of <code>span.Slice(start, span.Length - start)</code>. The JIT has made strides towards compiling this exactly how it would <code>span.Slice(start)</code>, but everyone is better off if the C# compiler just emits that in the first place. And it now does. The difference is clear in the IL for a method that returns <code>span[start..]</code>:</p>\n<pre><code>; Platform-independent IL\n-// Before: 21 bytes\n+// After: 9 bytes\n-.locals init ([0] System.Span&lt;char&gt;&amp;, [1] int32)\n ldarga.s span\n-stloc.0\n ldarg.1\n-stloc.1\n-ldloc.0\n-ldloc.1\n-ldloc.0\n-call instance int32 System.Span&lt;char&gt;::get_Length()\n-ldloc.1\n-sub\n-call instance System.Span&lt;char&gt; System.Span&lt;char&gt;::Slice(int32, int32)\n+call instance System.Span&lt;char&gt; System.Span&lt;char&gt;::Slice(int32)\n ret</code></pre>\n<h3 id=\"searching-and-comparing\">Searching and Comparing</h3>\n<p>Searching in one way, shape, or form is one of the most common things programs do. And when it comes to searching text, regular expressions are an extremely common and helpful way to specify and perform said search. .NET’s regex support has improved by leaps and bounds over the years, with significant investments in .NET 5 and .NET 7 and then every release since, including .NET 11.</p>\n<p>When a <code>Regex</code> instance is created, it needs to parse the incoming regular expression pattern and turn it into a form it can utilize for performing the actual searches. The regex language is very expressive and enables multiple ways of specifying the same pattern, some more efficient to process than others, so as part of parsing, <code>Regex</code> applies a variety of simplifications and optimizations over the parsed tree in order to put it into an ideal form, as well as to learn facts about the pattern to further optimize later processing (such as discovering a minimum and maximum length of any possible match). Each of these transformations can in turn expose more opportunity for other transformations, but based on the order the transformations are applied, sometimes those opportunities can be missed. In .NET 11, dotnet/runtime#125289 gives compiled and source-generated regexes one final cleanup pass after the whole-pattern optimizations have reshaped the pattern. Consider the pattern <code>[ab]+c[ab]+|[ab]+</code>. On input containing a long run of <code>a</code>s with no <code>c</code>, the .NET 10 source-generated matcher first scans the whole run for the first alternative, fails when it doesn’t find the <code>c</code>, and then scans the same run again for the second alternative. The final cleanup pass factors out the common <code>[ab]+</code>, leaving <code>c[ab]+</code> as an optional suffix:</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot;\n// dotnet run -c Release -f net11.0 --filter &quot;*&quot;\nusing System.Text.RegularExpressions;\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic partial class Benchmarks\n{\n    private readonly string _input = new(&#39;a&#39;, 4096);\n    [Benchmark]\n    public bool SharedPrefix() =&gt; SharedPrefixRegex().IsMatch(_input);\n    [GeneratedRegex(&quot;[ab]+c[ab]+|[ab]+&quot;)]\n    private static partial Regex SharedPrefixRegex();\n}</code></pre>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Mean</th><th>Ratio</th></tr></thead><tbody><tr><td>SharedPrefix</td><td>.NET 10.0</td><td>550.9 ns</td><td>1.00</td></tr><tr><td>SharedPrefix</td><td>.NET 11.0</td><td>282.8 ns</td><td>0.51</td></tr></tbody></table></div>\n<p>Beyond doing additional passes, several other changes improve what those\nanalysis passes can see. For example, for a pattern like <code>(http|https)</code> with\nordinal ignore-case matching, for uninteresting reasons previously the engine\nwould extract a prefix of <code>&quot;htt&quot;</code>, even though it could have extracted\n<code>&quot;http&quot;</code>. dotnet/runtime#124881\nimproves that, enabling the engine to skip far more false candidates. The\ninput here contains 25,000 <code>&quot;htt&quot;</code> prefixes that aren’t followed by a <code>p</code>\nbefore the final match:</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot;\n// dotnet run -c Release -f net11.0 --filter &quot;*&quot;\nusing System.Text.RegularExpressions;\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic partial class Benchmarks\n{\n    private readonly string _input = string.Concat(Enumerable.Repeat(&quot;httx&quot;, 25_000)) + &quot;https&quot;;\n    [Benchmark]\n    public bool IgnoreCaseAlternation() =&gt; Http.IsMatch(_input);\n    [GeneratedRegex(&quot;(http|https)&quot;, RegexOptions.IgnoreCase)]\n    private static partial Regex Http { get; }\n}</code></pre>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Mean</th><th>Ratio</th></tr></thead><tbody><tr><td>IgnoreCaseAlternation</td><td>.NET 10.0</td><td>415.0 μs</td><td>1.00</td></tr><tr><td>IgnoreCaseAlternation</td><td>.NET 11.0</td><td>7.012 μs</td><td>0.017</td></tr></tbody></table></div>\n<p>When those transformation passes are looking for various patterns, sometimes small\nthings obscure what they’re trying to see, and they miss optimizations.\ndotnet/runtime#124842\nimproves a case where captures were getting in the way of identifying a\nsearchable prefix. For a pattern like <code>\\b(in)\\b</code> with\n<code>RegexOptions.IgnoreCase</code>, it will now discover it can search for\nordinal-ignore-case <code>&quot;in&quot;</code>.</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot;\n// dotnet run -c Release -f net11.0 --filter &quot;*&quot;\nusing System.Text.RegularExpressions;\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic partial class Benchmarks\n{\n    private readonly string _input = string.Concat(Enumerable.Repeat(&quot;xn &quot;, 33_333)) + &quot;in&quot;;\n    [Benchmark]\n    public bool IgnoreCaseCapturedPrefix() =&gt; CapturedPrefix.IsMatch(_input);\n    [GeneratedRegex(@&quot;\\b(in)\\b&quot;, RegexOptions.IgnoreCase)]\n    private static partial Regex CapturedPrefix { get; }\n}</code></pre>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Mean</th><th>Ratio</th></tr></thead><tbody><tr><td>IgnoreCaseCapturedPrefix</td><td>.NET 10.0</td><td>277.7 μs</td><td>1.00</td></tr><tr><td>IgnoreCaseCapturedPrefix</td><td>.NET 11.0</td><td>7.286 μs</td><td>0.026</td></tr></tbody></table></div>\n<p>As these cases highlight, one of the most impactful things we can do for regular expression processing is improve the engine’s ability to find things to search for as the next possible place a match could apply, and to optimize that search. dotnet/runtime#124736 does that. For compiled, source-generated, and <code>NonBacktracking</code> regexes, it improves how the engine is able to search for one of several literal prefixes. For <code>agggtaaa|tttaccct</code>, for example, the .NET 10 source generator first searched for <code>[ag]</code> at offset 3 and then checked nearby characters for <code>[gt]</code>. That’s a weak filter for an input full of <code>a</code> characters, where almost every position becomes a candidate. The .NET 11 generator instead searches for the complete <code>agggtaaa</code> and <code>tttaccct</code> strings with <code>SearchValues&lt;string&gt;</code>. A frequency heuristic selects this approach only for case-sensitive alternatives where whole-string searching is expected to reject more false candidates than the available character-set filter.</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot;\n// dotnet run -c Release -f net11.0 --filter &quot;*&quot;\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Text.RegularExpressions;\nBenchmarkSwitcher.FromAssembly(typeof(RegexPrefixBenchmarks).Assembly).Run(args);\n[HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic partial class RegexPrefixBenchmarks\n{\n    private const string Pattern = &quot;agggtaaa|tttaccct&quot;;\n    private readonly string _match = new string(&#39;a&#39;, 100_000) + &quot;tttaccct&quot;;\n    private readonly string _miss = new(&#39;a&#39;, 100_000);\n    [Benchmark]\n    public bool Match() =&gt; Generated.IsMatch(_match);\n    [Benchmark]\n    public bool Miss() =&gt; Generated.IsMatch(_miss);\n    [GeneratedRegex(Pattern)]\n    private static partial Regex Generated { get; }\n}</code></pre>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Mean</th><th>Ratio</th></tr></thead><tbody><tr><td>Match</td><td>.NET 10.0</td><td>861.9 μs</td><td>1.00</td></tr><tr><td>Match</td><td>.NET 11.0</td><td>9.251 μs</td><td>0.011</td></tr><tr><td>Miss</td><td>.NET 10.0</td><td>861.4 μs</td><td>1.00</td></tr><tr><td>Miss</td><td>.NET 11.0</td><td>9.647 μs</td><td>0.011</td></tr></tbody></table></div>\n<p>Of course, searching for the next place to match isn’t the only opportunity for improvement. Once you’ve found that place, you need to try to match, and we want to optimize that further, too.</p>\n<p>Consider the pattern <code>\\b\\w+n\\b</code>. The <code>\\w+</code> can match <code>n</code>, which means we can’t automatically treat this loop as being atomic. Normally, after matching the loop greedily and failing to match <code>n</code>, we’d need to backtrack looking for the next viable place to match the <code>n</code>. But if what comes after the <code>n</code> (in this case, a boundary) can’t possibly match the loop, we can avoid doing that search. dotnet/runtime#125636 teaches the compiled and source-generated engines to prove that and test the final position directly rather than searching backward through the loop’s existing match. The same idea applies to other loops followed by a literal when the engine can prove that trying earlier positions can’t change the result.</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot;\n// dotnet run -c Release -f net11.0 --filter &quot;*&quot;\nusing BenchmarkDotNet.Running;\nusing System.Text.RegularExpressions;\nusing BenchmarkDotNet.Attributes;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic partial class Benchmarks\n{\n    private const int WordLength = 5000;\n    private readonly string _matchingWord = new string(&#39;a&#39;, WordLength - 1) + &quot;n&quot;;\n    private readonly string _nonMatchingWord = new string(&#39;a&#39;, WordLength - 1) + &quot;b&quot;;\n    [GeneratedRegex(@&quot;\\b\\w+n\\b&quot;)]\n    private static partial Regex Generated { get; }\n    [Benchmark]\n    public bool Matching() =&gt; Generated.IsMatch(_matchingWord);\n    [Benchmark]\n    public bool NonMatching() =&gt; Generated.IsMatch(_nonMatchingWord);\n}</code></pre>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Mean</th><th>Ratio</th></tr></thead><tbody><tr><td>Matching</td><td>.NET 10.0</td><td>3.441 μs</td><td>1.00</td></tr><tr><td>Matching</td><td>.NET 11.0</td><td>3.118 μs</td><td>0.91</td></tr><tr><td>NonMatching</td><td>.NET 10.0</td><td>83.656 μs</td><td>1.00</td></tr><tr><td>NonMatching</td><td>.NET 11.0</td><td>69.984 μs</td><td>0.84</td></tr></tbody></table></div>\n<p>A match can sometimes be ruled out before examining any of the input’s\ncharacters. When matching starts at position zero, a fixed-length pattern with\na leading <code>\\A</code> or non-multiline <code>^</code> and a trailing <code>\\z</code> can match only when the\nwhole input has exactly that length.\ndotnet/runtime#120916 emits\nthat length check up front for the compiled and source-generated engines when\nthe computed maximum length equals the minimum required length. Here, the\npattern requires exactly 512 characters while the input contains 513:</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot;\n// dotnet run -c Release -f net11.0 --filter &quot;*&quot;\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Text.RegularExpressions;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic partial class Benchmarks\n{\n    private readonly string _tooLong = new(&#39;a&#39;, 513);\n    [GeneratedRegex(@&quot;\\A[a-z]{512}\\z&quot;)]\n    private static partial Regex Generated { get; }\n    [Benchmark]\n    public bool AnchoredReject() =&gt; Generated.IsMatch(_tooLong);\n}</code></pre>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Mean</th><th>Ratio</th></tr></thead><tbody><tr><td>AnchoredReject</td><td>.NET 10.0</td><td>41.06 ns</td><td>1.00</td></tr><tr><td>AnchoredReject</td><td>.NET 11.0</td><td>16.05 ns</td><td>0.39</td></tr></tbody></table></div>\n<p>In general, we’ve tried to keep the compilers behind <code>RegexOptions.Compiled</code> (which emits IL) and the source generator (which emits C#) as close to 1:1 as possible. There are a few cases, however, where they have diverged from each other, generally where one was able to easily utilize some feature of the target language the other didn’t have. A good example is with alternations. If several left-to-right atomic branches each begin with a different literal character, the engine can read that character and jump straight to the matching branch rather than testing each branch in order. With C#, we emitted a <code>switch</code>, which the C# compiler could then lower to IL using various strategies. For IL, in .NET 10 and earlier, without the C# compiler to provide those optimizations, we just skipped the optimization. Now in .NET 11, dotnet/runtime#122959 emits a similar implementation to what the C# compiler would have, bringing this optimization to <code>RegexOptions.Compiled</code>.</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing System.Text.RegularExpressions;\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private readonly string _input = string.Concat(Enumerable.Repeat(&quot;p15&quot;, 10_000));\n    private readonly Regex _regex = new(@&quot;(?&gt;a0|b1|c2|d3|e4|f5|g6|h7|i8|j9|k10|l11|m12|n13|o14|p15)&quot;, RegexOptions.Compiled);\n    [Benchmark]\n    public int DispatchToFinalBranch() =&gt; _regex.Count(_input);\n}</code></pre>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Mean</th><th>Ratio</th></tr></thead><tbody><tr><td>DispatchToFinalBranch</td><td>.NET 10.0</td><td>233.9 μs</td><td>1.00</td></tr><tr><td>DispatchToFinalBranch</td><td>.NET 11.0</td><td>159.5 μs</td><td>0.68</td></tr></tbody></table></div>\n<p>Another of the few differences between compiled and source-generated regexes had to do with backreferences. A case-sensitive backreference, such as the <code>\\1</code> in <code>([a-z]+)-\\1</code>, asks whether the next input equals text that was previously captured in the match. Source-generated regexes were using the optimized <code>SequenceEqual</code> to do that comparison, whereas <code>RegexOptions.Compiled</code> wasn’t. With dotnet/runtime#123914 in .NET 11, now it does.</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Text.RegularExpressions;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private readonly string _input = new string(&#39;a&#39;, 256) + &quot;-&quot; + new string(&#39;a&#39;, 256);\n    private readonly Regex _regex = new(@&quot;^([a-z]{256})-\\1$&quot;, RegexOptions.Compiled);\n    [Benchmark]\n    public bool Backreference() =&gt; _regex.IsMatch(_input);\n}</code></pre>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Mean</th><th>Ratio</th></tr></thead><tbody><tr><td>Backreference</td><td>.NET 10.0</td><td>168.4 ns</td><td>1.00</td></tr><tr><td>Backreference</td><td>.NET 11.0</td><td>49.59 ns</td><td>0.29</td></tr></tbody></table></div>\n<p>Searching isn’t limited to <code>Regex</code>, of course. Many other methods in .NET help finding things and comparing things, some of which get notable bumps in .NET 11.</p>\n<p>The <code>Ascii</code> class provides optimized helpers for validating and manipulating ASCII text. Members like <code>Equals</code> are already vectorized in .NET 10, but in .NET 11, dotnet/runtime#123115 improves that implementation by ensuring that inputs of length 8 through 15 can be vectorized, as well.</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Running;\nusing BenchmarkDotNet.Attributes;\nusing System.Text;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    [Params(8, 15)]\n    public int Length { get; set; }\n    private byte[] _bytes = [];\n    private char[] _charsMatching = [];\n    [GlobalSetup]\n    public void Setup()\n    {\n        _bytes = new byte[Length];\n        _charsMatching = new char[Length];\n        for (int i = 0; i &lt; Length; i++)\n        {\n            byte b = (byte)(&#39;a&#39; + (i % 26));\n            _bytes[i] = b;\n            _charsMatching[i] = (char)b;\n        }\n    }\n    [Benchmark]\n    public bool Equals_Matching() =&gt; Ascii.Equals(_bytes, _charsMatching);\n}</code></pre>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Length</th><th>Mean</th><th>Ratio</th></tr></thead><tbody><tr><td>Equals_Matching</td><td>.NET 10.0</td><td>8</td><td>3.834 ns</td><td>1.00</td></tr><tr><td>Equals_Matching</td><td>.NET 11.0</td><td>8</td><td>1.966 ns</td><td>0.51</td></tr><tr><td>Equals_Matching</td><td>.NET 10.0</td><td>15</td><td>6.177 ns</td><td>1.00</td></tr><tr><td>Equals_Matching</td><td>.NET 11.0</td><td>15</td><td>2.398 ns</td><td>0.39</td></tr></tbody></table></div>\n<p>dotnet/runtime#130644\nalso improves equality performance, in this case with <code>SequenceEqual</code> over a\nspan of <code>Guid</code> or <code>Int128</code>. Previously, <code>SequenceEqual</code> treated these as\narbitrary structures and compared them one element at a time. The PR teaches\nthe runtime that their fixed bitwise representations are suitable for\ncomparison as raw bytes. That enables the same optimized memory-comparison\npath used for primitive types, including JIT unrolling and vectorization:</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private readonly Guid[] _guids1 = new Guid[2];\n    private readonly Guid[] _guids2 = new Guid[2];\n    private readonly Int128[] _int128s1 = new Int128[2];\n    private readonly Int128[] _int128s2 = new Int128[2];\n    [Benchmark]\n    public bool GuidEqual() =&gt; _guids1.AsSpan().SequenceEqual(_guids2);\n    [Benchmark]\n    public bool Int128Equal() =&gt; _int128s1.AsSpan().SequenceEqual(_int128s2);\n}</code></pre>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Mean</th><th>Ratio</th></tr></thead><tbody><tr><td>GuidEqual</td><td>.NET 10.0</td><td>2.960 ns</td><td>1.00</td></tr><tr><td>GuidEqual</td><td>.NET 11.0</td><td>2.077 ns</td><td>0.70</td></tr><tr><td>Int128Equal</td><td>.NET 10.0</td><td>3.547 ns</td><td>1.00</td></tr><tr><td>Int128Equal</td><td>.NET 11.0</td><td>2.077 ns</td><td>0.59</td></tr></tbody></table></div>\n<p>Another improvement in .NET 11 is to <code>string.Split</code>.\nBefore <code>string.Split</code> can produce the resulting strings, it first needs to\nfind the characters that separate them and record their positions. In .NET 10, that search is\nalready vectorized: rather than examine one UTF-16 character at a time, it\nloads a vector’s worth, compares all of its lanes against the separator in\nparallel, and turns the comparison result into a mask identifying any matches.\nIt then advances to the next vector, or uses the mask to record the matching\npositions. In .NET 11, on x86/x64, dotnet/runtime#125379\nfrom @hamarb123 makes the no-match path cheaper\nfor ASCII separators. It loads two vectors of UTF-16 characters, packs their\n16-bit elements into one vector of bytes, and checks that combined vector for\nthe separator. If there isn’t a match, it has skipped twice as much input with\none packed comparison; only a possible match requires the full 16-bit\ncomparisons needed to determine its exact position. (This same packing technique\nis already employed elsewhere, such as in various <code>SearchValues&lt;T&gt;</code> implementations.)</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private readonly string _input = new(&#39;a&#39;, 16_384);\n    [Benchmark]\n    public int SplitNoSeparators() =&gt; _input.Split(&#39;,&#39;).Length;\n}</code></pre>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Mean</th><th>Ratio</th></tr></thead><tbody><tr><td>SplitNoSeparators</td><td>.NET 10.0</td><td>786.0 ns</td><td>1.00</td></tr><tr><td>SplitNoSeparators</td><td>.NET 11.0</td><td>404.7 ns</td><td>0.51</td></tr></tbody></table></div>\n<p>A related Arm64 text-search improvement comes from\ndotnet/runtime#126678.\nA vector comparison produces a vector whose elements are all zero for\nnon-matches and all one bits for matches. Finding the first or last match then\nrequires condensing those bits into a scalar value and counting its leading or\ntrailing zeros. On x86, the runtime can use a movemask instruction for that\ncondensing step. Arm64 has no direct equivalent, and the old implementation\nneeded a sequence of shifts, widening operations, and a horizontal add to achieve it.\nThe .NET 11 implementation now uses <code>shrn</code>, Arm64’s shift-right-and-narrow\ninstruction, to pack the relevant bits directly. <code>SearchValues&lt;char&gt;</code> uses these helpers, so the following benchmark reaches\nthe affected code while searching for a match at the end of the input.</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Buffers;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[DisassemblyDiagnoser(maxDepth: 3)]\n[HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private const int Length = 8_192;\n    private static readonly SearchValues&lt;char&gt; s_vowels = SearchValues.Create(&quot;aeiouAEIOU&quot;);\n    private static readonly string s_input = new string(&#39;x&#39;, Length - 1) + &#39;e&#39;;\n    [Benchmark]\n    public int IndexOfAny() =&gt; s_input.AsSpan().IndexOfAny(s_vowels);\n}</code></pre>\n<p>The .NET 10 match-index path requires this sequence:</p>\n<pre><code>; Arm64\n; .NET 10\ncmeq    v16.16b, v16.16b, #0\nmovi    v17.16b, #0x80\nand     v16.16b, v16.16b, v17.16b\nldr     q17, [MASK]\nushl    v16.16b, v16.16b, v17.16b\nuxtl2   v17.8h, v16.16b\nshl     v17.8h, v17.8h, #8\nuaddw   v16.8h, v17.8h, v16.8b\naddv    h16, v16.8h\numov    w2, v16.h[0]\nmvn     w2, w2\nrbit    w2, w2\nclz     w2, w2</code></pre>\n<p>In .NET 11, the equivalent work is simpler:</p>\n<pre><code>; Arm64\n; .NET 11\ncmeq    v16.16b, v16.16b, #0\nmvn     v16.16b, v16.16b\nshrn    v16.8b, v16.8h, #4\nfmov    x2, d16\nrbit    x2, x2\nclz     x2, x2\nlsr     w2, w2, #2</code></pre>\n<p><code>MemoryExtensions</code> already provides span-based searches for one or more values\nwith <code>IndexOfAny</code>, and for contiguous ranges with <code>IndexOfAnyInRange</code>, along\nwith <code>Except</code>, <code>Contains</code>, and last-index variants of these operations. For\nexample, <code>span.IndexOfAnyInRange(&#39;0&#39;, &#39;9&#39;)</code> finds the next ASCII digit.\nWhitespace is also common to search for, but the characters recognized by\n<code>char.IsWhiteSpace</code> are spread across multiple parts of Unicode rather than\nforming one contiguous range. To avoid requiring every caller to construct\nthe same <code>SearchValues&lt;char&gt;</code>,\ndotnet/runtime#111439 from\n@AlexRadch adds\n<code>ContainsAnyWhiteSpace</code>, <code>IndexOfAnyWhiteSpace</code>,\n<code>IndexOfAnyExceptWhiteSpace</code>, <code>LastIndexOfAnyWhiteSpace</code>, and\n<code>LastIndexOfAnyExceptWhiteSpace</code> for <code>ReadOnlySpan&lt;char&gt;</code>. Their shared\n<code>SearchValues&lt;char&gt;</code>-based implementation vectorizes these searches for\nparsers, validators, trimming code, and other text-processing code.</p>\n<p>This is, however, a good example of how vectorization isn’t always a win. Take trimming. To trim leading whitespace, code needs to find the first character that isn’t whitespace. That character could be deep into the string, but in the most common case, there’s little or nothing to trim. A scalar loop can then return after inspecting just one or two characters, whereas the vectorized helper has fixed setup cost. It’s still worth vectorizing, because that overhead is small and the benefits when there is a lot to scan can be significant. Something to keep in mind.</p>\n<pre><code>// dotnet run -c Release -f net11.0 --filter &quot;*&quot;\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private readonly string _input = new(&#39; &#39;, 256);\n    [Benchmark(Baseline = true)]\n    public int Scalar()\n    {\n        ReadOnlySpan&lt;char&gt; input = _input;\n        for (int i = 0; i &lt; input.Length; i++)\n        {\n            if (!char.IsWhiteSpace(input[i]))\n                return i;\n        }\n        return -1;\n    }\n    [Benchmark]\n    public int Vectorized() =&gt; _input.AsSpan().IndexOfAnyExceptWhiteSpace();\n}</code></pre>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Mean</th><th>Ratio</th></tr></thead><tbody><tr><td>Scalar</td><td>127.87 ns</td><td>1.00</td></tr><tr><td>Vectorized</td><td>13.14 ns</td><td>0.10</td></tr></tbody></table></div>\n<p>This method is a particularly good fit when needing to validate that input does not contain any whitespace; that requires searching the entirety of input, which is where the vectorization in these methods shines. As an example of this, dotnet/runtime#127123 uses it to accelerate the parsing of the <code>&quot;X&quot;</code> GUID format:</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private static readonly string s_noWhitespace =\n        Guid.Parse(&quot;a8098c1a-f86e-11da-bd1a-00112444be1e&quot;).ToString(&quot;X&quot;);\n    [Benchmark]\n    public Guid ParseExactX() =&gt; Guid.ParseExact(s_noWhitespace, &quot;X&quot;);\n}</code></pre>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Mean</th><th>Ratio</th></tr></thead><tbody><tr><td>ParseExactX</td><td>.NET 10.0</td><td>120.6 ns</td><td>1.00</td></tr><tr><td>ParseExactX</td><td>.NET 11.0</td><td>87.10 ns</td><td>0.72</td></tr></tbody></table></div>\n<p>Closely related to searching is sorting. Years ago, sorting methods for <code>Span&lt;T&gt;</code> were added to <code>MemoryExtensions</code>. Interestingly, the method wasn’t added as <code>Sort&lt;T&gt;</code> but rather as <code>Sort&lt;T, TComparer&gt;</code> where <code>TComparer : IComparer&lt;T&gt;</code>. That signature enables a caller to provide a struct comparer without allocating a delegate or class-based comparer. Because the comparer is a constrained value type, the JIT should also be able to inline the comparison into the hot sorting loop. In practice, the implementation boxed the struct into an <code>IComparer&lt;T&gt;</code>, both allocating and turning every comparison back into an interface call. This was known at the time, but avoiding the box used generic implementation techniques that then carried too much runtime and code-size cost. Those supporting costs have since been addressed, so dotnet/runtime#116109 from @2A5F now carries a value-type comparer through <code>Span&lt;T&gt;.Sort</code> without boxing it. The JIT can specialize the sorting routine for that comparer and inline the comparison.</p>\n<p>The generic specialization does increase generated code and very large comparer structs can be more expensive to copy; this optimization is aimed at the small stateless or lightly stateful structs for which the API was designed.</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[MemoryDiagnoser(false), HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private readonly int[] _source = Enumerable.Range(0, 512).Select(i =&gt; (i * 257) % 512).ToArray();\n    private int[] _values = [];\n    [IterationSetup]\n    public void Setup() =&gt; _values = (int[])_source.Clone();\n    [Benchmark]\n    public void Sort() =&gt; _values.AsSpan().Sort(new DescendingComparer());\n    private readonly struct DescendingComparer : IComparer&lt;int&gt;\n    {\n        public int Compare(int x, int y) =&gt; y.CompareTo(x);\n    }\n}</code></pre>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Mean</th><th>Ratio</th><th>Allocated</th></tr></thead><tbody><tr><td>Sort</td><td>.NET 10.0</td><td>11.62 μs</td><td>1.00</td><td>88 B</td></tr><tr><td>Sort</td><td>.NET 11.0</td><td>3.533 μs</td><td>0.30</td><td>–</td></tr></tbody></table></div>\n<h2 id=\"collections-and-linq\">Collections and LINQ</h2>\n<p>Much of the collection and LINQ work in .NET 11 comes from taking better\nadvantage of information that’s already available. A collection often knows\nmuch more than an <code>IEnumerable&lt;T&gt;</code> can express: its count, its contiguous\nstorage, its comparer, or the layout of its hash table. Similarly, a LINQ\niterator can know how many elements it represents or how its operations were\ncomposed. Preserving that information can avoid enumeration, temporary\nstorage, repeated hashing, and other work a general-purpose implementation\nwould otherwise need to perform.</p>\n<p>dotnet/runtime#119896 from @prozolic changes <code>ImmutableArray.Create</code> to use <code>Array.Copy</code> rather than a hand-written element loop. A general element-by-element copy repeatedly performs indexing and assignment, while the runtime can specialize <code>Array.Copy</code> for the element type and size. For blittable data, it can use optimized bulk memory copies, and for reference types, which need GC write barriers, it performs the required write barriers in the runtime’s tuned copy helpers. The change therefore both simplifies the managed code and gives <code>ImmutableArray</code> access to those optimized implementations.</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Running;\nusing System.Collections.Immutable;\nusing BenchmarkDotNet.Attributes;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;RatioSD&quot;, &quot;Median&quot;)]\npublic class Benchmarks\n{\n    private readonly int[] _source = Enumerable.Range(0, 1_000).ToArray();\n    [Benchmark]\n    public ImmutableArray&lt;int&gt; CreateSlice() =&gt; ImmutableArray.Create(_source, 0, _source.Length);\n}</code></pre>\n<p>This in particular makes larger copies much faster.</p>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Mean</th><th>Ratio</th></tr></thead><tbody><tr><td>CreateSlice</td><td>.NET 10.0</td><td>552.5 ns</td><td>1.00</td></tr><tr><td>CreateSlice</td><td>.NET 11.0</td><td>277.7 ns</td><td>0.50</td></tr></tbody></table></div>\n<p>dotnet/runtime#118932 from @prozolic similarly keeps <code>ImmutableArrayExtensions.SequenceEqual</code> on optimized paths when the other sequence is an array, list, or another <code>ICollection&lt;T&gt;</code>.</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Running;\nusing System.Collections.Immutable;\nusing BenchmarkDotNet.Attributes;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;RatioSD&quot;, &quot;Median&quot;)]\npublic class Benchmarks\n{\n    private ImmutableArray&lt;int&gt; _immutable;\n    private List&lt;int&gt; _list = [];\n    [GlobalSetup]\n    public void Setup()\n    {\n        int[] values = Enumerable.Range(0, 1_000).ToArray();\n        _immutable = ImmutableArray.Create(values);\n        _list = [.. values];\n    }\n    [Benchmark]\n    public bool SequenceEqual() =&gt; _immutable.SequenceEqual(_list);\n}</code></pre>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Mean</th><th>Ratio</th></tr></thead><tbody><tr><td>SequenceEqual</td><td>.NET 10.0</td><td>925.8 ns</td><td>1.00</td></tr><tr><td>SequenceEqual</td><td>.NET 11.0</td><td>122.2 ns</td><td>0.13</td></tr></tbody></table></div>\n<p><code>Array.FindAll</code> has the opposite job: it produces a new collection. For a small result, its temporary storage used to cost more than the result itself.\ndotnet/runtime#120336 from @Henr1k80 has <code>Array.FindAll</code> collect its first four matches in an inline stack buffer rather than an intermediate <code>List&lt;T&gt;</code>:</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Running;\nusing BenchmarkDotNet.Attributes;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[MemoryDiagnoser(false), HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;RatioSD&quot;, &quot;Median&quot;)]\npublic class Benchmarks\n{\n    private int[] _data = [];\n    [Params(4, 5)]\n    public int Size { get; set; }\n    [GlobalSetup]\n    public void Setup() =&gt; _data = Enumerable.Range(0, Size).ToArray();\n    [Benchmark]\n    public int[] FindAllMatch() =&gt; Array.FindAll(_data, static _ =&gt; true);\n}</code></pre>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Size</th><th>Mean</th><th>Ratio</th><th>Allocated</th><th>Alloc Ratio</th></tr></thead><tbody><tr><td>FindAllMatch</td><td>.NET 10.0</td><td>4</td><td>27.61 ns</td><td>1.00</td><td>112 B</td><td>1.00</td></tr><tr><td>FindAllMatch</td><td>.NET 11.0</td><td>4</td><td>9.212 ns</td><td>0.33</td><td>40 B</td><td>0.36</td></tr><tr><td>FindAllMatch</td><td>.NET 10.0</td><td>5</td><td>36.93 ns</td><td>1.00</td><td>176 B</td><td>1.00</td></tr><tr><td>FindAllMatch</td><td>.NET 11.0</td><td>5</td><td>11.102 ns</td><td>0.30</td><td>48 B</td><td>0.27</td></tr></tbody></table></div>\n<p><code>Dictionary&lt;TKey, TValue&gt;.Remove</code> had also missed an optimization already used by lookup and insertion. dotnet/runtime#125884 gives value-type keys a streamlined loop for the common default-comparer case. Because that path doesn’t need a virtual comparer call, the JIT can keep more of the operation’s state in registers; reference-type keys and custom comparers continue to use the general path.</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private readonly Guid[] _keys = Enumerable.Range(0, 512).Select(i =&gt; new Guid(i, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0)).ToArray();\n    private Dictionary&lt;Guid, int&gt; _dictionary = [];\n    [IterationSetup]\n    public void Setup() =&gt; _dictionary = _keys.ToDictionary(key =&gt; key, key =&gt; key.GetHashCode());\n    [Benchmark(OperationsPerInvoke = 512)]\n    public void Remove()\n    {\n        foreach (Guid key in _keys)\n            _dictionary.Remove(key);\n    }\n}</code></pre>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Mean</th><th>Ratio</th></tr></thead><tbody><tr><td>Remove</td><td>.NET 10.0</td><td>5.285 ns</td><td>1.00</td></tr><tr><td>Remove</td><td>.NET 11.0</td><td>4.321 ns</td><td>0.82</td></tr></tbody></table></div>\n<p>dotnet/runtime#125893 changes <code>HashSet&lt;T&gt;</code>‘s internal chain walks to test the entry index against the array length with an unsigned comparison. That proves the subsequent array access is in range, allowing the JIT to remove its bounds check.</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private readonly HashSet&lt;int&gt; _set = Enumerable.Range(0, 4096).ToHashSet();\n    private readonly int[] _probes = Enumerable.Range(0, 4096).ToArray();\n    [Benchmark]\n    public int ContainsHits()\n    {\n        int count = 0;\n        foreach (int value in _probes)\n            count += _set.Contains(value) ? 1 : 0;\n        return count;\n    }\n}</code></pre>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Mean</th><th>Ratio</th></tr></thead><tbody><tr><td>ContainsHits</td><td>.NET 10.0</td><td>7.870 μs</td><td>1.00</td></tr><tr><td>ContainsHits</td><td>.NET 11.0</td><td>7.388 μs</td><td>0.94</td></tr></tbody></table></div>\n<p>dotnet/runtime#128988 from @prozolic removes a second hash-table lookup when removing a matching key-value pair from <code>OrderedDictionary&lt;TKey, TValue&gt;</code> through <code>ICollection&lt;KeyValuePair&lt;TKey, TValue&gt;&gt;</code>. That interface operation must first find the key and verify that its stored value equals the supplied value. Once both checks have succeeded, the implementation already has the entry index needed for removal. Looking up the key again unnecessarily repeats its hash computation and collision-chain walk, so the updated path removes the known entry directly.</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Running;\nusing System.Collections.Generic;\nusing BenchmarkDotNet.Attributes;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;RatioSD&quot;, &quot;Median&quot;)]\npublic class Benchmarks\n{\n    private const int N = 10_000;\n    private OrderedDictionary&lt;string, int&gt; _dict = [];\n    private KeyValuePair&lt;string, int&gt;[] _pairs = Enumerable.Range(0, N)\n        .Select(i =&gt; new KeyValuePair&lt;string, int&gt;($&quot;key{i}&quot;, i))\n        .ToArray();\n    [IterationSetup]\n    public void IterationSetup() =&gt; _dict = new OrderedDictionary&lt;string, int&gt;(_pairs);\n    [Benchmark]\n    public int Remove_ExplicitInterface()\n    {\n        ICollection&lt;KeyValuePair&lt;string, int&gt;&gt; col = _dict;\n        int removed = 0;\n        foreach (var pair in _pairs)\n            if (col.Remove(pair))\n                removed++;\n        return removed;\n    }\n}</code></pre>\n<p>For 10,000 entries:</p>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Mean</th><th>Ratio</th></tr></thead><tbody><tr><td>Remove_ExplicitInterface</td><td>.NET 10.0</td><td>220.7 ms</td><td>1.00</td></tr><tr><td>Remove_ExplicitInterface</td><td>.NET 11.0</td><td>179.0 ms</td><td>0.81</td></tr></tbody></table></div>\n<p>dotnet/runtime#122952 goes further when two hash tables have compatible layouts. Normally, <code>UnionWith</code> enumerates the source and inserts every element independently, recomputing hashes, checking for duplicates, and potentially resizing the destination along the way. If the destination is empty and both sets use compatible comparers, every source entry is already unique under exactly the equality rules the destination needs. <code>UnionWith</code> can therefore use the existing <code>HashSet&lt;T&gt;</code> copy-constructor fast path to clone the populated storage rather than rebuilding the same table entry by entry.</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Running;\nusing System.Collections.Generic;\nusing System.Linq;\nusing BenchmarkDotNet.Attributes;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[MemoryDiagnoser(false), HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;RatioSD&quot;, &quot;Median&quot;)]\npublic class Benchmarks\n{\n    private readonly HashSet&lt;int&gt; _source = new(Enumerable.Range(0, 4_096));\n    [Benchmark]\n    public HashSet&lt;int&gt; FreshDestinationUnionWith()\n    {\n        HashSet&lt;int&gt; destination = [];\n        destination.UnionWith(_source);\n        return destination;\n    }\n}</code></pre>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Mean</th><th>Ratio</th><th>Allocated</th><th>Alloc Ratio</th></tr></thead><tbody><tr><td>FreshDestinationUnionWith</td><td>.NET 10.0</td><td>46.207 μs</td><td>1.00</td><td>252.27 KB</td><td>1.00</td></tr><tr><td>FreshDestinationUnionWith</td><td>.NET 11.0</td><td>2.433 μs</td><td>0.05</td><td>76.07 KB</td><td>0.30</td></tr></tbody></table></div>\n<p>dotnet/runtime#128300\nfrom @AndrewP-GH also helps with collection construction. Building a <code>FrozenDictionary&lt;TKey, TValue&gt;</code> first requires collecting the input elements into a regular <code>Dictionary&lt;TKey, TValue&gt;</code> if they’re not already in one. That temporary dictionary resolves duplicate keys before the final frozen representation is chosen, but in .NET 10 it was growing incrementally even when the source’s count was readily available. This PR uses that count as the\ndictionary’s initial capacity, avoiding repeated allocation, copying, and\nrehashing as it’s populated.</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nusing System.Collections.Concurrent;\nusing System.Collections.Frozen;\nusing System.Collections.Generic;\nusing System.Collections.Immutable;\nusing System.Linq;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[MemoryDiagnoser(false), HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private readonly KeyValuePair&lt;int, int&gt;[] _array =\n        Enumerable.Range(0, 4096).Select(i =&gt; new KeyValuePair&lt;int, int&gt;(i, i)).ToArray();\n    [Benchmark]\n    public FrozenDictionary&lt;int, int&gt; FromArray() =&gt; _array.ToFrozenDictionary();\n}</code></pre>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Allocated</th><th>Alloc Ratio</th></tr></thead><tbody><tr><td>FromArray</td><td>.NET 10.0</td><td>347.17 KB</td><td>1.00</td></tr><tr><td>FromArray</td><td>.NET 11.0</td><td>127.16 KB</td><td>0.37</td></tr></tbody></table></div>\n<p><code>SetEquals</code> asks whether two sets contain the same values, regardless of insertion order. The general implementation needs a temporary mutable set so it can account for duplicates and arbitrary enumeration order. When the other input is already a hash set with a compatible comparer, though, that reconstruction is unnecessary. dotnet/runtime#126309 from @aw0lid adds to <code>ImmutableHashSet&lt;T&gt;.SetEquals</code> direct zero-allocation paths for compatible <code>ImmutableHashSet&lt;T&gt;</code> and <code>HashSet&lt;T&gt;</code> inputs; with an identical comparer, the sets can be considered equal if they have the same count and if every element from one is found in the other.</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing System.Collections.Immutable;\nusing BenchmarkDotNet.Attributes;\nusing BenchmarkDotNet.Running;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[MemoryDiagnoser(false)]\n[HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;Median&quot;, &quot;RatioSD&quot;)]\npublic class Benchmarks\n{\n    private ImmutableHashSet&lt;int&gt; _set = ImmutableHashSet&lt;int&gt;.Empty;\n    private ImmutableHashSet&lt;int&gt; _immutable = ImmutableHashSet&lt;int&gt;.Empty;\n    private HashSet&lt;int&gt; _mutable = [];\n    [GlobalSetup]\n    public void Setup()\n    {\n        int[] items = Enumerable.Range(0, 10_000).ToArray();\n        _set = ImmutableHashSet.CreateRange(items);\n        _immutable = ImmutableHashSet.CreateRange(items);\n        _mutable = new(items);\n    }\n    [Benchmark]\n    public bool EqualImmutableHashSet() =&gt; _set.SetEquals(_immutable);\n    [Benchmark]\n    public bool EqualHashSet() =&gt; _set.SetEquals(_mutable);\n}</code></pre>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Mean</th><th>Ratio</th><th>Allocated</th><th>Alloc Ratio</th></tr></thead><tbody><tr><td>EqualImmutableHashSet</td><td>.NET 10.0</td><td>775.6 μs</td><td>1.00</td><td>158.16 KB</td><td>1.00</td></tr><tr><td>EqualImmutableHashSet</td><td>.NET 11.0</td><td>559.7 μs</td><td>0.72</td><td>–</td><td>0</td></tr><tr><td>EqualHashSet</td><td>.NET 10.0</td><td>478.6 μs</td><td>1.00</td><td>157.99 KB</td><td>1.00</td></tr><tr><td>EqualHashSet</td><td>.NET 11.0</td><td>230.9 μs</td><td>0.48</td><td>–</td><td>0</td></tr></tbody></table></div>\n<p>Sorted sets have a related case. <code>SetEquals</code> can be passed any <code>IEnumerable&lt;T&gt;</code>. That sequence might be unordered and might contain duplicate values, so <code>ImmutableSortedSet&lt;T&gt;</code> previously copied it into a temporary <code>SortedSet&lt;T&gt;</code> before performing the comparison. However, when the input is another sorted set using the same ordering comparer, both sets contain unique values and enumerate those values\nin the same order. Equality can then be determined by first comparing their\ncounts and, if those match, advancing both enumerators together. The first\nunequal pair proves the sets are different, and reaching the end without finding\na difference proves they’re equal. dotnet/runtime#126549 from\n@aw0lid recognizes this case for\n<code>ImmutableSortedSet&lt;T&gt;</code>, avoiding the temporary <code>SortedSet&lt;T&gt;</code> and comparing the two sorted sequences directly in one linear pass:</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Running;\nusing System.Collections.Generic;\nusing System.Collections.Immutable;\nusing BenchmarkDotNet.Attributes;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[MemoryDiagnoser(false), HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;RatioSD&quot;, &quot;Median&quot;)]\npublic class Benchmarks\n{\n    private const int N = 10_000;\n    private ImmutableSortedSet&lt;int&gt; _set = ImmutableSortedSet&lt;int&gt;.Empty;\n    private ImmutableSortedSet&lt;int&gt; _equalSet = ImmutableSortedSet&lt;int&gt;.Empty;\n    [GlobalSetup]\n    public void Setup()\n    {\n        var items = new int[N];\n        for (int i = 0; i &lt; N; i++) items[i] = i;\n        _set = ImmutableSortedSet.CreateRange(items);\n        _equalSet = ImmutableSortedSet.CreateRange(items);\n    }\n    [Benchmark]\n    public bool SetEquals_EqualImmutableSortedSet() =&gt; _set.SetEquals(_equalSet);\n}</code></pre>\n<p>For 10,000 elements:</p>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Mean</th><th>Ratio</th><th>Allocated</th><th>Alloc Ratio</th></tr></thead><tbody><tr><td>SetEquals_EqualImmutableSortedSet</td><td>.NET 10.0</td><td>767.7 μs</td><td>1.00</td><td>430.02 KB</td><td>1.00</td></tr><tr><td>SetEquals_EqualImmutableSortedSet</td><td>.NET 11.0</td><td>117.6 μs</td><td>0.15</td><td>–</td><td>0</td></tr></tbody></table></div>\n<p><code>SortedSet&lt;T&gt;</code> already enjoyed an optimization for that case in .NET 10, but it’s not left out of .NET 11 improvements. <code>SortedSet&lt;T&gt;.GetViewBetween</code> returns a <code>SortedSet&lt;T&gt;</code> view, effectively a slice of another <code>SortedSet&lt;T&gt;</code>, a live window onto a range of another set: changes through the view affect the original set. Clearing a view therefore can’t replace the view with an empty collection; it must find and remove every original node in that range. dotnet/runtime#126410 from @prozolic reduces the temporary storage used for that operation. The implementation pre-sizes the list of elements to remove and walks it by index rather than repeatedly removing from and shrinking the temporary list.</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Running;\nusing BenchmarkDotNet.Attributes;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[MemoryDiagnoser(false), HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;RatioSD&quot;, &quot;Median&quot;)]\npublic class Benchmarks\n{\n    private const int N = 10_000;\n    private SortedSet&lt;int&gt; _fullSet = [];\n    [IterationSetup]\n    public void Setup() =&gt; _fullSet = new SortedSet&lt;int&gt;(Enumerable.Range(0, N));\n    [Benchmark]\n    public int GetViewBetweenThenClear()\n    {\n        SortedSet&lt;int&gt; view = _fullSet.GetViewBetween(0, N - 1);\n        view.Clear();\n        return _fullSet.Count;\n    }\n}</code></pre>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Allocated</th><th>Alloc Ratio</th></tr></thead><tbody><tr><td>GetViewBetweenThenClear</td><td>.NET 10.0</td><td>193.15 KB</td><td>1.00</td></tr><tr><td>GetViewBetweenThenClear</td><td>.NET 11.0</td><td>103.93 KB</td><td>0.54</td></tr></tbody></table></div>\n<p>Collections are frequently consumed through LINQ. Although its operators work\nin terms of the general <code>IEnumerable&lt;T&gt;</code> abstraction, LINQ’s internal\niterators can preserve useful facts about their sources and the operations\nalready applied. Those facts can sometimes answer a query without enumerating\nthe source at all.</p>\n<p>For example, consider <code>source.Append(x).Skip(10).LastOrDefault()</code>. LINQ queries are lazy, so the actual search begins only when <code>LastOrDefault</code> asks the <code>Skip</code> iterator for its last element. If <code>source.Append(x)</code> contains ten or fewer elements,\n<code>Skip(10)</code> necessarily removes all of them, leaving an empty sequence from\nwhich <code>LastOrDefault</code> must return the default value. <code>Append</code>, <code>Prepend</code>, and\n<code>Concat</code> iterators can cheaply report their total count when their underlying\nsources can do so. dotnet/runtime#123306 from @prozolic teaches the last-element path for <code>Skip</code> to compare that count with the number being skipped and immediately\nreport that there is no element, rather than searching a sequence it already\nknows is empty.</p>\n<pre><code>// dotnet run -c Release -f net10.0 --filter &quot;*&quot; --runtimes net10.0 net11.0\nusing BenchmarkDotNet.Running;\nusing System.Linq;\nusing BenchmarkDotNet.Attributes;\nBenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);\n[MemoryDiagnoser(false), HideColumns(&quot;Job&quot;, &quot;Error&quot;, &quot;StdDev&quot;, &quot;RatioSD&quot;, &quot;Median&quot;)]\npublic class Benchmarks\n{\n    private readonly int[] _source = [1, 2, 3, 4, 5];\n    [Benchmark]\n    public int AppendSkipLastOrDefault() =&gt; _source.Append(6).Skip(10).LastOrDefault();\n}</code></pre>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Runtime</th><th>Mean</th><th>Ratio</th><th>Allocated</th><th>Alloc Ratio</th></tr></thead><tbody><tr><td>AppendSkipLastOrDefault</td><td>.NET 10.0</td><td>44.89 ns</td><td>1.00</td><td>144 B</td><td>1.00</td></tr><tr><td>AppendSkipLastOrDefault</td><td>.NET 11.0</td><td>16.32 ns</td><td>0.36</td><td>112 B</td><td>0.78</td></tr></tbody></table></div>\n<p>.NET 11 also improve’s LINQ’s <code>Enumerable.Sum</code>. <code>Sum</code> already uses SIMD. The\nmain loop processes four vectors at a time, alternating between two\naccumulators so that the additions don’t require extra moves. However, <code>Sum</code>\nalso promises to throw if the result overflows. Alongside each vector addition,\nthe implementation uses the signs of the two inputs and the result to update\nanother vector that tracks whether any lane overflowed. After every\ngroup of four vectors, the loop tests that tracking vector and branches to the\nthrowing path if needed. Overflow is rare, though, so on the common path that test and branch almost always just\nconfirm that nothing happened. dotnet/runtime#127429\nremoves that repeated work in .NET 11. It accumulates the overflow information\nacross all of the vector processing and tests it once after the vector loops\nhave completed. The checked-overflow behavior remains the same, but the normal\npath no longer needs to stop and check after every four vectors. The PR also\nsimplifies how the method walks the input, replacing unsafe reference and index\narithmetic with span-based vector loads, progressively slicing off the elements\nalready processed, and using a <code>foreach</code> for the final scalar elements.</p>","headings":[{"level":2,"text":"Benchmarking Setup","id":"benchmarking-setup"},{"level":2,"text":"JIT","id":"jit"},{"level":3,"text":"Deabstraction","id":"deabstraction"},{"level":3,"text":"Runtime Async","id":"runtime-async"},{"level":3,"text":"Bounds Checks","id":"bounds-checks"},{"level":3,"text":"Assertion Propagation","id":"assertion-propagation"},{"level":3,"text":"Simplification","id":"simplification"},{"level":3,"text":"Vectorization","id":"vectorization"},{"level":3,"text":"Intrinsics","id":"intrinsics"},{"level":3,"text":"Register Allocation","id":"register-allocation"},{"level":3,"text":"Write Barriers and Garbage Collection","id":"write-barriers-and-garbage-collection"},{"level":3,"text":"Runtime Knowledge and Frozen Data","id":"runtime-knowledge-and-frozen-data"},{"level":3,"text":"JIT Throughput and Cleanup","id":"jit-throughput-and-cleanup"},{"level":2,"text":"Startup and Deployment","id":"startup-and-deployment"},{"level":2,"text":"Threading","id":"threading"},{"level":2,"text":"Numerics","id":"numerics"},{"level":2,"text":"Globalization","id":"globalization"},{"level":2,"text":"Strings and Spans","id":"strings-and-spans"},{"level":3,"text":"Searching and Comparing","id":"searching-and-comparing"},{"level":2,"text":"Collections and LINQ","id":"collections-and-linq"}]}}