{"article":{"slug":"arch-specific-simd-in-go","title":"Arch-specific SIMD in Go","subtitle":null,"summary":"The Go team introduces experimental architecture-specific SIMD APIs in Go 1.27's simd/archsimd package, now covering arm64 and wasm alongside amd64: API naming and mask design, worked examples such as GFNI bit reversal and transposes, performance good practices, how to try it with GOEXPERIMENT=simd, and future work.","content_type":"blog_post","language":"en","canonical_url":"https://go.dev/blog/archsimd","author":{"name":"Junyang Shao and David Chase","url":null,"person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"The Go Blog","url":"https://go.dev/blog/","listing_slug":null,"listing":null},"topics":[{"name":"Programming","slug":"programming","url":"https://listedarticles.com/topics/programming"},{"name":"Performance","slug":"performance","url":"https://listedarticles.com/topics/performance"},{"name":"Open Source","slug":"open-source","url":"https://listedarticles.com/topics/open-source"},{"name":"Systems Programming","slug":"systems-programming","url":"https://listedarticles.com/topics/systems-programming"}],"about_listings":[],"cover_image_url":null,"license":"CC-BY-4.0","word_count":2476,"reading_minutes":11,"published_at":"2026-10-02T00:00:00.000Z","added_at":"2026-10-05T05:08:05.948Z","updated_at":"2026-10-05T05:08:05.948Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":true},"profile_url":"https://listedarticles.com/articles/arch-specific-simd-in-go","markdown_url":"https://listedarticles.com/articles/arch-specific-simd-in-go.md","example":false,"citation":"Junyang Shao and David Chase, The Go Blog. \"Arch-specific SIMD in Go.\" 2 Oct 2026. https://go.dev/blog/archsimd (CC-BY-4.0)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://go.dev/blog/archsimd"},"body_markdown":"Since Go 1.26, we have supported an experimental `amd64` SIMD API. Now in Go 1.27, we support two more architectures: `arm64` and `wasm`.\n\nThe new APIs can be accessed by setting `GOEXPERIMENT=simd` when building a program. They are defined in the `simd/archsimd` package.\n\n`archsimd` is the lower-level infrastructure that `simd` builds on, essentially the intrinsics layer in other languages. For exotic operations that only exist on a specific architecture, users can access them by transitioning from `simd` to `archsimd` (via `ToArch` and `<Types>FromArch`). For details, please refer to the [simd blog](https://go.dev/blog/simd-experiment) post.\n\nFor readers with prior knowledge of SIMD intrinsics in other languages, you might find that our `archsimd` API does not directly mirror the underlying hardware instructions:\n\n- We give our APIs sensible names: e.g., instead of `_mm512_slli_epi64` , we just call it`ShiftAllLeft` .\n- We leave some architecture details to compiler optimizations rather than exposing them in the API: e.g., instead of `_mm512_maskz_add_ps(m, x, y)` , we optimize`x.Add(y).Masked(m)` into the same instruction.\n\nThese design decisions make our APIs smaller and more accessible, and they also happily make the `simd` implementation much easier.\n\nFor readers without prior knowledge of SIMD intrinsics, we hope our low-level API is a natural and pleasant journey for you.\n\nPlease share your feedback and comments. Anything about usability, performance, etc., is appreciated. The parent issue for this project is [#73787](https://go.dev/issue/73787).\n\n## What is SIMD\n\nSIMD stands for Single Instruction, Multiple Data. SIMD instructions operate on wide registers (128-bit, 256-bit, or more) containing multiple 8- to 64-bit elements, performing operations across all elements in parallel. Wider registers and parallel execution provide a significant performance boost to algorithms that can take advantage of them.\n\nUntil now, anyone who wanted to use SIMD in Go had to endure the friction of writing assembly language. The new `archsimd` package is here to change that.\n\n## Goals\n\nThe goal is to support as many SIMD instructions across different architectures as possible. Currently, we have `amd64`, `arm64`, and `wasm` available. Our `amd64` support has the largest API surface, covering AVX, AVX2, and many AVX-512 extensions. Our `arm64` support currently covers NEON, with SVE and some SVE2 already on the way. Our `wasm` support is also mostly complete, thanks to its well-defined 128-bit SIMD abstraction.\n\nSome users might be interested in matrix extensions on certain architectures, such as AMX and SME. We haven’t fully settled on how to represent matrices effectively in Go, so they are not yet supported, but we might support them in the future.\n\n## API Design and Naming\n\n### Types\n\nThe architecture-specific SIMD API is imported from `simd/archsimd`. `archsimd` uses distinct struct types, such as `Float32x4`, `Int32x8`, or `Uint8x16`, to signal the shape and element type of each vector, and defines operations as methods on those types. At this moment, `archsimd` supports only fixed-width vector extensions. Scalable vector extensions like `arm64` SVE and RVV will have different struct types to represent vectors. Mask types also come with vector-like shapes, e.g., `Mask32x4`.\n\n### Methods\n\nAcross different vector widths, element types, and architectures, we use the same method name whenever operations are semantically identical:\n\n- Regardless of vector width or architecture, element-wise addition is simply `x.Add(y)` ; the type of`x` determines which instruction to emit. Instead of 18 different function names for addition, there is only one method name.\n- When instructions on different architectures look similar at a glance but differ in edge-case behavior, we give them distinct names so differences don’t silently bite you. For example, 16-byte table lookup on `amd64` (`VPSHUFB` ) zeros the output byte when the index is negative (`index < 0` ) and wraps modulo 16 otherwise, so we call it`PermuteOrZero` . On`arm64` NEON (`VTBL` ) and`wasm` (`i8x16.swizzle` ), any out-of-range index (`index < 0 || index >= 16` ) zeros the output byte, so we call it`LookupOrZero` .\n- On `amd64` , many 256-bit and 512-bit instructions operate independently within 128-bit lanes rather than across the entire register. We explicitly suffix those methods with`Grouped` (such as`InterleaveLoGrouped` or`PermuteOrZeroGrouped` ) to make the 128-bit lane boundary obvious.\n\nMethods require a receiver value, which doesn’t work for creating a vector from memory or a scalar. For those, `archsimd` provides package-level functions named with the target vector type:\n\n- **Slices (the default)** : In Go 1.27, loading and storing slices use the shortest names:`archsimd.LoadFloat32x4(s []float32) Float32x4` and`x.Store(s []float32)` . For tail elements that may be shorter than a full vector,`archsimd.LoadFloat32x4Part(s []float32) (Float32x4, int)` zero-fills the remaining lanes and returns the number of elements loaded, paired with`x.StorePart(s []float32) int` .\n- **Arrays and Broadcasts** : Fixed-size array pointers use the`Array` suffix (`archsimd.LoadFloat32x4Array(y *[4]float32)` and`x.StoreArray(y *[4]float32)` ), and`archsimd.BroadcastFloat32x4(v float32)` broadcasts a scalar to all lanes.\n\n### Type Reinterpretation\n\nReinterpretations are type casts without any cost. SIMD algorithms frequently reinterpret bits across element types and widths.\n\nIn the Go 1.26 experiment, we implemented this with `As<Type>` methods, which suffered from a quadratic explosion of type pairs and couldn’t generalize to width-agnostic vectors in `simd` or SVE.\n\nIn Go 1.27, we replace them with composable zero-cost conversions:\n\n- `ToBits()` reinterprets a signed integer or float vector as an unsigned integer vector of the same element width (e.g.,`Int32x4.ToBits() -> Uint32x4` ), and`BitsToInt32()` /`BitsToFloat32()` converts back.\n- `ReshapeToUint<W>s()` on unsigned vectors changes the element width within the same register width. For example, if`x` is an`archsimd.Uint8x16` , then`x.ReshapeToUint32s().BitsToFloat32()` reinterprets the bits of`x` as an`archsimd.Float32x4` at zero runtime cost.\n\n### Masks\n\nDifferent architectures encode vector masks very differently, e.g. `k` mask registers on AVX-512, full vector bitmasks on AVX/AVX2, NEON, and `wasm`, and predicate registers on SVE. To hide this hardware detail, `archsimd` provides opaque `Mask` types (such as `Mask32x4`) matching each vector shape:\n\n- Comparisons like `x.Greater(y)` produce a`Mask` .\n- `x.Masked(m)` zeros elements where`m` is false, and`x.IfElse(m, y)` (which replaces`Merge` from Go 1.26) selects elements from`x` where`m` is true and from`y` where`m` is false.\n- The compiler peepholes mask operations with the surrounding instructions: on AVX-512, `x.Add(y).Masked(m)` or`x.Add(y).IfElse(m, z)` compiles into a single zero-masked or merge-masked`VPADD` instruction; on AVX2, NEON, or`wasm` , it lowers to the appropriate bitwise`AND` or blend/bitselect instruction.\n\n## Examples\n\nPortable algorithms like inner product can easily be implemented using the portable `simd` package, as shown in the companion [simd blog](https://go.dev/blog/simd-experiment) post. However, we still expose `archsimd` to tap into exotic, architecture-specific instructions that do not fit into a portable intersection API, or instructions that can simply do things faster than what `simd` provides.\n\nOn `amd64`, the GFNI extension exposes `GaloisFieldAffineTransform` operations. Putting aside what a Galois Field is, these operations do a simple thing:\n\n```\n// Each element of A is interpreted as a matrix of 8 rows of vectors of 8 bits.\n// Each element of x is interpreted as a vector of 8 bits.\n// The b argument is likewise a vector of 8 bits.\n// The result is z[i] = (A[i/8] * x[i]) + b.\n//\n// The matrix * and vector + follow the usual rules for linear algebra,\n// but based on AND and XOR for scalar * and +\n//\n// Asm: VGF2P8AFFINEQB, CPU Feature: AVX512GFNI\nfunc (x Uint8x16) GaloisFieldAffineTransform(A Uint64x2, b uint8) Uint8x16\n```\nEach 8x8 bit matrix `A` is packed row-by-row into a single `uint64`. Because any linear combination or permutation of the 8 bits within a byte can be written as an 8x8 bit matrix, this single instruction is a Swiss Army knife for byte-level bit manipulation.\n\nFor example, reversing the bit order of every byte (`bits.Reverse8`) normally requires nibble lookup tables, masks, shifts, and ORs in SIMD. With `GaloisFieldAffineTransform`, we can simply multiply each byte by the 8x8 anti-diagonal identity matrix (`0x8040201008040201`) to reverse all 8 bits of 64 bytes at a time in a single instruction:\n\n```\n/*\n    Anti-diagonal 8x8 bit matrix (0x8040201008040201):\n    [ 0 0 0 0 0 0 0 1 ]   [ x7 ]   [ x0 ]\n    [ 0 0 0 0 0 0 1 0 ]   [ x6 ]   [ x1 ]\n    [ 0 0 0 0 0 1 0 0 ]   [ x5 ]   [ x2 ]\n    [ 0 0 0 0 1 0 0 0 ] * [ x4 ] = [ x3 ]\n    [ 0 0 0 1 0 0 0 0 ]   [ x3 ]   [ x4 ]\n    [ 0 0 1 0 0 0 0 0 ]   [ x2 ]   [ x5 ]\n    [ 0 1 0 0 0 0 0 0 ]   [ x1 ]   [ x6 ]\n    [ 1 0 0 0 0 0 0 0 ]   [ x0 ]   [ x7 ]\n*/\nfunc ReverseBits(dst, src []uint8) {\n    if !archsimd.X86.AVX512GFNI() {\n        // A slow fallback emulation impl, details omitted.\n        slowReverseBits(dst, src)\n        return\n    }\n    if len(src) > len(dst) {\n        // Make sure src and dst are the same size\n        src = src[:len(dst)]\n    }\n    revMatrix := archsimd.BroadcastUint64x8(0x8040201008040201)\n    var v archsimd.Uint8x64\n    var i int\n    for i = 0; i < len(src)-v.Len()+1; i += v.Len() {\n        v = archsimd.LoadUint8x64(src[i : i+v.Len()])\n        v.GaloisFieldAffineTransform(revMatrix, 0).Store(dst[i : i+v.Len()])\n    }\n    if i < len(src) {\n        v, _ = archsimd.LoadUint8x64Part(src[i:])\n        v.GaloisFieldAffineTransform(revMatrix, 0).StorePart(dst[i:])\n    }\n}\n```\nAnother example is matrix transpose. On `amd64`, the permutation instructions are shaped in a non-portable way, so using `simd` is cumbersome and might incur emulations. An efficient implementation can be written more directly in `archsimd`. The example below transposes an 8-by-8 matrix of 32-bit integers in registers on `amd64` using AVX2’s 128-bit lane-grouped interleaves (`InterleaveLoGrouped`, `InterleaveHiGrouped`) and cross-lane permutations (`ConcatPermuteScalarsGrouped`, `ConcatPermute128Scalars`):\n\n```\nfunc Transpose8(a0, a1, a2, a3, a4, a5, a6, a7 archsimd.Int32x8) (\n    b0, b1, b2, b3, b4, b5, b6, b7 archsimd.Int32x8) {\n    if !archsimd.X86.AVX2() {\n        // A slow fallback emulation impl, details omitted.\n        return slowTranspose8(a0, a1, a2, a3, a4, a5, a6, a7)\n    }\n    /*\n            LOW  HIGH\n        a0: abcd efgh\n        a1: ijkl mnop\n        a2: qrst uvwx\n        a3: 0123 4567\n        a4: ABCD EFGH\n        a5: IJKL MNOP\n        a6: QRST UVWX\n        a7: 89yz YZ$@\n    */\n    t0 := a0.InterleaveLoGrouped(a1) // t0 = aibj emfn\n    t1 := a0.InterleaveHiGrouped(a1) // t1 = ckdl gohp\n    t2 := a2.InterleaveLoGrouped(a3) // t2 = q0r1 u4v5\n    t3 := a2.InterleaveHiGrouped(a3) // t3 = s2t3 w6x7\n    t4 := a4.InterleaveLoGrouped(a5) // t4 = AIBJ EMFN\n    t5 := a4.InterleaveHiGrouped(a5) // t5 = CKDL GOHP\n    t6 := a6.InterleaveLoGrouped(a7) // t6 = Q8R9 UYVZ\n    t7 := a6.InterleaveHiGrouped(a7) // t7 = SyTz W$X@\n    a0 = t0.ConcatPermuteScalarsGrouped(0, 1, 4, 5, t2) // a0 = aiq0 emu4\n    a1 = t0.ConcatPermuteScalarsGrouped(2, 3, 6, 7, t2) // a1 = bjr1 fnv5\n    a2 = t1.ConcatPermuteScalarsGrouped(0, 1, 4, 5, t3) // a2 = cks2 gow6\n    a3 = t1.ConcatPermuteScalarsGrouped(2, 3, 6, 7, t3) // a3 = dlt3 hpx7\n    a4 = t4.ConcatPermuteScalarsGrouped(0, 1, 4, 5, t6) // a4 = AIQ8 EMUY\n    a5 = t4.ConcatPermuteScalarsGrouped(2, 3, 6, 7, t6) // a5 = BJR9 FNVZ\n    a6 = t5.ConcatPermuteScalarsGrouped(0, 1, 4, 5, t7) // a6 = CKSy GOW$\n    a7 = t5.ConcatPermuteScalarsGrouped(2, 3, 6, 7, t7) // a7 = DLTz HPX@\n    b0 = a0.ConcatPermute128Scalars(0, 2, a4) // b0 = aiq0 AIQ8\n    b1 = a1.ConcatPermute128Scalars(0, 2, a5) // b1 = bjr1 BJR9\n    b2 = a2.ConcatPermute128Scalars(0, 2, a6) // b2 = cks2 CKSy\n    b3 = a3.ConcatPermute128Scalars(0, 2, a7) // b3 = dlt3 DLTz\n    b4 = a0.ConcatPermute128Scalars(1, 3, a4) // b4 = emu4 EMUY\n    b5 = a1.ConcatPermute128Scalars(1, 3, a5) // b5 = fnv5 FNVZ\n    b6 = a2.ConcatPermute128Scalars(1, 3, a6) // b6 = gow6 GOW$\n    b7 = a3.ConcatPermute128Scalars(1, 3, a7) // b7 = hpx7 HPX@\n    return\n}\n```\n## Good Practices\n\nWhen walking through the examples above, you might have a few questions:\n\n- What is `archsimd.X86.AVX512GFNI()` , and can I omit it?\n- Why is the strided loop written as `i = 0; i < len(src)-v.Len()+1; i += v.Len()` , instead of`i = 0; i + v.Len() <= len(src); i += v.Len()` ?\n- Why does `Transpose8` take 8 vector parameters rather than a struct or an`[8]archsimd.Int32x8` array?\n\nAt first glance these choices might look like minor stylistic preferences, but each of them is chosen for an important performance or correctness reason:\n\n-\nCPU feature checks ( `archsimd.X86.*` ,`archsimd.ARM64.*` )On `amd64` , we provide runtime feature checks like`archsimd.X86.AVX()` ,`AVX2()` ,`AVX512()` , and extension checks like`AVX512GFNI()` or`AVX512VNNI()` ; on`arm64` , NEON is always available, and we provide checks for optional extensions like`archsimd.ARM64.PMULL()` ,`SVE()` , and`SVE2()` . The compiler doesn’t know ahead of time whether the machine running your binary supports optional SIMD extensions. If you don’t guard your SIMD code with the appropriate feature check, your program may crash with`SIGILL` on hardware that lacks the instruction.Just as importantly, CPU feature checks act as **compiler optimization hints** . For example, AVX-512 supports hardware merge-masking on instruction outputs. How the compiler lowers`x.Add(y).IfElse(m, z)` on a 128-bit or 256-bit vector depends on which CPU features are known to be available. If the code is not guarded by`archsimd.X86.AVX512()` , the compiler must conservatively emit a vector`VPADD` followed by a`VPBLEND` instruction. Inside an`if archsimd.X86.AVX512()` block, however, the compiler knows AVX-512 is available and fuses the sequence into a single merge-masked AVX-512`VPADD` instruction.\n-\nBounds check elimination in strided loops How you write the loop bound determines whether the compiler’s `prove` pass can eliminate slice bounds checks inside the loop. In`i + v.Len() <= len(src)` , the addition`i + v.Len()` could theoretically overflow to a negative`int` if`len(src)` were near`math.MaxInt` , which prevents the compiler from proving that`i` is always non-negative and in-bounds. Writing`i < len(src)-v.Len()+1` avoids that potential overflow, allowing the loop body to compile with zero bounds checks.\n-\nKeeping vectors in registers (avoiding large composite types) `Transpose8` passes 8 vector parameters rather than a struct of 8 fields or an`[8]archsimd.Int32x8` array because Go’s ABI and SSA backend currently place large composite types in memory rather than in registers (a known issue tracked at[#24416](https://go.dev/issue/24416) ). Because SIMD vectors are 16 to 64 bytes each, spilling them to the stack hurts performance much more than spilling scalar values. This applies to both function parameters and local variables, so avoid wrapping SIMD vectors inside large arrays or structs in hot loops. There are ongoing efforts to promote composite memory accesses back into registers, which have not yet landed in Go 1.27.\n\n## Trying it out\n\nIn Go 1.27, `GOEXPERIMENT=simd` supports `amd64` (AVX, AVX2, and AVX-512), `arm64` (NEON), and `wasm` (WebAssembly 128-bit SIMD) out of the box. Whether you are on an Intel/AMD machine, an `arm64` server or Apple Silicon Mac (M1/M2/M3/M4), or targeting WebAssembly, you can try `archsimd` natively right away:\n\n```\nGOEXPERIMENT=simd go test simd/archsimd/...\n```\nYou can also cross-test `wasm` (or `amd64` via Rosetta emulation on Apple Silicon) by setting `GOARCH`:\n\n```\n# Run wasm SIMD tests using a WASI runtime:\nGOOS=wasip1 GOARCH=wasm GOEXPERIMENT=simd go test simd/archsimd/...\n# Or run amd64 AVX2 tests on Apple Silicon via Rosetta:\nGOARCH=amd64 GOEXPERIMENT=simd go test simd/archsimd/...\n```\nWe’re interested in all sorts of feedback! Early users have already helped us catch bugs (such as [#77582](https://go.dev/issue/77582)) and identify places where code generation or API ergonomics could be improved. We tried hard to pick clear, consistent names and cover the most useful instructions, but `archsimd` is a massive API surface and there is always room to improve.\n\n## Future work\n\nFor Go 1.28 and beyond, we are actively working on completing `arm64` SVE and SVE2 support in `archsimd` (introducing width-agnostic scalable vector types like `archsimd.Float32s` backed directly by hardware SVE registers and predicates) and wiring it into `simd`. We also plan to expand `archsimd` to additional architectures such as `riscv64`, `ppc64`, `s390x`, and `loong64`, improve register promotion for composite types, and continue filling in instructions and compiler optimizations based on community feedback.\n","body_html":"<p>Since Go 1.26, we have supported an experimental <code>amd64</code> SIMD API. Now in Go 1.27, we support two more architectures: <code>arm64</code> and <code>wasm</code>.</p>\n<p>The new APIs can be accessed by setting <code>GOEXPERIMENT=simd</code> when building a program. They are defined in the <code>simd/archsimd</code> package.</p>\n<p><code>archsimd</code> is the lower-level infrastructure that <code>simd</code> builds on, essentially the intrinsics layer in other languages. For exotic operations that only exist on a specific architecture, users can access them by transitioning from <code>simd</code> to <code>archsimd</code> (via <code>ToArch</code> and <code>&lt;Types&gt;FromArch</code>). For details, please refer to the <a href=\"https://go.dev/blog/simd-experiment\" rel=\"nofollow ugc noopener\">simd blog</a> post.</p>\n<p>For readers with prior knowledge of SIMD intrinsics in other languages, you might find that our <code>archsimd</code> API does not directly mirror the underlying hardware instructions:</p>\n<ul><li>We give our APIs sensible names: e.g., instead of <code>_mm512_slli_epi64</code> , we just call it<code>ShiftAllLeft</code> .</li><li>We leave some architecture details to compiler optimizations rather than exposing them in the API: e.g., instead of <code>_mm512_maskz_add_ps(m, x, y)</code> , we optimize<code>x.Add(y).Masked(m)</code> into the same instruction.</li></ul>\n<p>These design decisions make our APIs smaller and more accessible, and they also happily make the <code>simd</code> implementation much easier.</p>\n<p>For readers without prior knowledge of SIMD intrinsics, we hope our low-level API is a natural and pleasant journey for you.</p>\n<p>Please share your feedback and comments. Anything about usability, performance, etc., is appreciated. The parent issue for this project is <a href=\"https://go.dev/issue/73787\" rel=\"nofollow ugc noopener\">#73787</a>.</p>\n<h2 id=\"what-is-simd\">What is SIMD</h2>\n<p>SIMD stands for Single Instruction, Multiple Data. SIMD instructions operate on wide registers (128-bit, 256-bit, or more) containing multiple 8- to 64-bit elements, performing operations across all elements in parallel. Wider registers and parallel execution provide a significant performance boost to algorithms that can take advantage of them.</p>\n<p>Until now, anyone who wanted to use SIMD in Go had to endure the friction of writing assembly language. The new <code>archsimd</code> package is here to change that.</p>\n<h2 id=\"goals\">Goals</h2>\n<p>The goal is to support as many SIMD instructions across different architectures as possible. Currently, we have <code>amd64</code>, <code>arm64</code>, and <code>wasm</code> available. Our <code>amd64</code> support has the largest API surface, covering AVX, AVX2, and many AVX-512 extensions. Our <code>arm64</code> support currently covers NEON, with SVE and some SVE2 already on the way. Our <code>wasm</code> support is also mostly complete, thanks to its well-defined 128-bit SIMD abstraction.</p>\n<p>Some users might be interested in matrix extensions on certain architectures, such as AMX and SME. We haven’t fully settled on how to represent matrices effectively in Go, so they are not yet supported, but we might support them in the future.</p>\n<h2 id=\"api-design-and-naming\">API Design and Naming</h2>\n<h3 id=\"types\">Types</h3>\n<p>The architecture-specific SIMD API is imported from <code>simd/archsimd</code>. <code>archsimd</code> uses distinct struct types, such as <code>Float32x4</code>, <code>Int32x8</code>, or <code>Uint8x16</code>, to signal the shape and element type of each vector, and defines operations as methods on those types. At this moment, <code>archsimd</code> supports only fixed-width vector extensions. Scalable vector extensions like <code>arm64</code> SVE and RVV will have different struct types to represent vectors. Mask types also come with vector-like shapes, e.g., <code>Mask32x4</code>.</p>\n<h3 id=\"methods\">Methods</h3>\n<p>Across different vector widths, element types, and architectures, we use the same method name whenever operations are semantically identical:</p>\n<ul><li>Regardless of vector width or architecture, element-wise addition is simply <code>x.Add(y)</code> ; the type of<code>x</code> determines which instruction to emit. Instead of 18 different function names for addition, there is only one method name.</li><li>When instructions on different architectures look similar at a glance but differ in edge-case behavior, we give them distinct names so differences don’t silently bite you. For example, 16-byte table lookup on <code>amd64</code> (<code>VPSHUFB</code> ) zeros the output byte when the index is negative (<code>index &lt; 0</code> ) and wraps modulo 16 otherwise, so we call it<code>PermuteOrZero</code> . On<code>arm64</code> NEON (<code>VTBL</code> ) and<code>wasm</code> (<code>i8x16.swizzle</code> ), any out-of-range index (<code>index &lt; 0 || index &gt;= 16</code> ) zeros the output byte, so we call it<code>LookupOrZero</code> .</li><li>On <code>amd64</code> , many 256-bit and 512-bit instructions operate independently within 128-bit lanes rather than across the entire register. We explicitly suffix those methods with<code>Grouped</code> (such as<code>InterleaveLoGrouped</code> or<code>PermuteOrZeroGrouped</code> ) to make the 128-bit lane boundary obvious.</li></ul>\n<p>Methods require a receiver value, which doesn’t work for creating a vector from memory or a scalar. For those, <code>archsimd</code> provides package-level functions named with the target vector type:</p>\n<ul><li><strong>Slices (the default)</strong> : In Go 1.27, loading and storing slices use the shortest names:<code>archsimd.LoadFloat32x4(s []float32) Float32x4</code> and<code>x.Store(s []float32)</code> . For tail elements that may be shorter than a full vector,<code>archsimd.LoadFloat32x4Part(s []float32) (Float32x4, int)</code> zero-fills the remaining lanes and returns the number of elements loaded, paired with<code>x.StorePart(s []float32) int</code> .</li><li><strong>Arrays and Broadcasts</strong> : Fixed-size array pointers use the<code>Array</code> suffix (<code>archsimd.LoadFloat32x4Array(y *[4]float32)</code> and<code>x.StoreArray(y *[4]float32)</code> ), and<code>archsimd.BroadcastFloat32x4(v float32)</code> broadcasts a scalar to all lanes.</li></ul>\n<h3 id=\"type-reinterpretation\">Type Reinterpretation</h3>\n<p>Reinterpretations are type casts without any cost. SIMD algorithms frequently reinterpret bits across element types and widths.</p>\n<p>In the Go 1.26 experiment, we implemented this with <code>As&lt;Type&gt;</code> methods, which suffered from a quadratic explosion of type pairs and couldn’t generalize to width-agnostic vectors in <code>simd</code> or SVE.</p>\n<p>In Go 1.27, we replace them with composable zero-cost conversions:</p>\n<ul><li><code>ToBits()</code> reinterprets a signed integer or float vector as an unsigned integer vector of the same element width (e.g.,<code>Int32x4.ToBits() -&gt; Uint32x4</code> ), and<code>BitsToInt32()</code> /<code>BitsToFloat32()</code> converts back.</li><li><code>ReshapeToUint&lt;W&gt;s()</code> on unsigned vectors changes the element width within the same register width. For example, if<code>x</code> is an<code>archsimd.Uint8x16</code> , then<code>x.ReshapeToUint32s().BitsToFloat32()</code> reinterprets the bits of<code>x</code> as an<code>archsimd.Float32x4</code> at zero runtime cost.</li></ul>\n<h3 id=\"masks\">Masks</h3>\n<p>Different architectures encode vector masks very differently, e.g. <code>k</code> mask registers on AVX-512, full vector bitmasks on AVX/AVX2, NEON, and <code>wasm</code>, and predicate registers on SVE. To hide this hardware detail, <code>archsimd</code> provides opaque <code>Mask</code> types (such as <code>Mask32x4</code>) matching each vector shape:</p>\n<ul><li>Comparisons like <code>x.Greater(y)</code> produce a<code>Mask</code> .</li><li><code>x.Masked(m)</code> zeros elements where<code>m</code> is false, and<code>x.IfElse(m, y)</code> (which replaces<code>Merge</code> from Go 1.26) selects elements from<code>x</code> where<code>m</code> is true and from<code>y</code> where<code>m</code> is false.</li><li>The compiler peepholes mask operations with the surrounding instructions: on AVX-512, <code>x.Add(y).Masked(m)</code> or<code>x.Add(y).IfElse(m, z)</code> compiles into a single zero-masked or merge-masked<code>VPADD</code> instruction; on AVX2, NEON, or<code>wasm</code> , it lowers to the appropriate bitwise<code>AND</code> or blend/bitselect instruction.</li></ul>\n<h2 id=\"examples\">Examples</h2>\n<p>Portable algorithms like inner product can easily be implemented using the portable <code>simd</code> package, as shown in the companion <a href=\"https://go.dev/blog/simd-experiment\" rel=\"nofollow ugc noopener\">simd blog</a> post. However, we still expose <code>archsimd</code> to tap into exotic, architecture-specific instructions that do not fit into a portable intersection API, or instructions that can simply do things faster than what <code>simd</code> provides.</p>\n<p>On <code>amd64</code>, the GFNI extension exposes <code>GaloisFieldAffineTransform</code> operations. Putting aside what a Galois Field is, these operations do a simple thing:</p>\n<pre><code>// Each element of A is interpreted as a matrix of 8 rows of vectors of 8 bits.\n// Each element of x is interpreted as a vector of 8 bits.\n// The b argument is likewise a vector of 8 bits.\n// The result is z[i] = (A[i/8] * x[i]) + b.\n//\n// The matrix * and vector + follow the usual rules for linear algebra,\n// but based on AND and XOR for scalar * and +\n//\n// Asm: VGF2P8AFFINEQB, CPU Feature: AVX512GFNI\nfunc (x Uint8x16) GaloisFieldAffineTransform(A Uint64x2, b uint8) Uint8x16</code></pre>\n<p>Each 8x8 bit matrix <code>A</code> is packed row-by-row into a single <code>uint64</code>. Because any linear combination or permutation of the 8 bits within a byte can be written as an 8x8 bit matrix, this single instruction is a Swiss Army knife for byte-level bit manipulation.</p>\n<p>For example, reversing the bit order of every byte (<code>bits.Reverse8</code>) normally requires nibble lookup tables, masks, shifts, and ORs in SIMD. With <code>GaloisFieldAffineTransform</code>, we can simply multiply each byte by the 8x8 anti-diagonal identity matrix (<code>0x8040201008040201</code>) to reverse all 8 bits of 64 bytes at a time in a single instruction:</p>\n<pre><code>/*\n    Anti-diagonal 8x8 bit matrix (0x8040201008040201):\n    [ 0 0 0 0 0 0 0 1 ]   [ x7 ]   [ x0 ]\n    [ 0 0 0 0 0 0 1 0 ]   [ x6 ]   [ x1 ]\n    [ 0 0 0 0 0 1 0 0 ]   [ x5 ]   [ x2 ]\n    [ 0 0 0 0 1 0 0 0 ] * [ x4 ] = [ x3 ]\n    [ 0 0 0 1 0 0 0 0 ]   [ x3 ]   [ x4 ]\n    [ 0 0 1 0 0 0 0 0 ]   [ x2 ]   [ x5 ]\n    [ 0 1 0 0 0 0 0 0 ]   [ x1 ]   [ x6 ]\n    [ 1 0 0 0 0 0 0 0 ]   [ x0 ]   [ x7 ]\n*/\nfunc ReverseBits(dst, src []uint8) {\n    if !archsimd.X86.AVX512GFNI() {\n        // A slow fallback emulation impl, details omitted.\n        slowReverseBits(dst, src)\n        return\n    }\n    if len(src) &gt; len(dst) {\n        // Make sure src and dst are the same size\n        src = src[:len(dst)]\n    }\n    revMatrix := archsimd.BroadcastUint64x8(0x8040201008040201)\n    var v archsimd.Uint8x64\n    var i int\n    for i = 0; i &lt; len(src)-v.Len()+1; i += v.Len() {\n        v = archsimd.LoadUint8x64(src[i : i+v.Len()])\n        v.GaloisFieldAffineTransform(revMatrix, 0).Store(dst[i : i+v.Len()])\n    }\n    if i &lt; len(src) {\n        v, _ = archsimd.LoadUint8x64Part(src[i:])\n        v.GaloisFieldAffineTransform(revMatrix, 0).StorePart(dst[i:])\n    }\n}</code></pre>\n<p>Another example is matrix transpose. On <code>amd64</code>, the permutation instructions are shaped in a non-portable way, so using <code>simd</code> is cumbersome and might incur emulations. An efficient implementation can be written more directly in <code>archsimd</code>. The example below transposes an 8-by-8 matrix of 32-bit integers in registers on <code>amd64</code> using AVX2’s 128-bit lane-grouped interleaves (<code>InterleaveLoGrouped</code>, <code>InterleaveHiGrouped</code>) and cross-lane permutations (<code>ConcatPermuteScalarsGrouped</code>, <code>ConcatPermute128Scalars</code>):</p>\n<pre><code>func Transpose8(a0, a1, a2, a3, a4, a5, a6, a7 archsimd.Int32x8) (\n    b0, b1, b2, b3, b4, b5, b6, b7 archsimd.Int32x8) {\n    if !archsimd.X86.AVX2() {\n        // A slow fallback emulation impl, details omitted.\n        return slowTranspose8(a0, a1, a2, a3, a4, a5, a6, a7)\n    }\n    /*\n            LOW  HIGH\n        a0: abcd efgh\n        a1: ijkl mnop\n        a2: qrst uvwx\n        a3: 0123 4567\n        a4: ABCD EFGH\n        a5: IJKL MNOP\n        a6: QRST UVWX\n        a7: 89yz YZ$@\n    */\n    t0 := a0.InterleaveLoGrouped(a1) // t0 = aibj emfn\n    t1 := a0.InterleaveHiGrouped(a1) // t1 = ckdl gohp\n    t2 := a2.InterleaveLoGrouped(a3) // t2 = q0r1 u4v5\n    t3 := a2.InterleaveHiGrouped(a3) // t3 = s2t3 w6x7\n    t4 := a4.InterleaveLoGrouped(a5) // t4 = AIBJ EMFN\n    t5 := a4.InterleaveHiGrouped(a5) // t5 = CKDL GOHP\n    t6 := a6.InterleaveLoGrouped(a7) // t6 = Q8R9 UYVZ\n    t7 := a6.InterleaveHiGrouped(a7) // t7 = SyTz W$X@\n    a0 = t0.ConcatPermuteScalarsGrouped(0, 1, 4, 5, t2) // a0 = aiq0 emu4\n    a1 = t0.ConcatPermuteScalarsGrouped(2, 3, 6, 7, t2) // a1 = bjr1 fnv5\n    a2 = t1.ConcatPermuteScalarsGrouped(0, 1, 4, 5, t3) // a2 = cks2 gow6\n    a3 = t1.ConcatPermuteScalarsGrouped(2, 3, 6, 7, t3) // a3 = dlt3 hpx7\n    a4 = t4.ConcatPermuteScalarsGrouped(0, 1, 4, 5, t6) // a4 = AIQ8 EMUY\n    a5 = t4.ConcatPermuteScalarsGrouped(2, 3, 6, 7, t6) // a5 = BJR9 FNVZ\n    a6 = t5.ConcatPermuteScalarsGrouped(0, 1, 4, 5, t7) // a6 = CKSy GOW$\n    a7 = t5.ConcatPermuteScalarsGrouped(2, 3, 6, 7, t7) // a7 = DLTz HPX@\n    b0 = a0.ConcatPermute128Scalars(0, 2, a4) // b0 = aiq0 AIQ8\n    b1 = a1.ConcatPermute128Scalars(0, 2, a5) // b1 = bjr1 BJR9\n    b2 = a2.ConcatPermute128Scalars(0, 2, a6) // b2 = cks2 CKSy\n    b3 = a3.ConcatPermute128Scalars(0, 2, a7) // b3 = dlt3 DLTz\n    b4 = a0.ConcatPermute128Scalars(1, 3, a4) // b4 = emu4 EMUY\n    b5 = a1.ConcatPermute128Scalars(1, 3, a5) // b5 = fnv5 FNVZ\n    b6 = a2.ConcatPermute128Scalars(1, 3, a6) // b6 = gow6 GOW$\n    b7 = a3.ConcatPermute128Scalars(1, 3, a7) // b7 = hpx7 HPX@\n    return\n}</code></pre>\n<h2 id=\"good-practices\">Good Practices</h2>\n<p>When walking through the examples above, you might have a few questions:</p>\n<ul><li>What is <code>archsimd.X86.AVX512GFNI()</code> , and can I omit it?</li><li>Why is the strided loop written as <code>i = 0; i &lt; len(src)-v.Len()+1; i += v.Len()</code> , instead of<code>i = 0; i + v.Len() &lt;= len(src); i += v.Len()</code> ?</li><li>Why does <code>Transpose8</code> take 8 vector parameters rather than a struct or an<code>[8]archsimd.Int32x8</code> array?</li></ul>\n<p>At first glance these choices might look like minor stylistic preferences, but each of them is chosen for an important performance or correctness reason:</p>\n<p>-\nCPU feature checks ( <code>archsimd.X86.*</code> ,<code>archsimd.ARM64.*</code> )On <code>amd64</code> , we provide runtime feature checks like<code>archsimd.X86.AVX()</code> ,<code>AVX2()</code> ,<code>AVX512()</code> , and extension checks like<code>AVX512GFNI()</code> or<code>AVX512VNNI()</code> ; on<code>arm64</code> , NEON is always available, and we provide checks for optional extensions like<code>archsimd.ARM64.PMULL()</code> ,<code>SVE()</code> , and<code>SVE2()</code> . The compiler doesn’t know ahead of time whether the machine running your binary supports optional SIMD extensions. If you don’t guard your SIMD code with the appropriate feature check, your program may crash with<code>SIGILL</code> on hardware that lacks the instruction.Just as importantly, CPU feature checks act as <strong>compiler optimization hints</strong> . For example, AVX-512 supports hardware merge-masking on instruction outputs. How the compiler lowers<code>x.Add(y).IfElse(m, z)</code> on a 128-bit or 256-bit vector depends on which CPU features are known to be available. If the code is not guarded by<code>archsimd.X86.AVX512()</code> , the compiler must conservatively emit a vector<code>VPADD</code> followed by a<code>VPBLEND</code> instruction. Inside an<code>if archsimd.X86.AVX512()</code> block, however, the compiler knows AVX-512 is available and fuses the sequence into a single merge-masked AVX-512<code>VPADD</code> instruction.\n-\nBounds check elimination in strided loops How you write the loop bound determines whether the compiler’s <code>prove</code> pass can eliminate slice bounds checks inside the loop. In<code>i + v.Len() &lt;= len(src)</code> , the addition<code>i + v.Len()</code> could theoretically overflow to a negative<code>int</code> if<code>len(src)</code> were near<code>math.MaxInt</code> , which prevents the compiler from proving that<code>i</code> is always non-negative and in-bounds. Writing<code>i &lt; len(src)-v.Len()+1</code> avoids that potential overflow, allowing the loop body to compile with zero bounds checks.\n-\nKeeping vectors in registers (avoiding large composite types) <code>Transpose8</code> passes 8 vector parameters rather than a struct of 8 fields or an<code>[8]archsimd.Int32x8</code> array because Go’s ABI and SSA backend currently place large composite types in memory rather than in registers (a known issue tracked at<a href=\"https://go.dev/issue/24416\" rel=\"nofollow ugc noopener\">#24416</a> ). Because SIMD vectors are 16 to 64 bytes each, spilling them to the stack hurts performance much more than spilling scalar values. This applies to both function parameters and local variables, so avoid wrapping SIMD vectors inside large arrays or structs in hot loops. There are ongoing efforts to promote composite memory accesses back into registers, which have not yet landed in Go 1.27.</p>\n<h2 id=\"trying-it-out\">Trying it out</h2>\n<p>In Go 1.27, <code>GOEXPERIMENT=simd</code> supports <code>amd64</code> (AVX, AVX2, and AVX-512), <code>arm64</code> (NEON), and <code>wasm</code> (WebAssembly 128-bit SIMD) out of the box. Whether you are on an Intel/AMD machine, an <code>arm64</code> server or Apple Silicon Mac (M1/M2/M3/M4), or targeting WebAssembly, you can try <code>archsimd</code> natively right away:</p>\n<pre><code>GOEXPERIMENT=simd go test simd/archsimd/...</code></pre>\n<p>You can also cross-test <code>wasm</code> (or <code>amd64</code> via Rosetta emulation on Apple Silicon) by setting <code>GOARCH</code>:</p>\n<pre><code># Run wasm SIMD tests using a WASI runtime:\nGOOS=wasip1 GOARCH=wasm GOEXPERIMENT=simd go test simd/archsimd/...\n# Or run amd64 AVX2 tests on Apple Silicon via Rosetta:\nGOARCH=amd64 GOEXPERIMENT=simd go test simd/archsimd/...</code></pre>\n<p>We’re interested in all sorts of feedback! Early users have already helped us catch bugs (such as <a href=\"https://go.dev/issue/77582\" rel=\"nofollow ugc noopener\">#77582</a>) and identify places where code generation or API ergonomics could be improved. We tried hard to pick clear, consistent names and cover the most useful instructions, but <code>archsimd</code> is a massive API surface and there is always room to improve.</p>\n<h2 id=\"future-work\">Future work</h2>\n<p>For Go 1.28 and beyond, we are actively working on completing <code>arm64</code> SVE and SVE2 support in <code>archsimd</code> (introducing width-agnostic scalable vector types like <code>archsimd.Float32s</code> backed directly by hardware SVE registers and predicates) and wiring it into <code>simd</code>. We also plan to expand <code>archsimd</code> to additional architectures such as <code>riscv64</code>, <code>ppc64</code>, <code>s390x</code>, and <code>loong64</code>, improve register promotion for composite types, and continue filling in instructions and compiler optimizations based on community feedback.</p>","headings":[{"level":2,"text":"What is SIMD","id":"what-is-simd"},{"level":2,"text":"Goals","id":"goals"},{"level":2,"text":"API Design and Naming","id":"api-design-and-naming"},{"level":3,"text":"Types","id":"types"},{"level":3,"text":"Methods","id":"methods"},{"level":3,"text":"Type Reinterpretation","id":"type-reinterpretation"},{"level":3,"text":"Masks","id":"masks"},{"level":2,"text":"Examples","id":"examples"},{"level":2,"text":"Good Practices","id":"good-practices"},{"level":2,"text":"Trying it out","id":"trying-it-out"},{"level":2,"text":"Future work","id":"future-work"}]}}