{"article":{"slug":"hacking-the-go-compiler-to-efficiently-map-ipv4-to-ipv6","title":"Hacking the Go compiler to efficiently map IPv4 to IPv6","subtitle":null,"summary":"Go maintainers declined a netip.Addr.Map() method, saying the compiler should optimize netip.AddrFrom16(ip.As16()) instead. Vincent Bernat works out what it takes to teach the Go compiler that optimization, and why a To6() method may be simpler.","content_type":"blog_post","language":"en","canonical_url":"https://vincent.bernat.ch/en/blog/2026-go-netip-addrto6","author":{"name":"Vincent Bernat","url":"https://vincent.bernat.ch/","person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"vincent.bernat.ch","url":"https://vincent.bernat.ch/","listing_slug":null,"listing":null},"topics":[{"name":"Programming","slug":"programming","url":"https://listedarticles.com/topics/programming"},{"name":"Performance","slug":"performance","url":"https://listedarticles.com/topics/performance"},{"name":"Networking","slug":"networking","url":"https://listedarticles.com/topics/networking"},{"name":"Systems Programming","slug":"systems-programming","url":"https://listedarticles.com/topics/systems-programming"},{"name":"Open Source","slug":"open-source","url":"https://listedarticles.com/topics/open-source"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":4962,"reading_minutes":22,"published_at":"2026-10-04T00:00:00.000Z","added_at":"2026-10-05T02:15:03.574Z","updated_at":"2026-10-05T02:15:03.574Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":true},"profile_url":"https://listedarticles.com/articles/hacking-the-go-compiler-to-efficiently-map-ipv4-to-ipv6","markdown_url":"https://listedarticles.com/articles/hacking-the-go-compiler-to-efficiently-map-ipv4-to-ipv6.md","example":false,"citation":"Vincent Bernat, vincent.bernat.ch. \"Hacking the Go compiler to efficiently map IPv4 to IPv6.\" 4 Oct 2026. https://vincent.bernat.ch/en/blog/2026-go-netip-addrto6 (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://vincent.bernat.ch/en/blog/2026-go-netip-addrto6"},"body_markdown":"# Hacking the Go compiler to efficiently map IPv4 to IPv6\n\n\n[`netip.Addr`](https://pkg.go.dev/netip.addr) features an [`Unmap()`](https://pkg.go.dev/netip.addr.Unmap)\nmethod returning the unwrapped IPv4 contained in an [IPv4-mapped IPv6\naddress](https://www.rfc-editor.org/rfc/rfc4291#section-2.5.5.2): from `::ffff:203.0.113.10` or `::ffff:cb00:710a`, it returns\n`203.0.113.10`.<sup>[1](#sidenote-why)</sup> There is no `Map()` or `To6()` method for the reverse\ndirection. Such a method is trivial to implement, but Go maintainers have\n[rejected](https://github.com/golang/go/issues/54365#issuecomment-1570607144) it on the grounds that users should write\n`netip.AddrFrom16(ip.As16())` and let the compiler optimize it.<sup>[2](#sidenote-equiv)</sup> Today,\nthis pattern is eight times slower than a native method. How can we teach the\ncompiler to optimize this sequence?\n\n# The alternatives[#](#the-alternatives)\n\nLet’s explore three ways to implement the map semantics for `netip.Addr`. My\nfavorite is to add it to the Go standard library. Go maintainers prefer a small\nexternal helper chaining `netip.AddrFrom16()` and `netip.Addr.As16()`, hoping\nthe compiler eventually optimizes it. The [`unsafe` package](https://pkg.go.dev/unsafe)\nopens a third path, with the same performance as the first solution.\n\n## Modifying the Go standard library[#](#modifying-the-go-standard-library)\n\nInternally, [`netip.Addr`](https://pkg.go.dev/net/netip#Addr) stores any IP address as a\n128-bit value with an extra field `z` to encode the family and the zone:\n\n```\ntype Addr struct {\n    addr uint128\n    z unique.Handle[addrDetail]\n}\ntype addrDetail struct {\n    isV6   bool   // IPv4 is false, IPv6 is true.\n    zoneV6 string // != \"\" only if IsV6 is true.\n}\nvar (\n    z0    unique.Handle[addrDetail]\n    z4    = unique.Make(addrDetail{})\n    z6noz = unique.Make(addrDetail{isV6: true})\n)\n```\n[`AddrFrom4()`](https://pkg.go.dev/net/netip#AddrFrom4) encodes an IPv4 address as an\nIPv4-mapped IPv6 address and sets `z` to the unique value `z4`:\n\n```\n// AddrFrom4 returns the address of the IPv4 address given by the bytes in addr.\nfunc AddrFrom4(addr [4]byte) Addr {\n    return Addr{\n        addr: uint128{\n            0,\n            0xffff00000000 |\n                uint64(addr[0])<<24 | uint64(addr[1])<<16 |\n                uint64(addr[2])<<8 | uint64(addr[3])},\n        z: z4,\n    }\n}\n```\n[`Unmap()`](https://pkg.go.dev/net/netip#Addr.Unmap) turns an IPv4-mapped IPv6 address into an\nIPv4 address by setting the `z` field to `z4`:\n\n```\nfunc (ip Addr) Unmap() Addr {\n    if ip.Is4In6() {\n        ip.z = z4\n    }\n    return ip\n}\n```\nImplementing the reverse direction inside the Go standard library is trivial: we\nset the `z` field to `z6noz` if the address is IPv4.\n\n```\n// To6 maps an IPv4 address to an IPv4-mapped IPv6 address. It returns an\n// IPv6 address unmodified.\nfunc (ip Addr) To6() Addr {\n    if ip.Is4() {\n        ip.z = z6noz\n    }\n    return ip\n}\n```\n## As a helper[#](#as-a-helper)\n\nWe can’t access the `z` field from outside the `net/netip` package. Instead, we\nbuild a small helper around the `netip.AddrFrom16(ip.As16())` pattern:\n\n```\n// AddrTo6 maps an IPv4 address to an IPv4-mapped IPv6 address. It returns an\n// IPv6 address unmodified.\nfunc AddrTo6(ip netip.Addr) netip.Addr {\n    if ip.Is4() {\n        ip = netip.AddrFrom16(ip.As16())\n    }\n    return ip\n}\n```\n## As an unsafe function[#](#as-an-unsafe-function)\n\nAnother solution uses the `unsafe` package to alter the `Addr` struct through\na proxy with the same memory layout:[3](#sidenote-tests)\n\n```\n// addrProxy has the same memory layout as netip.Addr.\ntype addrProxy struct {\n    addr [2]uint64      // netip.uint128\n    z    unsafe.Pointer // unique.Handle[netip.addrDetail]\n}\nvar (\n    anyIPv6    = netip.IPv6Unspecified()\n    netipZ6noz = (*addrProxy)(unsafe.Pointer(&anyIPv6)).z\n)\n// AddrTo6 maps an IPv4 address to an IPv4-mapped IPv6 address. It returns an\n// IPv6 address unmodified.\nfunc AddrTo6(ip netip.Addr) netip.Addr {\n    if !ip.Is4() {\n        return ip\n    }\n    (*addrProxy)(unsafe.Pointer(&ip)).z = netipZ6noz\n    return ip\n}\n```\n## Benchmarks[#](#benchmarks)\n\nOn my computer, with Go 1.27.1, the standard library solution costs 0.88 ns per operation, while the solution favored by Go maintainers costs 7.14 ns. The unsafe solution matches the performance of the first one.\n\n```\ngoos: linux\ngoarch: amd64\npkg: github.com/vincentbernat/go-netip-addrto6\ncpu: AMD Ryzen 5 5600X 6-Core Processor\n                   │     sec/op     │\nAddrTo6/safe            7.137n ± 0%\nAddrTo6/unsafe         0.8682n ± 2%\nAddrTo6/builtin        0.8775n ± 2%\n```\n## Assembly code[#](#assembly-code)\n\nLet’s check the assembly code the compiler generates for each solution.<sup>[4](#sidenote-dis)</sup>\nThe one built into the standard library looks like this:[5](#sidenote-inline)\n\n```\n// AX = input.addr.hi, BX = input.addr.lo, CX = input.z\n CMPQ  net/netip·z4(SB), CX     ; check \"z\" if this is an IPv4 address\n JNE   end                      ; if not, stop here\n MOVQ  net/netip·z6noz(SB), CX  ; CX = netip.z6noz\nend:\n RET\n// return value = Addr{hi: AX, lo: BX, z: CX}\n```\nGo’s [assembly language](https://go.dev/doc/asm) is not a direct representation of the underlying\nmachine language: it operates on a semi-abstract instruction set derived from\n[Plan 9’s assembler](https://9p.io/sys/doc/asm.html). It has four pseudo-registers: FP (frame pointer\nfor function arguments), PC (program counter), SB (static base pointer for\nglobal symbols), and SP (stack pointer). It also has architecture-specific\nregisters like `AX`, `CX`, `DX`, `BX`, `SI`, `DI`, and `R8` to `R15`.\nInstructions storing data use their last argument as the destination.\nInstructions can carry an explicit size suffix: `MOVB` moves a byte, `MOVW` 16\nbits, `MOVL` 32 bits, and `MOVQ` 64 bits. In the example above, the first\ninstruction compares the 64-bit value `z4` with the `CX` register.\n\nThe unsafe solution looks almost the same:\n\n```\n// AX = input.addr.hi, BX = input.addr.lo, CX = input.z\n CMPQ  net/netip·z4(SB), CX  ; check \"z\" if this is an IPv4 address\n JNE   end                   ; if not, stop here\n MOVQ  netipZ6noz(SB), CX    ; CX = netip.z6noz\nend:\n RET\n// return value = Addr{hi: AX, lo: BX, z: CX}\n```\nThe helper solution has far more instructions. To understand why, let’s look at\nthe code for [`As16()`](https://pkg.go.dev/netip.addr.As16) and\n[`AddrFrom16()`](https://pkg.go.dev/netip.AddrFrom16). They are short enough for the compiler\nto inline them.\n\n```\nfunc (ip Addr) As16() (a16 [16]byte) {\n    byteorder.BEPutUint64(a16[:8], ip.addr.hi)\n    byteorder.BEPutUint64(a16[8:], ip.addr.lo)\n    return a16\n}\nfunc AddrFrom16(addr [16]byte) Addr {\n    return Addr{\n        addr: uint128{\n            byteorder.BEUint64(addr[:8]),\n            byteorder.BEUint64(addr[8:]),\n        },\n        z: z6noz,\n    }\n}\n```\nWe can already guess the pattern to optimize: the code packs the IP address into an array, copies it, then unpacks it. If we inline the Go code by hand, we get:\n\n```\nfunc AddrTo6(input netip.Addr) netip.Addr {\n    if !input.Is4() {\n        return input\n    }\n    var a16 [16]byte\n    byteorder.BEPutUint64(a16[:8], input.addr.hi)\n    byteorder.BEPutUint64(a16[8:], input.addr.lo)\n    addr := a16\n    var output netip.Addr\n    output.addr.hi = byteorder.BEUint64(addr[:8])\n    output.addr.lo = byteorder.BEUint64(addr[8:])\n    output.z = netip.z6noz\n    return output\n}\n```\nAs humans, we can mentally derive the optimized form:\n\n```\nfunc AddrTo6(input netip.Addr) netip.Addr {\n    if !input.Is4() {\n        return input\n    }\n    var output netip.Addr\n    output.addr.hi = input.addr.hi\n    output.addr.lo = input.addr.lo\n    output.z = netip.z6noz\n    return output\n}\n```\nUnfortunately, as of Go 1.26.8, the compiler is not smart enough to do the same:\n\n```\n// AX = input.addr.hi, BX = input.addr.lo, CX = input.z\n; Push the stack (32 bytes):\n;    0(SP) addr netip.uint128\n;   16(SP) a16 [16]byte\n PUSHQ   BP\n MOVQ    SP, BP\n SUBQ    $32, SP\n CMPQ    net/netip·z4(SB), CX  ; check \"z\" if this is an IPv4 address\n JNE     end                   ; if not, stop here\n; Pack: byteorder.BEPutUint64(a16[:8], input.addr.hi)\n;       byteorder.BEPutUint64(a16[8:], input.addr.lo)\n MOVBEQ  AX, net/netip·a16+16(SP)\n MOVBEQ  BX, net/netip·a16+24(SP)\n; addr = a16, 16 bytes at once through the vector register X0\n MOVUPS  net/netip·a16+16(SP), X0\n MOVUPS  X0, net/netip·addr(SP)\n; CX = netip.z6noz\n MOVQ    net/netip·z6noz(SB), CX\n; Unpack: output.addr.hi = byteorder.BEUint64(addr[:8])\n;         output.addr.lo = byteorder.BEUint64(addr[8:])\n MOVBEQ  net/netip·addr(SP), AX\n MOVBEQ  net/netip·addr+8(SP), BX\nend:\n ADDQ    $32, SP\n POPQ    BP\n RET\n// return value = Addr{hi: AX, lo: BX, z: CX}\n```\nThe compiler does a decent job on the byte shuffling: the eight byte stores of\n`BEPutUint64()` become a single `MOVBEQ`, which stores a register byte-swapped.\nThe eight byte loads of `BEUint64()` become a single `MOVBEQ` the other way\nround.<sup>[6](#sidenote-amd64v3)</sup> Three groups of instructions remain: a pack, a copy, and an\nunpack.\n\n# Hacking the Go compiler[#](#hacking-the-go-compiler)\n\nThe Go compiler has [several phases](https://github.com/golang/go/blob/1e963f82914ddfb47820034c5c85205a362ed73f/src/cmd/compile/README.md):\n\n- Parsing\n- The compiler [tokenizes and parses the source code](https://github.com/golang/go/tree/1e963f82914ddfb47820034c5c85205a362ed73f/src/cmd/compile/internal/syntax) . It builds a syntax tree\n  for each source file.\n- Type checking\n- The compiler [maps each identifier to the object it denotes](https://github.com/golang/go/tree/1e963f82914ddfb47820034c5c85205a362ed73f/src/cmd/compile/internal/types2) , folds\n  constants, and infers the type of every expression.\n- IR construction\n- The compiler converts the syntax tree and its types into its own [intermediate\n  representation](https://github.com/golang/go/tree/1e963f82914ddfb47820034c5c85205a362ed73f/src/cmd/compile/internal/ir) (IR). This process, called “[noding](https://github.com/golang/go/tree/1e963f82914ddfb47820034c5c85205a362ed73f/src/cmd/compile/internal/noder) ,” goes through\n  a serialization format named unified IR.\n- Middle end\n- The compiler performs several optimization passes on the IR,\n  such as [devirtualization](https://github.com/golang/go/tree/1e963f82914ddfb47820034c5c85205a362ed73f/src/cmd/compile/internal/devirtualize) ,[function call inlining](https://github.com/golang/go/tree/1e963f82914ddfb47820034c5c85205a362ed73f/src/cmd/compile/internal/inline) ,\n  and[escape analysis](https://github.com/golang/go/tree/1e963f82914ddfb47820034c5c85205a362ed73f/src/cmd/compile/internal/escape) .\n- Walk\n- This phase runs [two steps](https://github.com/golang/go/tree/1e963f82914ddfb47820034c5c85205a362ed73f/src/cmd/compile/internal/walk) : order of evaluation decomposes complex\n  statements into simpler ones, and desugaring transforms higher-level Go\n  constructs, like`switch` or channels, into more primitive instructions or\n  calls to the runtime.\n- Generic SSA\n- The compiler [converts the IR](https://github.com/golang/go/tree/1e963f82914ddfb47820034c5c85205a362ed73f/src/cmd/compile/internal/ssagen) into Static Single Assignment (SSA)\n  form, a lower-level intermediate representation suited for[machine-independent optimizations and rewrite rules](https://github.com/golang/go/tree/1e963f82914ddfb47820034c5c85205a362ed73f/src/cmd/compile/internal/ssa) .\n- Machine code generation\n- The compiler rewrites the SSA form into [machine-specific variants](https://github.com/golang/go/tree/1e963f82914ddfb47820034c5c85205a362ed73f/src/cmd/compile/internal/ssa/_gen) ,\n  allocates registers, and applies more optimization passes. At the end,[the\n  assembler](https://github.com/golang/go/tree/1e963f82914ddfb47820034c5c85205a362ed73f/src/cmd/internal/obj) turns the generated instructions into machine code.\n\n## The hammer[#](#the-hammer)\n\nMy first idea is to replace occurrences of `netip.AddrFrom16(ip.As16())` with\n`netip.Addr{addr: ip.addr, z: netip.z6noz}` as early as possible, during the\n“noding” process. Before that, the type checking phase prevents\nus from accessing unexported struct fields.\n\nGo 1.27 introduced a convenient debug option to dump the IR of a function at interesting points during compilation:\n\n```\n$ GOTOOLCHAIN=go1.27.1 GOAMD64=v3 go build -a -gcflags=\"-d=astdump=AddrTo6Safe\" .\nWriting text ast output for AddrTo6Safe to AddrTo6Safe.ast\nWriting html ast output for AddrTo6Safe to AddrTo6Safe.html\nWriting html syntax output for AddrTo6Safe to AddrTo6Safe.syntax.html\n```\nIn the HTML file, the first column shows the IR as it comes out of noding:\n\n```\nDCLFUNC addrto6.AddrTo6Safe ABI:ABIInternal FUNC-func(netip.Addr) netip.Addr\nDCLFUNC-Dcl\n. NAME-addrto6.ip Class:PPARAM Offset:0 OnStack Used netip.Addr\n. NAME-addrto6.~r0 Class:PPARAMOUT Offset:0 OnStack netip.Addr\nDCLFUNC-body\n. IF # ipv6_safe.go:11:2\n. IF-Cond\n. . CALLFUNC bool\n. . CALLFUNC-Fun\n. . . METHEXPR addrto6.Is4 FUNC-func(netip.Addr) bool\n. . . . TYPE netip.Addr Class:PEXTERN Offset:0 type netip.Addr\n. . CALLFUNC-Args\n. . . NAME-addrto6.ip Class:PPARAM Offset:0 OnStack Used netip.Addr\n. IF-Body\n. . AS # ipv6_safe.go:12:6\n. . . NAME-addrto6.ip Class:PPARAM Offset:0 OnStack Used netip.Addr\n. . . CALLFUNC netip.Addr\n. . . CALLFUNC-Fun\n. . . . NAME-netip.AddrFrom16 Class:PFUNC Offset:0 Used FUNC-func([16]byte) netip.Addr\n. . . CALLFUNC-Args\n. . . . CALLFUNC ARRAY-[16]byte\n. . . . CALLFUNC-Fun\n. . . . . METHEXPR addrto6.As16 FUNC-func(netip.Addr) [16]byte\n. . . . . . TYPE netip.Addr Class:PEXTERN Offset:0 type netip.Addr\n. . . . CALLFUNC-Args\n. . . . . NAME-addrto6.ip Class:PPARAM Offset:0 OnStack Used netip.Addr\n. RETURN # ipv6_safe.go:14:2\n. RETURN-Results\n. . NAME-addrto6.ip Class:PPARAM Offset:0 OnStack Used netip.Addr\n```\nIn the body of the `if` statement, we spot the calls to the method\n`netip.Addr.As16()` and to the function `netip.AddrFrom16()`. Our goal is to\npatch them with a struct literal:\n\n```\nIF-Body\n. AS # ipv6_safe.go:12:6\n. . NAME-addrto6.ip Class:PPARAM Offset:0 OnStack Used netip.Addr\n. . STRUCTLIT netip.Addr\n. . STRUCTLIT-List\n. . . STRUCTKEY netip.addr\n. . . . DOT netip.addr netip.uint128\n. . . . . NAME-addrto6.ip Class:PPARAM Offset:0 OnStack Used netip.Addr\n. . . STRUCTKEY netip.z\n. . . . NAME-netip.z6noz Class:PEXTERN Offset:0 unique.Handle[net/netip.addrDetail]\n```\nIn [noder’s `reader.go`](https://github.com/golang/go/blob/1e963f82914ddfb47820034c5c85205a362ed73f/src/cmd/compile/internal/noder/reader.go), the `expr()` method builds the IR tree for\nan expression. At the end of the `exprCall` case, we add a call to a\n`rewriteAddrFrom16As16()` function. It takes the current node and returns the\nstruct literal on success, or `nil` if the rewrite is not possible. First, we\ncheck that we have the expected pattern: a call to the `netip.AddrFrom16()`\nfunction with a call to the `netip.Addr.As16()` method as its only argument:\n\n```\nfunc rewriteAddrFrom16As16(n ir.Node) ir.Node {\n    call, ok := n.(*ir.CallExpr)\n    if !ok || call.Op() != ir.OCALLFUNC ||\n        len(call.Args) != 1 || len(call.Init()) != 0 ||\n        !isNetipFunc(call.Fun, \"AddrFrom16\") {\n        return nil\n    }\n    inner, ok := call.Args[0].(*ir.CallExpr)\n    if !ok || inner.Op() != ir.OCALLFUNC ||\n        len(inner.Args) != 1 || len(inner.Init()) != 0 ||\n        !isNetipFunc(inner.Fun, \"Addr.As16\") {\n        return nil\n    }\n    x := inner.Args[0]\n    // [...]\n}\n```\nThen, we fetch `netip.z6noz`:\n\n```\nz6noz, err := lookupVar(ir.StaticCalleeName(call.Fun).Sym().Pkg, \"z6noz\")\nif err != nil {\n    return nil\n}\n```\nAnd we build the struct literal:\n\n```\ntyp := call.Type()\npos := call.Pos()\nvar list []ir.Node\nfor i, f := range typ.Fields() {\n    var value ir.Node\n    switch f.Sym.Name {\n    case \"addr\":\n        value = typecheck.DotField(pos, x, i)\n    case \"z\":\n        value = z6noz\n    default:\n        return nil\n    }\n    list = append(list, ir.NewStructKeyExpr(pos, f, value))\n}\nlit := ir.NewCompLitExpr(pos, ir.OSTRUCTLIT, typ, list)\nlit.SetTypecheck(1)\nreturn lit\n```\nHave a look at the [complete patch](https://github.com/vincentbernat/go/commit/netipmap-fix2).<sup>[7](#sidenote-statement)</sup> We can test it\nwith the following commands:\n\n```\n$ cd src\n$ ./make.bash\nBuilding Go cmd/dist using /usr/lib/go-1.27. (go1.27.1 linux/amd64)\nBuilding Go toolchain1 and bootstrap cmd/go (go_bootstrap) using /usr/lib/go-1.27.\nBuilding Go toolchain2 using go_bootstrap and Go toolchain1.\nBuilding Go toolchain3 and commands using go_bootstrap and Go toolchain2.\nChecking command staleness for linux/amd64.\n---\nInstalled Go for linux/amd64 in /home/bernat/code/free/go\nInstalled commands in /home/bernat/code/free/go/bin\n*** You need to add /home/bernat/code/free/go/bin to your PATH.\n$ export PATH=$PWD/../bin:$PATH\n$ go version\ngo version go1.28-devel_9834516e20 Sat Sep 12 08:23:11 2026 -0700 linux/amd64\n$ go test net/netip/...\nok      net/netip   0.224s\n$ cd ../../go-netip-addrto6\n$ go test .\nok      github.com/vincentbernat/go-netip-addrto6   0.062s\n```\nThe generated code for the helper is now the shortest possible version!\n\n```\n// AX = input.addr.hi, BX = input.addr.lo, CX = input.z\n CMPQ    net/netip·z4(SB), CX     ; check \"z\" if this is an IPv4 address\n JNE     end                      ; if not, stop here\n MOVQ    net/netip·z6noz(SB), CX  ; CX = netip.z6noz\nend:\n RET\n// return value = Addr{hi: AX, lo: BX, z: CX}\n```\nGo maintainers are unlikely to accept this change. It relies on the internal\nstructure of `net/netip.Addr`. It’s an ugly hack in the noder, whose job is to\nfaithfully translate the type-checked AST into the IR. And it’s harder to\nmaintain than adding a `To6()` method.\n\n## The screwdriver[#](#the-screwdriver)\n\nThe right place for such an optimization is the [generic SSA phase](https://github.com/golang/go/blob/1e963f82914ddfb47820034c5c85205a362ed73f/src/cmd/compile/internal/ssa/README.md). One of\nthe last machine-independent passes is `memcombine`. With the appropriate debug\nflag, the compiler dumps the SSA form after this pass:[8](#sidenote-html)\n\n```\n$ GOTOOLCHAIN=go1.26.8 GOAMD64=v3 \\\n> go build -a -gcflags='-d=ssa/memcombine/dump=AddrTo6Safe' .\n$ head -5 AddrTo6Safe_01__memcombine.dump\nAddrTo6Safe func(netip.Addr) netip.Addr\n  b2:\n    (?) v1 = InitMem <mem>\n    (?) v2 = SP <uintptr>\n    (?) v3 = SB <uintptr>\n```\nThe result of the `memcombine` pass follows the same structure as the assembly\ncode for `AddrTo6Safe()` [we looked at earlier](#asm-addrto6safe): two stores,\none move, and two loads we would like to optimize away.\n\n```\n; […]\n  v502 = ArgIntReg <uint64> {ip+0} [0]              ; input.addr.hi\n  v490 = ArgIntReg <uint64> {ip+8} [1]              ; input.addr.lo\n  v466 = ArgIntReg <*netip.addrDetail> {ip+16} [2]  ; input.z\n; […]\n  v22 = LocalAddr <*[16]byte> {netip.a16} v2 v1     ; &a16\n  v173 = OffPtr <*byte> [8] v22                     ; &a16[8]\n  v442 = Bswap64 <uint64> v490                      ; bswap(input.addr.lo)\n  v542 = Bswap64 <uint64> v502                      ; bswap(input.addr.hi)\n  v161 = Store <mem> {uint64} v22 v542 v23          ; a16[:8] = bswap(hi)\n  v282 = Store <mem> {uint64} v173 v442 v161        ; a16[8:] = bswap(lo)\n  v285 = LocalAddr <*[16]byte> {netip.addr} v2 v282 ; &addr\n  v286 = Move <mem> {[16]byte} [16] v285 v22 v282   ; addr = a16\n  v415 = OffPtr <*byte> [8] v285                    ; &addr[8]\n  v416 = Load <uint64> v285 v286                    ; addr[:8]\n  v299 = Bswap64 <uint64> v416                      ; output.addr.hi\n  v174 = Load <uint64> v415 v286                    ; addr[8:]\n  v39  = Bswap64 <uint64> v174                      ; output.addr.lo\n; […]\n```\nEach line features a value identifier (`v442`), an operation with its type\n(`Bswap64 <uint64>`), and its arguments (`v490`).<sup>[9](#sidenote-lines)</sup> Values are the basic\nbuilding blocks of SSA and are defined exactly once. Square brackets enclose\ninteger parameters (`[8]`) and curly braces contain auxiliary arguments\n(`{netip.addr}`). Operations writing to memory produce a new memory state. Every\nmemory operation takes the current state as its last argument, which keeps them\nin order.\n\n### On paper[#](#on-paper)\n\nLet’s focus on `output.addr.lo`, aka `v39`:\n\n```\n  v490 = ArgIntReg <uint64> {ip+8} [1]              ; input.addr.lo\n  v22 = LocalAddr <*[16]byte> {netip.a16} v2 v1     ; &a16\n  v173 = OffPtr <*byte> [8] v22                     ; &a16[8]\n  v442 = Bswap64 <uint64> v490                      ; bswap(input.addr.lo)\n  v282 = Store <mem> {uint64} v173 v442 v161        ; a16[8:] = bswap(lo)\n  v285 = LocalAddr <*[16]byte> {netip.addr} v2 v282 ; &addr\n  v286 = Move <mem> {[16]byte} [16] v285 v22 v282   ; addr = a16\n  v415 = OffPtr <*byte> [8] v285                    ; &addr[8]\n  v174 = Load <uint64> v415 v286                    ; addr[8:]\n  v39  = Bswap64 <uint64> v174                      ; output.addr.lo\n```\nTo simplify this code, we could apply three rewriting rules:\n\n1. \nThe first one adds a shortcut when **loading through a move** :```\n(Load (OffPtr\n   [o] p) (Move p src mem)) => (Load (OffPtr [o] src) mem)\n```\n. This matches`v174` with its arguments`v415` and`v286` and creates a new value`v600` :```\n  v490 = ArgIntReg <uint64> {ip+8} [1]              ; input.addr.lo\n  v22 = LocalAddr <*[16]byte> {netip.a16} v2 v1     ; &a16\n  v173 = OffPtr <*byte> [8] v22                     ; &a16[8]\n  v442 = Bswap64 <uint64> v490                      ; bswap(input.addr.lo)\n  v282 = Store <mem> {uint64} v173 v442 v161        ; a16[8:] = bswap(lo)\n  v285 = LocalAddr <*[16]byte> {netip.addr} v2 v282 ; &addr\n  v286 = Move <mem> {[16]byte} [16] v285 v22 v282   ; addr = a16\n  v415 = OffPtr <*byte> [8] v285                    ; &addr[8]\n  v600 = OffPtr <*byte> [8] v22                     ; &a16[8]\n  v174 = Load <uint64> v600 v282                    ; a16[8:]\n  v39  = Bswap64 <uint64> v174                      ; output.addr.lo\n```\n2. \nThe second one simplifies a **load following a store** :```\n(Load p (Store p x\n   _)) => x\n```\n. The load is*forwarded* : the stored value replaces it and no memory\n   access remains. It matches`v174` . It notices that`v600` and`v173` are the\n   same address and replaces`v174` with a copy of`v442` :```\n  v490 = ArgIntReg <uint64> {ip+8} [1]              ; input.addr.lo\n  v22 = LocalAddr <*[16]byte> {netip.a16} v2 v1     ; &a16\n  v173 = OffPtr <*byte> [8] v22                     ; &a16[8]\n  v442 = Bswap64 <uint64> v490                      ; bswap(input.addr.lo)\n  v282 = Store <mem> {uint64} v173 v442 v161        ; a16[8:] = bswap(lo)\n  v285 = LocalAddr <*[16]byte> {netip.addr} v2 v282 ; &addr\n  v286 = Move <mem> {[16]byte} [16] v285 v22 v282   ; addr = a16\n  v415 = OffPtr <*byte> [8] v285                    ; &addr[8]\n  v600 = OffPtr <*byte> [8] v22                     ; &a16[8]\n  v174 = Copy <uint64> v442                         ; bswap(input.addr.lo)\n  v39  = Bswap64 <uint64> v174                      ; output.addr.lo\n```\n3. \nThe last step **cancels the two byte swaps** :`(Bswap64 (Bswap64 x)) => x` .`v39` becomes a copy of`v490` :```\n  v490 = ArgIntReg <uint64> {ip+8} [1]              ; input.addr.lo\n  v22 = LocalAddr <*[16]byte> {netip.a16} v2 v1     ; &a16\n  v173 = OffPtr <*byte> [8] v22                     ; &a16[8]\n  v442 = Bswap64 <uint64> v490                      ; bswap(input.addr.lo)\n  v282 = Store <mem> {uint64} v173 v442 v161        ; a16[8:] = bswap(lo)\n  v285 = LocalAddr <*[16]byte> {netip.addr} v2 v282 ; &addr\n  v286 = Move <mem> {[16]byte} [16] v285 v22 v282   ; addr = a16\n  v415 = OffPtr <*byte> [8] v285                    ; &addr[8]\n  v600 = OffPtr <*byte> [8] v22                     ; &a16[8]\n  v174 = Copy <uint64> v442                         ; bswap(input.addr.lo)\n  v39  = Copy <uint64> v490                         ; output.addr.lo = input.addr.lo\n```\n\nIf we ignore the values not needed to compute `v39`, only this SSA form remains:\n\n```\n  v490 = ArgIntReg <uint64> {ip+8} [1]  ; input.addr.lo\n  v39  = Copy <uint64> v490             ; output.addr.lo = input.addr.lo\n```\nLet’s switch to `output.addr.hi`, aka `v299`:\n\n```\n  v502 = ArgIntReg <uint64> {ip+0} [0]              ; input.addr.hi\n  v22 = LocalAddr <*[16]byte> {netip.a16} v2 v1     ; &a16\n  v173 = OffPtr <*byte> [8] v22                     ; &a16[8]\n  v542 = Bswap64 <uint64> v502                      ; bswap(input.addr.hi)\n  v161 = Store <mem> {uint64} v22 v542 v23          ; a16[:8] = bswap(hi)\n  v282 = Store <mem> {uint64} v173 v442 v161        ; a16[8:] = bswap(lo)\n  v285 = LocalAddr <*[16]byte> {netip.addr} v2 v282 ; &addr\n  v286 = Move <mem> {[16]byte} [16] v285 v22 v282   ; addr = a16\n  v416 = Load <uint64> v285 v286                    ; addr[:8]\n  v299 = Bswap64 <uint64> v416                      ; output.addr.hi\n```\nTo optimize it away, we also apply three rewriting rules:\n\n1. \nThe first one also adds a shortcut when **loading through a move** , but\n   without an offset:`(Load p (Move p src mem)) => (Load src mem)` . This rewrites`v416` to use arguments from`v286` :```\n  v502 = ArgIntReg <uint64> {ip+0} [0]              ; input.addr.hi\n  v22 = LocalAddr <*[16]byte> {netip.a16} v2 v1     ; &a16\n  v173 = OffPtr <*byte> [8] v22                     ; &a16[8]\n  v542 = Bswap64 <uint64> v502                      ; bswap(input.addr.hi)\n  v161 = Store <mem> {uint64} v22 v542 v23          ; a16[:8] = bswap(hi)\n  v282 = Store <mem> {uint64} v173 v442 v161        ; a16[8:] = bswap(lo)\n  v285 = LocalAddr <*[16]byte> {netip.addr} v2 v282 ; &addr\n  v286 = Move <mem> {[16]byte} [16] v285 v22 v282   ; addr = a16\n  v416 = Load <uint64> v22 v282                     ; a16[:8]\n  v299 = Bswap64 <uint64> v416                      ; output.addr.hi\n```\n2. \nThe second rule **forwards a value stored one step earlier** , skipping over a\n   store to another address:`(Load p (Store q _ (Store p x _))) => x` . This\n   matches`v416` :`x` is`v542` ,`p` is`v22` (`&a16` ),`q` is`v173` (`&a16[8]` ), and`p` and`q` do not overlap for`uint64` .```\n  v502 = ArgIntReg <uint64> {ip+0} [0]              ; input.addr.hi\n  v22 = LocalAddr <*[16]byte> {netip.a16} v2 v1     ; &a16\n  v173 = OffPtr <*byte> [8] v22                     ; &a16[8]\n  v542 = Bswap64 <uint64> v502                      ; bswap(input.addr.hi)\n  v161 = Store <mem> {uint64} v22 v542 v23          ; a16[:8] = bswap(hi)\n  v282 = Store <mem> {uint64} v173 v442 v161        ; a16[8:] = bswap(lo)\n  v285 = LocalAddr <*[16]byte> {netip.addr} v2 v282 ; &addr\n  v286 = Move <mem> {[16]byte} [16] v285 v22 v282   ; addr = a16\n  v416 = Copy <uint64> v542                         ; bswap(input.addr.hi)\n  v299 = Bswap64 <uint64> v416                      ; output.addr.hi\n```\n3. \nThe third rule **cancels two byte swaps** :`(Bswap64 (Bswap64 x)) => x` .`v299` becomes a copy of`v502` :```\n  v502 = ArgIntReg <uint64> {ip+0} [0]              ; input.addr.hi\n  v22 = LocalAddr <*[16]byte> {netip.a16} v2 v1     ; &a16\n  v173 = OffPtr <*byte> [8] v22                     ; &a16[8]\n  v542 = Bswap64 <uint64> v502                      ; bswap(input.addr.hi)\n  v161 = Store <mem> {uint64} v22 v542 v23          ; a16[:8] = bswap(hi)\n  v282 = Store <mem> {uint64} v173 v442 v161        ; a16[8:] = bswap(lo)\n  v285 = LocalAddr <*[16]byte> {netip.addr} v2 v282 ; &addr\n  v286 = Move <mem> {[16]byte} [16] v285 v22 v282   ; addr = a16\n  v416 = Copy <uint64> v542                         ; bswap(input.addr.hi)\n  v299 = Copy <uint64> v502                         ; output.addr.hi = input.addr.hi\n```\n\nIf we remove the values not used to compute `v299`, we get this SSA form:\n\n```\n  v502 = ArgIntReg <uint64> {ip+0} [0] ; input.addr.hi\n  v299 = Copy <uint64> v502            ; output.addr.hi = input.addr.hi\n```\n### In practice[#](#in-practice)\n\nMost of these rules already exist in [`generic.rules`](https://github.com/golang/go/blob/1e963f82914ddfb47820034c5c85205a362ed73f/src/cmd/compile/internal/ssa/_gen/generic.rules). They use\nconditions to validate their context: `ssa.IsSamePtr()` for the same address,\n`ssa.Disjoint()` for addresses that do not overlap. The rule forwarding a stored\nvalue to a load already exists with three variants looking through several other\nstores. Here are the two we need:\n\n```\n(Load <t1> p1 (Store {t2} p2 x _))\n    && ssa.IsSamePtr(p1, p2)\n    && copyCompatibleType(t1, x.Type)\n    && t1.Size() == t2.Size()\n    => x\n(Load <t1> p1 (Store {t2} p2 _ (Store {t3} p3 x _)))\n    && ssa.IsSamePtr(p1, p3)\n    && copyCompatibleType(t1, x.Type)\n    && t1.Size() == t3.Size()\n    && ssa.Disjoint(p3, t3, p2, t2)\n    => x\n```\nGo 1.27 added the rule loading through a move with [CL 748200](https://go-review.googlesource.com/c/go/+/748200) to fix\n[issue #77720](https://github.com/golang/go/issues/77720):\n\n```\n(Load <t1> op1:(OffPtr [o1] p1) move:(Move [n] p2 src mem))\n    && o1 >= 0 && o1+t1.Size() <= n && ssa.IsSamePtr(p1, p2)\n    && !ssa.IsVolatile(src)\n    => @move.Block (Load <t1> (OffPtr <op1.Type> [o1] src) mem)\n```\nIt lacks a variant without an offset:\n\n```\n(Load <t1> p1 move:(Move [n] p2 src mem))\n    && p1.Op != ssaop.OpOffPtr\n    && t1.Size() <= n && ssa.IsSamePtr(p1, p2)\n    && !ssa.IsVolatile(src)\n    => @move.Block (Load <t1> (OffPtr <p1.Type> [0] src) mem)\n```\nThere is no generic rule to cancel two byte swaps, but the [AMD64 lowering\npass](https://github.com/golang/go/blob/1e963f82914ddfb47820034c5c85205a362ed73f/src/cmd/compile/internal/ssa/_gen/AMD64.rules) includes this rule:\n\n```\n(BSWAP(Q|L) (BSWAP(Q|L) p)) => p\n```\nAfter switching to Go’s development branch and [adding the missing\nrule](https://github.com/vincentbernat/go/commit/netipmap-fix3), the generated assembly code is worse than with Go 1.26.8,\neven though our additional rule slightly improves the situation at the end:\n\n```\n// AX = input.addr.hi, BX = input.addr.lo, CX = input.z\n; Push the stack (16 bytes):\n;    0(SP) a16 [16]byte\n PUSHQ   BP\n MOVQ    SP, BP\n SUBQ    $16, SP\n CMPQ    net/netip·z4(SB), CX       ; check \"z\" if this is an IPv4 address\n JNE     end                        ; if not, stop here\n; The four forwarded bytes: the low half of input.addr.lo is taken apart and\n; put back together in registers\n MOVQ    BX, DX                     ; DX = input.addr.lo\n SHRQ    $24, BX                    ; BX = input.addr.lo >> 24\n MOVQ    DX, SI                     ; SI = input.addr.lo\n SHRQ    $16, DX                    ; DX = input.addr.lo >> 16\n MOVQ    SI, DI                     ; DI = input.addr.lo, kept for the pack\n SHRQ    $8, SI                     ; SI = input.addr.lo >> 8\n MOVBLZX DIB, R8                    ; R8 = byte(input.addr.lo)\n MOVBLZX SIB, SI                    ; SI = byte(input.addr.lo >> 8)\n SHLQ    $8, SI\n ORQ     R8, SI                     ; SI = two low bytes of input.addr.lo\n MOVBLZX DL, DX                     ; DX = byte(input.addr.lo >> 16)\n SHLQ    $16, DX\n ORQ     SI, DX\n MOVBLZX BL, BX                     ; BX = byte(input.addr.lo >> 24)\n SHLQ    $24, BX\n ORQ     DX, BX                     ; BX = input.addr.lo & 0xffffffff\n; Pack: byteorder.BEPutUint64(a16[:8], input.addr.hi)\n;       byteorder.BEPutUint64(a16[8:], input.addr.lo)\n MOVBEQ  AX, net/netip·a16(SP)\n MOVBEQ  DI, net/netip·a16+8(SP)\n; The four other bytes of input.addr.lo, read one by one from a16\n MOVBLZX net/netip·a16+11(SP), DX   ; a16[11]\n SHLQ    $32, DX\n ORQ     DX, BX\n MOVBLZX net/netip·a16+10(SP), DX   ; a16[10]\n SHLQ    $40, DX\n ORQ     DX, BX\n MOVBLZX net/netip·a16+9(SP), DX    ; a16[9]\n SHLQ    $48, DX\n ORQ     DX, BX\n MOVBLZX net/netip·a16+8(SP), DX    ; a16[8]\n SHLQ    $56, DX\n; output.z = netip.z6noz\n MOVQ    net/netip·z6noz(SB), CX\n; output.addr.hi = byteorder.BEUint64(a16[:8])\n MOVBEQ  net/netip·a16(SP), AX\n; output.addr.lo assembled from the previous steps\n ORQ     DX, BX\nend:\n LEAVEQ\n RET\n// return value = Addr{hi: AX, lo: BX, z: CX}\n```\nThe rule loading through a move, added in Go 1.27, introduced this regression.\n\n### Out of order[#](#out-of-order)\n\nLet’s not give up now! In reality, the rewriting rules run before `memcombine`,\nnotably in the `late opt` pass. At this point, the inlined versions of\n`BEPutUint64()` and `BEUint64()` still expand to sixteen byte stores and sixteen\nbyte loads, matching their [source code](https://pkg.go.dev/internal/byteorder#BEUint64):\n\n```\nfunc BEUint64(b []byte) uint64 {\n    _ = b[7] // bounds check hint to compiler; see golang.org/issue/14808\n    return uint64(b[7]) | uint64(b[6])<<8 | uint64(b[5])<<16 | uint64(b[4])<<24 |\n        uint64(b[3])<<32 | uint64(b[2])<<40 | uint64(b[1])<<48 | uint64(b[0])<<56\n}\n```\nLet’s follow two bytes of `output.addr.lo`: `addr[15]` and `addr[11]`. Here is\na simplified SSA form before `late opt`:\n\n```\n  v273 = Trunc64to8 <byte> v490                     ; byte(input.addr.lo)\n  v226 = Trunc64to8 <byte> v225                     ; byte(input.addr.lo >> 32)\n; […]\n  v235 = Store <mem> {byte} v233 v226 v223          ; a16[11] = byte(lo >> 32)\n  v247 = Store <mem> {byte} v245 v238 v235          ; a16[12] = …\n  v259 = Store <mem> {byte} v257 v250 v247          ; a16[13] = …\n  v271 = Store <mem> {byte} v269 v262 v259          ; a16[14] = …\n  v282 = Store <mem> {byte} v280 v273 v271          ; a16[15] = byte(lo)\n  v285 = LocalAddr <*[16]byte> {netip.addr} v2 v282 ; &addr\n  v286 = Move <mem> {[16]byte} [16] v285 v22 v282   ; addr = a16\n; […]\n  v433 = OffPtr <*byte> [15] v285                   ; &addr[15]\n  v435 = Load <byte> v433 v286                      ; addr[15]\n  v479 = OffPtr <*byte> [11] v285                   ; &addr[11]\n  v481 = Load <byte> v479 v286                      ; addr[11]\n```\nThe first rule loads through the move: ```\n(Load (OffPtr [o] p) (Move p src mem))\n=> (Load (OffPtr [o] src) mem)\n```\n. It matches both loads, which now read `a16`\nwith the memory state before the copy:\n\n```\n  v600 = OffPtr <*byte> [15] v22 ; &a16[15]\n  v435 = Load <byte> v600 v282   ; a16[15]\n  v601 = OffPtr <*byte> [11] v22 ; &a16[11]\n  v481 = Load <byte> v601 v282   ; a16[11]\n```\nThe second rule shortcuts a load following a store: ```\n(Load p (Store p x _)) =>\nx\n```\n. It matches `v435`, as `v282` stores `a16[15]`. It does not match `v481`:\n`v235` stores `a16[11]` four stores earlier in the chain, while the variants of\nthis rule look through three stores at most.\n\n```\n  v435 = Copy <byte> v273        ; byte(input.addr.lo)\n  v601 = OffPtr <*byte> [11] v22 ; &a16[11]\n  v481 = Load <byte> v601 v282   ; a16[11]\n```\nThe same happens to the other bytes: the rule forwards the four bytes stored\nlast, `a16[12]` to `a16[15]`. The twelve other loads now read `a16` instead of\n`addr`.\n\n`BEUint64()` becomes a chain of `Or64`, each one adding a byte shifted into\nplace. `memcombine` is a [pass written in Go](https://github.com/golang/go/blob/1e963f82914ddfb47820034c5c85205a362ed73f/src/cmd/compile/internal/ssacompile/memcombine.go), not a set of rewrite\nrules. It starts from the last `Or64` of the chain and collects up to eight\nterms. If each term is a byte load, extended to 64 bits and shifted, and if the\neight loads read consecutive addresses from the same pointer with the same\nmemory state, it replaces the whole chain with a single 64-bit load and a byte\nswap. Otherwise, it tries again with four, then two terms, and from each\nintermediate `Or64`. Here is the loop checking each term in a simplified version\nof `combineLoads()`:\n\n```\nfor i := int64(0); i < n; i++ {\n    v := a[i]\n    shift := int64(0)\n    if v.Op == shiftOp {\n        v, shift = peelShift(v)\n    }\n    if v.Op != extOp {\n        return false\n    }\n    load := v.Args[0]\n    if load.Op != ssaop.OpLoad {\n        return false\n    }\n    if load.Args[1] != mem {\n        return false\n    }\n    p, off := splitPtr(load.Args[0])\n    if p != base {\n        return false\n    }\n    r[i] = LoadRecord{load: load, offset: off, shift: shift}\n}\n```\nFor `output.addr.hi`, the eight loads read `a16` with the same memory state\n`v282`:\n\n```\n  v13  = Load <byte> v22 v282               ; a16[0]\n  v530 = Load <byte> v14 v282               ; a16[1]\n  v488 = Load <byte> v504 v282              ; a16[2]\n  v405 = Load <byte> v537 v282              ; a16[3]\n  v385 = Load <byte> v397 v282              ; a16[4]\n  v361 = Load <byte> v373 v282              ; a16[5]\n  v196 = Load <byte> v63 v282               ; a16[6]\n  v432 = Load <byte> v315 v282              ; a16[7]\n  v319 = ZeroExt8to64 <uint64> v432         ; uint64(a16[7])\n  v329 = ZeroExt8to64 <uint64> v196         ; uint64(a16[6])\n  v330 = Lsh64x64 <uint64> [true] v329 v138 ; uint64(a16[6]) << 8\n  v331 = Or64 <uint64> v319 v330            ; a16[7] | a16[6] << 8\n; […] same for a16[5] to a16[1]\n  v401 = ZeroExt8to64 <uint64> v13          ; uint64(a16[0])\n  v402 = Lsh64x64 <uint64> [true] v401 v55  ; uint64(a16[0]) << 56\n  v403 = Or64 <uint64> v402 v391            ; | a16[0] << 56 = output.addr.hi\n```\n`memcombine` merges them into one load and a swap:\n\n```\n  v286 = Load <uint64> v22 v282   ; a16[:8]\n  v285 = Bswap64 <uint64> v286    ; output.addr.hi\n```\nFor `output.addr.lo`, here is the chain `memcombine` sees after `late opt`:\n\n```\n  v436 = ZeroExt8to64 <uint64> v273 ; addr[15], forwarded\n  v448 = Or64 <uint64> v436 v447    ; | addr[14] << 8, forwarded\n  v460 = Or64 <uint64> v459 v448    ; | addr[13] << 16, forwarded\n  v472 = Or64 <uint64> v471 v460    ; | addr[12] << 24, forwarded\n  v484 = Or64 <uint64> v483 v472    ; | a16[11] << 32, loaded\n  v496 = Or64 <uint64> v495 v484    ; | a16[10] << 40, loaded\n  v508 = Or64 <uint64> v507 v496    ; | a16[9] << 48, loaded\n  v520 = Or64 <uint64> v519 v508    ; | a16[8] << 56, loaded\n```\nFrom `v520`, four of the eight terms are forwarded bytes, not loads from memory,\nand `memcombine` can’t combine them. It doesn’t merge the four remaining loads\neither, as they sit on top of the forwarded bytes.\n\n### Back in order[#](#back-in-order)\n\nIn summary, the rewriting rules run too early to be effective. A quick\nworkaround exists: run an earlier round of `memcombine` before `late opt`. After\n[this change](https://github.com/vincentbernat/go/commit/netipmap-fix4), the generated code for the helper is back to the\nshortest possible version:\n\n```\n// AX = input.addr.hi, BX = input.addr.lo, CX = input.z\n CMPQ    net/netip·z4(SB), CX     ; check \"z\" if this is an IPv4 address\n JNE     end                      ; if not, stop here\n MOVQ    net/netip·z6noz(SB), CX  ; CX = netip.z6noz\nend:\n RET\n// return value = Addr{hi: AX, lo: BX, z: CX}\n```\nAnd the benchmark confirms it! ✌️\n\n```\ngoos: linux\ngoarch: amd64\npkg: github.com/vincentbernat/go-netip-addrto6\ncpu: AMD Ryzen 5 5600X 6-Core Processor\n                   │   Go 1.26.8    │             Our branch              │\n                   │     sec/op     │    sec/op     vs base               │\nAddrTo6/safe          6.5470n ±  0%   0.8944n ± 4%  -86.34% (p=0.002 n=6)\nAddrTo6/unsafe        0.9071n ±  3%   0.8682n ± 1%   -4.28% (p=0.002 n=6)\nAddrTo6/builtin       0.8871n ±  2%   0.8785n ± 1%        ~ (p=0.310 n=6)\n```\n# Next steps[#](#next-steps)\n\nI think Go maintainers would reject this change because of the additional\n`memcombine` pass. Instead, I plan to publish this blog post and bring up the\nsubject again as a follow-up to [issue #54365](https://github.com/golang/go/issues/54365). Either the sheer complexity and\nthe Go 1.27 regression convince the maintainers that adding a `To6()` method is\nsimpler and more efficient, or they advise me on how to move forward. Either\nway, digging into this subject taught me a lot about the Go compiler! ⚙️\n\nUpdate (2026-10)\n\nI opened [issue #81994](https://github.com/golang/go/issues/81994) to propose `Addr.To6()`. Give it a 👍 if you want it in Go!\n","body_html":"<h1 id=\"hacking-the-go-compiler-to-efficiently-map-ipv4-to-ipv6\">Hacking the Go compiler to efficiently map IPv4 to IPv6</h1>\n<p><a href=\"https://pkg.go.dev/netip.addr\" rel=\"nofollow ugc noopener\"><code>netip.Addr</code></a> features an <a href=\"https://pkg.go.dev/netip.addr.Unmap\" rel=\"nofollow ugc noopener\"><code>Unmap()</code></a>\nmethod returning the unwrapped IPv4 contained in an <a href=\"https://www.rfc-editor.org/rfc/rfc4291#section-2.5.5.2\" rel=\"nofollow ugc noopener\">IPv4-mapped IPv6\naddress</a>: from <code>::ffff:203.0.113.10</code> or <code>::ffff:cb00:710a</code>, it returns\n<code>203.0.113.10</code>.&lt;sup&gt;<a href=\"#sidenote-why\">1</a>&lt;/sup&gt; There is no <code>Map()</code> or <code>To6()</code> method for the reverse\ndirection. Such a method is trivial to implement, but Go maintainers have\n<a href=\"https://github.com/golang/go/issues/54365#issuecomment-1570607144\" rel=\"nofollow ugc noopener\">rejected</a> it on the grounds that users should write\n<code>netip.AddrFrom16(ip.As16())</code> and let the compiler optimize it.&lt;sup&gt;<a href=\"#sidenote-equiv\">2</a>&lt;/sup&gt; Today,\nthis pattern is eight times slower than a native method. How can we teach the\ncompiler to optimize this sequence?</p>\n<h1 id=\"the-alternatives\">The alternatives<a href=\"#the-alternatives\">#</a></h1>\n<p>Let’s explore three ways to implement the map semantics for <code>netip.Addr</code>. My\nfavorite is to add it to the Go standard library. Go maintainers prefer a small\nexternal helper chaining <code>netip.AddrFrom16()</code> and <code>netip.Addr.As16()</code>, hoping\nthe compiler eventually optimizes it. The <a href=\"https://pkg.go.dev/unsafe\" rel=\"nofollow ugc noopener\"><code>unsafe</code> package</a>\nopens a third path, with the same performance as the first solution.</p>\n<h2 id=\"modifying-the-go-standard-library\">Modifying the Go standard library<a href=\"#modifying-the-go-standard-library\">#</a></h2>\n<p>Internally, <a href=\"https://pkg.go.dev/net/netip#Addr\" rel=\"nofollow ugc noopener\"><code>netip.Addr</code></a> stores any IP address as a\n128-bit value with an extra field <code>z</code> to encode the family and the zone:</p>\n<pre><code>type Addr struct {\n    addr uint128\n    z unique.Handle[addrDetail]\n}\ntype addrDetail struct {\n    isV6   bool   // IPv4 is false, IPv6 is true.\n    zoneV6 string // != &quot;&quot; only if IsV6 is true.\n}\nvar (\n    z0    unique.Handle[addrDetail]\n    z4    = unique.Make(addrDetail{})\n    z6noz = unique.Make(addrDetail{isV6: true})\n)</code></pre>\n<p><a href=\"https://pkg.go.dev/net/netip#AddrFrom4\" rel=\"nofollow ugc noopener\"><code>AddrFrom4()</code></a> encodes an IPv4 address as an\nIPv4-mapped IPv6 address and sets <code>z</code> to the unique value <code>z4</code>:</p>\n<pre><code>// AddrFrom4 returns the address of the IPv4 address given by the bytes in addr.\nfunc AddrFrom4(addr [4]byte) Addr {\n    return Addr{\n        addr: uint128{\n            0,\n            0xffff00000000 |\n                uint64(addr[0])&lt;&lt;24 | uint64(addr[1])&lt;&lt;16 |\n                uint64(addr[2])&lt;&lt;8 | uint64(addr[3])},\n        z: z4,\n    }\n}</code></pre>\n<p><a href=\"https://pkg.go.dev/net/netip#Addr.Unmap\" rel=\"nofollow ugc noopener\"><code>Unmap()</code></a> turns an IPv4-mapped IPv6 address into an\nIPv4 address by setting the <code>z</code> field to <code>z4</code>:</p>\n<pre><code>func (ip Addr) Unmap() Addr {\n    if ip.Is4In6() {\n        ip.z = z4\n    }\n    return ip\n}</code></pre>\n<p>Implementing the reverse direction inside the Go standard library is trivial: we\nset the <code>z</code> field to <code>z6noz</code> if the address is IPv4.</p>\n<pre><code>// To6 maps an IPv4 address to an IPv4-mapped IPv6 address. It returns an\n// IPv6 address unmodified.\nfunc (ip Addr) To6() Addr {\n    if ip.Is4() {\n        ip.z = z6noz\n    }\n    return ip\n}</code></pre>\n<h2 id=\"as-a-helper\">As a helper<a href=\"#as-a-helper\">#</a></h2>\n<p>We can’t access the <code>z</code> field from outside the <code>net/netip</code> package. Instead, we\nbuild a small helper around the <code>netip.AddrFrom16(ip.As16())</code> pattern:</p>\n<pre><code>// AddrTo6 maps an IPv4 address to an IPv4-mapped IPv6 address. It returns an\n// IPv6 address unmodified.\nfunc AddrTo6(ip netip.Addr) netip.Addr {\n    if ip.Is4() {\n        ip = netip.AddrFrom16(ip.As16())\n    }\n    return ip\n}</code></pre>\n<h2 id=\"as-an-unsafe-function\">As an unsafe function<a href=\"#as-an-unsafe-function\">#</a></h2>\n<p>Another solution uses the <code>unsafe</code> package to alter the <code>Addr</code> struct through\na proxy with the same memory layout:<a href=\"#sidenote-tests\">3</a></p>\n<pre><code>// addrProxy has the same memory layout as netip.Addr.\ntype addrProxy struct {\n    addr [2]uint64      // netip.uint128\n    z    unsafe.Pointer // unique.Handle[netip.addrDetail]\n}\nvar (\n    anyIPv6    = netip.IPv6Unspecified()\n    netipZ6noz = (*addrProxy)(unsafe.Pointer(&amp;anyIPv6)).z\n)\n// AddrTo6 maps an IPv4 address to an IPv4-mapped IPv6 address. It returns an\n// IPv6 address unmodified.\nfunc AddrTo6(ip netip.Addr) netip.Addr {\n    if !ip.Is4() {\n        return ip\n    }\n    (*addrProxy)(unsafe.Pointer(&amp;ip)).z = netipZ6noz\n    return ip\n}</code></pre>\n<h2 id=\"benchmarks\">Benchmarks<a href=\"#benchmarks\">#</a></h2>\n<p>On my computer, with Go 1.27.1, the standard library solution costs 0.88 ns per operation, while the solution favored by Go maintainers costs 7.14 ns. The unsafe solution matches the performance of the first one.</p>\n<pre><code>goos: linux\ngoarch: amd64\npkg: github.com/vincentbernat/go-netip-addrto6\ncpu: AMD Ryzen 5 5600X 6-Core Processor\n                   │     sec/op     │\nAddrTo6/safe            7.137n ± 0%\nAddrTo6/unsafe         0.8682n ± 2%\nAddrTo6/builtin        0.8775n ± 2%</code></pre>\n<h2 id=\"assembly-code\">Assembly code<a href=\"#assembly-code\">#</a></h2>\n<p>Let’s check the assembly code the compiler generates for each solution.&lt;sup&gt;<a href=\"#sidenote-dis\">4</a>&lt;/sup&gt;\nThe one built into the standard library looks like this:<a href=\"#sidenote-inline\">5</a></p>\n<pre><code>// AX = input.addr.hi, BX = input.addr.lo, CX = input.z\n CMPQ  net/netip·z4(SB), CX     ; check &quot;z&quot; if this is an IPv4 address\n JNE   end                      ; if not, stop here\n MOVQ  net/netip·z6noz(SB), CX  ; CX = netip.z6noz\nend:\n RET\n// return value = Addr{hi: AX, lo: BX, z: CX}</code></pre>\n<p>Go’s <a href=\"https://go.dev/doc/asm\" rel=\"nofollow ugc noopener\">assembly language</a> is not a direct representation of the underlying\nmachine language: it operates on a semi-abstract instruction set derived from\n<a href=\"https://9p.io/sys/doc/asm.html\" rel=\"nofollow ugc noopener\">Plan 9’s assembler</a>. It has four pseudo-registers: FP (frame pointer\nfor function arguments), PC (program counter), SB (static base pointer for\nglobal symbols), and SP (stack pointer). It also has architecture-specific\nregisters like <code>AX</code>, <code>CX</code>, <code>DX</code>, <code>BX</code>, <code>SI</code>, <code>DI</code>, and <code>R8</code> to <code>R15</code>.\nInstructions storing data use their last argument as the destination.\nInstructions can carry an explicit size suffix: <code>MOVB</code> moves a byte, <code>MOVW</code> 16\nbits, <code>MOVL</code> 32 bits, and <code>MOVQ</code> 64 bits. In the example above, the first\ninstruction compares the 64-bit value <code>z4</code> with the <code>CX</code> register.</p>\n<p>The unsafe solution looks almost the same:</p>\n<pre><code>// AX = input.addr.hi, BX = input.addr.lo, CX = input.z\n CMPQ  net/netip·z4(SB), CX  ; check &quot;z&quot; if this is an IPv4 address\n JNE   end                   ; if not, stop here\n MOVQ  netipZ6noz(SB), CX    ; CX = netip.z6noz\nend:\n RET\n// return value = Addr{hi: AX, lo: BX, z: CX}</code></pre>\n<p>The helper solution has far more instructions. To understand why, let’s look at\nthe code for <a href=\"https://pkg.go.dev/netip.addr.As16\" rel=\"nofollow ugc noopener\"><code>As16()</code></a> and\n<a href=\"https://pkg.go.dev/netip.AddrFrom16\" rel=\"nofollow ugc noopener\"><code>AddrFrom16()</code></a>. They are short enough for the compiler\nto inline them.</p>\n<pre><code>func (ip Addr) As16() (a16 [16]byte) {\n    byteorder.BEPutUint64(a16[:8], ip.addr.hi)\n    byteorder.BEPutUint64(a16[8:], ip.addr.lo)\n    return a16\n}\nfunc AddrFrom16(addr [16]byte) Addr {\n    return Addr{\n        addr: uint128{\n            byteorder.BEUint64(addr[:8]),\n            byteorder.BEUint64(addr[8:]),\n        },\n        z: z6noz,\n    }\n}</code></pre>\n<p>We can already guess the pattern to optimize: the code packs the IP address into an array, copies it, then unpacks it. If we inline the Go code by hand, we get:</p>\n<pre><code>func AddrTo6(input netip.Addr) netip.Addr {\n    if !input.Is4() {\n        return input\n    }\n    var a16 [16]byte\n    byteorder.BEPutUint64(a16[:8], input.addr.hi)\n    byteorder.BEPutUint64(a16[8:], input.addr.lo)\n    addr := a16\n    var output netip.Addr\n    output.addr.hi = byteorder.BEUint64(addr[:8])\n    output.addr.lo = byteorder.BEUint64(addr[8:])\n    output.z = netip.z6noz\n    return output\n}</code></pre>\n<p>As humans, we can mentally derive the optimized form:</p>\n<pre><code>func AddrTo6(input netip.Addr) netip.Addr {\n    if !input.Is4() {\n        return input\n    }\n    var output netip.Addr\n    output.addr.hi = input.addr.hi\n    output.addr.lo = input.addr.lo\n    output.z = netip.z6noz\n    return output\n}</code></pre>\n<p>Unfortunately, as of Go 1.26.8, the compiler is not smart enough to do the same:</p>\n<pre><code>// AX = input.addr.hi, BX = input.addr.lo, CX = input.z\n; Push the stack (32 bytes):\n;    0(SP) addr netip.uint128\n;   16(SP) a16 [16]byte\n PUSHQ   BP\n MOVQ    SP, BP\n SUBQ    $32, SP\n CMPQ    net/netip·z4(SB), CX  ; check &quot;z&quot; if this is an IPv4 address\n JNE     end                   ; if not, stop here\n; Pack: byteorder.BEPutUint64(a16[:8], input.addr.hi)\n;       byteorder.BEPutUint64(a16[8:], input.addr.lo)\n MOVBEQ  AX, net/netip·a16+16(SP)\n MOVBEQ  BX, net/netip·a16+24(SP)\n; addr = a16, 16 bytes at once through the vector register X0\n MOVUPS  net/netip·a16+16(SP), X0\n MOVUPS  X0, net/netip·addr(SP)\n; CX = netip.z6noz\n MOVQ    net/netip·z6noz(SB), CX\n; Unpack: output.addr.hi = byteorder.BEUint64(addr[:8])\n;         output.addr.lo = byteorder.BEUint64(addr[8:])\n MOVBEQ  net/netip·addr(SP), AX\n MOVBEQ  net/netip·addr+8(SP), BX\nend:\n ADDQ    $32, SP\n POPQ    BP\n RET\n// return value = Addr{hi: AX, lo: BX, z: CX}</code></pre>\n<p>The compiler does a decent job on the byte shuffling: the eight byte stores of\n<code>BEPutUint64()</code> become a single <code>MOVBEQ</code>, which stores a register byte-swapped.\nThe eight byte loads of <code>BEUint64()</code> become a single <code>MOVBEQ</code> the other way\nround.&lt;sup&gt;<a href=\"#sidenote-amd64v3\">6</a>&lt;/sup&gt; Three groups of instructions remain: a pack, a copy, and an\nunpack.</p>\n<h1 id=\"hacking-the-go-compiler\">Hacking the Go compiler<a href=\"#hacking-the-go-compiler\">#</a></h1>\n<p>The Go compiler has <a href=\"https://github.com/golang/go/blob/1e963f82914ddfb47820034c5c85205a362ed73f/src/cmd/compile/README.md\" rel=\"nofollow ugc noopener\">several phases</a>:</p>\n<ul><li>Parsing</li><li><p>The compiler <a href=\"https://github.com/golang/go/tree/1e963f82914ddfb47820034c5c85205a362ed73f/src/cmd/compile/internal/syntax\" rel=\"nofollow ugc noopener\">tokenizes and parses the source code</a> . It builds a syntax tree</p><p>for each source file.</p></li><li>Type checking</li><li><p>The compiler <a href=\"https://github.com/golang/go/tree/1e963f82914ddfb47820034c5c85205a362ed73f/src/cmd/compile/internal/types2\" rel=\"nofollow ugc noopener\">maps each identifier to the object it denotes</a> , folds</p><p>constants, and infers the type of every expression.</p></li><li>IR construction</li><li><p>The compiler converts the syntax tree and its types into its own [intermediate</p><p>representation](<a href=\"https://github.com/golang/go/tree/1e963f82914ddfb47820034c5c85205a362ed73f/src/cmd/compile/internal/ir\" rel=\"nofollow ugc noopener\">https://github.com/golang/go/tree/1e963f82914ddfb47820034c5c85205a362ed73f/src/cmd/compile/internal/ir</a>) (IR). This process, called “<a href=\"https://github.com/golang/go/tree/1e963f82914ddfb47820034c5c85205a362ed73f/src/cmd/compile/internal/noder\" rel=\"nofollow ugc noopener\">noding</a> ,” goes through\na serialization format named unified IR.</p></li><li>Middle end</li><li><p>The compiler performs several optimization passes on the IR,</p><p>such as <a href=\"https://github.com/golang/go/tree/1e963f82914ddfb47820034c5c85205a362ed73f/src/cmd/compile/internal/devirtualize\" rel=\"nofollow ugc noopener\">devirtualization</a> ,<a href=\"https://github.com/golang/go/tree/1e963f82914ddfb47820034c5c85205a362ed73f/src/cmd/compile/internal/inline\" rel=\"nofollow ugc noopener\">function call inlining</a> ,\nand<a href=\"https://github.com/golang/go/tree/1e963f82914ddfb47820034c5c85205a362ed73f/src/cmd/compile/internal/escape\" rel=\"nofollow ugc noopener\">escape analysis</a> .</p></li><li>Walk</li><li><p>This phase runs <a href=\"https://github.com/golang/go/tree/1e963f82914ddfb47820034c5c85205a362ed73f/src/cmd/compile/internal/walk\" rel=\"nofollow ugc noopener\">two steps</a> : order of evaluation decomposes complex</p><p>statements into simpler ones, and desugaring transforms higher-level Go\nconstructs, like<code>switch</code> or channels, into more primitive instructions or\ncalls to the runtime.</p></li><li>Generic SSA</li><li><p>The compiler <a href=\"https://github.com/golang/go/tree/1e963f82914ddfb47820034c5c85205a362ed73f/src/cmd/compile/internal/ssagen\" rel=\"nofollow ugc noopener\">converts the IR</a> into Static Single Assignment (SSA)</p><p>form, a lower-level intermediate representation suited for<a href=\"https://github.com/golang/go/tree/1e963f82914ddfb47820034c5c85205a362ed73f/src/cmd/compile/internal/ssa\" rel=\"nofollow ugc noopener\">machine-independent optimizations and rewrite rules</a> .</p></li><li>Machine code generation</li><li><p>The compiler rewrites the SSA form into <a href=\"https://github.com/golang/go/tree/1e963f82914ddfb47820034c5c85205a362ed73f/src/cmd/compile/internal/ssa/_gen\" rel=\"nofollow ugc noopener\">machine-specific variants</a> ,</p><p>allocates registers, and applies more optimization passes. At the end,<a href=\"https://github.com/golang/go/tree/1e963f82914ddfb47820034c5c85205a362ed73f/src/cmd/internal/obj\" rel=\"nofollow ugc noopener\">the\nassembler</a> turns the generated instructions into machine code.</p></li></ul>\n<h2 id=\"the-hammer\">The hammer<a href=\"#the-hammer\">#</a></h2>\n<p>My first idea is to replace occurrences of <code>netip.AddrFrom16(ip.As16())</code> with\n<code>netip.Addr{addr: ip.addr, z: netip.z6noz}</code> as early as possible, during the\n“noding” process. Before that, the type checking phase prevents\nus from accessing unexported struct fields.</p>\n<p>Go 1.27 introduced a convenient debug option to dump the IR of a function at interesting points during compilation:</p>\n<pre><code>$ GOTOOLCHAIN=go1.27.1 GOAMD64=v3 go build -a -gcflags=&quot;-d=astdump=AddrTo6Safe&quot; .\nWriting text ast output for AddrTo6Safe to AddrTo6Safe.ast\nWriting html ast output for AddrTo6Safe to AddrTo6Safe.html\nWriting html syntax output for AddrTo6Safe to AddrTo6Safe.syntax.html</code></pre>\n<p>In the HTML file, the first column shows the IR as it comes out of noding:</p>\n<pre><code>DCLFUNC addrto6.AddrTo6Safe ABI:ABIInternal FUNC-func(netip.Addr) netip.Addr\nDCLFUNC-Dcl\n. NAME-addrto6.ip Class:PPARAM Offset:0 OnStack Used netip.Addr\n. NAME-addrto6.~r0 Class:PPARAMOUT Offset:0 OnStack netip.Addr\nDCLFUNC-body\n. IF # ipv6_safe.go:11:2\n. IF-Cond\n. . CALLFUNC bool\n. . CALLFUNC-Fun\n. . . METHEXPR addrto6.Is4 FUNC-func(netip.Addr) bool\n. . . . TYPE netip.Addr Class:PEXTERN Offset:0 type netip.Addr\n. . CALLFUNC-Args\n. . . NAME-addrto6.ip Class:PPARAM Offset:0 OnStack Used netip.Addr\n. IF-Body\n. . AS # ipv6_safe.go:12:6\n. . . NAME-addrto6.ip Class:PPARAM Offset:0 OnStack Used netip.Addr\n. . . CALLFUNC netip.Addr\n. . . CALLFUNC-Fun\n. . . . NAME-netip.AddrFrom16 Class:PFUNC Offset:0 Used FUNC-func([16]byte) netip.Addr\n. . . CALLFUNC-Args\n. . . . CALLFUNC ARRAY-[16]byte\n. . . . CALLFUNC-Fun\n. . . . . METHEXPR addrto6.As16 FUNC-func(netip.Addr) [16]byte\n. . . . . . TYPE netip.Addr Class:PEXTERN Offset:0 type netip.Addr\n. . . . CALLFUNC-Args\n. . . . . NAME-addrto6.ip Class:PPARAM Offset:0 OnStack Used netip.Addr\n. RETURN # ipv6_safe.go:14:2\n. RETURN-Results\n. . NAME-addrto6.ip Class:PPARAM Offset:0 OnStack Used netip.Addr</code></pre>\n<p>In the body of the <code>if</code> statement, we spot the calls to the method\n<code>netip.Addr.As16()</code> and to the function <code>netip.AddrFrom16()</code>. Our goal is to\npatch them with a struct literal:</p>\n<pre><code>IF-Body\n. AS # ipv6_safe.go:12:6\n. . NAME-addrto6.ip Class:PPARAM Offset:0 OnStack Used netip.Addr\n. . STRUCTLIT netip.Addr\n. . STRUCTLIT-List\n. . . STRUCTKEY netip.addr\n. . . . DOT netip.addr netip.uint128\n. . . . . NAME-addrto6.ip Class:PPARAM Offset:0 OnStack Used netip.Addr\n. . . STRUCTKEY netip.z\n. . . . NAME-netip.z6noz Class:PEXTERN Offset:0 unique.Handle[net/netip.addrDetail]</code></pre>\n<p>In <a href=\"https://github.com/golang/go/blob/1e963f82914ddfb47820034c5c85205a362ed73f/src/cmd/compile/internal/noder/reader.go\" rel=\"nofollow ugc noopener\">noder’s <code>reader.go</code></a>, the <code>expr()</code> method builds the IR tree for\nan expression. At the end of the <code>exprCall</code> case, we add a call to a\n<code>rewriteAddrFrom16As16()</code> function. It takes the current node and returns the\nstruct literal on success, or <code>nil</code> if the rewrite is not possible. First, we\ncheck that we have the expected pattern: a call to the <code>netip.AddrFrom16()</code>\nfunction with a call to the <code>netip.Addr.As16()</code> method as its only argument:</p>\n<pre><code>func rewriteAddrFrom16As16(n ir.Node) ir.Node {\n    call, ok := n.(*ir.CallExpr)\n    if !ok || call.Op() != ir.OCALLFUNC ||\n        len(call.Args) != 1 || len(call.Init()) != 0 ||\n        !isNetipFunc(call.Fun, &quot;AddrFrom16&quot;) {\n        return nil\n    }\n    inner, ok := call.Args[0].(*ir.CallExpr)\n    if !ok || inner.Op() != ir.OCALLFUNC ||\n        len(inner.Args) != 1 || len(inner.Init()) != 0 ||\n        !isNetipFunc(inner.Fun, &quot;Addr.As16&quot;) {\n        return nil\n    }\n    x := inner.Args[0]\n    // [...]\n}</code></pre>\n<p>Then, we fetch <code>netip.z6noz</code>:</p>\n<pre><code>z6noz, err := lookupVar(ir.StaticCalleeName(call.Fun).Sym().Pkg, &quot;z6noz&quot;)\nif err != nil {\n    return nil\n}</code></pre>\n<p>And we build the struct literal:</p>\n<pre><code>typ := call.Type()\npos := call.Pos()\nvar list []ir.Node\nfor i, f := range typ.Fields() {\n    var value ir.Node\n    switch f.Sym.Name {\n    case &quot;addr&quot;:\n        value = typecheck.DotField(pos, x, i)\n    case &quot;z&quot;:\n        value = z6noz\n    default:\n        return nil\n    }\n    list = append(list, ir.NewStructKeyExpr(pos, f, value))\n}\nlit := ir.NewCompLitExpr(pos, ir.OSTRUCTLIT, typ, list)\nlit.SetTypecheck(1)\nreturn lit</code></pre>\n<p>Have a look at the <a href=\"https://github.com/vincentbernat/go/commit/netipmap-fix2\" rel=\"nofollow ugc noopener\">complete patch</a>.&lt;sup&gt;<a href=\"#sidenote-statement\">7</a>&lt;/sup&gt; We can test it\nwith the following commands:</p>\n<pre><code>$ cd src\n$ ./make.bash\nBuilding Go cmd/dist using /usr/lib/go-1.27. (go1.27.1 linux/amd64)\nBuilding Go toolchain1 and bootstrap cmd/go (go_bootstrap) using /usr/lib/go-1.27.\nBuilding Go toolchain2 using go_bootstrap and Go toolchain1.\nBuilding Go toolchain3 and commands using go_bootstrap and Go toolchain2.\nChecking command staleness for linux/amd64.\n---\nInstalled Go for linux/amd64 in /home/bernat/code/free/go\nInstalled commands in /home/bernat/code/free/go/bin\n*** You need to add /home/bernat/code/free/go/bin to your PATH.\n$ export PATH=$PWD/../bin:$PATH\n$ go version\ngo version go1.28-devel_9834516e20 Sat Sep 12 08:23:11 2026 -0700 linux/amd64\n$ go test net/netip/...\nok      net/netip   0.224s\n$ cd ../../go-netip-addrto6\n$ go test .\nok      github.com/vincentbernat/go-netip-addrto6   0.062s</code></pre>\n<p>The generated code for the helper is now the shortest possible version!</p>\n<pre><code>// AX = input.addr.hi, BX = input.addr.lo, CX = input.z\n CMPQ    net/netip·z4(SB), CX     ; check &quot;z&quot; if this is an IPv4 address\n JNE     end                      ; if not, stop here\n MOVQ    net/netip·z6noz(SB), CX  ; CX = netip.z6noz\nend:\n RET\n// return value = Addr{hi: AX, lo: BX, z: CX}</code></pre>\n<p>Go maintainers are unlikely to accept this change. It relies on the internal\nstructure of <code>net/netip.Addr</code>. It’s an ugly hack in the noder, whose job is to\nfaithfully translate the type-checked AST into the IR. And it’s harder to\nmaintain than adding a <code>To6()</code> method.</p>\n<h2 id=\"the-screwdriver\">The screwdriver<a href=\"#the-screwdriver\">#</a></h2>\n<p>The right place for such an optimization is the <a href=\"https://github.com/golang/go/blob/1e963f82914ddfb47820034c5c85205a362ed73f/src/cmd/compile/internal/ssa/README.md\" rel=\"nofollow ugc noopener\">generic SSA phase</a>. One of\nthe last machine-independent passes is <code>memcombine</code>. With the appropriate debug\nflag, the compiler dumps the SSA form after this pass:<a href=\"#sidenote-html\">8</a></p>\n<pre><code>$ GOTOOLCHAIN=go1.26.8 GOAMD64=v3 \\\n&gt; go build -a -gcflags=&#39;-d=ssa/memcombine/dump=AddrTo6Safe&#39; .\n$ head -5 AddrTo6Safe_01__memcombine.dump\nAddrTo6Safe func(netip.Addr) netip.Addr\n  b2:\n    (?) v1 = InitMem &lt;mem&gt;\n    (?) v2 = SP &lt;uintptr&gt;\n    (?) v3 = SB &lt;uintptr&gt;</code></pre>\n<p>The result of the <code>memcombine</code> pass follows the same structure as the assembly\ncode for <code>AddrTo6Safe()</code> <a href=\"#asm-addrto6safe\">we looked at earlier</a>: two stores,\none move, and two loads we would like to optimize away.</p>\n<pre><code>; […]\n  v502 = ArgIntReg &lt;uint64&gt; {ip+0} [0]              ; input.addr.hi\n  v490 = ArgIntReg &lt;uint64&gt; {ip+8} [1]              ; input.addr.lo\n  v466 = ArgIntReg &lt;*netip.addrDetail&gt; {ip+16} [2]  ; input.z\n; […]\n  v22 = LocalAddr &lt;*[16]byte&gt; {netip.a16} v2 v1     ; &amp;a16\n  v173 = OffPtr &lt;*byte&gt; [8] v22                     ; &amp;a16[8]\n  v442 = Bswap64 &lt;uint64&gt; v490                      ; bswap(input.addr.lo)\n  v542 = Bswap64 &lt;uint64&gt; v502                      ; bswap(input.addr.hi)\n  v161 = Store &lt;mem&gt; {uint64} v22 v542 v23          ; a16[:8] = bswap(hi)\n  v282 = Store &lt;mem&gt; {uint64} v173 v442 v161        ; a16[8:] = bswap(lo)\n  v285 = LocalAddr &lt;*[16]byte&gt; {netip.addr} v2 v282 ; &amp;addr\n  v286 = Move &lt;mem&gt; {[16]byte} [16] v285 v22 v282   ; addr = a16\n  v415 = OffPtr &lt;*byte&gt; [8] v285                    ; &amp;addr[8]\n  v416 = Load &lt;uint64&gt; v285 v286                    ; addr[:8]\n  v299 = Bswap64 &lt;uint64&gt; v416                      ; output.addr.hi\n  v174 = Load &lt;uint64&gt; v415 v286                    ; addr[8:]\n  v39  = Bswap64 &lt;uint64&gt; v174                      ; output.addr.lo\n; […]</code></pre>\n<p>Each line features a value identifier (<code>v442</code>), an operation with its type\n(<code>Bswap64 &lt;uint64&gt;</code>), and its arguments (<code>v490</code>).&lt;sup&gt;<a href=\"#sidenote-lines\">9</a>&lt;/sup&gt; Values are the basic\nbuilding blocks of SSA and are defined exactly once. Square brackets enclose\ninteger parameters (<code>[8]</code>) and curly braces contain auxiliary arguments\n(<code>{netip.addr}</code>). Operations writing to memory produce a new memory state. Every\nmemory operation takes the current state as its last argument, which keeps them\nin order.</p>\n<h3 id=\"on-paper\">On paper<a href=\"#on-paper\">#</a></h3>\n<p>Let’s focus on <code>output.addr.lo</code>, aka <code>v39</code>:</p>\n<pre><code>  v490 = ArgIntReg &lt;uint64&gt; {ip+8} [1]              ; input.addr.lo\n  v22 = LocalAddr &lt;*[16]byte&gt; {netip.a16} v2 v1     ; &amp;a16\n  v173 = OffPtr &lt;*byte&gt; [8] v22                     ; &amp;a16[8]\n  v442 = Bswap64 &lt;uint64&gt; v490                      ; bswap(input.addr.lo)\n  v282 = Store &lt;mem&gt; {uint64} v173 v442 v161        ; a16[8:] = bswap(lo)\n  v285 = LocalAddr &lt;*[16]byte&gt; {netip.addr} v2 v282 ; &amp;addr\n  v286 = Move &lt;mem&gt; {[16]byte} [16] v285 v22 v282   ; addr = a16\n  v415 = OffPtr &lt;*byte&gt; [8] v285                    ; &amp;addr[8]\n  v174 = Load &lt;uint64&gt; v415 v286                    ; addr[8:]\n  v39  = Bswap64 &lt;uint64&gt; v174                      ; output.addr.lo</code></pre>\n<p>To simplify this code, we could apply three rewriting rules:</p>\n<ol><li></li></ol>\n<p>The first one adds a shortcut when <strong>loading through a move</strong> :```\n(Load (OffPtr\n   [o] p) (Move p src mem)) =&gt; (Load (OffPtr [o] src) mem)</p>\n<pre><code>. This matches`v174` with its arguments`v415` and`v286` and creates a new value`v600` :```\n  v490 = ArgIntReg &lt;uint64&gt; {ip+8} [1]              ; input.addr.lo\n  v22 = LocalAddr &lt;*[16]byte&gt; {netip.a16} v2 v1     ; &amp;a16\n  v173 = OffPtr &lt;*byte&gt; [8] v22                     ; &amp;a16[8]\n  v442 = Bswap64 &lt;uint64&gt; v490                      ; bswap(input.addr.lo)\n  v282 = Store &lt;mem&gt; {uint64} v173 v442 v161        ; a16[8:] = bswap(lo)\n  v285 = LocalAddr &lt;*[16]byte&gt; {netip.addr} v2 v282 ; &amp;addr\n  v286 = Move &lt;mem&gt; {[16]byte} [16] v285 v22 v282   ; addr = a16\n  v415 = OffPtr &lt;*byte&gt; [8] v285                    ; &amp;addr[8]\n  v600 = OffPtr &lt;*byte&gt; [8] v22                     ; &amp;a16[8]\n  v174 = Load &lt;uint64&gt; v600 v282                    ; a16[8:]\n  v39  = Bswap64 &lt;uint64&gt; v174                      ; output.addr.lo</code></pre>\n<ol start=\"2\"><li></li></ol>\n<p>The second one simplifies a <strong>load following a store</strong> :```\n(Load p (Store p x\n   _)) =&gt; x</p>\n<pre><code>. The load is*forwarded* : the stored value replaces it and no memory\n   access remains. It matches`v174` . It notices that`v600` and`v173` are the\n   same address and replaces`v174` with a copy of`v442` :```\n  v490 = ArgIntReg &lt;uint64&gt; {ip+8} [1]              ; input.addr.lo\n  v22 = LocalAddr &lt;*[16]byte&gt; {netip.a16} v2 v1     ; &amp;a16\n  v173 = OffPtr &lt;*byte&gt; [8] v22                     ; &amp;a16[8]\n  v442 = Bswap64 &lt;uint64&gt; v490                      ; bswap(input.addr.lo)\n  v282 = Store &lt;mem&gt; {uint64} v173 v442 v161        ; a16[8:] = bswap(lo)\n  v285 = LocalAddr &lt;*[16]byte&gt; {netip.addr} v2 v282 ; &amp;addr\n  v286 = Move &lt;mem&gt; {[16]byte} [16] v285 v22 v282   ; addr = a16\n  v415 = OffPtr &lt;*byte&gt; [8] v285                    ; &amp;addr[8]\n  v600 = OffPtr &lt;*byte&gt; [8] v22                     ; &amp;a16[8]\n  v174 = Copy &lt;uint64&gt; v442                         ; bswap(input.addr.lo)\n  v39  = Bswap64 &lt;uint64&gt; v174                      ; output.addr.lo</code></pre>\n<ol start=\"3\"><li></li></ol>\n<p>The last step <strong>cancels the two byte swaps</strong> :<code>(Bswap64 (Bswap64 x)) =&gt; x</code> .<code>v39</code> becomes a copy of<code>v490</code> :```\n  v490 = ArgIntReg &lt;uint64&gt; {ip+8} [1]              ; input.addr.lo\n  v22 = LocalAddr &lt;*[16]byte&gt; {netip.a16} v2 v1     ; &amp;a16\n  v173 = OffPtr &lt;*byte&gt; [8] v22                     ; &amp;a16[8]\n  v442 = Bswap64 &lt;uint64&gt; v490                      ; bswap(input.addr.lo)\n  v282 = Store &lt;mem&gt; {uint64} v173 v442 v161        ; a16[8:] = bswap(lo)\n  v285 = LocalAddr &lt;*[16]byte&gt; {netip.addr} v2 v282 ; &amp;addr\n  v286 = Move &lt;mem&gt; {[16]byte} [16] v285 v22 v282   ; addr = a16\n  v415 = OffPtr &lt;*byte&gt; [8] v285                    ; &amp;addr[8]\n  v600 = OffPtr &lt;*byte&gt; [8] v22                     ; &amp;a16[8]\n  v174 = Copy &lt;uint64&gt; v442                         ; bswap(input.addr.lo)\n  v39  = Copy &lt;uint64&gt; v490                         ; output.addr.lo = input.addr.lo</p>\n<pre><code>\nIf we ignore the values not needed to compute `v39`, only this SSA form remains:\n</code></pre>\n<p>  v490 = ArgIntReg &lt;uint64&gt; {ip+8} [1]  ; input.addr.lo\n  v39  = Copy &lt;uint64&gt; v490             ; output.addr.lo = input.addr.lo</p>\n<pre><code>Let’s switch to `output.addr.hi`, aka `v299`:\n</code></pre>\n<p>  v502 = ArgIntReg &lt;uint64&gt; {ip+0} [0]              ; input.addr.hi\n  v22 = LocalAddr &lt;*[16]byte&gt; {netip.a16} v2 v1     ; &amp;a16\n  v173 = OffPtr &lt;*byte&gt; [8] v22                     ; &amp;a16[8]\n  v542 = Bswap64 &lt;uint64&gt; v502                      ; bswap(input.addr.hi)\n  v161 = Store &lt;mem&gt; {uint64} v22 v542 v23          ; a16[:8] = bswap(hi)\n  v282 = Store &lt;mem&gt; {uint64} v173 v442 v161        ; a16[8:] = bswap(lo)\n  v285 = LocalAddr &lt;*[16]byte&gt; {netip.addr} v2 v282 ; &amp;addr\n  v286 = Move &lt;mem&gt; {[16]byte} [16] v285 v22 v282   ; addr = a16\n  v416 = Load &lt;uint64&gt; v285 v286                    ; addr[:8]\n  v299 = Bswap64 &lt;uint64&gt; v416                      ; output.addr.hi</p>\n<pre><code>To optimize it away, we also apply three rewriting rules:\n\n1. \nThe first one also adds a shortcut when **loading through a move** , but\n   without an offset:`(Load p (Move p src mem)) =&gt; (Load src mem)` . This rewrites`v416` to use arguments from`v286` :```\n  v502 = ArgIntReg &lt;uint64&gt; {ip+0} [0]              ; input.addr.hi\n  v22 = LocalAddr &lt;*[16]byte&gt; {netip.a16} v2 v1     ; &amp;a16\n  v173 = OffPtr &lt;*byte&gt; [8] v22                     ; &amp;a16[8]\n  v542 = Bswap64 &lt;uint64&gt; v502                      ; bswap(input.addr.hi)\n  v161 = Store &lt;mem&gt; {uint64} v22 v542 v23          ; a16[:8] = bswap(hi)\n  v282 = Store &lt;mem&gt; {uint64} v173 v442 v161        ; a16[8:] = bswap(lo)\n  v285 = LocalAddr &lt;*[16]byte&gt; {netip.addr} v2 v282 ; &amp;addr\n  v286 = Move &lt;mem&gt; {[16]byte} [16] v285 v22 v282   ; addr = a16\n  v416 = Load &lt;uint64&gt; v22 v282                     ; a16[:8]\n  v299 = Bswap64 &lt;uint64&gt; v416                      ; output.addr.hi</code></pre>\n<ol start=\"2\"><li></li></ol>\n<p>The second rule <strong>forwards a value stored one step earlier</strong> , skipping over a\n   store to another address:<code>(Load p (Store q _ (Store p x _))) =&gt; x</code> . This\n   matches<code>v416</code> :<code>x</code> is<code>v542</code> ,<code>p</code> is<code>v22</code> (<code>&amp;a16</code> ),<code>q</code> is<code>v173</code> (<code>&amp;a16[8]</code> ), and<code>p</code> and<code>q</code> do not overlap for<code>uint64</code> .```\n  v502 = ArgIntReg &lt;uint64&gt; {ip+0} [0]              ; input.addr.hi\n  v22 = LocalAddr &lt;*[16]byte&gt; {netip.a16} v2 v1     ; &amp;a16\n  v173 = OffPtr &lt;*byte&gt; [8] v22                     ; &amp;a16[8]\n  v542 = Bswap64 &lt;uint64&gt; v502                      ; bswap(input.addr.hi)\n  v161 = Store &lt;mem&gt; {uint64} v22 v542 v23          ; a16[:8] = bswap(hi)\n  v282 = Store &lt;mem&gt; {uint64} v173 v442 v161        ; a16[8:] = bswap(lo)\n  v285 = LocalAddr &lt;*[16]byte&gt; {netip.addr} v2 v282 ; &amp;addr\n  v286 = Move &lt;mem&gt; {[16]byte} [16] v285 v22 v282   ; addr = a16\n  v416 = Copy &lt;uint64&gt; v542                         ; bswap(input.addr.hi)\n  v299 = Bswap64 &lt;uint64&gt; v416                      ; output.addr.hi</p>\n<pre><code>3. \nThe third rule **cancels two byte swaps** :`(Bswap64 (Bswap64 x)) =&gt; x` .`v299` becomes a copy of`v502` :```\n  v502 = ArgIntReg &lt;uint64&gt; {ip+0} [0]              ; input.addr.hi\n  v22 = LocalAddr &lt;*[16]byte&gt; {netip.a16} v2 v1     ; &amp;a16\n  v173 = OffPtr &lt;*byte&gt; [8] v22                     ; &amp;a16[8]\n  v542 = Bswap64 &lt;uint64&gt; v502                      ; bswap(input.addr.hi)\n  v161 = Store &lt;mem&gt; {uint64} v22 v542 v23          ; a16[:8] = bswap(hi)\n  v282 = Store &lt;mem&gt; {uint64} v173 v442 v161        ; a16[8:] = bswap(lo)\n  v285 = LocalAddr &lt;*[16]byte&gt; {netip.addr} v2 v282 ; &amp;addr\n  v286 = Move &lt;mem&gt; {[16]byte} [16] v285 v22 v282   ; addr = a16\n  v416 = Copy &lt;uint64&gt; v542                         ; bswap(input.addr.hi)\n  v299 = Copy &lt;uint64&gt; v502                         ; output.addr.hi = input.addr.hi</code></pre>\n<p>If we remove the values not used to compute <code>v299</code>, we get this SSA form:</p>\n<pre><code>  v502 = ArgIntReg &lt;uint64&gt; {ip+0} [0] ; input.addr.hi\n  v299 = Copy &lt;uint64&gt; v502            ; output.addr.hi = input.addr.hi</code></pre>\n<h3 id=\"in-practice\">In practice<a href=\"#in-practice\">#</a></h3>\n<p>Most of these rules already exist in <a href=\"https://github.com/golang/go/blob/1e963f82914ddfb47820034c5c85205a362ed73f/src/cmd/compile/internal/ssa/_gen/generic.rules\" rel=\"nofollow ugc noopener\"><code>generic.rules</code></a>. They use\nconditions to validate their context: <code>ssa.IsSamePtr()</code> for the same address,\n<code>ssa.Disjoint()</code> for addresses that do not overlap. The rule forwarding a stored\nvalue to a load already exists with three variants looking through several other\nstores. Here are the two we need:</p>\n<pre><code>(Load &lt;t1&gt; p1 (Store {t2} p2 x _))\n    &amp;&amp; ssa.IsSamePtr(p1, p2)\n    &amp;&amp; copyCompatibleType(t1, x.Type)\n    &amp;&amp; t1.Size() == t2.Size()\n    =&gt; x\n(Load &lt;t1&gt; p1 (Store {t2} p2 _ (Store {t3} p3 x _)))\n    &amp;&amp; ssa.IsSamePtr(p1, p3)\n    &amp;&amp; copyCompatibleType(t1, x.Type)\n    &amp;&amp; t1.Size() == t3.Size()\n    &amp;&amp; ssa.Disjoint(p3, t3, p2, t2)\n    =&gt; x</code></pre>\n<p>Go 1.27 added the rule loading through a move with <a href=\"https://go-review.googlesource.com/c/go/+/748200\" rel=\"nofollow ugc noopener\">CL 748200</a> to fix\n<a href=\"https://github.com/golang/go/issues/77720\" rel=\"nofollow ugc noopener\">issue #77720</a>:</p>\n<pre><code>(Load &lt;t1&gt; op1:(OffPtr [o1] p1) move:(Move [n] p2 src mem))\n    &amp;&amp; o1 &gt;= 0 &amp;&amp; o1+t1.Size() &lt;= n &amp;&amp; ssa.IsSamePtr(p1, p2)\n    &amp;&amp; !ssa.IsVolatile(src)\n    =&gt; @move.Block (Load &lt;t1&gt; (OffPtr &lt;op1.Type&gt; [o1] src) mem)</code></pre>\n<p>It lacks a variant without an offset:</p>\n<pre><code>(Load &lt;t1&gt; p1 move:(Move [n] p2 src mem))\n    &amp;&amp; p1.Op != ssaop.OpOffPtr\n    &amp;&amp; t1.Size() &lt;= n &amp;&amp; ssa.IsSamePtr(p1, p2)\n    &amp;&amp; !ssa.IsVolatile(src)\n    =&gt; @move.Block (Load &lt;t1&gt; (OffPtr &lt;p1.Type&gt; [0] src) mem)</code></pre>\n<p>There is no generic rule to cancel two byte swaps, but the <a href=\"https://github.com/golang/go/blob/1e963f82914ddfb47820034c5c85205a362ed73f/src/cmd/compile/internal/ssa/_gen/AMD64.rules\" rel=\"nofollow ugc noopener\">AMD64 lowering\npass</a> includes this rule:</p>\n<pre><code>(BSWAP(Q|L) (BSWAP(Q|L) p)) =&gt; p</code></pre>\n<p>After switching to Go’s development branch and <a href=\"https://github.com/vincentbernat/go/commit/netipmap-fix3\" rel=\"nofollow ugc noopener\">adding the missing\nrule</a>, the generated assembly code is worse than with Go 1.26.8,\neven though our additional rule slightly improves the situation at the end:</p>\n<pre><code>// AX = input.addr.hi, BX = input.addr.lo, CX = input.z\n; Push the stack (16 bytes):\n;    0(SP) a16 [16]byte\n PUSHQ   BP\n MOVQ    SP, BP\n SUBQ    $16, SP\n CMPQ    net/netip·z4(SB), CX       ; check &quot;z&quot; if this is an IPv4 address\n JNE     end                        ; if not, stop here\n; The four forwarded bytes: the low half of input.addr.lo is taken apart and\n; put back together in registers\n MOVQ    BX, DX                     ; DX = input.addr.lo\n SHRQ    $24, BX                    ; BX = input.addr.lo &gt;&gt; 24\n MOVQ    DX, SI                     ; SI = input.addr.lo\n SHRQ    $16, DX                    ; DX = input.addr.lo &gt;&gt; 16\n MOVQ    SI, DI                     ; DI = input.addr.lo, kept for the pack\n SHRQ    $8, SI                     ; SI = input.addr.lo &gt;&gt; 8\n MOVBLZX DIB, R8                    ; R8 = byte(input.addr.lo)\n MOVBLZX SIB, SI                    ; SI = byte(input.addr.lo &gt;&gt; 8)\n SHLQ    $8, SI\n ORQ     R8, SI                     ; SI = two low bytes of input.addr.lo\n MOVBLZX DL, DX                     ; DX = byte(input.addr.lo &gt;&gt; 16)\n SHLQ    $16, DX\n ORQ     SI, DX\n MOVBLZX BL, BX                     ; BX = byte(input.addr.lo &gt;&gt; 24)\n SHLQ    $24, BX\n ORQ     DX, BX                     ; BX = input.addr.lo &amp; 0xffffffff\n; Pack: byteorder.BEPutUint64(a16[:8], input.addr.hi)\n;       byteorder.BEPutUint64(a16[8:], input.addr.lo)\n MOVBEQ  AX, net/netip·a16(SP)\n MOVBEQ  DI, net/netip·a16+8(SP)\n; The four other bytes of input.addr.lo, read one by one from a16\n MOVBLZX net/netip·a16+11(SP), DX   ; a16[11]\n SHLQ    $32, DX\n ORQ     DX, BX\n MOVBLZX net/netip·a16+10(SP), DX   ; a16[10]\n SHLQ    $40, DX\n ORQ     DX, BX\n MOVBLZX net/netip·a16+9(SP), DX    ; a16[9]\n SHLQ    $48, DX\n ORQ     DX, BX\n MOVBLZX net/netip·a16+8(SP), DX    ; a16[8]\n SHLQ    $56, DX\n; output.z = netip.z6noz\n MOVQ    net/netip·z6noz(SB), CX\n; output.addr.hi = byteorder.BEUint64(a16[:8])\n MOVBEQ  net/netip·a16(SP), AX\n; output.addr.lo assembled from the previous steps\n ORQ     DX, BX\nend:\n LEAVEQ\n RET\n// return value = Addr{hi: AX, lo: BX, z: CX}</code></pre>\n<p>The rule loading through a move, added in Go 1.27, introduced this regression.</p>\n<h3 id=\"out-of-order\">Out of order<a href=\"#out-of-order\">#</a></h3>\n<p>Let’s not give up now! In reality, the rewriting rules run before <code>memcombine</code>,\nnotably in the <code>late opt</code> pass. At this point, the inlined versions of\n<code>BEPutUint64()</code> and <code>BEUint64()</code> still expand to sixteen byte stores and sixteen\nbyte loads, matching their <a href=\"https://pkg.go.dev/internal/byteorder#BEUint64\" rel=\"nofollow ugc noopener\">source code</a>:</p>\n<pre><code>func BEUint64(b []byte) uint64 {\n    _ = b[7] // bounds check hint to compiler; see golang.org/issue/14808\n    return uint64(b[7]) | uint64(b[6])&lt;&lt;8 | uint64(b[5])&lt;&lt;16 | uint64(b[4])&lt;&lt;24 |\n        uint64(b[3])&lt;&lt;32 | uint64(b[2])&lt;&lt;40 | uint64(b[1])&lt;&lt;48 | uint64(b[0])&lt;&lt;56\n}</code></pre>\n<p>Let’s follow two bytes of <code>output.addr.lo</code>: <code>addr[15]</code> and <code>addr[11]</code>. Here is\na simplified SSA form before <code>late opt</code>:</p>\n<pre><code>  v273 = Trunc64to8 &lt;byte&gt; v490                     ; byte(input.addr.lo)\n  v226 = Trunc64to8 &lt;byte&gt; v225                     ; byte(input.addr.lo &gt;&gt; 32)\n; […]\n  v235 = Store &lt;mem&gt; {byte} v233 v226 v223          ; a16[11] = byte(lo &gt;&gt; 32)\n  v247 = Store &lt;mem&gt; {byte} v245 v238 v235          ; a16[12] = …\n  v259 = Store &lt;mem&gt; {byte} v257 v250 v247          ; a16[13] = …\n  v271 = Store &lt;mem&gt; {byte} v269 v262 v259          ; a16[14] = …\n  v282 = Store &lt;mem&gt; {byte} v280 v273 v271          ; a16[15] = byte(lo)\n  v285 = LocalAddr &lt;*[16]byte&gt; {netip.addr} v2 v282 ; &amp;addr\n  v286 = Move &lt;mem&gt; {[16]byte} [16] v285 v22 v282   ; addr = a16\n; […]\n  v433 = OffPtr &lt;*byte&gt; [15] v285                   ; &amp;addr[15]\n  v435 = Load &lt;byte&gt; v433 v286                      ; addr[15]\n  v479 = OffPtr &lt;*byte&gt; [11] v285                   ; &amp;addr[11]\n  v481 = Load &lt;byte&gt; v479 v286                      ; addr[11]</code></pre>\n<p>The first rule loads through the move: ```\n(Load (OffPtr [o] p) (Move p src mem))\n=&gt; (Load (OffPtr [o] src) mem)</p>\n<pre><code>. It matches both loads, which now read `a16`\nwith the memory state before the copy:\n</code></pre>\n<p>  v600 = OffPtr &lt;*byte&gt; [15] v22 ; &amp;a16[15]\n  v435 = Load &lt;byte&gt; v600 v282   ; a16[15]\n  v601 = OffPtr &lt;*byte&gt; [11] v22 ; &amp;a16[11]\n  v481 = Load &lt;byte&gt; v601 v282   ; a16[11]</p>\n<pre><code>The second rule shortcuts a load following a store: ```\n(Load p (Store p x _)) =&gt;\nx</code></pre>\n<p>. It matches <code>v435</code>, as <code>v282</code> stores <code>a16[15]</code>. It does not match <code>v481</code>:\n<code>v235</code> stores <code>a16[11]</code> four stores earlier in the chain, while the variants of\nthis rule look through three stores at most.</p>\n<pre><code>  v435 = Copy &lt;byte&gt; v273        ; byte(input.addr.lo)\n  v601 = OffPtr &lt;*byte&gt; [11] v22 ; &amp;a16[11]\n  v481 = Load &lt;byte&gt; v601 v282   ; a16[11]</code></pre>\n<p>The same happens to the other bytes: the rule forwards the four bytes stored\nlast, <code>a16[12]</code> to <code>a16[15]</code>. The twelve other loads now read <code>a16</code> instead of\n<code>addr</code>.</p>\n<p><code>BEUint64()</code> becomes a chain of <code>Or64</code>, each one adding a byte shifted into\nplace. <code>memcombine</code> is a <a href=\"https://github.com/golang/go/blob/1e963f82914ddfb47820034c5c85205a362ed73f/src/cmd/compile/internal/ssacompile/memcombine.go\" rel=\"nofollow ugc noopener\">pass written in Go</a>, not a set of rewrite\nrules. It starts from the last <code>Or64</code> of the chain and collects up to eight\nterms. If each term is a byte load, extended to 64 bits and shifted, and if the\neight loads read consecutive addresses from the same pointer with the same\nmemory state, it replaces the whole chain with a single 64-bit load and a byte\nswap. Otherwise, it tries again with four, then two terms, and from each\nintermediate <code>Or64</code>. Here is the loop checking each term in a simplified version\nof <code>combineLoads()</code>:</p>\n<pre><code>for i := int64(0); i &lt; n; i++ {\n    v := a[i]\n    shift := int64(0)\n    if v.Op == shiftOp {\n        v, shift = peelShift(v)\n    }\n    if v.Op != extOp {\n        return false\n    }\n    load := v.Args[0]\n    if load.Op != ssaop.OpLoad {\n        return false\n    }\n    if load.Args[1] != mem {\n        return false\n    }\n    p, off := splitPtr(load.Args[0])\n    if p != base {\n        return false\n    }\n    r[i] = LoadRecord{load: load, offset: off, shift: shift}\n}</code></pre>\n<p>For <code>output.addr.hi</code>, the eight loads read <code>a16</code> with the same memory state\n<code>v282</code>:</p>\n<pre><code>  v13  = Load &lt;byte&gt; v22 v282               ; a16[0]\n  v530 = Load &lt;byte&gt; v14 v282               ; a16[1]\n  v488 = Load &lt;byte&gt; v504 v282              ; a16[2]\n  v405 = Load &lt;byte&gt; v537 v282              ; a16[3]\n  v385 = Load &lt;byte&gt; v397 v282              ; a16[4]\n  v361 = Load &lt;byte&gt; v373 v282              ; a16[5]\n  v196 = Load &lt;byte&gt; v63 v282               ; a16[6]\n  v432 = Load &lt;byte&gt; v315 v282              ; a16[7]\n  v319 = ZeroExt8to64 &lt;uint64&gt; v432         ; uint64(a16[7])\n  v329 = ZeroExt8to64 &lt;uint64&gt; v196         ; uint64(a16[6])\n  v330 = Lsh64x64 &lt;uint64&gt; [true] v329 v138 ; uint64(a16[6]) &lt;&lt; 8\n  v331 = Or64 &lt;uint64&gt; v319 v330            ; a16[7] | a16[6] &lt;&lt; 8\n; […] same for a16[5] to a16[1]\n  v401 = ZeroExt8to64 &lt;uint64&gt; v13          ; uint64(a16[0])\n  v402 = Lsh64x64 &lt;uint64&gt; [true] v401 v55  ; uint64(a16[0]) &lt;&lt; 56\n  v403 = Or64 &lt;uint64&gt; v402 v391            ; | a16[0] &lt;&lt; 56 = output.addr.hi</code></pre>\n<p><code>memcombine</code> merges them into one load and a swap:</p>\n<pre><code>  v286 = Load &lt;uint64&gt; v22 v282   ; a16[:8]\n  v285 = Bswap64 &lt;uint64&gt; v286    ; output.addr.hi</code></pre>\n<p>For <code>output.addr.lo</code>, here is the chain <code>memcombine</code> sees after <code>late opt</code>:</p>\n<pre><code>  v436 = ZeroExt8to64 &lt;uint64&gt; v273 ; addr[15], forwarded\n  v448 = Or64 &lt;uint64&gt; v436 v447    ; | addr[14] &lt;&lt; 8, forwarded\n  v460 = Or64 &lt;uint64&gt; v459 v448    ; | addr[13] &lt;&lt; 16, forwarded\n  v472 = Or64 &lt;uint64&gt; v471 v460    ; | addr[12] &lt;&lt; 24, forwarded\n  v484 = Or64 &lt;uint64&gt; v483 v472    ; | a16[11] &lt;&lt; 32, loaded\n  v496 = Or64 &lt;uint64&gt; v495 v484    ; | a16[10] &lt;&lt; 40, loaded\n  v508 = Or64 &lt;uint64&gt; v507 v496    ; | a16[9] &lt;&lt; 48, loaded\n  v520 = Or64 &lt;uint64&gt; v519 v508    ; | a16[8] &lt;&lt; 56, loaded</code></pre>\n<p>From <code>v520</code>, four of the eight terms are forwarded bytes, not loads from memory,\nand <code>memcombine</code> can’t combine them. It doesn’t merge the four remaining loads\neither, as they sit on top of the forwarded bytes.</p>\n<h3 id=\"back-in-order\">Back in order<a href=\"#back-in-order\">#</a></h3>\n<p>In summary, the rewriting rules run too early to be effective. A quick\nworkaround exists: run an earlier round of <code>memcombine</code> before <code>late opt</code>. After\n<a href=\"https://github.com/vincentbernat/go/commit/netipmap-fix4\" rel=\"nofollow ugc noopener\">this change</a>, the generated code for the helper is back to the\nshortest possible version:</p>\n<pre><code>// AX = input.addr.hi, BX = input.addr.lo, CX = input.z\n CMPQ    net/netip·z4(SB), CX     ; check &quot;z&quot; if this is an IPv4 address\n JNE     end                      ; if not, stop here\n MOVQ    net/netip·z6noz(SB), CX  ; CX = netip.z6noz\nend:\n RET\n// return value = Addr{hi: AX, lo: BX, z: CX}</code></pre>\n<p>And the benchmark confirms it! ✌️</p>\n<pre><code>goos: linux\ngoarch: amd64\npkg: github.com/vincentbernat/go-netip-addrto6\ncpu: AMD Ryzen 5 5600X 6-Core Processor\n                   │   Go 1.26.8    │             Our branch              │\n                   │     sec/op     │    sec/op     vs base               │\nAddrTo6/safe          6.5470n ±  0%   0.8944n ± 4%  -86.34% (p=0.002 n=6)\nAddrTo6/unsafe        0.9071n ±  3%   0.8682n ± 1%   -4.28% (p=0.002 n=6)\nAddrTo6/builtin       0.8871n ±  2%   0.8785n ± 1%        ~ (p=0.310 n=6)</code></pre>\n<h1 id=\"next-steps\">Next steps<a href=\"#next-steps\">#</a></h1>\n<p>I think Go maintainers would reject this change because of the additional\n<code>memcombine</code> pass. Instead, I plan to publish this blog post and bring up the\nsubject again as a follow-up to <a href=\"https://github.com/golang/go/issues/54365\" rel=\"nofollow ugc noopener\">issue #54365</a>. Either the sheer complexity and\nthe Go 1.27 regression convince the maintainers that adding a <code>To6()</code> method is\nsimpler and more efficient, or they advise me on how to move forward. Either\nway, digging into this subject taught me a lot about the Go compiler! ⚙️</p>\n<p>Update (2026-10)</p>\n<p>I opened <a href=\"https://github.com/golang/go/issues/81994\" rel=\"nofollow ugc noopener\">issue #81994</a> to propose <code>Addr.To6()</code>. Give it a 👍 if you want it in Go!</p>","headings":[{"level":1,"text":"Hacking the Go compiler to efficiently map IPv4 to IPv6","id":"hacking-the-go-compiler-to-efficiently-map-ipv4-to-ipv6"},{"level":1,"text":"The alternatives#","id":"the-alternatives"},{"level":2,"text":"Modifying the Go standard library#","id":"modifying-the-go-standard-library"},{"level":2,"text":"As a helper#","id":"as-a-helper"},{"level":2,"text":"As an unsafe function#","id":"as-an-unsafe-function"},{"level":2,"text":"Benchmarks#","id":"benchmarks"},{"level":2,"text":"Assembly code#","id":"assembly-code"},{"level":1,"text":"Hacking the Go compiler#","id":"hacking-the-go-compiler"},{"level":2,"text":"The hammer#","id":"the-hammer"},{"level":2,"text":"The screwdriver#","id":"the-screwdriver"},{"level":3,"text":"On paper#","id":"on-paper"},{"level":3,"text":"In practice#","id":"in-practice"},{"level":3,"text":"Out of order#","id":"out-of-order"},{"level":3,"text":"Back in order#","id":"back-in-order"},{"level":1,"text":"Next steps#","id":"next-steps"}]}}