{"article":{"slug":"hip-and-rocm-as-first-class-citizens","title":"HIP and ROCm as first-class citizens","subtitle":null,"summary":"Ludovic Courtès reports that Guix developers migrated the whole AMD HIP/ROCm stack from the Guix-HPC channels into Guix proper starting with 7.1.1: more than 40 packages covering the toolchain, linear algebra libraries, rocHPL, a ROCm Open MPI and profiling tools, plus HIP variants of scientific software.","content_type":"blog_post","language":"en","canonical_url":"https://hpc.guix.info/blog/2026/10/hip-and-rocm-as-first-class-citizens/","author":{"name":"Ludovic Courtès","url":null,"person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"Guix-HPC","url":"https://hpc.guix.info/","listing_slug":null,"listing":null},"topics":[{"name":"Open Source","slug":"open-source","url":"https://listedarticles.com/topics/open-source"},{"name":"Linux","slug":"linux","url":"https://listedarticles.com/topics/linux"},{"name":"Hardware","slug":"hardware","url":"https://listedarticles.com/topics/hardware"},{"name":"Performance","slug":"performance","url":"https://listedarticles.com/topics/performance"}],"about_listings":[],"cover_image_url":null,"license":"CC-BY-SA-4.0","word_count":1485,"reading_minutes":6,"published_at":"2026-10-09T00:00:00.000Z","added_at":"2026-10-09T14:41:14.817Z","updated_at":"2026-10-09T14:41:14.817Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":true},"profile_url":"https://listedarticles.com/articles/hip-and-rocm-as-first-class-citizens","markdown_url":"https://listedarticles.com/articles/hip-and-rocm-as-first-class-citizens.md","example":false,"citation":"Ludovic Courtès, Guix-HPC. \"HIP and ROCm as first-class citizens.\" 9 Oct 2026. https://hpc.guix.info/blog/2026/10/hip-and-rocm-as-first-class-citizens/ (CC-BY-SA-4.0)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://hpc.guix.info/blog/2026/10/hip-and-rocm-as-first-class-citizens/"},"body_markdown":"# HIP and ROCm as first-class citizens\n\nEarlier this year, [two years after AMD engineers contributed packages\nthe HIP/ROCm stack to the Guix-HPC\nchannels](https://hpc.guix.info/blog/2024/01/hip-and-rocm-come-to-guix/),\nGuix developers migrated the whole HIP/ROCm stack into Guix proper,\nstarting with version 7.1.1.  This was quite a milestone as it makes\nHIP/ROCm first-class citizens and gives them more exposure and better\nsupport in the community.\n\nThe package set consists of more than 40 packages covering many things:\n\n- [the toolchain\nitself](https://hpc.guix.info/package/rocm-toolchain) , which includes`hipcc` and related commands;\n- linear algebra libraries such as [`rocblas`](https://hpc.guix.info/package/rocblas) ,[`rocsparse`](https://hpc.guix.info/package/rocsparse) ,[`hipblas`](https://hpc.guix.info/package/hipblas) , and[`hipsparse`](https://hpc.guix.info/package/hipsparse) ;\n- the [rocHPL](https://hpc.guix.info/package/rochpl) benchmark;\n- a [ROCm-enabled variant of\nOpen MPI](https://hpc.guix.info/package/openmpi-rocm) ;\n- tools such as\n[`rocprofiler`](https://hpc.guix.info/package/rocprofiler) and[`roctracer`](https://hpc.guix.info/package/roctracer) .\n\nWe are also gradually adding HIP/ROCm variants of scientific software\nsuch as [CP2K](https://hpc.guix.info/package/cp2k-hip-rocm) and\n[Chameleon](https://hpc.guix.info/package/chameleon-hip-rocm), a dense\nlinear algebra solver developed at Inria.\n\n# Selecting target GPUs\n\nThe set of [AMD GPU architectures grows\nquickly](https://llvm.org/docs/AMDGPUUsage.html#amdgpu-processor-table).\nAs packagers, we choose a default set of target GPU architectures to\nbuild ROCm/HIP-enabled applications for, but that set of architectures\nmust be limited given the build time and size of resulting application\nbinaries.  It is crucial for users to be able to override this default\nset of target architectures to build specifically for the\narchitecture(s) they want.\n\nTo address that, we added a new package transformation option to Guix\n[called\n`--amd-gpu`](https://guix.gnu.org/manual/devel/en/html_node/Package-Transformation-Options.html#index-AMD-GPUs).\nJust like [`--tune` lets you build a package optimized for a specific\nCPU\nmicro-architecture](https://hpc.guix.info/blog/2022/01/tuning-packages-for-a-cpu-micro-architecture/),\n`--amd-gpu` creates, on the fly, a variant of the relevant packages\nbuilt specifically for the given GPU architecture(s).\n\nFor example, here is how you would run a variant of the [ROCm bandwidth\ntest](https://hpc.guix.info/package/rocm-bandwidth-test) built\nspecifically for AMD Instinct MI250 (`gfx90a`) and for AMD Instinct\nMI300 (`gfx942`):\n\n```\n$ srun --tasks-per-node=1 -N1 --exclusive … \\\n    guix shell rocm-bandwidth-test --amd-gpu=gfx90a,gfx942 -- \\\n    rocm-bandwidth-test plugin --run tb p2p\nTransferBench v1.64.00\n…\nBytes Per Direction 268435456\nUnidirectional copy peak bandwidth GB/s [Local read / Remote write] (GPU-Executor: GFX)\n SRC+EXE\\DST    CPU 00    CPU 01    CPU 02    CPU 03       GPU 00    GPU 01    GPU 02    GPU 03\n  CPU 00  ->     18.85     18.55     18.36     18.45        19.00     18.61     18.69     18.70\n  CPU 01  ->      7.87      8.82      8.42      8.63         7.99      8.16      7.40      7.64\n  CPU 02  ->      6.30      6.53      6.42      7.05         5.58      5.93      5.94      5.91\n  CPU 03  ->      6.34      7.09      7.50      7.35         6.17      6.23      6.37      6.23\n  GPU 00  ->   1307.15     90.22     90.85     91.45      1360.70     91.24     91.39     92.14\n  GPU 01  ->     90.58   1371.25     92.78     90.63        91.08   1431.93     92.38     90.84\n  GPU 02  ->     91.12     92.57   1355.95     91.82        90.68     92.22   1407.58     91.52\n  GPU 03  ->     91.41     91.08     91.66   1348.37        91.99     90.92     91.88   1468.65\n                           CPU->CPU  CPU->GPU  GPU->CPU  GPU->GPU\nAverages (During UniDir):     10.09      9.66    404.93     91.52\n…\n```\n(This particular run was on a node with MI300 GPUs.)\n\nOf course these GPU architecture identifiers are, well, hard to grasp.\nYou can find the full list [in the LLVM\ndocumentation](https://llvm.org/docs/AMDGPUUsage.html#amdgpu-processor-table);\nshould you make a typo or select an architecture that the toolchain at\nhand doesn’t support, Guix lets you know about it without going any\nfurther:\n\n```\n$ guix build rocm-bandwidth-test --amd-gpu=forgot-the-name \ngnu/packages/llvm.scm:2283:2: error: compiler rocm-toolchain@7.1.1 does not support AMD GPU target forgot-the-name\nhint: Compiler rocm-toolchain@7.1.1 supports the following AMD GPU targets:\n     bonaire, carrizo, fiji, generic, generic-hsa, gfx10-1-generic, gfx10-3-generic,\n     gfx1010, gfx1011, gfx1012, gfx1013, gfx1030, gfx1031, gfx1032, gfx1033, gfx1034,\n     gfx1035, gfx1036, gfx11-generic, gfx1100, gfx1101, gfx1102, gfx1103, gfx1150, gfx1151,\n     gfx1152, gfx1153, gfx12-generic, gfx1200, gfx1201, gfx600, gfx601, gfx602, gfx700,\n     gfx701, gfx702, gfx703, gfx704, gfx705, gfx801, gfx802, gfx803, gfx805, gfx810,\n     gfx9-4-generic, gfx9-generic, gfx900, gfx902, gfx904, gfx906, gfx908, gfx909, gfx90a,\n     gfx90c, gfx940, gfx941, gfx942, gfx950, hainan, hawaii, iceland, kabini, kaveri,\n     mullins, oland, pitcairn, polaris10, polaris11, stoney, tahiti, tonga, tongapro, verde\n```\n# Performance\n\nSo far, microbencharks are giving us green lights.  Let’s take a closer\nlook at benchmarks we ran on nodes with 8 MI250X GPUs of the [Adastra\nsupercomputer](https://dci.dci-gitlab.cines.fr/webextranet/architecture/index.html).\n\n## ROCm support for Open MPI\n\nFirst, there’s the bandwidth measured for MPI transfers among Graphics\nCompute Dies (GCDs)—specifically, using the ROCm-enabled Open MPI\npackage, [`openmpi-rocm`](https://hpc.guix.info/package/openmpi-rocm).\nFor this, we run the [OSU\nMicro-Benchmarks](https://hpc.guix.info/package/osu-micro-benchmarks-rocm)\nlinked against `openmpi-rocm`, asking it to measure device-to-device\ntransfers; we do that with large messages (16 MiB) and for all GCD pairs\nwithin a node, where each node has 8 GCDs:\n\n```\nexport HSA_ENABLE_SDMA=0\nfor first in $(seq 0 7)\ndo\n    for second in $(seq $(($first + 1)) 7)\n    do\n        export HIP_VISIBLE_DEVICES=\"$first,$second\"\n        echo \"# HIP_VISIBLE_DEVICES: $HIP_VISIBLE_DEVICES\"\n        guix time-machine -q --commit=f5c2937cddd4c8427f15b8b711f8a82211a09407 -- \\\n          shell openmpi-rocm osu-micro-benchmarks-rocm -- \\\n          mpirun -n 2 --mca pml ucx \\\n          osu_bw -m $((16*1024*1024)):$((16*1024*1024)) D D\n    done\ndone\n```\nSome explanations:\n\n- The `time-machine` part selects the commit, and thus*the entire\nsoftware stack* , that we have tested—fewer moving pieces.  The`shell` bit specifies the packages we need in our environment.  The\ncommit we selected here provides a stack with Open MPI 5.0.10 and\nHIP/ROCm 7.1.1.\n- Setting `HSA_ENABLE_SDMA=0` , which turns off use of System Direct\nMemory Access (SDMA) by the HIP/ROCm runtime,[gives higher\nthroughput](https://gpuopen.com/learn/amd-lab-notes/amd-lab-notes-gpu-aware-mpi-readme/) .\n- We create two MPI processes on the node (`mpirun -n 2` ).  The`--mca pml ucx` flag ensures Open MPI selects[ucx](https://hpc.guix.info/package/ucx-rocm) as its interconnect\nbackend.\n- Last, we run `osu_bw` , the bandwidth benchmark, for device-to-device\n(`D D` ) transfers with messages of 16 MiB.  Setting the`HIP_VISIBLE_DEVICES` right above allows us to ensure transfers are\nmade between these two GCDs.\n\nThis gives us the communication matrix below, showing the bandwidth for unidirectional copies from device to device:\n\nThe bandwidth we observe between each pair of GCDs matches [the GCD\ntopology and Infinity Fabric\nlinks](https://gpuopen.com/learn/amd-lab-notes/amd-lab-notes-gpu-aware-mpi-readme/#gpu-to-gpu-communication-options-);\nfor example, peak bandwidth between GCD 0 and GCD 1 is roughly four\ntimes that between GCD 0 and GCD 2, and two times that between GCD 0 and\nGCD 6.\n\n## Computing benchmark\n\nWhat about computing throughput?  A good test is\n[rocHPL](https://hpc.guix.info/package/rochpl), the ROCm-enabled variant\nof the classical [high-performance\nLINPACK](<https://en.wikipedia.org/wiki/HPL_(benchmark)>).\n\nFor double-precision (aka. “Binary64”) floating point operations, the theoretical peak performance is 23.93 TFlop/s; for nodes with 8 GCDs, we can thus expect at most 23.93 x 8 = 191.5 TFlop/s per node.\n\nTo get as close as possible to peak performance, we must arrange to let\nrocHPL work on a matrix that occupies almost all the GCD memory, which\nis 512 GiB per node here.  For a single 8-GCD node, a matrix of 256,000\nrows and columns fills 95% of device memory, making it a good choice—in\nline with what the [rocHPL\nwiki](https://github.com/ROCm/rocHPL/wiki/Common-rocHPL-run-configurations)\nsuggests.\n\nWe can run it with an incantation along these lines:\n\n```\nCOMMIT=bf5d83139d8d4d7aa2d737b1e13620941abff2a1\n# Workaround until <https://codeberg.org/guix/guix/pulls/11429>\n# is merged.\nexport ROCM_SMI_LIB_PATH=\"$(readlink -f $(guix time-machine \\\n  -q --commit=$COMMIT -- \\\n  build rocm-smi-lib | grep -v -e -bin$)/lib/librocm_smi64.so)\"\nguix time-machine -q --commit=$COMMIT -- \\\n     shell rochpl openmpi-rocm --tune=znver3 -- \\\n     sh -x mpirun_rochpl -P 2 -Q 4 -N 256000 --NB 512\n```\nExplanations:\n\n- The `time-machine` bit once again allows us to pin Guix to the\ncommit for which we’ve run this benchmark, while`shell` sets up the\nexecution environment.  For good measure, we use`--tune` to[tune\nCPU code for the micro-architecture we have at\nhand](https://hpc.guix.info/blog/2022/01/tuning-packages-for-a-cpu-micro-architecture/) .\n- `-P` and`-Q` specify the number of rows and columns of the MPI\ngrid; the product corresponds to the number of GCDs on the node.`-N` specifies the matrix size, as discussed above.\n\nThat gives us a throughput of 160 TFlop/s—below the theoretical peak, but to our knowledge comparable to what others observe on MI250X.\n\n# Future work\n\nThis post gives an overview of where the HIP/ROCm stack is in Guix and\nhow its performance can be validated.  There are a number of things we\nare planning to do, starting with a [minor-version upgrade of the\nHIP/ROCm stack](https://codeberg.org/guix/guix/pulls/10257) before we\ndive into more recent versions.\n\nMore importantly, we are working on consolidation the set of benchmarks\nwe want to use to validate the stack.  Ideally, we won’t limit ourselves\nto micro-benchmarks and instead look at scientific applications that are\nknown to exercise more of the supercomputer capabilities—such as GROMACS\nor CP2K.  Our goal would be able to run a set of benchmarks before any\npackage upgrade in the ROCm and MPI stacks, drawing from what admins [at\nInria](https://jdev26.sciencesconf.org/726473) and [at\nCINES](https://jdev26.sciencesconf.org/720413) have been doing for their\nown clusters.\n\nAll this is a much broader endeavor. To be continued!\n\n# Acknowledgments\n\nThese benchmarks and additional tests were performed on the Adastra supercomputer hosted by CINES. Many thanks to our colleagues at CINES and to Florent Pruvost at Inria for their help. Huge thanks to the people who contributed to the HIP/ROCm stack in Guix, in particular to David Elsing for handling the bulk of the migration from the Guix-HPC channel and for upgrading those packages.\n\nUnless otherwise stated, blog posts on this site are\ncopyrighted by their respective authors and published under the terms of\nthe [CC-BY-SA 4.0](https://creativecommons.org/licenses/by-sa/4.0/) license and those of the [GNU Free Documentation License](https://www.gnu.org/licenses/fdl-1.3.html) (version 1.3 or later, with no Invariant Sections, no\nFront-Cover Texts, and no Back-Cover Texts).\n","body_html":"<h1 id=\"hip-and-rocm-as-first-class-citizens\">HIP and ROCm as first-class citizens</h1>\n<p>Earlier this year, <a href=\"https://hpc.guix.info/blog/2024/01/hip-and-rocm-come-to-guix/\" rel=\"nofollow ugc noopener\">two years after AMD engineers contributed packages\nthe HIP/ROCm stack to the Guix-HPC\nchannels</a>,\nGuix developers migrated the whole HIP/ROCm stack into Guix proper,\nstarting with version 7.1.1.  This was quite a milestone as it makes\nHIP/ROCm first-class citizens and gives them more exposure and better\nsupport in the community.</p>\n<p>The package set consists of more than 40 packages covering many things:</p>\n<ul><li><p>[the toolchain</p><p>itself](<a href=\"https://hpc.guix.info/package/rocm-toolchain\" rel=\"nofollow ugc noopener\">https://hpc.guix.info/package/rocm-toolchain</a>) , which includes<code>hipcc</code> and related commands;</p></li><li>linear algebra libraries such as <a href=\"https://hpc.guix.info/package/rocblas\" rel=\"nofollow ugc noopener\"><code>rocblas</code></a> ,<a href=\"https://hpc.guix.info/package/rocsparse\" rel=\"nofollow ugc noopener\"><code>rocsparse</code></a> ,<a href=\"https://hpc.guix.info/package/hipblas\" rel=\"nofollow ugc noopener\"><code>hipblas</code></a> , and<a href=\"https://hpc.guix.info/package/hipsparse\" rel=\"nofollow ugc noopener\"><code>hipsparse</code></a> ;</li><li>the <a href=\"https://hpc.guix.info/package/rochpl\" rel=\"nofollow ugc noopener\">rocHPL</a> benchmark;</li><li><p>a [ROCm-enabled variant of</p><p>Open MPI](<a href=\"https://hpc.guix.info/package/openmpi-rocm\" rel=\"nofollow ugc noopener\">https://hpc.guix.info/package/openmpi-rocm</a>) ;</p></li><li><p>tools such as</p><p><a href=\"https://hpc.guix.info/package/rocprofiler\" rel=\"nofollow ugc noopener\"><code>rocprofiler</code></a> and<a href=\"https://hpc.guix.info/package/roctracer\" rel=\"nofollow ugc noopener\"><code>roctracer</code></a> .</p></li></ul>\n<p>We are also gradually adding HIP/ROCm variants of scientific software\nsuch as <a href=\"https://hpc.guix.info/package/cp2k-hip-rocm\" rel=\"nofollow ugc noopener\">CP2K</a> and\n<a href=\"https://hpc.guix.info/package/chameleon-hip-rocm\" rel=\"nofollow ugc noopener\">Chameleon</a>, a dense\nlinear algebra solver developed at Inria.</p>\n<h1 id=\"selecting-target-gpus\">Selecting target GPUs</h1>\n<p>The set of <a href=\"https://llvm.org/docs/AMDGPUUsage.html#amdgpu-processor-table\" rel=\"nofollow ugc noopener\">AMD GPU architectures grows\nquickly</a>.\nAs packagers, we choose a default set of target GPU architectures to\nbuild ROCm/HIP-enabled applications for, but that set of architectures\nmust be limited given the build time and size of resulting application\nbinaries.  It is crucial for users to be able to override this default\nset of target architectures to build specifically for the\narchitecture(s) they want.</p>\n<p>To address that, we added a new package transformation option to Guix\n<a href=\"https://guix.gnu.org/manual/devel/en/html_node/Package-Transformation-Options.html#index-AMD-GPUs\" rel=\"nofollow ugc noopener\">called\n<code>--amd-gpu</code></a>.\nJust like <a href=\"https://hpc.guix.info/blog/2022/01/tuning-packages-for-a-cpu-micro-architecture/\" rel=\"nofollow ugc noopener\"><code>--tune</code> lets you build a package optimized for a specific\nCPU\nmicro-architecture</a>,\n<code>--amd-gpu</code> creates, on the fly, a variant of the relevant packages\nbuilt specifically for the given GPU architecture(s).</p>\n<p>For example, here is how you would run a variant of the <a href=\"https://hpc.guix.info/package/rocm-bandwidth-test\" rel=\"nofollow ugc noopener\">ROCm bandwidth\ntest</a> built\nspecifically for AMD Instinct MI250 (<code>gfx90a</code>) and for AMD Instinct\nMI300 (<code>gfx942</code>):</p>\n<pre><code>$ srun --tasks-per-node=1 -N1 --exclusive … \\\n    guix shell rocm-bandwidth-test --amd-gpu=gfx90a,gfx942 -- \\\n    rocm-bandwidth-test plugin --run tb p2p\nTransferBench v1.64.00\n…\nBytes Per Direction 268435456\nUnidirectional copy peak bandwidth GB/s [Local read / Remote write] (GPU-Executor: GFX)\n SRC+EXE\\DST    CPU 00    CPU 01    CPU 02    CPU 03       GPU 00    GPU 01    GPU 02    GPU 03\n  CPU 00  -&gt;     18.85     18.55     18.36     18.45        19.00     18.61     18.69     18.70\n  CPU 01  -&gt;      7.87      8.82      8.42      8.63         7.99      8.16      7.40      7.64\n  CPU 02  -&gt;      6.30      6.53      6.42      7.05         5.58      5.93      5.94      5.91\n  CPU 03  -&gt;      6.34      7.09      7.50      7.35         6.17      6.23      6.37      6.23\n  GPU 00  -&gt;   1307.15     90.22     90.85     91.45      1360.70     91.24     91.39     92.14\n  GPU 01  -&gt;     90.58   1371.25     92.78     90.63        91.08   1431.93     92.38     90.84\n  GPU 02  -&gt;     91.12     92.57   1355.95     91.82        90.68     92.22   1407.58     91.52\n  GPU 03  -&gt;     91.41     91.08     91.66   1348.37        91.99     90.92     91.88   1468.65\n                           CPU-&gt;CPU  CPU-&gt;GPU  GPU-&gt;CPU  GPU-&gt;GPU\nAverages (During UniDir):     10.09      9.66    404.93     91.52\n…</code></pre>\n<p>(This particular run was on a node with MI300 GPUs.)</p>\n<p>Of course these GPU architecture identifiers are, well, hard to grasp.\nYou can find the full list <a href=\"https://llvm.org/docs/AMDGPUUsage.html#amdgpu-processor-table\" rel=\"nofollow ugc noopener\">in the LLVM\ndocumentation</a>;\nshould you make a typo or select an architecture that the toolchain at\nhand doesn’t support, Guix lets you know about it without going any\nfurther:</p>\n<pre><code>$ guix build rocm-bandwidth-test --amd-gpu=forgot-the-name \ngnu/packages/llvm.scm:2283:2: error: compiler rocm-toolchain@7.1.1 does not support AMD GPU target forgot-the-name\nhint: Compiler rocm-toolchain@7.1.1 supports the following AMD GPU targets:\n     bonaire, carrizo, fiji, generic, generic-hsa, gfx10-1-generic, gfx10-3-generic,\n     gfx1010, gfx1011, gfx1012, gfx1013, gfx1030, gfx1031, gfx1032, gfx1033, gfx1034,\n     gfx1035, gfx1036, gfx11-generic, gfx1100, gfx1101, gfx1102, gfx1103, gfx1150, gfx1151,\n     gfx1152, gfx1153, gfx12-generic, gfx1200, gfx1201, gfx600, gfx601, gfx602, gfx700,\n     gfx701, gfx702, gfx703, gfx704, gfx705, gfx801, gfx802, gfx803, gfx805, gfx810,\n     gfx9-4-generic, gfx9-generic, gfx900, gfx902, gfx904, gfx906, gfx908, gfx909, gfx90a,\n     gfx90c, gfx940, gfx941, gfx942, gfx950, hainan, hawaii, iceland, kabini, kaveri,\n     mullins, oland, pitcairn, polaris10, polaris11, stoney, tahiti, tonga, tongapro, verde</code></pre>\n<h1 id=\"performance\">Performance</h1>\n<p>So far, microbencharks are giving us green lights.  Let’s take a closer\nlook at benchmarks we ran on nodes with 8 MI250X GPUs of the <a href=\"https://dci.dci-gitlab.cines.fr/webextranet/architecture/index.html\" rel=\"nofollow ugc noopener\">Adastra\nsupercomputer</a>.</p>\n<h2 id=\"rocm-support-for-open-mpi\">ROCm support for Open MPI</h2>\n<p>First, there’s the bandwidth measured for MPI transfers among Graphics\nCompute Dies (GCDs)—specifically, using the ROCm-enabled Open MPI\npackage, <a href=\"https://hpc.guix.info/package/openmpi-rocm\" rel=\"nofollow ugc noopener\"><code>openmpi-rocm</code></a>.\nFor this, we run the <a href=\"https://hpc.guix.info/package/osu-micro-benchmarks-rocm\" rel=\"nofollow ugc noopener\">OSU\nMicro-Benchmarks</a>\nlinked against <code>openmpi-rocm</code>, asking it to measure device-to-device\ntransfers; we do that with large messages (16 MiB) and for all GCD pairs\nwithin a node, where each node has 8 GCDs:</p>\n<pre><code>export HSA_ENABLE_SDMA=0\nfor first in $(seq 0 7)\ndo\n    for second in $(seq $(($first + 1)) 7)\n    do\n        export HIP_VISIBLE_DEVICES=&quot;$first,$second&quot;\n        echo &quot;# HIP_VISIBLE_DEVICES: $HIP_VISIBLE_DEVICES&quot;\n        guix time-machine -q --commit=f5c2937cddd4c8427f15b8b711f8a82211a09407 -- \\\n          shell openmpi-rocm osu-micro-benchmarks-rocm -- \\\n          mpirun -n 2 --mca pml ucx \\\n          osu_bw -m $((16*1024*1024)):$((16*1024*1024)) D D\n    done\ndone</code></pre>\n<p>Some explanations:</p>\n<ul><li><p>The <code>time-machine</code> part selects the commit, and thus*the entire</p><p>software stack* , that we have tested—fewer moving pieces.  The<code>shell</code> bit specifies the packages we need in our environment.  The\ncommit we selected here provides a stack with Open MPI 5.0.10 and\nHIP/ROCm 7.1.1.</p></li><li><p>Setting <code>HSA_ENABLE_SDMA=0</code> , which turns off use of System Direct</p><p>Memory Access (SDMA) by the HIP/ROCm runtime,<a href=\"https://gpuopen.com/learn/amd-lab-notes/amd-lab-notes-gpu-aware-mpi-readme/\" rel=\"nofollow ugc noopener\">gives higher\nthroughput</a> .</p></li><li><p>We create two MPI processes on the node (<code>mpirun -n 2</code> ).  The<code>--mca pml ucx</code> flag ensures Open MPI selects<a href=\"https://hpc.guix.info/package/ucx-rocm\" rel=\"nofollow ugc noopener\">ucx</a> as its interconnect</p><p>backend.</p></li><li><p>Last, we run <code>osu_bw</code> , the bandwidth benchmark, for device-to-device</p><p>(<code>D D</code> ) transfers with messages of 16 MiB.  Setting the<code>HIP_VISIBLE_DEVICES</code> right above allows us to ensure transfers are\nmade between these two GCDs.</p></li></ul>\n<p>This gives us the communication matrix below, showing the bandwidth for unidirectional copies from device to device:</p>\n<p>The bandwidth we observe between each pair of GCDs matches <a href=\"https://gpuopen.com/learn/amd-lab-notes/amd-lab-notes-gpu-aware-mpi-readme/#gpu-to-gpu-communication-options-\" rel=\"nofollow ugc noopener\">the GCD\ntopology and Infinity Fabric\nlinks</a>;\nfor example, peak bandwidth between GCD 0 and GCD 1 is roughly four\ntimes that between GCD 0 and GCD 2, and two times that between GCD 0 and\nGCD 6.</p>\n<h2 id=\"computing-benchmark\">Computing benchmark</h2>\n<p>What about computing throughput?  A good test is\n<a href=\"https://hpc.guix.info/package/rochpl\" rel=\"nofollow ugc noopener\">rocHPL</a>, the ROCm-enabled variant\nof the classical <a href=\"https://en.wikipedia.org/wiki/HPL_(benchmark)\" rel=\"nofollow ugc noopener\">high-performance\nLINPACK</a>.</p>\n<p>For double-precision (aka. “Binary64”) floating point operations, the theoretical peak performance is 23.93 TFlop/s; for nodes with 8 GCDs, we can thus expect at most 23.93 x 8 = 191.5 TFlop/s per node.</p>\n<p>To get as close as possible to peak performance, we must arrange to let\nrocHPL work on a matrix that occupies almost all the GCD memory, which\nis 512 GiB per node here.  For a single 8-GCD node, a matrix of 256,000\nrows and columns fills 95% of device memory, making it a good choice—in\nline with what the <a href=\"https://github.com/ROCm/rocHPL/wiki/Common-rocHPL-run-configurations\" rel=\"nofollow ugc noopener\">rocHPL\nwiki</a>\nsuggests.</p>\n<p>We can run it with an incantation along these lines:</p>\n<pre><code>COMMIT=bf5d83139d8d4d7aa2d737b1e13620941abff2a1\n# Workaround until &lt;https://codeberg.org/guix/guix/pulls/11429&gt;\n# is merged.\nexport ROCM_SMI_LIB_PATH=&quot;$(readlink -f $(guix time-machine \\\n  -q --commit=$COMMIT -- \\\n  build rocm-smi-lib | grep -v -e -bin$)/lib/librocm_smi64.so)&quot;\nguix time-machine -q --commit=$COMMIT -- \\\n     shell rochpl openmpi-rocm --tune=znver3 -- \\\n     sh -x mpirun_rochpl -P 2 -Q 4 -N 256000 --NB 512</code></pre>\n<p>Explanations:</p>\n<ul><li><p>The <code>time-machine</code> bit once again allows us to pin Guix to the</p><p>commit for which we’ve run this benchmark, while<code>shell</code> sets up the\nexecution environment.  For good measure, we use<code>--tune</code> to<a href=\"https://hpc.guix.info/blog/2022/01/tuning-packages-for-a-cpu-micro-architecture/\" rel=\"nofollow ugc noopener\">tune\nCPU code for the micro-architecture we have at\nhand</a> .</p></li><li><p><code>-P</code> and<code>-Q</code> specify the number of rows and columns of the MPI</p><p>grid; the product corresponds to the number of GCDs on the node.<code>-N</code> specifies the matrix size, as discussed above.</p></li></ul>\n<p>That gives us a throughput of 160 TFlop/s—below the theoretical peak, but to our knowledge comparable to what others observe on MI250X.</p>\n<h1 id=\"future-work\">Future work</h1>\n<p>This post gives an overview of where the HIP/ROCm stack is in Guix and\nhow its performance can be validated.  There are a number of things we\nare planning to do, starting with a <a href=\"https://codeberg.org/guix/guix/pulls/10257\" rel=\"nofollow ugc noopener\">minor-version upgrade of the\nHIP/ROCm stack</a> before we\ndive into more recent versions.</p>\n<p>More importantly, we are working on consolidation the set of benchmarks\nwe want to use to validate the stack.  Ideally, we won’t limit ourselves\nto micro-benchmarks and instead look at scientific applications that are\nknown to exercise more of the supercomputer capabilities—such as GROMACS\nor CP2K.  Our goal would be able to run a set of benchmarks before any\npackage upgrade in the ROCm and MPI stacks, drawing from what admins <a href=\"https://jdev26.sciencesconf.org/726473\" rel=\"nofollow ugc noopener\">at\nInria</a> and <a href=\"https://jdev26.sciencesconf.org/720413\" rel=\"nofollow ugc noopener\">at\nCINES</a> have been doing for their\nown clusters.</p>\n<p>All this is a much broader endeavor. To be continued!</p>\n<h1 id=\"acknowledgments\">Acknowledgments</h1>\n<p>These benchmarks and additional tests were performed on the Adastra supercomputer hosted by CINES. Many thanks to our colleagues at CINES and to Florent Pruvost at Inria for their help. Huge thanks to the people who contributed to the HIP/ROCm stack in Guix, in particular to David Elsing for handling the bulk of the migration from the Guix-HPC channel and for upgrading those packages.</p>\n<p>Unless otherwise stated, blog posts on this site are\ncopyrighted by their respective authors and published under the terms of\nthe <a href=\"https://creativecommons.org/licenses/by-sa/4.0/\" rel=\"nofollow ugc noopener\">CC-BY-SA 4.0</a> license and those of the <a href=\"https://www.gnu.org/licenses/fdl-1.3.html\" rel=\"nofollow ugc noopener\">GNU Free Documentation License</a> (version 1.3 or later, with no Invariant Sections, no\nFront-Cover Texts, and no Back-Cover Texts).</p>","headings":[{"level":1,"text":"HIP and ROCm as first-class citizens","id":"hip-and-rocm-as-first-class-citizens"},{"level":1,"text":"Selecting target GPUs","id":"selecting-target-gpus"},{"level":1,"text":"Performance","id":"performance"},{"level":2,"text":"ROCm support for Open MPI","id":"rocm-support-for-open-mpi"},{"level":2,"text":"Computing benchmark","id":"computing-benchmark"},{"level":1,"text":"Future work","id":"future-work"},{"level":1,"text":"Acknowledgments","id":"acknowledgments"}]}}