{"article":{"slug":"janet-on-x32-32-bit-pointers-64-bit-speed-25-less-ram","title":"Janet on x32: 32-bit Pointers, 64-bit Speed, 25% Less RAM","subtitle":null,"summary":"Alex Alejandre tests whether compiling the Janet language for 32-bit pointers saves memory: an -m32 build on Arch cut memory about 20% but ran 50% slower, while the x32 ABI on an old Ubuntu server averaged about 25% less RAM with mixed speed changes, not enough to justify a broader call to action.","content_type":"tutorial","language":"en","canonical_url":"https://alexalejandre.com/programming/lisp/janet-for-the-x32-abi/","author":{"name":"Alex Alejandre","url":null,"person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"Alex Alejandre","url":"https://alexalejandre.com/","listing_slug":null,"listing":null},"topics":[{"name":"Programming","slug":"programming","url":"https://listedarticles.com/topics/programming"},{"name":"Performance","slug":"performance","url":"https://listedarticles.com/topics/performance"},{"name":"Linux","slug":"linux","url":"https://listedarticles.com/topics/linux"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":2564,"reading_minutes":11,"published_at":"2026-10-06T20:50:05.000Z","added_at":"2026-10-06T23:15:12.920Z","updated_at":"2026-10-06T23:15:12.920Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":true},"profile_url":"https://listedarticles.com/articles/janet-on-x32-32-bit-pointers-64-bit-speed-25-less-ram","markdown_url":"https://listedarticles.com/articles/janet-on-x32-32-bit-pointers-64-bit-speed-25-less-ram.md","example":false,"citation":"Alex Alejandre, Alex Alejandre. \"Janet on x32: 32-bit Pointers, 64-bit Speed, 25% Less RAM.\" 6 Oct 2026. https://alexalejandre.com/programming/lisp/janet-for-the-x32-abi/ (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://alexalejandre.com/programming/lisp/janet-for-the-x32-abi/"},"body_markdown":"# Janet on x32: 32-bit Pointers, 64-bit Speed, 25% Less RAM\n\nIt’s possible to save a dramatic amount of memory (close to half for programs with pointer heavy heaps) and also speed up programs by a modest amount (by way of more data fitting in cache) by using 32 bit pointers on 64 bit systems.\n\nThe Linux x32 ABI lets you do just this, but it’s tragically underused and overlooked. It’s disabled by default on Debian (though you can enable it with a boot flag) and not even compiled in on Arch Linux. Most software compiles just fine for x32, but there isn’t much packaging for it so you do have to compile nearly everything yourself.\n\nI experimented with deploying mastodon on x32 and it cut the app’s memory usage from 650mb to 350mb. There’s so much potential here, but the work is mostly of the thankless coordination and communication type and I don’t have the time or motivation to push it forward myself. - [Hailey](https://lobste.rs/s/gns27z/what_are_your_programming_hunches_you#c_tw3vok)\n\n\n### How to Compile for 32 Bits?\n\nJanet libraries like Spork supply their own build flags, but we can hijack them with a fake `cc` which starts the real one with the necessary `-m32 -msse2 -mfpmath=sse` flags. I modified the beginning of my default [Janet build script](https://codeberg.org/veqq/install-latest-janet/src/branch/master/install-latest-janet.sh):\n\n```\n#!/bin/sh\nset -eu\nTARGET=\"$PWD/janet32\"\nJANET=\"$TARGET/bin/janet\"\nmkdir -p \"$TARGET/cc32\"\nprintf '#!/bin/sh\\nexec gcc -m32 -msse2 -mfpmath=sse \"$@\"\\n' > \"$TARGET/cc32/cc\"\nchmod +x \"$TARGET/cc32/cc\"\nexport PATH=\"$TARGET/cc32:$PATH\"\nexport JANET_TOOLCHAIN=cc\nunset JANET_PATH JANET_TREE\nmkdir -p \"$TARGET\"\ncd \"$TARGET\"\nrm -rf build/spork janet\ngit clone https://github.com/janet-lang/janet\ncd janet\n```\nEasy enough! Now let’s try it!\n\n### Failing on Arch\n\ntl;dr: 20% memory reduction, 50% slower\n\n## See the Scripts and Results\n\nThis is on cachyos (arch, btw) with lib32-glibc lib32-gcc-libs. Adding this to a build script forces libraires like spork to build at 32 bit too!\n\nFor `mem.janet`:\n\n```\n(let [live @[]]\n  (for i 0 2_000_000\n    (array/push live [i (+ i 1)]))\n  (gccollect)\n  (print (length live)))\n```\nNote, these use Fish and I don’t want to rebuild my site with other language support:\n\n```\nsudo pacman -S time # different from just time\ncommand time -v ./bin/janet mem.janet # command skips the fish version\ncommand time -v ./janet32/bin/janet mem.janet\n2000000\n       Command being timed: \"./bin/janet mem.janet\"\n       User time (seconds): 0.24\n       System time (seconds): 0.05\n       Percent of CPU this job got: 100%\n       Elapsed (wall clock) time (h:mm:ss or m:ss): 0:00.30\n       Average shared text size (kbytes): 0\n       Average unshared data size (kbytes): 0\n       Average stack size (kbytes): 0\n       Average total size (kbytes): 0\n       Maximum resident set size (kbytes): 145196\n       Average resident set size (kbytes): 0\n       Major (requiring I/O) page faults: 0\n       Minor (reclaiming a frame) page faults: 33513\n       Voluntary context switches: 1\n       Involuntary context switches: 9\n2000000\n       Command being timed: \"./janet32/bin/janet mem.janet\"\n       User time (seconds): 0.47\n       System time (seconds): 0.03\n       Percent of CPU this job got: 99%\n       Elapsed (wall clock) time (h:mm:ss or m:ss): 0:00.50\n       Average shared text size (kbytes): 0\n       Average unshared data size (kbytes): 0\n       Average stack size (kbytes): 0\n       Average total size (kbytes): 0\n       Maximum resident set size (kbytes): 113716\n       Average resident set size (kbytes): 0\n       Major (requiring I/O) page faults: 0\n       Minor (reclaiming a frame) page faults: 25686\n       Voluntary context switches: 1\n       Involuntary context switches: 21\n```\nI added a 0 in the test\n\n```\ncommand time -v ./bin/janet mem.janet\n                  command time -v ./janet32/bin/janet mem.janet\n20000000\n        Command being timed: \"./bin/janet mem.janet\"\n        User time (seconds): 2.83\n        System time (seconds): 0.46\n        Percent of CPU this job got: 99%\n        Elapsed (wall clock) time (h:mm:ss or m:ss): 0:03.31\n        Average shared text size (kbytes): 0\n        Average unshared data size (kbytes): 0\n        Average stack size (kbytes): 0\n        Average total size (kbytes): 0\n        Maximum resident set size (kbytes): 1410668\n        Average resident set size (kbytes): 0\n        Major (requiring I/O) page faults: 0\n        Minor (reclaiming a frame) page faults: 316820\n        Voluntary context switches: 1\n        Involuntary context switches: 164\n        Exit status: 0\n20000000\n        Command being timed: \"./janet32/bin/janet mem.janet\"\n        User time (seconds): 4.64\n        System time (seconds): 0.35\n        Percent of CPU this job got: 99%\n        Elapsed (wall clock) time (h:mm:ss or m:ss): 0:05.00\n        Average shared text size (kbytes): 0\n        Average unshared data size (kbytes): 0\n        Average stack size (kbytes): 0\n        Average total size (kbytes): 0\n        Maximum resident set size (kbytes): 1098372\n        Average resident set size (kbytes): 0\n        Major (requiring I/O) page faults: 0\n        Minor (reclaiming a frame) page faults: 238633\n        Voluntary context switches: 1\n        Involuntary context switches: 193\n```\nFor e.g. `declarative-dsls/tests.janet`\n\n```\n98 passed\n        Command being timed: \"./bin/janet ./lib/janet/bin/judge declarative-dsls/tests.janet\"\n        User time (seconds): 0.42\n        System time (seconds): 0.02\n        Percent of CPU this job got: 100%\n        Elapsed (wall clock) time (h:mm:ss or m:ss): 0:00.45\n        Average shared text size (kbytes): 0\n        Average unshared data size (kbytes): 0\n        Average stack size (kbytes): 0\n        Average total size (kbytes): 0\n        Maximum resident set size (kbytes): 61900\ndeclarative-dsls/tests.janet\n98 passed\n        Command being timed: \"./bin/janet ./lib/janet/bin/judge declarative-dsls/tests.janet\"\n        User time (seconds): 0.70\n        System time (seconds): 0.02\n        Percent of CPU this job got: 99%\n        Elapsed (wall clock) time (h:mm:ss or m:ss): 0:00.72\n        Average shared text size (kbytes): 0\n        Average unshared data size (kbytes): 0\n        Average stack size (kbytes): 0\n        Average total size (kbytes): 0\n        Maximum resident set size (kbytes): 54108\n        Average resident set size (kbytes): 0\n        Major (requiring I/O) page faults: 0\n        Minor (reclaiming a frame) page faults: 12735\n```\nAs Hailey said, Arch does not come with the [x32 ABI](https://en.wikipedia.org/wiki/X32_ABI) compiled, so I had to use `-m32` and while we do get a 20% memory savings, we execute 50% slower because half the registers aren’t used! `-m32` is the old 32-bit x86 mode using only 8 registers. To see the speed-up, we need to use `-mx32` which uses the 64-bit instruction set with shrunken 32-bit pointers. But I don’t want to recompile Linux…\n\nUnfortunately, Ubuntu also dropped x32 support.\n\n### Winning on Deprecated Ubuntu\n\nLuckily, lazily, I have access to some deprecated servers (*of course* not in prod…) and OS installs, wasting hard disk space:\n\nUbuntu 20.04 LTS Focal Fossa has reached its end of standard support on 31 May 2025.\n\n\nUbuntu, so we have bash again:\n\n```\ngrep X86_X32 /boot/config-$(uname -r)  \nCONFIG_X86_X32=y\n```\nJust what we need, although `gcc version 9.4.0` is a bit old. We already have `build-essential git gcc-multilib libc6-dev-x32`. We’re almost ready to test this hunch. But first, we must change `src/include/janet.h`:\n\n```\n#if ((defined(__x86_64__) || defined(_M_X64)) \\\n     && (defined(JANET_POSIX) || defined(JANET_WINDOWS))) \\\n```\nbecomes:\n\n```\n#if ((defined(__x86_64__) || defined(_M_X64)) && !defined(__ILP32__) \\\n     && (defined(JANET_POSIX) || defined(JANET_WINDOWS))) \\\n```\nlest Janet choose the 64-bit value layout. Our build script will also pass a build flag to disable Janet’s FFI. Here is our full `32janet.sh`:\n\n```\n#!/bin/bash\nset -eu\nTARGET=\"$PWD/janetx32\"\nJANET=\"$TARGET/bin/janet\"\nmkdir -p \"$TARGET/cc32\"\n# this forces Janet C libraries like Spork to build with 32bit also\n# FFI build test fails on -mx32 build so -DJANET_NO_FFI\nprintf '#!/bin/sh\\nexec gcc -mx32 -DJANET_NO_FFI \"$@\"\\n' > \"$TARGET/cc32/cc\"\nchmod +x \"$TARGET/cc32/cc\"\nexport PATH=\"$TARGET/cc32:$PATH\"\nexport JANET_TOOLCHAIN=cc\nunset JANET_PATH JANET_TREE\nmkdir -p \"$TARGET\"\ncd \"$TARGET\"\ngit clone https://github.com/janet-lang/janet\ncd janet\n# make the change to Janet source\ngit checkout -q 0e5fdd53 # otherwise this sed will bit rot\nsed -i 's/^#if ((defined(__x86_64__) || defined(_M_X64)) \\\\$/#if ((defined(__x86_64__) || defined(_M_X64)) \\&\\& !defined(__ILP32__) \\\\/' src/include/janet.h\nexport CFLAGS='-fPIC -O3 -flto -fno-semantic-interposition -march=native' # speed gains\nPREFIX=\"$TARGET\" make clean\nPREFIX=\"$TARGET\" make\nPREFIX=\"$TARGET\" make test\nPREFIX=\"$TARGET\" make install\ncd ..\ngit clone --depth=1 https://github.com/janet-lang/spork build/spork\n\"$JANET\" -e '(if (bundle/installed? \"spork\") (bundle/replace \"spork\" \"build/spork\") (bundle/install \"build/spork\"))'\n# ### Packages\nPM=\"$TARGET/lib/janet/bin/janet-pm\"\ngit clone https://codeberg.org/veqq/declarative-dsls\n\"$JANET\" \"$PM\" install file::declarative-dsls # also does https://codeberg.org/veqq/varray\ngit clone https://github.com/ianthehenry/judge # also does https://github.com/ianthehenry/cmd\n\"$JANET\" \"$PM\" install file::judge\ntime \"$JANET\" \"$TARGET/lib/janet/bin/judge\" declarative-dsls/tests.janet # to show that it worked\n```\nAnd a `64janet.sh to compare it with:\n\n```\n#!/bin/bash\nset -eu\nTARGET=\"$PWD/janet64\"\nJANET=\"$TARGET/bin/janet\"\nunset JANET_PATH JANET_TREE\nmkdir -p \"$TARGET\"\ncd \"$TARGET\"\ngit clone https://github.com/janet-lang/janet\ncd janet\ngit checkout -q 0e5fdd53 # same commit as the x32 build\nexport CFLAGS='-fPIC -O3 -flto -fno-semantic-interposition -march=native' # speed gains\nPREFIX=\"$TARGET\" make clean\nPREFIX=\"$TARGET\" make\nPREFIX=\"$TARGET\" make test\nPREFIX=\"$TARGET\" make install\ncd ..\ngit clone --depth=1 https://github.com/janet-lang/spork build/spork\n\"$JANET\" -e '(if (bundle/installed? \"spork\") (bundle/replace \"spork\" \"build/spork\") (bundle/install \"build/spork\"))'\n# ### Packages\nPM=\"$TARGET/lib/janet/bin/janet-pm\"\ngit clone https://codeberg.org/veqq/declarative-dsls\n\"$JANET\" \"$PM\" install file::declarative-dsls # also does https://codeberg.org/veqq/varray\ngit clone https://github.com/ianthehenry/judge # also does https://github.com/ianthehenry/cmd\n\"$JANET\" \"$PM\" install file::judge\ntime \"$JANET\" \"$TARGET/lib/janet/bin/judge\" declarative-dsls/tests.janet # to show that it worked\n```\nNow we compare them:\n\n```\nsudo apt install time # different from just time\n/usr/bin/time -f '%e s  %M KB' janet64/bin/janet janet64/lib/janet/bin/judge janet64/declarative-dsls/tests.janet\n/usr/bin/time -f '%e s  %M KB' janetx32/bin/janet janetx32/lib/janet/bin/judge janetx32/declarative-dsls/tests.janet\n2.42 s  61668 KB # 64bit\n2.36 s  53472 KB # 32bit\n```\nBut [declarative-dsls](https://codeberg.org/veqq/declarative-dsls) is a worst case scenario, using [typed c arrays](https://codeberg.org/veqq/varray) for big speed ups, which don’t benefit much from this. Lets make some minimal benchmarks to really test the differences:\n\n## See the Benchmarks Scripts\n\n`words.janet` builds a text, splits it into words, counts them and sorts on that:\n\n```\n(let [start (os/clock :monotonic)\n      rng (math/rng 42)\n      vocab (seq [i :range [0 5000]] (string \"w\" (math/rng-int rng 100000)))\n      text @\"\"\n      counts @{}]\n  (repeat 600000 (buffer/push text (in vocab (math/rng-int rng 5000)) \" \"))\n  (each word (string/split \" \" text)\n    (put counts word (+ 1 (get counts word 0))))\n  (pp (take 3 (sort-by |(- (in $ 1)) (pairs counts))))\n  (print (length counts) \" distinct words\")\n  (print (- (os/clock :monotonic) start) \" s\"))\n```\n`records.janet` groups structs by department and summarizes those departments, with unused extra fields like real code:\n\n```\n(let [start (os/clock :monotonic)\n      rng (math/rng 7)\n      depts [:eng :ops :sales :legal :support]\n      people (seq [i :range [0 300000]]\n               {:id i\n               # :name and :tags not used\n                :name (string \"person-\" i)\n                :dept (in depts (math/rng-int rng 5))\n                :salary (+ 30000 (math/rng-int rng 90000))\n                :tags [(math/rng-int rng 10) (math/rng-int rng 10)]})\n      by-dept (group-by |(in $ :dept) people)]\n  (each dept (sorted (keys by-dept))\n    (let [staff (in by-dept dept)\n          salaries (map |(in $ :salary) staff)]\n      (pp [dept\n           (length staff)\n           (math/round (/ (sum salaries) (length staff)))\n           (length (filter |(> $ 100000) salaries))])))\n  (print (- (os/clock :monotonic) start) \" s\"))\n```\n`parse.janet` parses lines with a PEG and totals them:\n\n```\n(let [start (os/clock :monotonic)\n      rng (math/rng 3)\n      lines (seq [i :range [0 200000]]\n              (string i \",\" (math/rng-int rng 1000) \",item-\" (math/rng-int rng 500) \",\" (/ (math/rng-int rng 10000) 100)))\n      row (peg/compile\n            ~{:num (number (some (set \"0123456789.\")))\n              :txt (capture (some (if-not \",\" 1)))\n              :main (* :num \",\" :num \",\" :txt \",\" :num -1)})\n      rows (map |(peg/match row $) lines)\n      totals @{}]\n  (each [id qty item price] rows\n    (put totals item (+ (get totals item 0) (* qty price))))\n  (print (length rows) \" rows, \" (length totals) \" items, total \" (math/round (sum (values totals))))\n  (print (- (os/clock :monotonic) start) \" s\"))\n```\n`tree.janet` builds and walks binary trees, our pointer-heaviest example:\n\n```\n(let [start (os/clock :monotonic)\n      make (fn make [depth]\n             (if (= depth 0)\n               [nil nil]\n               [(make (- depth 1)) (make (- depth 1))]))\n      check (fn check [node]\n              (if (in node 0)\n                (+ 1 (check (in node 0)) (check (in node 1)))\n                1))\n      long-lived (make 19)]\n  (var total 0)\n  (repeat 40 (+= total (check (make 14))))\n  (print total \" \" (check long-lived))\n  (print (- (os/clock :monotonic) start) \" s\"))\n```\nWe benchmark:\n\n```\nfor f in words records tree parse; do  \n>   /usr/bin/time -f \"$f 64:  %e s  %M KB\" janet64/bin/janet $f.janet > /dev/null  \n>   /usr/bin/time -f \"$f x32: %e s  %M KB\" janetx32/bin/janet $f.janet > /dev/null  \n> done  \n> \nwords 64:  0.58 s  41904 KB  \nwords x32: 0.67 s  31856 KB  \nrecords 64:  6.40 s  138164 KB  \nrecords x32: 5.80 s  126916 KB  \ntree 64:  1.39 s  50504 KB  \ntree x32: 1.50 s  35532 KB  \nparse 64:  3.20 s  58504 KB  \nparse x32: 3.19 s  44428 KB\nwords 64:  0.63 s  41984 KB  \nwords x32: 0.70 s  31976 KB  \nrecords 64:  6.71 s  138240 KB  \nrecords x32: 5.77 s  126844 KB  \ntree 64:  1.40 s  50392 KB  \ntree x32: 1.48 s  35304 KB  \nparse 64:  3.37 s  58592 KB  \nparse x32: 3.09 s  44496 KB\nwords 64:  0.59 s  41756 KB  \nwords x32: 0.67 s  31788 KB  \nrecords 64:  6.50 s  138384 KB  \nrecords x32: 5.80 s  126960 KB  \ntree 64:  1.39 s  50288 KB  \ntree x32: 1.43 s  35560 KB  \nparse 64:  3.35 s  58520 KB  \nparse x32: 3.24 s  44572 KB\nwords 64:  0.60 s  42012 KB  \nwords x32: 0.67 s  31896 KB  \nrecords 64:  6.29 s  138448 KB  \nrecords x32: 6.00 s  126832 KB  \ntree 64:  1.38 s  50176 KB  \ntree x32: 1.43 s  35392 KB  \nparse 64:  3.36 s  58384 KB  \nparse x32: 3.13 s  44364 KB\n```\n### Conclusion\n\nWe see the highest memory savings from tiny objects, averaging about 25% with inconsistent speed ups (8%) or slowdowns (16%). Perhaps other projects are more ammenable to gains than Janet where nanboxing already packs values into 8 bytes and smaller headers see the main benefit. Midway through, I was hoping for performance gains due to more staying in the L1 cache etc. where, in another world with continued x32 support, we could use this as a default for scripts. Faced with the data, the only ~25% decrease in RAM usage is nice, but does not justify a call to action to get more out of our machines in this age of expensive RAM.\n\n### P.S. Implications for Janet\n\nIf Janet works with smaller object headers, why not shrink them in the implementation for 64-bit use? By my understanding, Janet [heap objects](https://github.com/janet-lang/janet/blob/c495f11f15a29033b5f52d66f82f13a105f98123/src/include/janet.h#L974) use 16 bytes to help the garbage collector:\n\n- 4 bytes for flags\n- 4 bytes of padding\n- 8 bytes linking to the next object\n\nReducing these would require a different garbage collection strategy e.g. our own allocator, which would hurt a core Janet use case: embedding in other projects.\n\nEach type then adds:\n\n- [string](https://github.com/janet-lang/janet/blob/c495f11f15a29033b5f52d66f82f13a105f98123/src/include/janet.h#L1081) : 8 bytes for length and hash\n- [tuple](https://github.com/janet-lang/janet/blob/c495f11f15a29033b5f52d66f82f13a105f98123/src/include/janet.h#L1061) : 8 bytes for length and hash, 8 bytes for source line and column info\n- [array](https://github.com/janet-lang/janet/blob/c495f11f15a29033b5f52d66f82f13a105f98123/src/include/janet.h#L1029) : 8 bytes for count and capacity, 8 bytes for a pointer to the data\n- [struct](https://github.com/janet-lang/janet/blob/c495f11f15a29033b5f52d66f82f13a105f98123/src/include/janet.h#L1071) : 12 bytes for length, hash and capacity, 4 bytes padding and 8 bytes for its prototype’s pointer\n- [table](https://github.com/janet-lang/janet/blob/c495f11f15a29033b5f52d66f82f13a105f98123/src/include/janet.h#L1045) : 12 bytes for count, capacity and how many deleted, 4 bytes padding, 8 bytes for a pointer and 8 for the prototype’s pointer\n\nFor structs and tables in particular, we could regain 8 bytes through rearranging (not needing the later padding). Onlt tuples parsed from code need the source information (for errors); a flag could save normal tuples 8 bytes.\n\nUnfortunately, glibc’s `malloc` adds 8 bytes and rounds to multiples of 16 so structs would remain the same. But tables and (some) tuples would shrunk more than the above indicates!\n","body_html":"<h1 id=\"janet-on-x32-32-bit-pointers-64-bit-speed-25-less-ram\">Janet on x32: 32-bit Pointers, 64-bit Speed, 25% Less RAM</h1>\n<p>It’s possible to save a dramatic amount of memory (close to half for programs with pointer heavy heaps) and also speed up programs by a modest amount (by way of more data fitting in cache) by using 32 bit pointers on 64 bit systems.</p>\n<p>The Linux x32 ABI lets you do just this, but it’s tragically underused and overlooked. It’s disabled by default on Debian (though you can enable it with a boot flag) and not even compiled in on Arch Linux. Most software compiles just fine for x32, but there isn’t much packaging for it so you do have to compile nearly everything yourself.</p>\n<p>I experimented with deploying mastodon on x32 and it cut the app’s memory usage from 650mb to 350mb. There’s so much potential here, but the work is mostly of the thankless coordination and communication type and I don’t have the time or motivation to push it forward myself. - <a href=\"https://lobste.rs/s/gns27z/what_are_your_programming_hunches_you#c_tw3vok\" rel=\"nofollow ugc noopener\">Hailey</a></p>\n<h3 id=\"how-to-compile-for-32-bits\">How to Compile for 32 Bits?</h3>\n<p>Janet libraries like Spork supply their own build flags, but we can hijack them with a fake <code>cc</code> which starts the real one with the necessary <code>-m32 -msse2 -mfpmath=sse</code> flags. I modified the beginning of my default <a href=\"https://codeberg.org/veqq/install-latest-janet/src/branch/master/install-latest-janet.sh\" rel=\"nofollow ugc noopener\">Janet build script</a>:</p>\n<pre><code>#!/bin/sh\nset -eu\nTARGET=&quot;$PWD/janet32&quot;\nJANET=&quot;$TARGET/bin/janet&quot;\nmkdir -p &quot;$TARGET/cc32&quot;\nprintf &#39;#!/bin/sh\\nexec gcc -m32 -msse2 -mfpmath=sse &quot;$@&quot;\\n&#39; &gt; &quot;$TARGET/cc32/cc&quot;\nchmod +x &quot;$TARGET/cc32/cc&quot;\nexport PATH=&quot;$TARGET/cc32:$PATH&quot;\nexport JANET_TOOLCHAIN=cc\nunset JANET_PATH JANET_TREE\nmkdir -p &quot;$TARGET&quot;\ncd &quot;$TARGET&quot;\nrm -rf build/spork janet\ngit clone https://github.com/janet-lang/janet\ncd janet</code></pre>\n<p>Easy enough! Now let’s try it!</p>\n<h3 id=\"failing-on-arch\">Failing on Arch</h3>\n<p>tl;dr: 20% memory reduction, 50% slower</p>\n<h2 id=\"see-the-scripts-and-results\">See the Scripts and Results</h2>\n<p>This is on cachyos (arch, btw) with lib32-glibc lib32-gcc-libs. Adding this to a build script forces libraires like spork to build at 32 bit too!</p>\n<p>For <code>mem.janet</code>:</p>\n<pre><code>(let [live @[]]\n  (for i 0 2_000_000\n    (array/push live [i (+ i 1)]))\n  (gccollect)\n  (print (length live)))</code></pre>\n<p>Note, these use Fish and I don’t want to rebuild my site with other language support:</p>\n<pre><code>sudo pacman -S time # different from just time\ncommand time -v ./bin/janet mem.janet # command skips the fish version\ncommand time -v ./janet32/bin/janet mem.janet\n2000000\n       Command being timed: &quot;./bin/janet mem.janet&quot;\n       User time (seconds): 0.24\n       System time (seconds): 0.05\n       Percent of CPU this job got: 100%\n       Elapsed (wall clock) time (h:mm:ss or m:ss): 0:00.30\n       Average shared text size (kbytes): 0\n       Average unshared data size (kbytes): 0\n       Average stack size (kbytes): 0\n       Average total size (kbytes): 0\n       Maximum resident set size (kbytes): 145196\n       Average resident set size (kbytes): 0\n       Major (requiring I/O) page faults: 0\n       Minor (reclaiming a frame) page faults: 33513\n       Voluntary context switches: 1\n       Involuntary context switches: 9\n2000000\n       Command being timed: &quot;./janet32/bin/janet mem.janet&quot;\n       User time (seconds): 0.47\n       System time (seconds): 0.03\n       Percent of CPU this job got: 99%\n       Elapsed (wall clock) time (h:mm:ss or m:ss): 0:00.50\n       Average shared text size (kbytes): 0\n       Average unshared data size (kbytes): 0\n       Average stack size (kbytes): 0\n       Average total size (kbytes): 0\n       Maximum resident set size (kbytes): 113716\n       Average resident set size (kbytes): 0\n       Major (requiring I/O) page faults: 0\n       Minor (reclaiming a frame) page faults: 25686\n       Voluntary context switches: 1\n       Involuntary context switches: 21</code></pre>\n<p>I added a 0 in the test</p>\n<pre><code>command time -v ./bin/janet mem.janet\n                  command time -v ./janet32/bin/janet mem.janet\n20000000\n        Command being timed: &quot;./bin/janet mem.janet&quot;\n        User time (seconds): 2.83\n        System time (seconds): 0.46\n        Percent of CPU this job got: 99%\n        Elapsed (wall clock) time (h:mm:ss or m:ss): 0:03.31\n        Average shared text size (kbytes): 0\n        Average unshared data size (kbytes): 0\n        Average stack size (kbytes): 0\n        Average total size (kbytes): 0\n        Maximum resident set size (kbytes): 1410668\n        Average resident set size (kbytes): 0\n        Major (requiring I/O) page faults: 0\n        Minor (reclaiming a frame) page faults: 316820\n        Voluntary context switches: 1\n        Involuntary context switches: 164\n        Exit status: 0\n20000000\n        Command being timed: &quot;./janet32/bin/janet mem.janet&quot;\n        User time (seconds): 4.64\n        System time (seconds): 0.35\n        Percent of CPU this job got: 99%\n        Elapsed (wall clock) time (h:mm:ss or m:ss): 0:05.00\n        Average shared text size (kbytes): 0\n        Average unshared data size (kbytes): 0\n        Average stack size (kbytes): 0\n        Average total size (kbytes): 0\n        Maximum resident set size (kbytes): 1098372\n        Average resident set size (kbytes): 0\n        Major (requiring I/O) page faults: 0\n        Minor (reclaiming a frame) page faults: 238633\n        Voluntary context switches: 1\n        Involuntary context switches: 193</code></pre>\n<p>For e.g. <code>declarative-dsls/tests.janet</code></p>\n<pre><code>98 passed\n        Command being timed: &quot;./bin/janet ./lib/janet/bin/judge declarative-dsls/tests.janet&quot;\n        User time (seconds): 0.42\n        System time (seconds): 0.02\n        Percent of CPU this job got: 100%\n        Elapsed (wall clock) time (h:mm:ss or m:ss): 0:00.45\n        Average shared text size (kbytes): 0\n        Average unshared data size (kbytes): 0\n        Average stack size (kbytes): 0\n        Average total size (kbytes): 0\n        Maximum resident set size (kbytes): 61900\ndeclarative-dsls/tests.janet\n98 passed\n        Command being timed: &quot;./bin/janet ./lib/janet/bin/judge declarative-dsls/tests.janet&quot;\n        User time (seconds): 0.70\n        System time (seconds): 0.02\n        Percent of CPU this job got: 99%\n        Elapsed (wall clock) time (h:mm:ss or m:ss): 0:00.72\n        Average shared text size (kbytes): 0\n        Average unshared data size (kbytes): 0\n        Average stack size (kbytes): 0\n        Average total size (kbytes): 0\n        Maximum resident set size (kbytes): 54108\n        Average resident set size (kbytes): 0\n        Major (requiring I/O) page faults: 0\n        Minor (reclaiming a frame) page faults: 12735</code></pre>\n<p>As Hailey said, Arch does not come with the <a href=\"https://en.wikipedia.org/wiki/X32_ABI\" rel=\"nofollow ugc noopener\">x32 ABI</a> compiled, so I had to use <code>-m32</code> and while we do get a 20% memory savings, we execute 50% slower because half the registers aren’t used! <code>-m32</code> is the old 32-bit x86 mode using only 8 registers. To see the speed-up, we need to use <code>-mx32</code> which uses the 64-bit instruction set with shrunken 32-bit pointers. But I don’t want to recompile Linux…</p>\n<p>Unfortunately, Ubuntu also dropped x32 support.</p>\n<h3 id=\"winning-on-deprecated-ubuntu\">Winning on Deprecated Ubuntu</h3>\n<p>Luckily, lazily, I have access to some deprecated servers (<em>of course</em> not in prod…) and OS installs, wasting hard disk space:</p>\n<p>Ubuntu 20.04 LTS Focal Fossa has reached its end of standard support on 31 May 2025.</p>\n<p>Ubuntu, so we have bash again:</p>\n<pre><code>grep X86_X32 /boot/config-$(uname -r)  \nCONFIG_X86_X32=y</code></pre>\n<p>Just what we need, although <code>gcc version 9.4.0</code> is a bit old. We already have <code>build-essential git gcc-multilib libc6-dev-x32</code>. We’re almost ready to test this hunch. But first, we must change <code>src/include/janet.h</code>:</p>\n<pre><code>#if ((defined(__x86_64__) || defined(_M_X64)) \\\n     &amp;&amp; (defined(JANET_POSIX) || defined(JANET_WINDOWS))) \\</code></pre>\n<p>becomes:</p>\n<pre><code>#if ((defined(__x86_64__) || defined(_M_X64)) &amp;&amp; !defined(__ILP32__) \\\n     &amp;&amp; (defined(JANET_POSIX) || defined(JANET_WINDOWS))) \\</code></pre>\n<p>lest Janet choose the 64-bit value layout. Our build script will also pass a build flag to disable Janet’s FFI. Here is our full <code>32janet.sh</code>:</p>\n<pre><code>#!/bin/bash\nset -eu\nTARGET=&quot;$PWD/janetx32&quot;\nJANET=&quot;$TARGET/bin/janet&quot;\nmkdir -p &quot;$TARGET/cc32&quot;\n# this forces Janet C libraries like Spork to build with 32bit also\n# FFI build test fails on -mx32 build so -DJANET_NO_FFI\nprintf &#39;#!/bin/sh\\nexec gcc -mx32 -DJANET_NO_FFI &quot;$@&quot;\\n&#39; &gt; &quot;$TARGET/cc32/cc&quot;\nchmod +x &quot;$TARGET/cc32/cc&quot;\nexport PATH=&quot;$TARGET/cc32:$PATH&quot;\nexport JANET_TOOLCHAIN=cc\nunset JANET_PATH JANET_TREE\nmkdir -p &quot;$TARGET&quot;\ncd &quot;$TARGET&quot;\ngit clone https://github.com/janet-lang/janet\ncd janet\n# make the change to Janet source\ngit checkout -q 0e5fdd53 # otherwise this sed will bit rot\nsed -i &#39;s/^#if ((defined(__x86_64__) || defined(_M_X64)) \\\\$/#if ((defined(__x86_64__) || defined(_M_X64)) \\&amp;\\&amp; !defined(__ILP32__) \\\\/&#39; src/include/janet.h\nexport CFLAGS=&#39;-fPIC -O3 -flto -fno-semantic-interposition -march=native&#39; # speed gains\nPREFIX=&quot;$TARGET&quot; make clean\nPREFIX=&quot;$TARGET&quot; make\nPREFIX=&quot;$TARGET&quot; make test\nPREFIX=&quot;$TARGET&quot; make install\ncd ..\ngit clone --depth=1 https://github.com/janet-lang/spork build/spork\n&quot;$JANET&quot; -e &#39;(if (bundle/installed? &quot;spork&quot;) (bundle/replace &quot;spork&quot; &quot;build/spork&quot;) (bundle/install &quot;build/spork&quot;))&#39;\n# ### Packages\nPM=&quot;$TARGET/lib/janet/bin/janet-pm&quot;\ngit clone https://codeberg.org/veqq/declarative-dsls\n&quot;$JANET&quot; &quot;$PM&quot; install file::declarative-dsls # also does https://codeberg.org/veqq/varray\ngit clone https://github.com/ianthehenry/judge # also does https://github.com/ianthehenry/cmd\n&quot;$JANET&quot; &quot;$PM&quot; install file::judge\ntime &quot;$JANET&quot; &quot;$TARGET/lib/janet/bin/judge&quot; declarative-dsls/tests.janet # to show that it worked</code></pre>\n<p>And a `64janet.sh to compare it with:</p>\n<pre><code>#!/bin/bash\nset -eu\nTARGET=&quot;$PWD/janet64&quot;\nJANET=&quot;$TARGET/bin/janet&quot;\nunset JANET_PATH JANET_TREE\nmkdir -p &quot;$TARGET&quot;\ncd &quot;$TARGET&quot;\ngit clone https://github.com/janet-lang/janet\ncd janet\ngit checkout -q 0e5fdd53 # same commit as the x32 build\nexport CFLAGS=&#39;-fPIC -O3 -flto -fno-semantic-interposition -march=native&#39; # speed gains\nPREFIX=&quot;$TARGET&quot; make clean\nPREFIX=&quot;$TARGET&quot; make\nPREFIX=&quot;$TARGET&quot; make test\nPREFIX=&quot;$TARGET&quot; make install\ncd ..\ngit clone --depth=1 https://github.com/janet-lang/spork build/spork\n&quot;$JANET&quot; -e &#39;(if (bundle/installed? &quot;spork&quot;) (bundle/replace &quot;spork&quot; &quot;build/spork&quot;) (bundle/install &quot;build/spork&quot;))&#39;\n# ### Packages\nPM=&quot;$TARGET/lib/janet/bin/janet-pm&quot;\ngit clone https://codeberg.org/veqq/declarative-dsls\n&quot;$JANET&quot; &quot;$PM&quot; install file::declarative-dsls # also does https://codeberg.org/veqq/varray\ngit clone https://github.com/ianthehenry/judge # also does https://github.com/ianthehenry/cmd\n&quot;$JANET&quot; &quot;$PM&quot; install file::judge\ntime &quot;$JANET&quot; &quot;$TARGET/lib/janet/bin/judge&quot; declarative-dsls/tests.janet # to show that it worked</code></pre>\n<p>Now we compare them:</p>\n<pre><code>sudo apt install time # different from just time\n/usr/bin/time -f &#39;%e s  %M KB&#39; janet64/bin/janet janet64/lib/janet/bin/judge janet64/declarative-dsls/tests.janet\n/usr/bin/time -f &#39;%e s  %M KB&#39; janetx32/bin/janet janetx32/lib/janet/bin/judge janetx32/declarative-dsls/tests.janet\n2.42 s  61668 KB # 64bit\n2.36 s  53472 KB # 32bit</code></pre>\n<p>But <a href=\"https://codeberg.org/veqq/declarative-dsls\" rel=\"nofollow ugc noopener\">declarative-dsls</a> is a worst case scenario, using <a href=\"https://codeberg.org/veqq/varray\" rel=\"nofollow ugc noopener\">typed c arrays</a> for big speed ups, which don’t benefit much from this. Lets make some minimal benchmarks to really test the differences:</p>\n<h2 id=\"see-the-benchmarks-scripts\">See the Benchmarks Scripts</h2>\n<p><code>words.janet</code> builds a text, splits it into words, counts them and sorts on that:</p>\n<pre><code>(let [start (os/clock :monotonic)\n      rng (math/rng 42)\n      vocab (seq [i :range [0 5000]] (string &quot;w&quot; (math/rng-int rng 100000)))\n      text @&quot;&quot;\n      counts @{}]\n  (repeat 600000 (buffer/push text (in vocab (math/rng-int rng 5000)) &quot; &quot;))\n  (each word (string/split &quot; &quot; text)\n    (put counts word (+ 1 (get counts word 0))))\n  (pp (take 3 (sort-by |(- (in $ 1)) (pairs counts))))\n  (print (length counts) &quot; distinct words&quot;)\n  (print (- (os/clock :monotonic) start) &quot; s&quot;))</code></pre>\n<p><code>records.janet</code> groups structs by department and summarizes those departments, with unused extra fields like real code:</p>\n<pre><code>(let [start (os/clock :monotonic)\n      rng (math/rng 7)\n      depts [:eng :ops :sales :legal :support]\n      people (seq [i :range [0 300000]]\n               {:id i\n               # :name and :tags not used\n                :name (string &quot;person-&quot; i)\n                :dept (in depts (math/rng-int rng 5))\n                :salary (+ 30000 (math/rng-int rng 90000))\n                :tags [(math/rng-int rng 10) (math/rng-int rng 10)]})\n      by-dept (group-by |(in $ :dept) people)]\n  (each dept (sorted (keys by-dept))\n    (let [staff (in by-dept dept)\n          salaries (map |(in $ :salary) staff)]\n      (pp [dept\n           (length staff)\n           (math/round (/ (sum salaries) (length staff)))\n           (length (filter |(&gt; $ 100000) salaries))])))\n  (print (- (os/clock :monotonic) start) &quot; s&quot;))</code></pre>\n<p><code>parse.janet</code> parses lines with a PEG and totals them:</p>\n<pre><code>(let [start (os/clock :monotonic)\n      rng (math/rng 3)\n      lines (seq [i :range [0 200000]]\n              (string i &quot;,&quot; (math/rng-int rng 1000) &quot;,item-&quot; (math/rng-int rng 500) &quot;,&quot; (/ (math/rng-int rng 10000) 100)))\n      row (peg/compile\n            ~{:num (number (some (set &quot;0123456789.&quot;)))\n              :txt (capture (some (if-not &quot;,&quot; 1)))\n              :main (* :num &quot;,&quot; :num &quot;,&quot; :txt &quot;,&quot; :num -1)})\n      rows (map |(peg/match row $) lines)\n      totals @{}]\n  (each [id qty item price] rows\n    (put totals item (+ (get totals item 0) (* qty price))))\n  (print (length rows) &quot; rows, &quot; (length totals) &quot; items, total &quot; (math/round (sum (values totals))))\n  (print (- (os/clock :monotonic) start) &quot; s&quot;))</code></pre>\n<p><code>tree.janet</code> builds and walks binary trees, our pointer-heaviest example:</p>\n<pre><code>(let [start (os/clock :monotonic)\n      make (fn make [depth]\n             (if (= depth 0)\n               [nil nil]\n               [(make (- depth 1)) (make (- depth 1))]))\n      check (fn check [node]\n              (if (in node 0)\n                (+ 1 (check (in node 0)) (check (in node 1)))\n                1))\n      long-lived (make 19)]\n  (var total 0)\n  (repeat 40 (+= total (check (make 14))))\n  (print total &quot; &quot; (check long-lived))\n  (print (- (os/clock :monotonic) start) &quot; s&quot;))</code></pre>\n<p>We benchmark:</p>\n<pre><code>for f in words records tree parse; do  \n&gt;   /usr/bin/time -f &quot;$f 64:  %e s  %M KB&quot; janet64/bin/janet $f.janet &gt; /dev/null  \n&gt;   /usr/bin/time -f &quot;$f x32: %e s  %M KB&quot; janetx32/bin/janet $f.janet &gt; /dev/null  \n&gt; done  \n&gt; \nwords 64:  0.58 s  41904 KB  \nwords x32: 0.67 s  31856 KB  \nrecords 64:  6.40 s  138164 KB  \nrecords x32: 5.80 s  126916 KB  \ntree 64:  1.39 s  50504 KB  \ntree x32: 1.50 s  35532 KB  \nparse 64:  3.20 s  58504 KB  \nparse x32: 3.19 s  44428 KB\nwords 64:  0.63 s  41984 KB  \nwords x32: 0.70 s  31976 KB  \nrecords 64:  6.71 s  138240 KB  \nrecords x32: 5.77 s  126844 KB  \ntree 64:  1.40 s  50392 KB  \ntree x32: 1.48 s  35304 KB  \nparse 64:  3.37 s  58592 KB  \nparse x32: 3.09 s  44496 KB\nwords 64:  0.59 s  41756 KB  \nwords x32: 0.67 s  31788 KB  \nrecords 64:  6.50 s  138384 KB  \nrecords x32: 5.80 s  126960 KB  \ntree 64:  1.39 s  50288 KB  \ntree x32: 1.43 s  35560 KB  \nparse 64:  3.35 s  58520 KB  \nparse x32: 3.24 s  44572 KB\nwords 64:  0.60 s  42012 KB  \nwords x32: 0.67 s  31896 KB  \nrecords 64:  6.29 s  138448 KB  \nrecords x32: 6.00 s  126832 KB  \ntree 64:  1.38 s  50176 KB  \ntree x32: 1.43 s  35392 KB  \nparse 64:  3.36 s  58384 KB  \nparse x32: 3.13 s  44364 KB</code></pre>\n<h3 id=\"conclusion\">Conclusion</h3>\n<p>We see the highest memory savings from tiny objects, averaging about 25% with inconsistent speed ups (8%) or slowdowns (16%). Perhaps other projects are more ammenable to gains than Janet where nanboxing already packs values into 8 bytes and smaller headers see the main benefit. Midway through, I was hoping for performance gains due to more staying in the L1 cache etc. where, in another world with continued x32 support, we could use this as a default for scripts. Faced with the data, the only ~25% decrease in RAM usage is nice, but does not justify a call to action to get more out of our machines in this age of expensive RAM.</p>\n<h3 id=\"p-s-implications-for-janet\">P.S. Implications for Janet</h3>\n<p>If Janet works with smaller object headers, why not shrink them in the implementation for 64-bit use? By my understanding, Janet <a href=\"https://github.com/janet-lang/janet/blob/c495f11f15a29033b5f52d66f82f13a105f98123/src/include/janet.h#L974\" rel=\"nofollow ugc noopener\">heap objects</a> use 16 bytes to help the garbage collector:</p>\n<ul><li>4 bytes for flags</li><li>4 bytes of padding</li><li>8 bytes linking to the next object</li></ul>\n<p>Reducing these would require a different garbage collection strategy e.g. our own allocator, which would hurt a core Janet use case: embedding in other projects.</p>\n<p>Each type then adds:</p>\n<ul><li><a href=\"https://github.com/janet-lang/janet/blob/c495f11f15a29033b5f52d66f82f13a105f98123/src/include/janet.h#L1081\" rel=\"nofollow ugc noopener\">string</a> : 8 bytes for length and hash</li><li><a href=\"https://github.com/janet-lang/janet/blob/c495f11f15a29033b5f52d66f82f13a105f98123/src/include/janet.h#L1061\" rel=\"nofollow ugc noopener\">tuple</a> : 8 bytes for length and hash, 8 bytes for source line and column info</li><li><a href=\"https://github.com/janet-lang/janet/blob/c495f11f15a29033b5f52d66f82f13a105f98123/src/include/janet.h#L1029\" rel=\"nofollow ugc noopener\">array</a> : 8 bytes for count and capacity, 8 bytes for a pointer to the data</li><li><a href=\"https://github.com/janet-lang/janet/blob/c495f11f15a29033b5f52d66f82f13a105f98123/src/include/janet.h#L1071\" rel=\"nofollow ugc noopener\">struct</a> : 12 bytes for length, hash and capacity, 4 bytes padding and 8 bytes for its prototype’s pointer</li><li><a href=\"https://github.com/janet-lang/janet/blob/c495f11f15a29033b5f52d66f82f13a105f98123/src/include/janet.h#L1045\" rel=\"nofollow ugc noopener\">table</a> : 12 bytes for count, capacity and how many deleted, 4 bytes padding, 8 bytes for a pointer and 8 for the prototype’s pointer</li></ul>\n<p>For structs and tables in particular, we could regain 8 bytes through rearranging (not needing the later padding). Onlt tuples parsed from code need the source information (for errors); a flag could save normal tuples 8 bytes.</p>\n<p>Unfortunately, glibc’s <code>malloc</code> adds 8 bytes and rounds to multiples of 16 so structs would remain the same. But tables and (some) tuples would shrunk more than the above indicates!</p>","headings":[{"level":1,"text":"Janet on x32: 32-bit Pointers, 64-bit Speed, 25% Less RAM","id":"janet-on-x32-32-bit-pointers-64-bit-speed-25-less-ram"},{"level":3,"text":"How to Compile for 32 Bits?","id":"how-to-compile-for-32-bits"},{"level":3,"text":"Failing on Arch","id":"failing-on-arch"},{"level":2,"text":"See the Scripts and Results","id":"see-the-scripts-and-results"},{"level":3,"text":"Winning on Deprecated Ubuntu","id":"winning-on-deprecated-ubuntu"},{"level":2,"text":"See the Benchmarks Scripts","id":"see-the-benchmarks-scripts"},{"level":3,"text":"Conclusion","id":"conclusion"},{"level":3,"text":"P.S. Implications for Janet","id":"p-s-implications-for-janet"}]}}