{"article":{"slug":"textbook-review-is-parallel-programming-hard-and-if-so-what-can-you-do-about-it","title":"Textbook review: Is Parallel Programming Hard, And, If So, What Can You Do About It?","subtitle":null,"summary":"Andrew Helwer reviews McKenney’s parallel programming textbook: what it covers well, where it frustrates, and who should actually read it.","content_type":"blog_post","language":"en","canonical_url":"https://ahelwer.ca/post/2026-09-21-concurrency-textbook/","author":{"name":"Andrew Helwer","url":null,"person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"ahelwer.ca","url":"https://ahelwer.ca/","listing_slug":null,"listing":null},"topics":[{"name":"Programming","slug":"programming","url":"https://listedarticles.com/topics/programming"},{"name":"Concurrency","slug":"concurrency","url":"https://listedarticles.com/topics/concurrency"},{"name":"Books","slug":"books","url":"https://listedarticles.com/topics/books"},{"name":"Systems Programming","slug":"systems-programming","url":"https://listedarticles.com/topics/systems-programming"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":1715,"reading_minutes":7,"published_at":"2026-09-21T12:00:00.000Z","added_at":"2026-09-24T21:17:46.362Z","updated_at":"2026-09-24T21:17:46.362Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":true},"profile_url":"https://listedarticles.com/articles/textbook-review-is-parallel-programming-hard-and-if-so-what-can-you-do-about-it","markdown_url":"https://listedarticles.com/articles/textbook-review-is-parallel-programming-hard-and-if-so-what-can-you-do-about-it.md","example":false,"citation":"Andrew Helwer, ahelwer.ca. \"Textbook review: Is Parallel Programming Hard, And, If So, What Can You Do About It?.\" 21 Sept 2026. https://ahelwer.ca/post/2026-09-21-concurrency-textbook/ (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://ahelwer.ca/post/2026-09-21-concurrency-textbook/"},"body_markdown":"Here are my thoughts on the free online textbook *Is Parallel Programming Hard, And, If So, What Can You Do About It?* by Paul E. McKenney, author of the Linux kernel’s RCU synchronization mechanism.\nI read a lot of this textbook, got my fill, and probably won’t read more of it in the near future so wanted to write this review while it is all still fresh.\n\n## Set & Setting\n\nHere I’ll talk about the mindset/life phase and physical setting I was in when I started reading this textbook. This might be like the tedious personal flavor preamble they have on recipe websites so skip ahead if that does not interest you.\n\nMy professional life had revolved around TLA⁺ & distributed systems for the past decade, and I was thinking it was time for a change.\nA transitional and emotionally tumultuous period!\nIn 2022 I had tried (and failed) to move to Lean, hoping to acquire an unbelievably niche & nonexistent job as the guy who formalizes researchers’ quantum information processing results for them.\nI burned out on that, which given recent advances in automated theorem proving might have been my temporarily-prescient nervous system dodging me a bullet.\nThus was the history & context in which I attended the 2026 *Software Should Work* conference in Columbia, Missouri.\n\nThe conference had a lot of good talks, but I especially enjoyed the one on Fil-C by Filip Pizlo:\n\nI also got to talk to Fil a fair bit, about interpreters and then about concurrency.\nI fancied myself pretty knowledgeable about concurrency from TLA⁺ & distributed systems, but Fil told me about the difficulty of writing a concurrent lock-free garbage collector and I realized I actually knew very little about concurrency (feeling that you know very little is the mark of a good conference).\nFil also mentioned TLA⁺ might not be useful (or at least ergonomic) for reasoning about events which happen *literally* concurrently (an actual possibility with a multicore CPU!) and the importance of analyzing concurrent algorithms for linearizability, a concept I sort of understood in the distributed systems sense.\n\nAll of this seemed very alluring, so I looked around for a textbook to read about concurrency that focused more on lock-free aspects as opposed to mutex-based or message-passing patterns.\n*Is Parallel Programming Hard, And, If So, What Can You Do About It?* seemed to fit the bill, focusing as it does on general concurrent programming & CPU cache effects instead of more specific textbooks about how to write lock-free datastructures.\nIt also had a few (2023, 2021, 2020, 2015, 2014, 2011) moderately interesting HN threads.\nI don’t think it’s useful spending time in analysis paralysis trying to find the exact “right” textbook (this is really just a clever way to procrastinate), so it seemed good enough.\n\nThe physical setting in which I read this textbook was a 1.5 week vacation to visit my family in a quiet, wooded part of Canada. 2026 also turned out to be a particularly horrific mosquito season. Thus I spent much of the time sitting in a cool screened-in patio, diligently watched over by hundreds of guards ensuring I did not leave my post:\n\nSatellites & cell towers have made distracting internet connectivity annoyingly good even in the more remote parts of the country, but otherwise this was an optimal textbook reading location.\n\n## The Textbook Format\n\nSome quick notes on the actual structure of the textbook; it is available in no fewer than *three* separate formats, all PDF:\n\n1. A dual-column format, like a scientific paper\n2. A single-column format with large margins\n3. A single-column format with no margins\n\nThe last one is perfect for reading on my Pine64 PineNote.\n\nThe textbook contains a huge number of internal links.\nSome of these links are used in quick knowledge-check question boxes, where clicking the link takes you to the question’s answer.\nOther links are used whenever a figure or section is mentioned, or for copious footnotes & citations.\nUnfortunately the latter are very annoying and should probably be reduced by at least 80%.\nIf your e-reader lacks physical page-turn buttons, then your experience of reading the book will consist of constantly accidentally pressing one of these links when you meant to turn the page and thus being sent who-knows-where.\nE-books are disorienting enough to navigate without this, and it pretty much meant it was impossible to quickly flip back & forth between two sections using repeated page-turn taps.\nFor times where I *did* want to click a link, some places had two links right next to each other; touch screens lack the precision to reliably click one link instead of the other.\nThe solution of simply disabling all links presents itself, but then you lose access to the quite nice knowledge-check questions.\nSo I just suffered through it.\n\n## The Textbook Content - Introductory Chapters\n\nThe textbook was more or less the perfect presentation of material for my level.\nThe book starts with nice light introduction & motivation chapters before heading off to the races in chapter 3, *Hardware and its Habits*.\nHere we learn about how modern CPUs work at a high level - what makes them fast, and what makes them slow.\nComplete with a bunch of humorous illustrations!\nSection 3.2.1 - *Hardware System Architecture* is where it really got interesting for me, as we are walked through a simplified account of a CPU core writing to a memory address that does not exist in its cache.\n\nHere is one missed opportunity: I would have really benefited from a basic explanation of the MESI protocol, possessing essentially no intuition about how CPU caches mediate concurrent reads & writes. I found out about MESI while searching online to better understand this section; MESI is only mentioned in the appendix of this book. But my understanding of the rest of the book was greatly improved by knowing about it.\n\nLearning about MESI also taught me that multiple CPU cores cannot write to the same data location literally concurrently! An x86 CPU doesn’t actually write directly to memory, it only writes to its cache (the cache value is then eventually flushed to memory). An x86 CPU core can only write to a particular address when it has exclusive ownership of the cacheline containing that address. If another core tries to write to that address at the same time, it has to wait for exclusive ownership of the cacheline to be moved to it. Thus literally concurrent writes do not actually happen. It is possible for writes to be torn if the data being written spans more than one cacheline, though.\n\nChapter 4, titled *Tools of the Trade*, was positively mind-bending.\nHere we learn that if you write parallel programs without due caution, the compiler will attempt unbelievably creative optimizations resulting in completely nonsensical behavior!\nSection 4.3.4.1, *Shared-Variable Shenanigans*, covers such horrifying mishaps as load tearing, store tearing, load fusing, store fusing, code reordering, invented loads, invented stores (particularly egregious), store-to-load transformations, and dead-code elimination.\n*Then*, assuming your code survived compilation unscathed, the chapter goes over the nonsense the CPU can pull when actually executing your program!\nThis double trouble made it difficult for me to think straight about parallel programs beyond nice familiar mutexes or message-passing.\n\nThe one issue I had with this chapter is that it is very specific to a Linux kernel context.\nI would have liked to have learned about the work C++11 and C11 did to formalize parallel programming with things like `std::memory_order`.\nThese were only given short paragraphs in sections 4.2.6 & 4.2.7, *Atomic Operations (C11)* and *Atomic Operations (Modern GCC)*.\nI escaped the chapter armed with a vague idea that the `ACCESS_ONCE()` and `WRITE_ONCE()` macros just cast things to `volatile*` and that was enough to scare away the compiler.\nVery interesting historical developments, like the debate over whether benign data races were errors, were skipped entirely.\n\n## The Textbook Content - Main Chapters\n\nChapter 5, *Counting*, is probably the marquee chapter of the book.\nIt was also the last chapter I read in any sort of depth.\nThe chapter covers 10 or so different ways of writing a program where several threads increment a counter.\nThe obvious non-broken implementation, where each thread uses atomic increment instructions, is dispatched early on by showing how terribly it performs.\nThis is where my extracurricular understanding of MESI really came in helpful.\n\nThe chapter culminates in something called a signal-theft limit counter, which to be honest I do not entirely understand. I think if I were writing a counter myself I probably would not go that far. I really liked array-based per-thread statistical counters, with their similarity to conflict-free replicated datatypes from the distributed systems world. They were also an excellent vehicle for learning about the performance impact of false sharing - you can’t just chuck all the thread-specific counters in a single contiguous array and call it good!\n\nAfter chapter 5 I mostly just skimmed the material looking for topics of interest.\nSome of the chapters were fairly conceptual, talking about ownership, partitioning, deferred processing, and other things readily translated from distributed systems.\nThe *Formal Verification* chapter used Promela and Spin, which I wasn’t motivated to learn as a TLA⁺ user.\nThe *Validation* chapter had a good section, 11.6.4, on *Hunting Heisenbugs*.\nLock-free programming doesn’t really get covered until chapter 14, *Advanced Synchronization*, and at that point I was ready to move to a textbook focusing on lock-free programming & data structures specifically.\nChapter 15 finally covers memory ordering, but I had already moved on to extracurricular sources trying to fix my confusion about it.\n\n## Overall review\n\nAlthough I only read the first five chapters in-depth, I think this textbook is excellent.\nI say this because it really gave me a thirst to learn more about parallel programming!\nMost lunches at work I struggle to keep myself from infodumping whatever nonsense I’ve learned onto my coworkers.\nLike did you know about the failure of formalizing release-consume ordering?\nOr how basic aligned loads & stores using `mov` are atomic on x86?\nOr out-of-thin-air values?\nOr the incredibly weak memory model of the DEC Alpha?\nOr how release-acquire semantics work?\nOr how deterministic simulation testing cannot (yet) meaningfully test lock-free algorithms?\nOr how atomics on ARM (pre-v8) can be pre-empted?\nThere’s a goldmine of comedically unintuitive nonsense here.\nI want to learn about high-performance garbage collection next.\nTaking recommendations!","body_html":"<p>Here are my thoughts on the free online textbook <em>Is Parallel Programming Hard, And, If So, What Can You Do About It?</em> by Paul E. McKenney, author of the Linux kernel’s RCU synchronization mechanism.\nI read a lot of this textbook, got my fill, and probably won’t read more of it in the near future so wanted to write this review while it is all still fresh.</p>\n<h2 id=\"set-setting\">Set &amp; Setting</h2>\n<p>Here I’ll talk about the mindset/life phase and physical setting I was in when I started reading this textbook. This might be like the tedious personal flavor preamble they have on recipe websites so skip ahead if that does not interest you.</p>\n<p>My professional life had revolved around TLA⁺ &amp; distributed systems for the past decade, and I was thinking it was time for a change.\nA transitional and emotionally tumultuous period!\nIn 2022 I had tried (and failed) to move to Lean, hoping to acquire an unbelievably niche &amp; nonexistent job as the guy who formalizes researchers’ quantum information processing results for them.\nI burned out on that, which given recent advances in automated theorem proving might have been my temporarily-prescient nervous system dodging me a bullet.\nThus was the history &amp; context in which I attended the 2026 <em>Software Should Work</em> conference in Columbia, Missouri.</p>\n<p>The conference had a lot of good talks, but I especially enjoyed the one on Fil-C by Filip Pizlo:</p>\n<p>I also got to talk to Fil a fair bit, about interpreters and then about concurrency.\nI fancied myself pretty knowledgeable about concurrency from TLA⁺ &amp; distributed systems, but Fil told me about the difficulty of writing a concurrent lock-free garbage collector and I realized I actually knew very little about concurrency (feeling that you know very little is the mark of a good conference).\nFil also mentioned TLA⁺ might not be useful (or at least ergonomic) for reasoning about events which happen <em>literally</em> concurrently (an actual possibility with a multicore CPU!) and the importance of analyzing concurrent algorithms for linearizability, a concept I sort of understood in the distributed systems sense.</p>\n<p>All of this seemed very alluring, so I looked around for a textbook to read about concurrency that focused more on lock-free aspects as opposed to mutex-based or message-passing patterns.\n<em>Is Parallel Programming Hard, And, If So, What Can You Do About It?</em> seemed to fit the bill, focusing as it does on general concurrent programming &amp; CPU cache effects instead of more specific textbooks about how to write lock-free datastructures.\nIt also had a few (2023, 2021, 2020, 2015, 2014, 2011) moderately interesting HN threads.\nI don’t think it’s useful spending time in analysis paralysis trying to find the exact “right” textbook (this is really just a clever way to procrastinate), so it seemed good enough.</p>\n<p>The physical setting in which I read this textbook was a 1.5 week vacation to visit my family in a quiet, wooded part of Canada. 2026 also turned out to be a particularly horrific mosquito season. Thus I spent much of the time sitting in a cool screened-in patio, diligently watched over by hundreds of guards ensuring I did not leave my post:</p>\n<p>Satellites &amp; cell towers have made distracting internet connectivity annoyingly good even in the more remote parts of the country, but otherwise this was an optimal textbook reading location.</p>\n<h2 id=\"the-textbook-format\">The Textbook Format</h2>\n<p>Some quick notes on the actual structure of the textbook; it is available in no fewer than <em>three</em> separate formats, all PDF:</p>\n<ol><li>A dual-column format, like a scientific paper</li><li>A single-column format with large margins</li><li>A single-column format with no margins</li></ol>\n<p>The last one is perfect for reading on my Pine64 PineNote.</p>\n<p>The textbook contains a huge number of internal links.\nSome of these links are used in quick knowledge-check question boxes, where clicking the link takes you to the question’s answer.\nOther links are used whenever a figure or section is mentioned, or for copious footnotes &amp; citations.\nUnfortunately the latter are very annoying and should probably be reduced by at least 80%.\nIf your e-reader lacks physical page-turn buttons, then your experience of reading the book will consist of constantly accidentally pressing one of these links when you meant to turn the page and thus being sent who-knows-where.\nE-books are disorienting enough to navigate without this, and it pretty much meant it was impossible to quickly flip back &amp; forth between two sections using repeated page-turn taps.\nFor times where I <em>did</em> want to click a link, some places had two links right next to each other; touch screens lack the precision to reliably click one link instead of the other.\nThe solution of simply disabling all links presents itself, but then you lose access to the quite nice knowledge-check questions.\nSo I just suffered through it.</p>\n<h2 id=\"the-textbook-content-introductory-chapters\">The Textbook Content - Introductory Chapters</h2>\n<p>The textbook was more or less the perfect presentation of material for my level.\nThe book starts with nice light introduction &amp; motivation chapters before heading off to the races in chapter 3, <em>Hardware and its Habits</em>.\nHere we learn about how modern CPUs work at a high level - what makes them fast, and what makes them slow.\nComplete with a bunch of humorous illustrations!\nSection 3.2.1 - <em>Hardware System Architecture</em> is where it really got interesting for me, as we are walked through a simplified account of a CPU core writing to a memory address that does not exist in its cache.</p>\n<p>Here is one missed opportunity: I would have really benefited from a basic explanation of the MESI protocol, possessing essentially no intuition about how CPU caches mediate concurrent reads &amp; writes. I found out about MESI while searching online to better understand this section; MESI is only mentioned in the appendix of this book. But my understanding of the rest of the book was greatly improved by knowing about it.</p>\n<p>Learning about MESI also taught me that multiple CPU cores cannot write to the same data location literally concurrently! An x86 CPU doesn’t actually write directly to memory, it only writes to its cache (the cache value is then eventually flushed to memory). An x86 CPU core can only write to a particular address when it has exclusive ownership of the cacheline containing that address. If another core tries to write to that address at the same time, it has to wait for exclusive ownership of the cacheline to be moved to it. Thus literally concurrent writes do not actually happen. It is possible for writes to be torn if the data being written spans more than one cacheline, though.</p>\n<p>Chapter 4, titled <em>Tools of the Trade</em>, was positively mind-bending.\nHere we learn that if you write parallel programs without due caution, the compiler will attempt unbelievably creative optimizations resulting in completely nonsensical behavior!\nSection 4.3.4.1, <em>Shared-Variable Shenanigans</em>, covers such horrifying mishaps as load tearing, store tearing, load fusing, store fusing, code reordering, invented loads, invented stores (particularly egregious), store-to-load transformations, and dead-code elimination.\n<em>Then</em>, assuming your code survived compilation unscathed, the chapter goes over the nonsense the CPU can pull when actually executing your program!\nThis double trouble made it difficult for me to think straight about parallel programs beyond nice familiar mutexes or message-passing.</p>\n<p>The one issue I had with this chapter is that it is very specific to a Linux kernel context.\nI would have liked to have learned about the work C++11 and C11 did to formalize parallel programming with things like <code>std::memory_order</code>.\nThese were only given short paragraphs in sections 4.2.6 &amp; 4.2.7, <em>Atomic Operations (C11)</em> and <em>Atomic Operations (Modern GCC)</em>.\nI escaped the chapter armed with a vague idea that the <code>ACCESS_ONCE()</code> and <code>WRITE_ONCE()</code> macros just cast things to <code>volatile*</code> and that was enough to scare away the compiler.\nVery interesting historical developments, like the debate over whether benign data races were errors, were skipped entirely.</p>\n<h2 id=\"the-textbook-content-main-chapters\">The Textbook Content - Main Chapters</h2>\n<p>Chapter 5, <em>Counting</em>, is probably the marquee chapter of the book.\nIt was also the last chapter I read in any sort of depth.\nThe chapter covers 10 or so different ways of writing a program where several threads increment a counter.\nThe obvious non-broken implementation, where each thread uses atomic increment instructions, is dispatched early on by showing how terribly it performs.\nThis is where my extracurricular understanding of MESI really came in helpful.</p>\n<p>The chapter culminates in something called a signal-theft limit counter, which to be honest I do not entirely understand. I think if I were writing a counter myself I probably would not go that far. I really liked array-based per-thread statistical counters, with their similarity to conflict-free replicated datatypes from the distributed systems world. They were also an excellent vehicle for learning about the performance impact of false sharing - you can’t just chuck all the thread-specific counters in a single contiguous array and call it good!</p>\n<p>After chapter 5 I mostly just skimmed the material looking for topics of interest.\nSome of the chapters were fairly conceptual, talking about ownership, partitioning, deferred processing, and other things readily translated from distributed systems.\nThe <em>Formal Verification</em> chapter used Promela and Spin, which I wasn’t motivated to learn as a TLA⁺ user.\nThe <em>Validation</em> chapter had a good section, 11.6.4, on <em>Hunting Heisenbugs</em>.\nLock-free programming doesn’t really get covered until chapter 14, <em>Advanced Synchronization</em>, and at that point I was ready to move to a textbook focusing on lock-free programming &amp; data structures specifically.\nChapter 15 finally covers memory ordering, but I had already moved on to extracurricular sources trying to fix my confusion about it.</p>\n<h2 id=\"overall-review\">Overall review</h2>\n<p>Although I only read the first five chapters in-depth, I think this textbook is excellent.\nI say this because it really gave me a thirst to learn more about parallel programming!\nMost lunches at work I struggle to keep myself from infodumping whatever nonsense I’ve learned onto my coworkers.\nLike did you know about the failure of formalizing release-consume ordering?\nOr how basic aligned loads &amp; stores using <code>mov</code> are atomic on x86?\nOr out-of-thin-air values?\nOr the incredibly weak memory model of the DEC Alpha?\nOr how release-acquire semantics work?\nOr how deterministic simulation testing cannot (yet) meaningfully test lock-free algorithms?\nOr how atomics on ARM (pre-v8) can be pre-empted?\nThere’s a goldmine of comedically unintuitive nonsense here.\nI want to learn about high-performance garbage collection next.\nTaking recommendations!</p>","headings":[{"level":2,"text":"Set & Setting","id":"set-setting"},{"level":2,"text":"The Textbook Format","id":"the-textbook-format"},{"level":2,"text":"The Textbook Content - Introductory Chapters","id":"the-textbook-content-introductory-chapters"},{"level":2,"text":"The Textbook Content - Main Chapters","id":"the-textbook-content-main-chapters"},{"level":2,"text":"Overall review","id":"overall-review"}]}}