{"article":{"slug":"union-vs-sum-types","title":"Union vs sum types","subtitle":null,"summary":"A clear comparison of union types and sum types: how they overlap, where their subtle differences change API design and safety, and when each model fits better in modern languages.","content_type":"essay","language":"en","canonical_url":"https://viralinstruction.com/posts/uniontypes/","author":{"name":"Viral Instruction","url":"https://viralinstruction.com/","person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"Viral Instruction","url":"https://viralinstruction.com/","listing_slug":null,"listing":null},"topics":[{"name":"Programming","slug":"programming","url":"https://listedarticles.com/topics/programming"},{"name":"Type Systems","slug":"type-systems","url":"https://listedarticles.com/topics/type-systems"},{"name":"Software","slug":"software","url":"https://listedarticles.com/topics/software"},{"name":"Engineering","slug":"engineering","url":"https://listedarticles.com/topics/engineering"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":1967,"reading_minutes":9,"published_at":"2021-10-06T00:00:00.000Z","added_at":"2026-09-20T09:07:51.994Z","updated_at":"2026-09-20T09:07:51.994Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":true},"profile_url":"https://listedarticles.com/articles/union-vs-sum-types","markdown_url":"https://listedarticles.com/articles/union-vs-sum-types.md","example":false,"citation":"Viral Instruction, Viral Instruction. \"Union vs sum types.\" 6 Oct 2021. https://viralinstruction.com/posts/uniontypes/ (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://viralinstruction.com/posts/uniontypes/"},"body_markdown":"*Written 2021-10-06*\n\nUnion types and sum types are programming language concepts that have been around for decades, but I think they're getting more popular these years. The two concepts are closely related but their subtle differences impacts their relative strengths. This post is an explanation of the concepts and a list of pros and cons of the two.\n\nUnion types, sum types and product types are all *algebraic data types*, which sound super complicated, but the basic concept is actually really simple.\n\nLet's begin somewhere familiar: With an ordinary struct. A database used by my job contains \"cases\" who are known by identifiers like this\n\n```\nstruct CaseID_V2 {\n    year: u16,\n    number: u32\n}\n```\nThis definition creates a new type `CaseID_V2`. We can think of a struct like an AND operator: `CaseID_V2` is a new type that is composed of a `u16` AND a `u32`.\n\nWhat is *type*, actually?\n\nWell, one can think about types as sets of possible values. Here for `u16`:\n\nWhat values of `CaseID_V2` are there? Well, if a `CaseID_V2` is a `u16` and a `u32`, then the set of possible `CaseID_V2` is simply the Cartesian product of the two types (e.g. all possible combinations of the two, denoted by ):\n\nAnd ta-da! That's why structs are called product types. That's really all there is to it.\n\nSometimes though, we want a new type which is not composed of one field AND another, but instead one field OR another. The same database at my work actually changed its `CaseID` in 2021, for some reason, hence the `_V2` suffix in the previous example. The old definition looked like this:\n\n```\nstruct CaseID_V1 {\n    numbers: u32,\n    letters: u32 // encoded in base36\n}\n```\nNow, any data type that contains a case ID must be able to have a notion of containing EITHER a `CaseID_V1` OR a `CaseID_V2`. We call such an either/or type a *union type*.\n\nIn pseudocode, it could look like:\n\n```\nunion type CaseID {\n    CaseID_V1,\n    CaseID_V2\n}\n```\nAnd we can then put *that* into a struct, if we want:\n\n```\nstruct Case {\n    id: CaseID,\n    creation: Date,\n    [ etc. ]\n}\n```\nWhy do we call it a union type? Well, similar to reason we call struct product types. The possible values in the new union type is the *union* of its members:\n\nSince its values are either `CaseID_V1` or `CaseID_V2`, clearly the set of possible values are just all the values that are in either set, or equivalently the union of the two sets.\n\nHere's a dilemma, though: What if we do this?\n\n```\nunion type MyType {\n    bool,\n    bool\n}\n```\nThis says that `MyType` is EITHER a `bool` OR a... `bool`? How many possible values is this?\n\nIt's still just the set ! In other words, `MyType` is equivalent to `bool`. Or one might even say it *is* `bool`.\n\nThat simplification is pretty neat, because it allows us to express uncertainty about types as union types, and do set operations on those. For example, suppose you have functions `f`, which returns the union , and `g` which returns  for four possible types total. If you now call either `f` OR `g`, what are your possible return types?\n\nIt's simply , \"deduplicated\" to just three types.\n\nA similar simplification happens if you union two types where one is a superset of the other. For example, suppose your language has a type `uint`, which just means \"any unsigned integer\", no matter its width. In that case  - after all, the set of values `uint`*contains* the set `u16`.\n\nSometimes when you program, you don't necessarily want that deduplication. Suppose you want to make a union type that contains *either* the year of the Gregorian calendar (stored in a `u16`), or the year according to the Hijri calendar (also stored in a `u16`). You can't express this as a union type , because in your case, these two `u16` are *different things*, that just happen to have the same representation, but shouldn't be conflated.\n\nThe solution is pretty straightforward: You create two new types that wrap the `u16`s, and serve as a \"type tag\" so the program knows how to interpret the data. Something like:\n\n```\nstruct Year_Gregorian {\n    val: u16\n}\nstruct Year_Hijri {\n    val: u16\n}\nunion type Year {\n    Year_Gregorian,\n    Year_Hijri\n}\n```\nThis kind of type - a union type with each member tagged - is called a *tagged union*. It's also called a *sum type*. By now you can guess why it's called a sum type: The number of values of type `Year` is exactly the sum of its members: .\n\nSum types are really useful when you want to be 100% sure you can distinguish all members of your union.\n\nRust calls sum types \"enums\" (a slight misnomer). You can make pretty complicated sum types very easily:\n\n```\nenum ComplicatedEnum {\n    IsEmpty,\n    Color(u8, u8, u8),\n    Name { given: String, sur: String }\n}\n```\nOne interesting catch about Rust's enums is this: Instead of defining three ordinary types `IsEmpty`, `Color` and `Name`, these three \"variants\" can only exist as part of an `ComplicatedEnum` and not on their own. This implies that no value can have the type `IsEmpty`: All values of `ComplicatedEnum` is just of the type `ComplicatedEnum`.\n\nI don't think there is any big theoretical reason for this \"forced wrapping\" of sum types in Rust, but it has important implications for practical use of Rust's sum types, which I'll get to in a bit.\n\nIn Julia, types matches perfectly well with the idea of \"types as sets of values\":\n\n```\njulia> 5 isa Int # check if 5 is an instance of Int\ntrue\njulia> 5 isa Union{Int, String}\ntrue\njulia> 5 isa Integer # Integer is a superset of Int\ntrue\njulia> 5 isa Union{String, Set, Char}\nfalse\njulia> Union{Int, Integer, Char, UInt, Int} # deduplication\nUnion{Char, Integer}\n```\nIn short, the value `5` belongs to both the types `Int`, `Union{Int, String}`, `Integer`, and an infinite number of other types.\n\nAnother difference from Rust is Julia is a dynamic language. Briefly, in static languages, expressions (e.g. code) has types, but types don't really exist at runtime since they are optimized away and everything is just a binary blob. In dynamic languages, values have types at runtime, and whatever type the compiler infer before runtime is immaterial: It has no impact on what values or types are actually produced at runtime.\n\nWhat this means is that, even if the compiler infers some value `x` to be of type `Union{A, B, C}`, at runtime, the type of `x` will be just `A`, `B` or `C`. Union types don't exist at runtime. They are only used to express the compiler's uncertainty about what is going to happen when the program runs.\n\nMany of the differences between Julia's and Rust's types actually come from the \"forced wrapping\" of Rust's sum types, not necessarily from the fact they are sum types instead of union types.\n\nIf you have an API that expects to be supplied with a `A`, then you can always change it to take a `Union{A, B}` without breakage, because all values of type `A` are also values of type `Union{A, B}`.\n\nSimilarly, if your function returns a `Union{A, B}`, you can change it to just return `A` without breakage.\n\nThis won't work in Rust: You can't change a function that took an `Option<usize>` to take a `usize` without breaking user's code, nor can you return `usize` where you previously returned `Option<usize>`.\n\nIn Rust, you can't access the variants of a sum type directly because they are always wrapped. This leads to a *lot* of boilerplate: Check the long list of methods for `Option` and `Result` which exist just for unwrapping and re-wrapping these types in various circumstances.\n\nWith Julia's system it's much easier: You don't unwrap and re-wrap because it's not wrapped in the first place. How do you add 1 to `x` if it's a `Union{Int, UInt}`? Just `x + 1`, like any normal integer.\n\nJust like it's not a breaking change to return a narrower union type or accept a broader one, it's also an allowed compiler change.\n\nSuppose you write a function `f` that returns `Union{A, B}` and you pass it into a function `g` expecting that. But now, in some code, you call `f` with one of the argument as a constant. The compiler will then check if that constant argument narrows down the return type of `f`. Let's say with the constant folded argument `f` is guaranteed to return `A`. If so, the compiler will then know `g` will be getting an `A`, not a `Union{A, B}` - so now `g` can be further optimized, for example by compiling away all branches that occur if the input is a `B`.\n\nJulia's union types may have less boilerplate because you can use them as if they were concrete types - but that's also a dangerous trap.\n\nConsider the Julia function `findfirst`, which returns `Union{Int, Nothing}` versus Rust's `iter.position`, returning `Option<usize>`: It's easy to forget `findfirst` can return nothing and not handle that case, introducing a bug. But it's not possible to mistaken an `Option<usize>` for a `usize`, because they're incompatible types and you *must* unwrap the sum type.\n\nThe intricate set operations possible with union types can also be pretty annoying when you're just trying to code. For example, suppose `f` is a function returning type `T`. What's the return type of this Rust code?\n\n`vec![f()]`\nYep, it's `Vec<T>`. Now what's the return type of this Julia code?\n\n`[f()]`\n`Vector{T}`, obviously! Right? Nope, not necessarily:\n\n```\njulia> f() = rand(Bool) ? 1 : nothing;\njulia> g() = [f()];\njulia> only(Core.Compiler.return_types(g, ()))\nUnion{Vector{Int64}, Vector{Nothing}}\n```\nInstead of a vector of unions, its a union of vectors. This must necessarily be true when you think about it, but it's just one of these examples where union types can \"pull the rug\" under you by suddenly doing something clever.\n\nThe Julia compiler optimizations mentioned above enabled by automatic restriction of union types are cute. But what if you have a union composed of, say 10 variants? If your language compiles specialized functions for every input type (\"monomorphization\"), as Julia and Rust does, this can cause an combinatorial explosion which leads to huge compilation times and bloated code. In fact, in Julia, this gets so bad that the compiler just gives up and emits code that checks the type at runtime if it infers that a value is a union with more than 4 members.\n\nIn this case, simply checking which variant you have with if/else statements is much more efficient than clever compiler tricks. Or even better than if/else statements...\n\nPrecisely because Rust's sum types don't do these clever type operations, the user can be confident that a sum type with variants `A`, `B` and `C` stays the same type with the same variants.\n\nThis enables *exhaustive pattern matching*: Pattern matching that will detect at compile time if you forget any edge cases. If you've used Rust for more than 5 minutes, you already know this is the best thing since sliced bread. If not, I *strongly* recommend you trying it out just so you know *how good* this would be to have in Your Favorite Language.\n\nThere are advantages to both union types and sum types. Quite fittingly, union types play to Julia's strengths: They enable expressive (low-boilerplate), generic and fast code. On the other hand, Rust's sum types enable code with predictable types, and much safer code through forced checking of edge cases.\n\nI'm not convinced this tradeoff between union and sum types is inherent. I think it may be possible to eat your cake and have it, too, but I'm not yet sure how such a system would look like.\n\nHopefully, that's a blog post - or a Julia package - for another time!","body_html":"<p><em>Written 2021-10-06</em></p>\n<p>Union types and sum types are programming language concepts that have been around for decades, but I think they&#39;re getting more popular these years. The two concepts are closely related but their subtle differences impacts their relative strengths. This post is an explanation of the concepts and a list of pros and cons of the two.</p>\n<p>Union types, sum types and product types are all <em>algebraic data types</em>, which sound super complicated, but the basic concept is actually really simple.</p>\n<p>Let&#39;s begin somewhere familiar: With an ordinary struct. A database used by my job contains &quot;cases&quot; who are known by identifiers like this</p>\n<pre><code>struct CaseID_V2 {\n    year: u16,\n    number: u32\n}</code></pre>\n<p>This definition creates a new type <code>CaseID_V2</code>. We can think of a struct like an AND operator: <code>CaseID_V2</code> is a new type that is composed of a <code>u16</code> AND a <code>u32</code>.</p>\n<p>What is <em>type</em>, actually?</p>\n<p>Well, one can think about types as sets of possible values. Here for <code>u16</code>:</p>\n<p>What values of <code>CaseID_V2</code> are there? Well, if a <code>CaseID_V2</code> is a <code>u16</code> and a <code>u32</code>, then the set of possible <code>CaseID_V2</code> is simply the Cartesian product of the two types (e.g. all possible combinations of the two, denoted by ):</p>\n<p>And ta-da! That&#39;s why structs are called product types. That&#39;s really all there is to it.</p>\n<p>Sometimes though, we want a new type which is not composed of one field AND another, but instead one field OR another. The same database at my work actually changed its <code>CaseID</code> in 2021, for some reason, hence the <code>_V2</code> suffix in the previous example. The old definition looked like this:</p>\n<pre><code>struct CaseID_V1 {\n    numbers: u32,\n    letters: u32 // encoded in base36\n}</code></pre>\n<p>Now, any data type that contains a case ID must be able to have a notion of containing EITHER a <code>CaseID_V1</code> OR a <code>CaseID_V2</code>. We call such an either/or type a <em>union type</em>.</p>\n<p>In pseudocode, it could look like:</p>\n<pre><code>union type CaseID {\n    CaseID_V1,\n    CaseID_V2\n}</code></pre>\n<p>And we can then put <em>that</em> into a struct, if we want:</p>\n<pre><code>struct Case {\n    id: CaseID,\n    creation: Date,\n    [ etc. ]\n}</code></pre>\n<p>Why do we call it a union type? Well, similar to reason we call struct product types. The possible values in the new union type is the <em>union</em> of its members:</p>\n<p>Since its values are either <code>CaseID_V1</code> or <code>CaseID_V2</code>, clearly the set of possible values are just all the values that are in either set, or equivalently the union of the two sets.</p>\n<p>Here&#39;s a dilemma, though: What if we do this?</p>\n<pre><code>union type MyType {\n    bool,\n    bool\n}</code></pre>\n<p>This says that <code>MyType</code> is EITHER a <code>bool</code> OR a... <code>bool</code>? How many possible values is this?</p>\n<p>It&#39;s still just the set ! In other words, <code>MyType</code> is equivalent to <code>bool</code>. Or one might even say it <em>is</em> <code>bool</code>.</p>\n<p>That simplification is pretty neat, because it allows us to express uncertainty about types as union types, and do set operations on those. For example, suppose you have functions <code>f</code>, which returns the union , and <code>g</code> which returns  for four possible types total. If you now call either <code>f</code> OR <code>g</code>, what are your possible return types?</p>\n<p>It&#39;s simply , &quot;deduplicated&quot; to just three types.</p>\n<p>A similar simplification happens if you union two types where one is a superset of the other. For example, suppose your language has a type <code>uint</code>, which just means &quot;any unsigned integer&quot;, no matter its width. In that case  - after all, the set of values <code>uint</code><em>contains</em> the set <code>u16</code>.</p>\n<p>Sometimes when you program, you don&#39;t necessarily want that deduplication. Suppose you want to make a union type that contains <em>either</em> the year of the Gregorian calendar (stored in a <code>u16</code>), or the year according to the Hijri calendar (also stored in a <code>u16</code>). You can&#39;t express this as a union type , because in your case, these two <code>u16</code> are <em>different things</em>, that just happen to have the same representation, but shouldn&#39;t be conflated.</p>\n<p>The solution is pretty straightforward: You create two new types that wrap the <code>u16</code>s, and serve as a &quot;type tag&quot; so the program knows how to interpret the data. Something like:</p>\n<pre><code>struct Year_Gregorian {\n    val: u16\n}\nstruct Year_Hijri {\n    val: u16\n}\nunion type Year {\n    Year_Gregorian,\n    Year_Hijri\n}</code></pre>\n<p>This kind of type - a union type with each member tagged - is called a <em>tagged union</em>. It&#39;s also called a <em>sum type</em>. By now you can guess why it&#39;s called a sum type: The number of values of type <code>Year</code> is exactly the sum of its members: .</p>\n<p>Sum types are really useful when you want to be 100% sure you can distinguish all members of your union.</p>\n<p>Rust calls sum types &quot;enums&quot; (a slight misnomer). You can make pretty complicated sum types very easily:</p>\n<pre><code>enum ComplicatedEnum {\n    IsEmpty,\n    Color(u8, u8, u8),\n    Name { given: String, sur: String }\n}</code></pre>\n<p>One interesting catch about Rust&#39;s enums is this: Instead of defining three ordinary types <code>IsEmpty</code>, <code>Color</code> and <code>Name</code>, these three &quot;variants&quot; can only exist as part of an <code>ComplicatedEnum</code> and not on their own. This implies that no value can have the type <code>IsEmpty</code>: All values of <code>ComplicatedEnum</code> is just of the type <code>ComplicatedEnum</code>.</p>\n<p>I don&#39;t think there is any big theoretical reason for this &quot;forced wrapping&quot; of sum types in Rust, but it has important implications for practical use of Rust&#39;s sum types, which I&#39;ll get to in a bit.</p>\n<p>In Julia, types matches perfectly well with the idea of &quot;types as sets of values&quot;:</p>\n<pre><code>julia&gt; 5 isa Int # check if 5 is an instance of Int\ntrue\njulia&gt; 5 isa Union{Int, String}\ntrue\njulia&gt; 5 isa Integer # Integer is a superset of Int\ntrue\njulia&gt; 5 isa Union{String, Set, Char}\nfalse\njulia&gt; Union{Int, Integer, Char, UInt, Int} # deduplication\nUnion{Char, Integer}</code></pre>\n<p>In short, the value <code>5</code> belongs to both the types <code>Int</code>, <code>Union{Int, String}</code>, <code>Integer</code>, and an infinite number of other types.</p>\n<p>Another difference from Rust is Julia is a dynamic language. Briefly, in static languages, expressions (e.g. code) has types, but types don&#39;t really exist at runtime since they are optimized away and everything is just a binary blob. In dynamic languages, values have types at runtime, and whatever type the compiler infer before runtime is immaterial: It has no impact on what values or types are actually produced at runtime.</p>\n<p>What this means is that, even if the compiler infers some value <code>x</code> to be of type <code>Union{A, B, C}</code>, at runtime, the type of <code>x</code> will be just <code>A</code>, <code>B</code> or <code>C</code>. Union types don&#39;t exist at runtime. They are only used to express the compiler&#39;s uncertainty about what is going to happen when the program runs.</p>\n<p>Many of the differences between Julia&#39;s and Rust&#39;s types actually come from the &quot;forced wrapping&quot; of Rust&#39;s sum types, not necessarily from the fact they are sum types instead of union types.</p>\n<p>If you have an API that expects to be supplied with a <code>A</code>, then you can always change it to take a <code>Union{A, B}</code> without breakage, because all values of type <code>A</code> are also values of type <code>Union{A, B}</code>.</p>\n<p>Similarly, if your function returns a <code>Union{A, B}</code>, you can change it to just return <code>A</code> without breakage.</p>\n<p>This won&#39;t work in Rust: You can&#39;t change a function that took an <code>Option&lt;usize&gt;</code> to take a <code>usize</code> without breaking user&#39;s code, nor can you return <code>usize</code> where you previously returned <code>Option&lt;usize&gt;</code>.</p>\n<p>In Rust, you can&#39;t access the variants of a sum type directly because they are always wrapped. This leads to a <em>lot</em> of boilerplate: Check the long list of methods for <code>Option</code> and <code>Result</code> which exist just for unwrapping and re-wrapping these types in various circumstances.</p>\n<p>With Julia&#39;s system it&#39;s much easier: You don&#39;t unwrap and re-wrap because it&#39;s not wrapped in the first place. How do you add 1 to <code>x</code> if it&#39;s a <code>Union{Int, UInt}</code>? Just <code>x + 1</code>, like any normal integer.</p>\n<p>Just like it&#39;s not a breaking change to return a narrower union type or accept a broader one, it&#39;s also an allowed compiler change.</p>\n<p>Suppose you write a function <code>f</code> that returns <code>Union{A, B}</code> and you pass it into a function <code>g</code> expecting that. But now, in some code, you call <code>f</code> with one of the argument as a constant. The compiler will then check if that constant argument narrows down the return type of <code>f</code>. Let&#39;s say with the constant folded argument <code>f</code> is guaranteed to return <code>A</code>. If so, the compiler will then know <code>g</code> will be getting an <code>A</code>, not a <code>Union{A, B}</code> - so now <code>g</code> can be further optimized, for example by compiling away all branches that occur if the input is a <code>B</code>.</p>\n<p>Julia&#39;s union types may have less boilerplate because you can use them as if they were concrete types - but that&#39;s also a dangerous trap.</p>\n<p>Consider the Julia function <code>findfirst</code>, which returns <code>Union{Int, Nothing}</code> versus Rust&#39;s <code>iter.position</code>, returning <code>Option&lt;usize&gt;</code>: It&#39;s easy to forget <code>findfirst</code> can return nothing and not handle that case, introducing a bug. But it&#39;s not possible to mistaken an <code>Option&lt;usize&gt;</code> for a <code>usize</code>, because they&#39;re incompatible types and you <em>must</em> unwrap the sum type.</p>\n<p>The intricate set operations possible with union types can also be pretty annoying when you&#39;re just trying to code. For example, suppose <code>f</code> is a function returning type <code>T</code>. What&#39;s the return type of this Rust code?</p>\n<p><code>vec![f()]</code>\nYep, it&#39;s <code>Vec&lt;T&gt;</code>. Now what&#39;s the return type of this Julia code?</p>\n<p><code>[f()]</code>\n<code>Vector{T}</code>, obviously! Right? Nope, not necessarily:</p>\n<pre><code>julia&gt; f() = rand(Bool) ? 1 : nothing;\njulia&gt; g() = [f()];\njulia&gt; only(Core.Compiler.return_types(g, ()))\nUnion{Vector{Int64}, Vector{Nothing}}</code></pre>\n<p>Instead of a vector of unions, its a union of vectors. This must necessarily be true when you think about it, but it&#39;s just one of these examples where union types can &quot;pull the rug&quot; under you by suddenly doing something clever.</p>\n<p>The Julia compiler optimizations mentioned above enabled by automatic restriction of union types are cute. But what if you have a union composed of, say 10 variants? If your language compiles specialized functions for every input type (&quot;monomorphization&quot;), as Julia and Rust does, this can cause an combinatorial explosion which leads to huge compilation times and bloated code. In fact, in Julia, this gets so bad that the compiler just gives up and emits code that checks the type at runtime if it infers that a value is a union with more than 4 members.</p>\n<p>In this case, simply checking which variant you have with if/else statements is much more efficient than clever compiler tricks. Or even better than if/else statements...</p>\n<p>Precisely because Rust&#39;s sum types don&#39;t do these clever type operations, the user can be confident that a sum type with variants <code>A</code>, <code>B</code> and <code>C</code> stays the same type with the same variants.</p>\n<p>This enables <em>exhaustive pattern matching</em>: Pattern matching that will detect at compile time if you forget any edge cases. If you&#39;ve used Rust for more than 5 minutes, you already know this is the best thing since sliced bread. If not, I <em>strongly</em> recommend you trying it out just so you know <em>how good</em> this would be to have in Your Favorite Language.</p>\n<p>There are advantages to both union types and sum types. Quite fittingly, union types play to Julia&#39;s strengths: They enable expressive (low-boilerplate), generic and fast code. On the other hand, Rust&#39;s sum types enable code with predictable types, and much safer code through forced checking of edge cases.</p>\n<p>I&#39;m not convinced this tradeoff between union and sum types is inherent. I think it may be possible to eat your cake and have it, too, but I&#39;m not yet sure how such a system would look like.</p>\n<p>Hopefully, that&#39;s a blog post - or a Julia package - for another time!</p>","headings":[]}}