{"article":{"slug":"english-a-vs-an","title":"English: a vs an","subtitle":"Why vowel letters are the wrong rule for indefinite articles in generated text","summary":"Amit Patel (Red Blob Games) digs into when English wants “a” versus “an”: the rule tracks spoken vowel sounds, not written vowel letters, and only a small set of common words need exceptions for procedural text generation.","content_type":"tutorial","language":"en","canonical_url":"https://www.redblobgames.com/blog/2026-09-16-english-a-vs-an/","author":{"name":"Amit Patel","url":"https://www.redblobgames.com/","person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"Red Blob Games","url":"https://www.redblobgames.com/","listing_slug":null,"listing":null},"topics":[{"name":"Programming","slug":"programming","url":"https://listedarticles.com/topics/programming"},{"name":"Writing","slug":"writing","url":"https://listedarticles.com/topics/writing"},{"name":"Algorithms","slug":"algorithms","url":"https://listedarticles.com/topics/algorithms"},{"name":"Education","slug":"education","url":"https://listedarticles.com/topics/education"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":274,"reading_minutes":1,"published_at":"2026-09-16T00:00:00.000Z","added_at":"2026-09-20T00:13:12.097Z","updated_at":"2026-09-20T00:13:12.097Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":true},"profile_url":"https://listedarticles.com/articles/english-a-vs-an","markdown_url":"https://listedarticles.com/articles/english-a-vs-an.md","example":false,"citation":"Amit Patel, Red Blob Games. \"English: a vs an.\" 16 Sept 2026. https://www.redblobgames.com/blog/2026-09-16-english-a-vs-an/ (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://www.redblobgames.com/blog/2026-09-16-english-a-vs-an/"},"body_markdown":"Blog post: 16 Sep 2026\n\nIn English, there is an “indefinite” article `a` that can go before a word. For example, `a raccoon`. But for some words, we use `an`. For example, `an apple`. \n\nWhen procedurally generating text, I want a function `a_or_an(\"apple\")` that tells me which article to use. That seems like it’d be easy. We can check the first letter to see if it’s a vowel. But that would mean we output `an unicorn`, not `a unicorn`. \n\nThe actual rule is not whether the _written_ word starts with a vowel letter, but whether the _spoken_ word starts with a vowel sound. The word `unicorn` starts with vowel letter (`u`) but a consonant sound (`Y`). The word `hour` starts with a consonant letter (`h`) but a vowel sound (`OW`). \n\n[![Tree style visualization of the first two letters of a word](/x/2635-a-vs-an/blog/level-2f.png?2026-08-25-11-10-15)](/x/2635-a-vs-an/blog/level-2f.png) Visualization showing whether the first two letters of a word are enough to determine whether it should have “a” or “an”\n\nI was curious how often these exceptions occurred, and whether they can be grouped together, so [I spent a day looking at the data and building some visualizations and wrote up the results.](/x/2635-a-vs-an/) I was surprised that only 129 of the 32,455 words in my list needed exceptions. \n\n[LLM note: I did _not_ use LLMs to write any of this code, but in hindsight, I should have. This is one-off code to answer a question. It doesn’t need to be clean or maintainable. It only needs to be correct. I would’ve spent more time on the trie simplification algorithm and less time on parsing cmudict and re-learning d3.js.]","body_html":"<p>Blog post: 16 Sep 2026</p>\n<p>In English, there is an “indefinite” article <code>a</code> that can go before a word. For example, <code>a raccoon</code>. But for some words, we use <code>an</code>. For example, <code>an apple</code>. </p>\n<p>When procedurally generating text, I want a function <code>a_or_an(&quot;apple&quot;)</code> that tells me which article to use. That seems like it’d be easy. We can check the first letter to see if it’s a vowel. But that would mean we output <code>an unicorn</code>, not <code>a unicorn</code>. </p>\n<p>The actual rule is not whether the <em>written</em> word starts with a vowel letter, but whether the <em>spoken</em> word starts with a vowel sound. The word <code>unicorn</code> starts with vowel letter (<code>u</code>) but a consonant sound (<code>Y</code>). The word <code>hour</code> starts with a consonant letter (<code>h</code>) but a vowel sound (<code>OW</code>). </p>\n<p><a href=\"/x/2635-a-vs-an/blog/level-2f.png\">Tree style visualization of the first two letters of a word</a> Visualization showing whether the first two letters of a word are enough to determine whether it should have “a” or “an”</p>\n<p>I was curious how often these exceptions occurred, and whether they can be grouped together, so <a href=\"/x/2635-a-vs-an/\">I spent a day looking at the data and building some visualizations and wrote up the results.</a> I was surprised that only 129 of the 32,455 words in my list needed exceptions. </p>\n<p>[LLM note: I did <em>not</em> use LLMs to write any of this code, but in hindsight, I should have. This is one-off code to answer a question. It doesn’t need to be clean or maintainable. It only needs to be correct. I would’ve spent more time on the trie simplification algorithm and less time on parsing cmudict and re-learning d3.js.]</p>","headings":[]}}