{"article":{"slug":"readable-regular-expressions-for-javascript-typescript-inspired-by-emacs-rx","title":"Readable Regular Expressions for JavaScript/TypeScript, Inspired by Emacs' rx","subtitle":null,"summary":"A tutorial bringing Emacs-style rx composable regex DSLs to JavaScript and TypeScript so complex patterns stay readable and maintainable.","content_type":"tutorial","language":"en","canonical_url":"https://rahuljuliato.com/posts/emacs-rx-in-typescript","author":{"name":"Rahul M. Juliato","url":"https://rahuljuliato.com/","person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"Rahul's Blog","url":"https://rahuljuliato.com/","listing_slug":null,"listing":null},"topics":[{"name":"Tutorials","slug":"tutorials","url":"https://listedarticles.com/topics/tutorials"},{"name":"Programming","slug":"programming","url":"https://listedarticles.com/topics/programming"},{"name":"Developer Tools","slug":"developer-tools","url":"https://listedarticles.com/topics/developer-tools"},{"name":"Open Source","slug":"open-source","url":"https://listedarticles.com/topics/open-source"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":4193,"reading_minutes":18,"published_at":"2026-09-30T00:00:00.000Z","added_at":"2026-10-02T18:14:59.256Z","updated_at":"2026-10-02T18:14:59.256Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":true},"profile_url":"https://listedarticles.com/articles/readable-regular-expressions-for-javascript-typescript-inspired-by-emacs-rx","markdown_url":"https://listedarticles.com/articles/readable-regular-expressions-for-javascript-typescript-inspired-by-emacs-rx.md","example":false,"citation":"Rahul M. Juliato, Rahul's Blog. \"Readable Regular Expressions for JavaScript/TypeScript, Inspired by Emacs' rx.\" 30 Sept 2026. https://rahuljuliato.com/posts/emacs-rx-in-typescript (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://rahuljuliato.com/posts/emacs-rx-in-typescript"},"body_markdown":"# Readable Regular Expressions for JavaScript/TypeScript, Inspired by Emacs' rx\n\n## Intro\n\nQuick, what does this match?\n\n```\n/^(0|[1-9]\\d*)\\.(0|[1-9]\\d*)\\.(0|[1-9]\\d*)(?:-((?:0|[1-9]\\d*|\\d*[a-zA-Z-][0-9a-zA-Z-]*)(?:\\.(?:0|[1-9]\\d*|\\d*[a-zA-Z-][0-9a-zA-Z-]*))*))?(?:\\+([0-9a-zA-Z-]+(?:\\.[0-9a-zA-Z-]+)*))?$/;\n// Take\n// your\n// time...\n//\n// ...still decoding?\n//\n// OK, keep reading :)\n```\nThat's the official regexp from semver.org. It validates version numbers like:\n\n```\n// matches\n\"1.2.3\"\n\"0.10.0\"\n\"2.0.0-rc.1\"\n\"1.0.0-alpha.1+build.5\"\n\"1.0.0+20260930\"\n// doesn't match\n\"01.2.3\"   // leading zero\n\"1.2\"      // missing patch\n\"v1.2.3\"   // no \"v\" prefix allowed\n\"1.0.0-01\" // numeric pre-release with a leading zero\n\"1.2.3-\"   // empty pre-release\n```\nDon't get me wrong, I love regexps, but in practice you probably spend a bunch of time writing one, testing it against some cases, and moving on, proud of your achievement!\n\nSome time passes and lucky future you (or unlucky someone else) has to change it. Dramatic pause here.\n\nI bet you've been there. Now your options are probably: decode it again from the start, rewrite the whole thing, or, in the age of AI, ask (and hopefully not blindly accept) an LLM for a new recipe.\n\nEmacs has had a nice answer for more readable regexps for a long time:\nthe `rx` macro. I started using it all the time in Emacs Lisp, as\nreviewers always suggested it to me. Later, I started missing this DSL\nin JavaScript and TypeScript, so I wrote a small version of it for my\nprojects.\n\nSo, what about reading that SemVer regexp like `semver` in the code\nbelow?\n\n```\nconst num = or(\"0\", seq(anyOf(\"1-9\"), zeroOrMore(digit)));\nconst idChar = anyOf(alnum, \"-\");\nconst preId = or(num, seq(zeroOrMore(digit), anyOf(alpha, \"-\"), zeroOrMore(idChar)));\nconst dotted = (x: Item) => seq(x, zeroOrMore(\".\", x));\nconst semver = RX(\n  start,\n  named(\"major\", num), \".\",\n  named(\"minor\", num), \".\",\n  named(\"patch\", num),\n  optional(\"-\", named(\"pre\", dotted(preId))),\n  optional(\"+\", named(\"build\", dotted(oneOrMore(idChar)))),\n  end,\n);\n```\nThe same strings match, and you get named groups as a bonus. By the end of this post you'll know every piece of it.\n\n**TL;DR:** jump straight to the cheat sheet,\nthe side-by-side examples, the full\nsource, or grab the\ngist\nto sneak a peek at the result.\n\n\n**NOTE:** the `RX` here has nothing to do with\nRxJS, which is an amazing library for reactive\nprogramming with observables.\n\n\n## A taste of rx in Emacs Lisp\n\nWith `rx` you describe a regexp as a tree of named forms, and Emacs\nturns it into the regexp string for you:\n\n```\n(rx bos (+ digit) eos)\n;; => \"\\\\`[[:digit:]]+\\\\'\"\n(rx bol \"colo\" (? \"u\") \"r\" eol)\n;; => \"^colou?r$\"\n(rx bos \"(\" (= 3 digit) \")\" space (= 3 digit) \"-\" (= 4 digit) eos)\n;; => \"\\\\`([[:digit:]]\\\\{3\\\\})[[:space:]][[:digit:]]\\\\{3\\\\}-[[:digit:]]\\\\{4\\\\}\\\\'\"\n```\nA few things to notice:\n\n1. **Strings are literals.**`\"(\"` means a parenthesis. You don't\nneed to escape anything by hand.\n2. **Sequence is implicit.** Every form takes a list of things and\nmatches them one after the other. You don't need to wrap them in a`seq` , even though`seq` exists.\n3. **Groups appear only when needed.**`(+ digit)` becomes`[[:digit:]]+` , not`\\(?:[[:digit:]]\\)+` .\n\nThe proposed JavaScript/TypeScript version in this post reads like this:\n\n```\nconst phone = RX(\n  start, \"(\", repeat(3, digit), \")\", space,\n  repeat(3, digit), \"-\", repeat(4, digit), end,\n);\n// => /^\\(\\d{3}\\)\\s\\d{3}-\\d{4}$/\n```\n## Under the hood\n\nIf you want strings to be literals, you can't represent a regexp piece\nas a plain `string`, otherwise you can't tell `\"(\"` (a literal\nparenthesis) apart from `\"(?:...)\"` (a group you built). So every\npiece is a small object:\n\n```\ntype Kind = \"atom\" | \"seq\" | \"alt\";\ninterface RxNode {\n  readonly src: string;\n  readonly kind: Kind;\n  readonly set?: string; // char sets only, see below\n  readonly neg?: boolean;\n}\ntype Item = string | RxNode;\n```\n`src` is the regexp text. `kind` records how that text behaves when\nyou glue it to other things:\n\n- `atom` : a single unit, like`a` ,`\\d` ,`[a-z]` or`(...)` . You can\nput a quantifier right after it.\n- `seq` : safe to concatenate, but a quantifier needs`(?:...)` around\nit.`abc` is a`seq` , and so is`a+` , since`a+?` would silently\nturn into a lazy quantifier.\n- `alt` : has a`|` at the top level, so it needs`(?:...)` almost\neverywhere.\n\nPlain strings go through `literal`, which escapes them:\n\n```\nconst esc = (s: string) => s.replace(/[.*+?^${}()|[\\]\\\\]/g, \"\\\\$&\");\nconst literal = (s: string): RxNode => ({\n  src: esc(s),\n  kind: s.length === 1 ? \"atom\" : \"seq\",\n});\nconst toNode = (x: Item): RxNode => (typeof x === \"string\" ? literal(x) : x);\n```\nWith that in place, `seq` joins nodes and only brackets alternations:\n\n```\nconst seq = (...xs: Item[]): RxNode => {\n  const nodes = xs.map(toNode).filter((n) => n.src !== \"\");\n  if (nodes.length === 0) return { src: \"\", kind: \"seq\" };\n  if (nodes.length === 1) return nodes[0];\n  let src = \"\";\n  for (const n of nodes) {\n\tconst part = n.kind === \"alt\" ? `(?:${n.src})` : n.src;\n\t// `\\1` followed by a literal `0` would read as `\\10`\n\tif (/\\\\\\d+$/.test(src) && /^\\d/.test(part)) src += \"(?:)\";\n\tsrc += part;\n  }\n  return { src, kind: \"seq\" };\n};\n```\n(That backreference check is one of those bugs you only find by\nwriting tests, or when it happens to you in prod. `backref(1)`\nfollowed by the literal `\"0\"` gives you backreference number ten.)\n\nEvery quantifier is a `seq` of its arguments plus a suffix, bracketed\nonly when the body isn't an atom:\n\n```\nconst quantifiable = (n: RxNode) =>\n  n.kind === \"atom\" ? n.src : `(?:${n.src})`;\nconst quantifier =\n  (suffix: string) =>\n  (...xs: Item[]): RxNode => ({\n\tsrc: quantifiable(seq(...xs)) + suffix,\n\tkind: \"seq\",\n  });\nconst zeroOrMore = quantifier(\"*\");\nconst oneOrMore = quantifier(\"+\");\nconst optional = quantifier(\"?\");\n```\nBecause each quantifier calls `seq` on its arguments, you get the\nimplicit sequence for free: `optional(\"-\", group(x))` becomes\n`(?:-(x))?`.\n\nAnd finally, the two entry points. As in Emacs, `rx` returns a\nstring. `RX` returns a `RegExp` you can use right away:\n\n```\nconst rx = (...xs: Item[]): string => seq(...xs).src;\nfunction RX(...xs: Item[]): RegExp {\n  return new RegExp(rx(...xs));\n}\nRX.flags = (flags: string, ...xs: Item[]): RegExp =>\n  new RegExp(rx(...xs), flags);\n```\n`RX.flags` exists because Emacs controls case folding through the\n`case-fold-search` variable, and JavaScript puts it on the regexp\nitself.\n\nThat's the whole engine! Now, let's build our vocabulary.\n\n## Character sets\n\nIn Emacs you write `(any \"a-z\" \"_\")`. Inside those strings, `a-z` is a\nrange, and a `-` at either end is a plain dash. I kept the same rule:\n\n```\nconst hexDigit = anyOf(\"0-9a-fA-F\");\nRX(start, \"#\", repeat(6, hexDigit), end);\n// => /^#[0-9a-fA-F]{6}$/\nRX(start, optional(anyOf(\"+-\")), oneOrMore(digit), end);\n// => /^[+\\-]?\\d+$/\n```\nThe dash comes out escaped because sets can merge. If you combine\n`anyOf(\"+-\")` with `anyOf(\"0-9\")`, an unescaped `-` would end up in\nthe middle and create a range from `+` to `0`. Escaping it costs one\nbackslash.\n\nAnd merging is the reason why `RxNode` has a `set` field. It is there\nto hold the text that goes between `[` and `]`, so `anyOf` can take\nother sets as arguments:\n\n```\nconst lower = anyOf(\"a-z\");\nconst upper = anyOf(\"A-Z\");\nconst alpha = anyOf(lower, upper);   // [a-zA-Z]\nconst alnum = anyOf(alpha, \"0-9\");   // [a-zA-Z0-9]\n```\n`not` negates a set, and it knows the shorthand classes:\n\n```\nnot(digit);                // \\D\nnot(anyOf(space, \"@\"));    // [^\\s@]\nnotChar(\",\");              // [^,]   (rx's not-char)\n```\nThe simple email check, which most of us have written as\n`/^[^\\s@]+@[^\\s@]+\\.[^\\s@]+$/` at some point, becomes:\n\n```\nconst part = oneOrMore(not(anyOf(space, \"@\")));\nRX(start, part, \"@\", part, \".\", part, end);\n// => /^[^\\s@]+@[^\\s@]+\\.[^\\s@]+$/\n```\nThe rest of the Emacs character classes are there too: `digit`,\n`hexDigit`, `space`, `blank`, `wordChar`, `notWordChar`, `alpha`,\n`alnum`, `lower`, `upper`, `punct`, `control`, `graphic`, `printing`,\n`ascii` and `nonascii`. One difference: in Emacs they understand\nUnicode, and mine are ASCII only. `alpha` won't match `é`.\n\nTwo more come from rx's symbol list, and people (me, many times) mix them up:\n\n```\nconst notNewline: RxNode = { src: \".\", kind: \"atom\" }; // rx: nonl\nconst anything = set(\"\\\\s\\\\S\");                        // rx: anything / anychar\n```\nIn rx, `anything` really means anything, newlines included. Here is\nwhere the difference shows up:\n\n```\nconst code = \"a = 1; /* first\\n   second */ b = 2;\";\nRX(\"/*\", zeroOrMoreLazy(notNewline), \"*/\").exec(code);\n// => null\nRX(\"/*\", zeroOrMoreLazy(anything), \"*/\").exec(code)?.[0];\n// => \"/* first\\n   second */\"\n```\n## Alternatives, and the longest match\n\n`or` works as you'd expect, and gets bracketed when it lands inside a\nsequence:\n\n```\nRX(start, or(\"cat\", \"dog\", \"bird\"), end);\n// => /^(?:bird|cat|dog)$/\n```\nDid you notice the order changed? I copied that behavior from\nEmacs. When every branch of an `or` is a plain string, rx hands them\nto `regexp-opt`, which builds a pattern that prefers the longest\nmatch:\n\n```\n(rx (or \"in\" \"int\" \"interface\"))\n;; => \"\\\\(?:in\\\\(?:t\\\\(?:erface\\\\)?\\\\)?\\\\)\"\n```\nJavaScript alternation takes the first branch that matches, going left to right. So the naive regexp for a list of keywords has a 'bug':\n\n```\n/in|int|interface/.exec(\"interface Foo\")?.[0];\n// => \"in\"\nRX(or(\"in\", \"int\", \"interface\")).exec(\"interface Foo\")?.[0];\n// => \"interface\"\n```\nI don't build a trie like `regexp-opt` does. Sorting the strings by\nlength, longest first, is enough to get the same behavior:\n\n```\nconst or = (...xs: Item[]): RxNode => {\n  if (xs.length === 0) return unmatchable;\n  if (xs.length === 1) return toNode(xs[0]);\n  const branches = xs.every((x) => typeof x === \"string\")\n\t? [...(xs as string[])].sort((a, b) => b.length - a.length)\n\t: xs;\n  return { src: branches.map((x) => toNode(x).src).join(\"|\"), kind: \"alt\" };\n};\n```\nAs in Emacs, `or()` with no branches returns `unmatchable`, which is\n`(?!)` here. It's handy when you build the branch list at runtime and\nit might come out empty.\n\n## Repetition, greedy and lazy\n\nEmacs has `(= n ...)`, `(>= n ...)` and `(** n m ...)`. Here they are\n`repeat`, `atLeast` and `between`:\n\n```\nRX(start, between(2, 4, digit), end);  // /^\\d{2,4}$/\nRX(atLeast(3, digit));                 // /\\d{3,}/\n```\nThe lazy versions `*?`, `+?` and `??` are `zeroOrMoreLazy`,\n`oneOrMoreLazy` and `optionalLazy`. The classic HTML tag example:\n\n```\nconst html = \"<b>bold</b> and <i>italic</i>\";\nRX(\"<\", oneOrMore(notNewline), \">\").exec(html)?.[0];\n// => \"<b>bold</b> and <i>italic</i>\"\nRX(\"<\", oneOrMoreLazy(notNewline), \">\").exec(html)?.[0];\n// => \"<b>\"\n```\n## Groups and backreferences\n\n`group` is a capturing group, and `backref` points back to it:\n\n```\nRX(start, group(oneOrMore(wordChar)), space, backref(1), end);\n// => /^(\\w+)\\s\\1$/   matches \"hello hello\", not \"hello world\"\n```\nEmacs also has `(group-n N ...)` to pick the group number. JavaScript\ncan't do that, but it has named groups, which serve the same purpose\nand read better:\n\n```\nconst date = RX(\n  named(\"y\", repeat(4, digit)), \"-\",\n  named(\"m\", repeat(2, digit)), \"-\",\n  named(\"d\", repeat(2, digit)),\n);\ndate.exec(\"2026-09-30\")?.groups;\n// => { y: '2026', m: '09', d: '30' }\n```\n`backref` accepts a name as well:\n\n```\nRX(\n  \"<\", named(\"tag\", oneOrMore(wordChar)), \">\",\n  zeroOrMoreLazy(notNewline),\n  \"</\", backref(\"tag\"), \">\",\n);\n// => /<(?<tag>\\w+)>.*?<\\/\\k<tag>>/\n```\n## Anchors\n\n`rx` distinguishes the start of the string (`bos`) from the start of a\nline (`bol`). In JavaScript both are `^`, and the `m` flag decides\nwhich one you get. I kept both names so the intent shows in the code:\n\n```\nconst text = \"TODO: write post\\nDONE: fix rx\\nTODO: publish\";\nconst todo = RX.flags(\n  \"gm\",\n  lineStart, \"TODO: \", named(\"task\", oneOrMore(notNewline)), lineEnd,\n);\n[...text.matchAll(todo)].map((m) => m.groups?.task);\n// => [ 'write post', 'publish' ]\n```\nWhy not add `start` and `end` automatically? Because you only want\nthem when validating a whole string. When searching inside a text, as\nin `split`, `replace` or `matchAll`, a hidden `^` and `$` would break\neverything. Emacs agrees: `bos` and `eos` are explicit in `rx` too.\n\n`wordBoundary` and `notWordBoundary` map straight to `\\b` and `\\B`.\nEmacs also has `bow` and `eow` (`\\<` and `\\>`), start and end of a\nword. JavaScript lacks those, so I combined `\\b` with a lookaround:\n\n```\nconst wordStart = zeroWidth(\"\\\\b(?=\\\\w)\");\nconst wordEnd = zeroWidth(\"\\\\b(?<=\\\\w)\");\n```\n## Literal and raw\n\nPlain strings are already literals, but `rx` has an explicit `literal`\nform for strings computed at runtime, and I kept it. It documents that\nthe value came from somewhere else:\n\n```\nconst userInput = \"1+1=2? (maybe)\";\nnew RegExp(userInput).test(userInput);  // false, oops\nRX(literal(userInput)).test(userInput); // true\n```\nThe opposite direction is `rx`'s `(regexp ...)` form, the escape hatch.\nHere it's `raw`, and it receives either a string or an existing\n`RegExp`. It lets you adopt the DSL in a codebase full of old regexps\nwithout rewriting all of them, like:\n\n```\nconst legacyZip = /\\d{5}(?:-\\d{4})?/;\nRX(start, repeat(2, upper), \" \", raw(legacyZip), end);\n// => /^[A-Z]{2} (?:\\d{5}(?:-\\d{4})?)$/\n```\n`raw` can't see inside the text it gets, so it adds brackets whenever\nit's combined with something else. It's an extra `(?:)`, and the\nregexp still works.\n\n## Shall we try SemVer again?\n\nBack to the regexp from the intro. In Emacs you would give names to\nthe pieces with `rx-define` or `rx-let`. In TypeScript those are just\n`const`s:\n\n```\nconst num = or(\"0\", seq(anyOf(\"1-9\"), zeroOrMore(digit)));\nconst idChar = anyOf(alnum, \"-\");\nconst preId = or(num, seq(zeroOrMore(digit), anyOf(alpha, \"-\"), zeroOrMore(idChar)));\nconst dotted = (x: Item) => seq(x, zeroOrMore(\".\", x));\nconst semver = RX(\n  start,\n  named(\"major\", num), \".\",\n  named(\"minor\", num), \".\",\n  named(\"patch\", num),\n  optional(\"-\", named(\"pre\", dotted(preId))),\n  optional(\"+\", named(\"build\", dotted(oneOrMore(idChar)))),\n  end,\n);\n```\nNow you can read the spec in the code. A numeric identifier is `0`, or\na non-zero digit followed by any number of digits. A pre-release is a\ndotted list of identifiers, and so is build metadata. `dotted` is a\nplain function returning a node, which is as far as abstraction needs\nto go here.\n\nIt matches the same strings as the official regexp, and the named groups give you a result like:\n\n```\nsemver.exec(\"1.0.0-alpha.1+build.5\")?.groups;\n// => {\n//   major: '1',\n//   minor: '0',\n//   patch: '0',\n//   pre: 'alpha.1',\n//   build: 'build.5'\n// }\n```\nNext time the spec changes, you can understand what the current regex does at a glance, instead of fighting an army of punctuation.\n\n## What's missing from Emacs rx\n\nI tried to map every `rx` form, and a few have no JavaScript equivalent:\n\n- `point` : JavaScript regexps don't know about a cursor.\n- `symbol-start` ,`symbol-end` ,`syntax` ,`category` : these depend on\nEmacs syntax tables.\n- `intersection` : possible with the`v` flag, but I haven't needed it.\n- `minimal-match` /`maximal-match` : these flip the greediness of\neverything inside them. Doable, but it would need a separate pass,\nand the`*Lazy` functions cover my use cases.\n- `eval` : TypeScript already evaluates expressions everywhere, so you\nget it for free.\n\nAnd one addition Emacs doesn't need: `RX.flags`.\n\n## Cheat sheet\n\n| Emacs rx | TypeScript | JS regexp (roughly) | \n|---|---|---|\n| `seq` ,`:` ,`and` | `seq(...)` , implicit in every form | `ab` | \n| `or` ,`\\|` | `or(...)` | `a\\|b` | \n| `any` ,`in` ,`char` | `anyOf(\"a-z\", \"_\", digit)` | `[a-z_\\d]` | \n| `not-char` | `notChar(...)` | `[^...]` | \n| `not` | `not(charset)` | `\\D` ,`[^...]` | \n| `*` ,`+` ,`?` | `zeroOrMore` ,`oneOrMore` ,`optional` | `x*` ,`x+` ,`x?` | \n| `*?` ,`+?` ,`??` | `zeroOrMoreLazy` ,`oneOrMoreLazy` ,`optionalLazy` | `x*?` ,`x+?` ,`x??` | \n| `=` ,`>=` ,`**` | `repeat` ,`atLeast` ,`between` | `x{n}` ,`x{n,}` ,`x{n,m}` | \n| `group` | `group(...)` | `(...)` | \n| `group-n` | `named(\"name\", ...)` | `(?<name>...)` | \n| `backref` | `backref(1)` ,`backref(\"name\")` | `\\1` ,`\\k<name>` | \n| `literal` | `literal(s)` | `s` , escaped:`1\\+1` | \n| `regexp` ,`regex` | `raw(\"...\")` ,`raw(/.../)` | `(?:...)` , as-is | \n| `rx-define` ,`rx-let` | `const` | (none) | \n| `bos` ,`eos` | `start` ,`end` | `^` ,`$` | \n| `bol` ,`eol` | `lineStart` ,`lineEnd` (with the`m` flag) | `^` ,`$` | \n| `bow` ,`eow` | `wordStart` ,`wordEnd` | `\\b(?=\\w)` ,`\\b(?<=\\w)` | \n| `word-boundary` | `wordBoundary` | `\\b` | \n| `not-word-boundary` | `notWordBoundary` | `\\B` | \n| `nonl` ,`not-newline` | `notNewline` | `.` | \n| `anychar` ,`anything` | `anything` | `[\\s\\S]` | \n| `unmatchable` | `unmatchable` | `(?!)` | \n| `digit` | `digit` | `\\d` | \n| `hex-digit` ,`xdigit` | `hexDigit` | `[0-9a-fA-F]` | \n| `space` ,`whitespace` | `space` | `\\s` | \n| `blank` | `blank` | `[ \\t]` | \n| `word` ,`wordchar` | `wordChar` | `\\w` | \n| `not-wordchar` | `notWordChar` | `\\W` | \n| `alpha` ,`letter` | `alpha` | `[a-zA-Z]` | \n| `alnum` | `alnum` | `[a-zA-Z0-9]` | \n| `lower` ,`upper` | `lower` ,`upper` | `[a-z]` ,`[A-Z]` | \n| `punct` ,`punctuation` | `punct` | ``[!-/:-@[-`{-~]`` | \n| `cntrl` ,`control` | `control` | `[\\x00-\\x1f\\x7f]` | \n| `graph` ,`graphic` | `graphic` | `[!-~]` | \n| `print` ,`printing` | `printing` | `[ -~]` | \n| `ascii` ,`nonascii` | `ascii` ,`nonascii` | `[\\x00-\\x7f]` ,`[\\u0080-\\uffff]` | \n\n## Examples: JS/TS regex vs RX\n\nEach example below shows the goal, the Emacs `rx` form in a comment,\nthe regexp you would write by hand, and the `RX` version. When `RX`\nproduces a different regexp text, the `// =>` line shows it. The\nresults at the bottom come from running both against the same strings.\n\n### Digits only\n\nThe whole string is digits.\n\n```\n// Emacs: (rx bos (+ digit) eos)\nconst regex = /^\\d+$/;\nconst dsl = RX(start, oneOrMore(digit), end);\n// \"123\" -> true\n// \"12a\" -> false\n```\n### Letters only\n\nThe whole string is ASCII letters.\n\n```\n// Emacs: (rx bos (+ alpha) eos)\nconst regex = /^[a-zA-Z]+$/;\nconst dsl = RX(start, oneOrMore(alpha), end);\n// \"Hello\" -> true\n// \"He11o\" -> false\n```\n### Optional letter\n\nBoth spellings, color and colour.\n\n```\n// Emacs: (rx bos \"colo\" (? \"u\") \"r\" eos)\nconst regex = /^colou?r$/;\nconst dsl = RX(start, \"colo\", optional(\"u\"), \"r\", end);\n// \"color\" -> true\n// \"colour\" -> true\n// \"colouur\" -> false\n```\n### Two words\n\nTwo words separated by a space.\n\n```\n// Emacs: (rx bos (+ wordchar) space (+ wordchar) eos)\nconst regex = /^\\w+\\s\\w+$/;\nconst word = oneOrMore(wordChar);\nconst dsl = RX(start, word, space, word, end);\n// \"hello world\" -> true\n// \"hello\" -> false\n```\n### Phone number\n\n(123) 456-7890, parentheses and all.\n\n```\n// Emacs: (rx bos \"(\" (= 3 digit) \")\" space (= 3 digit) \"-\" (= 4 digit) eos)\nconst regex = /^\\(\\d{3}\\)\\s\\d{3}-\\d{4}$/;\nconst dsl = RX(\n  start, \"(\", repeat(3, digit), \")\", space,\n  repeat(3, digit), \"-\", repeat(4, digit), end,\n);\n// \"(123) 456-7890\" -> true\n// \"123-456-7890\" -> false\n```\n### Hex color\n\n`#ff00aa`-style colors.\n\n```\n// Emacs: (rx bos \"#\" (= 6 hex-digit) eos)\nconst regex = /^#[0-9a-fA-F]{6}$/;\nconst dsl = RX(start, \"#\", repeat(6, hexDigit), end);\n// \"#ff00aa\" -> true\n// \"#ff00ag\" -> false\n```\n### Signed integer\n\nAn optional sign, then digits.\n\n```\n// Emacs: (rx bos (? (any \"+-\")) (+ digit) eos)\nconst regex = /^[+-]?\\d+$/;\nconst dsl = RX(start, optional(anyOf(\"+-\")), oneOrMore(digit), end);\n// => /^[+\\-]?\\d+$/\n// \"-42\" -> true\n// \"42\" -> true\n// \"*42\" -> false\n```\n### Simple email\n\nSomething@something.something, no spaces.\n\n```\n// Emacs: (rx-let ((part (+ (not (any space \"@\")))))\n//   (rx bos part \"@\" part \".\" part eos))\nconst regex = /^[^\\s@]+@[^\\s@]+\\.[^\\s@]+$/;\nconst part = oneOrMore(not(anyOf(space, \"@\")));\nconst dsl = RX(start, part, \"@\", part, \".\", part, end);\n// \"a@b.com\" -> true\n// \"a b@c.com\" -> false\n// \"a@b\" -> false\n```\n### No digits\n\nA string without any digit.\n\n```\n// Emacs: (rx bos (+ (not digit)) eos)\nconst regex = /^[^\\d]+$/;\nconst dsl = RX(start, oneOrMore(not(digit)), end);\n// => /^\\D+$/\n// \"abc\" -> true\n// \"a1b\" -> false\n```\n### CSV line\n\nExactly three comma-separated fields.\n\n```\n// Emacs: (rx-let ((field (+ (not-char \",\"))))\n//   (rx bos field \",\" field \",\" field eos))\nconst regex = /^[^,]+,[^,]+,[^,]+$/;\nconst field = oneOrMore(notChar(\",\"));\nconst dsl = RX(start, field, \",\", field, \",\", field, end);\n// \"a,b,c\" -> true\n// \"a,b,\" -> false\n```\n### One of many\n\nA fixed list of words.\n\n```\n// Emacs: (rx bos (or \"cat\" \"dog\" \"bird\") eos)\nconst regex = /^(?:cat|dog|bird)$/;\nconst dsl = RX(start, or(\"cat\", \"dog\", \"bird\"), end);\n// => /^(?:bird|cat|dog)$/\n// \"cat\" -> true\n// \"bird\" -> true\n// \"cow\" -> false\n```\n### Title and name\n\n`mr` or `ms`, then a name, keeping the title.\n\n```\n// Emacs: (rx bos (group (or \"mr\" \"ms\")) space (+ wordchar) eos)\nconst regex = /^(mr|ms)\\s\\w+$/;\nconst dsl = RX(start, group(or(\"mr\", \"ms\")), space, oneOrMore(wordChar), end);\n// \"mr john\" -> true\n// \"dr john\" -> false\n```\n### Between\n\nTwo to four digits.\n\n```\n// Emacs: (rx bos (** 2 4 digit) eos)\nconst regex = /^\\d{2,4}$/;\nconst dsl = RX(start, between(2, 4, digit), end);\n// \"12\" -> true\n// \"12345\" -> false\n```\n### At least\n\nThree or more digits, anywhere.\n\n```\n// Emacs: (rx (>= 3 digit))\nconst regex = /\\d{3,}/;\nconst dsl = RX(atLeast(3, digit));\n// \"12\" -> false\n// \"a123\" -> true\n```\n### Repeated word\n\nThe same word twice.\n\n```\n// Emacs: (rx bos (group (+ wordchar)) space (backref 1) eos)\nconst regex = /^(\\w+)\\s\\1$/;\nconst dsl = RX(start, group(oneOrMore(wordChar)), space, backref(1), end);\n// \"hello hello\" -> true\n// \"hello world\" -> false\n```\n### Matching tags\n\nAn open tag and its own closing tag.\n\n```\n// Emacs: (rx \"<\" (group-n 1 (+ wordchar)) \">\" (*? nonl) \"</\" (backref 1) \">\")\nconst regex = /<(?<tag>\\w+)>.*?<\\/\\k<tag>>/;\nconst dsl = RX(\n  \"<\", named(\"tag\", oneOrMore(wordChar)), \">\",\n  zeroOrMoreLazy(notNewline),\n  \"</\", backref(\"tag\"), \">\",\n);\n// \"<b>bold</b>\" -> true\n// \"<b>oops</i>\" -> false\n```\n### Whole word\n\n`cat` as a word, not inside another one.\n\n```\n// Emacs: (rx word-boundary \"cat\" word-boundary)\nconst regex = /\\bcat\\b/;\nconst dsl = RX(wordBoundary, \"cat\", wordBoundary);\n// \"the cat sat\" -> true\n// \"concatenate\" -> false\n```\n### Case-insensitive\n\n`hello`, in any case.\n\n```\n// Emacs: (let ((case-fold-search t))\n//   (string-match-p (rx bos \"hello\" eos) \"HeLLo\"))\nconst regex = /^hello$/i;\nconst dsl = RX.flags(\"i\", start, \"hello\", end);\n// \"HeLLo\" -> true\n// \"help\" -> false\n```\n## Full source\n\nIt's a single file with no dependencies. Copy it into your project and start deleting the forms you don't need, or adding the ones you miss.\n\nYou can check the same code, plus all the examples from this post (and a few more), in this gist. If you'd rather not set anything up, paste it into the TypeScript Playground, hit \"Run\", and check the \"Logs\" tab.\n\n```\n/* =========================================================\n * CORE\n * ========================================================= */\n// How a node behaves when combined with others:\n//   atom -> single unit, a quantifier can be glued right after it\n//   seq  -> safe to concatenate, needs (?:) to be quantified\n//   alt  -> has a top-level `|`, needs (?:) almost everywhere\ntype Kind = \"atom\" | \"seq\" | \"alt\";\ninterface RxNode {\n  readonly src: string;\n  readonly kind: Kind;\n  // char sets only: the text that goes inside [ ], so sets can merge\n  readonly set?: string;\n  readonly neg?: boolean;\n}\ntype Item = string | RxNode;\nconst esc = (s: string) => s.replace(/[.*+?^${}()|[\\]\\\\]/g, \"\\\\$&\");\nconst escSet = (s: string) => s.replace(/[\\]\\\\^-]/g, \"\\\\$&\");\n// rx: (literal EXPR) — a string computed at runtime, matched as-is\nconst literal = (s: string): RxNode => ({\n  src: esc(s),\n  kind: s.length === 1 ? \"atom\" : \"seq\",\n});\nconst toNode = (x: Item): RxNode => (typeof x === \"string\" ? literal(x) : x);\n// wrap a node so a quantifier applies to all of it\nconst quantifiable = (n: RxNode) =>\n  n.kind === \"atom\" ? n.src : `(?:${n.src})`;\n// rx: (regexp EXPR) — escape hatch, trust the regexp as-is\nconst raw = (re: string | RegExp): RxNode => ({\n  src: typeof re === \"string\" ? re : re.source,\n  kind: \"alt\",\n});\n/* =========================================================\n * COMPOSITION\n * ========================================================= */\nconst seq = (...xs: Item[]): RxNode => {\n  const nodes = xs.map(toNode).filter((n) => n.src !== \"\");\n  if (nodes.length === 0) return { src: \"\", kind: \"seq\" };\n  if (nodes.length === 1) return nodes[0];\n  let src = \"\";\n  for (const n of nodes) {\n\tconst part = n.kind === \"alt\" ? `(?:${n.src})` : n.src;\n\t// `\\1` followed by a literal `0` would read as `\\10`\n\tif (/\\\\\\d+$/.test(src) && /^\\d/.test(part)) src += \"(?:)\";\n\tsrc += part;\n  }\n  return { src, kind: \"seq\" };\n};\nconst unmatchable: RxNode = { src: \"(?!)\", kind: \"atom\" };\n// Like rx: when every branch is a plain string, try the longest first,\n// so or(\"in\", \"int\") matches \"int\" instead of stopping at \"in\".\nconst or = (...xs: Item[]): RxNode => {\n  if (xs.length === 0) return unmatchable;\n  if (xs.length === 1) return toNode(xs[0]);\n  const branches = xs.every((x) => typeof x === \"string\")\n\t? [...(xs as string[])].sort((a, b) => b.length - a.length)\n\t: xs;\n  return { src: branches.map((x) => toNode(x).src).join(\"|\"), kind: \"alt\" };\n};\n/* =========================================================\n * CHARACTER SETS\n * ========================================================= */\nconst set = (body: string, neg = false): RxNode => ({\n  src: neg ? `[^${body}]` : `[${body}]`,\n  kind: \"atom\",\n  set: body,\n  neg,\n});\n// class escapes are sets too, so they can go inside anyOf(...)\nconst classEscape = (e: string): RxNode => ({\n  src: e,\n  kind: \"atom\",\n  set: e,\n  neg: false,\n});\n// Same reading as rx: inside a string, \"a-z\" is a range, while a `-`\n// at the start or end is just a dash (\"+-\" is plus or minus).\nconst intervals = (s: string): string => {\n  let body = \"\";\n  let i = 0;\n  while (i < s.length) {\n\tif (i < s.length - 2 && s[i + 1] === \"-\") {\n\t  body += `${escSet(s[i])}-${escSet(s[i + 2])}`;\n\t  i += 3;\n\t} else {\n\t  body += escSet(s[i]);\n\t  i += 1;\n\t}\n  }\n  return body;\n};\n// rx: (any \"a-z\" \"_\" digit) — also known as `in` and `char`\nconst anyOf = (...xs: Item[]): RxNode => {\n  const body = xs\n\t.map((x) => {\n\t  if (typeof x === \"string\") return intervals(x);\n\t  if (x.set === undefined || x.neg)\n\t\tthrow new Error(`anyOf: not a positive char set: ${x.src}`);\n\t  return x.set;\n\t})\n\t.join(\"\");\n  return set(body);\n};\n// rx: (not charset) — not(digit) -> \\D, not(anyOf(\",;\")) -> [^,;]\nconst not = (x: Item): RxNode => {\n  const n = typeof x === \"string\" ? anyOf(x) : x;\n  if (n.set === undefined) throw new Error(`not: not a char set: ${n.src}`);\n  if (n.neg) return set(n.set);\n  if (/^\\\\[dswDSW]$/.test(n.src)) {\n\tconst c = n.src[1];\n\tconst flipped = c === c.toLowerCase() ? c.toUpperCase() : c.toLowerCase();\n\treturn classEscape(`\\\\${flipped}`);\n  }\n  return set(n.set, true);\n};\n// rx: (not-char \"a-z\" ...) — shorthand for (not (any ...))\nconst notChar = (...xs: Item[]) => not(anyOf(...xs));\n// rx char classes, `[[:name:]]` in Emacs\nconst digit = classEscape(\"\\\\d\");\nconst space = classEscape(\"\\\\s\");\nconst wordChar = classEscape(\"\\\\w\");\nconst notWordChar = not(wordChar);\nconst lower = anyOf(\"a-z\");\nconst upper = anyOf(\"A-Z\");\nconst alpha = anyOf(lower, upper);\nconst alnum = anyOf(alpha, \"0-9\");\nconst hexDigit = anyOf(\"0-9a-fA-F\");\nconst blank = set(\" \\\\t\");\nconst control = set(\"\\\\x00-\\\\x1f\\\\x7f\");\nconst punct = anyOf(\"!-/:-@[-`{-~\");\nconst graphic = anyOf(\"!-~\");\nconst printing = anyOf(\" -~\");\nconst ascii = set(\"\\\\x00-\\\\x7f\");\nconst nonascii = set(\"\\\\u0080-\\\\uffff\");\n// rx: `nonl` is any char but newline; `anything` really is anything\nconst notNewline: RxNode = { src: \".\", kind: \"atom\" };\nconst anything = set(\"\\\\s\\\\S\");\n/* =========================================================\n * ANCHORS (zero-width)\n * ========================================================= */\nconst zeroWidth = (src: string): RxNode => ({ src, kind: \"seq\" });\n// rx: bos / eos\nconst start = zeroWidth(\"^\");\nconst end = zeroWidth(\"$\");\n// rx: bol / eol — same symbols, only per line with the \"m\" flag\nconst lineStart = start;\nconst lineEnd = end;\nconst wordBoundary = zeroWidth(\"\\\\b\");\nconst notWordBoundary = zeroWidth(\"\\\\B\");\n// rx: bow / eow — JS has no \\< \\>, so a boundary plus a lookaround\nconst wordStart = zeroWidth(\"\\\\b(?=\\\\w)\");\nconst wordEnd = zeroWidth(\"\\\\b(?<=\\\\w)\");\n/* =========================================================\n * GROUPS & BACKREFERENCES\n * ========================================================= */\nconst group = (...xs: Item[]): RxNode => ({\n  src: `(${seq(...xs).src})`,\n  kind: \"atom\",\n});\n// rx has (group-n N ...); JS can't pick group numbers, but it can name them\nconst named = (name: string, ...xs: Item[]): RxNode => ({\n  src: `(?<${name}>${seq(...xs).src})`,\n  kind: \"atom\",\n});\nconst backref = (ref: number | string): RxNode => ({\n  src: typeof ref === \"number\" ? `\\\\${ref}` : `\\\\k<${ref}>`,\n  kind: \"atom\",\n});\n/* =========================================================\n * QUANTIFIERS\n * ========================================================= */\nconst quantifier =\n  (suffix: string) =>\n  (...xs: Item[]): RxNode => ({\n\tsrc: quantifiable(seq(...xs)) + suffix,\n\tkind: \"seq\",\n  });\n// greedy — rx: * + ?\nconst zeroOrMore = quantifier(\"*\");\nconst oneOrMore = quantifier(\"+\");\nconst optional = quantifier(\"?\");\n// lazy — rx: *? +? ??\nconst zeroOrMoreLazy = quantifier(\"*?\");\nconst oneOrMoreLazy = quantifier(\"+?\");\nconst optionalLazy = quantifier(\"??\");\n// rx: (= n ...) (>= n ...) (** n m ...)\nconst repeat = (n: number, ...xs: Item[]) => quantifier(`{${n}}`)(...xs);\nconst atLeast = (n: number, ...xs: Item[]) => quantifier(`{${n},}`)(...xs);\nconst between = (n: number, m: number, ...xs: Item[]) =>\n  quantifier(`{${n},${m}}`)(...xs);\n/* =========================================================\n * ENTRY POINTS\n * ========================================================= */\n// rx(...)  -> the regexp source string (like Emacs, rx returns a string)\n// RX(...)  -> a ready-to-use RegExp, no more `new RegExp(seq(...))`\nconst rx = (...xs: Item[]): string => seq(...xs).src;\nfunction RX(...xs: Item[]): RegExp {\n  return new RegExp(rx(...xs));\n}\n// Emacs uses `case-fold-search` for this; JS puts it on the regexp\nRX.flags = (flags: string, ...xs: Item[]): RegExp =>\n  new RegExp(rx(...xs), flags);\n```\n## Wrapping up\n\nNone of this is new. On the Emacs side, as I said before, `rx` has\nshipped for decades, and the Elisp version is more complete than mine.\n\nThe idea of describing patterns with a small DSL instead of raw syntax isn't new either. Plenty of people have tried it, each in their own way. One project I like a lot in this space is Zod, which I wrote about in my Zod quick tutorial. It's not a regexp builder: you compose small schema pieces, and Zod gives you back a parser and a TypeScript type from the same construction. It follows the same spirit, though: build big things out of small named pieces you can read.\n\nIf you write Elisp and have never tried `rx`, open `*scratch*`, type\n`(rx (+ digit))`, and `C-x C-e` it. If you write JavaScript or\nTypeScript, the file above is yours. And if you port it to another\nlanguage, send me a link.","body_html":"<h1 id=\"readable-regular-expressions-for-javascript-typescript-inspired-\">Readable Regular Expressions for JavaScript/TypeScript, Inspired by Emacs&#39; rx</h1>\n<h2 id=\"intro\">Intro</h2>\n<p>Quick, what does this match?</p>\n<pre><code>/^(0|[1-9]\\d*)\\.(0|[1-9]\\d*)\\.(0|[1-9]\\d*)(?:-((?:0|[1-9]\\d*|\\d*[a-zA-Z-][0-9a-zA-Z-]*)(?:\\.(?:0|[1-9]\\d*|\\d*[a-zA-Z-][0-9a-zA-Z-]*))*))?(?:\\+([0-9a-zA-Z-]+(?:\\.[0-9a-zA-Z-]+)*))?$/;\n// Take\n// your\n// time...\n//\n// ...still decoding?\n//\n// OK, keep reading :)</code></pre>\n<p>That&#39;s the official regexp from semver.org. It validates version numbers like:</p>\n<pre><code>// matches\n&quot;1.2.3&quot;\n&quot;0.10.0&quot;\n&quot;2.0.0-rc.1&quot;\n&quot;1.0.0-alpha.1+build.5&quot;\n&quot;1.0.0+20260930&quot;\n// doesn&#39;t match\n&quot;01.2.3&quot;   // leading zero\n&quot;1.2&quot;      // missing patch\n&quot;v1.2.3&quot;   // no &quot;v&quot; prefix allowed\n&quot;1.0.0-01&quot; // numeric pre-release with a leading zero\n&quot;1.2.3-&quot;   // empty pre-release</code></pre>\n<p>Don&#39;t get me wrong, I love regexps, but in practice you probably spend a bunch of time writing one, testing it against some cases, and moving on, proud of your achievement!</p>\n<p>Some time passes and lucky future you (or unlucky someone else) has to change it. Dramatic pause here.</p>\n<p>I bet you&#39;ve been there. Now your options are probably: decode it again from the start, rewrite the whole thing, or, in the age of AI, ask (and hopefully not blindly accept) an LLM for a new recipe.</p>\n<p>Emacs has had a nice answer for more readable regexps for a long time:\nthe <code>rx</code> macro. I started using it all the time in Emacs Lisp, as\nreviewers always suggested it to me. Later, I started missing this DSL\nin JavaScript and TypeScript, so I wrote a small version of it for my\nprojects.</p>\n<p>So, what about reading that SemVer regexp like <code>semver</code> in the code\nbelow?</p>\n<pre><code>const num = or(&quot;0&quot;, seq(anyOf(&quot;1-9&quot;), zeroOrMore(digit)));\nconst idChar = anyOf(alnum, &quot;-&quot;);\nconst preId = or(num, seq(zeroOrMore(digit), anyOf(alpha, &quot;-&quot;), zeroOrMore(idChar)));\nconst dotted = (x: Item) =&gt; seq(x, zeroOrMore(&quot;.&quot;, x));\nconst semver = RX(\n  start,\n  named(&quot;major&quot;, num), &quot;.&quot;,\n  named(&quot;minor&quot;, num), &quot;.&quot;,\n  named(&quot;patch&quot;, num),\n  optional(&quot;-&quot;, named(&quot;pre&quot;, dotted(preId))),\n  optional(&quot;+&quot;, named(&quot;build&quot;, dotted(oneOrMore(idChar)))),\n  end,\n);</code></pre>\n<p>The same strings match, and you get named groups as a bonus. By the end of this post you&#39;ll know every piece of it.</p>\n<p><strong>TL;DR:</strong> jump straight to the cheat sheet,\nthe side-by-side examples, the full\nsource, or grab the\ngist\nto sneak a peek at the result.</p>\n<p><strong>NOTE:</strong> the <code>RX</code> here has nothing to do with\nRxJS, which is an amazing library for reactive\nprogramming with observables.</p>\n<h2 id=\"a-taste-of-rx-in-emacs-lisp\">A taste of rx in Emacs Lisp</h2>\n<p>With <code>rx</code> you describe a regexp as a tree of named forms, and Emacs\nturns it into the regexp string for you:</p>\n<pre><code>(rx bos (+ digit) eos)\n;; =&gt; &quot;\\\\`[[:digit:]]+\\\\&#39;&quot;\n(rx bol &quot;colo&quot; (? &quot;u&quot;) &quot;r&quot; eol)\n;; =&gt; &quot;^colou?r$&quot;\n(rx bos &quot;(&quot; (= 3 digit) &quot;)&quot; space (= 3 digit) &quot;-&quot; (= 4 digit) eos)\n;; =&gt; &quot;\\\\`([[:digit:]]\\\\{3\\\\})[[:space:]][[:digit:]]\\\\{3\\\\}-[[:digit:]]\\\\{4\\\\}\\\\&#39;&quot;</code></pre>\n<p>A few things to notice:</p>\n<ol><li><p><strong>Strings are literals.</strong><code>&quot;(&quot;</code> means a parenthesis. You don&#39;t</p><p>need to escape anything by hand.</p></li><li><p><strong>Sequence is implicit.</strong> Every form takes a list of things and</p><p>matches them one after the other. You don&#39;t need to wrap them in a<code>seq</code> , even though<code>seq</code> exists.</p></li><li><strong>Groups appear only when needed.</strong><code>(+ digit)</code> becomes<code>[[:digit:]]+</code> , not<code>\\(?:[[:digit:]]\\)+</code> .</li></ol>\n<p>The proposed JavaScript/TypeScript version in this post reads like this:</p>\n<pre><code>const phone = RX(\n  start, &quot;(&quot;, repeat(3, digit), &quot;)&quot;, space,\n  repeat(3, digit), &quot;-&quot;, repeat(4, digit), end,\n);\n// =&gt; /^\\(\\d{3}\\)\\s\\d{3}-\\d{4}$/</code></pre>\n<h2 id=\"under-the-hood\">Under the hood</h2>\n<p>If you want strings to be literals, you can&#39;t represent a regexp piece\nas a plain <code>string</code>, otherwise you can&#39;t tell <code>&quot;(&quot;</code> (a literal\nparenthesis) apart from <code>&quot;(?:...)&quot;</code> (a group you built). So every\npiece is a small object:</p>\n<pre><code>type Kind = &quot;atom&quot; | &quot;seq&quot; | &quot;alt&quot;;\ninterface RxNode {\n  readonly src: string;\n  readonly kind: Kind;\n  readonly set?: string; // char sets only, see below\n  readonly neg?: boolean;\n}\ntype Item = string | RxNode;</code></pre>\n<p><code>src</code> is the regexp text. <code>kind</code> records how that text behaves when\nyou glue it to other things:</p>\n<ul><li><p><code>atom</code> : a single unit, like<code>a</code> ,<code>\\d</code> ,<code>[a-z]</code> or<code>(...)</code> . You can</p><p>put a quantifier right after it.</p></li><li><p><code>seq</code> : safe to concatenate, but a quantifier needs<code>(?:...)</code> around</p><p>it.<code>abc</code> is a<code>seq</code> , and so is<code>a+</code> , since<code>a+?</code> would silently\nturn into a lazy quantifier.</p></li><li><p><code>alt</code> : has a<code>|</code> at the top level, so it needs<code>(?:...)</code> almost</p><p>everywhere.</p></li></ul>\n<p>Plain strings go through <code>literal</code>, which escapes them:</p>\n<pre><code>const esc = (s: string) =&gt; s.replace(/[.*+?^${}()|[\\]\\\\]/g, &quot;\\\\$&amp;&quot;);\nconst literal = (s: string): RxNode =&gt; ({\n  src: esc(s),\n  kind: s.length === 1 ? &quot;atom&quot; : &quot;seq&quot;,\n});\nconst toNode = (x: Item): RxNode =&gt; (typeof x === &quot;string&quot; ? literal(x) : x);</code></pre>\n<p>With that in place, <code>seq</code> joins nodes and only brackets alternations:</p>\n<pre><code>const seq = (...xs: Item[]): RxNode =&gt; {\n  const nodes = xs.map(toNode).filter((n) =&gt; n.src !== &quot;&quot;);\n  if (nodes.length === 0) return { src: &quot;&quot;, kind: &quot;seq&quot; };\n  if (nodes.length === 1) return nodes[0];\n  let src = &quot;&quot;;\n  for (const n of nodes) {\n    const part = n.kind === &quot;alt&quot; ? `(?:${n.src})` : n.src;\n    // `\\1` followed by a literal `0` would read as `\\10`\n    if (/\\\\\\d+$/.test(src) &amp;&amp; /^\\d/.test(part)) src += &quot;(?:)&quot;;\n    src += part;\n  }\n  return { src, kind: &quot;seq&quot; };\n};</code></pre>\n<p>(That backreference check is one of those bugs you only find by\nwriting tests, or when it happens to you in prod. <code>backref(1)</code>\nfollowed by the literal <code>&quot;0&quot;</code> gives you backreference number ten.)</p>\n<p>Every quantifier is a <code>seq</code> of its arguments plus a suffix, bracketed\nonly when the body isn&#39;t an atom:</p>\n<pre><code>const quantifiable = (n: RxNode) =&gt;\n  n.kind === &quot;atom&quot; ? n.src : `(?:${n.src})`;\nconst quantifier =\n  (suffix: string) =&gt;\n  (...xs: Item[]): RxNode =&gt; ({\n    src: quantifiable(seq(...xs)) + suffix,\n    kind: &quot;seq&quot;,\n  });\nconst zeroOrMore = quantifier(&quot;*&quot;);\nconst oneOrMore = quantifier(&quot;+&quot;);\nconst optional = quantifier(&quot;?&quot;);</code></pre>\n<p>Because each quantifier calls <code>seq</code> on its arguments, you get the\nimplicit sequence for free: <code>optional(&quot;-&quot;, group(x))</code> becomes\n<code>(?:-(x))?</code>.</p>\n<p>And finally, the two entry points. As in Emacs, <code>rx</code> returns a\nstring. <code>RX</code> returns a <code>RegExp</code> you can use right away:</p>\n<pre><code>const rx = (...xs: Item[]): string =&gt; seq(...xs).src;\nfunction RX(...xs: Item[]): RegExp {\n  return new RegExp(rx(...xs));\n}\nRX.flags = (flags: string, ...xs: Item[]): RegExp =&gt;\n  new RegExp(rx(...xs), flags);</code></pre>\n<p><code>RX.flags</code> exists because Emacs controls case folding through the\n<code>case-fold-search</code> variable, and JavaScript puts it on the regexp\nitself.</p>\n<p>That&#39;s the whole engine! Now, let&#39;s build our vocabulary.</p>\n<h2 id=\"character-sets\">Character sets</h2>\n<p>In Emacs you write <code>(any &quot;a-z&quot; &quot;_&quot;)</code>. Inside those strings, <code>a-z</code> is a\nrange, and a <code>-</code> at either end is a plain dash. I kept the same rule:</p>\n<pre><code>const hexDigit = anyOf(&quot;0-9a-fA-F&quot;);\nRX(start, &quot;#&quot;, repeat(6, hexDigit), end);\n// =&gt; /^#[0-9a-fA-F]{6}$/\nRX(start, optional(anyOf(&quot;+-&quot;)), oneOrMore(digit), end);\n// =&gt; /^[+\\-]?\\d+$/</code></pre>\n<p>The dash comes out escaped because sets can merge. If you combine\n<code>anyOf(&quot;+-&quot;)</code> with <code>anyOf(&quot;0-9&quot;)</code>, an unescaped <code>-</code> would end up in\nthe middle and create a range from <code>+</code> to <code>0</code>. Escaping it costs one\nbackslash.</p>\n<p>And merging is the reason why <code>RxNode</code> has a <code>set</code> field. It is there\nto hold the text that goes between <code>[</code> and <code>]</code>, so <code>anyOf</code> can take\nother sets as arguments:</p>\n<pre><code>const lower = anyOf(&quot;a-z&quot;);\nconst upper = anyOf(&quot;A-Z&quot;);\nconst alpha = anyOf(lower, upper);   // [a-zA-Z]\nconst alnum = anyOf(alpha, &quot;0-9&quot;);   // [a-zA-Z0-9]</code></pre>\n<p><code>not</code> negates a set, and it knows the shorthand classes:</p>\n<pre><code>not(digit);                // \\D\nnot(anyOf(space, &quot;@&quot;));    // [^\\s@]\nnotChar(&quot;,&quot;);              // [^,]   (rx&#39;s not-char)</code></pre>\n<p>The simple email check, which most of us have written as\n<code>/^[^\\s@]+@[^\\s@]+\\.[^\\s@]+$/</code> at some point, becomes:</p>\n<pre><code>const part = oneOrMore(not(anyOf(space, &quot;@&quot;)));\nRX(start, part, &quot;@&quot;, part, &quot;.&quot;, part, end);\n// =&gt; /^[^\\s@]+@[^\\s@]+\\.[^\\s@]+$/</code></pre>\n<p>The rest of the Emacs character classes are there too: <code>digit</code>,\n<code>hexDigit</code>, <code>space</code>, <code>blank</code>, <code>wordChar</code>, <code>notWordChar</code>, <code>alpha</code>,\n<code>alnum</code>, <code>lower</code>, <code>upper</code>, <code>punct</code>, <code>control</code>, <code>graphic</code>, <code>printing</code>,\n<code>ascii</code> and <code>nonascii</code>. One difference: in Emacs they understand\nUnicode, and mine are ASCII only. <code>alpha</code> won&#39;t match <code>é</code>.</p>\n<p>Two more come from rx&#39;s symbol list, and people (me, many times) mix them up:</p>\n<pre><code>const notNewline: RxNode = { src: &quot;.&quot;, kind: &quot;atom&quot; }; // rx: nonl\nconst anything = set(&quot;\\\\s\\\\S&quot;);                        // rx: anything / anychar</code></pre>\n<p>In rx, <code>anything</code> really means anything, newlines included. Here is\nwhere the difference shows up:</p>\n<pre><code>const code = &quot;a = 1; /* first\\n   second */ b = 2;&quot;;\nRX(&quot;/*&quot;, zeroOrMoreLazy(notNewline), &quot;*/&quot;).exec(code);\n// =&gt; null\nRX(&quot;/*&quot;, zeroOrMoreLazy(anything), &quot;*/&quot;).exec(code)?.[0];\n// =&gt; &quot;/* first\\n   second */&quot;</code></pre>\n<h2 id=\"alternatives-and-the-longest-match\">Alternatives, and the longest match</h2>\n<p><code>or</code> works as you&#39;d expect, and gets bracketed when it lands inside a\nsequence:</p>\n<pre><code>RX(start, or(&quot;cat&quot;, &quot;dog&quot;, &quot;bird&quot;), end);\n// =&gt; /^(?:bird|cat|dog)$/</code></pre>\n<p>Did you notice the order changed? I copied that behavior from\nEmacs. When every branch of an <code>or</code> is a plain string, rx hands them\nto <code>regexp-opt</code>, which builds a pattern that prefers the longest\nmatch:</p>\n<pre><code>(rx (or &quot;in&quot; &quot;int&quot; &quot;interface&quot;))\n;; =&gt; &quot;\\\\(?:in\\\\(?:t\\\\(?:erface\\\\)?\\\\)?\\\\)&quot;</code></pre>\n<p>JavaScript alternation takes the first branch that matches, going left to right. So the naive regexp for a list of keywords has a &#39;bug&#39;:</p>\n<pre><code>/in|int|interface/.exec(&quot;interface Foo&quot;)?.[0];\n// =&gt; &quot;in&quot;\nRX(or(&quot;in&quot;, &quot;int&quot;, &quot;interface&quot;)).exec(&quot;interface Foo&quot;)?.[0];\n// =&gt; &quot;interface&quot;</code></pre>\n<p>I don&#39;t build a trie like <code>regexp-opt</code> does. Sorting the strings by\nlength, longest first, is enough to get the same behavior:</p>\n<pre><code>const or = (...xs: Item[]): RxNode =&gt; {\n  if (xs.length === 0) return unmatchable;\n  if (xs.length === 1) return toNode(xs[0]);\n  const branches = xs.every((x) =&gt; typeof x === &quot;string&quot;)\n    ? [...(xs as string[])].sort((a, b) =&gt; b.length - a.length)\n    : xs;\n  return { src: branches.map((x) =&gt; toNode(x).src).join(&quot;|&quot;), kind: &quot;alt&quot; };\n};</code></pre>\n<p>As in Emacs, <code>or()</code> with no branches returns <code>unmatchable</code>, which is\n<code>(?!)</code> here. It&#39;s handy when you build the branch list at runtime and\nit might come out empty.</p>\n<h2 id=\"repetition-greedy-and-lazy\">Repetition, greedy and lazy</h2>\n<p>Emacs has <code>(= n ...)</code>, <code>(&gt;= n ...)</code> and <code>(** n m ...)</code>. Here they are\n<code>repeat</code>, <code>atLeast</code> and <code>between</code>:</p>\n<pre><code>RX(start, between(2, 4, digit), end);  // /^\\d{2,4}$/\nRX(atLeast(3, digit));                 // /\\d{3,}/</code></pre>\n<p>The lazy versions <code>*?</code>, <code>+?</code> and <code>??</code> are <code>zeroOrMoreLazy</code>,\n<code>oneOrMoreLazy</code> and <code>optionalLazy</code>. The classic HTML tag example:</p>\n<pre><code>const html = &quot;&lt;b&gt;bold&lt;/b&gt; and &lt;i&gt;italic&lt;/i&gt;&quot;;\nRX(&quot;&lt;&quot;, oneOrMore(notNewline), &quot;&gt;&quot;).exec(html)?.[0];\n// =&gt; &quot;&lt;b&gt;bold&lt;/b&gt; and &lt;i&gt;italic&lt;/i&gt;&quot;\nRX(&quot;&lt;&quot;, oneOrMoreLazy(notNewline), &quot;&gt;&quot;).exec(html)?.[0];\n// =&gt; &quot;&lt;b&gt;&quot;</code></pre>\n<h2 id=\"groups-and-backreferences\">Groups and backreferences</h2>\n<p><code>group</code> is a capturing group, and <code>backref</code> points back to it:</p>\n<pre><code>RX(start, group(oneOrMore(wordChar)), space, backref(1), end);\n// =&gt; /^(\\w+)\\s\\1$/   matches &quot;hello hello&quot;, not &quot;hello world&quot;</code></pre>\n<p>Emacs also has <code>(group-n N ...)</code> to pick the group number. JavaScript\ncan&#39;t do that, but it has named groups, which serve the same purpose\nand read better:</p>\n<pre><code>const date = RX(\n  named(&quot;y&quot;, repeat(4, digit)), &quot;-&quot;,\n  named(&quot;m&quot;, repeat(2, digit)), &quot;-&quot;,\n  named(&quot;d&quot;, repeat(2, digit)),\n);\ndate.exec(&quot;2026-09-30&quot;)?.groups;\n// =&gt; { y: &#39;2026&#39;, m: &#39;09&#39;, d: &#39;30&#39; }</code></pre>\n<p><code>backref</code> accepts a name as well:</p>\n<pre><code>RX(\n  &quot;&lt;&quot;, named(&quot;tag&quot;, oneOrMore(wordChar)), &quot;&gt;&quot;,\n  zeroOrMoreLazy(notNewline),\n  &quot;&lt;/&quot;, backref(&quot;tag&quot;), &quot;&gt;&quot;,\n);\n// =&gt; /&lt;(?&lt;tag&gt;\\w+)&gt;.*?&lt;\\/\\k&lt;tag&gt;&gt;/</code></pre>\n<h2 id=\"anchors\">Anchors</h2>\n<p><code>rx</code> distinguishes the start of the string (<code>bos</code>) from the start of a\nline (<code>bol</code>). In JavaScript both are <code>^</code>, and the <code>m</code> flag decides\nwhich one you get. I kept both names so the intent shows in the code:</p>\n<pre><code>const text = &quot;TODO: write post\\nDONE: fix rx\\nTODO: publish&quot;;\nconst todo = RX.flags(\n  &quot;gm&quot;,\n  lineStart, &quot;TODO: &quot;, named(&quot;task&quot;, oneOrMore(notNewline)), lineEnd,\n);\n[...text.matchAll(todo)].map((m) =&gt; m.groups?.task);\n// =&gt; [ &#39;write post&#39;, &#39;publish&#39; ]</code></pre>\n<p>Why not add <code>start</code> and <code>end</code> automatically? Because you only want\nthem when validating a whole string. When searching inside a text, as\nin <code>split</code>, <code>replace</code> or <code>matchAll</code>, a hidden <code>^</code> and <code>$</code> would break\neverything. Emacs agrees: <code>bos</code> and <code>eos</code> are explicit in <code>rx</code> too.</p>\n<p><code>wordBoundary</code> and <code>notWordBoundary</code> map straight to <code>\\b</code> and <code>\\B</code>.\nEmacs also has <code>bow</code> and <code>eow</code> (<code>\\&lt;</code> and <code>\\&gt;</code>), start and end of a\nword. JavaScript lacks those, so I combined <code>\\b</code> with a lookaround:</p>\n<pre><code>const wordStart = zeroWidth(&quot;\\\\b(?=\\\\w)&quot;);\nconst wordEnd = zeroWidth(&quot;\\\\b(?&lt;=\\\\w)&quot;);</code></pre>\n<h2 id=\"literal-and-raw\">Literal and raw</h2>\n<p>Plain strings are already literals, but <code>rx</code> has an explicit <code>literal</code>\nform for strings computed at runtime, and I kept it. It documents that\nthe value came from somewhere else:</p>\n<pre><code>const userInput = &quot;1+1=2? (maybe)&quot;;\nnew RegExp(userInput).test(userInput);  // false, oops\nRX(literal(userInput)).test(userInput); // true</code></pre>\n<p>The opposite direction is <code>rx</code>&#39;s <code>(regexp ...)</code> form, the escape hatch.\nHere it&#39;s <code>raw</code>, and it receives either a string or an existing\n<code>RegExp</code>. It lets you adopt the DSL in a codebase full of old regexps\nwithout rewriting all of them, like:</p>\n<pre><code>const legacyZip = /\\d{5}(?:-\\d{4})?/;\nRX(start, repeat(2, upper), &quot; &quot;, raw(legacyZip), end);\n// =&gt; /^[A-Z]{2} (?:\\d{5}(?:-\\d{4})?)$/</code></pre>\n<p><code>raw</code> can&#39;t see inside the text it gets, so it adds brackets whenever\nit&#39;s combined with something else. It&#39;s an extra <code>(?:)</code>, and the\nregexp still works.</p>\n<h2 id=\"shall-we-try-semver-again\">Shall we try SemVer again?</h2>\n<p>Back to the regexp from the intro. In Emacs you would give names to\nthe pieces with <code>rx-define</code> or <code>rx-let</code>. In TypeScript those are just\n<code>const</code>s:</p>\n<pre><code>const num = or(&quot;0&quot;, seq(anyOf(&quot;1-9&quot;), zeroOrMore(digit)));\nconst idChar = anyOf(alnum, &quot;-&quot;);\nconst preId = or(num, seq(zeroOrMore(digit), anyOf(alpha, &quot;-&quot;), zeroOrMore(idChar)));\nconst dotted = (x: Item) =&gt; seq(x, zeroOrMore(&quot;.&quot;, x));\nconst semver = RX(\n  start,\n  named(&quot;major&quot;, num), &quot;.&quot;,\n  named(&quot;minor&quot;, num), &quot;.&quot;,\n  named(&quot;patch&quot;, num),\n  optional(&quot;-&quot;, named(&quot;pre&quot;, dotted(preId))),\n  optional(&quot;+&quot;, named(&quot;build&quot;, dotted(oneOrMore(idChar)))),\n  end,\n);</code></pre>\n<p>Now you can read the spec in the code. A numeric identifier is <code>0</code>, or\na non-zero digit followed by any number of digits. A pre-release is a\ndotted list of identifiers, and so is build metadata. <code>dotted</code> is a\nplain function returning a node, which is as far as abstraction needs\nto go here.</p>\n<p>It matches the same strings as the official regexp, and the named groups give you a result like:</p>\n<pre><code>semver.exec(&quot;1.0.0-alpha.1+build.5&quot;)?.groups;\n// =&gt; {\n//   major: &#39;1&#39;,\n//   minor: &#39;0&#39;,\n//   patch: &#39;0&#39;,\n//   pre: &#39;alpha.1&#39;,\n//   build: &#39;build.5&#39;\n// }</code></pre>\n<p>Next time the spec changes, you can understand what the current regex does at a glance, instead of fighting an army of punctuation.</p>\n<h2 id=\"what-s-missing-from-emacs-rx\">What&#39;s missing from Emacs rx</h2>\n<p>I tried to map every <code>rx</code> form, and a few have no JavaScript equivalent:</p>\n<ul><li><code>point</code> : JavaScript regexps don&#39;t know about a cursor.</li><li><p><code>symbol-start</code> ,<code>symbol-end</code> ,<code>syntax</code> ,<code>category</code> : these depend on</p><p>Emacs syntax tables.</p></li><li><code>intersection</code> : possible with the<code>v</code> flag, but I haven&#39;t needed it.</li><li><p><code>minimal-match</code> /<code>maximal-match</code> : these flip the greediness of</p><p>everything inside them. Doable, but it would need a separate pass,\nand the<code>*Lazy</code> functions cover my use cases.</p></li><li><p><code>eval</code> : TypeScript already evaluates expressions everywhere, so you</p><p>get it for free.</p></li></ul>\n<p>And one addition Emacs doesn&#39;t need: <code>RX.flags</code>.</p>\n<h2 id=\"cheat-sheet\">Cheat sheet</h2>\n<div class=\"table-wrap\"><table><thead><tr><th>Emacs rx</th><th>TypeScript</th><th>JS regexp (roughly)</th></tr></thead><tbody><tr><td><code>seq</code> ,<code>:</code> ,<code>and</code></td><td><code>seq(...)</code> , implicit in every form</td><td><code>ab</code></td></tr><tr><td><code>or</code> ,<code>|</code></td><td><code>or(...)</code></td><td><code>a|b</code></td></tr><tr><td><code>any</code> ,<code>in</code> ,<code>char</code></td><td><code>anyOf(&quot;a-z&quot;, &quot;_&quot;, digit)</code></td><td><code>[a-z_\\d]</code></td></tr><tr><td><code>not-char</code></td><td><code>notChar(...)</code></td><td><code>[^...]</code></td></tr><tr><td><code>not</code></td><td><code>not(charset)</code></td><td><code>\\D</code> ,<code>[^...]</code></td></tr><tr><td><code>*</code> ,<code>+</code> ,<code>?</code></td><td><code>zeroOrMore</code> ,<code>oneOrMore</code> ,<code>optional</code></td><td><code>x*</code> ,<code>x+</code> ,<code>x?</code></td></tr><tr><td><code>*?</code> ,<code>+?</code> ,<code>??</code></td><td><code>zeroOrMoreLazy</code> ,<code>oneOrMoreLazy</code> ,<code>optionalLazy</code></td><td><code>x*?</code> ,<code>x+?</code> ,<code>x??</code></td></tr><tr><td><code>=</code> ,<code>&gt;=</code> ,<code>**</code></td><td><code>repeat</code> ,<code>atLeast</code> ,<code>between</code></td><td><code>x{n}</code> ,<code>x{n,}</code> ,<code>x{n,m}</code></td></tr><tr><td><code>group</code></td><td><code>group(...)</code></td><td><code>(...)</code></td></tr><tr><td><code>group-n</code></td><td><code>named(&quot;name&quot;, ...)</code></td><td><code>(?&lt;name&gt;...)</code></td></tr><tr><td><code>backref</code></td><td><code>backref(1)</code> ,<code>backref(&quot;name&quot;)</code></td><td><code>\\1</code> ,<code>\\k&lt;name&gt;</code></td></tr><tr><td><code>literal</code></td><td><code>literal(s)</code></td><td><code>s</code> , escaped:<code>1\\+1</code></td></tr><tr><td><code>regexp</code> ,<code>regex</code></td><td><code>raw(&quot;...&quot;)</code> ,<code>raw(/.../)</code></td><td><code>(?:...)</code> , as-is</td></tr><tr><td><code>rx-define</code> ,<code>rx-let</code></td><td><code>const</code></td><td>(none)</td></tr><tr><td><code>bos</code> ,<code>eos</code></td><td><code>start</code> ,<code>end</code></td><td><code>^</code> ,<code>$</code></td></tr><tr><td><code>bol</code> ,<code>eol</code></td><td><code>lineStart</code> ,<code>lineEnd</code> (with the<code>m</code> flag)</td><td><code>^</code> ,<code>$</code></td></tr><tr><td><code>bow</code> ,<code>eow</code></td><td><code>wordStart</code> ,<code>wordEnd</code></td><td><code>\\b(?=\\w)</code> ,<code>\\b(?&lt;=\\w)</code></td></tr><tr><td><code>word-boundary</code></td><td><code>wordBoundary</code></td><td><code>\\b</code></td></tr><tr><td><code>not-word-boundary</code></td><td><code>notWordBoundary</code></td><td><code>\\B</code></td></tr><tr><td><code>nonl</code> ,<code>not-newline</code></td><td><code>notNewline</code></td><td><code>.</code></td></tr><tr><td><code>anychar</code> ,<code>anything</code></td><td><code>anything</code></td><td><code>[\\s\\S]</code></td></tr><tr><td><code>unmatchable</code></td><td><code>unmatchable</code></td><td><code>(?!)</code></td></tr><tr><td><code>digit</code></td><td><code>digit</code></td><td><code>\\d</code></td></tr><tr><td><code>hex-digit</code> ,<code>xdigit</code></td><td><code>hexDigit</code></td><td><code>[0-9a-fA-F]</code></td></tr><tr><td><code>space</code> ,<code>whitespace</code></td><td><code>space</code></td><td><code>\\s</code></td></tr><tr><td><code>blank</code></td><td><code>blank</code></td><td><code>[ \\t]</code></td></tr><tr><td><code>word</code> ,<code>wordchar</code></td><td><code>wordChar</code></td><td><code>\\w</code></td></tr><tr><td><code>not-wordchar</code></td><td><code>notWordChar</code></td><td><code>\\W</code></td></tr><tr><td><code>alpha</code> ,<code>letter</code></td><td><code>alpha</code></td><td><code>[a-zA-Z]</code></td></tr><tr><td><code>alnum</code></td><td><code>alnum</code></td><td><code>[a-zA-Z0-9]</code></td></tr><tr><td><code>lower</code> ,<code>upper</code></td><td><code>lower</code> ,<code>upper</code></td><td><code>[a-z]</code> ,<code>[A-Z]</code></td></tr><tr><td><code>punct</code> ,<code>punctuation</code></td><td><code>punct</code></td><td><code>[!-/:-@[-`{-~]</code></td></tr><tr><td><code>cntrl</code> ,<code>control</code></td><td><code>control</code></td><td><code>[\\x00-\\x1f\\x7f]</code></td></tr><tr><td><code>graph</code> ,<code>graphic</code></td><td><code>graphic</code></td><td><code>[!-~]</code></td></tr><tr><td><code>print</code> ,<code>printing</code></td><td><code>printing</code></td><td><code>[ -~]</code></td></tr><tr><td><code>ascii</code> ,<code>nonascii</code></td><td><code>ascii</code> ,<code>nonascii</code></td><td><code>[\\x00-\\x7f]</code> ,<code>[\\u0080-\\uffff]</code></td></tr></tbody></table></div>\n<h2 id=\"examples-js-ts-regex-vs-rx\">Examples: JS/TS regex vs RX</h2>\n<p>Each example below shows the goal, the Emacs <code>rx</code> form in a comment,\nthe regexp you would write by hand, and the <code>RX</code> version. When <code>RX</code>\nproduces a different regexp text, the <code>// =&gt;</code> line shows it. The\nresults at the bottom come from running both against the same strings.</p>\n<h3 id=\"digits-only\">Digits only</h3>\n<p>The whole string is digits.</p>\n<pre><code>// Emacs: (rx bos (+ digit) eos)\nconst regex = /^\\d+$/;\nconst dsl = RX(start, oneOrMore(digit), end);\n// &quot;123&quot; -&gt; true\n// &quot;12a&quot; -&gt; false</code></pre>\n<h3 id=\"letters-only\">Letters only</h3>\n<p>The whole string is ASCII letters.</p>\n<pre><code>// Emacs: (rx bos (+ alpha) eos)\nconst regex = /^[a-zA-Z]+$/;\nconst dsl = RX(start, oneOrMore(alpha), end);\n// &quot;Hello&quot; -&gt; true\n// &quot;He11o&quot; -&gt; false</code></pre>\n<h3 id=\"optional-letter\">Optional letter</h3>\n<p>Both spellings, color and colour.</p>\n<pre><code>// Emacs: (rx bos &quot;colo&quot; (? &quot;u&quot;) &quot;r&quot; eos)\nconst regex = /^colou?r$/;\nconst dsl = RX(start, &quot;colo&quot;, optional(&quot;u&quot;), &quot;r&quot;, end);\n// &quot;color&quot; -&gt; true\n// &quot;colour&quot; -&gt; true\n// &quot;colouur&quot; -&gt; false</code></pre>\n<h3 id=\"two-words\">Two words</h3>\n<p>Two words separated by a space.</p>\n<pre><code>// Emacs: (rx bos (+ wordchar) space (+ wordchar) eos)\nconst regex = /^\\w+\\s\\w+$/;\nconst word = oneOrMore(wordChar);\nconst dsl = RX(start, word, space, word, end);\n// &quot;hello world&quot; -&gt; true\n// &quot;hello&quot; -&gt; false</code></pre>\n<h3 id=\"phone-number\">Phone number</h3>\n<p>(123) 456-7890, parentheses and all.</p>\n<pre><code>// Emacs: (rx bos &quot;(&quot; (= 3 digit) &quot;)&quot; space (= 3 digit) &quot;-&quot; (= 4 digit) eos)\nconst regex = /^\\(\\d{3}\\)\\s\\d{3}-\\d{4}$/;\nconst dsl = RX(\n  start, &quot;(&quot;, repeat(3, digit), &quot;)&quot;, space,\n  repeat(3, digit), &quot;-&quot;, repeat(4, digit), end,\n);\n// &quot;(123) 456-7890&quot; -&gt; true\n// &quot;123-456-7890&quot; -&gt; false</code></pre>\n<h3 id=\"hex-color\">Hex color</h3>\n<p><code>#ff00aa</code>-style colors.</p>\n<pre><code>// Emacs: (rx bos &quot;#&quot; (= 6 hex-digit) eos)\nconst regex = /^#[0-9a-fA-F]{6}$/;\nconst dsl = RX(start, &quot;#&quot;, repeat(6, hexDigit), end);\n// &quot;#ff00aa&quot; -&gt; true\n// &quot;#ff00ag&quot; -&gt; false</code></pre>\n<h3 id=\"signed-integer\">Signed integer</h3>\n<p>An optional sign, then digits.</p>\n<pre><code>// Emacs: (rx bos (? (any &quot;+-&quot;)) (+ digit) eos)\nconst regex = /^[+-]?\\d+$/;\nconst dsl = RX(start, optional(anyOf(&quot;+-&quot;)), oneOrMore(digit), end);\n// =&gt; /^[+\\-]?\\d+$/\n// &quot;-42&quot; -&gt; true\n// &quot;42&quot; -&gt; true\n// &quot;*42&quot; -&gt; false</code></pre>\n<h3 id=\"simple-email\">Simple email</h3>\n<p>Something@something.something, no spaces.</p>\n<pre><code>// Emacs: (rx-let ((part (+ (not (any space &quot;@&quot;)))))\n//   (rx bos part &quot;@&quot; part &quot;.&quot; part eos))\nconst regex = /^[^\\s@]+@[^\\s@]+\\.[^\\s@]+$/;\nconst part = oneOrMore(not(anyOf(space, &quot;@&quot;)));\nconst dsl = RX(start, part, &quot;@&quot;, part, &quot;.&quot;, part, end);\n// &quot;a@b.com&quot; -&gt; true\n// &quot;a b@c.com&quot; -&gt; false\n// &quot;a@b&quot; -&gt; false</code></pre>\n<h3 id=\"no-digits\">No digits</h3>\n<p>A string without any digit.</p>\n<pre><code>// Emacs: (rx bos (+ (not digit)) eos)\nconst regex = /^[^\\d]+$/;\nconst dsl = RX(start, oneOrMore(not(digit)), end);\n// =&gt; /^\\D+$/\n// &quot;abc&quot; -&gt; true\n// &quot;a1b&quot; -&gt; false</code></pre>\n<h3 id=\"csv-line\">CSV line</h3>\n<p>Exactly three comma-separated fields.</p>\n<pre><code>// Emacs: (rx-let ((field (+ (not-char &quot;,&quot;))))\n//   (rx bos field &quot;,&quot; field &quot;,&quot; field eos))\nconst regex = /^[^,]+,[^,]+,[^,]+$/;\nconst field = oneOrMore(notChar(&quot;,&quot;));\nconst dsl = RX(start, field, &quot;,&quot;, field, &quot;,&quot;, field, end);\n// &quot;a,b,c&quot; -&gt; true\n// &quot;a,b,&quot; -&gt; false</code></pre>\n<h3 id=\"one-of-many\">One of many</h3>\n<p>A fixed list of words.</p>\n<pre><code>// Emacs: (rx bos (or &quot;cat&quot; &quot;dog&quot; &quot;bird&quot;) eos)\nconst regex = /^(?:cat|dog|bird)$/;\nconst dsl = RX(start, or(&quot;cat&quot;, &quot;dog&quot;, &quot;bird&quot;), end);\n// =&gt; /^(?:bird|cat|dog)$/\n// &quot;cat&quot; -&gt; true\n// &quot;bird&quot; -&gt; true\n// &quot;cow&quot; -&gt; false</code></pre>\n<h3 id=\"title-and-name\">Title and name</h3>\n<p><code>mr</code> or <code>ms</code>, then a name, keeping the title.</p>\n<pre><code>// Emacs: (rx bos (group (or &quot;mr&quot; &quot;ms&quot;)) space (+ wordchar) eos)\nconst regex = /^(mr|ms)\\s\\w+$/;\nconst dsl = RX(start, group(or(&quot;mr&quot;, &quot;ms&quot;)), space, oneOrMore(wordChar), end);\n// &quot;mr john&quot; -&gt; true\n// &quot;dr john&quot; -&gt; false</code></pre>\n<h3 id=\"between\">Between</h3>\n<p>Two to four digits.</p>\n<pre><code>// Emacs: (rx bos (** 2 4 digit) eos)\nconst regex = /^\\d{2,4}$/;\nconst dsl = RX(start, between(2, 4, digit), end);\n// &quot;12&quot; -&gt; true\n// &quot;12345&quot; -&gt; false</code></pre>\n<h3 id=\"at-least\">At least</h3>\n<p>Three or more digits, anywhere.</p>\n<pre><code>// Emacs: (rx (&gt;= 3 digit))\nconst regex = /\\d{3,}/;\nconst dsl = RX(atLeast(3, digit));\n// &quot;12&quot; -&gt; false\n// &quot;a123&quot; -&gt; true</code></pre>\n<h3 id=\"repeated-word\">Repeated word</h3>\n<p>The same word twice.</p>\n<pre><code>// Emacs: (rx bos (group (+ wordchar)) space (backref 1) eos)\nconst regex = /^(\\w+)\\s\\1$/;\nconst dsl = RX(start, group(oneOrMore(wordChar)), space, backref(1), end);\n// &quot;hello hello&quot; -&gt; true\n// &quot;hello world&quot; -&gt; false</code></pre>\n<h3 id=\"matching-tags\">Matching tags</h3>\n<p>An open tag and its own closing tag.</p>\n<pre><code>// Emacs: (rx &quot;&lt;&quot; (group-n 1 (+ wordchar)) &quot;&gt;&quot; (*? nonl) &quot;&lt;/&quot; (backref 1) &quot;&gt;&quot;)\nconst regex = /&lt;(?&lt;tag&gt;\\w+)&gt;.*?&lt;\\/\\k&lt;tag&gt;&gt;/;\nconst dsl = RX(\n  &quot;&lt;&quot;, named(&quot;tag&quot;, oneOrMore(wordChar)), &quot;&gt;&quot;,\n  zeroOrMoreLazy(notNewline),\n  &quot;&lt;/&quot;, backref(&quot;tag&quot;), &quot;&gt;&quot;,\n);\n// &quot;&lt;b&gt;bold&lt;/b&gt;&quot; -&gt; true\n// &quot;&lt;b&gt;oops&lt;/i&gt;&quot; -&gt; false</code></pre>\n<h3 id=\"whole-word\">Whole word</h3>\n<p><code>cat</code> as a word, not inside another one.</p>\n<pre><code>// Emacs: (rx word-boundary &quot;cat&quot; word-boundary)\nconst regex = /\\bcat\\b/;\nconst dsl = RX(wordBoundary, &quot;cat&quot;, wordBoundary);\n// &quot;the cat sat&quot; -&gt; true\n// &quot;concatenate&quot; -&gt; false</code></pre>\n<h3 id=\"case-insensitive\">Case-insensitive</h3>\n<p><code>hello</code>, in any case.</p>\n<pre><code>// Emacs: (let ((case-fold-search t))\n//   (string-match-p (rx bos &quot;hello&quot; eos) &quot;HeLLo&quot;))\nconst regex = /^hello$/i;\nconst dsl = RX.flags(&quot;i&quot;, start, &quot;hello&quot;, end);\n// &quot;HeLLo&quot; -&gt; true\n// &quot;help&quot; -&gt; false</code></pre>\n<h2 id=\"full-source\">Full source</h2>\n<p>It&#39;s a single file with no dependencies. Copy it into your project and start deleting the forms you don&#39;t need, or adding the ones you miss.</p>\n<p>You can check the same code, plus all the examples from this post (and a few more), in this gist. If you&#39;d rather not set anything up, paste it into the TypeScript Playground, hit &quot;Run&quot;, and check the &quot;Logs&quot; tab.</p>\n<pre><code>/* =========================================================\n * CORE\n * ========================================================= */\n// How a node behaves when combined with others:\n//   atom -&gt; single unit, a quantifier can be glued right after it\n//   seq  -&gt; safe to concatenate, needs (?:) to be quantified\n//   alt  -&gt; has a top-level `|`, needs (?:) almost everywhere\ntype Kind = &quot;atom&quot; | &quot;seq&quot; | &quot;alt&quot;;\ninterface RxNode {\n  readonly src: string;\n  readonly kind: Kind;\n  // char sets only: the text that goes inside [ ], so sets can merge\n  readonly set?: string;\n  readonly neg?: boolean;\n}\ntype Item = string | RxNode;\nconst esc = (s: string) =&gt; s.replace(/[.*+?^${}()|[\\]\\\\]/g, &quot;\\\\$&amp;&quot;);\nconst escSet = (s: string) =&gt; s.replace(/[\\]\\\\^-]/g, &quot;\\\\$&amp;&quot;);\n// rx: (literal EXPR) — a string computed at runtime, matched as-is\nconst literal = (s: string): RxNode =&gt; ({\n  src: esc(s),\n  kind: s.length === 1 ? &quot;atom&quot; : &quot;seq&quot;,\n});\nconst toNode = (x: Item): RxNode =&gt; (typeof x === &quot;string&quot; ? literal(x) : x);\n// wrap a node so a quantifier applies to all of it\nconst quantifiable = (n: RxNode) =&gt;\n  n.kind === &quot;atom&quot; ? n.src : `(?:${n.src})`;\n// rx: (regexp EXPR) — escape hatch, trust the regexp as-is\nconst raw = (re: string | RegExp): RxNode =&gt; ({\n  src: typeof re === &quot;string&quot; ? re : re.source,\n  kind: &quot;alt&quot;,\n});\n/* =========================================================\n * COMPOSITION\n * ========================================================= */\nconst seq = (...xs: Item[]): RxNode =&gt; {\n  const nodes = xs.map(toNode).filter((n) =&gt; n.src !== &quot;&quot;);\n  if (nodes.length === 0) return { src: &quot;&quot;, kind: &quot;seq&quot; };\n  if (nodes.length === 1) return nodes[0];\n  let src = &quot;&quot;;\n  for (const n of nodes) {\n    const part = n.kind === &quot;alt&quot; ? `(?:${n.src})` : n.src;\n    // `\\1` followed by a literal `0` would read as `\\10`\n    if (/\\\\\\d+$/.test(src) &amp;&amp; /^\\d/.test(part)) src += &quot;(?:)&quot;;\n    src += part;\n  }\n  return { src, kind: &quot;seq&quot; };\n};\nconst unmatchable: RxNode = { src: &quot;(?!)&quot;, kind: &quot;atom&quot; };\n// Like rx: when every branch is a plain string, try the longest first,\n// so or(&quot;in&quot;, &quot;int&quot;) matches &quot;int&quot; instead of stopping at &quot;in&quot;.\nconst or = (...xs: Item[]): RxNode =&gt; {\n  if (xs.length === 0) return unmatchable;\n  if (xs.length === 1) return toNode(xs[0]);\n  const branches = xs.every((x) =&gt; typeof x === &quot;string&quot;)\n    ? [...(xs as string[])].sort((a, b) =&gt; b.length - a.length)\n    : xs;\n  return { src: branches.map((x) =&gt; toNode(x).src).join(&quot;|&quot;), kind: &quot;alt&quot; };\n};\n/* =========================================================\n * CHARACTER SETS\n * ========================================================= */\nconst set = (body: string, neg = false): RxNode =&gt; ({\n  src: neg ? `[^${body}]` : `[${body}]`,\n  kind: &quot;atom&quot;,\n  set: body,\n  neg,\n});\n// class escapes are sets too, so they can go inside anyOf(...)\nconst classEscape = (e: string): RxNode =&gt; ({\n  src: e,\n  kind: &quot;atom&quot;,\n  set: e,\n  neg: false,\n});\n// Same reading as rx: inside a string, &quot;a-z&quot; is a range, while a `-`\n// at the start or end is just a dash (&quot;+-&quot; is plus or minus).\nconst intervals = (s: string): string =&gt; {\n  let body = &quot;&quot;;\n  let i = 0;\n  while (i &lt; s.length) {\n    if (i &lt; s.length - 2 &amp;&amp; s[i + 1] === &quot;-&quot;) {\n      body += `${escSet(s[i])}-${escSet(s[i + 2])}`;\n      i += 3;\n    } else {\n      body += escSet(s[i]);\n      i += 1;\n    }\n  }\n  return body;\n};\n// rx: (any &quot;a-z&quot; &quot;_&quot; digit) — also known as `in` and `char`\nconst anyOf = (...xs: Item[]): RxNode =&gt; {\n  const body = xs\n    .map((x) =&gt; {\n      if (typeof x === &quot;string&quot;) return intervals(x);\n      if (x.set === undefined || x.neg)\n        throw new Error(`anyOf: not a positive char set: ${x.src}`);\n      return x.set;\n    })\n    .join(&quot;&quot;);\n  return set(body);\n};\n// rx: (not charset) — not(digit) -&gt; \\D, not(anyOf(&quot;,;&quot;)) -&gt; [^,;]\nconst not = (x: Item): RxNode =&gt; {\n  const n = typeof x === &quot;string&quot; ? anyOf(x) : x;\n  if (n.set === undefined) throw new Error(`not: not a char set: ${n.src}`);\n  if (n.neg) return set(n.set);\n  if (/^\\\\[dswDSW]$/.test(n.src)) {\n    const c = n.src[1];\n    const flipped = c === c.toLowerCase() ? c.toUpperCase() : c.toLowerCase();\n    return classEscape(`\\\\${flipped}`);\n  }\n  return set(n.set, true);\n};\n// rx: (not-char &quot;a-z&quot; ...) — shorthand for (not (any ...))\nconst notChar = (...xs: Item[]) =&gt; not(anyOf(...xs));\n// rx char classes, `[[:name:]]` in Emacs\nconst digit = classEscape(&quot;\\\\d&quot;);\nconst space = classEscape(&quot;\\\\s&quot;);\nconst wordChar = classEscape(&quot;\\\\w&quot;);\nconst notWordChar = not(wordChar);\nconst lower = anyOf(&quot;a-z&quot;);\nconst upper = anyOf(&quot;A-Z&quot;);\nconst alpha = anyOf(lower, upper);\nconst alnum = anyOf(alpha, &quot;0-9&quot;);\nconst hexDigit = anyOf(&quot;0-9a-fA-F&quot;);\nconst blank = set(&quot; \\\\t&quot;);\nconst control = set(&quot;\\\\x00-\\\\x1f\\\\x7f&quot;);\nconst punct = anyOf(&quot;!-/:-@[-`{-~&quot;);\nconst graphic = anyOf(&quot;!-~&quot;);\nconst printing = anyOf(&quot; -~&quot;);\nconst ascii = set(&quot;\\\\x00-\\\\x7f&quot;);\nconst nonascii = set(&quot;\\\\u0080-\\\\uffff&quot;);\n// rx: `nonl` is any char but newline; `anything` really is anything\nconst notNewline: RxNode = { src: &quot;.&quot;, kind: &quot;atom&quot; };\nconst anything = set(&quot;\\\\s\\\\S&quot;);\n/* =========================================================\n * ANCHORS (zero-width)\n * ========================================================= */\nconst zeroWidth = (src: string): RxNode =&gt; ({ src, kind: &quot;seq&quot; });\n// rx: bos / eos\nconst start = zeroWidth(&quot;^&quot;);\nconst end = zeroWidth(&quot;$&quot;);\n// rx: bol / eol — same symbols, only per line with the &quot;m&quot; flag\nconst lineStart = start;\nconst lineEnd = end;\nconst wordBoundary = zeroWidth(&quot;\\\\b&quot;);\nconst notWordBoundary = zeroWidth(&quot;\\\\B&quot;);\n// rx: bow / eow — JS has no \\&lt; \\&gt;, so a boundary plus a lookaround\nconst wordStart = zeroWidth(&quot;\\\\b(?=\\\\w)&quot;);\nconst wordEnd = zeroWidth(&quot;\\\\b(?&lt;=\\\\w)&quot;);\n/* =========================================================\n * GROUPS &amp; BACKREFERENCES\n * ========================================================= */\nconst group = (...xs: Item[]): RxNode =&gt; ({\n  src: `(${seq(...xs).src})`,\n  kind: &quot;atom&quot;,\n});\n// rx has (group-n N ...); JS can&#39;t pick group numbers, but it can name them\nconst named = (name: string, ...xs: Item[]): RxNode =&gt; ({\n  src: `(?&lt;${name}&gt;${seq(...xs).src})`,\n  kind: &quot;atom&quot;,\n});\nconst backref = (ref: number | string): RxNode =&gt; ({\n  src: typeof ref === &quot;number&quot; ? `\\\\${ref}` : `\\\\k&lt;${ref}&gt;`,\n  kind: &quot;atom&quot;,\n});\n/* =========================================================\n * QUANTIFIERS\n * ========================================================= */\nconst quantifier =\n  (suffix: string) =&gt;\n  (...xs: Item[]): RxNode =&gt; ({\n    src: quantifiable(seq(...xs)) + suffix,\n    kind: &quot;seq&quot;,\n  });\n// greedy — rx: * + ?\nconst zeroOrMore = quantifier(&quot;*&quot;);\nconst oneOrMore = quantifier(&quot;+&quot;);\nconst optional = quantifier(&quot;?&quot;);\n// lazy — rx: *? +? ??\nconst zeroOrMoreLazy = quantifier(&quot;*?&quot;);\nconst oneOrMoreLazy = quantifier(&quot;+?&quot;);\nconst optionalLazy = quantifier(&quot;??&quot;);\n// rx: (= n ...) (&gt;= n ...) (** n m ...)\nconst repeat = (n: number, ...xs: Item[]) =&gt; quantifier(`{${n}}`)(...xs);\nconst atLeast = (n: number, ...xs: Item[]) =&gt; quantifier(`{${n},}`)(...xs);\nconst between = (n: number, m: number, ...xs: Item[]) =&gt;\n  quantifier(`{${n},${m}}`)(...xs);\n/* =========================================================\n * ENTRY POINTS\n * ========================================================= */\n// rx(...)  -&gt; the regexp source string (like Emacs, rx returns a string)\n// RX(...)  -&gt; a ready-to-use RegExp, no more `new RegExp(seq(...))`\nconst rx = (...xs: Item[]): string =&gt; seq(...xs).src;\nfunction RX(...xs: Item[]): RegExp {\n  return new RegExp(rx(...xs));\n}\n// Emacs uses `case-fold-search` for this; JS puts it on the regexp\nRX.flags = (flags: string, ...xs: Item[]): RegExp =&gt;\n  new RegExp(rx(...xs), flags);</code></pre>\n<h2 id=\"wrapping-up\">Wrapping up</h2>\n<p>None of this is new. On the Emacs side, as I said before, <code>rx</code> has\nshipped for decades, and the Elisp version is more complete than mine.</p>\n<p>The idea of describing patterns with a small DSL instead of raw syntax isn&#39;t new either. Plenty of people have tried it, each in their own way. One project I like a lot in this space is Zod, which I wrote about in my Zod quick tutorial. It&#39;s not a regexp builder: you compose small schema pieces, and Zod gives you back a parser and a TypeScript type from the same construction. It follows the same spirit, though: build big things out of small named pieces you can read.</p>\n<p>If you write Elisp and have never tried <code>rx</code>, open <code>*scratch*</code>, type\n<code>(rx (+ digit))</code>, and <code>C-x C-e</code> it. If you write JavaScript or\nTypeScript, the file above is yours. And if you port it to another\nlanguage, send me a link.</p>","headings":[{"level":1,"text":"Readable Regular Expressions for JavaScript/TypeScript, Inspired by Emacs' rx","id":"readable-regular-expressions-for-javascript-typescript-inspired-"},{"level":2,"text":"Intro","id":"intro"},{"level":2,"text":"A taste of rx in Emacs Lisp","id":"a-taste-of-rx-in-emacs-lisp"},{"level":2,"text":"Under the hood","id":"under-the-hood"},{"level":2,"text":"Character sets","id":"character-sets"},{"level":2,"text":"Alternatives, and the longest match","id":"alternatives-and-the-longest-match"},{"level":2,"text":"Repetition, greedy and lazy","id":"repetition-greedy-and-lazy"},{"level":2,"text":"Groups and backreferences","id":"groups-and-backreferences"},{"level":2,"text":"Anchors","id":"anchors"},{"level":2,"text":"Literal and raw","id":"literal-and-raw"},{"level":2,"text":"Shall we try SemVer again?","id":"shall-we-try-semver-again"},{"level":2,"text":"What's missing from Emacs rx","id":"what-s-missing-from-emacs-rx"},{"level":2,"text":"Cheat sheet","id":"cheat-sheet"},{"level":2,"text":"Examples: JS/TS regex vs RX","id":"examples-js-ts-regex-vs-rx"},{"level":3,"text":"Digits only","id":"digits-only"},{"level":3,"text":"Letters only","id":"letters-only"},{"level":3,"text":"Optional letter","id":"optional-letter"},{"level":3,"text":"Two words","id":"two-words"},{"level":3,"text":"Phone number","id":"phone-number"},{"level":3,"text":"Hex color","id":"hex-color"},{"level":3,"text":"Signed integer","id":"signed-integer"},{"level":3,"text":"Simple email","id":"simple-email"},{"level":3,"text":"No digits","id":"no-digits"},{"level":3,"text":"CSV line","id":"csv-line"},{"level":3,"text":"One of many","id":"one-of-many"},{"level":3,"text":"Title and name","id":"title-and-name"},{"level":3,"text":"Between","id":"between"},{"level":3,"text":"At least","id":"at-least"},{"level":3,"text":"Repeated word","id":"repeated-word"},{"level":3,"text":"Matching tags","id":"matching-tags"},{"level":3,"text":"Whole word","id":"whole-word"},{"level":3,"text":"Case-insensitive","id":"case-insensitive"},{"level":2,"text":"Full source","id":"full-source"},{"level":2,"text":"Wrapping up","id":"wrapping-up"}]}}