{"article":{"slug":"safere-1-0-released","title":"SafeRE 1.0 released","subtitle":null,"summary":"Eddie Aftandilian ships SafeRE 1.0, a linear-time Java regex library built with agents: differential testing vs the JDK, ReDoS resistance by construction, and performance that now beats JDK and RE2/J on Rebar workloads.","content_type":"blog_post","language":"en","canonical_url":"https://eaftan.github.io/safere-10/","author":{"name":"Eddie Aftandilian","url":"https://eaftan.github.io/","person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"Eddie Aftandilian","url":"https://eaftan.github.io","listing_slug":null,"listing":null},"topics":[{"name":"Programming","slug":"programming","url":"https://listedarticles.com/topics/programming"},{"name":"Java","slug":"java","url":"https://listedarticles.com/topics/java"},{"name":"Security","slug":"security","url":"https://listedarticles.com/topics/security"},{"name":"Open Source","slug":"open-source","url":"https://listedarticles.com/topics/open-source"},{"name":"AI Agents","slug":"ai-agents","url":"https://listedarticles.com/topics/ai-agents"},{"name":"Performance","slug":"performance","url":"https://listedarticles.com/topics/performance"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":1062,"reading_minutes":5,"published_at":"2026-09-28T00:00:00.000Z","added_at":"2026-10-01T06:14:14.662Z","updated_at":"2026-10-01T06:14:14.662Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":false},"profile_url":"https://listedarticles.com/articles/safere-1-0-released","markdown_url":"https://listedarticles.com/articles/safere-1-0-released.md","example":false,"citation":"Eddie Aftandilian, Eddie Aftandilian. \"SafeRE 1.0 released.\" 28 Sept 2026. https://eaftan.github.io/safere-10/ (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://eaftan.github.io/safere-10/"},"body_markdown":"# SafeRE 1.0 released\n\n*Eddie Aftandilian — September 28, 2026*\n\nI've just released SafeRE 1.0.0, a safe, correct, and fast regular expression library for Java. You can grab it from Maven Central.\n\nSafeRE guarantees linear-time matching for a fixed compiled pattern. SafeRE's finite automata prevent catastrophic backtracking by construction, blocking regular expression denial-of-service (ReDoS) attacks that exploit it.\n\n## What does 1.0 mean?\n\nMy initial goal in building SafeRE was to deliver a safe, production-grade regex library using agents to accelerate the process. It was an experiment to see whether I could build a technically complex library faster than previously possible, and a way for me to build experience developing in a serious (i.e., not vibecoded) codebase with agents.\n\nBut what does production-grade actually mean? At the beginning of the project, I thought it meant correctness and \"good enough\" performance, which to me meant rough parity with the JDK. So I spent most of my time on correctness, building a large amount of testing infrastructure — unit and regression tests, differential tests against the JDK, fuzz tests, and exhaustive tests — to make sure SafeRE was as correct as possible. SafeRE's CI runs over 10 million test cases on each commit, and there are 20 billion additional cases available through on-demand exhaustive sweeps. These days, the differences we find through fuzz testing more often turn out to be JDK bugs than SafeRE bugs. That's made me pretty confident that SafeRE is, if anything, more correct than `java.util.regex`.\n\nBut as people started testing SafeRE for use in their projects, it seemed that performance was now the adoption blocker. The people evaluating SafeRE expected better performance than whatever regex library they were using, which was often something faster than the JDK, such as RE2 via JNI/FFM or Joni.\n\nSo over the past 2-3 months, we've focused on performance optimizations. My collaborator Liam Miller-Cushon has been incredibly productive, delivering approximately 80 optimization PRs and sharing real-world workloads we can tune.\n\nSafeRE is now substantially faster than the JDK and RE2/J, and somewhat faster than native RE2, overall on the shared curated workloads in Rebar, a comprehensive regex benchmark suite created by @BurntSushi, the author of ripgrep and Rust's regex library. You can find the detailed results and methodology in the 1.0 benchmark report.\n\nRE2 sets a high bar; it's a highly tuned C++ implementation with decades of production use at Google, and its design was the foundation for SafeRE. I'm proud that SafeRE, written in Java, is now a little faster than RE2 on these benchmarks.\n\nThat's what 1.0 means to me: I'm confident SafeRE is ready for production adoption.\n\n## What differential testing taught me\n\nI built fuzz testing and exhaustive testing tools that compared SafeRE's results to the JDK and failed when their answers differed. I had assumed that the mature, battle-tested regex implementation in the JDK would be correct pretty much all the time.\n\nInitially, this worked well — I found and fixed a lot of SafeRE bugs. But eventually, matching the JDK's behavior started leading us in directions I wasn't comfortable with. For example, `hitEnd()` and `requireEnd()` report whether a match attempt reached the end of the input and whether more input could invalidate a successful match. While trying to reproduce the JDK's exact behavior, my agent started reimplementing parts of its backtracking engine! I concluded that reproducing the JDK's exact behavior likely required backtracking, which would compromise SafeRE's linear-time guarantee. I also didn't want SafeRE's API to depend on details of another engine's implementation, so I removed support for both methods.\n\nIn other cases, I found places where the JDK was returning the wrong answer. Consider this example:\n\n```\nvar matcher = java.util.regex.Pattern.compile(\"(?:(a))*$\").matcher(\"aaab\");\nmatcher.find();\n\nSystem.out.println(matcher.start());  // 4: end of the input\nSystem.out.println(matcher.group());  // \"\" (an empty match)\nSystem.out.println(matcher.group(1)); // \"a\" — should be null\n```\n\nThe pattern means \"zero or more `a`s, followed by the end of the input.\" Because the input ends with `b`, the only possible match is zero `a`s at the very end. The capturing group never participates in that match, so its value should be `null`. Instead, the JDK returns `\"a\"`, left over from an earlier attempt that failed. I reported this as JDK-8381696.\n\nSo now I'm in a situation where I'm finding differences, not bugs, via differential testing, and it requires judgement to determine whether there is a bug in SafeRE or the JDK. Right now, I typically fix reported bugs, even minor ones, but I also run benchmarks if the fix looks like it could affect performance. I won't ship a fix for a rare bug if it causes a performance regression.\n\n## On performance optimization\n\nMany of the recent performance optimizations in SafeRE come down to quickly skipping text that cannot match. For example, consider `[A-Z]+:\\s+[0-9]+`: uppercase letters, a colon, whitespace, and a number. If colons are rare in the input, it can be profitable to first scan for colons using Java's fast `String.indexOf` method, then check whether the surrounding text satisfies the rest of the pattern.\n\nHotSpot can accelerate `String.indexOf` using SIMD instructions. Liam's implementation of Teddy, an algorithm from Hyperscan that searches for multiple strings at once, uses SIMD operations to check many input positions in parallel. On UTF-8 alternation benchmarks with inputs from 1 KB to 100 KB, Liam measured roughly 5-6x speedups over SafeRE before this change.\n\n## Other production considerations\n\nProduction users also need a project they can depend on. Recently I've:\n\n- Increased the bus factor by adding Liam as a collaborator\n- Required code reviews before shipping changes\n- Started providing SNAPSHOT builds published on every push to main\n\n## What I've learned\n\nWhen I started this project, I naively thought I could build a production-grade safe regex library within a month or two with the help of agents. But it turned out to be much harder than that. The hard parts are the same parts that would always have been the hard parts: matching or exceeding battle-tested libraries on performance and correctness, while retaining the desired safety guarantees.\n\nOn the other hand, a project like this would definitely not be doable or maintainable in my spare time without agents. Not only that, but I've enjoyed building with agents and learning how to use them effectively.\n\n## Acknowledgements\n\nThanks to Liam Miller-Cushon, Mateusz \"Serafin\" Gajewski, Kevin Bourrillion, and my family.\n\n*Note: I wrote this post by hand. I used an agent for proofreading and feedback.*","body_html":"<h1 id=\"safere-1-0-released\">SafeRE 1.0 released</h1>\n<p><em>Eddie Aftandilian — September 28, 2026</em></p>\n<p>I&#39;ve just released SafeRE 1.0.0, a safe, correct, and fast regular expression library for Java. You can grab it from Maven Central.</p>\n<p>SafeRE guarantees linear-time matching for a fixed compiled pattern. SafeRE&#39;s finite automata prevent catastrophic backtracking by construction, blocking regular expression denial-of-service (ReDoS) attacks that exploit it.</p>\n<h2 id=\"what-does-1-0-mean\">What does 1.0 mean?</h2>\n<p>My initial goal in building SafeRE was to deliver a safe, production-grade regex library using agents to accelerate the process. It was an experiment to see whether I could build a technically complex library faster than previously possible, and a way for me to build experience developing in a serious (i.e., not vibecoded) codebase with agents.</p>\n<p>But what does production-grade actually mean? At the beginning of the project, I thought it meant correctness and &quot;good enough&quot; performance, which to me meant rough parity with the JDK. So I spent most of my time on correctness, building a large amount of testing infrastructure — unit and regression tests, differential tests against the JDK, fuzz tests, and exhaustive tests — to make sure SafeRE was as correct as possible. SafeRE&#39;s CI runs over 10 million test cases on each commit, and there are 20 billion additional cases available through on-demand exhaustive sweeps. These days, the differences we find through fuzz testing more often turn out to be JDK bugs than SafeRE bugs. That&#39;s made me pretty confident that SafeRE is, if anything, more correct than <code>java.util.regex</code>.</p>\n<p>But as people started testing SafeRE for use in their projects, it seemed that performance was now the adoption blocker. The people evaluating SafeRE expected better performance than whatever regex library they were using, which was often something faster than the JDK, such as RE2 via JNI/FFM or Joni.</p>\n<p>So over the past 2-3 months, we&#39;ve focused on performance optimizations. My collaborator Liam Miller-Cushon has been incredibly productive, delivering approximately 80 optimization PRs and sharing real-world workloads we can tune.</p>\n<p>SafeRE is now substantially faster than the JDK and RE2/J, and somewhat faster than native RE2, overall on the shared curated workloads in Rebar, a comprehensive regex benchmark suite created by @BurntSushi, the author of ripgrep and Rust&#39;s regex library. You can find the detailed results and methodology in the 1.0 benchmark report.</p>\n<p>RE2 sets a high bar; it&#39;s a highly tuned C++ implementation with decades of production use at Google, and its design was the foundation for SafeRE. I&#39;m proud that SafeRE, written in Java, is now a little faster than RE2 on these benchmarks.</p>\n<p>That&#39;s what 1.0 means to me: I&#39;m confident SafeRE is ready for production adoption.</p>\n<h2 id=\"what-differential-testing-taught-me\">What differential testing taught me</h2>\n<p>I built fuzz testing and exhaustive testing tools that compared SafeRE&#39;s results to the JDK and failed when their answers differed. I had assumed that the mature, battle-tested regex implementation in the JDK would be correct pretty much all the time.</p>\n<p>Initially, this worked well — I found and fixed a lot of SafeRE bugs. But eventually, matching the JDK&#39;s behavior started leading us in directions I wasn&#39;t comfortable with. For example, <code>hitEnd()</code> and <code>requireEnd()</code> report whether a match attempt reached the end of the input and whether more input could invalidate a successful match. While trying to reproduce the JDK&#39;s exact behavior, my agent started reimplementing parts of its backtracking engine! I concluded that reproducing the JDK&#39;s exact behavior likely required backtracking, which would compromise SafeRE&#39;s linear-time guarantee. I also didn&#39;t want SafeRE&#39;s API to depend on details of another engine&#39;s implementation, so I removed support for both methods.</p>\n<p>In other cases, I found places where the JDK was returning the wrong answer. Consider this example:</p>\n<pre><code>var matcher = java.util.regex.Pattern.compile(&quot;(?:(a))*$&quot;).matcher(&quot;aaab&quot;);\nmatcher.find();\n\nSystem.out.println(matcher.start());  // 4: end of the input\nSystem.out.println(matcher.group());  // &quot;&quot; (an empty match)\nSystem.out.println(matcher.group(1)); // &quot;a&quot; — should be null</code></pre>\n<p>The pattern means &quot;zero or more <code>a</code>s, followed by the end of the input.&quot; Because the input ends with <code>b</code>, the only possible match is zero <code>a</code>s at the very end. The capturing group never participates in that match, so its value should be <code>null</code>. Instead, the JDK returns <code>&quot;a&quot;</code>, left over from an earlier attempt that failed. I reported this as JDK-8381696.</p>\n<p>So now I&#39;m in a situation where I&#39;m finding differences, not bugs, via differential testing, and it requires judgement to determine whether there is a bug in SafeRE or the JDK. Right now, I typically fix reported bugs, even minor ones, but I also run benchmarks if the fix looks like it could affect performance. I won&#39;t ship a fix for a rare bug if it causes a performance regression.</p>\n<h2 id=\"on-performance-optimization\">On performance optimization</h2>\n<p>Many of the recent performance optimizations in SafeRE come down to quickly skipping text that cannot match. For example, consider <code>[A-Z]+:\\s+[0-9]+</code>: uppercase letters, a colon, whitespace, and a number. If colons are rare in the input, it can be profitable to first scan for colons using Java&#39;s fast <code>String.indexOf</code> method, then check whether the surrounding text satisfies the rest of the pattern.</p>\n<p>HotSpot can accelerate <code>String.indexOf</code> using SIMD instructions. Liam&#39;s implementation of Teddy, an algorithm from Hyperscan that searches for multiple strings at once, uses SIMD operations to check many input positions in parallel. On UTF-8 alternation benchmarks with inputs from 1 KB to 100 KB, Liam measured roughly 5-6x speedups over SafeRE before this change.</p>\n<h2 id=\"other-production-considerations\">Other production considerations</h2>\n<p>Production users also need a project they can depend on. Recently I&#39;ve:</p>\n<ul><li>Increased the bus factor by adding Liam as a collaborator</li><li>Required code reviews before shipping changes</li><li>Started providing SNAPSHOT builds published on every push to main</li></ul>\n<h2 id=\"what-i-ve-learned\">What I&#39;ve learned</h2>\n<p>When I started this project, I naively thought I could build a production-grade safe regex library within a month or two with the help of agents. But it turned out to be much harder than that. The hard parts are the same parts that would always have been the hard parts: matching or exceeding battle-tested libraries on performance and correctness, while retaining the desired safety guarantees.</p>\n<p>On the other hand, a project like this would definitely not be doable or maintainable in my spare time without agents. Not only that, but I&#39;ve enjoyed building with agents and learning how to use them effectively.</p>\n<h2 id=\"acknowledgements\">Acknowledgements</h2>\n<p>Thanks to Liam Miller-Cushon, Mateusz &quot;Serafin&quot; Gajewski, Kevin Bourrillion, and my family.</p>\n<p><em>Note: I wrote this post by hand. I used an agent for proofreading and feedback.</em></p>","headings":[{"level":1,"text":"SafeRE 1.0 released","id":"safere-1-0-released"},{"level":2,"text":"What does 1.0 mean?","id":"what-does-1-0-mean"},{"level":2,"text":"What differential testing taught me","id":"what-differential-testing-taught-me"},{"level":2,"text":"On performance optimization","id":"on-performance-optimization"},{"level":2,"text":"Other production considerations","id":"other-production-considerations"},{"level":2,"text":"What I've learned","id":"what-i-ve-learned"},{"level":2,"text":"Acknowledgements","id":"acknowledgements"}]}}