{"article":{"slug":"im-sorry-but-you-still-have-to-think","title":"I'm sorry, but you still have to think","subtitle":null,"summary":"Piotr Sarnacki examines DHH's AI-agent rewrites of Campfire Once from Rails into Rust, Elixir and Go, showing how unspecified prompts turn design trade-offs like backwards compatibility, latency versus throughput and crash behaviour into coin flips, and why benchmarks and reliability still require human critical thinking.","content_type":"opinion","language":"en","canonical_url":"https://itsallaboutthebit.com/i-am-sorry-but-you-still-have-to-think/","author":{"name":"Piotr Sarnacki","url":null,"person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"It's all about the bit","url":"https://itsallaboutthebit.com/","listing_slug":null,"listing":null},"topics":[{"name":"Software Engineering","slug":"software-engineering","url":"https://listedarticles.com/topics/software-engineering"},{"name":"AI","slug":"ai","url":"https://listedarticles.com/topics/ai"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":1742,"reading_minutes":8,"published_at":"2026-10-11T13:02:00.000Z","added_at":"2026-10-11T17:12:40.990Z","updated_at":"2026-10-11T17:12:40.990Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":true},"profile_url":"https://listedarticles.com/articles/im-sorry-but-you-still-have-to-think","markdown_url":"https://listedarticles.com/articles/im-sorry-but-you-still-have-to-think.md","example":false,"citation":"Piotr Sarnacki, It's all about the bit. \"I'm sorry, but you still have to think.\" 11 Oct 2026. https://itsallaboutthebit.com/i-am-sorry-but-you-still-have-to-think/ (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://itsallaboutthebit.com/i-am-sorry-but-you-still-have-to-think/"},"body_markdown":"Written by [Piotr Sarnacki](https://hachyderm.io/@drogus)  \non October 11, 2026\n\n# I'm sorry, but you still have to think\n\nDHH, the creator of the Ruby on Rails framework,\ndecided to rewrite [Campfire Once](https://github.com/basecamp/once-campfire) from Ruby on Rails to Rust. Or rather\nask a clanker to do it for him, as he can't stand reading or writing the Rust code himself. After that he added\n[rewrites in other languages](https://x.com/dhh/status/2106810173683851564?s=20), like Elixir and Go.\n\nThis is very interesting to me, as it's a very good insight into the results of someone\nusing AI agents without reading the code. In this case, that someone has also a lot of prior programming experience.\nAnd the results are... well, not very encouraging if you are hoping you can stop reading the code, or stop thinking altogether\nanytime soon. There are multiple problems with the rewrites, but let's start with non-functional differences cause\nthey nicely show something that a lot of people seem to not realize: if your prompt is not specific enough,\nmany decisions are a coin flip. This is because a whole lot of questions don't\nhave a single correct answer. Do we want to keep backwards compatibility? Do we care more about latency or throughput?\nHow much memory can the system use under load? Is loosing new content notifications after a crash acceptable?\nYou may not care about them, or at least some of them,\nbut they will be implicitly answered when an LLM implements what you think you want.\n\nIf you look into the rewrites closer, you will quickly see that they treat backwards compatibility and other constraints\ndifferently. The Rust version doesn't maintain 100% backwards compatibility, for example it drops the CSRF token so\nthat caching is easier. It also dropped Redis in favour of in-process queues, for example for notifications.\nElixir rewrite is much closer to the Rails original. This already makes the comparison pretty much useless,\nas these differences are language independent. It's not like a clanker heard \"rewrite in Elixir\" and chose to keep\n100% backwards compatibility because of the language that has been used. But it gets even better!\n\nIf you looked at the code, and yeah, I know, we should not be reading code anymore, you would quickly notice a lot\nof the things are not great. For example, I've seen people complaining that Elixir's version got a single process\nsequentially processing all of the SQL queries, even if these were reads that could be running concurrently. That's\na fair complaint, but DHH [seems to think](https://x.com/dhh/status/2106864061832991134?s=20) it's not the best look for Elixir that the clanker couldn't write\nperformant code. Not sure if I agree when I look at the Rust version. Cause in the Rust version some of the database\noperations are not async, which in some cases may be worse. You see, in Rust, when using an async runtime, the scheduling\nis not preemptive, but rather cooperative. If a task doesn't yield, no other task can run on the same worker thread.\nWhat this means in practice, is that the time spent in the task should be as short as possible. For example, when\nyou run an SQL query that takes 100ms, you don't want all the other tasks to wait as the task issuing the query is\nwaiting anyway. Thus, you should ideally use an async I/O operation that yields to the runtime\nwhile waiting for a response from the database. In the Rust rewrite some queries run on worker threads, but some run as\nblocking operations in async tasks. Same with locks. When using\nan async runtime, the safest bet is to use async locks like `tokio::sync::Mutex`. It's fine to use a non-async\nversion if you're sure that the lock is held for a very short time, but if you hold a sync lock for 10ms,\nall the tasks on the same thread wait for it, while the blocking task itself is not doing any work. Thus, I regret\nto inform you, that no, coding is not likely \"solved\", and you still need to know what you're doing.\n\nGoing further, it turned out that slop benchmarks are also, well, slop, as they only measure throughput, disregarding other properties of the system. For example, [Zach Daniels measured new post notifications\ndelivery rate under load](https://x.com/ZachSDaniel1/status/2106876400196059200), and discovered a 1% successful\ndelivery rate in the Rust version under heavy load. Not a great look. Benchmarks are hard.\nBut even the improved benchmark may not be actually testing what you want to test, depending on the characteristics of your system.\nYou see, when you test a system you may want to emphasize different properties under load.\nBoth DHH's and Zach's benchmarks were closed-loop benchmarks, so they were testing \"how many requests can the system handle in a specific amount of time?\".\nThe test was using N clients, each sending a new message as soon as it gets a response for the previous one.\nIn many cases, though, the increased load on the system\nmay come from a big number of users performing an action at the same time, who will not wait for other users to finish their actions. In this case you would prefer\n[constant arrival rate](https://grafana.com/docs/k6/latest/using-k6/scenarios/executors/constant-arrival-rate/), ie. you want to send requests at a constant\nrate rather than making the rate depend on how fast the system can respond. If you want to check how reliable the events delivery is,\nyou would ideally want to compare the same number of events.\nHere, if you look at the results, it shows that Rust delivered 1% of notifications out of 6-7k, whereas Elixir delivered 100%\nof notifications out of ~1.7k. What it shows is that the Elixir version handles backpressure better, but this only holds as long as the number\nof requests don't overload the system.\n\nBut let's leave benchmarks just for a moment and talk about trade-offs, cause programming is all about trade-offs. Sure, there are situations\nwhere a tool or a solution is clearly better, without any downsides, but it's quite rare, especially when you get to complex systems\nthat have to be reliable. I've seen quite a lot of\npeople from the Elixir community coming to various conclusions after they've seen the 1% delivery rate of the Rust version,\nwithout really trying to understand why it happens. The general consensus? Elixir is just better at concurrency! In Rust the\nscheduler is cooperative, how can you even live like this? I like Elixir, and I've successfully used it in production, but it's not a silver bullet. Yes, Elixir (or other BEAM based languages) are very good at concurrency, and admittedly\nwriting concurrent code in Elixir is, in general, easier than in Rust, but it comes at a cost of not having low level control,\nhigher memory usage, and oftentimes, speed. There is a reason why [rustler](https://github.com/rusterlium/rustler) exists. And proclaiming\nlanguage's superiority based on a single metric without knowing the root cause, may be misleading.\n\nRemember how I mentioned the closed-loop stress test shows just one of the properties of the system as Rust processed ~4 times more requests and thus\nhad to handle more events sent to the client through a WebSocket? I've rerun the test with constant delivery rate using the DHH's Rust version and Zach's Elixir version with various fixes. At 100 POSTs/s\nthe clients received ~14% events in Rust. In Elixir it was ~60%. Still better, right? Not so fast! At this traffic level Rust didn't have any\nHTTP errors. Elixir timed out on ~23% of HTTP POST requests. And what about latency? The worst event delivery latency was close to\n180s. In Rust when the deliveries are lagging, clients get disconnected. On reconnect, a browser client would fetch the latest messages,\nwhich largely invalidates the need for the missed events.\nWhat do you think is better UX: the client silently reconnecting in the background and fetching the new updates, or waiting for an\nupdate about a new message for 3 minutes? Which just shows that, again, a single metric doesn't tell the whole story.\n\nGetting back to trade-offs. Do you know *why* the Rust version drops so many messages under heavy load? It uses `tokio::sync::broadcast` channel\nto broadcast events to connected clients. An event informing about a new chat message may need to be sent to multiple connections,\nso that makes total sense. One of the properties of the broadcast channel is that it's bounded by a set capacity. If a receiver can't\nhandle the messages fast enough, the receiver will get a `RecvError::Lagged` error (refer to the docs to learn more about [lagging](https://docs.rs/tokio/latest/tokio/sync/broadcast/index.html#lagging)). The broadcast capacity in the Rust rewrite was set to 256. In Elixir, the `GenServer` process is handling events\ndelivery and by default `GenServer` mailboxes don't have a limit. A single line change in Rust:\n\n```\n- stream_capacity: 256\n+ stream_capacity: 16384\n```\n\nincreases the delivery rate at 100 req/s from ~14% to ~90%. It's better than Elixir now, isn't it? Not really. I think that\nthis version is actually worse because in this case it's better to fail fast and force the client to reconnect rather than handle things extremely slowly. You know\nwhat else changed after raising the capacity? The pMAX event delivery latency went up from 11s to >130s, similarly to how Elixir behaved, which I think\nis strictly worse than disconnecting the lagging clients. If you ask me, even 11s is too long and if the receiver can't pass the event faster\nthan that, it's better to drop the event. Which shows how important it is to set sensible constraints in the system, and that, in fact,\nElixir isn't just automatically fixing all of the concurrency issues out of the box. Also, if you don't set bounds yourself, you will likely run into\nexternal limits. During the 100 req/s stress test the Elixir version reached 1.8GB memory usage. Another thing to ponder on: is it better to drop some messages or\nget OOM killed? Trade-offs all the way. Elixir is great, but it won't magically solve all of your problems. No matter the language you're using, you have to think\nabout failure modes and trade-offs.\n\nSo what have we learned today? You still have to think critically. It's good to know what you're doing. Don't\nmake hasty assumptions based on a single metric. If you want a reliable system you should know when to fail. Also, benchmarks are hard.\n\n---\n\nIf you like this post please consider following me on [Twitter](http://twitter.com/drogus).","body_html":"<p>Written by <a href=\"https://hachyderm.io/@drogus\" rel=\"nofollow ugc noopener\">Piotr Sarnacki</a><br />\non October 11, 2026</p>\n<h1 id=\"i-m-sorry-but-you-still-have-to-think\">I&#39;m sorry, but you still have to think</h1>\n<p>DHH, the creator of the Ruby on Rails framework,\ndecided to rewrite <a href=\"https://github.com/basecamp/once-campfire\" rel=\"nofollow ugc noopener\">Campfire Once</a> from Ruby on Rails to Rust. Or rather\nask a clanker to do it for him, as he can&#39;t stand reading or writing the Rust code himself. After that he added\n<a href=\"https://x.com/dhh/status/2106810173683851564?s=20\" rel=\"nofollow ugc noopener\">rewrites in other languages</a>, like Elixir and Go.</p>\n<p>This is very interesting to me, as it&#39;s a very good insight into the results of someone\nusing AI agents without reading the code. In this case, that someone has also a lot of prior programming experience.\nAnd the results are... well, not very encouraging if you are hoping you can stop reading the code, or stop thinking altogether\nanytime soon. There are multiple problems with the rewrites, but let&#39;s start with non-functional differences cause\nthey nicely show something that a lot of people seem to not realize: if your prompt is not specific enough,\nmany decisions are a coin flip. This is because a whole lot of questions don&#39;t\nhave a single correct answer. Do we want to keep backwards compatibility? Do we care more about latency or throughput?\nHow much memory can the system use under load? Is loosing new content notifications after a crash acceptable?\nYou may not care about them, or at least some of them,\nbut they will be implicitly answered when an LLM implements what you think you want.</p>\n<p>If you look into the rewrites closer, you will quickly see that they treat backwards compatibility and other constraints\ndifferently. The Rust version doesn&#39;t maintain 100% backwards compatibility, for example it drops the CSRF token so\nthat caching is easier. It also dropped Redis in favour of in-process queues, for example for notifications.\nElixir rewrite is much closer to the Rails original. This already makes the comparison pretty much useless,\nas these differences are language independent. It&#39;s not like a clanker heard &quot;rewrite in Elixir&quot; and chose to keep\n100% backwards compatibility because of the language that has been used. But it gets even better!</p>\n<p>If you looked at the code, and yeah, I know, we should not be reading code anymore, you would quickly notice a lot\nof the things are not great. For example, I&#39;ve seen people complaining that Elixir&#39;s version got a single process\nsequentially processing all of the SQL queries, even if these were reads that could be running concurrently. That&#39;s\na fair complaint, but DHH <a href=\"https://x.com/dhh/status/2106864061832991134?s=20\" rel=\"nofollow ugc noopener\">seems to think</a> it&#39;s not the best look for Elixir that the clanker couldn&#39;t write\nperformant code. Not sure if I agree when I look at the Rust version. Cause in the Rust version some of the database\noperations are not async, which in some cases may be worse. You see, in Rust, when using an async runtime, the scheduling\nis not preemptive, but rather cooperative. If a task doesn&#39;t yield, no other task can run on the same worker thread.\nWhat this means in practice, is that the time spent in the task should be as short as possible. For example, when\nyou run an SQL query that takes 100ms, you don&#39;t want all the other tasks to wait as the task issuing the query is\nwaiting anyway. Thus, you should ideally use an async I/O operation that yields to the runtime\nwhile waiting for a response from the database. In the Rust rewrite some queries run on worker threads, but some run as\nblocking operations in async tasks. Same with locks. When using\nan async runtime, the safest bet is to use async locks like <code>tokio::sync::Mutex</code>. It&#39;s fine to use a non-async\nversion if you&#39;re sure that the lock is held for a very short time, but if you hold a sync lock for 10ms,\nall the tasks on the same thread wait for it, while the blocking task itself is not doing any work. Thus, I regret\nto inform you, that no, coding is not likely &quot;solved&quot;, and you still need to know what you&#39;re doing.</p>\n<p>Going further, it turned out that slop benchmarks are also, well, slop, as they only measure throughput, disregarding other properties of the system. For example, <a href=\"https://x.com/ZachSDaniel1/status/2106876400196059200\" rel=\"nofollow ugc noopener\">Zach Daniels measured new post notifications\ndelivery rate under load</a>, and discovered a 1% successful\ndelivery rate in the Rust version under heavy load. Not a great look. Benchmarks are hard.\nBut even the improved benchmark may not be actually testing what you want to test, depending on the characteristics of your system.\nYou see, when you test a system you may want to emphasize different properties under load.\nBoth DHH&#39;s and Zach&#39;s benchmarks were closed-loop benchmarks, so they were testing &quot;how many requests can the system handle in a specific amount of time?&quot;.\nThe test was using N clients, each sending a new message as soon as it gets a response for the previous one.\nIn many cases, though, the increased load on the system\nmay come from a big number of users performing an action at the same time, who will not wait for other users to finish their actions. In this case you would prefer\n<a href=\"https://grafana.com/docs/k6/latest/using-k6/scenarios/executors/constant-arrival-rate/\" rel=\"nofollow ugc noopener\">constant arrival rate</a>, ie. you want to send requests at a constant\nrate rather than making the rate depend on how fast the system can respond. If you want to check how reliable the events delivery is,\nyou would ideally want to compare the same number of events.\nHere, if you look at the results, it shows that Rust delivered 1% of notifications out of 6-7k, whereas Elixir delivered 100%\nof notifications out of ~1.7k. What it shows is that the Elixir version handles backpressure better, but this only holds as long as the number\nof requests don&#39;t overload the system.</p>\n<p>But let&#39;s leave benchmarks just for a moment and talk about trade-offs, cause programming is all about trade-offs. Sure, there are situations\nwhere a tool or a solution is clearly better, without any downsides, but it&#39;s quite rare, especially when you get to complex systems\nthat have to be reliable. I&#39;ve seen quite a lot of\npeople from the Elixir community coming to various conclusions after they&#39;ve seen the 1% delivery rate of the Rust version,\nwithout really trying to understand why it happens. The general consensus? Elixir is just better at concurrency! In Rust the\nscheduler is cooperative, how can you even live like this? I like Elixir, and I&#39;ve successfully used it in production, but it&#39;s not a silver bullet. Yes, Elixir (or other BEAM based languages) are very good at concurrency, and admittedly\nwriting concurrent code in Elixir is, in general, easier than in Rust, but it comes at a cost of not having low level control,\nhigher memory usage, and oftentimes, speed. There is a reason why <a href=\"https://github.com/rusterlium/rustler\" rel=\"nofollow ugc noopener\">rustler</a> exists. And proclaiming\nlanguage&#39;s superiority based on a single metric without knowing the root cause, may be misleading.</p>\n<p>Remember how I mentioned the closed-loop stress test shows just one of the properties of the system as Rust processed ~4 times more requests and thus\nhad to handle more events sent to the client through a WebSocket? I&#39;ve rerun the test with constant delivery rate using the DHH&#39;s Rust version and Zach&#39;s Elixir version with various fixes. At 100 POSTs/s\nthe clients received ~14% events in Rust. In Elixir it was ~60%. Still better, right? Not so fast! At this traffic level Rust didn&#39;t have any\nHTTP errors. Elixir timed out on ~23% of HTTP POST requests. And what about latency? The worst event delivery latency was close to\n180s. In Rust when the deliveries are lagging, clients get disconnected. On reconnect, a browser client would fetch the latest messages,\nwhich largely invalidates the need for the missed events.\nWhat do you think is better UX: the client silently reconnecting in the background and fetching the new updates, or waiting for an\nupdate about a new message for 3 minutes? Which just shows that, again, a single metric doesn&#39;t tell the whole story.</p>\n<p>Getting back to trade-offs. Do you know <em>why</em> the Rust version drops so many messages under heavy load? It uses <code>tokio::sync::broadcast</code> channel\nto broadcast events to connected clients. An event informing about a new chat message may need to be sent to multiple connections,\nso that makes total sense. One of the properties of the broadcast channel is that it&#39;s bounded by a set capacity. If a receiver can&#39;t\nhandle the messages fast enough, the receiver will get a <code>RecvError::Lagged</code> error (refer to the docs to learn more about <a href=\"https://docs.rs/tokio/latest/tokio/sync/broadcast/index.html#lagging\" rel=\"nofollow ugc noopener\">lagging</a>). The broadcast capacity in the Rust rewrite was set to 256. In Elixir, the <code>GenServer</code> process is handling events\ndelivery and by default <code>GenServer</code> mailboxes don&#39;t have a limit. A single line change in Rust:</p>\n<pre><code>- stream_capacity: 256\n+ stream_capacity: 16384</code></pre>\n<p>increases the delivery rate at 100 req/s from ~14% to ~90%. It&#39;s better than Elixir now, isn&#39;t it? Not really. I think that\nthis version is actually worse because in this case it&#39;s better to fail fast and force the client to reconnect rather than handle things extremely slowly. You know\nwhat else changed after raising the capacity? The pMAX event delivery latency went up from 11s to &gt;130s, similarly to how Elixir behaved, which I think\nis strictly worse than disconnecting the lagging clients. If you ask me, even 11s is too long and if the receiver can&#39;t pass the event faster\nthan that, it&#39;s better to drop the event. Which shows how important it is to set sensible constraints in the system, and that, in fact,\nElixir isn&#39;t just automatically fixing all of the concurrency issues out of the box. Also, if you don&#39;t set bounds yourself, you will likely run into\nexternal limits. During the 100 req/s stress test the Elixir version reached 1.8GB memory usage. Another thing to ponder on: is it better to drop some messages or\nget OOM killed? Trade-offs all the way. Elixir is great, but it won&#39;t magically solve all of your problems. No matter the language you&#39;re using, you have to think\nabout failure modes and trade-offs.</p>\n<p>So what have we learned today? You still have to think critically. It&#39;s good to know what you&#39;re doing. Don&#39;t\nmake hasty assumptions based on a single metric. If you want a reliable system you should know when to fail. Also, benchmarks are hard.</p>\n<hr />\n<p>If you like this post please consider following me on <a href=\"http://twitter.com/drogus\" rel=\"nofollow ugc noopener\">Twitter</a>.</p>","headings":[{"level":1,"text":"I'm sorry, but you still have to think","id":"i-m-sorry-but-you-still-have-to-think"}]}}