{"article":{"slug":"the-data-race-that-wasnt-a-bug-and-the-one-that-was","title":"The Data Race That Wasn't a Bug (and the One That Was)","subtitle":null,"summary":"Jesús Espino digs into a Go race detector warning inside net/http: how the HTTP client sends a request body across goroutines and can keep reading your buffer after Client.Do returns, why VictoriaMetrics kept that race in vmagent as harmless, and why the same behavior could corrupt data in vmauth, so they fixed it there.","content_type":"blog_post","language":"en","canonical_url":"https://victoriametrics.com/blog/http-race-condition/","author":{"name":"Jesús Espino","url":null,"person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"VictoriaMetrics","url":"https://victoriametrics.com/","listing_slug":null,"listing":null},"topics":[{"name":"Go","slug":"go","url":"https://listedarticles.com/topics/go"},{"name":"Programming","slug":"programming","url":"https://listedarticles.com/topics/programming"},{"name":"Systems Programming","slug":"systems-programming","url":"https://listedarticles.com/topics/systems-programming"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":3210,"reading_minutes":14,"published_at":"2026-10-06T00:00:00.000Z","added_at":"2026-10-07T14:27:26.953Z","updated_at":"2026-10-07T14:27:26.953Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":true},"profile_url":"https://listedarticles.com/articles/the-data-race-that-wasnt-a-bug-and-the-one-that-was","markdown_url":"https://listedarticles.com/articles/the-data-race-that-wasnt-a-bug-and-the-one-that-was.md","example":false,"citation":"Jesús Espino, VictoriaMetrics. \"The Data Race That Wasn't a Bug (and the One That Was).\" 6 Oct 2026. https://victoriametrics.com/blog/http-race-condition/ (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://victoriametrics.com/blog/http-race-condition/"},"body_markdown":"# The Data Race That Wasn't a Bug (and the One That Was)\n\nImagine this: you are testing the performance of some part of your application. Everything is going smoothly, the numbers look good, and as a last check you turn on Go’s race detector. Then, out of nowhere, it prints a warning you didn’t expect:\n\n```\nWARNING: DATA RACE\n```\nSo you look at it. You look at it again, and again, and you think: *“What the…?”* The race is between your code and a goroutine you never started, somewhere deep inside `net/http`. You have no idea how that is possible, or why. Time to dig in.\n\nThat is more or less what happened to [Vadim Alekseev](https://x.com/vadimaleksv) while running [vmagent](https://docs.victoriametrics.com/victoriametrics/vmagent/) under the race detector. He tracked the behavior down, reduced it to a small program, and reported it upstream as [golang/go#81445](https://github.com/golang/go/issues/81445). The same behavior also showed up in [vmauth](https://docs.victoriametrics.com/victoriametrics/vmauth/). In the end we left one of those races in place on purpose, and fixed the other. To see why, we first need to understand what the race actually is.\n\nThe Go internals in this post (function names, buffer sizes, timeouts) come from the Go 1.27 `net/http` source. The same mechanics have been there for many releases.\n\n## The code and the race\n\nHere is the client side of the program from [the upstream issue](https://github.com/golang/go/issues/81445) (the full runnable version is there). It sends a batch of logs to a server, and retries if the server answers with an error. To avoid allocating a new buffer for every attempt, it reuses a single global `bytes.Buffer`:\n\n```\ntype LogEntry struct {\n Timestamp int64\n Content   string\n}\nvar buf = bytes.NewBuffer(nil)\nfunc insertLogs(client *http.Client, logs []LogEntry) error {\n buf.Reset()\n if err := json.MarshalWrite(buf, logs); err != nil {\n  panic(err)\n }\n body := bytes.NewReader(buf.Bytes())\n resp, err := client.Post(\"http://example.com/insert/logs\", \"application/json\", body)\n if err != nil {\n  return err\n }\n defer resp.Body.Close()\n _, _ = io.Copy(io.Discard, resp.Body)\n if resp.StatusCode/100 != 2 {\n  return fmt.Errorf(\"unexpected status code %d\", resp.StatusCode)\n }\n return nil\n}\n// ...and in main, retry until it works:\nfor range 100 {\n err := insertLogs(client, logs)\n if err != nil {\n  fmt.Println(\"error inserting logs:\", err)\n  continue\n }\n break\n}\n```\n`logs` is 1024 entries with 256 bytes of content each, so the body is about 300 KB. On the other side there’s a small test server that always answers `502 Bad Gateway`, without even looking at the request body. The program sends, gets a `502`, and tries again. Nothing unusual. Run it with `-race` and, after a few attempts:\n\n```\nerror inserting logs: unexpected status code 502\nerror inserting logs: unexpected status code 502\n...\n==================\nWARNING: DATA RACE\nWrite at 0x00c000390f60 by main goroutine:\n  runtime.slicecopy()\n  encoding/json/v2.makeStructArshaler.func2()\n  ...\n  encoding/json/v2.MarshalWrite()\n  main.insertLogs()\n  main.main()\nPrevious read at 0x00c000390f60 by goroutine 49:\n  runtime.slicecopy()\n  bytes.(*Reader).Read()\n      bytes/reader.go:44\n  io.(*LimitedReader).Read()\n  ...\n  io.Copy()\n  net.genericReadFrom()\n  net.(*TCPConn).ReadFrom()\n  ...\n  net/http.persistConnWriter.ReadFrom()\n  bufio.(*Writer).ReadFrom()\n  ...\n  net/http.(*transferWriter).doBodyCopy()\n  net/http.(*transferWriter).writeBody()\n  net/http.(*Request).write()\n  net/http.(*persistConn).writeLoop()\n==================\nexit status 66\n```\nThe report describes two accesses to the same memory address, `0x00c000390f60`:\n\n- **The write** happens in the`main` goroutine, inside`insertLogs` . It’s`json.MarshalWrite` writing the JSON for a new attempt into`buf` .\n- **The read** happens in another goroutine, one we never started. It’s a`bytes.Reader` reading`buf` , called from a function named`writeLoop` inside`net/http` .\n\nSo something inside `net/http` was reading our buffer while our code was writing the next attempt into it. That’s surprising: by the time we write, `client.Post` for the previous attempt has already returned, and we have read and closed its response.\n\nTo understand how that can happen, we need to take a step back and look at how Go’s HTTP client really sends a request.\n\n## How the HTTP client sends a request: three goroutines\n\nWhen you call `client.Post`, it feels like a single operation: send the request, get the response. Under the hood, at least **three goroutines** are involved.\n\nFor every HTTP/1.1 connection it opens, the `http.Transport` creates a `persistConn` with two goroutines of its own:\n\n- **`writeLoop`** sends requests over the connection: the request line, the headers, then the body.\n- **`readLoop`** reads responses from the connection and delivers them.\n\nThe third goroutine is **yours**, the one that called `client.Post`.\n\nSo `writeLoop` and `readLoop` are already running, one pair per connection, waiting for work. When your goroutine calls `client.Post`, it goes down through the client and the transport until it reaches `persistConn.roundTrip`, the function that connects your request with those two goroutines. It does two things.\n\nFirst, it gives the request to **both** goroutines, one message on a channel to each:\n\n```\n// net/http/transport.go (simplified)\n// Write the request concurrently with waiting for a response,\n// in case the server decides to reply before reading our full\n// request body.\npc.writech <- writeRequest{req, writeErrCh, continueCh} // writeLoop: \"send this\"\npc.reqch <- requestAndChan{treq: req, ch: resc, ...}    // readLoop: \"wait for its answer\"\n```\nFrom this point on, `writeLoop` is sending the request and `readLoop` is waiting for the response, both at the same time.\n\nSecond, your goroutine waits to see what happens: either `writeLoop` reports that it finished writing the request, or `readLoop` delivers a response:\n\n```\nfor {\n select {\n case err := <-writeErrCh: // writeLoop finished writing the request\n  ...\n case re := <-resc: // readLoop got a response\n  return handleResponse(re)\n ...\n }\n}\n```\nLet’s see what this looks like visually:\n\nOn the left, your goroutine does the two channel sends we just saw: `writech` tells `writeLoop` to send the request, and `reqch` tells `readLoop` to wait for its answer. Then it parks in the `select`, waiting for whichever arrow comes back first.\n\nOn the right side there are two separate paths. The **request + body** arrow, from `writeLoop` to the server, is the request going out: first the request line and the headers, then the body, chunk after chunk. The arrow from the server back to `readLoop` is the response coming in. A TCP connection carries data in both directions at the same time, and each direction is independent. The response doesn’t have to wait for the request to finish before it can travel back.\n\nTypically, `writeLoop` sends the request line and the headers, and then starts sending the body. The server receives the request line and the headers, and based on them it calls the right handler. The handler reads and processes the body, and then sends a reply back, which `readLoop` picks up and passes to your goroutine along the **response** arrow.\n\nBut it isn’t always like that. Sometimes the headers alone are enough for the server to make a decision. The request may be unauthorized, or its `Content-Length` may say the body is bigger than the server accepts, or the server may be a proxy that can’t reach its backend. In those cases, the server can reply right away, without waiting for the body. The reply travels back on the other direction of the connection, `readLoop` passes it to your goroutine, your `select` takes the `resc` case, and `roundTrip` returns. So the client is still sending the body while the server has already sent its reply. That’s the interesting case for us here.\n\nAnd here’s the important consequence. Your `select` returns as soon as the response arrives, and it doesn’t wait for `writeLoop` to finish. So `client.Post` can return, and your code can move on with the response in hand, **while `writeLoop` is still running in the background, sending the rest of your body**.\n\nGo does document this. The [`http.RoundTripper`](https://pkg.go.dev/net/http#RoundTripper) interface says:\n\nRoundTrip must always close the body, including on errors, but depending on the implementation may do so in a separate goroutine even after RoundTrip returns. This means that callers wanting to reuse the body for subsequent requests must arrange to wait for the Close call before doing so.\n\nand [`Client.Do`](https://pkg.go.dev/net/http#Client.Do) warns that “the Body may be closed asynchronously after Do returns”.\n\nIf we go back to our `DATA RACE` report, the goroutine doing the read was running `writeLoop`. So now we have the *who*: it’s the goroutine that sends our request, and it can keep running after `client.Post` has returned. What we still don’t know is *what* exactly it reads from our buffer, and when. For that, we have to follow the bytes.\n\n## Following the bytes: your buffer, the reader, and the scratch buffer\n\nLet’s follow our JSON from the moment we write it until it leaves the machine:\n\nIt all starts inside the **your code** box, with **buf’s array**. Remember that `buf` is a global variable: it lives for the whole program, so every call to `insertLogs` uses the same buffer, and the same array inside it, which stays in memory between calls. `buf.Reset()` doesn’t throw that array away, it just marks it as empty, so every time we retry `insertLogs`, the new JSON goes into the same memory. That’s the point of reusing `buf`: no new allocation on every retry. The `bytes.Reader` we pass to `client.Post` doesn’t copy anything either: it reads straight from that same array.\n\nTo send the body, `writeLoop` follows the first arrow in our diagram: it copies 32 KiB chunks from our buffer into a buffer of its own, and sends them to the server through the network.\n\nThen the arrow at the bottom takes us back to buf’s array for the next chunk, and the whole thing **repeats until EOF or until the socket is closed**.\n\nNow we have all the pieces. Let’s put them together.\n\n## Putting it together: how the race happens\n\nNow let’s put all the pieces on one timeline:\n\nEach row is one goroutine, and the bottom row is the memory they share: **buf’s underlying array**. Let’s follow it from left to right.\n\nFirst, in **your code**, we call `buf.Reset()` and encode our logs, which writes the JSON into the array. Then we call `client.Do`, and our goroutine starts waiting.\n\nThat call sets the other two goroutines in motion, in parallel. `writeLoop` starts sending our request to the server: it copies the body out of the array, 32 KiB at a time, and writes each chunk to the socket. At the same time, `readLoop` starts waiting for the server to answer.\n\nThe answer comes back before the body is fully sent. The server only needed our headers to decide, so it replies `502` without reading the body. And because so much of the body is left unread, it adds **Connection: close**. `readLoop` hands that response back to us, and **Do returns 502**. For us, the request is done.\n\nBut `writeLoop` doesn’t know the server already answered. It’s **still reading!** from our array, sending the rest of the body.\n\nRight then, our code retries: `buf.Reset()` and encode again, and **WRITE again** goes into the same array `writeLoop` is still reading. That’s the **overlap window**: one goroutine writing new data into the array while another is still reading the old data from it. That’s our data race.\n\nFinally, because the server asked to close the connection, `readLoop` **closes the socket**. `writeLoop`’s next write fails, it closes our body, and it exits.\n\nThis race doesn’t show up every time: locally the body is sent almost instantly, and if the connection is going to be reused, Go waits up to 50 ms for `writeLoop` before letting us continue. But what is the actual problem when the race does happen?\n\n### What can actually go wrong\n\nFirst, the good news: our code writes the array and `writeLoop` only reads it, so our buffer itself is never corrupted. And in our example, every retry encodes the same logs, so the bytes `writeLoop` reads are the same either way.\n\nBut let’s do a small mental exercise. Imagine we reuse this buffer for *different* requests, like sending a new batch of logs each time. What would the leftover `writeLoop` send for the previous request?\n\nSay the previous request was:\n\n```\n[{\"Timestamp\":1790000000111111111,\"Content\":\"disk full\"}]\n```\nand the next one, written into the same array while `writeLoop` is still copying, is:\n\n```\n[{\"Timestamp\":1790000000999999999,\"Content\":\"disk ok!!\"}]\n```\nIf `writeLoop` had copied the first half when the new data landed, it sends the old first half and the new second half:\n\n```\n[{\"Timestamp\":1790000000111119999,\"Content\":\"disk ok!!\"}]\n```\nThat’s valid JSON, but with a timestamp that neither request ever had. Both payloads have the same shape, so the pieces fit together perfectly.\n\nNow imagine the next request is shorter than the previous one. `writeLoop` still thinks the body has the old length, so it copies the new data and then keeps going into what’s left of the old data, because `buf.Reset()` doesn’t clear anything:\n\n```\nprevious : [{\"Timestamp\":1,\"Content\":\"first\"},{\"Timestamp\":2,\"Content\":\"second\"}]\nnext     : [{\"Timestamp\":9,\"Content\":\"new\"}]\nsent     : [{\"Timestamp\":9,\"Content\":\"new\"}]},{\"Timestamp\":2,\"Content\":\"second\"}]\n```\nThis time it’s not even valid JSON.\n\nThat sounds scary. But whether it matters depends on one question: **who, if anyone, reads that mixed copy?**\n\n## vmagent: a real race we decided to keep\n\nvmagent’s remote write client ([#11507](https://github.com/VictoriaMetrics/VictoriaMetrics/issues/11507)) is the same pattern at a bigger scale. A worker pulls a block of compressed samples from its queue into a byte slice that it **reuses** on every iteration, and sends it with a fresh reader:\n\n```\n// app/vmagent/remotewrite/client.go (simplified)\nfunc (c *client) runWorker(readBlock func(dst []byte) ([]byte, bool)) {\n var block []byte\n for {\n  block, ok = readBlock(block[:0]) // <- overwrites the previous block's bytes\n  ...\n  c.sendBlock(block)\n }\n}\nfunc (c *client) newRequest(url string, body []byte) (*http.Request, error) {\n reqBody := bytes.NewBuffer(body) // <- a fresh reader for every request\n req, err := http.NewRequest(http.MethodPost, url, reqBody)\n ...\n}\n```\nPoint vmagent at a remote storage that answers `200` without reading the body (Vadim used `httpbin.org/status/200`), push a big import through it, and the race detector reports the queue writing the next block into `block` while `writeLoop` is still reading the previous one. It’s exactly the race we just dissected. We even merged a fix for it. Three days later we [reverted it](https://github.com/VictoriaMetrics/VictoriaMetrics/commit/8e9af3911f045ed69ed150393806cbebd37b4cd3), and later closed the issue without a fix. Here’s why.\n\n### Why we left it alone\n\nThe key to understanding this is how we build the request body: with `bytes.NewReader` in our example, or `bytes.NewBuffer` in vmagent. Both wrap our array with their own length and read position, so that part isn’t shared between requests. The only shared part is the array underneath. And the [Go memory model](https://go.dev/ref/mem#restrictions) guarantees that reading a byte while it’s being overwritten gives us either the old value or the new one, never some corrupted in-between state. So our readers are safe: the data they send can be mixed, but the program itself won’t break. Nothing crashes, and no other memory gets corrupted.\n\nSo the worst case is some mixed data, and in vmagent nobody uses it. It goes into a request the server has **already answered**, without reading the body, so the server didn’t want it anyway. Usually the connection is being closed too, so nothing on the other side ever reads them.\n\nWith nothing to protect, fixing it anyway would only cost us. A new buffer per request would undo the savings of reusing it. And waiting for the transport to finish with the body, which we tried and then [reverted](https://github.com/VictoriaMetrics/VictoriaMetrics/commit/8e9af3911f045ed69ed150393806cbebd37b4cd3), could stall workers, deadlock in rare cases, and still didn’t cover everything.\n\nPaying in performance and complexity to protect bytes nobody reads isn’t a good trade, so we kept the race.\n\n“Harmless” is not “free”. This is a known, accepted race-detector report, not a silent one. We’ll revisit it if Go gains an official way to wait for the transport to be done with a request body. In [the upstream issue](https://github.com/golang/go/issues/81445), Damien Neil suggested a new `Request.Close` method for exactly that.\n\n## vmauth: when the race is a real bug\n\nvmauth hit the same behavior ([#11508](https://github.com/VictoriaMetrics/VictoriaMetrics/issues/11508)), but here it was a real problem. vmauth is a proxy, and it can **retry**: if a backend fails or answers with an error like `503`, vmauth sends the same request to the next backend.\n\nTo send the same body twice, vmauth keeps it in memory in a type called `bufferedBody`. Before the fix, `bufferedBody` was also the reader: it held the bytes plus its own read position. On a retry, vmauth rewound that position to zero and handed **the same `bufferedBody`** to the next request.\n\nThat’s the key difference from vmagent. In vmagent, each request has its own reader, and only the bytes are shared. In vmauth, both requests shared **the reader itself**, including its read position. So the leftover `writeLoop` from the failed request could keep moving that position, and when it finished, it even reset it to zero, right in the middle of the retry’s upload.\n\nThe result: the next backend could receive the body with pieces missing or repeated. Sometimes the length came out wrong and the request failed. But sometimes it came out exactly right, and the backend accepted a corrupted request without anyone noticing. This time the mixed data doesn’t go nowhere: it goes to a healthy backend that processes it.\n\n### The fix: correctness first\n\nThe fix ([#11647](https://github.com/VictoriaMetrics/VictoriaMetrics/pull/11647)) does what vmagent already does: **never give two requests the same reader.** vmauth still keeps the bytes in `bufferedBody`, but every attempt now gets its own new `bytes.Buffer` over them. The leftover `writeLoop` can keep reading its own reader as long as it wants, and it can’t touch the retry’s. In the case of vmauth, the shared bytes are only ever read, never written, so there’s no race at all.\n\nHere, correctness wins easily: a proxy that can silently send corrupted data to a backend isn’t acceptable at any speed. And the cost is tiny: a couple of small allocations per attempt, with no copy of the data, on a path where the network round trip to the backend costs far more.\n\nSo the same race detector warning led to two opposite decisions: keep it in vmagent, fix it in vmauth. Let’s wrap up with what made the difference.\n\n## What to take away\n\nBoth warnings came from the same `net/http` behavior: the transport writes the request body in its own goroutine, and it can return the response to you before it’s done reading your body. The race detector was right both times. What it can’t tell you is whether the race matters, and that came down to two questions:\n\n1. **What exactly is shared?** Plain bytes that get copied somewhere, or*state* that decides behavior, like an offset, a length or a pointer? A race on bytes gives you stale bytes. A race on a cursor gives you wrong behavior.\n2. **Who consumes the result?** In vmagent, the mixed copy goes into a request the server has already answered, on a connection that’s being closed. In vmauth, the racing cursor decided what a*healthy backend* received and ingested.\n\nSo in vmagent we accepted the race and documented why, because fixing it would cost performance and complexity to protect bytes nobody reads. In vmauth we fixed it, by removing the sharing rather than adding synchronization, because correctness comes first.\n\nIf you use Go’s HTTP client, the rules are short:\n\n- Assume the transport may still be reading your request body after `Do` returns.\n- Avoid giving the same stateful `io.Reader` to two requests.\n- Be careful with reusing, pooling or mutating the memory behind a body you’ve already sent, until the transport has called `Close()` on it.\n\nAnd keep an eye on [golang/go#81445](https://github.com/golang/go/issues/81445): if Go gets an official way to wait until the transport is done with a request, most of this goes away.\n","body_html":"<h1 id=\"the-data-race-that-wasn-t-a-bug-and-the-one-that-was\">The Data Race That Wasn&#39;t a Bug (and the One That Was)</h1>\n<p>Imagine this: you are testing the performance of some part of your application. Everything is going smoothly, the numbers look good, and as a last check you turn on Go’s race detector. Then, out of nowhere, it prints a warning you didn’t expect:</p>\n<pre><code>WARNING: DATA RACE</code></pre>\n<p>So you look at it. You look at it again, and again, and you think: <em>“What the…?”</em> The race is between your code and a goroutine you never started, somewhere deep inside <code>net/http</code>. You have no idea how that is possible, or why. Time to dig in.</p>\n<p>That is more or less what happened to <a href=\"https://x.com/vadimaleksv\" rel=\"nofollow ugc noopener\">Vadim Alekseev</a> while running <a href=\"https://docs.victoriametrics.com/victoriametrics/vmagent/\" rel=\"nofollow ugc noopener\">vmagent</a> under the race detector. He tracked the behavior down, reduced it to a small program, and reported it upstream as <a href=\"https://github.com/golang/go/issues/81445\" rel=\"nofollow ugc noopener\">golang/go#81445</a>. The same behavior also showed up in <a href=\"https://docs.victoriametrics.com/victoriametrics/vmauth/\" rel=\"nofollow ugc noopener\">vmauth</a>. In the end we left one of those races in place on purpose, and fixed the other. To see why, we first need to understand what the race actually is.</p>\n<p>The Go internals in this post (function names, buffer sizes, timeouts) come from the Go 1.27 <code>net/http</code> source. The same mechanics have been there for many releases.</p>\n<h2 id=\"the-code-and-the-race\">The code and the race</h2>\n<p>Here is the client side of the program from <a href=\"https://github.com/golang/go/issues/81445\" rel=\"nofollow ugc noopener\">the upstream issue</a> (the full runnable version is there). It sends a batch of logs to a server, and retries if the server answers with an error. To avoid allocating a new buffer for every attempt, it reuses a single global <code>bytes.Buffer</code>:</p>\n<pre><code>type LogEntry struct {\n Timestamp int64\n Content   string\n}\nvar buf = bytes.NewBuffer(nil)\nfunc insertLogs(client *http.Client, logs []LogEntry) error {\n buf.Reset()\n if err := json.MarshalWrite(buf, logs); err != nil {\n  panic(err)\n }\n body := bytes.NewReader(buf.Bytes())\n resp, err := client.Post(&quot;http://example.com/insert/logs&quot;, &quot;application/json&quot;, body)\n if err != nil {\n  return err\n }\n defer resp.Body.Close()\n _, _ = io.Copy(io.Discard, resp.Body)\n if resp.StatusCode/100 != 2 {\n  return fmt.Errorf(&quot;unexpected status code %d&quot;, resp.StatusCode)\n }\n return nil\n}\n// ...and in main, retry until it works:\nfor range 100 {\n err := insertLogs(client, logs)\n if err != nil {\n  fmt.Println(&quot;error inserting logs:&quot;, err)\n  continue\n }\n break\n}</code></pre>\n<p><code>logs</code> is 1024 entries with 256 bytes of content each, so the body is about 300 KB. On the other side there’s a small test server that always answers <code>502 Bad Gateway</code>, without even looking at the request body. The program sends, gets a <code>502</code>, and tries again. Nothing unusual. Run it with <code>-race</code> and, after a few attempts:</p>\n<pre><code>error inserting logs: unexpected status code 502\nerror inserting logs: unexpected status code 502\n...\n==================\nWARNING: DATA RACE\nWrite at 0x00c000390f60 by main goroutine:\n  runtime.slicecopy()\n  encoding/json/v2.makeStructArshaler.func2()\n  ...\n  encoding/json/v2.MarshalWrite()\n  main.insertLogs()\n  main.main()\nPrevious read at 0x00c000390f60 by goroutine 49:\n  runtime.slicecopy()\n  bytes.(*Reader).Read()\n      bytes/reader.go:44\n  io.(*LimitedReader).Read()\n  ...\n  io.Copy()\n  net.genericReadFrom()\n  net.(*TCPConn).ReadFrom()\n  ...\n  net/http.persistConnWriter.ReadFrom()\n  bufio.(*Writer).ReadFrom()\n  ...\n  net/http.(*transferWriter).doBodyCopy()\n  net/http.(*transferWriter).writeBody()\n  net/http.(*Request).write()\n  net/http.(*persistConn).writeLoop()\n==================\nexit status 66</code></pre>\n<p>The report describes two accesses to the same memory address, <code>0x00c000390f60</code>:</p>\n<ul><li><strong>The write</strong> happens in the<code>main</code> goroutine, inside<code>insertLogs</code> . It’s<code>json.MarshalWrite</code> writing the JSON for a new attempt into<code>buf</code> .</li><li><strong>The read</strong> happens in another goroutine, one we never started. It’s a<code>bytes.Reader</code> reading<code>buf</code> , called from a function named<code>writeLoop</code> inside<code>net/http</code> .</li></ul>\n<p>So something inside <code>net/http</code> was reading our buffer while our code was writing the next attempt into it. That’s surprising: by the time we write, <code>client.Post</code> for the previous attempt has already returned, and we have read and closed its response.</p>\n<p>To understand how that can happen, we need to take a step back and look at how Go’s HTTP client really sends a request.</p>\n<h2 id=\"how-the-http-client-sends-a-request-three-goroutines\">How the HTTP client sends a request: three goroutines</h2>\n<p>When you call <code>client.Post</code>, it feels like a single operation: send the request, get the response. Under the hood, at least <strong>three goroutines</strong> are involved.</p>\n<p>For every HTTP/1.1 connection it opens, the <code>http.Transport</code> creates a <code>persistConn</code> with two goroutines of its own:</p>\n<ul><li><strong><code>writeLoop</code></strong> sends requests over the connection: the request line, the headers, then the body.</li><li><strong><code>readLoop</code></strong> reads responses from the connection and delivers them.</li></ul>\n<p>The third goroutine is <strong>yours</strong>, the one that called <code>client.Post</code>.</p>\n<p>So <code>writeLoop</code> and <code>readLoop</code> are already running, one pair per connection, waiting for work. When your goroutine calls <code>client.Post</code>, it goes down through the client and the transport until it reaches <code>persistConn.roundTrip</code>, the function that connects your request with those two goroutines. It does two things.</p>\n<p>First, it gives the request to <strong>both</strong> goroutines, one message on a channel to each:</p>\n<pre><code>// net/http/transport.go (simplified)\n// Write the request concurrently with waiting for a response,\n// in case the server decides to reply before reading our full\n// request body.\npc.writech &lt;- writeRequest{req, writeErrCh, continueCh} // writeLoop: &quot;send this&quot;\npc.reqch &lt;- requestAndChan{treq: req, ch: resc, ...}    // readLoop: &quot;wait for its answer&quot;</code></pre>\n<p>From this point on, <code>writeLoop</code> is sending the request and <code>readLoop</code> is waiting for the response, both at the same time.</p>\n<p>Second, your goroutine waits to see what happens: either <code>writeLoop</code> reports that it finished writing the request, or <code>readLoop</code> delivers a response:</p>\n<pre><code>for {\n select {\n case err := &lt;-writeErrCh: // writeLoop finished writing the request\n  ...\n case re := &lt;-resc: // readLoop got a response\n  return handleResponse(re)\n ...\n }\n}</code></pre>\n<p>Let’s see what this looks like visually:</p>\n<p>On the left, your goroutine does the two channel sends we just saw: <code>writech</code> tells <code>writeLoop</code> to send the request, and <code>reqch</code> tells <code>readLoop</code> to wait for its answer. Then it parks in the <code>select</code>, waiting for whichever arrow comes back first.</p>\n<p>On the right side there are two separate paths. The <strong>request + body</strong> arrow, from <code>writeLoop</code> to the server, is the request going out: first the request line and the headers, then the body, chunk after chunk. The arrow from the server back to <code>readLoop</code> is the response coming in. A TCP connection carries data in both directions at the same time, and each direction is independent. The response doesn’t have to wait for the request to finish before it can travel back.</p>\n<p>Typically, <code>writeLoop</code> sends the request line and the headers, and then starts sending the body. The server receives the request line and the headers, and based on them it calls the right handler. The handler reads and processes the body, and then sends a reply back, which <code>readLoop</code> picks up and passes to your goroutine along the <strong>response</strong> arrow.</p>\n<p>But it isn’t always like that. Sometimes the headers alone are enough for the server to make a decision. The request may be unauthorized, or its <code>Content-Length</code> may say the body is bigger than the server accepts, or the server may be a proxy that can’t reach its backend. In those cases, the server can reply right away, without waiting for the body. The reply travels back on the other direction of the connection, <code>readLoop</code> passes it to your goroutine, your <code>select</code> takes the <code>resc</code> case, and <code>roundTrip</code> returns. So the client is still sending the body while the server has already sent its reply. That’s the interesting case for us here.</p>\n<p>And here’s the important consequence. Your <code>select</code> returns as soon as the response arrives, and it doesn’t wait for <code>writeLoop</code> to finish. So <code>client.Post</code> can return, and your code can move on with the response in hand, <strong>while <code>writeLoop</code> is still running in the background, sending the rest of your body</strong>.</p>\n<p>Go does document this. The <a href=\"https://pkg.go.dev/net/http#RoundTripper\" rel=\"nofollow ugc noopener\"><code>http.RoundTripper</code></a> interface says:</p>\n<p>RoundTrip must always close the body, including on errors, but depending on the implementation may do so in a separate goroutine even after RoundTrip returns. This means that callers wanting to reuse the body for subsequent requests must arrange to wait for the Close call before doing so.</p>\n<p>and <a href=\"https://pkg.go.dev/net/http#Client.Do\" rel=\"nofollow ugc noopener\"><code>Client.Do</code></a> warns that “the Body may be closed asynchronously after Do returns”.</p>\n<p>If we go back to our <code>DATA RACE</code> report, the goroutine doing the read was running <code>writeLoop</code>. So now we have the <em>who</em>: it’s the goroutine that sends our request, and it can keep running after <code>client.Post</code> has returned. What we still don’t know is <em>what</em> exactly it reads from our buffer, and when. For that, we have to follow the bytes.</p>\n<h2 id=\"following-the-bytes-your-buffer-the-reader-and-the-scratch-buffe\">Following the bytes: your buffer, the reader, and the scratch buffer</h2>\n<p>Let’s follow our JSON from the moment we write it until it leaves the machine:</p>\n<p>It all starts inside the <strong>your code</strong> box, with <strong>buf’s array</strong>. Remember that <code>buf</code> is a global variable: it lives for the whole program, so every call to <code>insertLogs</code> uses the same buffer, and the same array inside it, which stays in memory between calls. <code>buf.Reset()</code> doesn’t throw that array away, it just marks it as empty, so every time we retry <code>insertLogs</code>, the new JSON goes into the same memory. That’s the point of reusing <code>buf</code>: no new allocation on every retry. The <code>bytes.Reader</code> we pass to <code>client.Post</code> doesn’t copy anything either: it reads straight from that same array.</p>\n<p>To send the body, <code>writeLoop</code> follows the first arrow in our diagram: it copies 32 KiB chunks from our buffer into a buffer of its own, and sends them to the server through the network.</p>\n<p>Then the arrow at the bottom takes us back to buf’s array for the next chunk, and the whole thing <strong>repeats until EOF or until the socket is closed</strong>.</p>\n<p>Now we have all the pieces. Let’s put them together.</p>\n<h2 id=\"putting-it-together-how-the-race-happens\">Putting it together: how the race happens</h2>\n<p>Now let’s put all the pieces on one timeline:</p>\n<p>Each row is one goroutine, and the bottom row is the memory they share: <strong>buf’s underlying array</strong>. Let’s follow it from left to right.</p>\n<p>First, in <strong>your code</strong>, we call <code>buf.Reset()</code> and encode our logs, which writes the JSON into the array. Then we call <code>client.Do</code>, and our goroutine starts waiting.</p>\n<p>That call sets the other two goroutines in motion, in parallel. <code>writeLoop</code> starts sending our request to the server: it copies the body out of the array, 32 KiB at a time, and writes each chunk to the socket. At the same time, <code>readLoop</code> starts waiting for the server to answer.</p>\n<p>The answer comes back before the body is fully sent. The server only needed our headers to decide, so it replies <code>502</code> without reading the body. And because so much of the body is left unread, it adds <strong>Connection: close</strong>. <code>readLoop</code> hands that response back to us, and <strong>Do returns 502</strong>. For us, the request is done.</p>\n<p>But <code>writeLoop</code> doesn’t know the server already answered. It’s <strong>still reading!</strong> from our array, sending the rest of the body.</p>\n<p>Right then, our code retries: <code>buf.Reset()</code> and encode again, and <strong>WRITE again</strong> goes into the same array <code>writeLoop</code> is still reading. That’s the <strong>overlap window</strong>: one goroutine writing new data into the array while another is still reading the old data from it. That’s our data race.</p>\n<p>Finally, because the server asked to close the connection, <code>readLoop</code> <strong>closes the socket</strong>. <code>writeLoop</code>’s next write fails, it closes our body, and it exits.</p>\n<p>This race doesn’t show up every time: locally the body is sent almost instantly, and if the connection is going to be reused, Go waits up to 50 ms for <code>writeLoop</code> before letting us continue. But what is the actual problem when the race does happen?</p>\n<h3 id=\"what-can-actually-go-wrong\">What can actually go wrong</h3>\n<p>First, the good news: our code writes the array and <code>writeLoop</code> only reads it, so our buffer itself is never corrupted. And in our example, every retry encodes the same logs, so the bytes <code>writeLoop</code> reads are the same either way.</p>\n<p>But let’s do a small mental exercise. Imagine we reuse this buffer for <em>different</em> requests, like sending a new batch of logs each time. What would the leftover <code>writeLoop</code> send for the previous request?</p>\n<p>Say the previous request was:</p>\n<pre><code>[{&quot;Timestamp&quot;:1790000000111111111,&quot;Content&quot;:&quot;disk full&quot;}]</code></pre>\n<p>and the next one, written into the same array while <code>writeLoop</code> is still copying, is:</p>\n<pre><code>[{&quot;Timestamp&quot;:1790000000999999999,&quot;Content&quot;:&quot;disk ok!!&quot;}]</code></pre>\n<p>If <code>writeLoop</code> had copied the first half when the new data landed, it sends the old first half and the new second half:</p>\n<pre><code>[{&quot;Timestamp&quot;:1790000000111119999,&quot;Content&quot;:&quot;disk ok!!&quot;}]</code></pre>\n<p>That’s valid JSON, but with a timestamp that neither request ever had. Both payloads have the same shape, so the pieces fit together perfectly.</p>\n<p>Now imagine the next request is shorter than the previous one. <code>writeLoop</code> still thinks the body has the old length, so it copies the new data and then keeps going into what’s left of the old data, because <code>buf.Reset()</code> doesn’t clear anything:</p>\n<pre><code>previous : [{&quot;Timestamp&quot;:1,&quot;Content&quot;:&quot;first&quot;},{&quot;Timestamp&quot;:2,&quot;Content&quot;:&quot;second&quot;}]\nnext     : [{&quot;Timestamp&quot;:9,&quot;Content&quot;:&quot;new&quot;}]\nsent     : [{&quot;Timestamp&quot;:9,&quot;Content&quot;:&quot;new&quot;}]},{&quot;Timestamp&quot;:2,&quot;Content&quot;:&quot;second&quot;}]</code></pre>\n<p>This time it’s not even valid JSON.</p>\n<p>That sounds scary. But whether it matters depends on one question: <strong>who, if anyone, reads that mixed copy?</strong></p>\n<h2 id=\"vmagent-a-real-race-we-decided-to-keep\">vmagent: a real race we decided to keep</h2>\n<p>vmagent’s remote write client (<a href=\"https://github.com/VictoriaMetrics/VictoriaMetrics/issues/11507\" rel=\"nofollow ugc noopener\">#11507</a>) is the same pattern at a bigger scale. A worker pulls a block of compressed samples from its queue into a byte slice that it <strong>reuses</strong> on every iteration, and sends it with a fresh reader:</p>\n<pre><code>// app/vmagent/remotewrite/client.go (simplified)\nfunc (c *client) runWorker(readBlock func(dst []byte) ([]byte, bool)) {\n var block []byte\n for {\n  block, ok = readBlock(block[:0]) // &lt;- overwrites the previous block&#39;s bytes\n  ...\n  c.sendBlock(block)\n }\n}\nfunc (c *client) newRequest(url string, body []byte) (*http.Request, error) {\n reqBody := bytes.NewBuffer(body) // &lt;- a fresh reader for every request\n req, err := http.NewRequest(http.MethodPost, url, reqBody)\n ...\n}</code></pre>\n<p>Point vmagent at a remote storage that answers <code>200</code> without reading the body (Vadim used <code>httpbin.org/status/200</code>), push a big import through it, and the race detector reports the queue writing the next block into <code>block</code> while <code>writeLoop</code> is still reading the previous one. It’s exactly the race we just dissected. We even merged a fix for it. Three days later we <a href=\"https://github.com/VictoriaMetrics/VictoriaMetrics/commit/8e9af3911f045ed69ed150393806cbebd37b4cd3\" rel=\"nofollow ugc noopener\">reverted it</a>, and later closed the issue without a fix. Here’s why.</p>\n<h3 id=\"why-we-left-it-alone\">Why we left it alone</h3>\n<p>The key to understanding this is how we build the request body: with <code>bytes.NewReader</code> in our example, or <code>bytes.NewBuffer</code> in vmagent. Both wrap our array with their own length and read position, so that part isn’t shared between requests. The only shared part is the array underneath. And the <a href=\"https://go.dev/ref/mem#restrictions\" rel=\"nofollow ugc noopener\">Go memory model</a> guarantees that reading a byte while it’s being overwritten gives us either the old value or the new one, never some corrupted in-between state. So our readers are safe: the data they send can be mixed, but the program itself won’t break. Nothing crashes, and no other memory gets corrupted.</p>\n<p>So the worst case is some mixed data, and in vmagent nobody uses it. It goes into a request the server has <strong>already answered</strong>, without reading the body, so the server didn’t want it anyway. Usually the connection is being closed too, so nothing on the other side ever reads them.</p>\n<p>With nothing to protect, fixing it anyway would only cost us. A new buffer per request would undo the savings of reusing it. And waiting for the transport to finish with the body, which we tried and then <a href=\"https://github.com/VictoriaMetrics/VictoriaMetrics/commit/8e9af3911f045ed69ed150393806cbebd37b4cd3\" rel=\"nofollow ugc noopener\">reverted</a>, could stall workers, deadlock in rare cases, and still didn’t cover everything.</p>\n<p>Paying in performance and complexity to protect bytes nobody reads isn’t a good trade, so we kept the race.</p>\n<p>“Harmless” is not “free”. This is a known, accepted race-detector report, not a silent one. We’ll revisit it if Go gains an official way to wait for the transport to be done with a request body. In <a href=\"https://github.com/golang/go/issues/81445\" rel=\"nofollow ugc noopener\">the upstream issue</a>, Damien Neil suggested a new <code>Request.Close</code> method for exactly that.</p>\n<h2 id=\"vmauth-when-the-race-is-a-real-bug\">vmauth: when the race is a real bug</h2>\n<p>vmauth hit the same behavior (<a href=\"https://github.com/VictoriaMetrics/VictoriaMetrics/issues/11508\" rel=\"nofollow ugc noopener\">#11508</a>), but here it was a real problem. vmauth is a proxy, and it can <strong>retry</strong>: if a backend fails or answers with an error like <code>503</code>, vmauth sends the same request to the next backend.</p>\n<p>To send the same body twice, vmauth keeps it in memory in a type called <code>bufferedBody</code>. Before the fix, <code>bufferedBody</code> was also the reader: it held the bytes plus its own read position. On a retry, vmauth rewound that position to zero and handed <strong>the same <code>bufferedBody</code></strong> to the next request.</p>\n<p>That’s the key difference from vmagent. In vmagent, each request has its own reader, and only the bytes are shared. In vmauth, both requests shared <strong>the reader itself</strong>, including its read position. So the leftover <code>writeLoop</code> from the failed request could keep moving that position, and when it finished, it even reset it to zero, right in the middle of the retry’s upload.</p>\n<p>The result: the next backend could receive the body with pieces missing or repeated. Sometimes the length came out wrong and the request failed. But sometimes it came out exactly right, and the backend accepted a corrupted request without anyone noticing. This time the mixed data doesn’t go nowhere: it goes to a healthy backend that processes it.</p>\n<h3 id=\"the-fix-correctness-first\">The fix: correctness first</h3>\n<p>The fix (<a href=\"https://github.com/VictoriaMetrics/VictoriaMetrics/pull/11647\" rel=\"nofollow ugc noopener\">#11647</a>) does what vmagent already does: <strong>never give two requests the same reader.</strong> vmauth still keeps the bytes in <code>bufferedBody</code>, but every attempt now gets its own new <code>bytes.Buffer</code> over them. The leftover <code>writeLoop</code> can keep reading its own reader as long as it wants, and it can’t touch the retry’s. In the case of vmauth, the shared bytes are only ever read, never written, so there’s no race at all.</p>\n<p>Here, correctness wins easily: a proxy that can silently send corrupted data to a backend isn’t acceptable at any speed. And the cost is tiny: a couple of small allocations per attempt, with no copy of the data, on a path where the network round trip to the backend costs far more.</p>\n<p>So the same race detector warning led to two opposite decisions: keep it in vmagent, fix it in vmauth. Let’s wrap up with what made the difference.</p>\n<h2 id=\"what-to-take-away\">What to take away</h2>\n<p>Both warnings came from the same <code>net/http</code> behavior: the transport writes the request body in its own goroutine, and it can return the response to you before it’s done reading your body. The race detector was right both times. What it can’t tell you is whether the race matters, and that came down to two questions:</p>\n<ol><li><strong>What exactly is shared?</strong> Plain bytes that get copied somewhere, or<em>state</em> that decides behavior, like an offset, a length or a pointer? A race on bytes gives you stale bytes. A race on a cursor gives you wrong behavior.</li><li><strong>Who consumes the result?</strong> In vmagent, the mixed copy goes into a request the server has already answered, on a connection that’s being closed. In vmauth, the racing cursor decided what a<em>healthy backend</em> received and ingested.</li></ol>\n<p>So in vmagent we accepted the race and documented why, because fixing it would cost performance and complexity to protect bytes nobody reads. In vmauth we fixed it, by removing the sharing rather than adding synchronization, because correctness comes first.</p>\n<p>If you use Go’s HTTP client, the rules are short:</p>\n<ul><li>Assume the transport may still be reading your request body after <code>Do</code> returns.</li><li>Avoid giving the same stateful <code>io.Reader</code> to two requests.</li><li>Be careful with reusing, pooling or mutating the memory behind a body you’ve already sent, until the transport has called <code>Close()</code> on it.</li></ul>\n<p>And keep an eye on <a href=\"https://github.com/golang/go/issues/81445\" rel=\"nofollow ugc noopener\">golang/go#81445</a>: if Go gets an official way to wait until the transport is done with a request, most of this goes away.</p>","headings":[{"level":1,"text":"The Data Race That Wasn't a Bug (and the One That Was)","id":"the-data-race-that-wasn-t-a-bug-and-the-one-that-was"},{"level":2,"text":"The code and the race","id":"the-code-and-the-race"},{"level":2,"text":"How the HTTP client sends a request: three goroutines","id":"how-the-http-client-sends-a-request-three-goroutines"},{"level":2,"text":"Following the bytes: your buffer, the reader, and the scratch buffer","id":"following-the-bytes-your-buffer-the-reader-and-the-scratch-buffe"},{"level":2,"text":"Putting it together: how the race happens","id":"putting-it-together-how-the-race-happens"},{"level":3,"text":"What can actually go wrong","id":"what-can-actually-go-wrong"},{"level":2,"text":"vmagent: a real race we decided to keep","id":"vmagent-a-real-race-we-decided-to-keep"},{"level":3,"text":"Why we left it alone","id":"why-we-left-it-alone"},{"level":2,"text":"vmauth: when the race is a real bug","id":"vmauth-when-the-race-is-a-real-bug"},{"level":3,"text":"The fix: correctness first","id":"the-fix-correctness-first"},{"level":2,"text":"What to take away","id":"what-to-take-away"}]}}