{"article":{"slug":"nobody-is-listening-on-port-8125","title":"Nobody Is Listening on Port 8125","subtitle":null,"summary":"There's a Fedora box on my desk named `stick`. An app on it has been firing StatsD datagrams at `127.0.0.1:8125` ten times a second for about a week now, the way apps have done since Etsy taught everyone to in 2011.","content_type":"blog_post","language":"en","canonical_url":"https://yeet.cx/blog/nobody-is-listening-on-8125","author":{"name":"Jacob Pradels","url":null,"person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"yeet","url":"https://yeet.cx","listing_slug":null,"listing":null},"topics":[{"name":"Infrastructure","slug":"infrastructure","url":"https://listedarticles.com/topics/infrastructure"},{"name":"Performance","slug":"performance","url":"https://listedarticles.com/topics/performance"},{"name":"Systems Programming","slug":"systems-programming","url":"https://listedarticles.com/topics/systems-programming"},{"name":"Engineering","slug":"engineering","url":"https://listedarticles.com/topics/engineering"},{"name":"Linux","slug":"linux","url":"https://listedarticles.com/topics/linux"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":2281,"reading_minutes":10,"published_at":"2026-09-30T00:00:00.000Z","added_at":"2026-10-03T09:17:31.192Z","updated_at":"2026-10-03T09:17:31.192Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":false},"profile_url":"https://listedarticles.com/articles/nobody-is-listening-on-port-8125","markdown_url":"https://listedarticles.com/articles/nobody-is-listening-on-port-8125.md","example":false,"citation":"Jacob Pradels, yeet. \"Nobody Is Listening on Port 8125.\" 30 Sept 2026. https://yeet.cx/blog/nobody-is-listening-on-8125 (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://yeet.cx/blog/nobody-is-listening-on-8125"},"body_markdown":"There's a Fedora box on my desk named `stick`. An app on it has been firing\nStatsD datagrams at `127.0.0.1:8125` ten times a second for about a week\nnow, the way apps have done since Etsy taught everyone to in 2011.\n\nNothing is listening on 8125.\n\n```\n$ ss -lun | grep 8125\n$\n```\nGrafana has the request rate anyway.\n\n\nThis is the first post in a series where we go down the list of exporters\neveryone runs next to Prometheus and replace them, one at a time, with a\nyeet script. The pitch for the series is short: you do\nnot need fifty exporters in six languages, each with its own port, its own\nconfig format, and its own opinion about vendoring. You need one binary\nthat can read the kernel. We're starting with `statsd_exporter` because it's\nthe one where the replacement stops being the same thing and turns into\nsomething kind of better.\n\n`statsd_exporter` is a translator. Your app pushes StatsD lines at it over\nUDP, it holds them as Prometheus metrics, Prometheus pulls. It's maintained\nby the Prometheus project itself, so this is an official piece of the stack\nwe're pulling out.\n\nWe replaced the listener with an eBPF program that watches port 8125 at the\ntraffic control layer of every interface, and a JavaScript decoder that\nturns the lines into Prometheus families. The app keeps doing fire-and-forget\nUDP. Nothing changes on the app. Nothing gets installed on the host except\n`yeet`. And the metrics show up whether or not anyone is there to receive\nthe datagrams.\n\nYou'll learn:\n\n`yeet:telemetry` turns a\nscript into a Prometheus target without the script ever touching a socket.\nThe whole thing is open source. To run it on your own boxes:\n\nLess than you'd think. It binds a UDP socket, reads datagrams, splits them on newlines, and parses lines that look like this:\n\n```\napp.requests:1|c\napp.queue.depth:22|g\napp.request.time:143|ms\napp.users:20|s\napp.errors:1|c|@0.1|#service:api,region:us\n```\nThen it keeps a counter, gauge, histogram or set per name, runs the names\nthrough a YAML mapping file so your dots turn into labels, and serves the\nwhole thing on `:9102/metrics`. That's it. That's the exporter. It's a\nperfectly good exporter and I've run it for years without thinking about it\nonce, which is about the nicest thing you can say about infrastructure.\n\nBut look at what it *eats*. UDP datagrams. On a well-known port. In a\nplaintext line format. Usually on the same host as the app. By the time\nthat socket gets the bytes, the kernel has already had them, looked at\nthem, and decided where they go. The listener is the **second** reader of\nevery packet.\n\nSo we asked the first one.\n\nThe kernel half is a tcx program clipped onto the ingress and egress hooks of every interface that's up, loopback included. It does three things: find the transport header, check if either port is one we care about, copy a bounded chunk of the payload into a ring buffer. Done.\n\n```\n/* The ports worth capturing, filled by the script. */\nstruct {\n    __uint(type, BPF_MAP_TYPE_HASH);\n    __uint(max_entries, 64);\n    __type(key, __u16);\n    __type(value, __u8);\n} ports SEC(\".maps\");\nstatic __always_inline int wanted(__u16 sport, __u16 dport)\n{\n    return bpf_map_lookup_elem(&ports, &sport) != NULL\n        || bpf_map_lookup_elem(&ports, &dport) != NULL;\n}\nstatic __always_inline int handle(struct __sk_buff *skb, __u8 dir)\n{\n    /* ... walk Ethernet or raw IP, then IPv4 or IPv6, to the L4 header ... */\n    __u16 sport = 0, dport = 0;\n    bpf_skb_load_bytes(skb, l4,     &sport, 2);\n    bpf_skb_load_bytes(skb, l4 + 2, &dport, 2);\n    sport = bpf_ntohs(sport);\n    dport = bpf_ntohs(dport);\n    if (!wanted(sport, dport))\n        return TCX_NEXT;\n    /* ... TCP data offset or the 8-byte UDP header, then the payload ... */\n    struct wire_event *e = bpf_ringbuf_reserve(&events, sizeof(*e), 0);\n    if (!e)\n        return TCX_NEXT;\n    e->ts        = bpf_ktime_get_ns();\n    e->sport     = sport;\n    e->dport     = dport;\n    e->ifindex   = skb->ifindex;\n    e->dir       = dir;\n    e->total_len = plen;\n    e->captured  = cap;\n    if (bpf_skb_load_bytes(skb, poff, e->data, cap) < 0) {\n        bpf_ringbuf_discard(e, 0);\n        return TCX_NEXT;\n    }\n    bpf_ringbuf_submit(e, 0);\n    return TCX_NEXT;\n}\n```\nThat `ifindex` is there for a reason. A packet on loopback goes through\nthe egress hook on `lo`, and then, because the kernel just hands it right\nback to itself, through the ingress hook on `lo`. Same bytes, two events.\nTCP you can dedupe on the sequence number. UDP has no sequence number. So\non loopback the collector keeps the egress copy and drops the other one,\nand every datagram counts exactly once:\n\n```\nif (ev.dir === DIR_INGRESS && ev.ifindex === lo) return;\n```\nLook at the `ports` map too. There's no port number anywhere in the C.\nThe script fills the map at startup from its arguments, and 8125 is contentionally\nwhat statsd used, but there's no reason why it couldn't be anything else.\nBecause the kernel side has no idea what StatsD is, the same mechanism covers anything else that speaks\na line protocol on a known port. It matches a port and copies up to 512\nbytes. Parsing is a JavaScript problem, and JavaScript is cheap. The\ncollector on `stick` already runs the Redis, memcached and Graphite\ndecoders off this one probe, each with its own flag, and those get their\nown posts. The next protocol costs a port number and a decoder.\n\nEverything else stays in the kernel. ACKs, TLS, your Zoom call, none of it\never crosses into userspace. And the program returns `TCX_NEXT` on every\npath, so it never drops or rewrites a packet. It reads. That's the whole\njob.\n\nyeet loads it from a script with\n`yeet:bpf`. The attach target is\nwhatever the system graph says is up:\n\n```\nimport { BpfObject, HashMap, RingBuf } from \"yeet:bpf\";\nconst { data } = await yeet.graph.query(`{ network_interfaces { index is_up } }`);\nconst tcx = { kind: \"tcx\", ifindex: data.network_interfaces.filter((i) => i.is_up).map((i) => i.index) };\nconst control = await new BpfObject({ exe: \"../bpf/bin/wire.bpf.o\", base: import.meta.dirname })\n  .bind(\"events\", { kind: \"ringbuf\", btf_struct: \"wire_event\" })\n  .bind(\"ports\", { kind: \"hash\" })\n  .attach(\"on_ingress\", tcx)\n  .attach(\"on_egress\", tcx)\n  .start();\nconst ports = new HashMap(control, \"ports\");\nfor (const port of [8125, 8126]) await ports.update(port, 1);   // from yeet.args, really\nawait new RingBuf(control, \"events\").subscribe(onEvent, () => dropped.inc());\n```\n`btf_struct: \"wire_event\"` is the entire deserializer. The daemon reads the\nstruct layout out of the object's BTF and hands `onEvent` a plain object\nwith `sport`, `dport`, `data` and friends already on it. Same trick in the\nother direction for the map: `update(8125, 1)` is a `__u16` key and a\n`__u8` value because the C said so. We did not write a byte parser. We're\nnever going to write a byte parser.\n\nHere's the StatsD half with the plumbing cut out:\n\n```\nrequest(ev, bytes) {\n  this.datagrams.inc();\n  for (const line of latin1(bytes).split(\"\\n\")) {\n    const m = /^([^:]+):([^|]+)\\|([a-z]+)(?:\\|@([0-9.]+))?(?:\\|#(.*))?$/.exec(line.trim());\n    if (!m) { this.lines.labels({ status: \"invalid\" }).inc(); continue; }\n    const [, rawName, rawValue, type, rate, tags] = m;\n    const name = rawName.replace(/\\./g, \"_\");\n    const labels = parseTags(tags);          // #service:api,region:us\n    const sample = 1 / (Number(rate) || 1);  // @0.1 counts for ten\n    this.#apply(name, type, rawValue, Number(rawValue), labels, sample);\n  }\n}\n```\n`#apply` is a switch on the type letter. `c` bumps a counter by\n`value * sample`. `g` sets a gauge, or nudges it if the value came with a\nsign. `ms`, `h` and `d` observe a histogram, with milliseconds scaled to\nseconds. `s` throws the value in a set and publishes the size. That's the\nStatsD spec. All of it. It fits on a napkin.\n\nFamilies get created the first time a name shows up, which is the part\n`statsd_exporter` makes you do in YAML ahead of time:\n\n```\n#family(name, type, labels) {\n  const key = `${type}:${name}`;\n  let f = this.#series.get(key);\n  if (!f) {\n    const opts = { help: `StatsD ${type} ${name}, read off the wire.`, labels: labels.names };\n    if (type === \"histogram\") opts.buckets = LATENCY_BUCKETS;\n    f = this.telemetry[type](name, opts);\n    this.#series.set(key, f);\n  }\n  return f;\n}\n```\n`this.telemetry` is a\n`Telemetry` registry. The\nfirst `app.errors` line that shows up with `#service:api,region:us` on it\nlocks in the label set for `app_errors_total`. A later line with different\ntags gets counted under `statsd_exporter_lines_total{status=\"rejected\"}`\ninstead of taking the collector down. Same rule the real exporter has, we\njust didn't make you write it in a config file.\n\nCorrect. The isolate a yeet script runs in cannot bind a port. It also\ncan't `fetch`, which we went on about last time\nand still think is the best thing about it. The only listeners anywhere in\nthis picture belong to the daemon, and the daemon speaks HTTP and\nWebSocket and that's the list.\n\nSo how does Prometheus scrape the thing?\n\nTwo pieces. The collector runs as a `SharedWorker` and calls\n`telemetry.serve()`, which answers snapshot requests over a message port.\nThe scrape endpoint is a separate script, and it's tiny:\n\n```\nimport { render } from \"yeet:telemetry\";\nimport { emit } from \"../../lib/emit.js\";\nimport { pick, upAs } from \"../../lib/snapshot.js\";\n// Everything the wire collector holds that is not another protocol's.\nconst OTHERS = [\"redis_\", \"memcached_\", \"graphite_\"];\nemit(await render({ worker: \"../../lib/collectors/wire.js\" }, ({ worker: w }) => ({\n  ...upAs(w, \"statsd_exporter_up\", \"1 when the wire collector answered.\"),\n  ...Object.fromEntries(Object.entries(w).filter(([k, f]) =>\n    k !== \"yeet_worker_up\" && !OTHERS.some((p) => k.startsWith(p)) && !/^Graphite/.test(f?.help ?? \"\"))),\n})));\n```\n`render({ worker })` connects to the shared worker, asks for its registry,\nand hands the families to a function that reshapes them. This one throws\nout the series the other decoders on the same probe produced and keeps\nthe rest under `statsd_exporter`'s name. What comes out gets printed to the\nconsole as Prometheus text format, and the encoder validates it on the way\nout so you can't ship a malformed document by accident.\n\nThen `yeet service` does the bit where a console becomes an HTTP response:\n\n```\nyeet service unit add exporter-swap/gateway -W \"http://0.0.0.0:9100\"\nyeet service unit add exporter-swap/statsd_exporter -I exporters/statsd_exporter/metrics.js --lazy\nyeet service mount exporter-swap/gateway -L /statsd_exporter/metrics -t statsd_exporter -p console -T per-connection\n```\nA plain GET on `/statsd_exporter/metrics` spawns a fresh isolate, runs the\nscript above, streams its console into the response body, and exits. A\nscrape is a process that lives for a few milliseconds and never has\nnetwork access at any point in its life. Prometheus cannot tell the\ndifference:\n\n```\nscrape_configs:\n  - job_name: statsd\n    metrics_path: /statsd_exporter/metrics\n    static_configs:\n      - targets: [\"stick:9100\"]\n```\n```\n$ curl -s stick:9100/statsd_exporter/metrics\n# HELP statsd_exporter_up 1 when the wire collector answered.\n# TYPE statsd_exporter_up gauge\nstatsd_exporter_up 1\n# HELP statsd_exporter_udp_packets_total StatsD datagrams seen on the wire.\n# TYPE statsd_exporter_udp_packets_total counter\nstatsd_exporter_udp_packets_total 215702\n# HELP app_requests_total StatsD counter app_requests, read off the wire.\n# TYPE app_requests_total counter\napp_requests_total 129798\n# HELP app_errors_total StatsD counter app_errors, read off the wire.\n# TYPE app_errors_total counter\napp_errors_total{service=\"api\",region=\"us\"} 1296520\n# HELP app_queue_depth StatsD gauge app_queue_depth, read off the wire.\n# TYPE app_queue_depth gauge\napp_queue_depth 22\n# HELP app_request_time StatsD histogram app_request_time, read off the wire.\n# TYPE app_request_time histogram\napp_request_time_bucket{le=\"0.0005\"} 0\napp_request_time_bucket{le=\"0.001\"} 0\napp_request_time_bucket{le=\"0.002\"} 462\n...\napp_request_time_sum 19534.094000000434\napp_request_time_count 129174\n```\n`app_errors_total` is ten times `app_requests_total` because the app sends\nit with `@0.1`. The tags turned into labels. The timer turned into a\nhistogram in seconds. Nobody wrote a mapping file.\n\ntl;dr: the datagram was always the metric. We just stopped needing a process to catch it.\n\n\nNothing here is magic, so here's the receipt.\n\n**You see what this host sends.** A real listener sees what arrives. If\nyour app on host A ships StatsD to a collector on host B, this runs on A\nand counts at the source, and Prometheus's `instance` label tells you\nwhich A. I'd argue that's better. It is definitely different.\n\n**Names follow the default mapping.** Dots become underscores and that's\nit. If you lean on `statsd_exporter`'s mapping language to turn\n`app.api.us.requests` into `app_requests{service=\"api\", region=\"us\"}`,\nyou'd write that in the decoder instead.  Not a lengthy addition but\nwas outside the scope of this post. Feel free to submit a PR if you\nwant to add it.\n\n**Counters get `_total`.** The registry appends the suffix the way every\nPrometheus client library does. `statsd_exporter` leaves it off unless you\nask. Fix your dashboards or fix the decoder, whichever annoys you less.\n\n**512 bytes per datagram.** StatsD datagrams are\nsupposed to fit in an MTU and usually clients send batches way smaller, but\n`DATA_MAX` is a `#define` if yours don't.\n\n**It's plaintext.** But so is StatsD. When this series\ngets to the HTTP exporters that stops being true and we'll talk about the\nTLS uprobe then.\n\nThe listener was never the point. It was the thing that got the datagram out of the kernel and into a process that knew how to count. Now the counting happens right next to the kernel: a program the verifier has signed off on copies the line into a ring buffer, a decoder that fits on one screen does the math, and Prometheus pulls the answer out of a sandbox that has no sockets and gets scraped anyway.\n\nNobody is listening on 8125. The numbers are fine.\n\n*Next up: the nginx exporter you never installed, and why it knows more\nthan the one you did.*","body_html":"<p>There&#39;s a Fedora box on my desk named <code>stick</code>. An app on it has been firing\nStatsD datagrams at <code>127.0.0.1:8125</code> ten times a second for about a week\nnow, the way apps have done since Etsy taught everyone to in 2011.</p>\n<p>Nothing is listening on 8125.</p>\n<pre><code>$ ss -lun | grep 8125\n$</code></pre>\n<p>Grafana has the request rate anyway.</p>\n<p>This is the first post in a series where we go down the list of exporters\neveryone runs next to Prometheus and replace them, one at a time, with a\nyeet script. The pitch for the series is short: you do\nnot need fifty exporters in six languages, each with its own port, its own\nconfig format, and its own opinion about vendoring. You need one binary\nthat can read the kernel. We&#39;re starting with <code>statsd_exporter</code> because it&#39;s\nthe one where the replacement stops being the same thing and turns into\nsomething kind of better.</p>\n<p><code>statsd_exporter</code> is a translator. Your app pushes StatsD lines at it over\nUDP, it holds them as Prometheus metrics, Prometheus pulls. It&#39;s maintained\nby the Prometheus project itself, so this is an official piece of the stack\nwe&#39;re pulling out.</p>\n<p>We replaced the listener with an eBPF program that watches port 8125 at the\ntraffic control layer of every interface, and a JavaScript decoder that\nturns the lines into Prometheus families. The app keeps doing fire-and-forget\nUDP. Nothing changes on the app. Nothing gets installed on the host except\n<code>yeet</code>. And the metrics show up whether or not anyone is there to receive\nthe datagrams.</p>\n<p>You&#39;ll learn:</p>\n<p><code>yeet:telemetry</code> turns a\nscript into a Prometheus target without the script ever touching a socket.\nThe whole thing is open source. To run it on your own boxes:</p>\n<p>Less than you&#39;d think. It binds a UDP socket, reads datagrams, splits them on newlines, and parses lines that look like this:</p>\n<pre><code>app.requests:1|c\napp.queue.depth:22|g\napp.request.time:143|ms\napp.users:20|s\napp.errors:1|c|@0.1|#service:api,region:us</code></pre>\n<p>Then it keeps a counter, gauge, histogram or set per name, runs the names\nthrough a YAML mapping file so your dots turn into labels, and serves the\nwhole thing on <code>:9102/metrics</code>. That&#39;s it. That&#39;s the exporter. It&#39;s a\nperfectly good exporter and I&#39;ve run it for years without thinking about it\nonce, which is about the nicest thing you can say about infrastructure.</p>\n<p>But look at what it <em>eats</em>. UDP datagrams. On a well-known port. In a\nplaintext line format. Usually on the same host as the app. By the time\nthat socket gets the bytes, the kernel has already had them, looked at\nthem, and decided where they go. The listener is the <strong>second</strong> reader of\nevery packet.</p>\n<p>So we asked the first one.</p>\n<p>The kernel half is a tcx program clipped onto the ingress and egress hooks of every interface that&#39;s up, loopback included. It does three things: find the transport header, check if either port is one we care about, copy a bounded chunk of the payload into a ring buffer. Done.</p>\n<pre><code>/* The ports worth capturing, filled by the script. */\nstruct {\n    __uint(type, BPF_MAP_TYPE_HASH);\n    __uint(max_entries, 64);\n    __type(key, __u16);\n    __type(value, __u8);\n} ports SEC(&quot;.maps&quot;);\nstatic __always_inline int wanted(__u16 sport, __u16 dport)\n{\n    return bpf_map_lookup_elem(&amp;ports, &amp;sport) != NULL\n        || bpf_map_lookup_elem(&amp;ports, &amp;dport) != NULL;\n}\nstatic __always_inline int handle(struct __sk_buff *skb, __u8 dir)\n{\n    /* ... walk Ethernet or raw IP, then IPv4 or IPv6, to the L4 header ... */\n    __u16 sport = 0, dport = 0;\n    bpf_skb_load_bytes(skb, l4,     &amp;sport, 2);\n    bpf_skb_load_bytes(skb, l4 + 2, &amp;dport, 2);\n    sport = bpf_ntohs(sport);\n    dport = bpf_ntohs(dport);\n    if (!wanted(sport, dport))\n        return TCX_NEXT;\n    /* ... TCP data offset or the 8-byte UDP header, then the payload ... */\n    struct wire_event *e = bpf_ringbuf_reserve(&amp;events, sizeof(*e), 0);\n    if (!e)\n        return TCX_NEXT;\n    e-&gt;ts        = bpf_ktime_get_ns();\n    e-&gt;sport     = sport;\n    e-&gt;dport     = dport;\n    e-&gt;ifindex   = skb-&gt;ifindex;\n    e-&gt;dir       = dir;\n    e-&gt;total_len = plen;\n    e-&gt;captured  = cap;\n    if (bpf_skb_load_bytes(skb, poff, e-&gt;data, cap) &lt; 0) {\n        bpf_ringbuf_discard(e, 0);\n        return TCX_NEXT;\n    }\n    bpf_ringbuf_submit(e, 0);\n    return TCX_NEXT;\n}</code></pre>\n<p>That <code>ifindex</code> is there for a reason. A packet on loopback goes through\nthe egress hook on <code>lo</code>, and then, because the kernel just hands it right\nback to itself, through the ingress hook on <code>lo</code>. Same bytes, two events.\nTCP you can dedupe on the sequence number. UDP has no sequence number. So\non loopback the collector keeps the egress copy and drops the other one,\nand every datagram counts exactly once:</p>\n<pre><code>if (ev.dir === DIR_INGRESS &amp;&amp; ev.ifindex === lo) return;</code></pre>\n<p>Look at the <code>ports</code> map too. There&#39;s no port number anywhere in the C.\nThe script fills the map at startup from its arguments, and 8125 is contentionally\nwhat statsd used, but there&#39;s no reason why it couldn&#39;t be anything else.\nBecause the kernel side has no idea what StatsD is, the same mechanism covers anything else that speaks\na line protocol on a known port. It matches a port and copies up to 512\nbytes. Parsing is a JavaScript problem, and JavaScript is cheap. The\ncollector on <code>stick</code> already runs the Redis, memcached and Graphite\ndecoders off this one probe, each with its own flag, and those get their\nown posts. The next protocol costs a port number and a decoder.</p>\n<p>Everything else stays in the kernel. ACKs, TLS, your Zoom call, none of it\never crosses into userspace. And the program returns <code>TCX_NEXT</code> on every\npath, so it never drops or rewrites a packet. It reads. That&#39;s the whole\njob.</p>\n<p>yeet loads it from a script with\n<code>yeet:bpf</code>. The attach target is\nwhatever the system graph says is up:</p>\n<pre><code>import { BpfObject, HashMap, RingBuf } from &quot;yeet:bpf&quot;;\nconst { data } = await yeet.graph.query(`{ network_interfaces { index is_up } }`);\nconst tcx = { kind: &quot;tcx&quot;, ifindex: data.network_interfaces.filter((i) =&gt; i.is_up).map((i) =&gt; i.index) };\nconst control = await new BpfObject({ exe: &quot;../bpf/bin/wire.bpf.o&quot;, base: import.meta.dirname })\n  .bind(&quot;events&quot;, { kind: &quot;ringbuf&quot;, btf_struct: &quot;wire_event&quot; })\n  .bind(&quot;ports&quot;, { kind: &quot;hash&quot; })\n  .attach(&quot;on_ingress&quot;, tcx)\n  .attach(&quot;on_egress&quot;, tcx)\n  .start();\nconst ports = new HashMap(control, &quot;ports&quot;);\nfor (const port of [8125, 8126]) await ports.update(port, 1);   // from yeet.args, really\nawait new RingBuf(control, &quot;events&quot;).subscribe(onEvent, () =&gt; dropped.inc());</code></pre>\n<p><code>btf_struct: &quot;wire_event&quot;</code> is the entire deserializer. The daemon reads the\nstruct layout out of the object&#39;s BTF and hands <code>onEvent</code> a plain object\nwith <code>sport</code>, <code>dport</code>, <code>data</code> and friends already on it. Same trick in the\nother direction for the map: <code>update(8125, 1)</code> is a <code>__u16</code> key and a\n<code>__u8</code> value because the C said so. We did not write a byte parser. We&#39;re\nnever going to write a byte parser.</p>\n<p>Here&#39;s the StatsD half with the plumbing cut out:</p>\n<pre><code>request(ev, bytes) {\n  this.datagrams.inc();\n  for (const line of latin1(bytes).split(&quot;\\n&quot;)) {\n    const m = /^([^:]+):([^|]+)\\|([a-z]+)(?:\\|@([0-9.]+))?(?:\\|#(.*))?$/.exec(line.trim());\n    if (!m) { this.lines.labels({ status: &quot;invalid&quot; }).inc(); continue; }\n    const [, rawName, rawValue, type, rate, tags] = m;\n    const name = rawName.replace(/\\./g, &quot;_&quot;);\n    const labels = parseTags(tags);          // #service:api,region:us\n    const sample = 1 / (Number(rate) || 1);  // @0.1 counts for ten\n    this.#apply(name, type, rawValue, Number(rawValue), labels, sample);\n  }\n}</code></pre>\n<p><code>#apply</code> is a switch on the type letter. <code>c</code> bumps a counter by\n<code>value * sample</code>. <code>g</code> sets a gauge, or nudges it if the value came with a\nsign. <code>ms</code>, <code>h</code> and <code>d</code> observe a histogram, with milliseconds scaled to\nseconds. <code>s</code> throws the value in a set and publishes the size. That&#39;s the\nStatsD spec. All of it. It fits on a napkin.</p>\n<p>Families get created the first time a name shows up, which is the part\n<code>statsd_exporter</code> makes you do in YAML ahead of time:</p>\n<pre><code>#family(name, type, labels) {\n  const key = `${type}:${name}`;\n  let f = this.#series.get(key);\n  if (!f) {\n    const opts = { help: `StatsD ${type} ${name}, read off the wire.`, labels: labels.names };\n    if (type === &quot;histogram&quot;) opts.buckets = LATENCY_BUCKETS;\n    f = this.telemetry[type](name, opts);\n    this.#series.set(key, f);\n  }\n  return f;\n}</code></pre>\n<p><code>this.telemetry</code> is a\n<code>Telemetry</code> registry. The\nfirst <code>app.errors</code> line that shows up with <code>#service:api,region:us</code> on it\nlocks in the label set for <code>app_errors_total</code>. A later line with different\ntags gets counted under <code>statsd_exporter_lines_total{status=&quot;rejected&quot;}</code>\ninstead of taking the collector down. Same rule the real exporter has, we\njust didn&#39;t make you write it in a config file.</p>\n<p>Correct. The isolate a yeet script runs in cannot bind a port. It also\ncan&#39;t <code>fetch</code>, which we went on about last time\nand still think is the best thing about it. The only listeners anywhere in\nthis picture belong to the daemon, and the daemon speaks HTTP and\nWebSocket and that&#39;s the list.</p>\n<p>So how does Prometheus scrape the thing?</p>\n<p>Two pieces. The collector runs as a <code>SharedWorker</code> and calls\n<code>telemetry.serve()</code>, which answers snapshot requests over a message port.\nThe scrape endpoint is a separate script, and it&#39;s tiny:</p>\n<pre><code>import { render } from &quot;yeet:telemetry&quot;;\nimport { emit } from &quot;../../lib/emit.js&quot;;\nimport { pick, upAs } from &quot;../../lib/snapshot.js&quot;;\n// Everything the wire collector holds that is not another protocol&#39;s.\nconst OTHERS = [&quot;redis_&quot;, &quot;memcached_&quot;, &quot;graphite_&quot;];\nemit(await render({ worker: &quot;../../lib/collectors/wire.js&quot; }, ({ worker: w }) =&gt; ({\n  ...upAs(w, &quot;statsd_exporter_up&quot;, &quot;1 when the wire collector answered.&quot;),\n  ...Object.fromEntries(Object.entries(w).filter(([k, f]) =&gt;\n    k !== &quot;yeet_worker_up&quot; &amp;&amp; !OTHERS.some((p) =&gt; k.startsWith(p)) &amp;&amp; !/^Graphite/.test(f?.help ?? &quot;&quot;))),\n})));</code></pre>\n<p><code>render({ worker })</code> connects to the shared worker, asks for its registry,\nand hands the families to a function that reshapes them. This one throws\nout the series the other decoders on the same probe produced and keeps\nthe rest under <code>statsd_exporter</code>&#39;s name. What comes out gets printed to the\nconsole as Prometheus text format, and the encoder validates it on the way\nout so you can&#39;t ship a malformed document by accident.</p>\n<p>Then <code>yeet service</code> does the bit where a console becomes an HTTP response:</p>\n<pre><code>yeet service unit add exporter-swap/gateway -W &quot;http://0.0.0.0:9100&quot;\nyeet service unit add exporter-swap/statsd_exporter -I exporters/statsd_exporter/metrics.js --lazy\nyeet service mount exporter-swap/gateway -L /statsd_exporter/metrics -t statsd_exporter -p console -T per-connection</code></pre>\n<p>A plain GET on <code>/statsd_exporter/metrics</code> spawns a fresh isolate, runs the\nscript above, streams its console into the response body, and exits. A\nscrape is a process that lives for a few milliseconds and never has\nnetwork access at any point in its life. Prometheus cannot tell the\ndifference:</p>\n<pre><code>scrape_configs:\n  - job_name: statsd\n    metrics_path: /statsd_exporter/metrics\n    static_configs:\n      - targets: [&quot;stick:9100&quot;]</code></pre>\n<pre><code>$ curl -s stick:9100/statsd_exporter/metrics\n# HELP statsd_exporter_up 1 when the wire collector answered.\n# TYPE statsd_exporter_up gauge\nstatsd_exporter_up 1\n# HELP statsd_exporter_udp_packets_total StatsD datagrams seen on the wire.\n# TYPE statsd_exporter_udp_packets_total counter\nstatsd_exporter_udp_packets_total 215702\n# HELP app_requests_total StatsD counter app_requests, read off the wire.\n# TYPE app_requests_total counter\napp_requests_total 129798\n# HELP app_errors_total StatsD counter app_errors, read off the wire.\n# TYPE app_errors_total counter\napp_errors_total{service=&quot;api&quot;,region=&quot;us&quot;} 1296520\n# HELP app_queue_depth StatsD gauge app_queue_depth, read off the wire.\n# TYPE app_queue_depth gauge\napp_queue_depth 22\n# HELP app_request_time StatsD histogram app_request_time, read off the wire.\n# TYPE app_request_time histogram\napp_request_time_bucket{le=&quot;0.0005&quot;} 0\napp_request_time_bucket{le=&quot;0.001&quot;} 0\napp_request_time_bucket{le=&quot;0.002&quot;} 462\n...\napp_request_time_sum 19534.094000000434\napp_request_time_count 129174</code></pre>\n<p><code>app_errors_total</code> is ten times <code>app_requests_total</code> because the app sends\nit with <code>@0.1</code>. The tags turned into labels. The timer turned into a\nhistogram in seconds. Nobody wrote a mapping file.</p>\n<p>tl;dr: the datagram was always the metric. We just stopped needing a process to catch it.</p>\n<p>Nothing here is magic, so here&#39;s the receipt.</p>\n<p><strong>You see what this host sends.</strong> A real listener sees what arrives. If\nyour app on host A ships StatsD to a collector on host B, this runs on A\nand counts at the source, and Prometheus&#39;s <code>instance</code> label tells you\nwhich A. I&#39;d argue that&#39;s better. It is definitely different.</p>\n<p><strong>Names follow the default mapping.</strong> Dots become underscores and that&#39;s\nit. If you lean on <code>statsd_exporter</code>&#39;s mapping language to turn\n<code>app.api.us.requests</code> into <code>app_requests{service=&quot;api&quot;, region=&quot;us&quot;}</code>,\nyou&#39;d write that in the decoder instead.  Not a lengthy addition but\nwas outside the scope of this post. Feel free to submit a PR if you\nwant to add it.</p>\n<p><strong>Counters get <code>_total</code>.</strong> The registry appends the suffix the way every\nPrometheus client library does. <code>statsd_exporter</code> leaves it off unless you\nask. Fix your dashboards or fix the decoder, whichever annoys you less.</p>\n<p><strong>512 bytes per datagram.</strong> StatsD datagrams are\nsupposed to fit in an MTU and usually clients send batches way smaller, but\n<code>DATA_MAX</code> is a <code>#define</code> if yours don&#39;t.</p>\n<p><strong>It&#39;s plaintext.</strong> But so is StatsD. When this series\ngets to the HTTP exporters that stops being true and we&#39;ll talk about the\nTLS uprobe then.</p>\n<p>The listener was never the point. It was the thing that got the datagram out of the kernel and into a process that knew how to count. Now the counting happens right next to the kernel: a program the verifier has signed off on copies the line into a ring buffer, a decoder that fits on one screen does the math, and Prometheus pulls the answer out of a sandbox that has no sockets and gets scraped anyway.</p>\n<p>Nobody is listening on 8125. The numbers are fine.</p>\n<p>*Next up: the nginx exporter you never installed, and why it knows more\nthan the one you did.*</p>","headings":[]}}