{"article":{"slug":"modern-fs-benchmark","title":"modern-fs-benchmark","subtitle":null,"summary":"Bartosz Fenski’s continuous benchmark suite for multi-device CoW filesystems (btrfs, ZFS, bcachefs) measures snapshot aging, compression, rebuild, ENOSPC, and other workloads classic single-disk fio tests miss.","content_type":"research","language":"en","canonical_url":"https://bartosz.fenski.pl/modern-fs-benchmark/","author":{"name":"Bartosz Fenski","url":"https://bartosz.fenski.pl/","person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"Bartosz Fenski","url":"https://bartosz.fenski.pl/","listing_slug":null,"listing":null},"topics":[{"name":"Systems Programming","slug":"systems-programming","url":"https://listedarticles.com/topics/systems-programming"},{"name":"Benchmarks","slug":"benchmarks","url":"https://listedarticles.com/topics/benchmarks"},{"name":"Open Source","slug":"open-source","url":"https://listedarticles.com/topics/open-source"},{"name":"Infrastructure","slug":"infrastructure","url":"https://listedarticles.com/topics/infrastructure"},{"name":"Linux","slug":"linux","url":"https://listedarticles.com/topics/linux"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":3195,"reading_minutes":14,"published_at":"2026-09-20T06:17:08.979Z","added_at":"2026-09-20T06:17:08.979Z","updated_at":"2026-09-20T06:17:08.979Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":true},"profile_url":"https://listedarticles.com/articles/modern-fs-benchmark","markdown_url":"https://listedarticles.com/articles/modern-fs-benchmark.md","example":false,"citation":"Bartosz Fenski, Bartosz Fenski. \"modern-fs-benchmark.\" 20 Sept 2026. https://bartosz.fenski.pl/modern-fs-benchmark/ (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://bartosz.fenski.pl/modern-fs-benchmark/"},"body_markdown":"# modern-fs-benchmark\n\nContinuous benchmarks for **multi-device, copy-on-write filesystems** — btrfs,\nZFS, bcachefs — measuring the things single-device ext4-style benchmarks\n(Phoronix et al.) never touch: redundancy layouts, snapshot aging and scaling,\ntransparent compression, encryption (native vs LUKS), reflinks, fsync tail\nlatency, degraded operation and rebuild, corruption self-healing, and\nnear-full/ENOSPC behavior — with ext4/xfs over md/LVM as the classic-stack\nbaselines.\n\n## Why\n\nClassic filesystem benchmarks run fio on one device with default mkfs options.\nThat says nothing about what modern filesystems are actually deployed for.\nThis suite benchmarks the *machinery*:\n\n| Phase | What it measures |\n|---|---|\n| host calibration | fio on the runner's own disk *before* any filesystem exists — a VM-noise anchor |\n| seq / rand write, rand read | baseline throughput on the chosen redundancy layout |\n| trivial-op latency under load | \"how long until my prompt comes back\": a 4k write+fsync every 200ms (shell history, editor swap), p99 and worst case — idle, then while a 1M streaming writer floods the filesystem; CoW commit storms live here |\n| source-tree ops | create / cold `cp -r` / `rm -rf` of a 20k-small-file tree — the \"copy a kernel tree\" test |\n| large-directory scalability | create 100k empty files in one directory, enumerate names cold, stat every entry cold and warm, then delete — directory indexing and inode-cache behavior that tree-shaped workloads miss (`LARGEDIR_FILES=1000000` reproduces the [million-file variant](https://paste.sr.ht/~arya_elfren/31e435822ca401cdf4c64de8d13c45f56973ec0f)) |\n| parallel random read | same cold-cache read with 4 concurrent threads — a mirror can only serve from both copies under concurrency, so this is where replica read-scaling shows (on real hardware; CI loop devices share one disk and physically can't) |\n| fsync tail latency | p99 / p99.9 fdatasync completion latency from the random-write phase — CoW transaction commits (ZFS txg, btrfs commit interval) spike periodically in ways the IOPS average hides |\n| snapshot aging | random-overwrite bandwidth as snapshots accumulate (CoW fragmentation cost) — **100 snapshots** where the technology allows; ZFS at 128K recordsize pins ~the whole file per snapshot so its default-recordsize layouts run 10, and old-style LVM snapshots amplify every origin write per snapshot so lvm layouts run 8 (both caps are findings, not shortcuts) |\n| snapshot create | metadata cost of taking a snapshot |\n| snapshot delete + reclaim | delete latency, foreground write bandwidth while background cleaning runs, time until the space actually returns |\n| compression | zstd ratio + write throughput on 75%-compressible data |\n| reflink | `cp --reflink=always` of a large file |\n| clone divergence | the unshare penalty: the same 4k-overwrite workload into a plain file, a fresh reflink clone (btrfs/bcachefs/ZFS/xfs), and a freshly-snapshotted file (CoW filesystems and LVM) |\n| degraded + rebuild | fail one device: IO while degraded, then time the rebuild onto a spare |\n| snapshot-count scaling | 500 snapshots with no churn between them: create latency at the tail, snapshot-list time, remount time, bulk delete (native-snapshot filesystems) |\n| near-full / ENOSPC | on a fresh small array of the same layout: write throughput near 95% and 99% full, then fill to hard ENOSPC — can you still delete (CoW needs free space to delete), and does deleting make the fs writable again? Caveat: btrfs hits its chunk-allocation wall *before* df crosses the target on small devices (1G data chunks are a big fraction of a CI-sized array — on multi-TB disks the same wall sits at 99.9%), so its probes run at the wall; the actual fullness at each probe is recorded in the JSON (`nearfull*_pct`) |\n| corruption + scrub | write a 2G raw range onto one device behind the filesystem's back, scrub, then compare one tracked test-file hash; this probes checksum/redundancy recovery but is not a whole-filesystem health check, and the overwritten range can include allocated or unused space |\n\nResults are published as a dashboard: **<https://bartosz.fenski.pl/modern-fs-benchmark/>**\n— per-metric charts sorted best-first, aging curves, and trends across runs,\nfilterable by filesystem family and layout class (e.g. \"btrfs vs bcachefs,\nmulti-device only\"), with linear/log scale switching and a sortable table.\nRun history lives on the `results-data` branch.\n\nThe suite uses the established `fio` tool and conventional MiB/s, IOPS, and\nlatency units, but the exact workload recipes and composite summary index are\nproject-specific rather than an industry-standard benchmark. It does not use\n`O_DIRECT`: documented cold-read phases drop caches before normal buffered\nreads, while other phases use normal buffered I/O with explicit durability\nbarriers where noted.\n\nEvery result records the exact tools *and kernel-module* versions tested —\nessential for ZFS and bcachefs, which are out-of-tree, where the kernel\nversion alone doesn't identify what actually ran. Shown in the dashboard\ntable, stored in the JSON.\n\nDefault matrix — 26 configurations (4 devices, plus baselines; the\nauthoritative list is the matrix in `.github/workflows/bench.yml`):\n\n- **ext4 single** — one device, the \"what does any of this cost\" anchor\n- **ext4 on md raid10** — the classic layered stack\n- **ext4 on LVM raid10** — layered stack with block-layer CoW snapshots,\n  so the snapshot-aging phase is comparable with the native-CoW filesystems\n- **xfs single / on md raid10 / on LVM raid10** — the same three stacks again;\n  XFS additionally has reflink, unlike ext4\n- **btrfs / bcachefs / ZFS single-device** — the CoW filesystems without\n  redundancy, head-to-head with ext4/xfs single: the pure cost (and features)\n  of CoW itself. btrfs uses `-m single`: mkfs defaults to DUP metadata on a\n  single device, which would double its metadata writes vs every other\n  single-device row (community catch)\n- **Encryption variants** — ZFS native per-dataset AES-256-GCM\n  (`mirror-enc`), bcachefs native whole-fs ChaCha20/Poly1305\n  (`replicas2-enc`), btrfs over one LUKS layer *per device*\n  (`raid1-luks` — no native option, so every replica is encrypted\n  separately), and ext4 over a single LUKS layer on top of md\n  (`md-raid10-luks` — the classic stack encrypts once, above the raid).\n  Compression runs on all of them, so encrypt-after-compress vs\n  opaque-blocks falls out of the existing zstd phase\n- **btrfs** — `-d raid1 -m raid1`\n- **Single-parity** — zfs `raidz1`, plus `raidz1-enc` with native encryption\n  on top (community request)\n- **Dual-parity (raid6-class)** — zfs `raidz2` (and `raidz2-enc`,\n  community request), btrfs `-d raid6 -m raid1c3`\n  (parity metadata is discouraged — write hole), ext4 on md raid6, and\n  bcachefs `--erasure_code --replicas=3` (stable since 1.37; write-hole-free\n  by design — writes replicate first, background reconcile stripes them)\n  (community request, incl. the correction that EC is no longer experimental)\n- **xfs on a ZFS zvol** — the Franken-stack people actually run: XFS\n  semantics on top; ZFS snapshots (fsfreeze-consistent), self-healing,\n  and compression underneath (community request)\n- **xfs on LVM raid10 + dm-integrity** (`--raidintegrity y`) — per-sector\n  checksums give the classic stack detection AND correction: the fairest\n  classic-vs-CoW comparison in the corruption phase, with the performance\n  tax quantified (community request)\n- **ZFS** — striped mirror pairs (raid10-like), at the default 128K recordsize\n  and again at `recordsize=8k` — one-variable proof of how much of ZFS's\n  small-random-write cost is configuration, not design\n- **bcachefs** — `--replicas=2` (kernel module built via DKMS from\n  [apt.bcachefs.org](https://apt.bcachefs.org/) since bcachefs left mainline in 6.17)\n\n  If you want to try bcachefs as a working storage system rather than just\n  benchmark it, [NASty](https://github.com/nasty-project/nasty) is a NixOS-based\n  NAS appliance built around it and a practical place to start.\n\n## The point is data integrity, not the winner's podium\n\nBenchmark charts invite \"which is fastest\". For long-term storage that is\nthe wrong question — the right one is **which stack tells you the truth\nabout your data**, and it's why this suite exists (the corruption phase\nre-proves it every couple of hours):\n\n- **ext4/xfs on md or LVM raid — the default \"safe\" Linux setup — has no\n  data checksums.** Raid protects against a *missing* disk, not a *lying*\n  one: when a copy goes bad (disk firmware, cable, controller, bad RAM,\n  power cut mid-write), the array cannot tell which copy is right. In our\n  corruption test these stacks have returned garbage to the application\n  **with no error whatsoever** — reads succeed, exit codes are 0, and the\n  damage can propagate into backups silently. A scrub *counts* mismatches;\n  it cannot say which side is correct. An intact hash in one run is not proof\n  of repair: read balancing may simply select the good copy.\n- **btrfs, ZFS, and bcachefs verify every read against checksums** and,\n  given appropriate redundancy, can reconstruct a bad allocated block from a\n  good replica or parity. Published redundant-layout runs have kept the tracked\n  test file readable and hash-identical, including native-encryption and parity\n  layouts. Nonzero found counts show detected damage, while nonzero repaired\n  counts provide the strongest evidence of reconstruction. Zero or unavailable\n  counts do not prove that the raw overwrite hit the tracked file.\n- **The classic stack *can* buy the same guarantee** — LVM raid with\n  `--raidintegrity y` (dm-integrity) is the first classic layout to pass\n  our corruption phase — but almost nobody runs it, and the performance\n  tax is measurable (that's the `xfs/lvm-raid10-int` row).\n\nSpeed matters and we measure it honestly. But if you keep data you care\nabout — photos, archives, the family's one copy of anything — on a\nnon-checksumming stack, no benchmark number compensates for corruption you\nwon't discover until years later. That risk is invisible in every classic\nfilesystem benchmark; here it's a first-class result\n(*corruption + scrub* on the dashboard).\n\nThe dashboard's corruption badge is deliberately scoped. `SURVIVED` means the\nsingle tracked `read.dat` file was still readable and had the same MD5 after\nscrub; `FAIL` means it differed or could not be read; `UNPROVEN` means the hash\nmatched on a stack without end-to-end data checksums. The test does not hash the\nentire filesystem, prove every overwritten byte was allocated, or perform a\npost-corruption remount for every backend.\n\n## How it runs\n\n### CI (GitHub Actions, loop devices)\n\nEvery push/2-hourly cron builds each filesystem across 4 loop devices backed by\nsparse files, runs the suite, and publishes a results table in the job summary\nplus JSON artifacts. Each job's artifact also contains a **full command trace**\n(`raw/<config>-trace.log`) — every command executed, arguments fully expanded,\nwith source file and line — so \"what exactly was run\" is never a question.\n(`BENCH_TRACE=1` mirrors it into the live log instead.)\n\n**Interpret CI numbers carefully.** Runners are shared VMs and all \"devices\"\nlive on one virtual disk, so absolute MiB/s is meaningless and RAID striping\ngains are fiction. Matrix jobs also run in parallel, **each on its own\nephemeral VM** — so comparing filesystem A against filesystem B compares two\ndifferent machines. Mitigations, from strongest signal to weakest:\n\n1. *Within-job* ratios and shapes (aging curve slope, compression on/off,\n   degraded vs healthy) — same VM, same disk, directly meaningful.\n2. Every job runs a **host calibration** first (fio on the runner's disk,\n   before any filesystem exists). Jobs on VMs below the calibration floor\n   (`CALIB_MIN_*`, ~25% of runners' normal disk speed margin) **fail fast\n   and are automatically rerun on a fresh runner** (up to 3 attempts) —\n   junk numbers from an unlucky VM never enter the results.\n3. Cross-filesystem deltas within one run — treat small differences (tens of\n   percent) as noise; large ones (2×+) are usually real.\n4. Trends over repeated runs (2-hourly cron + every push) average the VM\n   lottery out — this is where cross-filesystem conclusions belong.\n\n### Real hardware\n\nThe same scripts take real block devices — this is where absolute numbers\nbecome valid:\n\n```sh\nsudo BENCH_DEVICES=\"/dev/sdb /dev/sdc /dev/sdd /dev/sde\" BENCH_WIPE=1 \\\n  scripts/run-bench.sh btrfs raid1\n```\n\nSafety: devices must be unmounted, and anything carrying a filesystem\nsignature is refused unless `BENCH_WIPE=1`. **Listed devices are wiped.**\n\nFor an unmanaged hardware run, invoke `scripts/run-bench.sh` directly with\n`BENCH_DEVICES`, `BENCH_SPARE_DEVICE`, and `BENCH_WIPE=1`. The dedicated\nself-hosted GitHub workflow instead requires the NixOS module below; it never\naccepts device paths from workflow inputs. Workload sizes can be adjusted with\n`SEQ_SIZE`, `AGING_SIZE`, `AGING_ITERS`, and the other documented environment\nvariables. CI defaults are sized for 4×16 GB loop files.\n\n#### NixOS / deploy-rs\n\nThis repository is also a flake with a reusable NixOS module and benchmark\npackage. Cluster configurations can import `nixosModules.modern-fs-benchmark`;\nthe module installs a dedicated Actions\nrunner, the filesystem tools and matching out-of-tree modules for the\ncluster-selected kernel, and a restricted root wrapper with fixed device paths.\nIt deliberately does not select a kernel or configure machine-wide boot,\nnetworking, users, or partitioning.\n\n```nix\n{\n  inputs.modern-fs-benchmark.url =\n    \"github:fenio/modern-fs-benchmark\";\n\n  # In the target node's modules list:\n  services.modern-fs-benchmark = {\n    enable = true;\n    repository = \"https://github.com/fenio/modern-fs-benchmark\";\n    tokenFile = \"/run/secrets/modern-fs-benchmark-runner\";\n    runnerName = \"farm3\";\n    hardwareProfile = \"farm3\";\n    runnerLabels = [ \"fs-benchmark\" ];\n    devices = [\n      \"/dev/disk/by-partlabel/fsbench-nvme0-a\"\n      \"/dev/disk/by-partlabel/fsbench-nvme1-a\"\n      \"/dev/disk/by-partlabel/fsbench-nvme0-b\"\n      \"/dev/disk/by-partlabel/fsbench-nvme1-b\"\n    ];\n    spareDevice = \"/dev/disk/by-partlabel/fsbench-nvme0-spare\";\n    zfsSingleDevice = \"/dev/disk/by-partlabel/fsbench-nvme0-zfs-single\";\n  };\n}\n```\n\nThe four member devices and spare must each be exactly 16 GiB. The dedicated\n`zfsSingleDevice` must be exactly 32 GiB, matching the hosted-runner matrix.\nFor an unregistered manual run, the same immutable package is available as\n`nix run .#manual -- <fs> <layout>`; provide the documented `BENCH_*`\nenvironment variables and run it as root.\nSet `BENCH_HARDWARE_RANDOM_SCALING=1` to include the optional 8- and 16-worker\nrandom read/write measurements and the 4/8/16-worker shard-aware write series\nthat the managed hardware wrapper enables.\n\nThe master cluster flake owns the node assignment and deploy-rs deployment, so\nthe runner can move to another machine without changing benchmark code. The\ndedicated `bench-real-hw.yml` workflow targets the `fs-benchmark` label and\nuses only the module's fixed devices. It publishes the farm3 history to\n`results-real-hw` and the dashboard under `/real-hw/`; the existing `bench.yml`\nworkflow remains hosted-only and continues publishing `results-data` at the\nroot dashboard. Hardware runs can be dispatched manually. The weekly schedule\nis enabled only when the repository variable `ENABLE_HARDWARE_BENCHMARKS` is\nset to `true`. The token file should contain a fine-grained PAT because\nephemeral runners re-register after every job.\n\nThe rotational `sas-hdd` profile is routed independently through the\n`fs-benchmark-sas-hdd` runner label. Its workflow, enable variable, history,\nand dashboard are respectively `bench-real-hw-sas-hdd.yml`,\n`ENABLE_SAS_HDD_BENCHMARKS`, `results-real-hw-sas-hdd`, and\n`/sas-hdd/`. Every result carries `hardware_profile: \"sas-hdd\"`, and\npublication rejects a missing or mismatched profile so results from different\nmachines cannot enter the same trend series.\n\nThe separate [`hybrid-tier-v1`](docs/sas-hdd-hybrid-tier.md) scenario compares\nBtrfs over mirrored `dm-cache`, ZFS special/L2ARC classes, and native bcachefs\nforeground/background/promote targets. It publishes independently under\n`/sas-hdd/hybrid-tier/` and never expands the default hosted matrix.\n\n**The plan is bigger than loop devices.** CI is the regression-tracking\nharness; the goal is to gather dedicated hardware and run the REAL tests\nthere — including the tiered topologies these filesystems were built for and\nthat no publication benchmarks today: NVMe cache/metadata in front of\nrotational data disks (bcachefs foreground/background targets, ZFS\nspecial/log/cache vdevs, LVM dm-cache with writeback and writethrough),\nmixed-rotational RAID, and how each setup behaves degraded and while\nrebuilding. Same suite, same JSON, same dashboard — only the device lists and\ntopology descriptions change.\n\n#### Supporting real-hardware runs\n\nRegular mixed-media runs depend on access to a machine whose benchmark devices\nmay be wiped. A suitable dedicated server currently costs approximately\n**€70 per month**, depending on availability and its exact disk configuration.\nIf recurring sponsorship covers that cost, the server specification,\nconfiguration, raw results, and command traces will all remain public.\n\nI am also open to hardware support in other forms:\n\n- remote root access to an isolated Linux machine with clearly identified,\n  wipeable HDD, SSD, or NVMe devices;\n- donated disks, components, or a complete server for a self-hosted runner;\n- hosting for donated hardware.\n\nIf a suitable remote server does not become available, I may eventually build\nand host a runner myself as the budget allows. Financial support is available\nthrough [GitHub Sponsors](https://github.com/sponsors/fenio) and\n[Ko-fi](https://ko-fi.com/fenio). To offer hardware or discuss a topology,\n[open an issue](https://github.com/fenio/modern-fs-benchmark/issues). Sponsors\nand hardware providers can be acknowledged if they wish, but do not receive\neditorial control over the methodology or results.\n\n### Locally (Linux, loop devices)\n\n```sh\nsudo scripts/install-deps.sh btrfs\nsudo scripts/run-bench.sh btrfs raid1\nscripts/summarize.sh results/result-*.json\n```\n\n## Layout\n\n```\nscripts/run-bench.sh       orchestrates the phases, emits results/result-<fs>-<layout>.json\nscripts/lib/common.sh      device layer (loop files or BENCH_DEVICES), LUKS helpers,\n                           corruption injection, fio helpers, default fs hooks\nscripts/lib/layered.sh     shared md/LVM assembly, snapshots, degrade/repair (ext4 + xfs)\nscripts/fs/<fs>.sh         per-filesystem backend\nscripts/install-deps.sh    Debian/Ubuntu package setup per filesystem\nscripts/summarize.sh       JSON results → markdown table (job summaries)\nscripts/make-dashboard.py  results history → the static dashboard page\nscripts/audit-results.py   anomaly scan over the results history — impossible\n                           orderings, self-healing failures, ENOSPC regressions,\n                           unexpected nulls (daily via the results-audit workflow)\nscripts/result-schema.json machine-readable result keys, types, capabilities, and display metadata\nscripts/result_schema.py   shared result schema loading and validation\nscripts/validate-result.py validates result JSON against that contract\n```\n\nResult documents carry `schema_version`; historical unversioned documents are\ntreated as version 1 so new metrics do not invalidate the stored history.\n\nAdding a filesystem = one file in `scripts/fs/` implementing `fs_setup`,\n`fs_snapshot`, and `fs_teardown`; everything else (`fs_setup_compression`,\n`fs_compress_ratio`, `fs_snapshot_delete_all`, `fs_remount`, `fs_snap_list`,\n`fs_snapscale_delete`, `fs_degrade`, `fs_rebuild`, `fs_scrub`, `fs_version`,\n`fs_drop_caches`, `fs_free_bytes`) has safe defaults in `lib/common.sh` and is\noptional — unimplemented hooks simply record null for their metrics.\n\n## Ideas, hints, and requests welcome\n\nThis suite is deliberately open-ended — if you have opinions on **what to\ntest and how**, please open an issue or PR:\n\n- workloads that would expose behavior the current phases miss\n  (databases, VM images, send/receive, metadata-heavy trees, …)\n- extra configurations and tuning you want measured: mount options,\n  recordsize/extent knobs, compression algorithms and levels, RAID\n  profiles, SLOG/special vdevs, `nodatacow`, …\n- fairness problems in the methodology — if a filesystem is being\n  measured in a way that misrepresents it, that's a bug here\n- additional filesystems or layered stacks (a backend is one small file\n  in `scripts/fs/`)\n\nTuned variants sit next to the defaults in the same matrix (see\n`zfs mirror-8k`), so every suggestion becomes a directly comparable row.\n\n## Roadmap\n\nCoW-specific phases (the behaviors nothing mainstream benchmarks):\n\n- [ ] **send/receive**: full + incremental stream throughput (btrfs, ZFS);\n      rsync over the classic stack as the contrast; bcachefs: not available\n- [ ] **Partial device loss**: device disappears briefly and returns — md\n      write-intent bitmaps, ZFS delta resilver, bcachefs journal catch-up;\n      a different (and common) recovery scenario than full-device rebuild\n- [ ] **lvm-thin layouts**: our lvm rows use old-style snapshots, which are\n      the known-bad strawman — thin pools are what modern LVM users run, with\n      proper CoW snapshots that should survive the aging and snapshot-scaling\n      phases at full count. Next up.\n- [ ] **\"The Tower\"**: ext4/xfs on lvm-thin on dm-vdo on LUKS on raid on\n      dm-integrity — the full feature-parity classic stack (checksums +\n      redundancy + encryption + compression/dedup + CoW snapshots from five\n      dm layers), versus the integrated filesystems that do it in one.\n      Requested with a 😜 but taken seriously: dm-vdo is mainline since 6.9\n      and thin-on-VDO is a documented configuration.\n- [ ] **Stratis** (XFS on dm-thin/dm-integrity/dm-crypt, managed) — arguably\n      the closest classic-stack analogue to btrfs (community suggestion)\n- [ ] **NVMe cache tiers** (community request): bcachefs\n      foreground/background targets, ZFS special/log/cache vdevs, LVM\n      dm-cache — needs real mixed hardware; on CI loop devices both \"tiers\"\n      are the same cloud SSD, so an honest cache benchmark is impossible\n      there (see *Real hardware* above)\n- [ ] **ext4 fscrypt** variant (directory-level encryption — the third model\n      next to native and block-layer)\n\nInfrastructure:\n\n- [ ] Kernel matrix: boot mainline kernels in qemu (runners support nested KVM)\n      and track behavioral regressions per kernel release\n- [ ] Device add/remove/rebalance timing\n- [ ] btrfs/raid1-luks degraded phase (loop-detach can't fail a dm-crypt\n      mapper — needs the dm-error wrapper trick the lvm layouts use)\n- [ ] Parse bcachefs scrub found/repaired counts (verdict via md5 works;\n      the counts aren't in 1.38-tools output)\n- [ ] Normalize cross-job comparisons by the calibration anchor in the dashboard\n\n## License\n\nCopyright 2026 Bartosz Fenski.\n\nSource code, configuration, workflows, and documentation are licensed under the\n[Apache License 2.0](LICENSE). Published benchmark result datasets, including\nthe `results-data`, `results-real-hw`, and `results-real-hw-sas-hdd` history\nbranches, are licensed under\n[Creative Commons Attribution 4.0 International](LICENSE-DATA).","body_html":"<h1 id=\"modern-fs-benchmark\">modern-fs-benchmark</h1>\n<p>Continuous benchmarks for <strong>multi-device, copy-on-write filesystems</strong> — btrfs,\nZFS, bcachefs — measuring the things single-device ext4-style benchmarks\n(Phoronix et al.) never touch: redundancy layouts, snapshot aging and scaling,\ntransparent compression, encryption (native vs LUKS), reflinks, fsync tail\nlatency, degraded operation and rebuild, corruption self-healing, and\nnear-full/ENOSPC behavior — with ext4/xfs over md/LVM as the classic-stack\nbaselines.</p>\n<h2 id=\"why\">Why</h2>\n<p>Classic filesystem benchmarks run fio on one device with default mkfs options.\nThat says nothing about what modern filesystems are actually deployed for.\nThis suite benchmarks the <em>machinery</em>:</p>\n<div class=\"table-wrap\"><table><thead><tr><th>Phase</th><th>What it measures</th></tr></thead><tbody><tr><td>host calibration</td><td>fio on the runner&#39;s own disk <em>before</em> any filesystem exists — a VM-noise anchor</td></tr><tr><td>seq / rand write, rand read</td><td>baseline throughput on the chosen redundancy layout</td></tr><tr><td>trivial-op latency under load</td><td>&quot;how long until my prompt comes back&quot;: a 4k write+fsync every 200ms (shell history, editor swap), p99 and worst case — idle, then while a 1M streaming writer floods the filesystem; CoW commit storms live here</td></tr><tr><td>source-tree ops</td><td>create / cold <code>cp -r</code> / <code>rm -rf</code> of a 20k-small-file tree — the &quot;copy a kernel tree&quot; test</td></tr><tr><td>large-directory scalability</td><td>create 100k empty files in one directory, enumerate names cold, stat every entry cold and warm, then delete — directory indexing and inode-cache behavior that tree-shaped workloads miss (<code>LARGEDIR_FILES=1000000</code> reproduces the <a href=\"https://paste.sr.ht/~arya_elfren/31e435822ca401cdf4c64de8d13c45f56973ec0f\" rel=\"nofollow ugc noopener\">million-file variant</a>)</td></tr><tr><td>parallel random read</td><td>same cold-cache read with 4 concurrent threads — a mirror can only serve from both copies under concurrency, so this is where replica read-scaling shows (on real hardware; CI loop devices share one disk and physically can&#39;t)</td></tr><tr><td>fsync tail latency</td><td>p99 / p99.9 fdatasync completion latency from the random-write phase — CoW transaction commits (ZFS txg, btrfs commit interval) spike periodically in ways the IOPS average hides</td></tr><tr><td>snapshot aging</td><td>random-overwrite bandwidth as snapshots accumulate (CoW fragmentation cost) — <strong>100 snapshots</strong> where the technology allows; ZFS at 128K recordsize pins ~the whole file per snapshot so its default-recordsize layouts run 10, and old-style LVM snapshots amplify every origin write per snapshot so lvm layouts run 8 (both caps are findings, not shortcuts)</td></tr><tr><td>snapshot create</td><td>metadata cost of taking a snapshot</td></tr><tr><td>snapshot delete + reclaim</td><td>delete latency, foreground write bandwidth while background cleaning runs, time until the space actually returns</td></tr><tr><td>compression</td><td>zstd ratio + write throughput on 75%-compressible data</td></tr><tr><td>reflink</td><td><code>cp --reflink=always</code> of a large file</td></tr><tr><td>clone divergence</td><td>the unshare penalty: the same 4k-overwrite workload into a plain file, a fresh reflink clone (btrfs/bcachefs/ZFS/xfs), and a freshly-snapshotted file (CoW filesystems and LVM)</td></tr><tr><td>degraded + rebuild</td><td>fail one device: IO while degraded, then time the rebuild onto a spare</td></tr><tr><td>snapshot-count scaling</td><td>500 snapshots with no churn between them: create latency at the tail, snapshot-list time, remount time, bulk delete (native-snapshot filesystems)</td></tr><tr><td>near-full / ENOSPC</td><td>on a fresh small array of the same layout: write throughput near 95% and 99% full, then fill to hard ENOSPC — can you still delete (CoW needs free space to delete), and does deleting make the fs writable again? Caveat: btrfs hits its chunk-allocation wall <em>before</em> df crosses the target on small devices (1G data chunks are a big fraction of a CI-sized array — on multi-TB disks the same wall sits at 99.9%), so its probes run at the wall; the actual fullness at each probe is recorded in the JSON (<code>nearfull*_pct</code>)</td></tr><tr><td>corruption + scrub</td><td>write a 2G raw range onto one device behind the filesystem&#39;s back, scrub, then compare one tracked test-file hash; this probes checksum/redundancy recovery but is not a whole-filesystem health check, and the overwritten range can include allocated or unused space</td></tr></tbody></table></div>\n<p>Results are published as a dashboard: <strong><a href=\"https://bartosz.fenski.pl/modern-fs-benchmark/\" rel=\"nofollow ugc noopener\">https://bartosz.fenski.pl/modern-fs-benchmark/</a></strong>\n— per-metric charts sorted best-first, aging curves, and trends across runs,\nfilterable by filesystem family and layout class (e.g. &quot;btrfs vs bcachefs,\nmulti-device only&quot;), with linear/log scale switching and a sortable table.\nRun history lives on the <code>results-data</code> branch.</p>\n<p>The suite uses the established <code>fio</code> tool and conventional MiB/s, IOPS, and\nlatency units, but the exact workload recipes and composite summary index are\nproject-specific rather than an industry-standard benchmark. It does not use\n<code>O_DIRECT</code>: documented cold-read phases drop caches before normal buffered\nreads, while other phases use normal buffered I/O with explicit durability\nbarriers where noted.</p>\n<p>Every result records the exact tools <em>and kernel-module</em> versions tested —\nessential for ZFS and bcachefs, which are out-of-tree, where the kernel\nversion alone doesn&#39;t identify what actually ran. Shown in the dashboard\ntable, stored in the JSON.</p>\n<p>Default matrix — 26 configurations (4 devices, plus baselines; the\nauthoritative list is the matrix in <code>.github/workflows/bench.yml</code>):</p>\n<ul><li><strong>ext4 single</strong> — one device, the &quot;what does any of this cost&quot; anchor</li><li><strong>ext4 on md raid10</strong> — the classic layered stack</li><li><p><strong>ext4 on LVM raid10</strong> — layered stack with block-layer CoW snapshots,</p><p>so the snapshot-aging phase is comparable with the native-CoW filesystems</p></li><li><p><strong>xfs single / on md raid10 / on LVM raid10</strong> — the same three stacks again;</p><p>XFS additionally has reflink, unlike ext4</p></li><li><p><strong>btrfs / bcachefs / ZFS single-device</strong> — the CoW filesystems without</p><p>redundancy, head-to-head with ext4/xfs single: the pure cost (and features)\nof CoW itself. btrfs uses <code>-m single</code>: mkfs defaults to DUP metadata on a\nsingle device, which would double its metadata writes vs every other\nsingle-device row (community catch)</p></li><li><p><strong>Encryption variants</strong> — ZFS native per-dataset AES-256-GCM</p><p>(<code>mirror-enc</code>), bcachefs native whole-fs ChaCha20/Poly1305\n(<code>replicas2-enc</code>), btrfs over one LUKS layer <em>per device</em>\n(<code>raid1-luks</code> — no native option, so every replica is encrypted\nseparately), and ext4 over a single LUKS layer on top of md\n(<code>md-raid10-luks</code> — the classic stack encrypts once, above the raid).\nCompression runs on all of them, so encrypt-after-compress vs\nopaque-blocks falls out of the existing zstd phase</p></li><li><strong>btrfs</strong> — <code>-d raid1 -m raid1</code></li><li><p><strong>Single-parity</strong> — zfs <code>raidz1</code>, plus <code>raidz1-enc</code> with native encryption</p><p>on top (community request)</p></li><li><p><strong>Dual-parity (raid6-class)</strong> — zfs <code>raidz2</code> (and <code>raidz2-enc</code>,</p><p>community request), btrfs <code>-d raid6 -m raid1c3</code>\n(parity metadata is discouraged — write hole), ext4 on md raid6, and\nbcachefs <code>--erasure_code --replicas=3</code> (stable since 1.37; write-hole-free\nby design — writes replicate first, background reconcile stripes them)\n(community request, incl. the correction that EC is no longer experimental)</p></li><li><p><strong>xfs on a ZFS zvol</strong> — the Franken-stack people actually run: XFS</p><p>semantics on top; ZFS snapshots (fsfreeze-consistent), self-healing,\nand compression underneath (community request)</p></li><li><p><strong>xfs on LVM raid10 + dm-integrity</strong> (<code>--raidintegrity y</code>) — per-sector</p><p>checksums give the classic stack detection AND correction: the fairest\nclassic-vs-CoW comparison in the corruption phase, with the performance\ntax quantified (community request)</p></li><li><p><strong>ZFS</strong> — striped mirror pairs (raid10-like), at the default 128K recordsize</p><p>and again at <code>recordsize=8k</code> — one-variable proof of how much of ZFS&#39;s\nsmall-random-write cost is configuration, not design</p></li><li><p><strong>bcachefs</strong> — <code>--replicas=2</code> (kernel module built via DKMS from</p><p><a href=\"https://apt.bcachefs.org/\" rel=\"nofollow ugc noopener\">apt.bcachefs.org</a> since bcachefs left mainline in 6.17)</p>\n<p>If you want to try bcachefs as a working storage system rather than just\nbenchmark it, <a href=\"https://github.com/nasty-project/nasty\" rel=\"nofollow ugc noopener\">NASty</a> is a NixOS-based\nNAS appliance built around it and a practical place to start.</p></li></ul>\n<h2 id=\"the-point-is-data-integrity-not-the-winner-s-podium\">The point is data integrity, not the winner&#39;s podium</h2>\n<p>Benchmark charts invite &quot;which is fastest&quot;. For long-term storage that is\nthe wrong question — the right one is <strong>which stack tells you the truth\nabout your data</strong>, and it&#39;s why this suite exists (the corruption phase\nre-proves it every couple of hours):</p>\n<ul><li><p>**ext4/xfs on md or LVM raid — the default &quot;safe&quot; Linux setup — has no</p><p>data checksums.** Raid protects against a <em>missing</em> disk, not a <em>lying</em>\none: when a copy goes bad (disk firmware, cable, controller, bad RAM,\npower cut mid-write), the array cannot tell which copy is right. In our\ncorruption test these stacks have returned garbage to the application\n<strong>with no error whatsoever</strong> — reads succeed, exit codes are 0, and the\ndamage can propagate into backups silently. A scrub <em>counts</em> mismatches;\nit cannot say which side is correct. An intact hash in one run is not proof\nof repair: read balancing may simply select the good copy.</p></li><li><p><strong>btrfs, ZFS, and bcachefs verify every read against checksums</strong> and,</p><p>given appropriate redundancy, can reconstruct a bad allocated block from a\ngood replica or parity. Published redundant-layout runs have kept the tracked\ntest file readable and hash-identical, including native-encryption and parity\nlayouts. Nonzero found counts show detected damage, while nonzero repaired\ncounts provide the strongest evidence of reconstruction. Zero or unavailable\ncounts do not prove that the raw overwrite hit the tracked file.</p></li><li><p><strong>The classic stack <em>can</em> buy the same guarantee</strong> — LVM raid with</p><p><code>--raidintegrity y</code> (dm-integrity) is the first classic layout to pass\nour corruption phase — but almost nobody runs it, and the performance\ntax is measurable (that&#39;s the <code>xfs/lvm-raid10-int</code> row).</p></li></ul>\n<p>Speed matters and we measure it honestly. But if you keep data you care\nabout — photos, archives, the family&#39;s one copy of anything — on a\nnon-checksumming stack, no benchmark number compensates for corruption you\nwon&#39;t discover until years later. That risk is invisible in every classic\nfilesystem benchmark; here it&#39;s a first-class result\n(<em>corruption + scrub</em> on the dashboard).</p>\n<p>The dashboard&#39;s corruption badge is deliberately scoped. <code>SURVIVED</code> means the\nsingle tracked <code>read.dat</code> file was still readable and had the same MD5 after\nscrub; <code>FAIL</code> means it differed or could not be read; <code>UNPROVEN</code> means the hash\nmatched on a stack without end-to-end data checksums. The test does not hash the\nentire filesystem, prove every overwritten byte was allocated, or perform a\npost-corruption remount for every backend.</p>\n<h2 id=\"how-it-runs\">How it runs</h2>\n<h3 id=\"ci-github-actions-loop-devices\">CI (GitHub Actions, loop devices)</h3>\n<p>Every push/2-hourly cron builds each filesystem across 4 loop devices backed by\nsparse files, runs the suite, and publishes a results table in the job summary\nplus JSON artifacts. Each job&#39;s artifact also contains a <strong>full command trace</strong>\n(<code>raw/&lt;config&gt;-trace.log</code>) — every command executed, arguments fully expanded,\nwith source file and line — so &quot;what exactly was run&quot; is never a question.\n(<code>BENCH_TRACE=1</code> mirrors it into the live log instead.)</p>\n<p><strong>Interpret CI numbers carefully.</strong> Runners are shared VMs and all &quot;devices&quot;\nlive on one virtual disk, so absolute MiB/s is meaningless and RAID striping\ngains are fiction. Matrix jobs also run in parallel, <strong>each on its own\nephemeral VM</strong> — so comparing filesystem A against filesystem B compares two\ndifferent machines. Mitigations, from strongest signal to weakest:</p>\n<ol><li><p><em>Within-job</em> ratios and shapes (aging curve slope, compression on/off,</p><p> degraded vs healthy) — same VM, same disk, directly meaningful.</p></li><li><p>Every job runs a <strong>host calibration</strong> first (fio on the runner&#39;s disk,</p><p> before any filesystem exists). Jobs on VMs below the calibration floor\n (<code>CALIB_MIN_*</code>, ~25% of runners&#39; normal disk speed margin) <strong>fail fast\n and are automatically rerun on a fresh runner</strong> (up to 3 attempts) —\n junk numbers from an unlucky VM never enter the results.</p></li><li><p>Cross-filesystem deltas within one run — treat small differences (tens of</p><p> percent) as noise; large ones (2×+) are usually real.</p></li><li><p>Trends over repeated runs (2-hourly cron + every push) average the VM</p><p> lottery out — this is where cross-filesystem conclusions belong.</p></li></ol>\n<h3 id=\"real-hardware\">Real hardware</h3>\n<p>The same scripts take real block devices — this is where absolute numbers\nbecome valid:</p>\n<pre><code class=\"language-sh\">sudo BENCH_DEVICES=&quot;/dev/sdb /dev/sdc /dev/sdd /dev/sde&quot; BENCH_WIPE=1 \\\n  scripts/run-bench.sh btrfs raid1</code></pre>\n<p>Safety: devices must be unmounted, and anything carrying a filesystem\nsignature is refused unless <code>BENCH_WIPE=1</code>. <strong>Listed devices are wiped.</strong></p>\n<p>For an unmanaged hardware run, invoke <code>scripts/run-bench.sh</code> directly with\n<code>BENCH_DEVICES</code>, <code>BENCH_SPARE_DEVICE</code>, and <code>BENCH_WIPE=1</code>. The dedicated\nself-hosted GitHub workflow instead requires the NixOS module below; it never\naccepts device paths from workflow inputs. Workload sizes can be adjusted with\n<code>SEQ_SIZE</code>, <code>AGING_SIZE</code>, <code>AGING_ITERS</code>, and the other documented environment\nvariables. CI defaults are sized for 4×16 GB loop files.</p>\n<h4 id=\"nixos-deploy-rs\">NixOS / deploy-rs</h4>\n<p>This repository is also a flake with a reusable NixOS module and benchmark\npackage. Cluster configurations can import <code>nixosModules.modern-fs-benchmark</code>;\nthe module installs a dedicated Actions\nrunner, the filesystem tools and matching out-of-tree modules for the\ncluster-selected kernel, and a restricted root wrapper with fixed device paths.\nIt deliberately does not select a kernel or configure machine-wide boot,\nnetworking, users, or partitioning.</p>\n<pre><code class=\"language-nix\">{\n  inputs.modern-fs-benchmark.url =\n    &quot;github:fenio/modern-fs-benchmark&quot;;\n\n  # In the target node&#39;s modules list:\n  services.modern-fs-benchmark = {\n    enable = true;\n    repository = &quot;https://github.com/fenio/modern-fs-benchmark&quot;;\n    tokenFile = &quot;/run/secrets/modern-fs-benchmark-runner&quot;;\n    runnerName = &quot;farm3&quot;;\n    hardwareProfile = &quot;farm3&quot;;\n    runnerLabels = [ &quot;fs-benchmark&quot; ];\n    devices = [\n      &quot;/dev/disk/by-partlabel/fsbench-nvme0-a&quot;\n      &quot;/dev/disk/by-partlabel/fsbench-nvme1-a&quot;\n      &quot;/dev/disk/by-partlabel/fsbench-nvme0-b&quot;\n      &quot;/dev/disk/by-partlabel/fsbench-nvme1-b&quot;\n    ];\n    spareDevice = &quot;/dev/disk/by-partlabel/fsbench-nvme0-spare&quot;;\n    zfsSingleDevice = &quot;/dev/disk/by-partlabel/fsbench-nvme0-zfs-single&quot;;\n  };\n}</code></pre>\n<p>The four member devices and spare must each be exactly 16 GiB. The dedicated\n<code>zfsSingleDevice</code> must be exactly 32 GiB, matching the hosted-runner matrix.\nFor an unregistered manual run, the same immutable package is available as\n<code>nix run .#manual -- &lt;fs&gt; &lt;layout&gt;</code>; provide the documented <code>BENCH_*</code>\nenvironment variables and run it as root.\nSet <code>BENCH_HARDWARE_RANDOM_SCALING=1</code> to include the optional 8- and 16-worker\nrandom read/write measurements and the 4/8/16-worker shard-aware write series\nthat the managed hardware wrapper enables.</p>\n<p>The master cluster flake owns the node assignment and deploy-rs deployment, so\nthe runner can move to another machine without changing benchmark code. The\ndedicated <code>bench-real-hw.yml</code> workflow targets the <code>fs-benchmark</code> label and\nuses only the module&#39;s fixed devices. It publishes the farm3 history to\n<code>results-real-hw</code> and the dashboard under <code>/real-hw/</code>; the existing <code>bench.yml</code>\nworkflow remains hosted-only and continues publishing <code>results-data</code> at the\nroot dashboard. Hardware runs can be dispatched manually. The weekly schedule\nis enabled only when the repository variable <code>ENABLE_HARDWARE_BENCHMARKS</code> is\nset to <code>true</code>. The token file should contain a fine-grained PAT because\nephemeral runners re-register after every job.</p>\n<p>The rotational <code>sas-hdd</code> profile is routed independently through the\n<code>fs-benchmark-sas-hdd</code> runner label. Its workflow, enable variable, history,\nand dashboard are respectively <code>bench-real-hw-sas-hdd.yml</code>,\n<code>ENABLE_SAS_HDD_BENCHMARKS</code>, <code>results-real-hw-sas-hdd</code>, and\n<code>/sas-hdd/</code>. Every result carries <code>hardware_profile: &quot;sas-hdd&quot;</code>, and\npublication rejects a missing or mismatched profile so results from different\nmachines cannot enter the same trend series.</p>\n<p>The separate <code>hybrid-tier-v1</code> scenario compares\nBtrfs over mirrored <code>dm-cache</code>, ZFS special/L2ARC classes, and native bcachefs\nforeground/background/promote targets. It publishes independently under\n<code>/sas-hdd/hybrid-tier/</code> and never expands the default hosted matrix.</p>\n<p><strong>The plan is bigger than loop devices.</strong> CI is the regression-tracking\nharness; the goal is to gather dedicated hardware and run the REAL tests\nthere — including the tiered topologies these filesystems were built for and\nthat no publication benchmarks today: NVMe cache/metadata in front of\nrotational data disks (bcachefs foreground/background targets, ZFS\nspecial/log/cache vdevs, LVM dm-cache with writeback and writethrough),\nmixed-rotational RAID, and how each setup behaves degraded and while\nrebuilding. Same suite, same JSON, same dashboard — only the device lists and\ntopology descriptions change.</p>\n<h4 id=\"supporting-real-hardware-runs\">Supporting real-hardware runs</h4>\n<p>Regular mixed-media runs depend on access to a machine whose benchmark devices\nmay be wiped. A suitable dedicated server currently costs approximately\n<strong>€70 per month</strong>, depending on availability and its exact disk configuration.\nIf recurring sponsorship covers that cost, the server specification,\nconfiguration, raw results, and command traces will all remain public.</p>\n<p>I am also open to hardware support in other forms:</p>\n<ul><li><p>remote root access to an isolated Linux machine with clearly identified,</p><p>wipeable HDD, SSD, or NVMe devices;</p></li><li>donated disks, components, or a complete server for a self-hosted runner;</li><li>hosting for donated hardware.</li></ul>\n<p>If a suitable remote server does not become available, I may eventually build\nand host a runner myself as the budget allows. Financial support is available\nthrough <a href=\"https://github.com/sponsors/fenio\" rel=\"nofollow ugc noopener\">GitHub Sponsors</a> and\n<a href=\"https://ko-fi.com/fenio\" rel=\"nofollow ugc noopener\">Ko-fi</a>. To offer hardware or discuss a topology,\n<a href=\"https://github.com/fenio/modern-fs-benchmark/issues\" rel=\"nofollow ugc noopener\">open an issue</a>. Sponsors\nand hardware providers can be acknowledged if they wish, but do not receive\neditorial control over the methodology or results.</p>\n<h3 id=\"locally-linux-loop-devices\">Locally (Linux, loop devices)</h3>\n<pre><code class=\"language-sh\">sudo scripts/install-deps.sh btrfs\nsudo scripts/run-bench.sh btrfs raid1\nscripts/summarize.sh results/result-*.json</code></pre>\n<h2 id=\"layout\">Layout</h2>\n<pre><code>scripts/run-bench.sh       orchestrates the phases, emits results/result-&lt;fs&gt;-&lt;layout&gt;.json\nscripts/lib/common.sh      device layer (loop files or BENCH_DEVICES), LUKS helpers,\n                           corruption injection, fio helpers, default fs hooks\nscripts/lib/layered.sh     shared md/LVM assembly, snapshots, degrade/repair (ext4 + xfs)\nscripts/fs/&lt;fs&gt;.sh         per-filesystem backend\nscripts/install-deps.sh    Debian/Ubuntu package setup per filesystem\nscripts/summarize.sh       JSON results → markdown table (job summaries)\nscripts/make-dashboard.py  results history → the static dashboard page\nscripts/audit-results.py   anomaly scan over the results history — impossible\n                           orderings, self-healing failures, ENOSPC regressions,\n                           unexpected nulls (daily via the results-audit workflow)\nscripts/result-schema.json machine-readable result keys, types, capabilities, and display metadata\nscripts/result_schema.py   shared result schema loading and validation\nscripts/validate-result.py validates result JSON against that contract</code></pre>\n<p>Result documents carry <code>schema_version</code>; historical unversioned documents are\ntreated as version 1 so new metrics do not invalidate the stored history.</p>\n<p>Adding a filesystem = one file in <code>scripts/fs/</code> implementing <code>fs_setup</code>,\n<code>fs_snapshot</code>, and <code>fs_teardown</code>; everything else (<code>fs_setup_compression</code>,\n<code>fs_compress_ratio</code>, <code>fs_snapshot_delete_all</code>, <code>fs_remount</code>, <code>fs_snap_list</code>,\n<code>fs_snapscale_delete</code>, <code>fs_degrade</code>, <code>fs_rebuild</code>, <code>fs_scrub</code>, <code>fs_version</code>,\n<code>fs_drop_caches</code>, <code>fs_free_bytes</code>) has safe defaults in <code>lib/common.sh</code> and is\noptional — unimplemented hooks simply record null for their metrics.</p>\n<h2 id=\"ideas-hints-and-requests-welcome\">Ideas, hints, and requests welcome</h2>\n<p>This suite is deliberately open-ended — if you have opinions on <strong>what to\ntest and how</strong>, please open an issue or PR:</p>\n<ul><li><p>workloads that would expose behavior the current phases miss</p><p>(databases, VM images, send/receive, metadata-heavy trees, …)</p></li><li><p>extra configurations and tuning you want measured: mount options,</p><p>recordsize/extent knobs, compression algorithms and levels, RAID\nprofiles, SLOG/special vdevs, <code>nodatacow</code>, …</p></li><li><p>fairness problems in the methodology — if a filesystem is being</p><p>measured in a way that misrepresents it, that&#39;s a bug here</p></li><li><p>additional filesystems or layered stacks (a backend is one small file</p><p>in <code>scripts/fs/</code>)</p></li></ul>\n<p>Tuned variants sit next to the defaults in the same matrix (see\n<code>zfs mirror-8k</code>), so every suggestion becomes a directly comparable row.</p>\n<h2 id=\"roadmap\">Roadmap</h2>\n<p>CoW-specific phases (the behaviors nothing mainstream benchmarks):</p>\n<ul><li><p><input type=\"checkbox\" disabled /> <strong>send/receive</strong>: full + incremental stream throughput (btrfs, ZFS);</p><pre><code>rsync over the classic stack as the contrast; bcachefs: not available</code></pre></li><li><p><input type=\"checkbox\" disabled /> <strong>Partial device loss</strong>: device disappears briefly and returns — md</p><pre><code>write-intent bitmaps, ZFS delta resilver, bcachefs journal catch-up;\na different (and common) recovery scenario than full-device rebuild</code></pre></li><li><p><input type=\"checkbox\" disabled /> <strong>lvm-thin layouts</strong>: our lvm rows use old-style snapshots, which are</p><pre><code>the known-bad strawman — thin pools are what modern LVM users run, with\nproper CoW snapshots that should survive the aging and snapshot-scaling\nphases at full count. Next up.</code></pre></li><li><p><input type=\"checkbox\" disabled /> <strong>&quot;The Tower&quot;</strong>: ext4/xfs on lvm-thin on dm-vdo on LUKS on raid on</p><pre><code>dm-integrity — the full feature-parity classic stack (checksums +\nredundancy + encryption + compression/dedup + CoW snapshots from five\ndm layers), versus the integrated filesystems that do it in one.\nRequested with a 😜 but taken seriously: dm-vdo is mainline since 6.9\nand thin-on-VDO is a documented configuration.</code></pre></li><li><p><input type=\"checkbox\" disabled /> <strong>Stratis</strong> (XFS on dm-thin/dm-integrity/dm-crypt, managed) — arguably</p><pre><code>the closest classic-stack analogue to btrfs (community suggestion)</code></pre></li><li><p><input type=\"checkbox\" disabled /> <strong>NVMe cache tiers</strong> (community request): bcachefs</p><pre><code>foreground/background targets, ZFS special/log/cache vdevs, LVM\ndm-cache — needs real mixed hardware; on CI loop devices both &quot;tiers&quot;\nare the same cloud SSD, so an honest cache benchmark is impossible\nthere (see *Real hardware* above)</code></pre></li><li><p><input type=\"checkbox\" disabled /> <strong>ext4 fscrypt</strong> variant (directory-level encryption — the third model</p><pre><code>next to native and block-layer)</code></pre></li></ul>\n<p>Infrastructure:</p>\n<ul><li><p><input type=\"checkbox\" disabled /> Kernel matrix: boot mainline kernels in qemu (runners support nested KVM)</p><pre><code>and track behavioral regressions per kernel release</code></pre></li><li><input type=\"checkbox\" disabled /> Device add/remove/rebalance timing</li><li><p><input type=\"checkbox\" disabled /> btrfs/raid1-luks degraded phase (loop-detach can&#39;t fail a dm-crypt</p><pre><code>mapper — needs the dm-error wrapper trick the lvm layouts use)</code></pre></li><li><p><input type=\"checkbox\" disabled /> Parse bcachefs scrub found/repaired counts (verdict via md5 works;</p><pre><code>the counts aren&#39;t in 1.38-tools output)</code></pre></li><li><input type=\"checkbox\" disabled /> Normalize cross-job comparisons by the calibration anchor in the dashboard</li></ul>\n<h2 id=\"license\">License</h2>\n<p>Copyright 2026 Bartosz Fenski.</p>\n<p>Source code, configuration, workflows, and documentation are licensed under the\nApache License 2.0. Published benchmark result datasets, including\nthe <code>results-data</code>, <code>results-real-hw</code>, and <code>results-real-hw-sas-hdd</code> history\nbranches, are licensed under\nCreative Commons Attribution 4.0 International.</p>","headings":[{"level":1,"text":"modern-fs-benchmark","id":"modern-fs-benchmark"},{"level":2,"text":"Why","id":"why"},{"level":2,"text":"The point is data integrity, not the winner's podium","id":"the-point-is-data-integrity-not-the-winner-s-podium"},{"level":2,"text":"How it runs","id":"how-it-runs"},{"level":3,"text":"CI (GitHub Actions, loop devices)","id":"ci-github-actions-loop-devices"},{"level":3,"text":"Real hardware","id":"real-hardware"},{"level":3,"text":"Locally (Linux, loop devices)","id":"locally-linux-loop-devices"},{"level":2,"text":"Layout","id":"layout"},{"level":2,"text":"Ideas, hints, and requests welcome","id":"ideas-hints-and-requests-welcome"},{"level":2,"text":"Roadmap","id":"roadmap"},{"level":2,"text":"License","id":"license"}]}}