modern-fs-benchmark

Continuous benchmarks for multi-device, copy-on-write filesystems — btrfs, ZFS, bcachefs — measuring the things single-device ext4-style benchmarks (Phoronix et al.) never touch: redundancy layouts, snapshot aging and scaling, transparent compression, encryption (native vs LUKS), reflinks, fsync tail latency, degraded operation and rebuild, corruption self-healing, and near-full/ENOSPC behavior — with ext4/xfs over md/LVM as the classic-stack baselines.

Why

Classic filesystem benchmarks run fio on one device with default mkfs options. That says nothing about what modern filesystems are actually deployed for. This suite benchmarks the machinery:

PhaseWhat it measures
host calibrationfio on the runner's own disk before any filesystem exists — a VM-noise anchor
seq / rand write, rand readbaseline throughput on the chosen redundancy layout
trivial-op latency under load"how long until my prompt comes back": a 4k write+fsync every 200ms (shell history, editor swap), p99 and worst case — idle, then while a 1M streaming writer floods the filesystem; CoW commit storms live here
source-tree opscreate / cold cp -r / rm -rf of a 20k-small-file tree — the "copy a kernel tree" test
large-directory scalabilitycreate 100k empty files in one directory, enumerate names cold, stat every entry cold and warm, then delete — directory indexing and inode-cache behavior that tree-shaped workloads miss (LARGEDIR_FILES=1000000 reproduces the million-file variant)
parallel random readsame cold-cache read with 4 concurrent threads — a mirror can only serve from both copies under concurrency, so this is where replica read-scaling shows (on real hardware; CI loop devices share one disk and physically can't)
fsync tail latencyp99 / p99.9 fdatasync completion latency from the random-write phase — CoW transaction commits (ZFS txg, btrfs commit interval) spike periodically in ways the IOPS average hides
snapshot agingrandom-overwrite bandwidth as snapshots accumulate (CoW fragmentation cost) — 100 snapshots where the technology allows; ZFS at 128K recordsize pins ~the whole file per snapshot so its default-recordsize layouts run 10, and old-style LVM snapshots amplify every origin write per snapshot so lvm layouts run 8 (both caps are findings, not shortcuts)
snapshot createmetadata cost of taking a snapshot
snapshot delete + reclaimdelete latency, foreground write bandwidth while background cleaning runs, time until the space actually returns
compressionzstd ratio + write throughput on 75%-compressible data
reflinkcp --reflink=always of a large file
clone divergencethe unshare penalty: the same 4k-overwrite workload into a plain file, a fresh reflink clone (btrfs/bcachefs/ZFS/xfs), and a freshly-snapshotted file (CoW filesystems and LVM)
degraded + rebuildfail one device: IO while degraded, then time the rebuild onto a spare
snapshot-count scaling500 snapshots with no churn between them: create latency at the tail, snapshot-list time, remount time, bulk delete (native-snapshot filesystems)
near-full / ENOSPCon a fresh small array of the same layout: write throughput near 95% and 99% full, then fill to hard ENOSPC — can you still delete (CoW needs free space to delete), and does deleting make the fs writable again? Caveat: btrfs hits its chunk-allocation wall before df crosses the target on small devices (1G data chunks are a big fraction of a CI-sized array — on multi-TB disks the same wall sits at 99.9%), so its probes run at the wall; the actual fullness at each probe is recorded in the JSON (nearfull*_pct)
corruption + scrubwrite a 2G raw range onto one device behind the filesystem's back, scrub, then compare one tracked test-file hash; this probes checksum/redundancy recovery but is not a whole-filesystem health check, and the overwritten range can include allocated or unused space