Zephon: Fast, Flexible, Elastically Deterministic Data Loading

If you want to know whether a change to your training data made a model better, everything else about the training run has to hold still. In practice, the data loader often doesn't: resume a run on a different number of GPUs, and most loaders will quietly change what each GPU sees, which in our experiments has been enough to muddy comparisons between data curation methods.

At Datology, we run thousands of data experiments, and more and more of them are designed and run by agents, which (much like us humans) won't notice when a result depends on the GPU count.