The assumption that costs you GPU hours

If you are training robot policies with imitation learning, you are almost certainly using open-source datasets. Bridge V2, Open X-Embodiment, ALOHA, LeRobot community datasets -- these are the foundations that labs and startups build on. The implicit assumption is that these datasets are clean, well-structured, and ready for training. After all, they come from top research groups. They have papers. They have thousands of GitHub stars.

We decided to test that assumption. We ran automated quality checks across 10 popular open-source robotics datasets, examining structural integrity, metadata correctness, and semantic completeness. The checks are not exotic -- they look for things like consistent schemas, valid camera intrinsics, correct action space descriptions, and decodable video frames. The kind of validation you would expect any dataset to pass before release.