We are building celld, a runtime for Cloudflare Workers and Durable Objects applications on your own machines. It can run with an S3-compatible object store as its only external service dependency.

celld is a distributed system. Making distributed systems reliable is hard, partly because we can unintentionally rely on assumptions that don’t hold in practice, even when we know the pitfalls. Peter Deutsch’s The Eight Fallacies of Distributed Computing lists eight such assumptions, including “The network is reliable” and “Latency is zero.”

A bug can depend on a particular sequence of delayed messages, failed writes, and node restarts. Those events can happen in a different order on the next test run, making the failure difficult to reproduce. We need to be able to repeat the failing run so we can investigate the cause and check whether a proposed fix really resolves the problem.