---
title: "Deterministic simulation testing in celld: exploring execution orders and reproducing bugs"
slug: deterministic-simulation-testing-in-celld-exploring-execution-orders-and-reproducing-bugs
url: https://listedarticles.com/articles/deterministic-simulation-testing-in-celld-exploring-execution-orders-and-reproducing-bugs
canonical_url: https://celld.dev/docs/engineering/deterministic-simulation-testing/
content_type: blog_post
language: en
published_at: 2026-10-02T00:00:00.000Z
updated_at: 2026-10-11T05:06:05.364Z
author: "Yusuke Tanaka"
authored_by: human
publisher: "celld"
publisher_url: https://celld.dev/
topics: ["Distributed Systems", "Testing", "Software Engineering"]
about: ["https://listedstartups.com/companies/deno"]
license: all-rights-reserved
word_count: 1426
reading_minutes: 6
citation: "Yusuke Tanaka, celld. \"Deterministic simulation testing in celld: exploring execution orders and reproducing bugs.\" 2 Oct 2026. https://celld.dev/docs/engineering/deterministic-simulation-testing/ (all-rights-reserved)"
# The full text follows. The web page shows an extract and sends readers
# to the source above; quote the citation and link the canonical URL.
---

# Deterministic simulation testing in celld: exploring execution orders and reproducing bugs

> Yusuke Tanaka explains how the Deno team uses deterministic simulation testing for celld, a self-hostable runtime for Cloudflare Workers and Durable Objects apps, to explore message orderings, storage failures and node restarts, replay failing runs exactly, and turn discovered bugs into regression tests.

We are building [celld](https://github.com/denoland/celld/blob/12d5b6333fe52717325addcfe1e99e9fd4f77bcd/README.md), a runtime for Cloudflare Workers and Durable Objects applications on your own machines. It can run with an S3-compatible object store as its only external service dependency.

celld is a distributed system. Making distributed systems reliable is hard, partly because we can unintentionally rely on assumptions that don’t hold in practice, even when we know the pitfalls. Peter Deutsch’s [The Eight Fallacies of Distributed Computing](https://nighthacks.com/jag/res/Fallacies.html) lists eight such assumptions, including “The network is reliable” and “Latency is zero.”

A bug can depend on a particular sequence of delayed messages, failed writes, and node restarts. Those events can happen in a different order on the next test run, making the failure difficult to reproduce. We need to be able to repeat the failing run so we can investigate the cause and check whether a proposed fix really resolves the problem.

That is why we use **deterministic simulation testing (DST)**. Our simulator is still under development and is not included in celld’s public repository, but it has already found previously unknown bugs. In this article, we’ll walk through how DST works in celld and how it helped us find, reproduce, and fix one of those bugs.

## How DST works in celld

DST runs celld’s production code in an environment controlled by a simulator.

A [**cell**](https://github.com/denoland/celld/blob/12d5b6333fe52717325addcfe1e99e9fd4f77bcd/README.md) in celld runs application code and has its own SQLite database. As cells do their work, celld handles events such as incoming requests, completed storage operations, and timer firings. The code that selects the next event is separate from the code that handles it. This lets the simulator control the order of events while running the [same event-handling code as in production](https://github.com/denoland/celld/blob/12d5b6333fe52717325addcfe1e99e9fd4f77bcd/crates/celld/actor.rs#L2461-L2491).

The simulator also controls when [asynchronous tasks run](https://github.com/denoland/celld/blob/12d5b6333fe52717325addcfe1e99e9fd4f77bcd/crates/celld/lib.rs#L309-L320) and how [object storage](https://github.com/denoland/celld/blob/12d5b6333fe52717325addcfe1e99e9fd4f77bcd/crates/celld/bucket.rs#L533-L551) responds. It can delay a write or make it fail. It can advance simulated time without waiting for real time to pass. We use this control to test how time-dependent behavior, such as retries and alarms, interacts with other events.

We configure which requests and faults the simulator can explore. Within those limits, it uses random choices to generate requests, select what runs next, decide whether storage operations succeed or fail, and choose which faults occur and when.

A **seed** is the starting value for its random-number generator. With the same code, settings, and seed, we get the same sequence, intermediate states, and result. We can explore different runs by changing the seed, then repeat a failing run to trace where things went wrong.

After each action, a checker uses observed responses and stored data to test **invariants**: conditions we define that the system must uphold throughout a run. The test initializes celld with the simulated environment, then explores actions up to a configured limit. In pseudocode:

```
const simulation = createSimulation({ settings, seed });

for (let step = 0; step < settings.maxActions; step++) {
  const candidates = simulation.availableActions();
  if (candidates.length === 0) {
    checkCompletion(simulation);
    break;
  }

  const action = simulation.choose(candidates);
  simulation.run(action);
  checkInvariants(simulation);
}
```

Here, an action is something the simulator can advance: delivering a request, letting a task run, advancing the simulated clock, or completing a storage operation. If no actions are available, the test checks its completion conditions, such as whether all planned requests have finished.

The simplified replay below starts after Cell A has inserted a row into SQLite. Watch how the data is saved to object storage before celld receives confirmation. Other requests or background tasks can run in that gap, while celld is still waiting to return success to the client. Bugs can emerge when these operations interleave in unexpected ways. The simulator controls these steps separately so we can explore different execution orders and replay those that expose a bug.

Press **Next** to step through choosing an action, animating it, and revealing its result.

**Candidates**

* Send the SQLite changes; let object storage receive them.

→

**Chosen**

1. 1Choose
2. 2Run
3. 3Result

Starting state

The row is in SQLite. The client is still waiting for success.

←Next→

Step **0** / 12

The replay above shows a successful request. Next, we’ll look at an execution order that exposed a bug in celld.

## An alarm race the simulator found

An alarm schedules a cell to run application code at a specified time. The cell need not stay in memory until then: celld can [unload an idle cell](https://github.com/denoland/celld/blob/12d5b6333fe52717325addcfe1e99e9fd4f77bcd/crates/logic/lib.rs#L5694-L5716) to make room for others and load it again in time to run its alarm.

The alarm's scheduled time is [stored in the cell's SQLite database](https://github.com/denoland/celld/blob/12d5b6333fe52717325addcfe1e99e9fd4f77bcd/crates/celld/storage.rs#L466-L472). If the cell has been unloaded, finding its alarm from SQLite alone would mean opening its database again. Doing that for every cell would be expensive, so celld instead scans [**wake entries**](https://github.com/denoland/celld/blob/12d5b6333fe52717325addcfe1e99e9fd4f77bcd/crates/celld/wake.rs#L33-L89) in object storage. Each entry identifies a cell and when to reactivate it. Once the cell is active, celld reads the scheduled time from SQLite to determine when to run the alarm.

When we found this bug, celld reused wake entries to reduce the number of writes to object storage. For example, an application could delete a 10:00 alarm and then set a new one for 10:05. If the 10:00 wake entry was still present, celld could reuse it to reactivate the cell at 10:00, early enough for the new alarm. Once the cell was active, celld would read 10:05 from SQLite and wait until then to run the alarm.1

Deleting the 10:00 alarm clears its scheduled time from SQLite. Once that change has been saved to object storage, celld returns success to the client. A separate cleanup task removes the 10:00 wake entry later, so the client does not have to wait for that extra storage operation. The client can therefore set the 10:05 alarm while the old cleanup is still waiting to run.

After celld confirms that the 10:05 alarm is set, a wake entry must remain that can reactivate the cell by 10:05. The simulator checks that this remains true until the alarm runs or is canceled.

The figure below shows how this invariant is violated when the old cleanup runs after the 10:05 alarm has been set.

**Candidates**

* Delete Cell A's 10:00 alarm.

→

**Chosen**

1. 1Choose
2. 2Run
3. 3Result
4. 4Check

Starting state

Cell A has a 10:00 alarm in SQLite and a 10:00 wake entry in object storage.

←Next→

Step **0** / 21

In this run, both deleting the 10:00 alarm and setting the new 10:05 alarm returned success while the old cleanup waited. When the simulator finally advanced that cleanup, it removed the 10:00 wake entry that the new alarm still needed. The checker reported an invariant violation: the 10:05 alarm was still set in SQLite, but no wake entry remained to reactivate the cell.

If the node restarted before 10:05, the missing wake entry could prevent the alarm from running, despite the earlier success response.

The simulator found this failing order by exploring requests and pending work. We had not written a test that prescribed this sequence.

## Reproducing and fixing the race

We replayed the failure with the same code, settings, and seed. Each run let us inspect the moment the old cleanup deleted the needed wake entry, without searching for the failing order again.

Since then, we have changed the wake-entry design. Each new alarm setting gets its own entry in object storage. A delayed deletion for an older alarm can only remove that older entry, leaving the new one intact. Old entries are removed once celld has confirmed they are no longer needed.

We turned the failing run into a regression test that repeats the requests and cleanup order shown above. It checks that the new alarm keeps a usable wake entry.

The regression test checks the invariant during the run, not just at the end. If the wake entry disappeared and was later restored, checking only the final state would miss the violation.

## What's next

The simulator is still under development, but it has already helped us find and fix previously unknown bugs in celld. Next, we will test more combinations of requests, storage failures, and node restarts. For example, a node could restart during a storage outage and receive new requests before storage recovers.

We will also add checks for more of celld’s guarantees. A write reported as successful, for example, must survive a node restart. When these tests find a bug, we can replay the failure, fix it, and keep the sequence as a regression test.

1. If celld needed to evict the cell while waiting for 10:05, it would first ensure that a wake entry for 10:05 was in place. ↩
