001
DeepSeek Elastic Compute (DSec):
A Sandbox Infrastructure for Effective Agentic Training at Scale
Abstract
Large-scale agentic training and evaluation with large language models (LLMs) rely on isolated, stateful execution environments in which models inspect repositories, invoke tools, execute commands, and interact with task-specific services. These workloads create sandboxes in large bursts, span heterogeneous functionality and isolation requirements, retain state across long interactions, and draw from large image corpora with limited reuse. Supporting them therefore requires an elastic execution platform rather than a single sandbox runtime. This report presents DeepSeek Elastic Compute (DSec), a production sandbox platform that exposes FnCall, container, microVM, and full-VM sandbox backends through a unified SDK. DSec coordinates placement and lifecycle management across the cluster, composes environments from independently versioned layers, combines memory sharing, reclamation, and CPU scheduling for high-density execution, and loads image data on demand from Fire-Flyer File System (3FS), a cluster-wide distributed filesystem. DSec is co-designed with the reinforcement learning (RL) framework, decouples stateful rollout execution from preemptible GPU training, coordinates sandbox lifecycle with training to preserve rollout state while reclaiming idle resources, and mitigates agent misbehavior such as reward hacking. A single production-scale unit of DSec spans around 160 nodes, serving about 3 million sandboxes per day; in production, it supports over 380,000 concurrent sandboxes and sustains over 5,000 sandbox creations per second. Our evaluation and deployment experience show that these mechanisms reduce environment setup and image-distribution overhead, improve memory efficiency, and preserve latency-sensitive performance under high-density overcommit.
1 Introduction
Recent advances in frontier LLMs have made agentic workflows practical and widely adopted (Guo et al., 2025; OpenAI et al., 2024; Jimenez et al., 2024). Instead of producing a single text answer, an agentic model interacts with an execution environment: it may navigate codebases, call tools, execute commands, inspect failures, and modify files, or operate browsers and desktop applications through graphical interfaces in computer-use tasks (Xie et al., 2024; Zhou et al., 2024). Across these workloads, the model iterates based on feedback until a task is solved. This execution model has led to a growing ecosystem of agent tools and orchestration harnesses, such as DeepSeek Harness (DSH) (Shi et al., 2026), OpenCode (Anomaly, 2025), and multi-agent training harnesses. Training reliable agents requires reinforcement learning (RL) at scale, in which models learn through interaction with real, isolated execution environments rather than solely from static input-output examples.
The agentic training pipeline encompasses environment and data construction, RL rollouts, reward computation, policy updates, and periodic evaluation. Among these stages, RL rollout and evaluation impose the highest pressure on the sandbox platform because they are large-scale, concurrent, and tightly coupled with the training loop. In RL (Ouyang et al., 2022; Guo et al., 2025), training proceeds as a feedback loop with three stages. First, during rollout, the current model interacts with the sandboxed environment: it reads files, issues tool calls, executes commands, observes outputs, and produces a trajectory for each task. Second, during reward computation, the framework scores the trajectory using native execution signals such as exit codes, stdout, test pass rates, or task-specific verifiers. Third, during policy update, the RL algorithm updates the model parameters from the collected trajectories and rewards. Periodic evaluation follows a similar execution path, except that the resulting trajectories are used to measure model capability rather than to update parameters. Recent systems further pipeline generation and policy optimization through asynchronous rollouts, continuously replenishing completed samples to maintain high concurrency and mitigate long-tail stragglers (DeepSeek-AI, 2026). For agentic workloads, this design keeps many stateful sandbox sessions in flight and may interrupt and resume their associated rollouts across policy updates or scheduler preemptions, further increasing the platform’s concurrency, lifecycle-management, and state-consistency requirements.
For each rollout or evaluation task, the platform must materialize an isolated task-specific environment, including its repositories, dependencies, services, evaluation scripts, and coding harnesses. The environment must be close enough to a real machine to run unmodified software stacks, package managers, build tools, browsers, emulators, and task-specific services. A robust, high-throughput sandbox runtime is therefore foundational for obtaining accurate and verifiable RL and evaluation results.
Agentic sandbox workloads have several properties that shape the platform design:
(1)
Rollout and evaluation jobs create sandboxes in a bursty manner. A single job may request up to 32K sandbox instances, so the platform must accept and place many sandboxes concurrently. Such bursts make horizontal scalability a system-wide requirement and require shared services, such as scheduling and image distribution, to avoid centralized bottlenecks.
(2)
Sandboxes must run at high density. During agent interaction, a sandbox often waits for the LLM to generate the next action, so CPU usage is sparse and naturally suitable for overcommit. For instance, in production, this allows a single node to host up to 800 microVMs or 3,200 containers, but only if the platform can safely overcommit resources and manage lifecycle pressure at node scale.
(3)
Agent sandboxes are stateful and long-lived. The model may modify files, install dependencies, and start services, and later tool calls depend on this accumulated state. Since a sandbox can stay alive across many LLM interaction turns, memory footprint, guest page cache, host page cache, and writable state may remain pinned long after the CPU becomes idle. Under high-density overcommit, these resident costs directly limit cluster capacity, so memory sharing and reclamation become important platform requirements.
(4)
Agent workloads are highly heterogeneous. The platform must cover OJ-like script execution, software-engineering tasks over full repositories, security tasks, computer-use workloads, mobile development environments (e.g., Android), and other full-system environments. These workloads differ substantially in CPU and memory demand, dependency footprint, required system functionality, and isolation strength. A single sandbox abstraction cannot cover all of them efficiently. For example, lightweight function calls are preferable for short stateless tasks, whereas virtual machines (VMs) are better suited to workloads that require a complete commercial off-the-shelf operating system.
(5)
Environment diversity is high even within the same workload class. Training and evaluation corpora contain many tasks, and each task may require its own repository, dependency versions, services, toolkits, evaluation scripts, or VM snapshots. As a result, the platform must serve a large number of distinct images and environment artifacts, with limited reuse for many of them. Under bursty startup, fetching these diverse task images from a registry would concentrate load on the distribution path, inflate startup latency, and introduce extra I/O that interferes with already-running sandboxes. In our ablation, eager image pulling stretches completion time by 1.7, while on-demand loading reduces cumulative disk writes by 57%.
(6)
Agent execution is untrustworthy. Agents may corrupt filesystems, exhaust resources, or interfere with system components, potentially disrupting rollouts or other co-located workloads. The platform therefore requires fine-grained access control and misbehavior analysis to contain and diagnose agent-induced failures.
(7)
Agent execution is interruptible. GPU training jobs may be preempted while long-running rollouts are still in progress. The platform must therefore preserve execution state and support efficient recovery across interruptions.
These properties define the role of an agent sandbox platform. DSec provides elastic service scaling, high-density resource management, memory sharing and reclamation, multiple isolation mechanisms for different workload classes, scalable image distribution, and explicit integration with the training framework for preemption-safe resumption, task-specific network policy, and agent misbehaving mitigation.
The rest of this report presents DSec from platform abstraction to implementation and evaluation. introduces DSec from the user perspective, including supported workloads, sandbox backends, and operating scale. describes the end-to-end platform architecture. characterizes the production workload and the platform challenges it creates. presents the core system mechanisms for environment composition, image distribution, and high-density resource management. describes co-design with the RL framework for environment construction, state preservation, resource reclamation across preemption, and the analysis of agent misbehavior with targeted access-control mitigations. summarizes additional implementation details. evaluates the effectiveness of the design, and discusses related works.
2 Overview of DSec
This chapter presents the user-facing view of DSec. From the platform’s perspective, users are the training frameworks, evaluation frameworks, and data-construction pipelines that call the software development kit (SDK) on behalf of researchers; we refer to them collectively as users throughout the report. It covers the SDK entry point, the sandbox backends exposed by the platform and the workload classes they serve, the lifecycle of a sandbox session, and the operating scale of the production deployment.
2.1 SDK Entry Point
Users access DSec through libdsec, a Python client library for the sandbox service. libdsec gives users a unified SDK entry point for creating and operating sandboxes, while still requiring them to choose the sandbox backend appropriate for the task. A typical request specifies the sandbox type, image or environment identifier, CPU and memory limits, lifetime settings, network rules, and initial user context. After creation, the user can execute shell commands or tool calls and collect command outputs and return status. shows a minimal container session: the client connects to the service endpoint, requests a sandbox with the desired resource and network policy, runs a command, and releases it.
In this example, the network rules allow access to PyPI (pypi=True) but deny access to NPM (npm=False). This fine-grained network control is discussed in detail in . This interface is intentionally not a full semantic abstraction over all backends. Function calls, containers, microVMs, and full VMs have different startup costs, isolation boundaries, filesystem semantics, and operating-system capabilities. libdsec provides a unified access path and a similar operational model, but the caller remains responsible for selecting a backend that matches the workload.
2.2 Sandbox Backends
Sandbox runtimes face a fundamental tension: stronger isolation and more complete system functionality usually come with higher startup latency and resource overhead. Since no single sandbox abstraction fits all agentic tasks, DSec supports multiple backends spanning this tradeoff space. summarizes their typical fit.
| Characteristic | FnCall | Container | MicroVM | Full VM |
|---|---|---|---|---|
| Runtime Performance | ● ● ● | ● ● ○ | ● ◐ ○ | ● ○ ○ |
| Dependency footprint | ○ ○ ○ | ● ● ● | ● ● ● | ● ● ○ |
| Isolation level | ○ ○ ○ | ● ● ○ | ● ● ● | ● ● ● |
| Full OS functionality | ○ ○ ○ | ● ○ ○ | ● ● ○ | ● ● ● |
| Resource overhead | ○ ○ ○ | ● ○ ○ | ● ● ○ | ● ● ● |
| Scenarios | OJ-like tasks | SWE | Security | COTS OS |
| GPU kernel exec | Tool use | Computer use | Graphics |
More ●= higher demand.
FnCall targets short, stateless tasks such as OJ workloads, code compilation, serverless programs, GPU kernels, and utility code. FnCall tasks run in reusable precreated CPU or GPU containers, avoiding per-invocation provisioning overhead. For GPU workloads, FnCall supports (i) shared mode, where multiple containers share a GPU instance, maximizing utilization for lightweight workloads, and (ii) exclusive mode, where one container reserves a GPU instance during its lifecycle for performance-sensitive tasks (e.g., operator evaluation). Containers are the main backend for software-engineering and general tool-use workloads. They provide fast startup and high packing density, and they run the Linux software stacks used by most repository-level tasks. Their main limitation is that they share the host kernel, which is not always appropriate for security-sensitive tasks. Firecracker microVMs (Agache et al., 2020) provide a stronger isolation boundary while retaining Linux compatibility. They are useful for security-sensitive tasks, stronger tenant isolation, and workloads that need a VM boundary with Linux compatibility. This comes at higher memory overhead and slower startup than containers. Full VM backends cover workloads that require a complete commercial off-the-shelf operating system environment, such as Android VMs through QEMU (Bellard, 2005), as well as those that require a GUI or graphics rendering. These backends have the highest resource overhead, but they are necessary for tasks that depend on OS-specific APIs, mobile runtime behavior, or full-system execution.
In production, containers and microVMs dominate both instance count and resource consumption. FnCall serves a large number of lightweight invocations with a small set of resident environments, while full VM backends cover specialized but important workload classes.
2.3 User-Visible Lifecycle
Although the supported backends differ internally, users see a unified high-level lifecycle. First, the caller creates a sandbox by selecting a backend and specifying the environment artifact, resource limits, lifetime policy, and network policy. The environment artifact varies by backend and workload. For containers and microVMs, it is a base image together with task-specific workspace and toolkit layers, which the platform composes into the running environment. For full VM workloads, it is a prepared VM image or snapshot. For FnCall, it is a task specification containing the task type, dependency files, and the code or script to run. These artifacts become the basis for the environment composition and image-distribution mechanisms discussed later in the report. Second, the platform prepares the environment and makes it ready for interaction. Third, the user issues commands or tool calls, observes outputs, and runs task-specific checks or tests. A sandbox is stateful throughout its lifetime: file edits, installed dependencies, and started services persist across calls, so later commands observe the effects of earlier ones. Because a sandbox stays alive across many interaction turns while its CPU is often idle between them, its resident state remains pinned long after the last command, one of the high-density challenges characterized in . Finally, the sandbox is stopped explicitly or reclaimed once its time-to-live elapses, so that idle or abandoned sessions do not hold resources indefinitely.
2.4 Deployment Scale
DSec is deployed across multiple scale units that share a 3FS (DeepSeek-AI, ) distributed file system deployment for base images and workspace storage. Within one scale unit, the platform spans nearly 160 CPU nodes with 30K cores and 250 TB of DRAM. It manages petabytes of layers and images. On a typical day, a single scale unit serves about 3 M sandbox instances, with peak concurrency reaching 380K and a creation rate exceeding 5,000 instances per second.
These numbers are important for understanding the rest of the report. DSec is not a single sandbox runtime or a thin wrapper around containers. It is a production execution platform that must combine user-facing sandbox abstractions, backend-specific runtimes, scalable image storage, high-density resource management, and training-framework integration.
3 Platform Architecture
presented DSec as users see it: an SDK, a set of sandbox backends, and a session lifecycle. This chapter turns to the platform behind that interface and describes how a request travels from the SDK to a running sandbox and which components it passes through. We describe the architecture in terms of cluster-level services and the sandbox runtime. Cluster-level services provide request ingress, identity and access management, sandbox placement, and a view of cluster health and load. The sandbox runtime handles node-local admission, sandbox creation, execution, and resource reclamation, relying on 3FS for image data.
3.1 Overview
At a high level, a sandbox creation request is first sent to IAM for authentication and authorization. Once authorized, the request proceeds to the placement engine, which selects a target node using health and load information collected by the watcher. After placement, the apiserver forwards the request to the edge on that node. The edge then checks local capacity, creating the sandbox with the requested backend if capacity permits and rejecting the request otherwise. Image data needed by the sandbox is stored in 3FS and fetched on demand during startup and execution. Container, microVM, and full VM sandboxes run a per-sandbox proxy (aether) and one or more chronus instances for command execution, filesystem access, and other runtime operations. After one of these sandboxes is running, its operations are routed through the apiserver, edge, aether, and chronus. FnCall, by contrast, uses neither aether nor chronus and follows a separate request path: the submitted task is executed directly in a precreated container, followed by best-effort cleanup of task state.
3.2 Cluster-Level Services
Cluster-level services manage access to the platform and coordinate sandbox requests across compute nodes. They comprise IAM, the apiserver, the placement engine, and the watcher.
IAM. Identity and Access Management (IAM) authenticates callers and authorizes all management requests to DSec. For example, requests to create or delete sandboxes or change a user’s resource or concurrency limits must pass IAM checks before execution. A principal is the user or service identity associated with a management request. IAM uses projects to define scopes for resource management and access control. Within a project, access policies specify which principals may perform which management operations on its resources, while resource quotas limit resource consumption.
We support multi-level project nesting rather than the flat or two-level hierarchies common in cloud platforms. Authorized principals, including agents and harnesses, can create subprojects, delegate part of the parent quota, and grant management permissions within them. Delegation is bounded by the parent: a principal cannot grant permissions it does not hold, and subproject policies and quotas cannot exceed the parent’s access-control or resource limits. Humans and agents use the same management API and authorization model.
API Server. The apiserver serves as the ingress proxy for the sandbox cluster. Training and evaluation code invokes libdsec from trusted GPU servers, while sandboxes execute untrusted model-generated code and may access external networks. The two sides are therefore network-isolated, with the apiserver as the only permitted communication path. All sandbox requests, including creation, command execution, and streaming I/O, pass through this ingress. The apiserver maintains no per-sandbox state. It periodically refreshes the set of edge nodes from the watcher, while each sandbox ID encodes its owning edge. Any apiserver instance can therefore resolve and forward a request directly to the target edge, enabling the ingress tier to scale horizontally.
Placement Engine. The placement engine selects a host node for each new sandbox. Placement proceeds in two stages: filtering and ranking. The filtering stage retains only healthy nodes that provide the backend and hardware capabilities required by the request. For example, a request for a GPU-enabled sandbox is restricted to nodes equipped with the required GPUs. The ranking stage randomly samples a few eligible nodes and selects the least loaded among them.
Watcher. The placement engine’s decisions are only as good as its view of the fleet, which the watcher provides. The watcher periodically probes the health of each edge and host and collects scheduling-relevant state, such as the number of running sandboxes across backend types, broken down per edge, per user, and per task. The placement engine periodically pulls this state from the watcher and uses the latest view when evaluating new creation requests.
Please note that neither the placement engine nor the watcher requires durable state. The placement engine keeps no sandbox execution state, and the watcher can rebuild its fleet view after a restart by polling the edges again. This makes placement engine and watcher instances easy to add or replace without a costly recovery step.
3.3 Sandbox Runtime
The sandbox runtime creates and operates individual sandboxes and manages their resources. It includes edge, aether, and chronus, and relies on 3FS for shared image storage.
Edge. Each node runs an edge, a per-machine component that handles creation requests from the apiserver for container, microVM, QEMU-based full VM, and FnCall backends. Before accepting a creation request, the edge checks the node’s current capacity and rejects the request if capacity is insufficient. This node-local admission check complements the placement engine’s placement decision, which is based on periodically refreshed cluster state. During creation, the edge provisions storage, applies the eBPF-based network policy, and launches the runtime.
FnCall and containers run inside QEMU/libvirt VMs rather than directly on the host. The VM provides an isolated kernel and network stack and serves as an additional security boundary between untrusted containers and the bare metal. To better support graphics-intensive workloads, such as computer-use GUI applications, browsers, video games, and 3D rendering, we leverage para-virtualized GPU interfaces of the host hypervisor (e.g., virtio-gpu). Within the full VM, we support both workloads whose graphics APIs are natively compatible with the host OS, as well as those whose rendering stacks can be translated into host-native APIs through compatibility layers such as DXVK (DXVK, 2018).
Besides tracking the sandbox lifecycle, edge coordinates disk and memory snapshots and releases node-local resources when the sandbox stops or its TTL expires.
Aether. Container and VM sandboxes run aether, a cross-platform proxy that establishes a communication channel with the edge. This edge-to-aether channel uses a platform-specific transport, such as a Unix domain socket for Linux containers or vsock for VM backends. The edge monitors sandbox health through the channel and marks the sandbox as failed if the channel closes. For each operation, aether uses the operation’s terminal-session identifier to create or locate the corresponding chronus instance, then forwards the operation over the local channel. When the session ends, aether terminates the corresponding chronus process tree.
Chronus. chronus provides a shell-session abstraction inside the sandbox, with each instance representing one independent shell session. It exposes cross-platform interfaces for command execution, filesystem operations, HTTP requests, and streaming I/O. Multiple chronus instances can run concurrently within the same sandbox. Together, aether and chronus allow libdsec to expose a unified interface for container and VM sandbox operations.
Base image and workspace storage. The sandbox runtime uses 3FS as a shared backing store for base images and workspace images. Container images are converted offline from OCI into EROFS, which separates metadata from data so that metadata is kept local while image data remains in 3FS. MicroVM disk images use an OverlayBD (Li et al., 2020) format over the same storage. Together, these image formats support on-demand loading and incremental snapshots over a shared base, allowing an edge to start a sandbox without first pulling a full image. characterizes the workload pressures that make scalable image distribution necessary, while describes the corresponding on-demand loading mechanism.
3.4 Cloud Bursting with Selective Offloading
DSec uses cloud VMs to absorb transient peaks in sandbox demand while serving the steady-state workload on-premise. When on-premise utilization exceeds 80%, the placement engine offloads a portion of eligible incoming sandbox creation requests to cloud VMs.
Instead of combining a managed container service with object storage, we reuse the on-premise container runtime and EROFS-based image-loading path on cloud VMs. The EROFS images reside in a cloud-hosted distributed filesystem and are mounted by the cloud VMs.
Production file-access traces show that a compact, de-duplicated EROFS image set totaling 30 TB covers the image files accessed by 70% of container tasks. We synchronize this shared image set to the cloud filesystem offline. Container tasks whose image dependencies are fully contained in this set are classified as cloud-eligible. Other tasks remain on-premise. In production, 200 cloud VMs in one scale unit absorb 30% of peak overflow, increasing capacity without over-provisioning the on-premise cluster.
(Paper continues on arXiv; extract covers abstract and opening sections. Full text: https://arxiv.org/abs/2609.22978)