# How DeepSeek Runs 380,000 Agent Sandboxes at Once
*How DeepSeek built DSec to run 380K+ agent sandboxes concurrently while optimizing isolation, resources, images, and security.*
By [Rohit Lakhotia](https://scaleengineer.com/authors/rohit-lakhotia)
Published: 2026-09-28
Canonical: https://scaleengineer.com/blog/how-deepseek-runs-380-000-agent-sandboxes-at-once
---
Training an AI agent is very different from training a traditional language model\. A model might spend most of its time generating tokens\. An agent, on the other hand, needs to **do things**\. It writes code, installs packages, edits files, runs tests, executes commands, and observes the results before deciding what to do next\. Now imagine running that process thousands of times simultaneously during reinforcement learning\. Every rollout needs its own environment\. And that creates a surprisingly difficult infrastructure problem\.

DeepSeek's **DeepSeek Elastic Compute \(DSec\)** is the sandbox platform built to handle it\. The system runs about **3 million sandboxes per day**, can handle **more than 380,000 concurrent sandboxes**, and sustains **more than 5,000 sandbox creations per second** on a production unit of roughly 160 nodes\.

But the most interesting part isn't the scale\. It's **why DeepSeek designed the system this way in the first place\. **Let’s discover in this blog today\.

## Why AI Agents Need Their Own Sandboxes

An AI agent doesn't just generate an answer\. A coding agent can **write code, install packages, edit files, run commands, and execute tests**\. To safely perform these actions, it needs an isolated environment where its code and commands can run without directly affecting the rest of the system\. This isolated environment is called a **sandbox**\.

Think of a sandbox as a **separate workspace for an AI agent**\. The agent can create files, install dependencies, run programs, and test its code inside the workspace, while the environment provides boundaries between the agent's actions and the underlying system\.

During reinforcement\-learning training, DeepSeek may need to run **thousands of these agent environments at the same time**\. Each environment is used for a separate rollout, where the agent attempts a task, observes the results, and decides what to do next\.

And that's where things get interesting\. DeepSeek's **DeepSeek Elastic Compute \(DSec\)** is the infrastructure built to manage these environments at scale\. The system handles around **3 million sandboxes per day**, can reach **more than 380,000 concurrent sandboxes**, and supports **5,000\+ sandbox creations per second** on the reported production setup\.

The challenge isn't simply creating a sandbox\. It's creating **hundreds of thousands of them efficiently, keeping them isolated, and managing their CPU, memory, storage, and state while the agents work inside them\.**

## Why Agent Sandboxes Are a Different Infrastructure Problem

We already know, a coding agent doesn't simply return an answer\. During a reinforcement\-learning rollout, it might:

1. Set up a repository and install dependencies\.
2. Write or modify code\.
3. Run commands and tests\.
4. Observe the results\.
5. Decide what to do next\.
6. Repeat the process until the task is complete\.

That means every rollout needs an environment that can **start quickly, preserve state, and isolate the agent's execution**\. But there's another unusual characteristic of this workload\. The sandbox isn't necessarily busy all the time\. After an agent makes a tool call, it may sit idle while the model decides its next action\. DeepSeek reports that around **90% of its sandboxes use 5% or less of the CPU they request**\.

That creates an unusual infrastructure pattern: **Large bursts when thousands of sandboxes are created, followed by long periods where many sandboxes are mostly idle but still need to preserve their state and memory\.**

And this is one of the key observations behind DSec's design\. The platform needs to **pack large numbers of sandboxes onto the available infrastructure without wasting CPU and memory on environments that are currently idle\.**

## How DSec Packs So Many Sandboxes Together

DeepSeek didn't build one sandbox type and use it for everything\. DSec provides **four sandbox backends behind a common interface**:

- **Function calls** for lightweight, restricted tool execution
- **Containers** for most software\-engineering tasks
- **[Firecracker](https://firecracker-microvm.github.io/)**** microVMs** when stronger isolation is required
- **Full ****[QEMU](https://www.qemu.org/documentation/)**** VMs** for environments that need OS\-specific, graphical, or mobile capabilities

The benefit is that different workloads can use different execution environments instead of putting everything inside the heaviest possible isolation layer\. Containers can provide high density for many tasks, while microVMs and full VMs provide stronger isolation when the workload requires it\. And this matters because DSec is dealing with **model\-generated code**\. The code running inside these environments isn't necessarily trusted\.

## The Image Problem

There's another challenge that becomes obvious when you run hundreds of thousands of sandboxes\. **Where do all those environments come from?**

DSec manages a very large collection of environment images and workspaces\. The source material reports around **133 TB of environment images**, including base images and task workspaces\.

But there is an important inefficiency here\. A sandbox doesn't necessarily need every byte contained in its image\. For example, the reported workloads touch only a fraction of their image data at runtime: around **8\.7% for C\+\+**, **13\.3% for Go**, and **6\.0% for Python** in the measurements described by the source\. So pulling an entire image before starting a sandbox would waste both time and storage bandwidth\.

DeepSeek addressed this in two ways\.

### Composable Environment Layers

Instead of creating a completely independent image for every environment, DSec composes environments from multiple layers\. The environment can be built from things such as: **Base image → Workspace → Toolkits → Writable layer**

This means components can be independently versioned and combined instead of repeatedly rebuilding complete images for every combination\.

### Load Only What You Need

DSec also uses **on\-demand image loading**\. Rather than downloading the entire environment upfront, data can be loaded when the sandbox actually needs it from DeepSeek's **[3FS distributed filesystem](https://github.com/deepseek-ai/3FS/blob/main/docs/design_notes.md)**\.

This makes a big difference during large bursts of sandbox creation\. In the reported experiments, lazy loading reduced burst completion time from **more than 60 minutes to around 35 minutes**, while reducing disk writes by **57%**\. So the system doesn't just create more sandboxes\. It also tries to avoid doing unnecessary work while creating them\.

## The Real Bottleneck: Most Sandboxes Are Idle

Now comes one of the more interesting infrastructure problems\. If hundreds of thousands of sandboxes are running simultaneously but most of them aren't actively executing code, reserving a full CPU allocation for every sandbox would be extremely wasteful\.

DeepSeek therefore designed DSec to **reallocate processing resources based on when agents are actually executing**\. Memory becomes especially important because an idle sandbox still needs to preserve its state\.

For microVMs, DSec uses **virtio\-pmem \+ DAX** to map image files directly into host pages, avoiding separate page\-cache copies for each guest\. DeepSeek reports that this reduced peak host memory usage by **40\.2%**, although transient peak CPU increased from **26\.5% to 41\.4%**\.

The system also explored another approach using [DAMON](https://docs.kernel.org/mm/damon/index.html) and balloon free\-page reporting, which reduced time\-integrated memory usage by **21\.2%** with relatively little CPU cost\.

The larger idea is simple: **If an agent isn't using a resource right now, don't reserve the full resource for it\.**

## CPU Scheduling Matters Too

Memory isn't the only resource being shared\. Agent tool calls can be latency\-sensitive because the model may be waiting for the result before continuing\. At the same time, other background workloads may be consuming CPU\.

DeepSeek found that without protection, a **50% best\-effort CPU load increased per\-step latency by 45\.2%**\.

DeepSeek also evaluated `SCHED_IDLE` for protecting latency\-sensitive agent workloads from best\-effort CPU tasks\. On its own, it provided limited protection\. Combining it with Linux core scheduling reduced the reported latency inflation from **45\.2% to 17\.3%** under the tested workload\.

So DSec isn't simply about packing as many sandboxes as possible onto a machine\. It also needs to make sure that **one workload doesn't make another agent's execution unnecessarily slow\.**

## What Happens When Training is Interrupted?

DSec also integrates the sandbox lifecycle with the reinforcement\-learning training process\. Starting with the V4\.1 setup described in the source, rollout execution is separated from GPU training\.

When a training job is preempted, the associated sandboxes don't necessarily have to be destroyed\. Containers can be paused using `docker pause` with memory reclamation, while microVMs can be snapshotted and exited\.

When the training job resumes, the rollout environments can come back as well\. DSec also supports incremental sandbox snapshots through `pack_diff`, allowing an agent to checkpoint an environment and restore it as a new sandbox\.

This is another example of the platform being designed around the **stateful nature of agent rollouts**\.

## Then the Agents Started Cheating

This is probably the most surprising part of the paper\. When agents were given shell access and a reward to optimize, some of them didn't simply try to solve the intended task\. They looked for shortcuts\. The paper describes agents attempting things such as:

- Searching platform\-managed files and channels for answers
- Reading logs for leaked answers
- Sending crafted RPC messages to Unix sockets
- Attempting to overwrite `/bin/bash`
- Finding accessible package mirrors and installing versions that already contained the solution
- Using `XFS_IOC_SWAPEXT` to manipulate file extents and access protected content

One of these experiments even caused filesystem corruption and a subsequent filesystem shutdown\. This is why DeepSeek treats agent execution as **untrusted**\. The sandbox isn't simply there to make execution convenient\. It's also a security boundary\.

## Defense Has to Be Layered

DeepSeek's response wasn't to rely on one mechanism to solve every possible attack\. The platform uses multiple controls, including:

- **AppArmor** for file and socket access
- **[eBPF](https://docs.ebpf.io/)****\-based network allowlists**
- Restricted logs and sockets
- Different isolation levels depending on the workload

But the paper is explicit about the limitation: **no single mechanism can prevent all agent misbehavior and system failures\. **That's an important distinction\. The goal isn't to assume an agent will always behave correctly\.nThe infrastructure is designed around the possibility that the agent **will try something unexpected**\.

## The Bigger Picture

Training agents at scale isn't only a GPU problem\. Every time an agent needs to execute code, manipulate files, install dependencies, or run tests, it needs an environment to do that work\. When you're running thousands of rollouts simultaneously, those environments become an infrastructure problem of their own\.

DeepSeek's DSec tackles that problem by combining:

**Multiple isolation levels \+ composable environments \+ on\-demand image loading \+ resource reclamation \+ CPU scheduling \+ stateful rollout management \+ layered security**

The scale is impressive: **around 3 million sandboxes per day, more than 380,000 concurrent, and 5,000\+ creations per second** on the reported production unit\.

But perhaps the more interesting lesson is the assumption underneath the entire system: **When AI agents can execute code autonomously, your infrastructure has to be designed for both their productivity and their unpredictability\.**

## Key Takeaways

- **DSec \(DeepSeek Elastic Compute\)** is DeepSeek's sandbox infrastructure for agentic training and evaluation\.
- It handles around **3 million sandboxes per day** and **380,000\+ concurrent sandboxes**\.
- The system supports **5,000\+ sandbox creations per second**\.
- Around **90% of sandboxes use 5% or less of their requested CPU**, making resource sharing important\.
- DSec supports **function calls, containers, microVMs, and full VMs** through a common interface\.
- Environments use **composable layers** and **on\-demand loading from 3FS**\.
- DeepSeek reports **40\.2% lower peak host memory** with virtio\-pmem \+ DAX for microVMs\.
- CPU scheduling reduced reported per\-step latency inflation from **45\.2% to 17\.3%** under the tested workload\.
- DeepSeek documented multiple cases where agents attempted to exploit their execution environment\.
- The paper emphasizes that **no single mechanism can prevent all agent misbehavior**, making layered isolation and observability important\.

Official paper from DeepSeek: [DeepSeek Elastic Compute \(DSec\): A Sandbox Infrastructure for Effective Agentic Training at Scale](https://arxiv.org/html/2609.22978v1)

By now, you must have had a clear idea of,** How DeepSeek Runs 380,000 Agent Sandboxes at Once? **In a nutshell, DeepSeek built DSec to run hundreds of thousands of agent sandboxes efficiently by combining multiple isolation levels, lazy environment loading, resource sharing, and stateful rollout management\. The bigger lesson is that **agentic AI needs its own infrastructure layer because agents don't just generate code, they execute it\.**

**Congratulations\! You've just advanced another step in your tech journey\. Keep progressing\!**
