DeepSeek Expands Elastic Computing Team as DSec Sandbox Infrastructure Scales
DeepSeek is hiring senior engineers for its elastic computing team and detailed DSec, its sandbox infrastructure for agent training.
DSec currently runs one scaling shard with about 160 servers, 30,000 CPU cores and 250 terabytes of memory, serving roughly 3 million sandboxes per day, QbitAI reported. At peak times, more than 380,000 sandboxes are online simultaneously, and the system can create more than 5,000 sandboxes per second. DeepSeek has deployed multiple such shards in production, supporting millions of concurrent sandboxes. The next goal is to expand the number and types of agent runtime environments by hundreds or thousands of times.
The infrastructure supports all of DeepSeek-V4's training, evaluation and data preprocessing, according to the report. From DeepSeek-V3.2 to V4.1, all sandbox workloads generated by agent training, evaluation and data preprocessing have run on DSec. The workloads differ from ordinary cloud computing: creation requests arrive in bursts; CPUs are idle most of the time but memory must be retained for files and processes; different agent tasks require very different execution environments; base images have low reuse rates; and training may be interrupted when GPU resources are preempted.
To reduce the cost of supplying agent environments, DSec splits each environment into three layers: a base image for the operating system and basic software, a workspace for task code and dependencies, and a toolkit for tools such as DeepSeek Harness. Versions are managed separately and combined when a sandbox is created, so updating a tool only requires rebuilding the toolkit layer. In one week of production data in 2026, the Container backend used 11,266 base images, 102,171 workspaces and hundreds of toolkits. DeepSeek found that agents actually access only 4.2 percent to 13.3 percent of an image, so it stores image data in its 3FS distributed file system, keeps metadata locally and reads data on demand. In a test creating 8,192 containers, full image pulls took more than 60 minutes, while on-demand loading cut the time to about 35 minutes, a roughly 1.71-fold speedup, and reduced disk writes by about 57 percent. In another workspace experiment, replacing per-sandbox tar.gz extraction with direct EROFS layer mounting reduced a task from 79 minutes to 45 minutes and cut disk writes to about one-fifth of the original.
DeepSeek also found that agent sandboxes are sparse. About 90 percent of sandboxes use less than 5 percent of their requested CPU resources on average, because agents often wait for the model to generate the next action while files and processes must remain in memory. In production, DSec's resource oversubscription rate exceeds 50 times. To prevent memory from becoming a bottleneck, DSec uses virtio-pmem and DAX to let MicroVMs on the same host share the host page cache, lowering peak host memory use by 40.2 percent in an experiment. DAMON and balloon idle page reporting reclaim unused memory, reducing cumulative memory consumption by another 21.2 percent. DSec also prioritizes latency-sensitive agent tasks and lets less time-sensitive tasks use remaining CPU capacity. In an experiment where other tasks on the same machine had already consumed 50 percent of the node's CPU capacity, scheduling optimization reduced the latency impact on sensitive tasks from 45.2 percent to 17.3 percent.
DSec supports four execution backends: FnCall for short tasks such as online evaluation, Container for software engineering and tool calls, MicroVM for stronger isolation, and Full VM for a complete operating system that can support GUI, graphics rendering and Android applications. DeepSeek is also using agents to build environments for agents. A mechanism called pack_diff lets an agent generate an incremental snapshot after configuring an environment in a sandbox, so the same environment can be restored later. Agents can also save the state after each step as a reusable environment. This enables trajectory forking: at step k, the system can save a snapshot and restore multiple sandboxes from the identical state, letting different branches continue while sharing earlier environment and data and recording only later changes.
The report said DSec also decouples agent execution from GPU training. Previously, the agent execution loop and GPU training ran in the same Pod, so GPU preemption could interrupt the agent loop, and recovery required replaying command records to align training progress with the actual sandbox state. Starting with DeepSeek-V4.1, that execution logic moved into DSec. Agent sandboxes and worker containers that advance interactions now run outside the preemptible GPU resource pool, so environment state can continue to be saved even when GPU tasks are interrupted, and training can resume from the interruption when GPUs return.
DeepSeek has observed agents trying to exploit environments in production, including reading residual answers, forging RPC requests, overwriting /bin/bash to inject commands, and calling XFS_IOC_SWAPEXT to bypass access controls, according to QbitAI. DSec restricts file and socket access with AppArmor, which remains effective even if an agent gains administrative privileges, and uses eBPF to set a per-sandbox network whitelist for addresses, ports and protocols. DeepSeek acknowledges these measures do not solve every problem; for example, if an agent finds a kernel vulnerability, there is still no general defense. The company expects attack and defense between models and agents to continue as model capabilities improve.