← All writing

What Are GPU Sandboxes? How Isolation Works When Agents Need a GPU (2026)

Carl Peterson · October 1, 2026 · 11 min read

A GPU sandbox is an isolated environment that gives code a real GPU without exposing the host. Isolation serves two purposes: it contains untrusted code, and it gives every reinforcement learning (RL) or eval rollout a clean, reproducible environment. Most AI sandbox platforms are CPU-only, because the isolation technology many of them use can't hand a GPU to a workload. Three managed platforms offer GPU sandboxes today: Thunder Compute, Modal and Daytona.

GPU sandboxes allow agents to run local models, score reinforcement learning (RL) rollouts or execute evals against self-hosted LLMs. This guide covers how GPU isolation works, the three ways platforms add a GPU to a sandbox, and how the providers compare.

Takeaways

  • GPU sandboxes give untrusted, agent-generated code a real GPU behind an isolation boundary.
  • Most AI sandboxes are CPU-only, because stock Firecracker doesn't support GPU passthrough.
  • Three isolation paths exist: VFIO passthrough (Daytona), gVisor (Modal) and Firecracker microVMs with custom GPU support (Thunder Compute).
  • Preemption, GPU range and persistence separate the three managed GPU sandbox platforms.
  • Thunder Compute's H100 sandboxes cost $2.95/hr and are never preempted.

What Is a GPU Sandbox?

A GPU sandbox is a short-lived, API-created environment that provides GPUs behind an isolation boundary. This protection stops the code from reading host memory, touching other tenants' data or taking over the machine. Agents typically create one sandbox per task, run code in it and destroy it when the task ends.

Isolation is what separates a GPU sandbox from a rented GPU. Agent-generated code is untrusted by definition, because no human reviews it before it runs. A GPU sandbox lets that code use the GPU while containing whatever the code does.

GPU Sandbox vs GPU Instance vs GPU Container

A GPU sandbox has a stronger boundary than a GPU container and a much shorter lifecycle than a GPU instance:

  • GPU instance: a virtual machine you control for hours or days, usually for one trusted user.
  • GPU container: an environment that shares the host's kernel and NVIDIA driver, so a container escape reaches the host.
  • GPU sandbox: an isolated environment that code creates, uses and destroys, often once per task.

A regular GPU instance is simpler and usually cheaper for long jobs when you are the only user and trust your code. GPU sandboxes fit untrusted code, code generated on the fly and code submitted by many users.

Why Most AI Sandboxes Don't Support GPUs

Most AI sandboxes are CPU-only because isolating a GPU is harder than isolating a CPU. CPU isolation is a solved software problem. On the other hand, GPU isolation depends on hardware features, hypervisor support and driver behavior.

CPU Isolation Is a Kernel Problem, GPU Isolation Is a Device Problem

CPU isolation requires memory separation, process separation and a controlled path for system calls. MicroVMs provide it by giving each workload its own kernel. User-space kernels like gVisor (Google's sandboxed container runtime) provide it by intercepting system calls.

GPU isolation requires hardware support, because a GPU is a PCIe device with direct memory access, its own memory and a complex driver. Passing a GPU to an isolated workload safely needs three pieces:

  • an IOMMU (the unit that controls a device's memory access) to stop the GPU from touching memory it shouldn't
  • a way to detach the GPU from the host driver
  • a hypervisor that can pass the GPU into the guest

Every host has to configure all three correctly.

Why Multi-Tenant Isolation Is Harder With GPUs

Multi-tenant isolation is harder with GPUs because a GPU holds state that outlives any single process. Model weights, intermediate tensors and kernel code sit in GPU memory. Sharing one card between untrusted tenants widens the attack surface. For that reason, GPU sandbox providers generally assign whole GPUs to one sandbox at a time rather than slicing a card between tenants.

Whole-GPU assignment means you pay for the full GPU for the life of the sandbox. The cost is the same whether the agent keeps the GPU busy or spends most of the session waiting on an LLM call.

Can Firecracker MicroVMs Use GPUs?

Stock Firecracker microVMs cannot use GPUs, because Firecracker doesn't support GPU passthrough. Firecracker (AWS's open-source microVM hypervisor) keeps its device model deliberately minimal to reduce its attack surface. PCIe device passthrough falls outside that model, so many Firecracker-based code sandboxes, including E2B and Vercel Sandbox, are CPU-only.

Thunder Compute takes a different approach and uses custom software to support GPUs on Firecracker MicroVMs.

The Nested Virtualization Constraint

Nested virtualization decides whether a platform can run hardware-isolated microVMs at all. MicroVMs need KVM (Linux's built-in hypervisor), which requires bare metal or a cloud VM that exposes virtualization extensions. Many cloud instances don't expose those extensions, so platforms running on them must use a different isolation approach.

Platforms with KVM access can pass a GPU into a VM. Platforms without KVM access typically fall back to syscall-level isolation such as gVisor.

Three Ways to Give a Sandbox a GPU

GPU sandbox platforms generally use one of three approaches, and each places the isolation boundary somewhere different. The right approach depends on how much you trust the workload and which trade-offs you accept.

VFIO Passthrough Into a Sandbox VM

VFIO passthrough (Linux's framework for handing physical devices to VMs) gives a virtual machine direct control of a physical GPU. Daytona's GPU sandboxes use VFIO passthrough to assign one dedicated physical GPU per sandbox VM. Because Daytona doesn't share cards, one sandbox cannot observe another sandbox's GPU work.

gVisor With a Proxied GPU Driver

gVisor isolates workloads with a user-space kernel that intercepts system calls instead of running a full guest kernel. For GPUs, gVisor proxies NVIDIA driver calls to the host driver. Modal Sandboxes, part of Modal's serverless AI platform, use gVisor for isolation and support GPUs from T4 to B300.

Firecracker MicroVMs With Custom GPU Support

Thunder Compute runs GPU sandboxes in Firecracker microVMs, the same hypervisor behind many CPU-only code sandboxes. Each Thunder Compute sandbox gets its own guest kernel and Firecracker's small device model. Thunder Compute's custom software makes a GPU available inside the microVM, so GPU workloads keep microVM-level isolation.

gVisor vs Firecracker vs VFIO Passthrough: How the Three Paths Compare

The three paths differ mainly in where the isolation boundary sits. MicroVMs and VFIO passthrough put a hardware-virtualized guest kernel between the workload and the host. gVisor filters the workload's system calls in user space instead.

Approach Isolation Boundary Guest Kernel GPU Assignment Example Platform
VFIO passthrough into a sandbox VM Hardware virtualization Yes Dedicated physical GPUs, never shared Daytona
gVisor with proxied GPU driver User-space kernel (syscall interception) No, a user-space kernel runs in its place GPUs reserved per sandbox Modal
Firecracker microVM with custom GPU support Hardware virtualization Yes One GPU per sandbox Thunder Compute

When Do You Need a GPU Sandbox?

You need a GPU sandbox when untrusted or generated code has to run on a GPU. An agent that only runs ordinary Python and calls a hosted LLM API needs only a CPU sandbox. Four common workloads need the GPU inside the isolation boundary.

Agents Running Local Models or Embeddings

Agents that call a local model instead of an external API need the GPU inside the sandbox. Teams choose local models for privacy, cost or latency. Running the model inside the sandbox keeps the agent's data and generated code behind one boundary, with no network hop to a separate inference service.

RL Rollouts and Reward Evaluation

RL pipelines need GPU sandboxes when rollouts or reward functions require GPU compute. RL pipelines execute model-generated actions and score them, often thousands of times per training step. Each rollout should run in isolation so a bad rollout can't corrupt the others. Non-preemptible sandboxes matter for RL, because an evicted rollout has to be rerun.

Coding-Agent Evals Against Self-Hosted Models

Coding-agent evals against self-hosted models need a GPU in the same environment as the generated code. Coding-agent benchmarks run generated code against test suites. One sandbox per eval task keeps runs reproducible and independent, and lets teams run many evals in parallel.

Multi-Tenant Code Execution With GPU Notebooks

Products that let users run their own code on GPUs need a boundary between customers. Hosted notebooks and AI app builders are common examples. Each user's session gets its own sandbox and its own GPU, and nothing persists between tenants.

Which Platforms Offer GPU Sandboxes in 2026?

Three managed platforms offer GPUs inside the sandbox itself: Thunder Compute, Modal and Daytona. Separate GPU inference APIs don't qualify, because the GPU sits outside the sandbox. The three platforms differ enough on GPU range, preemption and persistence to decide most choices. Details in the comparison table were verified in October 2026.

Feature Thunder Compute Modal Daytona
GPU types H100 (A100 80GB and RTX A6000 coming soon) T4, L4, A10, L40S, A100 40GB and 80GB, RTX PRO 6000, H100, H200, B200, B300 H100, H200, B200, B300, RTX PRO 6000, RTX 4090, RTX 5090, AMD MI355X
Isolation Firecracker microVM gVisor Sandbox VM with VFIO passthrough
Preemption Non-preemptible Preemptible only On-demand or spot
Max GPUs per sandbox 1 Not documented for Sandboxes 8
Persistence Ephemeral Volumes and filesystem snapshots Ephemeral, deleted on stop; filesystem snapshots
Deployment Managed Managed Managed, or BYOC

How Thunder Compute Runs GPU Sandboxes on Firecracker

Thunder Compute runs every GPU sandbox in a Firecracker microVM, so agent code gets a GPU without giving up microVM isolation. Thunder Compute sandboxes are short-lived, API-driven environments built for untrusted code that needs a GPU.

Why Firecracker Isolation and a GPU Aren't Mutually Exclusive

Firecracker doesn't support GPU passthrough, but Thunder Compute uses custom software to support GPUs in Firecracker microVMs. Each Thunder Compute sandbox keeps Firecracker's minimal device model and its own guest kernel. Teams don't have to choose between microVM isolation for untrusted code and GPU access.

Non-Preemptible GPU Sandboxes with H100s at $2.95 per Hour

Thunder Compute's H100 sandboxes cost $2.95 per hour, billed per second. An H100 costs $3.95/hr at Modal and as an on-demand H100 at Daytona. Every Thunder Compute sandbox is non-preemptible, so long evals and RL runs aren't evicted partway through.

Requesting Access and Getting Started

Thunder Compute enables sandbox access per organization, so the first step is requesting access. Teams create sandboxes through a Python SDK or REST API. Commands run as durable jobs that survive client disconnects, and outbound network rules range from fully open to blocked or allowlisted. Each sandbox currently supports one H100 and is ephemeral, so download results before terminating it.

Request GPU sandbox access on Thunder Compute.

Last Thoughts on GPU Sandboxes

GPU sandboxes give untrusted, agent-generated code a GPU without giving it the host, yet most sandbox platforms still don't offer them. The platforms that do take different paths: VFIO passthrough, gVisor or Firecracker microVMs with custom GPU support.

Choose by workload first, since preemption, GPU range, persistence and hourly price narrow the field fastest.

FAQ

What Is a GPU Sandbox?

A GPU sandbox is a short-lived, isolated environment that gives untrusted code, such as AI agent output, access to a real GPU. The isolation boundary keeps the code from reaching the host or other tenants.

Can Firecracker MicroVMs Use GPUs?

Stock Firecracker cannot pass a GPU into a microVM. Thunder Compute uses custom software to support GPUs, so its sandboxes run as Firecracker microVMs with a GPU.

How Is a GPU Sandbox Different From a GPU Container?

A GPU sandbox has a stronger isolation boundary than a GPU container. A GPU container shares the host's kernel and NVIDIA driver, so a container escape reaches the host.

Which Platforms Offer GPU Sandboxes in 2026?

Thunder Compute, Modal and Daytona offer managed GPU sandboxes as of October 2026. Beam and NVIDIA OpenShell also support GPUs in sandboxed environments.

Does Sandbox Isolation Slow Down GPU Workloads?

Isolation mostly affects the CPU side of a workload, such as system calls, file I/O and startup. GPU kernels run on the GPU itself. Benchmark your own job, including sandbox startup time, before committing.