Give your agent a GPU.
Firecracker microVMs with a GPU attached. Create one from Python or MCP, run anything, and it's gone when you're done.
A full virtual machine
Every sandbox boots its own Ubuntu microVM on Firecracker.
Compatibility
systemd, docker compose and the rest of a normal Linux box work like they do on a server.
Isolation
Each sandbox runs its own kernel, not one shared with other workloads.
Bring your own image
Build from a Dockerfile or pull from a registry.
Scale with your agents
Start a sandbox per task and run them side by side. Each one lives and ends with the agent that started it.
Agent-first by design
Every action is an SDK or MCP call, and the logs from every sandbox land in one place.
import thunder_sandbox as thunder
sandbox = thunder.Sandbox.create(
cpu=4, memory=32, storage=50,
gpu_type=thunder.GPUType.H100, gpu_count=1,
timeout=900,
)
sandbox.wait_until_ready()
process = sandbox.exec("nvidia-smi")
Run your code and let the sandbox expire
Every sandbox has a lifetime. Stop it yourself or let the TTL clean up.
Pay by the second
H100 from $2.95/GPU-hr, billed per second
Decide what your sandbox can reach
Pick a policy when you create it. Change it while it runs.
- 01
Open
Full outbound internet. Private and cloud metadata networks stay blocked.
{ "internet_access": "open" } - 02
Restricted
Only the domains and IP ranges you allow. Everything else is refused.
{ "internet_access": "restricted", "domain_allowlist": ["pypi.org"] } - 03
Closed
No outbound traffic at all. Upload what you need and run fully offline.
{ "internet_access": "closed" }
A sandbox for every task
Give every task its own machine, then throw it away.
RL environments
Give every rollout its own sandbox with a GPU and a real kernel. Run them side by side, and each one expires when its episode ends.
- One sandbox per rollout, each with its own GPU
- Its own kernel, so environments can’t see each other
- A TTL cleans up anything your trainer forgets
envs = [
thunder.Sandbox.create(
"python", "rollout.py", "--episodes", "64",
gpu_type=thunder.GPUType.H100, gpu_count=1,
timeout=900,
)
for _ in range(8)
]
GPU CI/CD
Run your CUDA tests on an NVIDIA GPU for every push, each in a fresh VM. Billing is per second, so a four-minute job costs four minutes.
- A clean VM per run, with nothing left over from the last
- Build the image from a Dockerfile in your repo
- A TTL ends hung jobs before they run up a bill
ci = thunder.Sandbox.create(
gpu_type=thunder.GPUType.H100, gpu_count=1,
timeout=1800,
)
ci.wait_until_ready()
ci.upload("repo", "/home/ubuntu/repo", recursive=True)
tests = ci.exec("pytest", "-x", "tests/gpu",
workdir="/home/ubuntu/repo")
exit_code = tests.wait()
ci.terminate()
Evals
Score each task in its own sandbox, so one bad sample can’t affect the rest. Close the network and the model can’t look answers up.
- One sandbox per task, with nothing shared between them
- Network closed or allowlisted, per sandbox
- Per-second billing, so short tasks stay cheap
def score(task):
box = thunder.Sandbox.create(
gpu_type=thunder.GPUType.H100, gpu_count=1,
timeout=600, block_network=True,
)
box.wait_until_ready()
box.upload(task, "/home/ubuntu/task", recursive=True)
run = box.exec("python", "task/grade.py")
return run.wait()
Drive it from code or from your agent
A typed Python SDK, MCP tools for your coding agent, and a REST API
$ pip install thunder-sandbox
>>> sb = thunder.Sandbox.create(
... gpu_type=thunder.GPUType.H100,
... gpu_count=1)
>>> sb.exec("nvidia-smi")
Pick the right tool for the job
Sandboxes for short, automated work. Instances for long-running development.
| Sandboxes | Instances | |
|---|---|---|
| Lifetime | Minutes to hours, ends on a TTL | Runs until you delete it |
| Created from | Python SDK, MCP or REST API | Console, CLI or VS Code |
| Disk | Wiped when it ends | Kept, with snapshots |
| Billing | Per second | Per minute |
| Best for | Agents, evals, CI and batch jobs | Development, training, notebooks |