Thunder Compute
on your own GPUs
License Thunder Compute’s orchestration software for your VPC or on-prem cluster. Every team draws GPUs from one shared pool.
node-01
- train
- train
- train
- train
- train
- train
- eval
- idle
node-02
- infer
- infer
- infer
- train
- train
- train
- train
- notebook
node-03
- eval
- eval
- train
- train
- infer
- infer
- notebook
- idle
node-04
- notebook
- notebook
- notebook
- eval
- eval
- infer
- infer
- train
ml-research · inference · evals · notebooks, one pool30/32 busy
Terminal
$ sudo thunder up
saved auth token in /etc/thunder/thunderd.env
started thunderd.service
$ thunder monitor
Zone Peers
roster=1 status=self latency=-
roster=2 status=ok latency=1ms
roster=3 status=ok latency=1ms
● runningthunderd · node-01 · 8 GPUs
How it works
One command per machine
Install the thunder CLI on each GPU server and run thunder up. The machine enrolls, finds its peers and joins the pool.
Drivers handled
thunder ensure-driver installs the NVIDIA driver Thunder needs
Updates itself
Nodes update thunderd when a new release is rolled out
Architecture
No control plane to run
thunderd runs on every machine. The nodes elect one manager and fail over to a standby if it goes down.
Peer to peer
Nodes track each other directly, with no scheduler to host on site
Enrolled with Thunder
Each node registers with Thunder Central to join its zone
thunder central
your VPC or data center
- node-01● manager
- node-02● standby
- node-03● member
- node-04● member
Deployment and security
Your cloud or your data center
Runs inside your VPC or on your own servers. Your workloads stay there, and here is exactly what Thunder sees.
Your VPC or data centeryou run
- Runs on
- cloud GPU instances or your servers
- You manage
- hardware or cloud account, network
- CUDA calls
- client process to GPU server, direct
- Datasets and weights
- on your storage
- Gossip between nodes
- encrypted with your zone’s key
Thunderwhat we do and see
- Setup
- installed with our team
- → Enrollment
- hostname, IP, ports, GPU type, count
- → Check-ins
- node health, on a regular cadence
- → Logs
- thunderd and its components, for support
- ← Updates
- new thunderd releases
Put every GPU in one pool
Tell us about your cluster and we’ll walk you through a deployment.