Skip to main content
Use these guides to launch open models on Thunder Compute, pick a model format that fits the GPU you selected, and expose the model through a local or public API endpoint.
GPU availability, supported GPU counts, vCPU choices, and pricing can change. Use tnr create or the pricing page to confirm the current options before launching a longer run.

Available Guides

How To Choose A Runtime

For Qwen3.6 27B, start with the dedicated guide below. It uses a configuration that was tested on a Thunder Compute RTX A6000 base instance: UD-Q4_K_XL, full GPU offload, 32K context, and a public OpenAI-compatible endpoint.

Run Qwen3.6 27B

Launch Qwen3.6 27B dense with the tested A6000 setup, then expose it through the /v1/chat/completions API.

Run GPT-OSS 120B

Use the existing GPT-OSS guide from the Models section.

Run DeepSeek R1

Use the existing DeepSeek R1 guide from the Models section.