Ship on distributed GPUs without learning a new stack.
A Docker image and an API key are all you need. Everything below, from your first deployment to production patterns for interruptible hardware, is written for engineers who want to get to running instances quickly.
# Python SDK from salad_cloud_sdk import SaladCloudSdk sdk = SaladCloudSdk(api_key="...") groups = sdk.container_groups.list_container_groups( organization_name="acme", project_name="prod") for g in groups.items: print(g.name, g.current_state.status, g.replicas)
Your first deployment.
API reference
REST API with Python, JavaScript / TypeScript, Go, and Java SDKs, plus an OpenAPI spec.
Reference →Build for a fleet that churns.
Individual nodes run at 90–95% reliability. These patterns turn that into a system that is more reliable than any single machine.
Real-time inference
Container Gateway for HTTP with TLS, load balancing, and health-probe-gated routing. Run at least three replicas.
Guide →Long-running tasks
Job Queues, RabbitMQ, SQS, or the open-source Kelpie API with checkpointing so interruption costs seconds, not hours.
Guide →High-performance storage
Patterns for model weights, datasets, and checkpoints on a distributed network, and what's coming with Salad storage products.
Guide →How-tos, benchmarks, and case studies.
Where the code lives.
Before your first deploy.
How does SaladCloud work?
Providers share spare compute time; some full-time, some a couple of hours a day. Nodes behave like spot instances: interruptible, bid for at different priority levels. You specify image, GPU class, and replicas; SaladCloud fills the request from available, compatible nodes and keeps it filled.
What are the unique traits I should design for?
Longer cold starts than a data center, interruption without warning, and network performance that varies by node. Build for amd64, use IPv6 for the Container Gateway, run at least three replicas, and retry failed requests.
What about performance and offline nodes?
Every GPU is tested before joining and scored continuously by the trust rating. When a node goes offline, SCE reallocates the workload to another of the same class automatically.
Which GPUs and CUDA versions?
NVIDIA RTX-class GPUs from the 3060 to the 5090, plus selected AMD classes. The RTX 5090 requires CUDA 12.8 or later; maintain a separate image for older GPUs if you target both.