Use case · voice AI
Save up to 90% on voice AI inference.
Speech-to-text, text-to-speech, translation, and captioning are a near-perfect fit for consumer GPUs: high volume, small models, and latency budgets measured in seconds. Managed APIs charge for the opposite.
47,638minutes transcribed per dollar, Parakeet TDT 1.1B
6M+TTS words per dollar, OpenVoice
230words per second, OpenVoice on RTX 3080 Ti
$1,260to transcribe 1,000,000 hours of audio
Speech-to-text
A 1,000-fold cost reduction against popular APIs.
Self-managed Whisper and Parakeet on SaladCloud transcribe for a fraction of a cent per hour of audio. If you would rather not run the model, the Salad Transcription API gives you the top benchmark accuracy from $0.16 per audio hour.
- Parakeet TDT 1.1B: 47,638 minutes per dollar
- Distil-Whisper Large v2: 29,994 minutes per dollar
- Whisper Large v3: 11,736 minutes per dollar
- Around 60x real-time on RTX 3090s
Text-to-speech
Millions of words per dollar.
RTX and GTX GPUs deliver the best speed-to-cost ratio for TTS inference. Older, cheaper cards often win on cost per word.
- OpenVoice: almost 6,000,000 words per dollar on RTX 2070 and GTX 1650 class GPUs
- OpenVoice: 230 words per second on RTX 3080 Ti at $0.20/hr, the best speed-to-cost ratio
- Bark: 39,000 words per dollar on RTX 3060 and GTX 1060
- XTTS-v2, MetaVoice, and custom voices via your own container
Benchmarks and guides