Use case · voice AI

Save up to 90% on voice AI inference.

Speech-to-text, text-to-speech, translation, and captioning are a near-perfect fit for consumer GPUs: high volume, small models, and latency budgets measured in seconds. Managed APIs charge for the opposite.

Bar chart of minutes of audio per dollar: Salad 375, Deepgram 231, Azure 167, Assembly AI 162, Google 62, Amazon 42
47,638minutes transcribed per dollar, Parakeet TDT 1.1B
6M+TTS words per dollar, OpenVoice
230words per second, OpenVoice on RTX 3080 Ti
$1,260to transcribe 1,000,000 hours of audio
Speech-to-text

A 1,000-fold cost reduction against popular APIs.

Self-managed Whisper and Parakeet on SaladCloud transcribe for a fraction of a cent per hour of audio. If you would rather not run the model, the Salad Transcription API gives you the top benchmark accuracy from $0.16 per audio hour.

  • Parakeet TDT 1.1B: 47,638 minutes per dollar
  • Distil-Whisper Large v2: 29,994 minutes per dollar
  • Whisper Large v3: 11,736 minutes per dollar
  • Around 60x real-time on RTX 3090s
Text-to-speech

Millions of words per dollar.

RTX and GTX GPUs deliver the best speed-to-cost ratio for TTS inference. Older, cheaper cards often win on cost per word.

  • OpenVoice: almost 6,000,000 words per dollar on RTX 2070 and GTX 1650 class GPUs
  • OpenVoice: 230 words per second on RTX 3080 Ti at $0.20/hr, the best speed-to-cost ratio
  • Bark: 39,000 words per dollar on RTX 3060 and GTX 1060
  • XTTS-v2, MetaVoice, and custom voices via your own container
Benchmarks and guides

Read the methodology.

Speech-to-text: 47,638 minutes/$

Parakeet TDT 1.1B on SaladCloud.

Read →

Speech-to-text: 29,994 minutes/$

Distil-Whisper Large v2.

Read →

Speech-to-text: 11,736 minutes/$

Whisper Large v3.

Read →

Text-to-speech: 6M+ words/$

OpenVoice across 20+ GPU classes.

Read →

Text-to-speech: 39,000 words/$

Bark TTS benchmark.

Read →

Your own ChatGPT

Ollama, ChatUI, and Salad for voice-enabled assistants.

Read →

Stop renting data-center GPUs for voice.