Skip to content
r/LocalLLaMA

21 GPUs benchmarked running a small TTS model, with 5GB peak VRAM

21 GPU's benchmarked running a small TTS model (vram peak: 5GB)

A Reddit user rented 21 GPUs on vast.ai to benchmark OmniVoice, a small TTS model with about 5GB peak VRAM, using xRT as the audio generation speed metric and averaging 3 voice-cloning runs with reference audio.

Why it matters: HKR-H/K/R pass: a 21-GPU TTS benchmark with 5GB peak VRAM and 3-run xRT averaging is useful to local-inference builders. Scope is niche, so it sits at the low featured band.

Read the original ↗Export Markdown