Ultrafast Qwen3-TTS at 34 ms Time-to-First-Audio, Handling 10 Requests Per Second [OSS]
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| Hey locallama! We recently open sourced a Qwen3-TTS 1.7B implementation that achieves 10 requests per second (RPS) and sub-50 ms p95 time-to-first-audio (TTFA) while maintaining real-time playback on 1 x H100. This extends to 20 RPS at sub-100 ms p95 TTFA. By adjusting settings, you can get a 4090 to perform at ~50 ms p95 TTFA as well. We were frustrated with locally runnable models having slow response speeds - even slower than some cloud ones (which include network times). We achieve a big speedup compared to popular engines such as vLLM-Omni and SGLang-Omni. We open source the implementation and benchmark. Our methodology is explained in our blog! Hope you enjoy and would love to get feedback :) [link] [comments] |
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.