Abstract: The trade-off between robustness, latency, and prosody critically challenges text-to-speech (TTS) systems. Autoregressive models, despite fidelity, are slow and error-prone; non-autoregressive (NAR) alternatives, while fast, often sacrifice prosodic naturalness via rigid alignments. This paper introduces StellarTTS, a novel mobile-optimized NAR TTS framework based on a sparse temporal embedding strategy, enabling granular control of phoneme duration, pronunciation, and prosody. Furthermore, we propose a semantic-aware codec that facilitates efficient single-stage decoding. Conditioned on the sparse temporal embedding, our 83M-parameter lightweight masked generative transformer achieves a real-time factor (RTF) of 0.08. Experiments demonstrate that StellarTTS attains lower latency and stronger robustness compared to state-of-the-art TTS systems, while maintaining competitive performance in audio quality, prosodic naturalness, and speaker similarity.
| Prompt | text | StellarTTS (Ours) | CosyVoice | F5-TTS | MaskGCT | ||||||
| 共同建设面向未来的交通,和出行服务新生态。 | |||||||||||
| 据了解,目前公安县已成立专班,妥善处理此事件。 | |||||||||||
| 那松弛的肌肉在她的手中,如同棉花一样软弱无力。 | |||||||||||
| 好吧,我们别别别别别别别耽搁耽搁耽搁耽搁耽搁耽搁时间了,收拾收拾收拾收拾收拾收拾收拾东西,干干干干干干正经事吧。 |