Sihang Nie, Jinxin Ji, Xiaofen Xing, Deyi Tuo, Chengbin Jin, Jialong Mai, Xiangmin Xu
This paper proposes a framework for explicit and decoupled control over word-level acoustic attributes (duration, boundary, energy, pitch, tone) in LLM-based TTS.
State-of-the-art LLM-based TTS systems lack the fine-grained, word-level control necessary for precise stylistic interventions and strict temporal alignment required in applications like audiobook narration and dubbing. This bottleneck is due to the scarcity of fine-grained annotated datasets and the architectural challenge of integrating multi-dimensional control signals into discrete autoregressive generation.
Extensive experiments show that WordVoice achieves superior, decoupled control over multiple acoustic dimensions while maintaining competitive zero-shot synthesis stability. The key contributions are the release of a large-scale word-level annotated dataset and the proposal of a new paradigm for precise control in LLM TTS.