通过声学和文本上下文改善有声读物文本到语音综合的语音韵律

论文标题

通过声学和文本上下文改善有声读物文本到语音综合的语音韵律

Improving Speech Prosody of Audiobook Text-to-Speech Synthesis with Acoustic and Textual Contexts

论文作者

Xin, Detai, Adavanne, Sharath, Ang, Federico, Kulkarni, Ashish, Takamichi, Shinnosuke, Saruwatari, Hiroshi

论文摘要

储层计算是预测湍流的有力工具，其简单的架构具有处理大型系统的计算效率。然而，其实现通常需要完整的状态向量测量和系统非线性知识。我们使用非线性投影函数将系统测量扩展到高维空间，然后将其输入到储层中以获得预测。我们展示了这种储层计算网络在时空混沌系统上的应用，该系统模拟了湍流的若干特征。我们表明，使用径向基函数作为非线性投影器，即使只有部分观测并且不知道控制方程，也能稳健地捕捉复杂的系统非线性。最后，我们表明，当测量稀疏、不完整且带有噪声，甚至控制方程变得不准确时，我们的网络仍然可以产生相当准确的预测，从而为实际湍流系统的无模型预测铺平了道路。

We present a multi-speaker Japanese audiobook text-to-speech (TTS) system that leverages multimodal context information of preceding acoustic context and bilateral textual context to improve the prosody of synthetic speech. Previous work either uses unilateral or single-modality context, which does not fully represent the context information. The proposed method uses an acoustic context encoder and a textual context encoder to aggregate context information and feeds it to the TTS model, which enables the model to predict context-dependent prosody. We conducted comprehensive objective and subjective evaluations on a multi-speaker Japanese audiobook dataset. Experimental results demonstrate that the proposed method significantly outperforms two previous works. Additionally, we present insights about the different choices of context - modalities, lateral information and length - for audiobook TTS that have never been discussed in the literature before.

下载PDF全文

下载文献需遵守相关版权规定

论文标题