Speaker and Expression Factorization for Audiobook Data: Expressiveness and Transplantation

Chen, L; Braunschweiler, N; Gales, MJF

Speaker and Expression Factorization for Audiobook Data: Expressiveness and Transplantation

Repository URI

https://www.repository.cam.ac.uk/handle/1810/247404

Files

Chen et al 2015 IEEE Transactions on Audio Speech and Language Processing.pdf (765.34 KB)

Type

Article

Authors

Chen, L

Braunschweiler, N

Gales, MJF

Abstract

Expressive synthesis from text is a challenging problem. There are two issues. First, read text is often highly expressive to convey the emotion and scenario in the text. Second, since the expressive training speech is not always available for different speakers, it is necessary to develop methods to share the expressive information over speakers. This paper investigates the approach of using very expressive, highly diverse audiobook data from multiple speakers to build an expressive speech synthesis system. Both of two problems are addressed by considering a factorized framework where speaker and emotion are modelled in separate sub-spaces of a cluster adaptive training (CAT) parametric speech synthesis system. The sub-spaces for the expressive state of a speaker and the characteristics of the speaker are jointly trained using a set of audiobooks. In this work, the expressive speech synthesis system works in two distinct modes. In the first mode, the expressive information is given by audio data and the adaptation method is used to extract the expressive information in the audio data. In the second mode, the input of the synthesis system is plain text and a full expressive synthesis system is examined where the expressive state is predicted from the text. In both modes, the expressive information is shared and transplanted over different speakers. Experimental results show that in both modes, the expressive speech synthesis method proposed in this work significantly improves the expressiveness of the synthetic speech for different speakers. Finally, this paper also examines whether it is possible to predict the expressive states from text for multiple speakers using a single model, or whether the prediction process needs to be speaker specific.

Keywords

Audiobook, cluster adaptive training, expressive speech synthesis, factorization, hidden Markov model, neural network

Journal Title

IEEE Transactions on Audio, Speech and Language Processing

Journal ISSN

1558-7916
2329-9304

Volume Title

23

Publisher

Institute of Electrical and Electronics Engineers (IEEE)

Publisher DOI

https://doi.org/10.1109/TASLP.2014.2385478

Rights

http://www.rioxx.net/licenses/all-rights-reserved

Collections

Scholarly Works - Engineering
Symplectic mapped items for data match