Corpus Generation for Voice Command in Smart Home and the Effect of Speech Synthesis on End-to-End SLU

Thierry Desot, François Portet, Michel Vacher


Abstract
Massive amounts of annotated data greatly contributed to the advance of the machine learning field. However such large data sets are often unavailable for novel tasks performed in realistic environments such as smart homes. In this domain, semantically annotated large voice command corpora for Spoken Language Understanding (SLU) are scarce, especially for non-English languages. We present the automatic generation process of a synthetic semantically-annotated corpus of French commands for smart-home to train pipeline and End-to-End (E2E) SLU models. SLU is typically performed through Automatic Speech Recognition (ASR) and Natural Language Understanding (NLU) in a pipeline. Since errors at the ASR stage reduce the NLU performance, an alternative approach is End-to-End (E2E) SLU to jointly perform ASR and NLU. To that end, the artificial corpus was fed to a text-to-speech (TTS) system to generate synthetic speech data. All models were evaluated on voice commands acquired in a real smart home. We show that artificial data can be combined with real data within the same training set or used as a stand-alone training corpus. The synthetic speech quality was assessedby comparing it to real data using dynamic time warping (DTW).
Anthology ID:
2020.lrec-1.786
Volume:
Proceedings of the Twelfth Language Resources and Evaluation Conference
Month:
May
Year:
2020
Address:
Marseille, France
Editors:
Nicoletta Calzolari, Frédéric Béchet, Philippe Blache, Khalid Choukri, Christopher Cieri, Thierry Declerck, Sara Goggi, Hitoshi Isahara, Bente Maegaard, Joseph Mariani, Hélène Mazo, Asuncion Moreno, Jan Odijk, Stelios Piperidis
Venue:
LREC
SIG:
Publisher:
European Language Resources Association
Note:
Pages:
6395–6404
Language:
English
URL:
https://aclanthology.org/2020.lrec-1.786
DOI:
Bibkey:
Cite (ACL):
Thierry Desot, François Portet, and Michel Vacher. 2020. Corpus Generation for Voice Command in Smart Home and the Effect of Speech Synthesis on End-to-End SLU. In Proceedings of the Twelfth Language Resources and Evaluation Conference, pages 6395–6404, Marseille, France. European Language Resources Association.
Cite (Informal):
Corpus Generation for Voice Command in Smart Home and the Effect of Speech Synthesis on End-to-End SLU (Desot et al., LREC 2020)
Copy Citation:
PDF:
https://aclanthology.org/2020.lrec-1.786.pdf