EDBT 2026 Demo / reviewers in the wild / expert
Alan W. Black
dblp:b/AlanWBlack
· DBLP profile ↗
236ranked-venue papers
23as first author
33since 2021 · last 2025
0000-0001-8820-8831ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 182 · 17 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 160 · 17 first-author · 25 since 2021Human-computer interaction and ubiquitous computing · 4Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Sylber: Syllabic Embedding Representation of Speech from Raw AudioabstractSyllables are compositional units of spoken language that efficiently structure human speech perception and production. However, current neural speech representations lack such structure, resulting in dense token sequences that are costly to process. To bridge this gap, we propose a new model, Sylber, that produces speech representations with clean and robust syllabic structure. Specifically, we propose a self-supervised learning (SSL) framework that bootstraps syllabic embeddings by distilling from its own initial unsupervised syllabic segmentation. This results in a highly structured representation of speech features, offering three key benefits: 1) a fast, linear-time syllable segmentation algorithm, 2) efficient syllabic tokenization with an average of 4.27 tokens per second, and 3) novel phonological units suited for efficient spoken language modeling. Our proposed segmentation method is highly robust and generalizes to out-of-domain data and unseen languages without any tuning. By training token-to-speech generative models, fully intelligible speech can be reconstructed from Sylber tokens with a significantly lower bitrate than baseline SSL tokens. This suggests that our model effectively compresses speech into a compact sequence of tokens with minimal information loss. Lastly, we demonstrate that categorical perception—a linguistic phenomenon in speech perception—emerges naturally in Sylber, making the embedding space more categorical and sparse than previous speech features and thus supporting the high efficiency of our tokenization. Together, we present a novel SSL approach for representing speech as syllables, with significant potential for efficient speech tokenization and spoken language modeling. Cheol Jun Cho, Nicholas Lee, Akshat Gupta, Dhruv Agarwal 0005, Ethan Chen, Alan W. Black, Gopala Krishna Anumanchipalli |
ICLR | 6 |
| 2024 | SD-HuBERT: Sentence-Level Self-Distillation Induces Syllabic Organization in HubertabstractData-driven unit discovery in self-supervised learning (SSL) of speech has embarked on a new era of spoken language processing. Yet, the discovered units often remain in phonetic space and speech units beyond phonemes are largely underexplored. Here, we demonstrate that a syllabic organization emerges in learning sentence-level representation of speech. In particular, we adopt "self-distillation" objective to fine-tune the pretrained HuBERT with an aggregator token that summarizes the entire sentence. Without any supervision, the resulting model draws definite boundaries in speech, and the representations across frames exhibit salient syllabic structures. We demonstrate that this emergent structure largely corresponds to the ground truth syllables. Furthermore, we propose a new benchmark task, Spoken Speech ABX, for evaluating sentence-level representation of speech. When compared to previous models, our model outperforms in both unsupervised syllable discovery and learning sentence-level representation. Together, we demonstrate that the self-distillation of HuBERT gives rise to syllabic organization without relying on external labels or modalities, and potentially provides novel data-driven units for spoken language modeling. Cheol Jun Cho, Abdel-rahman Mohamed, Shang-Wen Li 0001, Alan W. Black, Gopala Krishna Anumanchipalli |
ICASSP | 4 |
| 2024 | Self-Supervised Models of Speech Infer Universal Articulatory KinematicsabstractSelf-Supervised Learning (SSL) based models of speech have shown remarkable performance on a range of downstream tasks. These state-of-the-art models have remained blackboxes, but many recent studies have begun “probing” models like HuBERT, to correlate their internal representations to different aspects of speech. In this paper, we show “inference of articulatory kinematics” as fundamental property of SSL models, i.e., the ability of these models to transform acoustics into the causal articulatory dynamics underlying the speech signal. We also show that this abstraction is largely overlapping across the language of the data used to train the model, with preference to the language with similar phonological system. Furthermore, we show that with simple affine transformations, Acoustic-to-Articulatory inversion (AAI) is transferrable across speakers, even across genders, languages, and dialects, showing the generalizability of this property. Together, these results shed new light on the internals of SSL models that are critical to their superior performance, and open up new avenues into language-agnostic universal models for speech engineering, that are interpretable and grounded in speech science. Cheol Jun Cho, Abdel-rahman Mohamed, Alan W. Black, Gopala Krishna Anumanchipalli |
ICASSP | 3 |
| 2024 | Towards EMG-to-Speech with Necklace Form Factor
Peter Wu, Ryan Kaveh, Raghav Nautiyal, Christine Zhang, Albert Guo, Anvitha Kachinthaya, Tavish Mishra, Bohan Yu, Alan W. Black, Rikky Muller, Gopala Krishna Anumanchipalli |
INTERSPEECH | 9 |
| 2023 | CTC Alignments Improve Autoregressive TranslationabstractBrian Yan, Siddharth Dalmia, Yosuke Higuchi, Graham Neubig, Florian Metze, Alan W Black, Shinji Watanabe. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. Brian Yan, Siddharth Dalmia, Yosuke Higuchi, Graham Neubig, Florian Metze, Alan W. Black, Shinji Watanabe 0001 |
EACL | 6 |
| 2023 | Articulatory Representation Learning via Joint Factor Analysis and Neural Matrix FactorizationabstractArticulatory representation learning is the fundamental research in modeling neural speech production system. Our previous work has established a deep paradigm to decompose the articulatory kinematics data into gestures, which explicitly model the phonological and linguistic structure encoded with human speech production mechanism, and corresponding gestural scores. We continue with this line of work by raising two concerns: (1) The articulators are entangled together in the original algorithm such that some of the articulators do not leverage effective moving patterns, which limits the interpretability of both gestures and gestural scores; (2) The EMA data is sparsely sampled from articulators, which limits the intelligibility of learned representations. In this work, we propose a novel articulatory representation decomposition algorithm that takes the advantage of guided factor analysis to derive the articulatory-specific factors and factor scores. A neural convolutive matrix factorization algorithm is then employed on the factor scores to derive the new gestures and gestural scores. We experiment with the rtMRI corpus that captures the fine-grained vocal tract contours. Both subjective and objective evaluation results suggest that the newly proposed system delivers the articulatory representations that are intelligible, generalizable, efficient and interpretable. Jiachen Lian, Alan W. Black, Yijing Lu, Louis Goldstein, Shinji Watanabe 0001, Gopala Krishna Anumanchipalli |
ICASSP | 2 |
| 2023 | A Fast and Accurate Pitch Estimation Algorithm Based on the Pseudo Wigner-Ville DistributionabstractEstimation of fundamental frequency (F0) in voiced segments of speech signals, also known as pitch tracking, plays a crucial role in pitch synchronous speech analysis, speech synthesis, and speech manipulation. In this paper, we capitalize on the high time and frequency resolution of the pseudo Wigner-Ville distribution (PWVD) and propose a new PWVD-based pitch estimation method. We devise an efficient algorithm to compute PWVD faster and use cepstrum-based pre-filtering to avoid cross-term interference. Evaluating our approach on databases with speech and electroglottograph (EGG) recordings yields a state-of-the-art mean absolute error (MAE) of around 4Hz. Our approach is also effective at voiced/unvoiced classification and handling sudden frequency changes. Yisi Liu, Peter Wu, Alan W. Black, Gopala Krishna Anumanchipalli |
ICASSP | 3 |
| 2023 | Speaker-Independent Acoustic-to-Articulatory Speech InversionabstractTo build speech processing methods that can handle speech as naturally as humans, researchers have explored multiple ways of building an invertible mapping from speech to an interpretable space. The articulatory space is a promising inversion target, since this space captures the mechanics of speech production. To this end, we build an acoustic-to-articulatory inversion (AAI) model that leverages autoregression, adversarial training, and self supervision to generalize to unseen speakers. Our approach obtains 0.784 correlation on an electromagnetic articulography (EMA) dataset, improving the state-of-the-art by 12.5%. Additionally, we show the interpretability of these representations through directly com-paring the behavior of estimated representations with speech production behavior. Finally, we propose a resynthesis-based AAI evaluation metric that does not rely on articulatory labels, demonstrating its efficacy with an 18-speaker dataset. Peter Wu, Cheol Jun Cho, Shinji Watanabe 0001, Louis Goldstein, Alan W. Black, Gopala Krishna Anumanchipalli |
ICASSP | 6 |
| 2023 | Deep Speech Synthesis from MRI-Based Articulatory Representations
Peter Wu, Tingle Li, Yijing Lu, Yubin Zhang, Jiachen Lian, Alan W. Black, Louis Goldstein, Shinji Watanabe 0001, Gopala Krishna Anumanchipalli |
INTERSPEECH | 6 |
| 2022 | ESPnet-SLU: Advancing Spoken Language Understanding Through ESPnetabstractAs Automatic Speech Processing (ASR) systems are getting better, there is an increasing interest of using the ASR output to do downstream Natural Language Processing (NLP) tasks. However, there are few open source toolkits that can be used to generate reproducible results on different Spoken Language Understanding (SLU) benchmarks. Hence, there is a need to build an open source standard that can be used to have a faster start into SLU research. We present ESPnet-SLU, which is designed for quick development of spoken language understanding in a single framework. ESPnet-SLU is a project inside end-to-end speech processing toolkit, ESPnet, which is a widely used open-source standard for various speech processing tasks like ASR, Text to Speech (TTS) and Speech Translation (ST). We enhance the toolkit to provide implementations for various SLU benchmarks that enable researchers to seamlessly mix-and-match different ASR and NLU models. We also provide pretrained models with intensively tuned hyper-parameters that can match or even outperform the current state-of-the-art performances. The toolkit is publicly available at https://github.com/espnet/espnet. Siddhant Arora, Siddharth Dalmia, Pavel Denisov, Xuankai Chang, Yushi Ueda, Yifan Peng 0003, Yuekai Zhang, Sujay Kumar, Karthik Ganesan 0003, Brian Yan, Ngoc Thang Vu, Alan W. Black, Shinji Watanabe 0001 |
ICASSP | 12 |
| 2022 | End-to-End Speech Summarization Using Restricted Self-AttentionabstractSpeech summarization is typically performed by using a cascade of speech recognition and text summarization models. End-to-end modeling of speech summarization models is challenging due to memory and compute constraints arising from long input audio sequences. Recent work in document summarization has inspired methods to reduce the complexity of self-attentions, which enables transformer models to handle long sequences. In this work, we introduce a single model optimized end-to-end for speech summarization. We apply the restricted self-attention technique from text-based models to speech models to address the memory and compute constraints. We demonstrate that the proposed model learns to directly summarize speech for the How-2 corpus of instructional videos. The proposed end-to-end model outperforms the previously proposed cascaded model by 3 points absolute on ROUGE. Further, we consider the spoken language understanding task of predicting concepts from speech inputs and show that the proposed end-to-end model outperforms the cascade model by 4 points absolute F-1. Shruti Palaskar, Alan W. Black, Florian Metze |
ICASSP | 3 |
| 2022 | Two-Pass Low Latency End-to-End Spoken Language Understanding
Siddhant Arora, Siddharth Dalmia, Xuankai Chang, Brian Yan, Alan W. Black, Shinji Watanabe 0001 |
INTERSPEECH | 5 |
| 2022 | ASR2K: Speech Recognition for Around 2000 Languages without Audio
Florian Metze, David R. Mortensen, Alan W. Black, Shinji Watanabe 0001 |
INTERSPEECH | 4 |
| 2022 | Deep Neural Convolutive Matrix Factorization for Articulatory Representation DecompositionabstractMost of the research on data-driven speech representation learning has focused on raw audios in an end-to-end manner, paying little attention to their internal phonological or gestural structure.This work, investigating the speech representations derived from articulatory kinematics signals, uses a neural implementation of convolutive sparse matrix factorization to decompose the articulatory data into interpretable gestures and gestural scores.By applying sparse constraints, the gestural scores leverage the discrete combinatorial properties of phonological gestures.Phoneme recognition experiments were additionally performed to show that gestural scores indeed code phonological information successfully.The proposed work thus makes a bridge between articulatory phonology and deep neural networks to leverage informative, intelligible, interpretable,and efficient speech representations. Jiachen Lian, Alan W. Black, Louis Goldstein, Gopala Krishna Anumanchipalli |
INTERSPEECH | 2 |
| 2022 | Building African Voices
Perez Ogayo, Graham Neubig, Alan W. Black |
INTERSPEECH | 3 |
| 2022 | Deep Speech Synthesis from Articulatory Representations
Peter Wu, Shinji Watanabe 0001, Louis Goldstein, Alan W. Black, Gopala Krishna Anumanchipalli |
INTERSPEECH | 4 |
| 2022 | Intent classification using pre-trained language agnostic embeddings for low resource languagesabstractBuilding Spoken Language Understanding (SLU) systems that do not rely on language specific Automatic Speech Recognition (ASR) is an important yet less explored problem in language processing.In this paper, we present a comparative study aimed at employing a pre-trained language agnostic acoustic model to perform SLU in low resource scenarios.Specifically, we use three different embedding settings extracted using Allosaurus, a pre-trained universal phone decoder: (1) Phonelabels (2) Panphone, and (3) Allo embeddings (proposed by us).These embeddings are then used in identifying the spoken intent.We perform experiments across three different languages: English, Sinhala, and Tamil each with different data sizes to simulate high, medium, and low resource scenarios.Our system improves on the state-of-the-art (SOTA) intent classification accuracy by absolute 2.11% for Sinhala and 7.00% for Tamil and achieves competitive results in English.Furthermore, we also present a quantitative analysis to show how the performance scales with the number of training examples. Hemant Yadav, Akshat Gupta, Sai Krishna Rallabandi, Alan W. Black, Rajiv Ratn Shah |
INTERSPEECH | 4 |
| 2022 | Phone Inventories and Recognition for Every LanguageabstractIdentifying phone inventories is a crucial component in language documentation and the preservation of endangered languages. However, even the largest collection of phone inventory only covers about 2000 languages, which is only 1/4 of the total number of languages in the world. A majority of the remaining languages are endangered. In this work, we attempt to solve this problem by estimating the phone inventory for any language listed in Glottolog, which contains phylogenetic information regarding 8000 languages. In particular, we propose one probabilistic model and one non-probabilistic model, both using phylogenetic trees (“language family trees”) to measure the distance between languages. We show that our best model outperforms baseline models by 6.5 F1. Furthermore, we demonstrate that, with the proposed inventories, the phone recognition model can be customized for every language in the set, which improved the PER (phone error rate) in phone recognition by 25%. Florian Metze, David R. Mortensen, Alan W. Black, Shinji Watanabe 0001 |
LREC | 4 |
| 2021 | Breaking Down Walls of Text: How Can NLP Benefit Consumer Privacy?abstractAbhilasha Ravichander, Alan W Black, Thomas Norton, Shomir Wilson, Norman Sadeh. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Abhilasha Ravichander, Alan W. Black, Thomas B. Norton, Shomir Wilson, Norman M. Sadeh |
ACL/IJCNLP (1) | 2 |
| 2021 | Intent Recognition and Unsupervised Slot Identification for Low-Resourced Spoken Dialog SystemsabstractIntent Recognition and Slot Identification are crucial components in spoken language understanding (SLU) systems. In this paper, we present a novel approach towards both these tasks in the context of low-resourced and unwritten languages. We use an acoustic based SLU system that converts speech to its phonetic transcription using a universal phone recognition system. We build a word-free natural language understanding module that does intent recognition and slot identification from these phonetic transcription. Our proposed SLU system performs competitively for resource rich scenarios and significantly outperforms existing approaches as the amount of available data reduces. We train both recurrent and transformer based neural networks and test our system on five natural speech datasets in five different languages. We observe more than 10% improvement for intent classification in Tamil and more than 5% improvement for intent classification in Sinhala. Additionally, we present a novel approach towards unsupervised slot identification using normalized attention scores. This approach can be used for unsupervised slot labelling, data augmentation and to generate data for a new slot in a one-shot way with only one speech recording. Akshat Gupta, Olivia Deng, Akruti Kushwaha, Saloni Mittal, William Zeng, Sai Krishna Rallabandi, Alan W. Black |
ASRU | 7 |
| 2021 | Towards Using Heterogeneous Relation Graphs for End-to-End TTSabstractNeural models for end-to-end text-to-speech (TTS) synthe-sis are increasingly outperforming traditional approaches in statistical parametric speech synthesis. Speech generation in these neural models predominantly relies on using free-form text as the input modality. However, the earlier statistical parametric models were built on encoded phonetic and syn-tactic features. In this work, we explore the possibility of explicitly feeding deterministic linguistic structure to a neural TTS system in the form of Heterogeneous Relational Graphs (HRGs), an expressive formalism capable of representing pho-netic and syntactic information. Specifically, we use Graph Convolutional Networks to learn structurally informed contin-uous representations of the HRGs, which can be seamlessly passed to the encoders of popular neural TTS models like TransformerTTS or Tacotron. Furthermore, our simple HRG based text-to-speech synthesis leverages the syntactic bias in HRGs as demonstrated by improvements in automated met-rics and human evaluation on i) the single speaker dataset LJSpeech; ii) the multi-speaker dataset Arctic; and iii) out-of-domain test sets from the Blizzard challenge.11The code, trained models, and our dataset of HRGs will be released at https://github.com/ars22/GraphNeuralTTS/. Amrith Setlur, Aman Madaan, Tanmay Parekh, Yiming Yang 0002, Alan W. Black |
ASRU | 5 |
| 2021 | Cross-Lingual Transfer for Speech Processing Using Acoustic Language SimilarityabstractSpeech processing systems currently do not support the vast majority of languages, in part due to the lack of data in low-resource languages. Cross-lingual transfer offers a compelling way to help bridge this digital divide by incorporating high-resource data into low-resource systems. Current cross-lingual algorithms have shown success in text-based tasks and speech-related tasks over some low-resource languages. However, scaling up speech systems to support hundreds of low-resource languages remains unsolved. To help bridge this gap, we propose a language similarity approach that can efficiently identify acoustic cross-lingual transfer pairs across hundreds of languages. We demonstrate the effectiveness of our approach in language family classification, speech recognition, and speech synthesis tasks. Peter Wu, Jiatong Shi, Yifan Zhong, Shinji Watanabe 0001, Alan W. Black |
ASRU | 5 |
| 2021 | NoiseQA: Challenge Set Evaluation for User-Centric Question AnsweringabstractAbhilasha Ravichander, Siddharth Dalmia, Maria Ryskina, Florian Metze, Eduard Hovy, Alan W Black. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Abhilasha Ravichander, Siddharth Dalmia, Maria Ryskina, Florian Metze, Eduard H. Hovy, Alan W. Black |
EACL | 6 |
| 2021 | Acoustics Based Intent Recognition Using Discovered Phonetic Units for Low Resource LanguagesabstractWith recent advancements in language technologies, humans are now speaking to devices. Increasing the reach of spoken language technologies requires building systems in local languages. A major bottleneck here are the underlying data-intensive parts that make up such systems, including automatic speech recognition (ASR) systems that require large amounts of labelled data. With the aim of aiding development of spoken dialog systems in low resourced languages, we propose a novel acoustics based intent recognition system that uses discovered phonetic units for intent classification. The system is made up of two blocks - the first block is a universal phone recognition system that generates a transcript of discovered phonetic units for the input audio, and the second block performs intent classification from the generated phonetic transcripts. We propose a CNN+LSTM based architecture and present results for two languages families - Indic languages and Romance languages, for two different intent recognition tasks. We also perform multilingual training of our intent classifier and show improved cross-lingual transfer and zero-shot performance on an unknown language within the same language family. Akshat Gupta, Sai Krishna Rallabandi, Alan W. Black |
ICASSP | 4 |
| 2021 | Phone Distribution Estimation for Low Resource Languages
Juncheng Li 0001, Jiali Yao, Alan W. Black, Florian Metze |
ICASSP | 4 |
| 2021 | Multilingual Phonetic Dataset for Low Resource Speech RecognitionabstractPhone Recognition is one of the most important tasks in the field of multilingual speech recognition, especially for low-resource languages whose orthographies are not available. However, most speech recognition datasets so far only focus on high-resource languages, there are very few datasets available for low-resource languages, especially datasets with detailed phone annotation. In this work, we present a large multilingual phonetic dataset, which is preprocessed and aligned from the UCLA phonetic dataset. The dataset contains around 100 low-resource languages and 7000 utterances in total. This dataset would provide an ideal training/evaluation set for universal phone recognition. David R. Mortensen, Florian Metze, Alan W. Black |
ICASSP | 4 |
| 2021 | DialoGraph: Incorporating Interpretable Strategy-Graph Networks into Negotiation Dialogues
Rishabh Joshi, Vidhisha Balachandran, Shikhar Vashishth, Alan W. Black, Yulia Tsvetkov |
ICLR | 4 |
| 2021 | Rethinking End-to-End Evaluation of Decomposable Tasks: A Case Study on Spoken Language UnderstandingabstractDecomposable tasks are complex and comprise of a hierarchy of sub-tasks. Spoken intent prediction, for example, combines automatic speech recognition and natural language understanding. Existing benchmarks, however, typically hold out examples for only the surface-level sub-task. As a result, models with similar performance on these benchmarks may have unobserved performance differences on the other sub-tasks. To allow insightful comparisons between competitive end-to-end architectures, we propose a framework to construct robust test sets using coordinate ascent over sub-task specific utility functions. Given a dataset for a decomposable task, our method optimally creates a test set for each sub-task to individually assess sub-components of the end-to-end model. Using spoken language understanding as a case study, we generate new splits for the Fluent Speech Commands and Snips SmartLights datasets. Each split has two test sets: one with held-out utterances assessing natural language understanding abilities, and one with held-out speakers to test speech processing skills. Our splits identify performance gaps up to 10% between end-to-end systems that were within 1% of each other on the original test sets. These performance gaps allow more realistic and actionable comparisons between different architectures, driving future model development. We release our splits and tools for the community. Siddhant Arora, Alissa Ostapenko, Vijay Viswanathan 0002, Siddharth Dalmia, Florian Metze, Shinji Watanabe 0001, Alan W. Black |
Interspeech | 7 |
| 2021 | Hierarchical Phone Recognition with Compositional Phonetics
Juncheng Li 0001, Florian Metze, Alan W. Black |
Interspeech | 4 |
| 2021 | Multimodal Speech Summarization Through Semantic Concept Learning
Shruti Palaskar, Ruslan Salakhutdinov, Alan W. Black, Florian Metze |
Interspeech | 3 |
| 2021 | Case Study: Deontological Ethics in NLPabstractShrimai Prabhumoye, Brendon Boldt, Ruslan Salakhutdinov, Alan W Black. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Shrimai Prabhumoye, Brendon Boldt, Ruslan Salakhutdinov, Alan W. Black |
NAACL-HLT | 4 |
| 2021 | Focused Attention Improves Document-Grounded GenerationabstractShrimai Prabhumoye, Kazuma Hashimoto, Yingbo Zhou, Alan W Black, Ruslan Salakhutdinov. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Shrimai Prabhumoye, Kazuma Hashimoto, Yingbo Zhou 0002, Alan W. Black, Ruslan Salakhutdinov |
NAACL-HLT | 4 |
| 2021 | Towards Automatic Route Description Unification in Spoken Dialog SystemsabstractIn telephone-based dialog navigation systems, scheduling and direction information are typically collected from routing APIs in text, and then delivered to users via speech. These systematic directions may be augmented with human descriptions to provide more accurate and personalized routes and cover broader user needs. However, manually collecting, transcribing, correcting, and rewriting human descriptions is time-consuming. Also its inconsistency with systematic directions can be confusing to users when delivered orally. This paper describes the construction of a pipeline to automate the route description unification process which also renders the resulting direction delivery more concise and consistent. Yulan Feng, Alan W. Black, Maxine Eskénazi |
SLT | 2 |
| 2020 | Towards Minimal Supervision BERT-Based Grammar Error Correction (Student Abstract)abstractCurrent grammatical error correction (GEC) models typically consider the task as sequence generation, which requires large amounts of annotated data and limit the applications in data-limited settings. We try to incorporate contextual information from pre-trained language model to leverage annotation and benefit multilingual scenarios. Results show strong potential of Bidirectional Encoder Representations from Transformers (BERT) in grammatical error correction task. Yiyuan Li, Antonios Anastasopoulos, Alan W. Black |
AAAI | 3 |
| 2020 | Towards Zero-Shot Learning for Automatic Phonemic TranscriptionabstractAutomatic phonemic transcription tools are useful for low-resource language documentation. However, due to the lack of training sets, only a tiny fraction of languages have phonemic transcription tools. Fortunately, multilingual acoustic modeling provides a solution given limited audio training data. A more challenging problem is to build phonemic transcribers for languages with zero training data. The difficulty of this task is that phoneme inventories often differ between the training languages and the target language, making it infeasible to recognize unseen phonemes. In this work, we address this problem by adopting the idea of zero-shot learning. Our model is able to recognize unseen phonemes in the target language without any training data. In our model, we decompose phonemes into corresponding articulatory attributes such as vowel and consonant. Instead of predicting phonemes directly, we first predict distributions over articulatory attributes, and then compute phoneme distributions with a customized acoustic model. We evaluate our model by training it using 13 languages and testing it using 7 unseen languages. We find that it achieves 7.7% better phoneme error rate on average over a standard multilingual model. Siddharth Dalmia, David R. Mortensen, Juncheng Li 0001, Alan W. Black, Florian Metze |
AAAI | 5 |
| 2020 | ClarQ: A large-scale and diverse dataset for Clarification Question GenerationabstractQuestion answering and conversational systems are often baffled and need help clarifying certain ambiguities.However, limitations of existing datasets hinder the development of large-scale models capable of generating and utilising clarification questions.In order to overcome these limitations, we devise a novel bootstrapping framework (based on self-supervision) that assists in the creation of a diverse, large-scale dataset of clarification questions based on post-comment tuples extracted from stackexchange.The framework utilises a neural network based architecture for classifying clarification questions.It is a two-step method where the first aims to increase the precision of the classifier and second aims to increase its recall.We quantitatively demonstrate the utility of the newly created dataset by applying it to the downstream task of question-answering.The final dataset, ClarQ, consists of ∼2M examples distributed across 173 domains of stackexchange.We release this dataset 1 in order to foster research into the field of clarification question generation with the larger goal of enhancing dialog and question answering systems. Vaibhav Kumar, Alan W. Black |
ACL | 2 |
| 2020 | Politeness Transfer: A Tag and Generate ApproachabstractAman Madaan, Amrith Setlur, Tanmay Parekh, Barnabas Poczos, Graham Neubig, Yiming Yang, Ruslan Salakhutdinov, Alan W Black, Shrimai Prabhumoye. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. Aman Madaan, Amrith Setlur, Tanmay Parekh, Barnabás Póczos, Graham Neubig, Yiming Yang 0002, Ruslan Salakhutdinov, Alan W. Black, Shrimai Prabhumoye |
ACL | 8 |
| 2020 | Topological Sort for Sentence OrderingabstractSentence ordering is the task of arranging the sentences of a given text in the correct order.Recent work using deep neural networks for this task has framed it as a sequence prediction problem.In this paper, we propose a new framing of this task as a constraint solving problem and introduce a new technique to solve it.Additionally, we propose a human evaluation for this task.The results on both automatic and human metrics across four different datasets show that this new technique is better at capturing coherence in documents. Shrimai Prabhumoye, Ruslan Salakhutdinov, Alan W. Black |
ACL | 3 |
| 2020 | Phone Features Improve Speech TranslationabstractEnd-to-end models for speech translation (ST) more tightly couple speech recognition (ASR) and machine translation (MT) than a traditional cascade of separate ASR and MT models, with simpler model architectures and the potential for reduced error propagation.Their performance is often assumed to be superior, though in many conditions this is not yet the case.We compare cascaded and end-to-end models across high, medium, and low-resource conditions, and show that cascades remain stronger baselines.Further, we introduce two methods to incorporate phone features into ST models.We show that these features improve both architectures, closing the gap between end-to-end models and cascades, and outperforming previous academic work -by up to 9 BLEU on our low-resource setting. Elizabeth Salesky, Alan W. Black |
ACL | 2 |
| 2020 | A Corpus for Large-Scale Phonetic TypologyabstractA major hurdle in data-driven research on typology is having sufficient data in many languages to draw meaningful conclusions. We present VoxClamantis v1.0, the first large-scale corpus for phonetic typology, with aligned segments and estimated phoneme-level labels in 690 readings spanning 635 languages, along with acoustic-phonetic measures of vowels and sibilants. Access to such data can greatly facilitate investigation of phonetic typology at a large scale and across many languages. However, it is non-trivial and computationally intensive to obtain such alignments for hundreds of languages, many of which have few to no resources presently available. We describe the methodology to create our corpus, discuss caveats with current methods and their impact on the utility of this data, and illustrate possible research directions through a series of case studies on the 48 highest-quality readings. Our corpus and scripts are publicly available for non-commercial use at https://voxclamantisproject.github.io. Elizabeth Salesky, Eleanor Chodroff, Tiago Pimentel, Matthew Wiesner, Ryan Cotterell, Alan W. Black, Jason Eisner |
ACL | 6 |
| 2020 | Exploring Controllable Text Generation TechniquesabstractNeural controllable text generation is an important area gaining attention due to its plethora of applications.Although there is a large body of prior work in controllable text generation, there is no unifying theme.In this work, we provide a new schema of the pipeline of the generation process by classifying it into five modules.The control of attributes in the generation process requires modification of these modules.We present an overview of different techniques used to perform the modulation of these modules.We also provide an analysis on the advantages and disadvantages of these techniques.We further pave ways to develop new architectures based on the combination of the modules described in this paper. Shrimai Prabhumoye, Alan W. Black, Ruslan Salakhutdinov |
COLING | 2 |
| 2020 | Understanding Linguistic Accommodation in Code-Switched Human-Machine DialoguesabstractCode-switching is a ubiquitous phenomenon in multilingual communities. Natural language technologies that wish to communicate like humans must therefore adaptively incorporate code-switching techniques when they are deployed in multilingual settings. To this end, we propose a Hindi-English human-machine dialogue system that elicits code-switching conversations in a controlled setting. It uses different code-switching agent strategies to understand how users respond and accommodate to the agent's language choice. Through this system, we collect and release a new dataset CommonDost, comprising of 439 human-machine multilingual conversations. We adapt pre-defined metrics to discover linguistic accommodation from users to agents. Finally, we compare these dialogues with Spanish-English dialogues collected in a similar setting, and analyze the impact of linguistic and socio-cultural factors on code-switching patterns across the two language pairs. Tanmay Parekh, Emily P. Ahn, Yulia Tsvetkov, Alan W. Black |
CoNLL | 4 |
| 2020 | Reading Between the Lines: Exploring Infilling in Visual NarrativesabstractGenerating long form narratives such as stories and procedures from multiple modalities has been a long standing dream for artificial intelligence.In this regard, there is often crucial subtext that is derived from the surrounding contexts.The general seq2seq training methods render the models shorthanded while attempting to bridge the gap between these neighbouring contexts.In this paper, we tackle this problem by using infilling techniques involving prediction of missing steps in a narrative while generating textual descriptions from a sequence of images.We also present a new large scale visual procedure telling (ViPT) dataset with a total of 46,200 procedures and around 340k pairwise images and textual descriptions that is rich in such contextual dependencies.Generating steps using infilling technique demonstrates the effectiveness in visual procedures with more coherent texts.We conclusively show a ME-TEOR score of 27.51 on procedures which is higher than the state-of-the-art on visual storytelling.We also demonstrate the effects of interposing new text with missing images during inference.The code and the dataset will be publicly available at https://visual- narratives.github.io/Visual-Narratives. Khyathi Raghavi Chandu, Ruo-Ping Dong, Alan W. Black |
EMNLP (1) | 3 |
| 2020 | Universal Phone Recognition with a Multilingual Allophone SystemabstractMultilingual models can improve language processing, particularly for low resource situations, by sharing parameters across languages. Multilingual acoustic models, however, generally ignore the difference between phonemes (sounds that can support lexical contrasts in a particular language) and their corresponding phones (the sounds that are actually spoken, which are language independent). This can lead to performance degradation when combining a variety of training languages, as identically annotated phonemes can actually correspond to several different underlying phonetic realizations. In this work, we propose a joint model of both language-independent phone and language-dependent phoneme distributions. In multilingual ASR experiments over 11 languages, we find that this model improves testing performance by 2% phoneme error rate absolute in low-resource conditions. Additionally, because we are explicitly modeling language-independent phones, we can build a (nearly-)universal phone recognizer that, when combined with the PHOIBLE [1] large, manually curated database of phone inventories, can be customized into 2,000 language dependent recognizers. Experiments on two low-resourced indigenous languages, Inuktitut and Tusom, show that our recognizer achieves phone accuracy improvements of more than 17%, moving a step closer to speech recognition for all languages in the world.1 Siddharth Dalmia, Juncheng Li 0001, Matthew Lee 0012, Patrick Littell, Jiali Yao, Antonios Anastasopoulos, David R. Mortensen, Graham Neubig, Alan W. Black, Florian Metze |
ICASSP | 10 |
| 2020 | Augmenting Non-Collaborative Dialog Systems with Explicit Semantic and Strategic Dialog History
Yiheng Zhou, Yulia Tsvetkov, Alan W. Black, Zhou Yu 0005 |
ICLR | 3 |
| 2020 | Style Variation as a Vantage Point for Code-SwitchingabstractCode-Switching (CS) is a common phenomenon observed in several bilingual and multilingual communities, thereby attaining prevalence in digital and social media platforms. This increasing prominence demands the need to model CS languages for critical downstream tasks. A major problem in this domain is the dearth of annotated data and a substantial corpora to train large scale neural models. Generating vast amounts of quality text assists several down stream tasks that heavily rely on language modeling such as speech recognition, text-to-speech synthesis etc,. We present a novel vantage point of CS to be style variations between both the participating languages. Our approach does not need any external annotations such as lexical language ids. It mainly relies on easily obtainable monolingual corpora without any parallel alignment and a limited set of naturally CS sentences. We propose a two-stage generative adversarial training approach where the first stage generates competitive negative examples for CS and the second stage generates more realistic CS sentences. We present our experiments on the following pairs of languages: Spanish-English, Mandarin-English, Hindi-English and Arabic-French. We show that the trends in metrics for generated CS move closer to real CS data in each of the above language pairs through the dual stage training process. We believe this viewpoint of CS as style variations opens new perspectives for modeling various tasks in CS text. Khyathi Raghavi Chandu, Alan W. Black |
INTERSPEECH | 2 |
| 2020 | Nonlinear ISA with Auxiliary Variables for Learning Speech RepresentationsabstractThis paper extends recent work on nonlinear Independent Component Analysis (ICA) by introducing a theoretical framework for nonlinear Independent Subspace Analysis (ISA) in the presence of auxiliary variables. Observed high dimensional acoustic features like log Mel spectrograms can be considered as surface level manifestations of nonlinear transformations over individual multivariate sources of information like speaker characteristics, phonological content etc. Under assumptions of energy based models we use the theory of nonlinear ISA to propose an algorithm that learns unsupervised speech representations whose subspaces are independent and potentially highly correlated with the original non-stationary multivariate sources. We show how nonlinear ICA with auxiliary variables can be extended to a generic identifiable model for subspaces as well while also providing sufficient conditions for the identifiability of these high dimensional subspaces. Our proposed methodology is generic and can be integrated with standard unsupervised approaches to learn speech representations with subspaces that can theoretically capture independent higher order speech signals. We evaluate the gains of our algorithm when integrated with the Autoregressive Predictive Decoding (APC) model by showing empirical results on the speaker verification and phoneme recognition tasks. Amrith Setlur, Barnabás Póczos, Alan W. Black |
INTERSPEECH | 3 |
| 2020 | A Resource for Computational Experiments on MapudungunabstractWe present a resource for computational experiments on Mapudungun, a polysynthetic indigenous language spoken in Chile with upwards of 200 thousand speakers. We provide 142 hours of culturally significant conversations in the domain of medical treatment. The conversations are fully transcribed and translated into Spanish. The transcriptions also include annotations for code-switching and non-standard pronunciations. We also provide baseline results on three core NLP tasks: speech recognition, speech synthesis, and machine translation between Spanish and Mapudungun. We further explore other applications for which the corpus will be suitable, including the study of code-switching, historical orthography change, linguistic structure, and sociological and anthropological studies. Mingjun Duan, Carlos Fasola, Sai Krishna Rallabandi, Rodolfo Vega, Antonios Anastasopoulos, Lori S. Levin, Alan W. Black |
LREC | 7 |
| 2020 | AlloVera: A Multilingual Allophone DatabaseabstractWe introduce a new resource, AlloVera, which provides mappings from 218 allophones to phonemes for 14 languages. Phonemes are contrastive phonological units, and allophones are their various concrete realizations, which are predictable from phonological context. While phonemic representations are language specific, phonetic representations (stated in terms of (allo)phones) are much closer to a universal (language-independent) transcription. AlloVera allows the training of speech recognition models that output phonetic transcriptions in the International Phonetic Alphabet (IPA), regardless of the input language. We show that a “universal” allophone model, Allosaurus, built with AlloVera, outperforms “universal” phonemic models and language-specific models on a speech-transcription task. We explore the implications of this technology (and related technologies) for the documentation of endangered and minority languages. We further explore other applications for which AlloVera will be suitable as it grows, including phonological typology. David R. Mortensen, Patrick Littell, Alexis Michaud, Shruti Rijhwani, Antonios Anastasopoulos, Alan W. Black, Florian Metze, Graham Neubig |
LREC | 7 |
| 2020 | Speech Technology for Unwritten LanguagesabstractSpeech technology plays an important role in our everyday life. Among others, speech is used for human-computer interaction, for instance for information retrieval and on-line shopping. In the case of an unwritten language, however, speech technology is unfortunately difficult to create, because it cannot be created by the standard combination of pre-trained speech-to-text and text-to-speech subsystems. The research presented in this article takes the first steps towards speech technology for unwritten languages. Specifically, the aim of this work was 1) to learn speech-to-meaning representations without using text as an intermediate representation, and 2) to test the sufficiency of the learned representations to regenerate speech or translated text, or to retrieve images that depict the meaning of an utterance in an unwritten language. The results suggest that building systems that go directly from speech-to-meaning and from meaning-to-speech, bypassing the need for text, is possible. Odette Scharenborg, Lucas Ondel Yang, Shruti Palaskar, Philip Arthur, Francesco Ciannella, Mingxing Du, Elin Larsen, Danny Merkx, Rachid Riad, Liming Wang 0003, Emmanuel Dupoux, Laurent Besacier, Alan W. Black, Mark Hasegawa-Johnson, Florian Metze, Graham Neubig, Sebastian Stüker, Pierre Godard, Markus Müller 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 13 |
| 2019 | Storyboarding of Recipes: Grounded Contextual GenerationabstractInformation need of humans is essentially multimodal in nature, enabling maximum exploitation of situated context.We introduce a dataset for sequential procedural (how-to) text generation from images in cooking domain.The dataset consists of 16,441 cooking recipes with 160,479 photos associated with different steps.We setup a baseline motivated by the best performing model in terms of human evaluation for the Visual Story Telling (ViST) task.In addition, we introduce two models to incorporate high level structure learnt by a Finite State Machine (FSM) in neural sequential generation process by: (1) Scaffolding Structure in Decoder (SSiD) (2) Scaffolding Structure in Loss (SSiL).Our best performing model (SSiL) achieves a METEOR score of 0.31, which is an improvement of 0.6 over the baseline model.We also conducted human evaluation of the generated grounded recipes, which reveal that 61% found that our proposed (SSiL) model is better than the baseline model in terms of overall recipes.We also discuss analysis of the output highlighting key important NLP issues for prospective directions. Khyathi Raghavi Chandu, Eric Nyberg, Alan W. Black |
ACL (1) | 3 |
| 2019 | Boosting Dialog Response GenerationabstractNeural models have become one of the most important approaches to dialog response generation.However, they still tend to generate the most common and generic responses in the corpus all the time.To address this problem, we designed an iterative training process and ensemble method based on boosting.We combined our method with different training and decoding paradigms as the base model, including mutual-information-based decoding and reward-augmented maximum likelihood learning.Empirical results show that our approach can significantly improve the diversity and relevance of the responses generated by all base models, backed by objective measurements and human evaluation. Wenchao Du, Alan W. Black |
ACL (1) | 2 |
| 2019 | Exploring Phoneme-Level Speech Representations for End-to-End Speech TranslationabstractPrevious work on end-to-end translation from speech has primarily used frame-level features as speech representations, which creates longer, sparser sequences than text.We show that a naïve method to create compressed phoneme-like speech representations is far more effective and efficient for translation than traditional frame-level speech features.Specifically, we generate phoneme labels for speech frames and average consecutive frames with the same label to create shorter, higher-level source sequences for translation.We see improvements of up to 5 BLEU on both our high and low resource language pairs, with a reduction in training time of 60%.Our improvements hold across multiple data sizes and two language pairs. Elizabeth Salesky, Matthias Sperber, Alan W. Black |
ACL (1) | 3 |
| 2019 | Question Answering for Privacy Policies: Combining Computational and Legal PerspectivesabstractAbhilasha Ravichander, Alan W Black, Shomir Wilson, Thomas Norton, Norman Sadeh. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Abhilasha Ravichander, Alan W. Black, Shomir Wilson, Thomas B. Norton, Norman M. Sadeh |
EMNLP/IJCNLP (1) | 2 |
| 2019 | CMU Wilderness Multilingual Speech DatasetabstractThis paper describes the CMU Wilderness Multilingual Speech Dataset. A dataset of over 700 different languages providing audio, aligned text and word pronunciations. On average each language provides around 20 hours of sentence-lengthed transcriptions. We describe our multi-pass alignment techniques and evaluate the results by building speech synthesizers on the aligned data. Most of the resulting synthesizers are good enough for deployment and use. The tools to do this work are released as open source, and instructions on how to apply such alignment for novel languages are given. Alan W. Black |
ICASSP | 1 |
| 2019 | Phoneme Level Language Models for Sequence Based Low Resource ASRabstractBuilding multilingual and crosslingual models help bring different languages together in a language universal space. It allows models to share parameters and transfer knowledge across languages, enabling faster and better adaptation to a new language. These approaches are particularly useful for low resource languages. In this paper, we propose a phoneme-level language model that can be used multilingually and for crosslingual adaptation to a target language. We show that our model performs almost as well as the monolingual models by using six times fewer parameters, and is capable of better adaptation to languages not seen during training in a low resource scenario. We show that these phoneme-level language models can be used to decode sequence based Connectionist Temporal Classification (CTC) acoustic model outputs to obtain comparable word error rates with Weighted Finite State Transducer (WFST) based decoding in Babel languages. We also show that these phoneme-level language models outperform WFST decoding in various low-resource conditions like adapting to a new language and domain mismatch between training and testing data. Siddharth Dalmia, Alan W. Black, Florian Metze |
ICASSP | 3 |
| 2019 | Learning Disentangled Representation in Latent Stochastic Models: A Case Study with Image CaptioningabstractMultimodal tasks require learning joint representation across modalities. In this paper, we present an approach to employ latent stochastic models for a multimodal task image captioning. Encoder Decoder models with stochastic latent variables are often faced with optimization issues such as latent collapse preventing them from realizing their full potential of rich representation learning and disentanglement. We present an approach to train such models by incorporating joint continuous and discrete representation in the prior distribution. We evaluate the performance of proposed approach on a multitude of metrics against vanilla latent stochastic models. We also perform a qualitative assessment and observe that the proposed approach indeed has the potential to learn composite information and explain novel combinations not seen in the training data. Nidhi Vyas, Sai Krishna Rallabandi, Lalitesh Morishetti, Eduard H. Hovy, Alan W. Black |
ICASSP | 5 |
| 2019 | Bag-of-Acoustic-Words for Mental Health Assessment: A Deep Autoencoding Approach
Wenchao Du, Louis-Philippe Morency, Jeffrey F. Cohn, Alan W. Black |
INTERSPEECH | 4 |
| 2019 | The Zero Resource Speech Challenge 2019: TTS Without TabstractWe present the Zero Resource Speech Challenge 2019, which proposes to build a speech synthesizer without any text or phonetic labels: hence, TTS without T (text-to-speech without text). We provide raw audio for a target voice in an unknown language (the Voice dataset), but no alignment, text or labels. Participants must discover subword units in an unsupervised way (using the Unit Discovery dataset) and align them to the voice recordings in a way that works best for the purpose of synthesizing novel utterances from novel speakers, similar to the target speaker's voice. We describe the metrics used for evaluation, a baseline system consisting of unsupervised subword unit discovery plus a standard TTS system, and a topline TTS using gold phoneme transcriptions. We present an overview of the 19 submitted systems from 10 teams and discuss the main results. Ewan Dunbar, Robin Algayres, Julien Karadayi, Mathieu Bernard, Juan Benjumea, Xuan-Nga Cao, Lucie Miskic, Charlotte Dugrain, Lucas Ondel Yang, Alan W. Black, Laurent Besacier, Sakriani Sakti, Emmanuel Dupoux |
INTERSPEECH | 10 |
| 2019 | Unsupervised Phonetic and Word Level Discovery for Speech to Speech Translation for Unwritten Languages
Steven Hillis, Anushree Prasanna Kumar, Alan W. Black |
INTERSPEECH | 3 |
| 2019 | Multilingual Speech Recognition with Corpus Relatedness SamplingabstractMultilingual acoustic models have been successfully applied to low-resource speech recognition.Most existing works have combined many small corpora together, and pretrained a multilingual model by sampling from each corpus uniformly.The model is eventually fine-tuned on each target corpus.This approach, however, fails to exploit the relatedness and similarity among corpora in the training set.For example, the target corpus might benefit more from a corpus in the same domain or a corpus from a close language.In this work, we propose a simple but useful sampling strategy to take advantage of this relatedness.We first compute the corpus-level embeddings and estimate the similarity between each corpus.Next we start training the multilingual model with uniform-sampling from each corpus at first, then we gradually increase the probability to sample from related corpora based on its similarity with the target corpus.Finally the model would be fine-tuned automatically on the target corpus.Our sampling strategy outperforms the baseline multilingual model on 16 low-resource tasks.Additionally, we demonstrate that our corpus embeddings capture the language and domain information of each corpus. Siddharth Dalmia, Alan W. Black, Florian Metze |
INTERSPEECH | 3 |
| 2019 | SANTLR: Speech Annotation Toolkit for Low Resource Languages
Zhong Zhou, Siddharth Dalmia, Alan W. Black, Florian Metze |
INTERSPEECH | 4 |
| 2019 | Variational Attention Using Articulatory Priors for Generating Code Mixed Speech Using Monolingual Corpora
Sai Krishna Rallabandi, Alan W. Black |
INTERSPEECH | 2 |
| 2019 | Ordinal Triplet Loss: Investigating Sleepiness Detection from Speech
Peter Wu, Sai Krishna Rallabandi, Alan W. Black, Eric Nyberg |
INTERSPEECH | 3 |
| 2019 | A Dynamic Strategy Coach for Effective NegotiationabstractNegotiation is a complex activity involving strategic reasoning, persuasion, and psychology.An average person is often far from an expert in negotiation.Our goal is to assist humans to become better negotiators through a machine-in-the-loop approach that combines machine's advantage at data-driven decisionmaking and human's language generation ability.We consider a bargaining scenario where a seller and a buyer negotiate the price of an item for sale through a text-based dialog.Our negotiation coach monitors messages between them and recommends tactics in real time to the seller to get a better deal (e.g., "reject the proposal and propose a price", "talk about your personal experience with the product").The best strategy and tactics largely depend on the context (e.g., the current price, the buyer's attitude).Therefore, we first identify a set of negotiation tactics, then learn to predict the best strategy and tactics in a given dialog context from a set of human-human bargaining dialogs.Evaluation on human-human dialogs shows that our coach increases the profits of the seller by almost 60%. 1 Yiheng Zhou, He He 0001, Alan W. Black, Yulia Tsvetkov |
SIGdial | 3 |
| 2019 | Analyzing Wikipedia Deletion Debates with a Group Decision-Making Forecast ModelabstractIn this work we show that machine learning with natural language processing can accurately forecast the outcomes of group decision-making in online discussions. Specifically, we study Articles for Deletion, a Wikipedia forum for determining which content should be included on the site. Applying this model, we replicate several findings from prior work on the factors that predict debate outcomes; we then extend this prior work and present new avenues for study, particularly in the use of policy citation during discussion. Alongside these findings, we introduce a structured corpus and source code for analyzing over 400,000 deletion debates spanning Wikipedia's history, enabling future large-scale studies of group decision-making discourse Elijah Mayfield, Alan W. Black |
Proc. ACM Hum. Comput. Interact. | 2 |
| 2018 | Style Transfer Through Back-TranslationabstractStyle transfer is the task of rephrasing the text to contain specific stylistic properties without changing the intent or affect within the context.This paper introduces a new method for automatic style transfer.We first learn a latent representation of the input sentence which is grounded in a language translation model in order to better preserve the meaning of the sentence while reducing stylistic properties.Then adversarial generation techniques are used to make the output match the desired style.We evaluate this technique on three different style transformations: sentiment, gender and political slant.Compared to two state-of-the-art style transfer modeling techniques we show improvements both in automatic evaluation of style transfer and in manual evaluation of meaning preservation and fluency. Shrimai Prabhumoye, Yulia Tsvetkov, Ruslan Salakhutdinov, Alan W. Black |
ACL (1) | 4 |
| 2018 | A Dataset for Document Grounded ConversationsabstractThis paper introduces a document grounded dataset for conversations.We define "Document Grounded Conversations" as conversations that are about the contents of a specified document.In this dataset the specified documents were Wikipedia articles about popular movies.The dataset contains 4112 conversations with an average of 21.43 turns per conversation.This positions this dataset to not only provide a relevant chat history while generating responses but also provide a source of information that the models could use.We describe two neural architectures that provide benchmark performance on the task of generating the next response.We also evaluate our models for engagement and fluency, and find that the information from the document helps in generating more engaging and fluent responses. Kangyan Zhou, Shrimai Prabhumoye, Alan W. Black |
EMNLP | 3 |
| 2018 | Sequence-Based Multi-Lingual Low Resource Speech RecognitionabstractTechniques for multi-lingual and cross-lingual speech recognition can help in low resource scenarios, to bootstrap systems and enable analysis of new languages and domains. End-to-end approaches, in particular sequence-based techniques, are attractive because of their simplicity and elegance. While it is possible to integrate traditional multi-lingual bottleneck feature extractors as front-ends, we show that end-to-end multi-lingual training of sequence models is effective on context independent models trained using Connectionist Temporal Classification (CTC) loss. We show that our model improves performance on Babel languages by over 6% absolute in terms of word/phoneme error rate when compared to mono-lingual systems built in the same setting for these languages. We also show that the trained model can be adapted cross-lingually to an unseen language using just 25% of the target data. We show that training on multiple languages is important for very low resource cross-lingual target scenarios, but not for multi-lingual testing scenarios. Here, it appears beneficial to include large well prepared datasets. Siddharth Dalmia, Ramon Sanabria, Florian Metze, Alan W. Black |
ICASSP | 4 |
| 2018 | Linguistic Unit Discovery from Multi-Modal Inputs in Unwritten Languages: Summary of the "Speaking Rosetta" JSALT 2017 WorkshopabstractWe summarize the accomplishments of a multi-disciplinary workshop exploring the computational and scientific issues surrounding the discovery of linguistic units (subwords and words) in a language without orthography. We study the replacement of orthographic transcriptions by images and/or translated text in a well-resourced language to help unsupervised discovery from raw speech. Odette Scharenborg, Laurent Besacier, Alan W. Black, Mark Hasegawa-Johnson, Florian Metze, Graham Neubig, Sebastian Stüker, Pierre Godard, Markus Müller 0001, Lucas Ondel Yang, Shruti Palaskar, Philip Arthur, Francesco Ciannella, Mingxing Du, Elin Larsen, Danny Merkx, Rachid Riad, Liming Wang 0003, Emmanuel Dupoux |
ICASSP | 3 |
| 2018 | An Investigation of Convolution Attention Based Models for Multilingual Speech Synthesis of Indian Languages
Pallavi Baljekar, Sai Krishna Rallabandi, Alan W. Black |
INTERSPEECH | 3 |
| 2018 | Multimodal Polynomial Fusion for Detecting Driver DistractionabstractDistracted driving is deadly, claiming 3,477 lives in the U.S. in 2015 alone. Although there has been a considerable amount of research on modeling the distracted behavior of drivers under various conditions, accurate automatic detection using multiple modalities and especially the contribution of using the speech modality to improve accuracy has received little attention. This paper introduces a new multimodal dataset for distracted driving behavior and discusses automatic distraction detection using features from three modalities: facial expression, speech and car signals. Detailed multimodal feature analysis shows that adding more modalities monotonically increases the predictive accuracy of the model. Finally, a simple and effective multimodal fusion technique using a polynomial fusion layer shows superior distraction detection results compared to the baseline SVM and neural network models. Yulun Du, Alan W. Black, Louis-Philippe Morency, Maxine Eskénazi |
INTERSPEECH | 2 |
| 2018 | Investigating Utterance Level Representations for Detecting Intent from Acoustics
Sai Krishna Rallabandi, Bhavya Karki, Carla Viegas, Eric Nyberg, Alan W. Black |
INTERSPEECH | 5 |
| 2018 | DialCrowd: A toolkit for easy dialog system assessmentabstractWhen creating a dialog system, developers need to test each version to ensure that it is performing correctly.Recently the trend has been to test on large datasets or to ask many users to try out a system.Crowdsourcing has solved the issue of finding users, but it presents new challenges such as how to use a crowdsourcing platform and what type of test is appropriate.Di-alCrowd makes system assessment using crowdsourcing easier by providing tools, templates and analytics.This paper describes the services that DialCrowd provides and how it works.It also describes a test of DialCrowd by a group of dialog system developers. Kyusong Lee, Alan W. Black, Maxine Eskénazi |
SIGDIAL Conference | 3 |
| 2018 | An Empirical Study of Self-Disclosure in Spoken Dialogue SystemsabstractSelf-disclosure is a key social strategy employed in conversation to build relations and increase conversational depth.It has been heavily studied in psychology and linguistic literature, particularly for its ability to induce self-disclosure from the recipient, a phenomena known as reciprocity.However, we know little about how self-disclosure manifests in conversation with automated dialog systems, especially as any self-disclosure on the part of a dialog system is patently disingenuous.In this work, we run a large-scale quantitative analysis on the effect of selfdisclosure by analyzing interactions between real-world users and a spoken dialog system in the context of social conversation.We find that indicators of reciprocity occur even in human-machine dialog, with far-reaching implications for chatbots in a variety of domains including education, negotiation and social dialog. Abhilasha Ravichander, Alan W. Black |
SIGDIAL Conference | 2 |
| 2018 | Domain Robust Feature Extraction for Rapid Low Resource ASR DevelopmentabstractDeveloping a practical speech recognizer for a low resource language is challenging, not only because of the (potentially unknown) properties of the language, but also because test data may not be from the same domain as the available training data.In this paper, we focus on the latter challenge, i.e. domain mismatch, for systems trained using a sequence-based criterion. We demonstrate the effectiveness of using a pre-trained English recognizer, which is robust to such mismatched conditions, as a domain normalizing feature extractor on a low resource language. In our example, we use Turkish Conversational Speech and Broadcast News data.This enables rapid development of speech recognizers for new languages which can easily adapt to any domain. Testing in various cross-domain scenarios, we achieve relative improvements of around 25% in phoneme error rate, with improvements being around 50% for some domains. Siddharth Dalmia, Florian Metze, Alan W. Black |
SLT | 4 |
| 2018 | The ARIEL-CMU situation frame detection pipeline for LoReHLT16: a model translation approach
Patrick Littell, Ruochen Xu, Zaid Sheikh, David R. Mortensen, Lori S. Levin, Francis M. Tyers, Hiroaki Hayashi, Graham Horwood, Steve Sloto, Emily Tagtow, Alan W. Black, Yiming Yang 0002, Teruko Mitamura, Eduard H. Hovy |
Mach. Transl. | 12 |
| 2017 | Integrating Verbal and Nonvebval Input into a Dynamic Response Spoken Dialogue SystemabstractIn this work, we present a dynamic response spoken dialogue system (DRSDS). It is capable of understanding the verbal and nonverbal language of users and making instant, situation-aware response. Incorporating with two external systems, MultiSense and email summarization, we built an email reading agent on mobile device to show the functionality of DRSDS. Ting-Yao Hu, Chirag Raman, Salvador Medina Maza, Liangke Gui, Tadas Baltrusaitis, Robert E. Frederking, Louis-Philippe Morency, Alan W. Black, Maxine Eskénazi |
AAAI | 8 |
| 2017 | The CMU entry to blizzard machine learning challengeabstractThe paper describes Carnegie Mellon University's (CMU) entry to the ES-1 sub-task of the Blizzard Machine Learning Speech Synthesis Challenge 2017. The submitted system is a parametric model trained to predict vocoder parameters given linguistic features. The task in this year's challenge was to synthesize speech from children's audiobooks. Linguistic and acoustic features were provided by the organizers and the task was to find the best performing model. The paper explores various RNN architectures that were investigated and describes the final model that was submitted. Pallavi Baljekar, Sai Krishna Rallabandi, Alan W. Black |
ASRU | 3 |
| 2017 | The blizzard machine learning challenge 2017abstractThis paper describes the Blizzard Machine Learning Challenge (BMLC) 2017, which is a spin-off of the Blizzard Challenge. The annual Blizzard Challenges 2005-2017 were held to better understand and compare research techniques in building corpus-based text-to-speech (TTS) systems on the same data. The series of Blizzard Challenges has helped us measure progress in TTS technology. However, to get competitive performance, a lot time has to be spent on skilled tasks. This may make the Blizzard Challenge unattractive to machine learning researchers from other fields. Therefore, we recommend that the BMLC not involve these speech-specific tasks and that it allow participants to concentrate on the acoustic modeling task, framed as a straightforward machine learning problem, with a fixed dataset. In the BMLC 2017, two types of datasets consisting of four hours of speech data suitable for machine learning problems were distributed. This paper summarizes the purpose, design, and whole process of the challenge and its results. Kei Sawada, Keiichi Tokuda, Simon King 0001, Alan W. Black |
ASRU | 4 |
| 2017 | Learning Conversational Systems that Interleave Task and Non-Task ContentabstractTask-oriented dialog systems have been applied in various tasks, such as automated personal assistants, customer service providers and tutors. These systems work well when users have clear and explicit intentions that are well-aligned to the systems' capabilities. However, they fail if users intentions are not explicit.To address this shortcoming, we propose a framework to interleave non-task content (i.e.everyday social conversation) into task conversations. When the task content fails, the system can still keep the user engaged with the non-task content. We trained a policy using reinforcement learning algorithms to promote long-turn conversation coherence and consistency, so that the system can have smooth transitions between task and non-task content.To test the effectiveness of the proposed framework, we developed a movie promotion dialog system. Experiments with human users indicate that a system that interleaves social and task content achieves a better task success rate and is also rated as more engaging compared to a pure task-oriented system. Zhou Yu 0005, Alexander I. Rudnicky, Alan W. Black |
IJCAI | 3 |
| 2017 | Speech Synthesis for Mixed-Language Navigation InstructionsabstractText-to-Speech (TTS) systems that can read navigation instructions are one of the most widely used speech interfaces today. Text in the navigation domain may contain named entities such as location names that are not in the language that the TTS database is recorded in. Moreover, named entities can be compound words where individual lexical items belong to different languages. These named entities may be transliterated into the script that the TTS system is trained on. This may result in incorrect pronunciation rules being used for such words. We describe experiments to extend our previous work in generating code-mixed speech to synthesize navigation instructions, with a mixed-lingual TTS system. We conduct subjective listening tests with two sets of users, one being students who are native speakers of an Indian language and very proficient in English, and the other being drivers with low English literacy, but familiarity with location names. We find that in both sets of users, there is a significant preference for our proposed system over a baseline system that synthesizes instructions in English. Khyathi Raghavi Chandu, Sai Krishna Rallabandi, Sunayana Sitaram, Alan W. Black |
INTERSPEECH | 4 |
| 2017 | On Building Mixed Lingual Speech Synthesis Systems
Sai Krishna Rallabandi, Alan W. Black |
INTERSPEECH | 2 |
| 2017 | Segment Level Voice Conversion with Recurrent Neural Networks
Miguel Varela Ramos, Alan W. Black, Ramón Fernandez Astudillo, Isabel Trancoso, Nuno Fonseca |
INTERSPEECH | 2 |
| 2016 | Towards building an attentive artificial listener: on the perception of attentiveness in audio-visual feedback tokensabstractCurrent dialogue systems typically lack a variation of audio-visual feedback tokens. Either they do not encompass feedback tokens at all, or only support a limited set of stereotypical functions. However, this does not mirror the subtleties of spontaneous conversations. If we want to be able to build an artificial listener, as a first step towards building an empathetic artificial agent, we also need to be able to synthesize more subtle audio-visual feedback tokens. In this study, we devised an array of monomodal and multimodal binary comparison perception tests and experiments to understand how different realisations of verbal and visual feedback tokens influence third-party perception of the degree of attentiveness. This allowed us to investigate i) which features (amplitude, frequency, duration...) of the visual feedback influences attentiveness perception; ii) whether visual or verbal backchannels are perceived to be more attentive iii) whether the fusion of unimodal tokens with low perceived attentiveness increases the degree of perceived attentiveness compared to unimodal tokens with high perceived attentiveness taken alone; iv) the automatic ranking of audio-visual feedback token in terms of conveyed degree of attentiveness. Catharine Oertel, José Lopes 0001, Yu Yu 0003, Kenneth Alberto Funes Mora, Joakim Gustafson, Alan W. Black, Jean-Marc Odobez |
ICMI | 6 |
| 2016 | Towards Building an Attentive Artificial Listener: On the Perception of Attentiveness in Feedback UtterancesabstractTowards Building an Attentive Artificial Listener: On the Perception of Attentiveness in Feedback Utterances Catharine Oertel, Joakim Gustafson, Alan W. Black |
INTERSPEECH | 3 |
| 2016 | Deriving Phonetic Transcriptions and Discovering Word Segmentations for Speech-to-Speech Translation in Low-Resource Settings
Andrew Wilkinson, Alan W. Black |
INTERSPEECH | 3 |
| 2016 | User Engagement Study with Virtual Agents Under Different Cultural Contexts
Zhou Yu 0005, Xinrui He, Alan W. Black, Alexander I. Rudnicky |
IVA | 3 |
| 2016 | Socially-Aware Virtual Agents: Automatically Assessing Dyadic Rapport from Temporal Patterns of Behavior
Tanmay Sinha, Alan W. Black, Justine Cassell |
IVA | 3 |
| 2016 | Speech Synthesis of Code-Mixed Text
Sunayana Sitaram, Alan W. Black |
LREC | 2 |
| 2016 | Polyglot Neural Language Models: A Case Study in Cross-Lingual Phonetic Representation LearningabstractYulia Tsvetkov, Sunayana Sitaram, Manaal Faruqui, Guillaume Lample, Patrick Littell, David Mortensen, Alan W Black, Lori Levin, Chris Dyer. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Yulia Tsvetkov, Sunayana Sitaram, Manaal Faruqui, Guillaume Lample, Patrick Littell, David R. Mortensen, Alan W. Black, Lori S. Levin, Chris Dyer |
HLT-NAACL | 7 |
| 2016 | Initiations and Interruptions in a Spoken Dialog System
Leah Nicolich-Henkin, Carolyn P. Rosé, Alan W. Black |
SIGDIAL Conference | 3 |
| 2016 | A Wizard-of-Oz Study on A Non-Task-Oriented Dialog Systems That Reacts to User EngagementabstractIn this paper, we describe a system that reacts to both possible system breakdowns and low user engagement with a set of conversational strategies.These general strategies reduce the number of inappropriate responses and produce better user engagement.We also found that a system that reacts to both possible system breakdowns and low user engagement is rated by both experts and non-experts as having better overall user engagement compared to a system that only reacts to possible system breakdowns.We argue that for non-task-oriented systems we should optimize on both system response appropriateness and user engagement.We also found that apart from making the system response appropriate, funny and provocative responses can also lead to better user engagement.On the other hand, short appropriate responses, such as "Yes" or "No" can lead to decreased user engagement.We will use these findings to further improve our system. Zhou Yu 0005, Leah Nicolich-Henkin, Alan W. Black, Alexander I. Rudnicky |
SIGDIAL Conference | 3 |
| 2016 | Strategy and Policy Learning for Non-Task-Oriented Conversational SystemsabstractWe propose a set of generic conversational strategies to handle possible system breakdowns in non-task-oriented dialog systems.We also design policies to select these strategies according to dialog context.We combine expert knowledge and the statistical findings derived from data in designing these policies.The policy learned via reinforcement learning outperforms the random selection policy and the locally greedy policy in both simulated and real-world settings.In addition, we propose three metrics for conversation quality evaluation which consider both the local and global quality of the conversation. Zhou Yu 0005, Ziyu Xu 0001, Alan W. Black, Alexander I. Rudnicky |
SIGDIAL Conference | 3 |
| 2016 | Automatic Recognition of Conversational Strategies in the Service of a Socially-Aware Dialog SystemabstractIn this work, we focus on automatically recognizing social conversational strategies that in human conversation contribute to building, maintaining or sometimes destroying a budding relationship.These conversational strategies include self-disclosure, reference to shared experience, praise and violation of social norms.By including rich contextual features drawn from verbal, visual and vocal modalities of the speaker and interlocutor in the current and previous turn, we can successfully recognize these dialog phenomena with an accuracy of over 80% and kappa ranging from 60-80%.Our findings have been successfully integrated into an end-to-end socially aware dialog system, with implications for virtual agents that can use rapport between user and system to improve task-oriented assistance. Tanmay Sinha, Alan W. Black, Justine Cassell |
SIGDIAL Conference | 3 |
| 2016 | This Table is Different: A WordNet-Based Approach to Identifying References to Document EntitiesabstractWriting intended to inform frequently contains references to document entities (DEs), a mixed class that includes orthographically structured items (e.g., illustrations, sections, lists) and discourse entities (arguments, suggestions, points).Such references are vital to the interpretation of documents, but they often eschew identifiers such as "Figure 1" for inexplicit phrases like "in this figure " or "from these premises".We examine inexplicit references to DEs, termed DE references, and recast the problem of their automatic detection into the determination of relevant word senses.We then show the feasibility of machine learning for the detection of DErelevant word senses, using a corpus of human-labeled synsets from WordNet.We test cross-domain performance by gathering lemmas and synsets from three corpora: website privacy policies, Wikipedia articles, and Wikibooks textbooks.Identifying DE references will enable language technologies to use the information encoded by them, permitting the automatic generation of finely-tuned descriptions of DEs and the presentation of richly-structured information to readers. Shomir Wilson, Alan W. Black, Jon Oberlander |
GWC | 2 |
| 2016 | Mining Parallel Corpora from Sina Weibo and TwitterabstractMicroblogs such as Twitter, Facebook, and Sina Weibo (China's equivalent of Twitter) are a remarkable linguistic resource. In contrast to content from edited genres such as newswire, microblogs contain discussions of virtually every topic by numerous individuals in different languages and dialects and in different styles. In this work, we show that some microblog users post “self-translated” messages targeting audiences who speak different languages, either by writing the same message in multiple languages or by retweeting translations of their original posts in a second language. We introduce a method for finding and extracting this naturally occurring parallel data. Identifying the parallel content requires solving an alignment problem, and we give an optimally efficient dynamic programming algorithm for this. Using our method, we extract nearly 3M Chinese–English parallel segments from Sina Weibo using a targeted crawl of Weibo users who post in multiple languages. Additionally, from a random sample of Twitter, we obtain substantial amounts of parallel data in multiple language pairs. Evaluation is performed by assessing the accuracy of our extraction approach relative to a manual annotation as well as in terms of utility as training data for a Chinese–English machine translation system. Relative to traditional parallel data resources, the automatically extracted parallel data yield substantial translation quality improvements in translating microblog text and modest improvements in translating edited news content. Wang Ling, Luís Marujo, Chris Dyer, Alan W. Black, Isabel Trancoso |
Comput. Linguistics | 4 |
| 2016 | Postfilters to Modify the Modulation Spectrum for Statistical Parametric Speech SynthesisabstractThis paper presents novel approaches based on modulation spectrum (MS) for high-quality statistical parametric speech synthesis, including text-to-speech (TTS) and voice conversion (VC). Although statistical parametric speech synthesis offers various advantages over concatenative speech synthesis, the synthetic speech quality is still not as good as that of concatenative speech synthesis or the quality of natural speech. One of the biggest issues causing the quality degradation is the over-smoothing effect often observed in the generated speech parameter trajectories. Global variance (GV) is known as a feature well correlated with the over-smoothing effect, and the effectiveness of keeping the GV of the generated speech parameter trajectories similar to those of natural speech has been confirmed. However, the quality gap between natural speech and synthetic speech is still large. In this paper, we propose using the MS of the generated speech parameter trajectories as a new feature to effectively quantify the over-smoothing effect. Moreover, we propose postfilters to modify the MS utterance by utterance or segment by segment to make the MS of synthetic speech close to that of natural speech. The proposed postfilters are applicable to various synthesizers based on statistical parametric speech synthesis. We first perform an evaluation of the proposed method in the framework of hidden Markov model (HMM)-based TTS, examining its properties from different perspectives. Furthermore, effectiveness of the proposed postfilters are also evaluated in Gaussian mixture model (GMM)-based VC and classification and regression trees (CART)-based TTS (a.k.a., CLUSTERGEN). The experimental results demonstrate that 1) the proposed utterance-level postfilter achieves quality comparable to the conventional generation algorithm considering the GV, and yields significant improvements by applying to the GV-based generation algorithm in HMM-based TTS, 2) the proposed segment-level postfilter capable of achieving low-delay synthesis also yields significant improvements in synthetic speech quality, and 3) the proposed postfilters are also effective in not only HMM-based TTS but also GMM-based VC and CLUSTERGEN. Shinnosuke Takamichi, Tomoki Toda, Alan W. Black, Graham Neubig, Sakriani Sakti, Satoshi Nakamura 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2015 | Utterance classification in speech-to-speech translation for zero-resource languages in the hospital administration domainabstractAlthough substantial progress has been achieved in speech-to-speech translation systems over the last few years, such systems still require that the speech be written in some appropriate orthography. As speech may differ greatly from the standardized written form of a language, it can be non-trivial to collect written data when there is no standard way for it to be represented. This project addresses the problem from the other end and expects that speech alone is available in the target language, and that no (standard or non-standard) orthography exists. It, therefore, treats the acoustic representation of the language as primary and uses language-independent methods to produce a phonetically-related symbolic representation that is then used in the translation system. Thus, the speech translation system is created for the target language as defined by the recording of that language rather than some body of orthographic transcripts. In this work, we are creating an application called APT (Acoustic Patient Translator), which uses a novel scheme of speech recognition and translation within a targeted domain. By working with a set of predefined sentences appropriately chosen to fit a scenario, we use utterance classification as a speech recognition algorithm. The utterance classification is achieved using cross-lingual, language-independent phonetic labeling. Since we are working with a set of select phrases, the translation part is trivial. We are concentrating on communication with hospital staff, such as scheduling a doctor's appointment, as our domain. In addition to English, we also run experiments on Tamil. Lara J. Martin, Andrew Wilkinson, Sai Sumanth Miryala, Vivian Robison, Alan W. Black |
ASRU | 5 |
| 2015 | Finding Function in Form: Compositional Character Models for Open Vocabulary Word RepresentationabstractWang Ling, Chris Dyer, Alan W Black, Isabel Trancoso, Ramón Fermandez, Silvio Amir, Luís Marujo, Tiago Luís. Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing. 2015. Wang Ling, Chris Dyer, Alan W. Black, Isabel Trancoso, Ramon Fermandez, Silvio Amir, Luís Marujo, Tiago Luís |
EMNLP | 3 |
| 2015 | Not All Contexts Are Created Equal: Better Word Representations with Variable AttentionabstractWang Ling, Yulia Tsvetkov, Silvio Amir, Ramón Fermandez, Chris Dyer, Alan W Black, Isabel Trancoso, Chu-Cheng Lin. Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing. 2015. Wang Ling, Yulia Tsvetkov, Silvio Amir, Ramon Fermandez, Chris Dyer, Alan W. Black, Isabel Trancoso, Chu-Cheng Lin |
EMNLP | 6 |
| 2015 | Parameter generation algorithm considering Modulation Spectrum for HMM-based speech synthesisabstractThis paper proposes a novel parameter generation algorithm for high-quality speech generation in Hidden Markov Model (HMM)-based speech synthesis. One of the biggest issues causing significant quality degradation is the over-smoothing effect often observed in generated parameter trajectories. Global Variance (GV) is known as a feature well correlated with the over-smoothing effect and a metric on the GV of the generated parameters is effectively used as a penalty term in the conventional parameter generation. However, the quality of the synthetic speech is far from that of the natural speech. Recently, we have found that a Modulation Spectrum (MS) of the generated parameters, which is also regarded as an extension of the GV, is more sensitively correlated with the over-smoothing effect than the GV. This paper incorporates a metric on the MS as a new penalty term in the proposed parameter generation algorithm. The experimental results demonstrate that the proposed parameter generation algorithm considering the MS yields significant improvements in synthetic speech quality compared to the conventional parameter generation algorithm considering the GV. Shinnosuke Takamichi, Tomoki Toda, Alan W. Black, Satoshi Nakamura 0001 |
ICASSP | 3 |
| 2015 | Modulation spectrum-constrained trajectory training algorithm for GMM-based Voice ConversionabstractThis paper presents a novel training algorithm for Gaussian Mixture Model (GMM)-based Voice Conversion (VC). One of the advantages of GMM-based VC is computationally efficient conversion processing enabling to achieve real-time VC applications. On the other hand, the quality of the converted speech is still significantly worse than that of natural speech. In order to address this problem while preserving the computationally efficient conversion processing, the proposed training method enables 1) to use a consistent optimization criterion between training and conversion and 2) to compensate a Modulation Spectrum (MS) of the converted parameter trajectory as a feature sensitively correlated with over-smoothing effects causing quality degradation of the converted speech. The experimental results demonstrate that the proposed algorithm yields significant improvements in term of both the converted speech quality and the conversion accuracy for speaker individuality compared to the basic training algorithm. Shinnosuke Takamichi, Tomoki Toda, Alan W. Black, Satoshi Nakamura 0001 |
ICASSP | 3 |
| 2015 | Using articulatory features and inferred phonological segments in zero resource speech processing
Pallavi Baljekar, Sunayana Sitaram, Prasanna Kumar Muthukumar, Alan W. Black |
INTERSPEECH | 4 |
| 2015 | Random forests for statistical speech synthesis
Alan W. Black, Prasanna Kumar Muthukumar |
INTERSPEECH | 1 |
| 2015 | Distributed representation-based spoken word sense inductionabstractSpoken Term Detection (STD) or Keyword Search (KWS) techniques can locate keyword instances but do not differentiate between meanings. Spoken Word Sense Induction (SWSI) differentiates target instances by clustering according to context, providing a more useful result. In this paper we present a fully unsupervised SWSI approach based on distributed representations of spoken utterances. We compare this approach to several others, including the state-of-the-art Hierarchical Dirichlet Process (HDP). To determine how ASR performance affects SWSI, we used three different levels of Word Error Rate (WER), 40%, 20 % and 0%; 40 % WER is representative of online video, 0 % of text. We show that the distributed representation approach outperforms all other approaches, regardless of the WER. Although LDA-based approaches do well on clean data, they degrade significantly with WER. Paradoxically, lower WER does not guarantee better SWSI performance, due to the influence of common locutions. Justin T. Chiu, Yajie Miao, Alan W. Black, Alexander I. Rudnicky |
INTERSPEECH | 3 |
| 2015 | Using acoustics to improve pronunciation for synthesis of low resource languages
Sunayana Sitaram, Serena Jeblee, Alan W. Black |
INTERSPEECH | 3 |
| 2015 | Universal grapheme-based speech synthesis
Sunayana Sitaram, Alok Parlikar, Gopala Krishna Anumanchipalli, Alan W. Black |
INTERSPEECH | 4 |
| 2015 | Modulation spectrum-constrained trajectory training algorithm for HMM-based speech synthesis
Shinnosuke Takamichi, Tomoki Toda, Alan W. Black, Satoshi Nakamura 0001 |
INTERSPEECH | 3 |
| 2015 | Two/Too Simple Adaptations of Word2Vec for Syntax ProblemsabstractWang Ling, Chris Dyer, Alan W. Black, Isabel Trancoso. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015. Wang Ling, Chris Dyer, Alan W. Black, Isabel Trancoso |
HLT-NAACL | 3 |
| 2015 | The Real Challenge 2014: Progress and ProspectsabstractThe REAL Challenge took place for the first time in 2014, with a long term goal of creating streams of real data that the research community can use, by fostering the creation of systems that are capable of attracting real users.A novel approach is to have high school and undergraduate students devise the types of applications that would attract many real users and that need spoken interaction.The projects are presented to researchers from the spoken dialog research community and the researchers and students work together to refine and develop the ideas.Eleven projects were presented at the first workshop.Many of them have found mentors to help in the next stages of the projects.The students have also brought out issues in the use of speech for real applications.Those issues involve privacy and significant personalization of the applications.While long-term impact of the challenge remains to be seen, the challenge has already been a success at its immediate aims of bringing new ideas and new researchers into the community, and serves as a model for related outreach efforts. Maxine Eskénazi, Alan W. Black, David R. Traum |
SIGDIAL Conference | 2 |
| 2015 | An Incremental Turn-Taking Model with Active System Barge-in for Spoken Dialog SystemsabstractThis paper deals with an incremental turntaking model that provides a novel solution for end-of-turn detection.It includes a flexible framework that enables active system barge-in.In order to accomplish this, a systematic procedure of teaching a dialog system to produce meaningful system barge-in is presented.This procedure improves system robustness and success rate.It includes constructing cost models and learning optimal policy using reinforcement learning.Results show that our model reduces false cut-in rate by 37.1% and response delay by 32.5% compared to the baseline system.Also the learned system barge-in strategy yields a 27.7% increase in average reward from user responses. Alan W. Black, Maxine Eskénazi |
SIGDIAL Conference | 2 |
| 2014 | Automatic discovery of a phonetic inventory for unwritten languages for statistical speech synthesisabstractSpeech synthesis systems are typically built with speech data and transcriptions. In this paper, we try to build synthesis systems when no transcriptions or knowledge about the language are available. It is usually necessary to at least possess phonetic knowledge about the language. In this paper, we propose an automated way of obtaining phones and phonetic knowledge about the corpus at hand by making use of Articulatory Features (AFs). An Articulatory Feature predictor is trained on a bootstrap corpus in an arbitrary other language using a three-hidden layer neural network. This neural network is run on the speech corpus to extract AFs. Hierarchical clustering is used to cluster the AFs into categories i.e. phones. Phonetic information about each of these inferred phones is obtained by computing the mean of the AFs in each cluster. Results of systems built with this framework in multiple languages are reported. Prasanna Kumar Muthukumar, Alan W. Black |
ICASSP | 2 |
| 2013 | Microblogs as Parallel Corpora
Wang Ling, Guang Xiang, Chris Dyer, Alan W. Black, Isabel Trancoso |
ACL (1) | 4 |
| 2013 | Improved punctuation recovery through combination of multiple speech streamsabstractIn this paper, we present a technique to use the information in multiple parallel speech streams, which are approximate translations of each other, in order to improve performance in a punctuation recovery task. We first build a phraselevel alignment of these multiple streams, using phrase tables to link the phrase pairs together. The information so collected is then used to make it more likely that sentence units are equivalent across streams. We applied this technique to a number of simultaneously interpreted speeches of the European Parliament Committees, for the recovery of the full stop, in four different languages (English, Italian, Portuguese and Spanish). We observed an average improvement in SER of 37% when compared to an existing baseline, in Portuguese and English. João Miranda, João Paulo da Silva Neto, Alan W. Black |
ASRU | 3 |
| 2013 | Paraphrasing 4 Microblog NormalizationabstractCompared to the edited genres that have played a central role in NLP research, microblog texts use a more informal register with nonstandard lexical items, abbreviations, and free orthographic variation.When confronted with such input, conventional text analysis tools often perform poorly.Normalization -replacing orthographically or lexically idiosyncratic forms with more standard variants -can improve performance.We propose a method for learning normalization rules from machine translations of a parallel corpus of microblog messages.To validate the utility of our approach, we evaluate extrinsically, showing that normalizing English tweets and then translating improves translation quality (compared to translating unnormalized text) using three standard web translation services as well as a phrase-based translation system trained on parallel microblog data. Wang Ling, Chris Dyer, Alan W. Black, Isabel Trancoso |
EMNLP | 3 |
| 2013 | Accent Group modeling for improved prosody in statistical parameteric speech synthesisabstractThis paper presents an `Accent Group' based intonation model for statistical parametric speech synthesis. We propose an approach to automatically model phonetic realizations of fundamental frequency(F0) contours as a sequence of intonational events anchored to a group of syllables (an Accent Group). We train an accent grouping model specific to that of the speaker, using a stochastic context free grammar and contextual decision trees on the syllables. This model is used to `parse' an unseen text into its constituent accent groups over each of which appropriate intonation is predicted. The performance of the model is shown objectively and subjectively on a variety of prosodically diverse tasks- read speech, news broadcast and audio books. Gopala Krishna Anumanchipalli, Luís C. Oliveira, Alan W. Black |
ICASSP | 3 |
| 2013 | A style capturing approach to F0 transformation in voice conversionabstractIn this paper, we present a new approach to F0 transformation, that can capture aspects of speaking style. Instead of using the traditional 5ms frames as units in transformation, we propose a method that looks at longer phonological regions such as metrical feet. We automatically detect metrical feet in the source speech, and for each of source speaker's feet, we find its phonological correspondence in target speech. We use a statistical phrase accent model to represent the F0 contour, where a 4-dimensional TILT representation is used for the F0 is parameterized over each feet region for the source and target speakers. This forms the parallel data that is the training data for our transformation. We transform the phrase component using simple z-score mapping. We use a joint density Gaussian mixture model to transform the accent contours. Our transformation method generates F0 contours that are significantly more correlated with the target speech than a baseline, frame-based method. Gopala Krishna Anumanchipalli, Luís C. Oliveira, Alan W. Black |
ICASSP | 3 |
| 2013 | Improving ASR by integrating lecture audio and slidesabstractWe propose a method to combine audio of a lecture with its supporting slides in order to improve automatic speech recognition performance. We view both the lecture speech and the slides as parallel streams which contain redundant information. We integrate both streams in order to bias the recognizer's language model towards the words in the slides, by first aligning the speech with the slide words, thus correcting errors on the ASR transcripts. We obtain a 5.9% relative WER improvement on a lecture test set, when compared to a speech recognition only system. João Miranda, João Paulo da Silva Neto, Alan W. Black |
ICASSP | 3 |
| 2013 | Bootstrapping Text-to-Speech for speech processing in languages without an orthographyabstractSpeech synthesis technology has reached the stage where given a well-designed corpus of audio and accurate transcription an at least understandable synthesizer can be built without necessarily resorting to new innovations. However many languages do not have a well-defined writing system but such languages could still greatly benefit from speech systems. In this paper we consider the case where we have a (potentially large) single speaker database but have no transcriptions and no standardized way to write transcriptions. To address this scenario we propose a method that allows us to bootstrap synthetic voices purely from speech data. We use a novel combination of automatic speech recognition and automatic word segmentation for the bootstrapping. Our experimental results on speech corpora in two languages, English and German, show that synthetic voices that are built using this method are close to understandable. Our method is language-independent and can thus be used to build synthetic voices from a speech corpus in any new language. Sunayana Sitaram, Sukhada Palkar, Yun-Nung Chen, Alok Parlikar, Alan W. Black |
ICASSP | 5 |
| 2013 | Analysis and modeling of "focus" in contextabstractThis paper uses a crowd-sourced definition of a speech phe-nomenon we have called “focus”. Given sentences, text and speech, in isolation and in context, we asked annotators to iden-tify what we term the “focus ” word. We present their consis-tency in identifying the focused word, when presented with text or speech stimuli. We then build models to show how well we predict that focus word from lexical (and higher) level features. Also, using spectral and prosodic information, we show the dif-ferences in these focus words when spoken with and without context. Finally, we show how we can improve speech synthe-sis of these utterances given focus information. Dirk Hovy, Gopala Krishna Anumanchipalli, Alok Parlikar, Caroline Vaughn, Adam C. Lammert, Eduard H. Hovy, Alan W. Black |
INTERSPEECH | 7 |
| 2013 | Optimizations and fitting procedures for the liljencrants-fant model for statistical parametric speech synthesis
Prasanna Kumar Muthukumar, Alan W. Black, H. Timothy Bunnell |
INTERSPEECH | 2 |
| 2013 | The Dialog State Tracking Challenge
Jason D. Williams, Antoine Raux, Deepak Ramachandran, Alan W. Black |
SIGDIAL Conference | 4 |
| 2013 | Automatic Prediction of Friendship via Multi-model Dyadic Features
Zhou Yu 0005, David Gerritsen, Amy Ogan, Alan W. Black, Justine Cassell |
SIGDIAL Conference | 4 |
| 2012 | Entropy-based Pruning for Phrase-based Machine Translation
Wang Ling, João Graça, Isabel Trancoso, Alan W. Black |
EMNLP-CoNLL | 4 |
| 2012 | Articulatory features for expressive speech synthesisabstractThis paper describes some of the results from the project entitled “New Parameterization for Emotional Speech Synthesis” held at the Summer 2011 JHU CLSP workshop. We describe experiments on how to use articulatory features as a meaningful intermediate representation for speech synthesis. This parameterization not only allows us to reproduce natural sounding speech but also allows us to generate stylistically varying speech. Alan W. Black, H. Timothy Bunnell, Ying Dou, Prasanna Kumar Muthukumar, Florian Metze, Daniel Perry 0002, Tim Polzehl, Kishore Prahallad, Stefan Steidl, Callie Vaughn |
ICASSP | 1 |
| 2012 | Data-driven phrasing for speech synthesis in low-resource languagesabstractWe present an approach to build phrase break prediction models when synthesizing text in low resource languages. This method allows building models without depending on the availability of part of speech taggers, or corpus with hand annotated breaks. We use the same speech data used for building a synthetic voice, to deduce acoustic phrase breaks. We perform unsupervised part of speech induction over a small text corpus in the language at hand. We use these tags and train a grammar based phrasing model. In this paper, we show results for the languages: English, Portuguese and Marathi, which suggest that we can quickly build very reasonable phrasing models for new languages using very little data. Alok Parlikar, Alan W. Black |
ICASSP | 2 |
| 2012 | Text-dependent pathological voice detectionabstractWhile global characteristics of the speaker’s source and spectral features have been successfully employed in pathological voice detection, the underlying text has largely been ignored. In this work, we focus on experiments that exploit the text stimulus that is read by the subject. Features derived from text include the mean cepstral distortion of the subject from an average intelligible speaker, and prosodic features include the speaking rate, statistics of phoneme durations, etc. The phonetic labeling information is also exploited to ignore all the unvoiced regions of the speech samples to improve the discriminability between intelligible and pathological voices. We also designed features that capture the speaker’s overall closeness to intelligible instances of the same text stimulus from other speakers. Our experiments show that the proposed text-derived features improve the detection of pathological voices by 20%. Index Terms: Pathological voices, example based detection, text-driven features, fusion of classification methods. Gopala Krishna Anumanchipalli, Hugo Meinedo, Miguel M. F. Bugalho, Isabel Trancoso, Luís C. Oliveira, Alan W. Black |
INTERSPEECH | 6 |
| 2012 | Modelling a Noisy-channel for Voice Conversion Using Articulatory FeaturesabstractIn this paper, we propose modeling a noisy-channel for the task of voice conversion (VC). We have used the artificial neural networks (ANN) to capture speaker-specific characteristics of a target speaker which avoid the need for any training utterance from a source speaker. We use articulatory features (AFs) as a canonical form or speaker-independent representation of a speech signal. Our studies show that AFs contain a significant amount of speaker information in their trajectories. Suitable techniques are proposed to normalize the speaker-specific information in AF trajectories and the resultant AFs are used in voice conversion. The results of voice conversion evaluated using objective and subjective measures confirm that AFs can be used as a canonical form in nosiy-channel to capture speakerspecific characteristics of a target speaker. Bajibabu Bollepalli, Alan W. Black, Kishore Prahallad |
INTERSPEECH | 2 |
| 2012 | Parallel combination of multilingual speech streams for improved ASR
João Miranda, João Paulo da Silva Neto, Alan W. Black |
INTERSPEECH | 3 |
| 2012 | Modeling Pause-Duration for Style-Specific Speech SynthesisabstractA major contribution to speaking style comes from both the location of phrase breaks in an utterance, as well as the duration of these breaks. This paper is about modeling the duration of style specific breaks. We look at six styles of speech here. We present analysis that shows that these styles differ in the duration of pauses in natural speech. We have built CART models to predict the pause duration in these corpora and have integrated them into the Festival speech synthesis system. Our objective results show that if we have sufficient training data, we can build style specific models. Our subjective tests show that people can perceive the difference between different models and that they prefer style specific models over simple pause duration models. Alok Parlikar, Alan W. Black |
INTERSPEECH | 2 |
| 2012 | The IIIT-H Indic Speech DatabasesabstractThis paper discusses the efforts in collecting speech databases for Indian languages – Bengali, Hindi, Kannada, Malayalam, Marathi, Tamil and Telugu. We discuss relevant design considerations in collecting these databases, and demonstrate their usage in speech synthesis. By releasing these speech databases in the public domain without any restrictions for non commercial and commercial purposes, we hope to promote research and developmental activities in building speech synthesis systems in Indian languages. Index Terms:speech databases, speech synthesis, Indian languages Kishore Prahallad, Naresh Kumar Elluru, Venkatesh Keri, Rajendran S, Alan W. Black |
INTERSPEECH | 5 |
| 2012 | Practical Evaluation of Human and Synthesized Speech for Virtual Human Dialogue Systems
Kallirroi Georgila, Alan W. Black, Kenji Sagae, David R. Traum |
LREC | 2 |
| 2012 | "Love ya, jerkface": Using Sparse Log-Linear Models to Build Positive and Impolite Relationships with Teens
William Yang Wang, Samantha L. Finkelstein, Amy Ogan, Alan W. Black, Justine Cassell |
SIGDIAL Conference | 4 |
| 2012 | Intent transfer in speech-to-speech machine translationabstractThis paper presents an approach for transfer of speaker intent in speech-to-speech machine translation (S2SMT). Specifically, we describe techniques to retain the prominence patterns of the source language utterance through the translation pipeline and impose this information during speech synthesis in the target language. We first present an analysis of word focus across languages to motivate the problem of transfer. We then propose an approach for training an appropriate transfer function for intonation on a parallel speech corpus in the two languages within which the translation is carried out. We present our analysis and experiments on English↔Portuguese and English↔German language pairs and evaluate the proposed transformation techniques through objective measures. Gopala Krishna Anumanchipalli, Luís Caldas de Oliveira, Alan W. Black |
SLT | 3 |
| 2012 | Recovery of acronyms, out-of-lattice words and pronunciations from parallel multilingual speechabstractIn this work we present a set of techniques which explore information from multiple, different language versions of the same speech, to improve Automatic Speech Recognition (ASR) performance. Using this redundant information we are able to recover acronyms, words that cannot be found in the multiple hypotheses produced by the ASR systems, and pronunciations absent from their pronunciation dictionaries. When used together, the three techniques yield a relative improvement of 5.0% over the WER of our baseline system, and 24.8% relative when compared with standard speech recognition, in an Europarl Committee dataset with three different languages (Portuguese, Spanish and English). One full iteration of the system has a parallel Real Time Factor (RTF) of 3.08 and a sequential RTF of 6.44. João Miranda, João Paulo da Silva Neto, Alan W. Black |
SLT | 3 |
| 2011 | Discriminative Phrase-based Lexicalized Reordering Models using Weighted Reordering Graphs
Wang Ling, João Graça, David Martins de Matos, Isabel Trancoso, Alan W. Black |
IJCNLP | 5 |
| 2011 | A Statistical Phrase/Accent Model for Intonation ModelingabstractThis paper proposes a statistical phrase/accent model of voice fundamental frequency(F0) for speech synthesis. It presents an approach for automatic extraction and modeling of phrase and accent phenomena from F0 contours by taking into account their overall trends in the training data. An iterative optimization algorithm is described to extract these components, minimizing the reconstruction error of the F0 contour. This method of modeling local and global components of F0 separately is shown to be better than conventional F0 models used in Statistical Parametric Speech Synthesis (SPSS). Perceptual evaluations confirm that the proposed model is significantly better than baseline SPSS F0 models in 3 prosodically diverse tasks – read speech, radio broadcast speech and audio book speech. Gopala Krishna Anumanchipalli, Luís C. Oliveira, Alan W. Black |
INTERSPEECH | 3 |
| 2011 | Using Speaker ID to Discover Repeat Callers of a Spoken Dialog SystemabstractThis paper describes using speaker ID techniques to identify repeat callers in a spoken dialog system, using only acoustic features. Often it is useful to know if a dialog user is a novice or is experienced, and it can be the case that identifying data such as Caller ID is either unreliable or unavailable. Our approach attempts to remedy this by determining user identity in a dialog session using the acoustic information in the dialog. We optimize the audio content of each call by removing artifacts not relevant to modeling speech. This technique is applied to finding consecutive callers and creating unique user identities over all calls over a larger time frame, with the aim of tuning or adapting the dialog system based on the user identity. Our results show that the technique is effective in recognizing consecutive callers and in identifying a unique user identities in a large set of calls. Index Terms: spoken dialog, speaker ID Andrew Fandrianto, Brian Langner, Alan W. Black |
INTERSPEECH | 3 |
| 2011 | A Grammar Based Approach to Style Specific Phrase PredictionabstractWe present an approach to style specific phrasing for Text-to-Speech (TTS) systems. We formulate the problem of phrase break prediction (or phrasing) as generation of a sequence of breaks (B) and non-breaks (NB) after each word in a sentence. We use prosodic breaks in speech data to build shallow parses over corresponding text. We then learn a grammar that can predict these shallow prosodic parses from text. We then combine this prosodic phrasing information with other word level features in a CART tree to predict where phrase breaks should be inserted in new text. We show that a model built to target a specific reading style can predict phrase breaks more accurately than the standard generic model. Alok Parlikar, Alan W. Black |
INTERSPEECH | 2 |
| 2011 | Spoken Dialog Challenge 2010: Comparison of Live and Control Test Results
Alan W. Black, Susanne Burger, Alistair Conkie, Helen Hastie, Simon Keizer, Oliver Lemon, Nicolas Merigaud, Gabriel Parent, Gabriel Schubiner, Blaise Thomson, Jason D. Williams, Kai Yu 0004, Steve J. Young, Maxine Eskénazi |
SIGDIAL Conference | 1 |
| 2011 | Segmentation of Monologues in Audio Books for Building Synthetic VoicesabstractOne of the issues in using audio books for building a synthetic voice is the segmentation of large speech files. The use of the Viterbi algorithm to obtain phone boundaries on large audio files fails primarily because of huge memory requirements. Earlier works have attempted to resolve this problem by using large vocabulary speech recognition system employing restricted dictionary and language model. In this paper, we propose suitable modifications to the Viterbi algorithm and demonstrate its usefulness for segmentation of large speech files in audio books. The utterances obtained from large speech files in audio books are used to build synthetic voices. We show that synthetic voices built from audio books in the public domain have Mel-cepstral distortion scores in the range of 4-7, which is similar to voices built from studio quality recordings such as CMU ARCTIC. Kishore Prahallad, Alan W. Black |
IEEE Trans. Speech Audio Process. | 2 |
| 2010 | Evaluating a dialog language generation system: comparing the mountain system to other NLG approachesabstractThis paper describes the MOUNTAIN language generation system, a fully-automatic, data-driven approach to natural language generation aimed at spoken dialog applications. MOUN-TAIN uses statistical machine translation techniques and natural corpora to generate human-like language from a structured internal language, such as a representation of the dialog state. We briefly describe the training process for the MOUNTAIN approach, and show results of automatic evaluation in a standard language generation domain: the METEO weather forecasting corpus. Further, we compare output from the MOUNTAIN system to several other NLG systems in the same domain, using both automatic and human-based evaluation metrics; our results show our approach is comparable in quality to other advanced approaches. Finally, we discuss potential extensions, improvements, and other planned tests. Index Terms: natural language generation, evaluation, translation-based generation, weather forecasts, spoken dialog Brian Langner, Stephan Vogel, Alan W. Black |
INTERSPEECH | 3 |
| 2010 | Improving speech synthesis of machine translation outputabstractSpeech synthesizers are optimized for fluent natural text. However, in a speech to speech translation system, they have to process machine translation output, which is often not fluent. Rendering machine translations as speech makes them even harder to understand than the synthesis of natural text. A speech synthesizer must deal with the disfluencies in translations in order to be comprehensible and communicate the content. In this paper, we explore three synthesis strategies that address different problems found in translation output. By carrying out listening tasks and measuring transcription accuracies, we find that these methods can make the synthesis of translations more intelligible. Alok Parlikar, Alan W. Black, Stephan Vogel |
INTERSPEECH | 2 |
| 2010 | Towards Improving the Naturalness of Social Conversations with Dialogue Systems
Matthew Marge, João Miranda, Alan W. Black, Alexander I. Rudnicky |
SIGDIAL Conference | 3 |
| 2010 | Spoken Dialog Challenge 2010abstractThis paper describes the 2010 Spoken Dialog Challenge, a multi-site challenge to help investigate different spoken dialog system and evaluation techniques. The aim of the Challenge is to bring together multiple implementations of the same dialog task and deploy them in uncontrolled real user conditions and then make the results available for common evaluation techniques. This paper gives an overview of teh Challenge itself and the task, and presents the results for the “controlled task” part of the evaluation. The paper also discusses the infrastructure and organizational issues encountered and the solutions that made this challenge possible. Alan W. Black, Susanne Burger, Brian Langner, Gabriel Parent, Maxine Eskénazi |
SLT | 1 |
| 2010 | Spectral Mapping Using Artificial Neural Networks for Voice ConversionabstractIn this paper, we use artificial neural networks (ANNs) for voice conversion and exploit the mapping abilities of an ANN model to perform mapping of spectral features of a source speaker to that of a target speaker. A comparative study of voice conversion using an ANN model and the state-of-the-art Gaussian mixture model (GMM) is conducted. The results of voice conversion, evaluated using subjective and objective measures, confirm that an ANN-based VC system performs as good as that of a GMM-based VC system, and the quality of the transformed speech is intelligible and possesses the characteristics of a target speaker. In this paper, we also address the issue of dependency of voice conversion techniques on parallel data between the source and the target speakers. While there have been efforts to use nonparallel data and speaker adaptation techniques, it is important to investigate techniques which capture speaker-specific characteristics of a target speaker, and avoid any need for source speaker's data either for training or for adaptation. In this paper, we propose a voice conversion approach using an ANN model to capture speaker-specific characteristics of a target speaker and demonstrate that such a voice conversion approach can perform monolingual as well as cross-lingual voice conversion of an arbitrary source speaker. Srinivas Desai, Alan W. Black, Bayya Yegnanarayana, Kishore Prahallad |
IEEE Trans. Speech Audio Process. | 2 |
| 2009 | Pronunciation modeling for dialectal arabic speech recognitionabstractShort vowels in Arabic are normally omitted in written text which leads to ambiguity in the pronunciation. This is even more pronounced for dialectal Arabic where a single word can be pronounced quite differently based on the speaker's nationality, level of education, social class and religion. In this paper we focus on pronunciation modeling for Iraqi-Arabic speech. We introduce multiple pronunciations into the Iraqi speech recognition lexicon, and compare the performance, when weights computed via forced alignment are assigned to the different pronunciations of a word. Incorporating multiple pronunciations improved recognition accuracy compared to a single pronunciation baseline and introducing pronunciation weights further improved performance. Using these techniques an absolute reduction in word-error-rate of 2.4% was obtained compared to the baseline system. Hassan Al-Haj, Roger Hsiao, Ian Lane, Alan W. Black, Alex Waibel |
ASRU | 4 |
| 2009 | Speaker de-identification via voice transformationabstractIt is a common feature of modern automated voice-driven applications and services to record and transmit a user's spoken request. At the same time, several domains and applications may require keeping the content of the user's request confidential and at the same time preserving the speaker's identity. This requires a technology that allows the speaker's voice to be de-identified in the sense that the voice sounds natural and intelligible but does not reveal the identity of the speaker. In this paper we investigate different voice transformation strategies on a large population of speakers to disguise the speakers' identities while preserving the intelligibility of the voices. We apply two automatic speaker identification approaches to verify the success of de-identification with voice transformation, a GMM-based and a phonetic approach. The evaluation based on the automatic speaker identification systems verifies that the proposed voice transformation technique enables transmission of the content of the users' spoken requests while successfully preserving their identities. Also, the results indicate that different speakers still sound distinct after the transformation. Furthermore, we carried out a human listening test that proved the transformed speech to be both intelligible and securely de-identified, as it hid the identity of the speakers even to listeners who knew the speakers very well. Qin Jin, Arthur R. Toth, Tanja Schultz, Alan W. Black |
ASRU | 4 |
| 2009 | Optimizing segment label boundaries for statistical speech synthesisabstractThis paper introduces a new optimization technique for moving segment labels (phone and subphonetic) to optimize statistical parametric speech synthesis models. The choice of objective measures is investigated thoroughly and listening tests show the results to significantly improve the quality of the generated speech equivalent to increasing the database size by 3 fold. Alan W. Black, John Kominek |
ICASSP | 1 |
| 2009 | Voice conversion using Artificial Neural NetworksabstractIn this paper, we propose to use artificial neural networks (ANN) for voice conversion. We have exploited the mapping abilities of ANN to perform mapping of spectral features of a source speaker to that of a target speaker. A comparative study of voice conversion using ANN and the state-of-the-art Gaussian mixture model (GMM) is conducted. The results of voice conversion evaluated using subjective and objective measures confirm that ANNs perform better transformation than GMMs and the quality of the transformed speech is intelligible and has the characteristics of the target speaker. Srinivas Desai, E. Veera Raghavendra, Bayya Yegnanarayana, Alan W. Black, Kishore Prahallad |
ICASSP | 4 |
| 2009 | Voice convergin: Speaker de-identification by voice transformationabstractSpeaker identification might be a suitable answer to prevent unauthorized access to personal data. However we also need to provide solutions to secure transmission of spoken information. This challenge divides into two major aspects. First, the secure transmission of the content of the spoken input and second the secure transmission of the identity of the speaker. In this paper we concentrate on the latter, i.e. how to securely transmit information via voice without revealing the identity of the speaker to unauthorized listeners. In order to make the first steps toward solving this problem we study in this paper the potential of voice transformation for speaker de-identification. We use two speaker identification approaches to verify the success of de-identification with voice transformation, a GMM-based and a Phonetic approach, and study different voice transformation strategies to disguise speaker identity information while preserving understandability. Qin Jin, Arthur R. Toth, Tanja Schultz, Alan W. Black |
ICASSP | 4 |
| 2009 | The Spoken Dialogue Challenge
Alan W. Black, Maxine Eskénazi |
SIGDIAL Conference | 1 |
| 2009 | Statistical parametric speech synthesis
Heiga Zen, Keiichi Tokuda, Alan W. Black |
Speech Commun. | 3 |
| 2008 | Significance of early tagged contextual graphemes in grapheme based speech synthesis and recognition systemsabstractIn this paper we present our argument that context information could be used in early stages i.e., during the definition of mapping of the words into sequence of graphemes. We show that the early tagged contextual graphemes play a significant role in improving the performance of grapheme based speech synthesis and speech recognition systems. Gopala Krishna Anumanchipalli, Kishore Prahallad, Alan W. Black |
ICASSP | 3 |
| 2008 | Is voice transformation a threat to speaker identification?abstractWith the development of voice transformation and speech synthesis technologies, speaker identification systems are likely to face attacks from imposters who use voice transformed or synthesized speech to mimic a particular speaker. Therefore, we investigated in this paper how speaker identification systems perform on voice transformed speech. We conducted experiments with two different approaches, the classical GMM-based speaker identification system and the Phonetic speaker identification system. Our experimental results showed that current standard voice transformation techniques are able to fool the GMM-based system but not the Phonetic speaker identification system. These findings imply that future speaker identification systems should include idiosyncratic knowledge in order to successfully distinguish transformed speech from natural speech and thus be armed against imposter attacks. Qin Jin, Arthur R. Toth, Alan W. Black, Tanja Schultz |
ICASSP | 3 |
| 2008 | Let's go lab: a platform for evaluation of spoken dialog systems with real world users
Maxine Eskénazi, Alan W. Black, Antoine Raux, Brian Langner |
INTERSPEECH | 2 |
| 2008 | Improving speech systems built from very little dataabstractThis paper studies two ways for helping non-specialist users develop speech systems from limited data for new languages. Focused web re-crawling finds additional examples of text matching the domain as specified by the user. This improves the language model and cuts word error rate nearly in half. Iterative voice building with interleaved lexicon construction uses the voice from a previous iteration to help construct an improved voice. 4.5 hours of the user’s time reduces transcription error rate from 32 % to 4%. 1. John Kominek, Sameer Badaskar, Tanja Schultz, Alan W. Black |
INTERSPEECH | 4 |
| 2008 | Building sleek synthesizers for multi-lingual screen readerabstractIn this paper, we are investigating the unit size: syllable, half-phone and quarter-phone to be used for speech synthesis in multi-lingual screen reader in phonetic languages such as Telugu and non-phonetic language English. Perceptual studies show that syllable-level unit performs better for Telugu and half-phone units perform better for English. While syllable based synthesizers produce better sounding speech, the coverage of all syllables is a non-trivial issue. We address the issue of coverage of syllables through approximate matching of syllable and show that such approximation produces intelligible and better quality speech than diphone units. In this paper, we also propose a hybrid synthesizer within the framework of unit selection and also show that the hybrid synthesizer built from pruned database performs as well as hybrid synthesizer built from unpruned database. Index Terms: speech synthesis, unit selection, unit size, database pruning, and hybrid speech synthesis. E. Veera Raghavendra, Bayya Yegnanarayana, Alan W. Black, Kishore Prahallad |
INTERSPEECH | 3 |
| 2008 | Incorporating durational modification in voice transformationabstractVoice transformation is the process of using a small amount of speech data from a target speaker to build a transformation model that can be used to generate arbitrary speech that sounds like the target speaker. One common current technique is building Gausian Mixture Models to map spectral aspects from source to target speakers. This paper proposes the use of duration models to improve the transformation models and output speech quality. Testing across seven target speakers shows a statistically significant improvement in a popular objective metric when duration modification is performed both during training and testing of a Gaussian Mixture Model mapping based voice transformation system. Index Terms: voice transformation, speech synthesis 1. Arthur R. Toth, Alan W. Black |
INTERSPEECH | 2 |
| 2008 | NineOneOne: Recognizing and Classifying Speech for Handling Minority Language Emergency Calls
Udhyakumar Nallasamy, Alan W. Black, Tanja Schultz, Robert E. Frederking |
LREC | 2 |
| 2008 | Global syllable set for building speech synthesis in Indian languagesabstractIndian languages are syllabic in nature where many syllables are found common across its languages. This motivates us to build a global syllable set by combining multiple language syllables to build a synthesizer which can borrow units from a different language when the required syllable is not found. Such synthesizer make use of speech database in different languages spoken by different speakers, whose output is likely to pick units from multiple languages and hence the synthesized utterance contains units spoken by multiple speakers which would annoy the user. We intend to use a cross lingual Voice Conversion framework using Artificial Neural Networks (ANN) to transform such an utterance to a single target speaker. E. Veera Raghavendra, Srinivas Desai, Bayya Yegnanarayana, Alan W. Black, Kishore Prahallad |
SLT | 4 |
| 2008 | Statistical mapping between articulatory movements and acoustic spectrum using a Gaussian mixture model
Tomoki Toda, Alan W. Black, Keiichi Tokuda |
Speech Commun. | 2 |
| 2007 | Statistical Parametric Speech SynthesisabstractThis paper gives a general overview of techniques in statistical parametric speech synthesis. One of the instances of these techniques, called HMM-based generation synthesis (or simply HMM-based synthesis), has recently been shown to be very effective in generating acceptable speech synthesis. This paper also contrasts these techniques with the more conventional unit selection technology that has dominated speech synthesis over the last ten years. Advantages and disadvantages of statistical parametric synthesis are highlighted as well as identifying where we expect the key developments to appear in the immediate future. Alan W. Black, Heiga Zen, Keiichi Tokuda |
ICASSP (4) | 1 |
| 2007 | ugloss: a framework for improving spoken language generation understandabilityabstractUnderstandable spoken presentation of structured and complex information is a difficult task to do well. As speech synthesis is used in more applications, there is likely to be an increasing requirement to present complex information in an understandable manner. This paper introduces uGloss, a language generation framework designed to influence the understandability of spoken output. We describe relevant factors to its design and provide a general description of our algorithm. We compare our approach to human performance for a straightforward task, and discuss areas of improvement and our future goals for this work. Brian Langner, Alan W. Black |
INTERSPEECH | 2 |
| 2007 | Automatic building of synthetic voices from large multi-paragraph speech databasesabstractLarge multi paragraph speech databases encapsulate prosodic and contextual information beyond the sentence level which could be exploited to build natural sounding voices. This paper discusses our efforts on automatic building of synthetic voices from large multi-paragraph speech databases. We show that the primary issue of segmentation of large speech file could be addressed with modifications to forced-alignment technique and that the proposed technique is independent of the duration of the audio file. We also discuss how this framework could be extended to build a large number of voices from public domain large multi-paragraph recordings. Index Terms: speech synthesis, large multi-paragraph speech databases, forced-alignment, public domain recordings Kishore Prahallad, Arthur R. Toth, Alan W. Black |
INTERSPEECH | 3 |
| 2007 | SPICE: web-based tools for rapid language adaptation in speech processing systemsabstractIn this paper we describe the design and implementation of a user interface for SPICE, a web-based toolkit for rapid prototyping of speech and language processing components.We report on the challenges and experiences gathered from testing these tools in an advanced graduate hands-on course, in which we created speech recognition, speech synthesis, and smalldomain translation components for 10 different languages within only 6 weeks. Tanja Schultz, Alan W. Black, Sameer Badaskar, Matthew Hornyak, John Kominek |
INTERSPEECH | 2 |
| 2007 | Voice Conversion Based on Maximum-Likelihood Estimation of Spectral Parameter TrajectoryabstractIn this paper, we describe a novel spectral conversion method for voice conversion (VC). A Gaussian mixture model (GMM) of the joint probability density of source and target features is employed for performing spectral conversion between speakers. The conventional method converts spectral parameters frame by frame based on the minimum mean square error. Although it is reasonably effective, the deterioration of speech quality is caused by some problems: 1) appropriate spectral movements are not always caused by the frame-based conversion process, and 2) the converted spectra are excessively smoothed by statistical modeling. In order to address those problems, we propose a conversion method based on the maximum-likelihood estimation of a spectral parameter trajectory. Not only static but also dynamic feature statistics are used for realizing the appropriate converted spectrum sequence. Moreover, the oversmoothing effect is alleviated by considering a global variance feature of the converted spectra. Experimental results indicate that the performance of VC can be dramatically improved by the proposed method in view of both speech quality and conversion accuracy for speaker individuality. Tomoki Toda, Alan W. Black, Keiichi Tokuda |
IEEE Trans. Speech Audio Process. | 2 |
| 2006 | Pocketsphinx: A Free, Real-Time Continuous Speech Recognition System for Hand-Held DevicesabstractThe availability of real-time continuous speech recognition on mobile and embedded devices has opened up a wide range of research opportunities in human-computer interactive applications. Unfortunately, most of the work in this area to date has been confined to proprietary software, or has focused on limited domains with constrained grammars. In this paper, we present a preliminary case study on the porting and optimization of CMU Sphinx-11, a popular open source large vocabulary continuous speech recognition (LVCSR) system, to hand-held devices. The resulting system operates in an average 0.87 times real-time on a 206 MHz device, 8.03 times faster than the baseline system. To our knowledge, this is the first hand-held LVCSR system available under an open-source license David Huggins-Daines, Arthur Chan, Alan W. Black, Mosur Ravishankar, Alexander I. Rudnicky |
ICASSP (1) | 4 |
| 2006 | Sub-Phonetic Modeling For Capturing Pronunciation Variations For Conversational Speech SynthesisabstractIn this paper we address the issue of pronunciation modeling for conversational speech synthesis. We experiment with two different HMM topologies (fully connected state model and forward connected state model) for sub-phonetic modeling to capture the deletion and insertion of sub-phonetic states during speech production process. We show that the experimented HMM topologies have higher log likelihood than the traditional 5-state sequential model. We also study the first and second mentions of content words and their influence on the pronunciation variation. Finally we report phone recognition experiments using the modified HMM topologies Kishore Prahallad, Alan W. Black, Mosur Ravishankar |
ICASSP (1) | 2 |
| 2006 | Challenges with Rapid Adaptation of Speech Translation Systems to New Language PairsabstractAlthough we have far from solved the issues in porting speech translation systems to new languages, we have gathered sufficient experience by now to identify a number of major challenges in the process. Although well-defined processes exist for building speech recognition, speech synthesis and statistical machine translation models, they still require both significant native speaker involvement and linguistic expertise. As the core technology improves we believe we will see increasing cultural and social issues in contributions from native speakers. This paper identifies some of these issues and presents our initial attempts to build tools that we hope will eventually allow linguistically naive native informants build complete speech translation systems Tanja Schultz, Alan W. Black |
ICASSP (5) | 2 |
| 2006 | Text-Independent Voice Conversion Based on Unit SelectionabstractSo far, most of the voice conversion training procedures are text-dependent, i.e., they are based on parallel training utterances of source and large speaker. Since several applications (e.g. speech-to-speech translation or dubbing) require text-independent training, over the last two years, training techniques that use non-parallel data were proposed In this paper, we present a new approach that applies unit selection to find corresponding time frames in source and target speech. By means of a subjective experiment it is shown that this technique achieves the same performance as the conventional text-dependent training David Suendermann-Oeft, Harald Höge, Antonio Bonafonte, Hermann Ney, Alan W. Black, Shri Narayanan |
ICASSP (1) | 5 |
| 2006 | Visual Evaluation of Voice Transformation Based on Knowledge of SpeakerabstractVoice transformation techniques are maturing. The ability to automatically change a source voice to a target voice, using a model built from a small amount of target speaker data, has brought with it the need to better understand how to evaluate the quality of a transformation model. This paper presents a simple experiment to measure how familiarity with the particular source and target speakers affects perception of the transformation. The results show that listeners' views of the transformation are not affected by familiarity with the speakers. In addition to these results, we also introduce transformation triangle diagrams, a graphical mechanism to better display certain relationships that are important in the evaluation of voice transformation Arthur R. Toth, Alan W. Black |
ICASSP (1) | 2 |
| 2006 | CLUSTERGEN: a statistical parametric synthesizer using trajectory modelingabstractUnit selection synthesis has shown itself to be capable of producing high quality natural sounding synthetic speech when constructed from large databases of well-recorded, well-labeled speech. However, the cost in time and expertise of building such voices is still too expensive and specialized to be able to build individual voices for everyone. The quality in unit selection synthesis is directly related to the quality and size of the database used. As we require our speech synthesizers to have more variation, style and emotion, for unit selection synthesis, much larger databases will be required. As an alternative, more recently we have started looking for parametric models for speech synthesis, that are still trained from databases of natural speech but are more robust to errors and allow for better modeling of variation. This paper presents the CLUSTERGEN synthesizer which is implemented within the Festival/FestVox voice building environment. As well as the basic technique, three methods of modeling dynamics in the signal are presented and compared: a simple point model, a basic trajectory model and a trajectory model with overlap and add. Index Terms: speech synthesis, statistical parametric synthesis, trajectory HMMs. Alan W. Black |
INTERSPEECH | 1 |
| 2006 | Optimizing components for handheld two-way speech translation for an English-iraqi Arabic systemabstractThis paper described our handheld two-way speech translation system for English and Iraqi. The focus is on developing a field usable handheld device for speech-to-speech translation. The computation and memory limitations on the handheld impose critical constraints on the ASR, SMT, and TTS components. In this paper we discuss our approaches to optimize these components for the handheld device and present performance numbers from the evaluations that were an integral part of the project. Since one major aspect of the TransTac program is to build fieldable systems, we spent significant effort on developing an intuitive interface that minimizes the training time for users but also provides useful information such as back translations for translation quality feedback. Roger Hsiao, Ashish Venugopal, Thilo Köhler, Ying Zhang 0048, Paisarn Charoenpornsawat, Andreas Zollmann, Stephan Vogel, Alan W. Black, Tanja Schultz, Alex Waibel |
INTERSPEECH | 8 |
| 2006 | Generating time-constrained audio presentations of structured informationabstractPresenting complex information in an understandable manner using speech is a challenging task to do well. Significant limitations, both in the generation process and from the human listeners’ capabilities, typically make for poorly understood speech. This work examines possible strategies for producing understandable spoken complex information working within those limitations, as well as identifying ways to improve systems to reduce the limitations’ impact. We discuss a simple user study that explores these strategies with complex structured information, and describe a spoken dialog system that will make use of this work to provide a speech interface to structured information in a more understandable manner. Index Terms: speech synthesis, natural language generation, information presentation. Brian Langner, Rohit Kumar 0001, Arthur Chan, Lingyun Gu, Alan W. Black |
INTERSPEECH | 5 |
| 2006 | Doing research on a deployed spoken dialogue system: one year of let's go! experienceabstractThis paper describes our work with Let’s Go, a telephonebased bus schedule information system that has been in use by the Pittsburgh population since March 2005. Results from several studies show that while task success correlates strongly with speech recognition accuracy, other aspects of dialogue such as turn-taking, the set of error recovery strategies, and the initiative style also significantly impact system performance and user behavior. Index Terms: spoken dialogue systems, real-world applications, speech recognition Antoine Raux, Dan Bohus, Brian Langner, Alan W. Black, Maxine Eskénazi |
INTERSPEECH | 4 |
| 2006 | Intelligibility of machine translation output in speech synthesisabstractOne use of text-to-speech synthesis (TTS) is as a component of speech-to-speech translation systems. The output of automatic machine translation (MT) can vary widely in quality, however. A synthetic voice that is extremely intelligible on naturally-occurring text may be far less intelligible when asked to render text that is automatically generated. In this paper, we compare the quality of synthesis of naturally-occurring text and its MT counterpart. We find that intelligibility of TTS on MT output is significantly lower than on either naturally-occurring text or semantically unpredictable sentences, and explore the reasons why. Laura Mayfield Tomokiyo, Kay Peterson, Alan W. Black, Kevin A. Lenzo |
INTERSPEECH | 3 |
| 2006 | Learning Pronunciation Dictionaries: Language Complexity and Word Selection Strategies
John Kominek, Alan W. Black |
HLT-NAACL | 2 |
| 2006 | Online Supervised Learning of Non-Understanding Recovery PoliciesabstractSpoken dialog systems typically use a limited number of non- understanding recovery strategies and simple heuristic policies to engage them (e.g. first ask user to repeat, then give help, then transfer to an operator). We propose a supervised, online method for learning a non-understanding recovery policy over a large set of recovery strategies. The approach consists of two steps: first, we construct runtime estimates for the likelihood of success of each recovery strategy, and then we use these estimates to construct a policy. An experiment with a publicly available spoken dialog system shows that the learned policy produced a 12.5% relative improvement in the non-understanding recovery rate. Dan Bohus, Brian Langner, Antoine Raux, Alan W. Black, Maxine Eskénazi, Alexander I. Rudnicky |
SLT | 4 |
| 2006 | Flexible speech translation systemsabstractSpeech translation research has made significant progress over the years with many high-visibility efforts showing that translation of spontaneously spoken speech from and to diverse languages is possible and applicable in a variety of domains. As language and domains continue to expand, practical concerns such as portability and reconfigurability of speech come into play: system maintenance becomes a key issue and data is never sufficient to cover the changing domains over varying languages. In this paper, we discuss strategies to overcome the limits of today's speech translation systems. In the first part, we describe our layered system architecture that allows for easy component integration, resource sharing across components, comparison of alternative approaches, and the migration toward hybrid desktop/PDA or stand-alone PDA systems. In the second part, we show how flexibility and reconfigurability is implemented by more radically relying on learning approaches and use our English–Thai two-way speech translation system as a concrete example. Tanja Schultz, Alan W. Black, Stephan Vogel, Monika Woszczyna |
IEEE Trans. Speech Audio Process. | 2 |
| 2005 | Prediction of Pronunciation Variations for Speech Synthesis: A Data-Driven ApproachabstractThe fact that speakers vary pronunciations of the same word within their own speech is well known, but little has been done to categorize and predict a speaker's pronunciation distribution automatically for unit selection speech synthesis. Recent work demonstrated how to identify automatically a speaker's choice between full and reduced pronunciations using acoustic modeling techniques from speech recognition. We extend this approach and show how its results can be used to predict a speaker's choice of pronunciations for synthesis. We apply machine learning techniques to the automatically categorized data to produce a pronunciation variation prediction model given only the utterance text - allowing the system to synthesize novel phrases with variations like those the speaker would make. Empirical studies emphasize that we can improve automatic pronunciation labels and successfully utilize the results for prediction of future synthesized examples. The prediction results based on these automatic labels are very similar to those trained from human labeled data - allowing us to reduce manual effort while still achieving comparable results. Christina L. Bennett, Alan W. Black |
ICASSP (1) | 2 |
| 2005 | Improving the Understandability of Speech Synthesis by Modeling Speech in NoiseabstractAlthough the quality of synthetic speech has increased dramatically in the past several years, many people still have difficulty understanding speech produced by even the highest quality synthesizers. We describe an approach to improve understandability of synthetic speech using speech in noise. Natural speech in noise is a change in the style of speech that is used by people to improve the understandability to the listener when speaking in poor channel conditions. We show that altering the presentation of synthetic speech in similar ways also improves understandability. Further, we discuss methods of obtaining speech in noise for use in speech synthesis, as well as the results of an evaluation of several synthetic voices that "speak in noise". Brian Langner, Alan W. Black |
ICASSP (1) | 2 |
| 2005 | Thai Automatic Speech RecognitionabstractWe describe the development of a robust and flexible Thai speech recognizer as integrated into our English-Thai speech-to-speech translation system. We focus on the discussion of the rapid deployment of ASR for Thai under limited time and data resources, including rapid data collection issues, acoustic model bootstrap, and automatic generation of pronunciations. Issues relating to the translation and overall system will be reported elsewhere. Sinaporn Suebvisai, Paisarn Charoenpornsawat, Alan W. Black, Monika Woszczyna, Tanja Schultz |
ICASSP (1) | 3 |
| 2005 | Spectral Conversion Based on Maximum Likelihood Estimation Considering Global Variance of Converted ParameterabstractThe paper describes a novel spectral conversion method for voice transformation. We perform spectral conversion between speakers using a Gaussian mixture model (GMM) on the joint probability density of source and target features. A smooth spectral sequence can be estimated by applying maximum likelihood (ML) estimation to the GMM-based mapping using dynamic features. However, there is still degradation of the converted speech quality due to an over-smoothing of the converted spectra, which is inevitable in conventional ML-based parameter estimation. In order to alleviate the over-smoothing, we propose an ML-based conversion taking account of the global variance of the converted parameter in each utterance. Experimental results show that the performance of the voice conversion can be improved by using the global variance information. Moreover, it is demonstrated that the proposed algorithm is more effective than spectral enhancement by postfiltering. Tomoki Toda, Alan W. Black, Keiichi Tokuda |
ICASSP (1) | 2 |
| 2005 | The blizzard challenge - 2005: evaluating corpus-based speech synthesis on common datasetsabstractIn order to better understand different speech synthesis techniques on a common dataset, we devised a challenge that will help us better compare research techniques in building corpusbased speech synthesizers. In 2004, we released the first two 1200-utterance single-speaker databases from the CMU ARC-TIC speech databases, and challenged current groups working in speech synthesis around the world to build their best voices from these databases. In January of 2005, we released two further databases and a set of 50 utterance texts from each of five genres and asked the participants to synthesize these utterances. Their resulting synthesized utterances were then presented to three groups of listeners: speech experts, volunteers, and US English-speaking undergraduates. This paper summarizes the purpose, design, and whole process of the challenge. Alan W. Black, Keiichi Tokuda |
INTERSPEECH | 1 |
| 2005 | Measuring unsupervised acoustic clustering through phoneme pair merge-and-split testsabstractSubphonetic discovery through segmental clustering is a central step in building a corpus-based synthesizer. To help decide what clustering algorithm to use we employed mergeand-split tests on English fricatives. Compared to reference of 2%, Gaussian EM achieved a misclassification rate of 6%, Kmeans 10%, while predictive CART trees performed poorly. John Kominek, Alan W. Black |
INTERSPEECH | 2 |
| 2005 | Let's go public! taking a spoken dialog system to the real worldabstractIn this paper, we describe how a research spoken dialog system was made available to the general public. The Let’s Go Public spoken dialog system provides bus schedule information to the Pittsburgh population during off-peak times. This paper describes the changes necessary to make the system usable for the general public and presents analysis of the calls and strategies we have used to ensure high performance. 1. Antoine Raux, Brian Langner, Dan Bohus, Alan W. Black, Maxine Eskénazi |
INTERSPEECH | 4 |
| 2005 | Foreign accents in synthetic speech: development and evaluationabstractThis paper addresses the generation and evaluation of foreign-accented speech in concatenative text-to-speech (TTS) synthesis. We describe three possible methods of building a Spanish-accented English voice, and evaluate and compare them with respect to preference, intelligibility, and smoothness. Effects of speaking rate and content are also examined. It is found that although using an unmodified Spanish voice to read English text is possible, the result is not highly intelligible. With some modifications to the linguistic model, a relatively high level of comprehensibility and smoothness can be achieved, not differing widely from ratings given to a native voice at a comparable stage of development. Listeners in perceptual experiments were very consistent in their preference rankings of the three voices, showing that differences in voicebuilding method are both detectable and contribute to synthesis quality. 1. Laura Mayfield Tomokiyo, Alan W. Black, Kevin A. Lenzo |
INTERSPEECH | 2 |
| 2005 | Cross-speaker articulatory position data for phonetic feature predictionabstractThrough the use of a device called an Electromagnetic Articulograph, it is possible to measure the locations of a person’s articulators during speech. As more of this data becomes available, one important question is how it can be used. In this paper, we demonstrate that it can improve performance for the recognition of some phonetic features. As articulatory position data is scarce, we also describe experiments that use articulatory position data from one speaker with another and provide results. These experiments use cross-speaker articulatory positions to predict phonetic features. Arthur R. Toth, Alan W. Black |
INTERSPEECH | 2 |
| 2004 | Multilingual text-to-speech synthesisabstractThe paper presents a framework for building multilingual text-to-speech systems. It addresses the issue at three levels. First it discusses the necessary steps required to build a synthetic voice from scratch in a new language. The second concerns the building of a new voice without recording any new acoustic data, and the restrictions that imposes. The third more speculative part discusses the steps that would be necessary to allow high quality synthesis of new languages by recording only minimal amounts in that language. Alan W. Black, Kevin A. Lenzo |
ICASSP (3) | 1 |
| 2004 | A family-of-models approach to HMM-based segmentation for unit selection speech synthesisabstractFor segmenting a speech database, using a family of acoustic models provides multiple estimates of each boundary point. This is more robust than a single estimate because by taking consensus values, large labeling errors are less prevalent in the synthesis catalog, which improves the resulting voice. This paper describes HMM-based segmentation in which up to 500 related models are applied to each wavefile. In a listening test of twelve utterances, human judges preferred the proposed technique over the baseline by a tally of 6 to 2, with 4 ties. John Kominek, Alan W. Black |
INTERSPEECH | 2 |
| 2004 | Boostrapping phonetic lexicons for new languagesabstractAlthough phonetic lexicons are critical for many speech applications, the process of building one for a new language can take a significant amount of time and effort. We present a bootstrapping algorithm to build phonetic lexicons for new languages. Our method relies on a large amount of unlabeled text, a small set of ’seed words’ with their phonetic transcription, and the proficiency of a native speaker in correctly inspecting the generated pronunciations of the words. The method proceeds by automatically building Letter-to-Sound (LTS) rules from a small set of the most commonly occurring words in a large corpus of a given language. These LTS rules are retrained as new words are added to the lexicon in an Active Learning step. This procedure is repeated until we have a lexicon that can predict the pronunciation of any word in the target language with the accuracy desired. We tested our approach for three languages: English, Sameer Maskey, Alan W. Black, Laura Tomokiya |
INTERSPEECH | 2 |
| 2004 | Acoustic-to-articulatory inversion mapping with Gaussian mixture modelabstractThis paper describes the acoustic-to-articulatory inversion mapping using a Gaussian Mixture Model (GMM).Correspondence of an acoustic parameter and an articulatory parameter is modeled by the GMM trained using the parallel acousticarticulatory data.We measure the performance of the GMMbased mapping and investigate the effectiveness of using multiple acoustic frames as an input feature and using multiple mixtures.As a result, it is shown that although increasing the number of mixtures is useful for reducing the estimation error, it causes many discontinuities in the estimated articulatory trajectories.In order to address this problem, we apply maximum likelihood estimation (MLE) considering articulatory dynamic features to the GMM-based mapping.Experimental results demonstrate that the MLE using dynamic features can estimate more appropriate articulatory movements compared with the GMM-based mapping applied smoothing by lowpass filter. Tomoki Toda, Alan W. Black, Keiichi Tokuda |
INTERSPEECH | 2 |
| 2003 | Using acoustic models to choose pronunciation variations for synthetic voices
Christina L. Bennett, Alan W. Black |
INTERSPEECH | 2 |
| 2003 | Unit selection and emotional speechabstractUnit Selection Synthesis, where appropriate units are selected from large databases of natural speech, has greatly improved the quality of speech synthesis. But the quality improvement has come at a cost. The quality of the synthesis relies on the fact that little or no signal processing is done on the selected units, thus the style of the recording is maintained in the quality of the synthesis. The synthesis style is implicitly the style of the database. If we want more general flexibility we have to record more data of the desired style. Which means that our already large unit databases must be made even larger. This paper Alan W. Black |
INTERSPEECH | 1 |
| 2003 | Unit size in unit selection speech synthesisabstractIn this paper, we address the issue of choice of unit size in unit selection speech synthesis. We discuss the development of a Hindi speech synthesizer and our experiments with different choices of units: syllable, diphone, phone and half phone. Perceptual tests conducted to evaluate the quality of the synthesizers with different unit size indicate that the syllable synthesizer performs better than the phone, diphone and half phone synthesizers, and the half phone synthesizer performs better than diphone and phone synthesizers. Kishore Prahallad, Alan W. Black |
INTERSPEECH | 2 |
| 2003 | Evaluating and correcting phoneme segmentation for unit selection synthesisabstractAs part of improved support for building unit selection voices, the Festival speech synthesis system now includes two algorithms for automatic labeling of wavefile data. The two methods are based on dynamic time warping and HMM-based acoustic modeling. Our experiments show that DTW is more accurate 70% of the time, but is also more prone to gross labeling errors. HMM modeling exhibits a systematic bias of 15 ms. Combining both methods directs human labelers towards data most likely to be problematic. John Kominek, Christina L. Bennett, Alan W. Black |
INTERSPEECH | 3 |
| 2003 | LET's GO: improving spoken dialog systems for the elderly and non-nativesabstractWith the recent improvements in speech technology, it is now possible to build spoken dialog systems that basically work. However, such systems are designed and tailored for the general population. When users come from less general sections of the population, such as the elderly and non-native speakers of English, the accuracy of dialog systems degrades. This paper describes Let’s Go, a dialog system specifically designed to allow dialog experiments to be carried out on the elderly and non-native speakers in order to better tune such systems for these important populations. Let’s Go is designed to provide Pittsburgh area bus information. The basic system is described and our initial experiments are outlined. Antoine Raux, Brian Langner, Alan W. Black, Maxine Eskénazi |
INTERSPEECH | 3 |
| 2003 | Arabic in my hand: small-footprint synthesis of egyptian arabicabstract... of synthesis of Arabic, a language that has shot to prominence in the past few years, and synthesis on a handheld device, realization of which presents difficult software engineering problems. Our system was developed in conjunction with the DARPA BABYLON project, and has been integrated with English synthesis, English and Arabic ASR, and machine translation on a single off-the-shelf PDA. We present Laura Mayfield Tomokiyo, Alan W. Black, Kevin A. Lenzo |
INTERSPEECH | 2 |
| 2003 | Speechalator: two-way speech-to-speech translation on a consumer PDAabstractThis paper describes a working two-way speech-to-speech translation system that runs in near real-time on a consumer handheld computer. It can translate from English to Arabic and Arabic to English in the domain of medical interviews. We describe the general architecture and frameworks within which we developed each of the components: HMM-based recognition, interlingua translation (both rule and statistically based), and unit selection synthesis. Alex Waibel, Ahmed Badran, Alan W. Black, Robert E. Frederking, Donna Gates, Alon Lavie, Lori S. Levin, Kevin A. Lenzo, Laura Mayfield Tomokiyo, Jürgen Reichert, Tanja Schultz, Dorcas Wallace, Monika Woszczyna, Jing Zhang 0011 |
INTERSPEECH | 3 |
| 2003 | Identifying speakers in children's stories for speech synthesisabstractChoosing appropriate voices for synthesizing children’s stories requires text analysis techniques that can identify which portions of the text should be read by which speakers. Our work presents techniques to take raw text stories and automatically identify the quoted speech, identify the characters within the stories and assign characters to each quote. The resulting marked-up story may then be rendered with a standard speech synthesizer with appropriate voices for the characters. This paper presents each of the basic stages in identification, and the algorithms, both rule-driven and data-driven, used to achieve this. A variety of story texts are used to test our system. Results are presented with a discussion of the limitations and recommendations on how to improve speaker assignment in further texts. Jason Y. Zhang 0002, Alan W. Black, Richard Sproat |
INTERSPEECH | 2 |
| 2003 | Speechalator: Two-Way Speech-to-Speech Translation in Your Hand
Alex Waibel, Ahmed Badran, Alan W. Black, Robert E. Frederking, Donna Gates, Alon Lavie, Lori S. Levin, Kevin A. Lenzo, Laura Mayfield Tomokiyo, Jürgen Reichert, Tanja Schultz, Dorcas Wallace, Monika Woszczyna, Jing Zhang 0011 |
HLT-NAACL | 3 |
| 2002 | Building voiceXML-based applicationsabstractThe Language Technologies Institute (LTI) at Carnegie Mellon University has, for the past several years, conducted a lab course in building spoken-language dialog systems. In the most recent versions of the course, we have used (commercial) web-based development environments to build systems. This paper describes our experiences and discusses the characteristics of applications that are developed within this framework. Christina L. Bennett, Ariadna Font Llitjós, Stefanie Shriver, Alexander I. Rudnicky, Alan W. Black |
INTERSPEECH | 5 |
| 2002 | Rapid development of speech-to-speech translation systems
Alan W. Black, Ralf D. Brown, Robert E. Frederking, Kevin A. Lenzo, John Moody, Alexander I. Rudnicky, Rita Singh, Eric Steinbrecher |
INTERSPEECH | 1 |
| 2002 | Field Testing the Tongues Speech-to-Speech Machine Translation System
Robert E. Frederking, Alan W. Black, Ralf D. Brown, John Moody, Eric Steinbrecher |
LREC | 2 |
| 2002 | Evaluation and collection of proper name pronunciations online
Ariadna Font Llitjós, Alan W. Black |
LREC | 2 |
| 2001 | A study on speech over the telephone and agingabstractWe describe an experiment to show how the comprehensibility of speech over the telephone is related to the age of the listener. Our intention is to show figures to prove the commonly-held belief that as we get older our hearing of information over the telephone degrades. The study was set up to determine, for all age groups from 20-29 to 80-89, whether comprehension degrades with age and with the type of speech (synthetic or natural) . We gave subjects sentences containing target word pairs that they were to write down. The pairs contained more or less predictable words. Maxine Eskénazi, Alan W. Black |
INTERSPEECH | 2 |
| 2001 | Knowledge of language origin improves pronunciation accuracy of proper namesabstractAs it is impossible to have a lexicon with complete coverage, and a high proportion of unknown words are proper names, this paper addresses the issue of automatically finding pronunciations of unseen proper names in US English. Proper names, especially in the US, may come from a large range of ethnic backgrounds. We present a model and results showing that including ethnic origin of words in a statistical model can improve pronunciation results. Ariadna Font Llitjós, Alan W. Black |
INTERSPEECH | 2 |
| 2001 | Normalization of non-standard words
Richard Sproat, Alan W. Black, Stanley F. Chen, Shankar Kumar, Mari Ostendorf, Christopher Richards |
Comput. Speech Lang. | 2 |
| 2001 | Heterogeneous relation graphs as a formalism for representing linguistic information
Paul Taylor 0001, Alan W. Black, Richard Caley |
Speech Commun. | 2 |
| 2000 | Limited domain synthesis
Alan W. Black, Kevin A. Lenzo |
INTERSPEECH | 1 |
| 2000 | Statistically trained orthographic to sound models for Thai
Ananlada Chotimongkol, Alan W. Black |
INTERSPEECH | 2 |
| 2000 | Diphone collection and synthesisabstractIn this paper, we describe the design and collection of corpora for diphone synthesis, the voice building process, and our experience in the creation of a new, publically available database of ten diphone sets of one American English speaker for the Festival Speech Synthesis System [3], using the FestVox document and tools [1]. In support of our goal to make the tools and techniques available for anyone to build their own synthetic voices, we have generalized and streamlined the tasks involved from what were once arcane anecdotes, half-written one-off scripts, and partial descriptions, to detailed, complete instructions that others have followed with good results. 1. INTRODUCTION The FestVox [1] document is a growing, publically available resource that contains tools, data, and text about building complete synthetic voices in English and other languages. That work covers everything from building text analyzers, lexicons, prosodic models as well as various waveform synthesis technique... Kevin A. Lenzo, Alan W. Black |
INTERSPEECH | 2 |
| 2000 | Non-standard word and homograph resolution for asian language text analysisabstractIn this paper we present a general model for text analysis of Asian languages (Chinese and Japanese).That is a method for mapping strings of characters to strings of identified trivially pronounceable words.This work is based on the English Non-Standard Word analysis model suitably augmented to deal with both the lack of spaces between words in Japanese and Chinese and addressing the issues of homographs.Results are present for the sub-components of the process. Craig Olinsky, Alan W. Black |
INTERSPEECH | 2 |
| 2000 | Towards a universal speech interfaceabstractWe discuss our ongoing attempt to design and evaluate universal human-machine speech-based interfaces. We describe one such initial design suitable for database retrieval applications, and discuss its implementation in a movie information application prototype. Initial user studies provided encouraging results regarding the usability of the design, as well as suggest some questions for further investigation. 1. INTRODUCTION Speech recognition technology has made spoken interaction with machines feasible. However, no suitable universal interaction paradigm has yet been proposed for humans to communicate effectively, efficiently and effortlessly by voice with machines. On one hand, natural language applications have been demonstrated in narrow domains, but building such systems is data-, labor- and expertise-intensive. Perhaps more importantly, unconstrained natural language severely strains recognition technology, and fails to delineate the functional limitations of the machine. On th... Ronald Rosenfeld, Xiaojin Zhu 0001, Arthur R. Toth, Stefanie Shriver, Kevin A. Lenzo, Alan W. Black |
INTERSPEECH | 6 |
| 2000 | Task and domain specific modelling in the Carnegie Mellon communicator systemabstractThe Carnegie Mellon Communicator is a telephone-based dialog system that supports planning in a travel domain. The implementation of such a system requires two complimentary components, an architecture capable of managing interaction and the task, as well as a knowledge base that captures the speech, language and task characteristics specific to the domain. Given a suitable architecture, the principal effort in development in taken up in the acquisition and processing of a domain knowledge base. This paper describes a variety of techniques we have applied to modeling in acoustic, language, task, generation and synthesis components of the system. 1. INTRODUCTION System development involves a great deal of knowledge engineering, which is both time-consuming and requires a variety of experts to participate in the process. Therefore methods that seek to minimize this resource, for example through training based on domain-specific corpora are preferred. Effective use of corpora, however, ... Alexander I. Rudnicky, Christina L. Bennett, Alan W. Black, Ananlada Chotimongkol, Kevin A. Lenzo, Alice Oh, Rita Singh |
INTERSPEECH | 3 |
| 2000 | Audio signals in speech interfacesabstractThis paper discusses a variety of types of non-lexical signals such as beeps, prosodic variation and speaker style changes, and we consider four cases in which such signals might be used to good effect. We discuss the results of user tests to determine if specific types of non-lexical signals are better in some situations than in others, and we discuss the advantages and disadvantages of using such signals. 1. INTRODUCTION This paper discusses options for including non-lexical cues in audio output in order to convey pertinent information to users of speech interface systems. By non-lexical cues we mean any noises or supra-lexical features such as prosody or pitch which can be inserted or altered in an otherwise lexical string. These non-lexical cues can be arbitrary (e.g. a beep to suggest that an item is optional) or non-arbitrary (e.g. a ticking clock sound to indicate that the item in question is a time value). We believe that non-lexical cues can be particularly useful for applic... Stefanie Shriver, Alan W. Black, Ronald Rosenfeld |
INTERSPEECH | 2 |
| 1999 | Using decision trees within the tilt intonation model to predict F0 contoursabstractA major obstacle for the migration of automatic speech recognition into every-day life products is environmental robustness. Automatic speech recognition systems work reasonably well under clean (laboratory) conditions but degrade seriously under real world conditions (e.g. out-door, car). A lot of research work is devoted to increase the environmental robustness of automatic speech recognition systems. A common method is to use clean (office) data as a starting point and simulate the degraded environmental situation by additive artificial (e.g. Gaussian) or recorded noise from the real environment [1]. We study the validity of such additive noise experiments with regard to a real noisy environment. With regard to a previously published work on database adaptation we also examine the possible benefit when using models trained in the simulated environment as a starting point for adaptation ([2]). We present experimental results on data recorded for task-dependent whole word and phoneme modeling in the car environment on data from the the MoTiV Car Speech Data Collection (CSDC) [3]. Kurt E. Dusterhoff, Alan W. Black, Paul Taylor 0001 |
EUROSPEECH | 2 |
| 1999 | Speech synthesis by phonological structure matchingabstractThis paper presents a new technique for speech synthesis by unit selection. The technique works by specifying the synthesis target and the speech database as phonological trees, and using a selection algorithm which finds the largest parts of trees in the database which match parts of the target tree. The technique avoids many of the errors made by prosody generation modules by incorporating their operation in the selection implicitly. A technique for using signal processing only when it is needed most is also described. The technique produces better quality speech than previous approaches and is also significantly faster. 1. INTRODUCTION It is common in any overview of a speech synthesis system (e.g. [13], [8]) to see the system broken down into a number of components, which nearly always include things such as text normalisation, lexical lookup, intonation, duration, diphone concatenation and signal processing. A standard model of waveform generation over the last years has been fo... Paul Taylor 0001, Alan W. Black |
EUROSPEECH | 2 |
| 1998 | On the use of automatically generated discourse-level information in a concept-to-speech synthesis systemabstractThis paper describes the latest version of the SOLE concept-tospeech system, which uses linguistic information provided by a natural language generation system to improve the prosody of synthetic speech. We discuss the types of linguistic information that prove most useful and the implications for text-to-speech systems. 1. INTRODUCTION The purpose of the SOLE project is to make use of automaticallygenerated, high-level linguistic information to improve the quality of the intonation of synthetic speech. After choosing an initial set of linguistic constructs thought to have some influence on prosody, we developed an SGML-based mark-up language to serve as a general interface between NLG and speech synthesis systems, and trained our synthesis system to recognise correlations between the mark-up and intonational contours so that it can make use of this mark-up when synthesising. As a result, many of the errors that the synthesiser makes with regard to knowing when to accent or deaccent ... Janet Hitzeman, Alan W. Black, Paul Taylor 0001, Chris Mellish, Jon Oberlander |
ICSLP | 2 |
| 1998 | Letter to sound rules for accented lexicon compressionabstractThis paper presents trainable methods for generating letter to sound rules from a given lexicon for use in pronouncing out-of-vocabulary words and as a method for lexicon compression. As the relationship between a string of letters and a string of phonemes representing its pronunciation for many languages is not trivial, we discuss two alignment procedures, one fully automatic and one hand-seeded which produce reasonable alignments of letters to phones. Top Down Induction Tree models are trained on the aligned entries. We show how combined phoneme/stress prediction is better than separate prediction processes, and still better when including in the model the last phonemes transcribed and part of speech information. For the lexicons we have tested, our models have a word accuracy (including stress) of 78% for OALD, 62% for CMU and 94% for BRULEX. The extremely high scores on the training sets allow substantial size reductions (more than 1/20). WWW site: http://tcts.fpms.ac.be/synthesis/mbrdico Vincent Pagel, Kevin A. Lenzo, Alan W. Black |
ICSLP | 3 |
| 1998 | SABLE: a standard for TTS markupabstractCurrently, speech synthesizers are controlled by a multitude of proprietary tag sets. These tag sets vary substantially across synthesizers and are an inhibitor to the adoption of speech synthesis technology by developers. SABLE is an XML/SGML-based markup scheme for text-to-speech synthesis, developed to address the need for a common TTS control paradigm. This paper presents an overview of the SABLE specification, and provides links to sites where further information on SABLE can be accessed. Richard Sproat, Andrew J. Hunt, Mari Ostendorf, Paul Taylor 0001, Alan W. Black, Kevin A. Lenzo, Mike Edgington |
ICSLP | 5 |
| 1998 | Assigning phrase breaks from part-of-speech sequences
Paul Taylor 0001, Alan W. Black |
Comput. Speech Lang. | 2 |
| 1997 | Automatically clustering similar units for unit selection in speech synthesisabstractThis paper describes a new method for synthesizing speech by concatenating sub-word units from a database of labelled speech. A large unit inventory is created by automatically clustering units of the same phone class based on their phonetic and prosodic context. The appropriate cluster is then selected for a target unit offering a small set of candidate units. An optimal path is found through the candidate units based on their distance from the cluster center and an acoustically based join cost. Details of the method and justification are presented. The results of experiments using two different databases are given, optimising various parameters within the system. Also a comparison with other existing selection based synthesis techniques is given showing the advantages this method has over existing ones. The method is implemented within a full text-to-speech system offering efficient natural sounding speech synthesis. 1. BACKGROUND Speech synthesis by concatenation of sub-word units ... Alan W. Black, Paul Taylor 0001 |
EUROSPEECH | 1 |
| 1997 | Assigning phrase breaks from part-of-speech sequencesabstractOne of the important stages in the process of turning unmarked text into speech is the assignment of appropriate phrase break boundaries. Phrase break boundaries are important to later modules including accent assignment, duration control and pause insertion. A number of different algorithms have been proposed for such a task, ranging from the simple to the complex. These different algorithms require different information such as part of speech tags, syntax and even semantic understanding of the text. Obviously these requirements come at differing costs and it is important to trade off difficulty in finding particular input features versus accuracy of the model. The simplest models are deterministic rules. A model simply inserting phrase breaks after punctuation Alan W. Black, Paul Taylor 0001 |
EUROSPEECH | 1 |
| 1996 | Unit selection in a concatenative speech synthesis system using a large speech databaseabstractOne approach to the generation of natural-sounding synthesized speech waveforms is to select and concatenate units from a large speech database. Units (in the current work, phonemes) are selected to produce a natural realisation of a target phoneme sequence predicted from text which is annotated with prosodic and phonetic context information. We propose that the units in a synthesis database can be considered as a state transition network in which the state occupancy cost is the distance between a database unit and a target, and the transition cost is an estimate of the quality of concatenation of two consecutive units. This framework has many similarities to HMM-based speech recognition. A pruned Viterbi search is used to select the best units for synthesis from the database. This approach to waveform synthesis permits training from natural speech: two methods for training from speech are presented which provide weights which produce more natural speech than can be obtained by hand-tuning. Andrew J. Hunt, Alan W. Black |
ICASSP | 2 |
| 1996 | Generating F0 contours from toBI labels using linear regressionabstractThis paper describes a method for generating F 0 contours from ToBI labelled utterances. The method uses linear regression to predict F 0 target values for the start, mid-vowel and end of every syllable, using features representing the ToBI labels, stress and syllable position. Contours generated by this method for an English database have a correlation of 0.62 and 34.8 Hz RMS error when compared with originals from test data. These results are significant improvements on a previous rule driven method (0.40 and 44.7), and the new method contours are preferred by human listeners. The technique has also been successfully applied to Japanese ToBI with similar improvements. 1. INTRODUCTION One problem in the process of synthesizing natural sounding speech is the prediction of an F 0 contour which adequately reflects the desired prosodic tune. In most synthesizers the task of generating a prosodic tune consists of two sub-tasks, the prediction of intonation labels (accents, tones, etc) fr... Alan W. Black, Andrew J. Hunt |
ICSLP | 1 |
| 1995 | Optimising selection of units from speech databases for concatenative synthesisabstractConcatenating units of natural speech is one method of speech synthesis1. Most such systems use an inventory of fixed length units, typically diphones or triphones with one instance of each type. An alternative is to use more varied, non-uniform units extracted from large speech databases containing multiple instances of each. The greater variability in such natural speech segments allows closer modeling of naturalness and differences in speaking styles, and eliminates the need for specially-recorded, single-use databases. However, with the greater variability comes the problem of how to select between the many instances of units in the database. This paper addresses that issue and presents a general method for unit selection. Alan W. Black, Nick Campbell 0001 |
EUROSPEECH | 1 |
| 1994 | CHATR: a generic speech synthesis system
Alan W. Black, Paul A. Taylor |
COLING | 1 |
| 1994 | Assigning intonation elements and prosodic phrasing for English speech synthesis from high level linguistic inputabstractThis paper describes a method for generating intonation events and prosodic phrasing from a high level linguistic description. Specifically, the input consists of information normally available from linguistic processing: part of speech, constituent structure, and, importantly, speech act. The generated output contains explicit intonational events from which an F0 contour may be generated. Prosody can be controlled via features in the input describing the function of words and phrases without direct reference to intonation.The results are evaluated against natural spoken sentences. Alan W. Black, Paul Taylor 0001 |
ICSLP | 1 |
| 1992 | Embedding DRT in a Situation Theoretic Framework
Alan W. Black |
COLING | 1 |
| 1991 | Analysis of Unknown Words through Morphological Decomposition
Alan W. Black, Joke van de Plassche, Briony Williams |
EACL | 1 |
| 1987 | Formalisms For Morphographemic Description
Alan W. Black, Graeme D. Ritchie, Stephen G. Pulman, Graham Russell |
EACL | 1 |
| 1987 | A Computational Framework for Lexical Description
Graeme D. Ritchie, Stephen G. Pulman, Alan W. Black, Graham Russell |
Comput. Linguistics | 3 |
| 1986 | A Dictionary and Morphological Analyser for English
Graham Russell, Stephen G. Pulman, Graeme D. Ritchie, Alan W. Black |
COLING | 4 |