VLDB 2026 Research / reviewers in the wild / expert
Alex Waibel
dblp:08/2456 · also Alexander H. Waibel, Alexander Waibel
· DBLP profile ↗
353ranked-venue papers
26as first author
32since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 251 · 15 first-author · 22 since 2021Artificial intelligence and machine learning · 225 · 11 first-author · 23 since 2021Human-computer interaction and ubiquitous computing · 12 · 3 first-authorSystems, architecture and hardware · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond Transcripts: A Renewed Perspective on Audio ChapteringabstractAudio chaptering, the task of segmenting longform audio into coherent sections, is increasingly important for navigating podcasts, lectures, and videos.Despite its relevance, research remains limited and text-based, leaving key questions unresolved about leveraging audio information, handling ASR errors, and transcript-free evaluation.We address these gaps through three contributions: (1) a systematic comparison between text-based models with acoustic features, a novel audio-only architecture (AudioSeg) operating on learned audio representations, and multimodal LLMs;(2) empirical analysis of factors affecting performance, including transcript quality, acoustic features, duration, and speaker composition; and (3) formalized evaluation protocols contrasting transcript-dependent text-space protocols with transcript-invariant time-space protocols.Our experiments on YTSeg reveal that AudioSeg substantially outperforms text-based approaches, pauses provide the largest acoustic gains, and MLLMs remain limited by context length and weak instruction following, yet MLLMs are promising on shorter audio. 1 Fabian Retkowski, Maike Züfle, Thai-Binh Nguyen, Jan Niehues, Alex Waibel |
ACL (1) | 5 |
| 2026 | Paragraph Segmentation Revisited: Towards a Standard Task for Structuring Speech
Fabian Retkowski, Alex Waibel |
LREC | 2 |
| 2026 | MUSCAT: MUltilingual, SCientific ConversATion Benchmark
Supriti Sinhamahapatra, Thai-Binh Nguyen, Yigit Oguz, Enes Yavuz Ugan, Jan Niehues, Alex Waibel |
LREC | 6 |
| 2026 | CAPE: A CLIP-Aware Pointing Ensemble of Complementary Heatmap Cues for Embodied Reference UnderstandingabstractWe address Embodied Reference Understanding, the task of predicting the object a person in the scene refers to through pointing gesture and language. This requires multimodal reasoning over text, visual pointing cues, and scene context, yet existing methods often fail to fully exploit visual disambiguation signals. We also observe that while the referent often aligns with the head-to-fingertip direction, in many cases it aligns more closely with the wrist-to-fingertip direction, making a single-line assumption overly limiting. To address this, we propose a dual-model framework, where one model learns from the head-to-fingertip direction and the other from the wrist-to-fingertip direction. We introduce a Gaussian ray heatmap representation of these lines and use them as input to provide a strong supervisory signal that encourages the model to better attend to pointing cues. To fuse their complementary strengths, we present the CLIP-Aware Pointing Ensemble module, which performs a hybrid ensemble guided by CLIP features. We further incorporate an auxiliary object center prediction head to enhance referent localization. We validate our approach on YouRefIt, achieving 75.0 mAP at 0.25 IoU, alongside state-of-the-art CLIP and CDscores, and demonstrate its generality on unseen CAESAR and ISL Pointing, showing robust performance across benchmarks. Fevziye Irem Eyiokur, Doggucan Yaman, Hazim Kemal Ekenel, Alex Waibel |
WACV | 4 |
| 2025 | Few-Shot Learning Translation from New LanguagesabstractRecent work shows strong transfer learning capability to unseen languages in sequence-tosequence neural networks, under the assumption that we have high-quality word representations for the target language.We evaluate whether this direction is a viable path forward for translation from low-resource languages by investigating how much data is required to learn such high-quality word representations.We first show that learning word embeddings separately from a translation model can enable rapid adaptation to new languages with only a few hundred sentences of parallel data.To see whether the current bottleneck in transfer to low-resource languages lies mainly with learning the word representations, we then train word embeddings models on varying amounts of data, to then plug them into a machine translation model.We show that in this simulated low-resource setting with only 500 parallel sentences and 31,250 sentences of monolingual data we can exceed 15 BLEU on Flores on unseen languages.Finally, we investigate why on a real low-resource language the results are less favorable and find fault with the publicly available multilingual language modelling datasets.Lng MADLAD Fineweb CulturaX HPLT clean noisy cs 782,001K hr 48,765K 503,047K kk 46,234K 78,582K 69,353K 51,000K 81,006K is 33,625K 73,760K 48,105K 39,527K 69,643K af 24,339K 64,576K 51,437K 19,537K 37,737K tl 23,639K 97,038K 41,619K 4,516K 52,879K mk 22,537K 48,877K 42,075K 38,494K 57,008K gl 22,214K 78,345K 31,112K 24,524K 61,177K ka 21,008K 56,944K 55,733K 48,299K 63,722K uz 16,581K 28,099K 19,873K 1,152K 14,800K bs 16,561K 217,725K 253,877K 1.2K 268,156K sw 12,839K 27,954K 18,004K 576K 34,308K gu 10,849K 22,527K 20,395K 18,774K 20,639K ur 10,830K 24,907K 43,720K 38,004K 50,629K kn 10,787K 26,427K 24,693K 20,243K 24,929K si 10,777K 20,775K 15,238K 16,332K 33,707K ne 9,535K 18,476K 38,949K 35,581K 37,138K ky 7,402K 14,054K 13,508K 9,295K 10,041K ga 7,155K 124,945K 11,156K 6,108K 10,993K mt 6,442K 18,627K 7,224K 3,337K 8,675K ha 3,560K 7,868K --5,688K ceb 1,677K 10,756K 2,906K 3,375K 2,864K zu 1,320K 8,093K 2,023K -2,710K war 72K 26,042K 2.8K 48K 87k Carlos Mullov, Alex Waibel |
EMNLP | 2 |
| 2025 | Summarizing Speech: A Comprehensive SurveyabstractFabian Retkowski, Maike Züfle, Andreas Sudmann, Dinah Pfau, Shinji Watanabe, Jan Niehues, Alexander Waibel. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Fabian Retkowski, Maike Züfle, Andreas Sudmann, Dinah Pfau, Shinji Watanabe 0001, Jan Niehues, Alex Waibel |
EMNLP | 7 |
| 2025 | Continuously Learning New Words in Automatic Speech RecognitionabstractDespite recent advances, Automatic Speech Recognition (ASR) systems are still far from perfect. Typical errors include acronyms, named entities, and domain-specific special words for which little or no labeled data is available. To address the problem of recognizing these words, we propose a self-supervised continual learning approach: Given the audio of a lecture talk with the corresponding slides, we bias the model towards decoding new words from the slides by using a memory-enhanced ASR model from the literature. Then, we perform inference on the talk, collecting utterances that contain detected new words into an adaptation data set. Continual learning is then performed by training adaptation weights added to the model on this data set. The whole procedure is iterated for many talks. We show that with this approach, we obtain increasing performance on the new words when they occur more frequently (>80% recall) while preserving the general performance of the model. Christian Huber, Alex Waibel |
ICASSP | 2 |
| 2025 | Factorized-VITS: Decoupling Prosody and Text in End-to-End Speech Synthesis without External or Secondary AlignerabstractWe propose Factorized-VITS, an advanced end-to-end text-to-speech model that incorporates explicit text-side prosody modeling control into VITS while achieving a clean factorization of the audio prior hidden space into text and prosody subspaces. Unlike previous works that rely on external or secondary aligners, Factorized-VITS is the first work attempting to do on-the-fly alignment in the factorized text subspace without introducing extra parameters, which not only simplifies the training procedure but also enables the use of a more complex prosody prior. Our experiments demonstrate the accuracy and effectiveness of this approximation strategy. Furthermore, we implement an in-context learning joint predictor for pitch, energy, and duration, which offers a flexible streaming deployment option. Alex Waibel |
ICASSP | 2 |
| 2025 | Improving Pronunciation and Accent Conversion through Knowledge Distillation And Synthetic Ground-Truth from Native TTSabstractPrevious approaches on accent conversion (AC) mainly aimed at making non-native speech sound more native while maintaining the original content and speaker identity. However, non-native speakers sometimes have pronunciation issues, which can make it difficult for listeners to understand them. Hence, we developed a new AC approach that not only focuses on accent conversion but also improves pronunciation of non-native accented speaker. By providing the non-native audio and the corresponding transcript, we generate the ideal ground-truth audio with native-like pronunciation with original duration and prosody. This ground-truth data aids the model in learning a direct mapping between accented and native speech. We utilize the end-to-end VITS framework to achieve high-quality waveform reconstruction for the AC task. As a result, our system not only produces audio that closely resembles native accents and while retaining the original speaker’s identity but also improve pronunciation, as demonstrated by evaluation results. Tuan Nam Nguyen, Seymanur Akti, Ngoc-Quan Pham, Alex Waibel |
ICASSP | 4 |
| 2025 | MSA-ASR: Efficient Multilingual Speaker Attribution with frozen ASR Models
Thai-Binh Nguyen, Alex Waibel |
ICASSP | 2 |
| 2025 | PIER: A Novel Metric for Evaluating What Matters in Code-SwitchingabstractCode-switching, the alternation of languages within a single discourse, presents a significant challenge for Automatic Speech Recognition. Despite the unique nature of the task, performance is commonly measured with established metrics such as Word-Error-Rate (WER). However, in this paper, we question whether these general metrics accurately assess performance on code-switching. Specifically, using both Connectionist-Temporal-Classification and Encoder-Decoder models, we show fine-tuning on non-code-switched data from both matrix and embedded language improves classical metrics on code-switching test sets, although actual code-switched words worsen (as expected). Therefore, we propose Point-of-Interest Error Rate (PIER), a variant of WER that focuses only on specific words of interest. We instantiate PIER on code-switched utterances and show that this more accurately describes the code-switching performance, showing huge room for improvement in future work. This focused evaluation allows for a more precise assessment of model performance, particularly in challenging aspects such as inter-word and intra-word code-switching. Enes Yavuz Ugan, Ngoc-Quan Pham, Leonard Bärmann, Alex Waibel |
ICASSP | 4 |
| 2025 | Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion
Seymanur Akti, Tuan-Nam Nguyen, Alex Waibel |
INTERSPEECH | 3 |
| 2025 | Streaming Non-Autoregressive Model for Accent Conversion and Pronunciation Improvement
Tuan-Nam Nguyen, Ngoc-Quan Pham, Seymanur Akti, Alex Waibel |
INTERSPEECH | 4 |
| 2025 | Cocktail-Party Audio-Visual Speech Recognition
Thai-Binh Nguyen, Ngoc-Quan Pham, Alex Waibel |
INTERSPEECH | 3 |
| 2025 | Weight Factorization and Centralization for Continual Learning in Speech Recognition
Enes Yavuz Ugan, Ngoc-Quan Pham, Alex Waibel |
INTERSPEECH | 3 |
| 2025 | From Speech Science to Language Transparence
Alex Waibel |
INTERSPEECH | 1 |
| 2024 | Decoupled Vocabulary Learning Enables Zero-Shot Translation from Unseen LanguagesabstractMultilingual neural machine translation systems learn to map sentences of different languages into a common representation space.Intuitively, with a growing number of seen languages the encoder sentence representation grows more flexible and easily adaptable to new languages.In this work, we test this hypothesis by zero-shot translating from unseen languages.To deal with unknown vocabularies from unknown languages we propose a setup where we decouple learning of vocabulary and syntax, i.e. for each language we learn word representations in a separate step (using cross-lingual word embeddings), and then train to translate while keeping those word representations frozen.We demonstrate that this setup enables zero-shot translation from entirely unseen languages.Zero-shot translating with a model trained on Germanic and Romance languages we achieve scores of 42.6 BLEU for Portuguese-English and 20.7 BLEU for Russian-English on TED domain.We explore how this zero-shot translation capability develops with varying number of languages seen by the encoder.Lastly, we explore the effectiveness of our decoupled learning strategy for unsupervised machine translation.By exploiting our model's zero-shot translation capability for iterative back-translation we attain near parity with a supervised setting. Carlos Mullov, Ngoc-Quan Pham, Alex Waibel |
ACL (1) | 3 |
| 2024 | DECM: Evaluating Bilingual ASR Performance on a Code-switching/mixing BenchmarkabstractAutomatic Speech Recognition has made significant progress, but challenges persist. Code-switched (CSW) Speech presents one such challenge, involving the mixing of multiple languages by a speaker. Even when multilingual ASR models are trained, each utterance on its own usually remains monolingual. We introduce an evaluation dataset for German-English CSW, with German as the matrix language and English as the embedded language. The dataset comprises spontaneous speech from diverse domains, enabling realistic CSW evaluation in German-English. It includes splits with varying degrees of CSW to facilitate specialized model analysis. As it is difficult to collect CSW data for all language pairs, the provision of such evaluation data, is crucial for developing and analyzing ASR models capable of generalizing across unseen pairs. Detailed data statistics are presented, and state-of-the-art (SOTA) multilingual models are evaluated showing challanges of CSW speech. Enes Yavuz Ugan, Ngoc-Quan Pham, Alex Waibel |
LREC/COLING | 3 |
| 2024 | From Text Segmentation to Smart Chaptering: A Novel Benchmark for Structuring Video TranscriptionsabstractText segmentation is a fundamental task in natural language processing, where documents are split into contiguous sections.However, prior research in this area has been constrained by limited datasets, which are either small in scale, synthesized, or only contain well-structured documents.In this paper, we address these limitations by introducing a novel benchmark YTSEG focusing on spoken content that is inherently more unstructured and both topically and structurally diverse.As part of this work, we introduce an efficient hierarchical segmentation model MiniSeg, that outperforms state-ofthe-art baselines.Lastly, we expand the notion of text segmentation to a more practical "smart chaptering" task that involves the segmentation of unstructured content, the generation of meaningful segment titles, and a potential real-time application of the models. Fabian Retkowski, Alex Waibel |
EACL (1) | 2 |
| 2024 | Audio-Driven Talking Face Generation with Stabilized Synchronization Loss
Doggucan Yaman, Fevziye Irem Eyiokur, Leonard Bärmann, Hazim Kemal Ekenel, Alex Waibel |
ECCV (19) | 5 |
| 2024 | SciEx: Benchmarking Large Language Models on Scientific Exams with Human Expert Grading and Automatic GradingabstractTu Anh Dinh, Carlos Mullov, Leonard Bärmann, Zhaolin Li, Danni Liu, Simon Reiß, Jueun Lee, Nathan Lerzer, Jianfeng Gao, Fabian Peller-Konrad, Tobias Röddiger, Alexander Waibel, Tamim Asfour, Michael Beigl, Rainer Stiefelhagen, Carsten Dachsbacher, Klemens Böhm, Jan Niehues. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Tu Anh Dinh, Carlos Mullov, Leonard Bärmann, Zhaolin Li, Simon Reiß, Jueun Lee, Nathan Lerzer, Jianfeng Gao 0002, Fabian Tërnava, Tobias Röddiger, Alex Waibel, Tamim Asfour, Michael Beigl, Rainer Stiefelhagen, Carsten Dachsbacher, Klemens Böhm, Jan Niehues |
EMNLP | 12 |
| 2024 | Synthetic Conversations Improve Multi-Talker ASRabstractIn recent times, automatic speech recognition (ASR) has seen remarkable progress, particularly in recognizing dominant speakers. Nevertheless, the challenge of multi-talker scenarios involving distinguishing between speakers and transcribing their speech accurately remains unsolved due to limited data constraining model effectiveness. In this study, We propose a novel methodology called Systematic Synthetic Conversations (SSC), which leverages conventional ASR datasets to help an end-to-end (E2E) multi-talker ASR model establish new state-of-the-art results across synthetic and authentic multi-talker datasets. Notably, we achieved a 3.47% word error rate (WER) for the Libri2Mix [1] set, and WERs of 13.96% and 19.51% for the AMI-IHM and AMI-SDM [2] sets, respectively. These outcomes underscore the hidden potential of existing resources in tackling the complicated multi-talker problems within the domain of ASR. Thai-Binh Nguyen, Alex Waibel |
ICASSP | 2 |
| 2023 | AdapITN: A Fast, Reliable, and Dynamic Adaptive Inverse Text NormalizationabstractInverse text normalization (ITN) is the task that transforms text in spoken-form into written-form. While automatic speech recognition (ASR) produces text in spoken-form, human and natural language understanding systems prefer to consume text in written-form. ITN generally deals with semiotic phrases (e.g., numbers, date, time). However, lack of studies to deal with phonetization phrases, which is ASR’s output when it handles unseen data (e.g., foreign-named entities, domain names), although these exist in the same form in the spoken-form text. The reason is that phonetization phrases are infinite patterns and language-dependent. In this study, we introduce a novel end2end model that can handle both semiotic phrases (SEP) and phonetization phrases (PHP), named AdapITN. We call it "Adap" because it allows for handling unseen PHP. The model performs only when necessary by providing a mechanism to narrow normalized regions and external query knowledge, reducing the runtime significantly. Le Duc Minh Nhat, Quang Minh Nguyen, Quoc Truong Do, Alex Waibel |
ICASSP | 6 |
| 2023 | SYNTACC : Synthesizing Multi-Accent Speech By Weight FactorizationabstractConventional multi-speaker text-to-speech synthesis (TTS) is known to be capable of synthesizing speech for multiple voices, yet it cannot generate speech in different accents. This limitation has motivated us to develop SYNTACC (Synthesizing speech with accents) which adapts conventional multi-speaker TTS to produce multi-accent speech. Our method uses the YourTTS model and involves a novel multi-accent training mechanism. The method works by decomposing each weight matrix into a shared component and an accent-dependent component, with the former being initialized by the pretrained multi-speaker TTS model and the latter being factorized into vectors using rank-1 matrices to reduce the number of training parameters per accent. This weight factorization method proves to be effective in fine-tuning the SYNTACC on multi-accent data sets in a low-resource condition. Our SYNTACC model eventually allows speech synthesis in not only different voices but also in different accents. Tuan-Nam Nguyen, Ngoc-Quan Pham, Alex Waibel |
ICASSP | 3 |
| 2023 | Towards continually learning new languages
Ngoc-Quan Pham, Jan Niehues, Alex Waibel |
INTERSPEECH | 3 |
| 2023 | Incremental Blockwise Beam Search for Simultaneous Speech Translation with Controllable Quality-Latency TradeoffabstractBlockwise self-attentional encoder models have recently emerged as one promising end-to-end approach to simultaneous speech translation. These models employ a blockwise beam search with hypothesis reliability scoring to determine when to wait for more input speech before translating further. However, this method maintains multiple hypotheses until the entire speech input is consumed -- this scheme cannot directly show a single \textit{incremental} translation to users. Further, this method lacks mechanisms for \textit{controlling} the quality vs. latency tradeoff. We propose a modified incremental blockwise beam search incorporating local agreement or hold-$n$ policies for quality-latency control. We apply our framework to models trained for online or offline translation and demonstrate that both types can be effectively used in online mode. Experimental results on MuST-C show 0.6-3.6 BLEU improvement without changing latency or 0.8-1.4 s latency improvement without changing quality. Peter Polak, Brian Yan, Shinji Watanabe 0001, Alex Waibel, Ondrej Bojar |
INTERSPEECH | 4 |
| 2023 | A survey on computer vision based human analysis in the COVID-19 era
Fevziye Irem Eyiokur, Alperen Kantarci, Mustafa Ekrem Erakin, Naser Damer, Ferda Ofli, Muhammad Imran 0002, Janez Krizaj, Albert Ali Salah, Alex Waibel, Vitomir Struc, Hazim Kemal Ekenel |
Image Vis. Comput. | 9 |
| 2022 | Accent Conversion using Pre-trained Model and Synthesized Data from Voice ConversionabstractAccent conversion (AC) aims to generate synthetic audios by changing the pronunciation pattern and prosody of source speakers (in source audios) while preserving voice quality and linguistic content.There has not been a parallel corpus that contains pairs of audios having the same contents yet coming from the same speakers in different accents, the authors hence work on a solution to synthesize one as training input.The training pipeline is conducted via two steps.First, a voice conversion (VC) model is constructed to synthesize a training data set, containing pairs of audios in the same voice but two different accents.Second, an AC model is trained with the synthesized data to convert a source accented speech to a target accented speech.Given the recognized success of self-supervised learning speech representation (wav2vec 2.0) on certain speech problems such as VC, speech recognition, speech translation, and speech-tospeech translation, we adopt this architecture with some customization to train the AC model in the second step.With just 9-hour synthesized training data, the encoder initialized by the weight of the pre-trained wav2vec 2.0 model outperforms the LSTM-based encoder. Tuan-Nam Nguyen, Ngoc-Quan Pham, Alex Waibel |
INTERSPEECH | 3 |
| 2022 | Adaptive multilingual speech recognition with pretrained modelsabstractMultilingual speech recognition with supervised learning has achieved great results as reflected in recent research.With the development of pretraining methods on audio and text data, it is imperative to transfer the knowledge from unsupervised multilingual models to facilitate recognition, especially in many languages with limited data.Our work investigated the effectiveness of using two pretrained models for two modalities: wav2vec 2.0 for audio and MBART50 for text, together with the adaptive weight techniques to massively improve the recognition quality on the public datasets containing CommonVoice and Europarl.Overall, we noticed an 44% improvement over purely supervised learning, and more importantly, each technique provides a different reinforcement in different languages.We also explore other possibilities to potentially obtain the best model by slightly adding either depth or relative attention to the architecture. Ngoc-Quan Pham, Alex Waibel, Jan Niehues |
INTERSPEECH | 2 |
| 2021 | Instant One-Shot Word-Learning for Context-Specific Neural Sequence-to-Sequence Speech RecognitionabstractNeural sequence-to-sequence systems deliver state-of-the-art performance for automatic speech recognition (ASR). When using appropriate modeling units, e.g., byte-pair encoded characters, these systems are in principal open vocabulary systems. In practice, however, they often fail to recognize words not seen during training, e.g., named entities, numbers or technical terms. To alleviate this problem we supplement an end-to-end ASR system with a word/phrase memory and a mechanism to access this memory to recognize the words and phrases correctly. After the training of the ASR system, and when it has already been deployed, a relevant word can be added or subtracted instantly without the need for further training. In this paper we demonstrate that through this mechanism our system is able to recognize more than 85% of newly added words that it previously failed to recognize compared to a strong baseline. Christian Huber, Juan Hussain, Sebastian Stüker, Alex Waibel |
ASRU | 4 |
| 2021 | Super-Human Performance in Online Low-Latency Recognition of Conversational SpeechabstractAchieving super-human performance in recognizing human speech has been a goal for several decades, as researchers have worked on increasingly challenging tasks. In the 1990's it was discovered, that conversational speech between two humans turns out to be considerably more difficult than read speech as hesitations, disfluencies, false starts and sloppy articulation complicate acoustic processing and require robust handling of acoustic, lexical and language context, jointly. Early attempts with statistical models could only reach error rates over 50% and far from human performance (WER of around 5.5%). Neural hybrid models and recent attention-based encoder-decoder models have considerably improved performance as such contexts can now be learned in an integral fashion. However, processing such contexts requires an entire utterance presentation and thus introduces unwanted delays before a recognition result can be output. In this paper, we address performance as well as latency. We present results for a system that can achieve super-human performance (at a WER of 5.0%, over the Switchboard conversational benchmark) at a word based latency of only 1 second behind a speaker's speech. The system uses multiple attention-based encoder-decoder networks integrated within a novel low latency incremental inference approach. Thai Son Nguyen, Sebastian Stüker, Alex Waibel |
Interspeech | 3 |
| 2021 | Efficient Weight Factorization for Multilingual Speech RecognitionabstractEnd-to-end multilingual speech recognition involves using a single model training on a compositional speech corpus including many languages, resulting in a single neural network to handle transcribing different languages. Due to the fact that each language in the training data has different characteristics, the shared network may struggle to optimize for all various languages simultaneously. In this paper we propose a novel multilingual architecture that targets the core operation in neural networks: linear transformation functions. The key idea of the method is to assign fast weight matrices for each language by decomposing each weight matrix into a shared component and a language dependent component. The latter is then factorized into vectors using rank-1 assumptions to reduce the number of parameters per language. This efficient factorization scheme is proved to be effective in two multilingual settings with $7$ and $27$ languages, reducing the word error rates by $26\%$ and $27\%$ rel. for two popular architectures LSTM and Transformer, respectively. Ngoc-Quan Pham, Tuan-Nam Nguyen, Sebastian Stüker, Alex Waibel |
Interspeech | 4 |
| 2020 | ELITR: European Live TranslatorabstractELITR (European Live Translator) project aims to create a speech translation system for simultaneous subtitling of conferences and online meetings targetting up to 43 languages. The technology is tested by the Supreme Audit Office of the Czech Republic and by alfaview®, a German online conferencing system. Other project goals are to advance document-level and multilingual machine translation, automatic speech recognition, and automatic minuting. Ondrej Bojar, Dominik Machácek, Sangeet Sagar, Otakar Smrz, Jonás Kratochvíl, Ebrahim Ansari, Dario Franceschini, Chiara Canton, Ivan Simonini, Thai Son Nguyen, Sebastian Stüker, Alex Waibel, Barry Haddow, Rico Sennrich, Philip Williams |
EAMT | 13 |
| 2020 | Incorporating External Annotation to improve Named Entity Translation in NMTabstractThe correct translation of named entities (NEs) still poses a challenge for conventional neural machine translation (NMT) systems. This study explores methods incorporating named entity recognition (NER) into NMT with the aim to improve named entity translation. It proposes an annotation method that integrates named entities and inside–outside–beginning (IOB) tagging into the neural network input with the use of source factors. Our experiments on English→German and English→ Chinese show that just by including different NE classes and IOB tagging, we can increase the BLEU score by around 1 point using the standard test set from WMT2019 and achieve up to 12% increase in NE translation rates over a strong baseline. Maciej Modrzejewski, Miriam Exel, Bianka Buschbeck-Wolf, Thanh-Le Ha, Alex Waibel |
EAMT | 5 |
| 2020 | Improving Sequence-To-Sequence Speech Recognition Training with On-The-Fly Data AugmentationabstractSequence-to-Sequence (S2S) models recently started to show state-of-the-art performance for automatic speech recognition (ASR). With these large and deep models overfitting remains the largest problem, outweighing performance improvements that can be obtained from better architectures. One solution to the overfitting problem is increasing the amount of available training data and the variety exhibited by the training data with the help of data augmentation. In this paper we examine the influence of three data augmentation methods on the performance of two S2S model architectures. One of the data augmentation method comes from literature, while two other methods are our own development - a time perturbation in the frequency domain and sub-sequence sampling. Our experiments on Switchboard and Fisher data show state-of-the-art performance for S2S models that are trained solely on the speech training data and do not use additional text data. Thai Son Nguyen, Sebastian Stüker, Jan Niehues, Alex Waibel |
ICASSP | 4 |
| 2020 | High Performance Sequence-to-Sequence Model for Streaming Speech RecognitionabstractRecently sequence-to-sequence models have started to achieve state-of-the-art performance on standard speech recognition tasks when processing audio data in batch mode, i.e., the complete audio data is available when starting processing. However, when it comes to performing run-on recognition on an input stream of audio data while producing recognition results in real-time and with low word-based latency, these models face several challenges. For many techniques, the whole audio sequence to be decoded needs to be available at the start of the processing, e.g., for the attention mechanism or the bidirectional LSTM (BLSTM). In this paper, we propose several techniques to mitigate these problems. We introduce an additional loss function controlling the uncertainty of the attention mechanism, a modified beam search identifying partial, stable hypotheses, ways of working with BLSTM in the encoder, and the use of chunked BLSTM. Our experiments show that with the right combination of these techniques, it is possible to perform run-on speech recognition with low word-based latency without sacrificing in word error rate performance. Thai Son Nguyen, Ngoc-Quan Pham, Sebastian Stüker, Alex Waibel |
INTERSPEECH | 4 |
| 2020 | Relative Positional Encoding for Speech Recognition and Direct TranslationabstractTransformer models are powerful sequence-to-sequence architectures that are capable of directly mapping speech inputs to transcriptions or translations. However, the mechanism for modeling positions in this model was tailored for text modeling, and thus is less ideal for acoustic inputs. In this work, we adapt the relative position encoding scheme to the Speech Transformer, where the key addition is relative distance between input states in the self-attention network. As a result, the network can better adapt to the variable distributions present in speech data. Our experiments show that our resulting model achieves the best recognition result on the Switchboard benchmark in the non-augmentation condition, and the best published result in the MuST-C speech translation benchmark. We also show that this model is able to better utilize synthetic data than the Transformer, and adapts better to variable sentence segmentation quality for speech translation. Ngoc-Quan Pham, Thanh-Le Ha, Tuan-Nam Nguyen, Thai Son Nguyen, Elizabeth Salesky, Sebastian Stüker, Jan Niehues, Alex Waibel |
INTERSPEECH | 8 |
| 2020 | DaCToR: A Data Collection Tool for the RELATER ProjectabstractCollecting domain-specific data for under-resourced languages, e.g., dialects of languages, can be very expensive, potentially financially prohibitive and taking long time. Moreover, in the case of rarely written languages, the normalization of non-canonical transcription might be another time consuming but necessary task. In order to collect domain-specific data in such circumstances in a time and cost-efficient way, collecting read data of pre-prepared texts is often a viable option. In order to collect data in the domain of psychiatric diagnosis in Arabic dialects for the project RELATER, we have prepared the data collection tool DaCToR for collecting read texts by speakers in the respective countries and districts in which the dialects are spoken. In this paper we describe our tool, its purpose within the project RELATER and the dialects which we have started to collect with the tool. Juan Hussain, Oussama Zenkri, Sebastian Stüker, Alex Waibel |
LREC | 4 |
| 2019 | Self-Attentional Models for Lattice InputsabstractLattices are an efficient and effective method to encode ambiguity of upstream systems in natural language processing tasks, for example to compactly capture multiple speech recognition hypotheses, or to represent multiple linguistic analyses.Previous work has extended recurrent neural networks to model lattice inputs and achieved improvements in various tasks, but these models suffer from very slow computation speeds.This paper extends the recently proposed paradigm of self-attention to handle lattice inputs.Self-attention is a sequence modeling technique that relates inputs to one another by computing pairwise similarities and has gained popularity for both its strong results and its computational efficiency.To extend such models to handle lattices, we introduce probabilistic reachability masks that incorporate lattice structure into the model and support lattice scores if available.We also propose a method for adapting positional embeddings to lattice structures.We apply the proposed model to a speech translation task and find that it outperforms all examined baselines while being much faster to compute than previous neural lattice models during both training and inference. Matthias Sperber, Graham Neubig, Ngoc-Quan Pham, Alex Waibel |
ACL (1) | 4 |
| 2019 | Neural Codes to Factor Language in Multilingual Speech RecognitionabstractIn the past, we adapted neural network based multilingual acoustic models using language codes. In this work, we study the extracted language codes and the language properties they encode: We use the codes to generate language prototype vectors, which represent the features of a language. Computing distances between prototype vectors shows that languages from the same family have smaller distances. This structure found within the feature representation supports the assumption that language codes do encode language information and not other properties like, e.g. channel characteristics, and in addition providing a richer language representation than the language identity alone.The network architecture of our system is based on a factorized model, which consists of multiple language dependent subnets. While we recently demonstrated that this approach enables multilingual setups to outperform monolingual ones, we here propose further optimizations. We evaluated using a) more language dependent subnets and b) wider BiLSTM layers. Our results indicate that using a larger number of language dependent subnets increases the system performance and renders phonetic pretraining superfluous. In addition, increasing the size of the hidden layers further improved the performance, with the system now outperforming the monolingual baseline by 6.3% relative. Markus Müller 0001, Sebastian Stüker, Alex Waibel |
ICASSP | 3 |
| 2019 | Connecting Humans with Humans: Multimodal, Multilingual, Multiparty MediationabstractBehind much of my research work over 4 decades has been the simple observation that people like people and love interacting with other people more than they like interacting with machines. Technologies that truly support such social desires are more likely to be adopted broadly. Consider email, texting, chat rooms, social media, video conferencing, the internet, speech translation, even videogames with a social element (e.g., Fortnite): we enjoy the technology whenever it brings us closer to our fellow humans, instead of imposing attention-grabbing clutter. If so, how then can we build better technologies that improve, encourage, support human-human interaction? In this talk, I will recount my own story along this journey. When I began, building technologies for the human-human experience, presented formidable challenges: Computer interfaces would need to anticipate and understand the way humans interact, but in 1976, a typical computer had only two instructions to interact with humans: character-in & character-out, and both only supported human-computer interaction. Over the decades that followed, we began to develop interfaces that can process the various modalities of human communication and we built systems that used several modalities in services to improve human-human interaction. These included: Alex Waibel |
ICMI | 1 |
| 2019 | Very Deep Self-Attention Networks for End-to-End Speech RecognitionabstractRecently, end-to-end sequence-to-sequence models for speech recognition have gained significant interest in the research community. While previous architecture choices revolve around time-delay neural networks (TDNN) and long short-term memory (LSTM) recurrent neural networks, we propose to use self-attention via the Transformer architecture as an alternative. Our analysis shows that deep Transformer networks with high learning capacity are able to exceed performance from previous end-to-end approaches and even match the conventional hybrid systems. Moreover, we trained very deep models with up to 48 Transformer layers for both encoder and decoders combined with stochastic residual connections, which greatly improve generalizability and training efficiency. The resulting models outperform all previous end-to-end ASR approaches on the Switchboard benchmark. An ensemble of these models achieve 9.9% and 17.7% WER on Switchboard and CallHome test sets respectively. This finding brings our end-to-end models to competitive levels with previous hybrid systems. Further, with model ensembling the Transformers can outperform certain hybrid systems, which are more complicated in terms of both structure and training procedure. Ngoc-Quan Pham, Thai Son Nguyen, Jan Niehues, Markus Müller 0001, Alex Waibel |
INTERSPEECH | 5 |
| 2019 | An Interactive Indoor Drone AssistantabstractWith the rapid advance of sophisticated control algorithms, the capabilities of drones to stabilise, fly and manoeuvre autonomously have dramatically improved, enabling us to pay greater attention to entire missions and the interaction of a drone with humans and with its environment during the course of such a mission. In this paper, we present an indoor office drone assistant that is tasked to run errands and carry out simple tasks at our laboratory, while given instructions from and interacting with humans in the space. To accomplish its mission, the system has to be able to understand verbal instructions from humans, and perform subject to constraints from control and hardware limitations, uncertain localisation information, unpredictable and uncertain obstacles and environmental factors. We combine and evaluate the dialogue, navigation, flight control, depth perception and collision avoidance components. We discuss performance and limitations of our assistant at the component as well as the mission level. A 78% mission success rate was obtained over the course of 27 missions. Tino Fuhrman, David Schneider 0006, Felix Altenberg, Simon Blasen, Stefan Constantin, Alex Waibel |
IROS | 7 |
| 2019 | Attention-Passing Models for Robust and Data-Efficient End-to-End Speech TranslationabstractSpeech translation has traditionally been approached through cascaded models consisting of a speech recognizer trained on a corpus of transcribed speech, and a machine translation system trained on parallel texts. Several recent works have shown the feasibility of collapsing the cascade into a single, direct model that can be trained in an end-to-end fashion on a corpus of translated speech. However, experiments are inconclusive on whether the cascade or the direct model is stronger, and have only been conducted under the unrealistic assumption that both are trained on equal amounts of data, ignoring other available speech recognition and machine translation corpora. In this paper, we demonstrate that direct speech translation models require more data to perform well than cascaded models, and although they allow including auxiliary data through multi-task training, they are poor at exploiting such data, putting them at a severe disadvantage. As a remedy, we propose the use of end- to-end trainable models with two attention mechanisms, the first establishing source speech to source text alignments, the second modeling source to target text alignment. We show that such models naturally decompose into multi-task–trainable recognition and translation tasks and propose an attention-passing technique that alleviates error propagation issues in a previous formulation of a model with two attention stages. Our proposed model outperforms all examined baselines and is able to exploit auxiliary training data much more effectively than direct attentional models. Matthias Sperber, Graham Neubig, Jan Niehues, Alex Waibel |
Trans. Assoc. Comput. Linguistics | 4 |
| 2018 | Multilingual Adaptation of RNN Based ASR SystemsabstractIn this work, we focus on multilingual systems based on recurrent neural networks (RNNs), trained using the Connectionist Temporal Classification (CTC) loss function. Using a multilingual set of acoustic units poses difficulties. To address this issue, we proposed Language Feature Vectors (LFV s) to train language adaptive multilingual systems. Language adaptation, in contrast to speaker adaptation, needs to be applied not only on the feature level, but also to deeper layers of the network. In this work, we therefore extended our previous approach by introducing a novel technique which we call “modulation”. Based on this method, we modulated the hidden layers of RNNs using LFVs. We evaluated this approach in both full and low resource conditions, as well as for grapheme and phone based systems. Lower error rates throughout the different conditions could be achieved by the use of the modulation. Markus Müller 0001, Sebastian Stüker, Alex Waibel |
ICASSP | 3 |
| 2018 | Exploring Ctc-Network Derived Features with Conventional Hybrid SystemabstractRecently in automatic speech recognition (ASR) a lot of attention has been given to decoding optimization to boost the performance of connectionist temporal classification criterion (CTC) systems in all neural setups. Different from that, we investigated the use of the output of CTC network as input features to traditional HMM/ANN hybrid systems. By doing so, we benefit from the strengths of the CTC network at label discrimination and the highly optimized decoding stack of conventional hybrid systems. In a Switchboard setup, a feedforward network system using our proposed CTC-network derived features with cross-entropy training outperforms a strong CTC baseline by a margin of 5% rel. in word error rate. With the same model, we achieved further improvements of 9% rel. when combining them with bottleneck features. Additionally, we revealed the possible elimination of the blank label during decoding and the alignment relationship between the CTC model and the traditional HMM system. Thai Son Nguyen, Sebastian Stiiker, Alex Waibel |
ICASSP | 3 |
| 2018 | Neural Language Codes for Multilingual Acoustic ModelsabstractMultilingual Speech Recognition is one of the most costly AI problems, because each language (7,000+) and even different accents require their own acoustic models to obtain best recognition performance. Even though they all use the same phoneme symbols, each language and accent imposes its own coloring or "twang". Many adaptive approaches have been proposed, but they require further training, additional data and generally are inferior to monolingually trained models. In this paper, we propose a different approach that uses a large multilingual model that is \emph{modulated} by the codes generated by an ancillary network that learns to code useful differences between the "twangs" or human language. We use Meta-Pi networks to have one network (the language code net) gate the activity of neurons in another (the acoustic model nets). Our results show that during recognition multilingual Meta-Pi networks quickly adapt to the proper language coloring without retraining or new data, and perform better than monolingually trained networks. The model was evaluated by training acoustic modeling nets and modulating language code nets jointly and optimize them for best recognition performance. Markus Müller 0001, Sebastian Stüker, Alex Waibel |
INTERSPEECH | 3 |
| 2018 | Term Extraction via Neural Sequence Labeling a Comparative Evaluation of Strategies Using Recurrent Neural Networks
Maren Kucza, Jan Niehues, Thomas Zenkel, Alex Waibel, Sebastian Stüker |
INTERSPEECH | 4 |
| 2018 | Low-Latency Neural Speech TranslationabstractThrough the development of neural machine translation, the quality of machine translation systems has been improved significantly.By exploiting advancements in deep learning, systems are now able to better approximate the complex mapping from source sentences to target sentences.But with this ability, new challenges also arise.An example is the translation of partial sentences in low-latency speech translation.Since the model has only seen complete sentences in training, it will always try to generate a complete sentence, though the input may only be a partial sentence.We show that NMT systems can be adapted to scenarios where no task-specific training data is available.Furthermore, this is possible without losing performance on the original training data.We achieve this by creating artificial data and by using multi-task learning.After adaptation, we are able to reduce the number of corrections displayed during incremental output construction by 45%, without a decrease in translation quality. Jan Niehues, Ngoc-Quan Pham, Thanh-Le Ha, Matthias Sperber, Alex Waibel |
INTERSPEECH | 5 |
| 2018 | Self-Attentional Acoustic ModelsabstractSelf-attention is a method of encoding sequences of vectors by relating these vectors to each-other based on pairwise similarities. These models have recently shown promising results for modeling discrete sequences, but they are non-trivial to apply to acoustic modeling due to computational and modeling issues. In this paper, we apply self-attention to acoustic modeling, proposing several improvements to mitigate these issues: First, self-attention memory grows quadratically in the sequence length, which we address through a downsampling technique. Second, we find that previous approaches to incorporate position information into the model are unsuitable and explore other representations and hybrid models to this end. Third, to stress the importance of local context in the acoustic signal, we propose a Gaussian biasing approach that allows explicit control over the context range. Experiments find that our model approaches a strong baseline based on LSTMs with network-in-network connections while being much faster to compute. Besides speed, we find that interpretability is a strength of self-attentional acoustic models, and demonstrate that self-attention heads learn a linguistically plausible division of labor. Matthias Sperber, Jan Niehues, Graham Neubig, Sebastian Stüker, Alex Waibel |
INTERSPEECH | 5 |
| 2018 | Subword and Crossword Units for CTC Acoustic ModelsabstractThis paper proposes a novel approach to create an unit set for CTC based speech recognition systems. By using Byte Pair Encoding we learn an unit set of an arbitrary size on a given training text. In contrast to using characters or words as units this allows us to find a good trade-off between the size of our unit set and the available training data. We evaluate both Crossword units, that may span multiple word, and Subword units. By combining this approach with decoding methods using a separate language model we are able to achieve state of the art results for grapheme based CTC systems. Thomas Zenkel, Ramon Sanabria, Florian Metze, Alex Waibel |
INTERSPEECH | 4 |
| 2018 | KIT-Multi: A Translation-Oriented Multilingual Embedding Corpus
Thanh-Le Ha, Jan Niehues, Matthias Sperber, Ngoc-Quan Pham, Alex Waibel |
LREC | 5 |
| 2018 | BULBasaa: A Bilingual Basaa-French Speech Corpus for the Evaluation of Language Documentation Tools
Fatima Hamlaoui, Emmanuel-Moselly Makasso, Markus Müller 0001, Jonas Engelmann, Gilles Adda, Alex Waibel, Sebastian Stüker |
LREC | 6 |
| 2018 | Automated Evaluation of Out-of-Context Errors
Patrick Huber, Jan Niehues, Alex Waibel |
LREC | 3 |
| 2018 | Towards Fluent Translations From Disfluent SpeechabstractWhen translating from speech, special consideration for conversational speech phenomena such as disfluencies is necessary. Most machine translation training data consists of well-formed written texts, causing issues when translating spontaneous speech. Previous work has introduced an intermediate step between speech recognition (ASR) and machine translation (MT) to remove disfluencies, making the data better-matched to typical translation text and significantly improving performance. However, with the rise of end-to-end speech translation systems, this intermediate step must be incorporated into the sequence-to-sequence architecture. Further, though translated speech datasets exist, they are typically news or rehearsed speech without many disfluencies (e.g. TED), or the disfluencies are translated into the references (e.g. Fisher). To generate clean translations from disfluent speech, cleaned references are necessary for evaluation. We introduce a corpus of cleaned target data for the Fisher Spanish-English dataset for this task. We compare how different architectures handle disfluencies and provide a baseline for removing disfluencies in end-to-end translation. Elizabeth Salesky, Susanne Burger, Jan Niehues, Alex Waibel |
SLT | 4 |
| 2017 | DBLSTM based multilingual articulatory feature extraction for language documentationabstractWith more than 7,000 living languages in the world and many of them facing extinction, the need for language documentation is now more pressing than ever. This process is time-consuming, requiring linguists as each language features peculiarities that need to be addressed. While automating the whole process is difficult, we aim at providing methods to support linguists during documentation. One important step in the workflow is the discovery of the phonetic inventory. In the past, we proposed a first approach of first automatically segmenting recordings into phone-line units and second clustering these segments based on acoustic similarity, determined by articulatory features (AFs). We now propose a refined method using Deep Bi-directional LSTMs (DBLSTMs) over DNNs. Additionally, we use Language Feature Vectors (LFVs) which encode language specific peculiarities in a low dimensional representation. In contrast to adding LFVs to the acoustic input features, we modulated the output of the last hidden LSTM layer, forcing groups of LSTM cells to adapt to language related features. We evaluated our approach multilingually, using data from multiple languages. Results show an improvement in recognition accuracy across AF types: While LFVs improved the performance of DNNs, the gain is even bigger when using DBLSTMs. Markus Müller 0001, Sebastian Stüker, Alex Waibel |
ASRU | 3 |
| 2017 | Neural Lattice-to-Sequence Models for Uncertain InputsabstractThe input to a neural sequence-tosequence model is often determined by an up-stream system, e.g. a word segmenter, part of speech tagger, or speech recognizer.These up-stream models are potentially error-prone.Representing inputs through word lattices allows making this uncertainty explicit by capturing alternative sequences and their posterior probabilities in a compact form.In this work, we extend the TreeLSTM (Tai et al., 2015) into a LatticeLSTM that is able to consume word lattices, and can be used as encoder in an attentional encoderdecoder model.We integrate lattice posterior scores into this architecture by extending the TreeLSTM's child-sum and forget gates and introducing a bias term into the attention mechanism.We experiment with speech translation lattices and report consistent improvements over baselines that translate either the 1-best hypothesis or the lattice without posterior scores. Matthias Sperber, Graham Neubig, Jan Niehues, Alex Waibel |
EMNLP | 4 |
| 2017 | Keynote TalkabstractNo abstract available. Alex Waibel |
HAI | 1 |
| 2017 | Towards phoneme inventory discovery for documentation of unwritten languagesabstractDocumenting unwritten languages is a challenging task, even for trained specialists. To help linguists in better and faster documenting new languages is the goal of the French-German ANR-DFG project BULB. To discover the phonetic inventory of a language the project follows three steps: estimating phoneme boundaries, classifying articulatory features (AFs) for each individual segment and clustering the segments into a phoneme inventory. In this work, we focus on estimating the phoneme boundaries and the extraction of AFs, but also perform a first simple clustering based on the recognized AFs. We demonstrate that our Deep Bidirectional LSTM-based approach for identifying phoneme boundaries achieves state-of-the-art performance and evaluate AF extraction based on feed forward neural networks. Markus Müller 0001, Jörg Franke, Alex Waibel, Sebastian Stüker |
ICASSP | 3 |
| 2017 | NMT-Based Segmentation and Punctuation Insertion for Real-Time Spoken Language Translation
Eunah Cho, Jan Niehues, Alex Waibel |
INTERSPEECH | 3 |
| 2017 | Enhancing Backchannel Prediction Using Word Embeddings
Robin Ruede, Markus Müller 0001, Sebastian Stüker, Alex Waibel |
INTERSPEECH | 4 |
| 2017 | Comparison of Decoding Strategies for CTC Acoustic ModelsabstractConnectionist Temporal Classification has recently attracted a lot of interest as it offers an elegant approach to building acoustic models (AMs) for speech recognition. The CTC loss function maps an input sequence of observable feature vectors to an output sequence of symbols. Output symbols are conditionally independent of each other under CTC loss, so a language model (LM) can be incorporated conveniently during decoding, retaining the traditional separation of acoustic and linguistic components in ASR. For fixed vocabularies, Weighted Finite State Transducers provide a strong baseline for efficient integration of CTC AMs with n-gram LMs. Character-based neural LMs provide a straight forward solution for open vocabulary speech recognition and all-neural models, and can be decoded with beam search. Finally, sequence-to-sequence models can be used to translate a sequence of individual sounds into a word string. We compare the performance of these three approaches, and analyze their error patterns, which provides insightful guidance for future research and development in this important area. Thomas Zenkel, Ramon Sanabria, Florian Metze, Jan Niehues, Matthias Sperber, Sebastian Stüker, Alex Waibel |
INTERSPEECH | 7 |
| 2017 | Transcribing against time
Matthias Sperber, Graham Neubig, Jan Niehues, Satoshi Nakamura 0001, Alex Waibel |
Speech Commun. | 5 |
| 2016 | Pre-Translation for Neural Machine TranslationabstractRecently, the development of neural machine translation (NMT) has significantly improved the translation quality of automatic machine translation. While most sentences are more accurate and fluent than translations by statistical machine translation (SMT)-based systems, in some cases, the NMT system produces translations that have a completely different meaning. This is especially the case when rare words occur. When using statistical machine translation, it has already been shown that significant gains can be achieved by simplifying the input in a preprocessing step. A commonly used example is the pre-reordering approach. In this work, we used phrase-based machine translation to pre-translate the input into the target language. Then a neural machine translation system generates the final hypothesis using the pre-translation. Thereby, we use either only the output of the phrase-based machine translation (PBMT) system or a combination of the PBMT output and the source sentence. We evaluate the technique on the English to German translation task. Using this approach we are able to outperform the PBMT system as well as the baseline neural MT system by up to 2 BLEU points. We analyzed the influence of the quality of the initial system on the final result. Jan Niehues, Eunah Cho, Thanh-Le Ha, Alex Waibel |
COLING | 4 |
| 2016 | Lightly Supervised Quality EstimationabstractEvaluating the quality of output from language processing systems such as machine translation or speech recognition is an essential step in ensuring that they are sufficient for practical use. However, depending on the practical requirements, evaluation approaches can differ strongly. Often, reference-based evaluation measures (such as BLEU or WER) are appealing because they are cheap and allow rapid quantitative comparison. On the other hand, practitioners often focus on manual evaluation because they must deal with frequently changing domains and quality standards requested by customers, for which reference-based evaluation is insufficient or not possible due to missing in-domain reference data (Harris et al., 2016). In this paper, we attempt to bridge this gap by proposing a framework for lightly supervised quality estimation. We collect manually annotated scores for a small number of segments in a test corpus or document, and combine them with automatically predicted quality scores for the remaining segments to predict an overall quality estimate. An evaluation shows that our framework estimates quality more reliably than using fully automatic quality estimation approaches, while keeping annotation effort low by not requiring full references to be available for the particular domain. Matthias Sperber, Graham Neubig, Jan Niehues, Sebastian Stüker, Alex Waibel |
COLING | 5 |
| 2016 | An empirical exploration of CTC acoustic modelsabstractThe connectionist temporal classification (CTC) loss function has several interesting properties relevant for automatic speech recognition (ASR): applied on top of deep recurrent neural networks (RNNs), CTC learns the alignments between speech frames and label sequences automatically, which removes the need for pre-generated frame-level labels. CTC systems also do not require context decision trees for good performance, using context-independent (CI) phonemes or characters as targets. This paper presents an extensive exploration of CTC-based acoustic models applied to a variety of ASR tasks, including an empirical study of the optimal configuration and architectural variants for CTC. We observe that on large amounts of training data, CTC models tend to outperform state-of-the-art hybrid approach. Further experiments reveal that CTC can be readily ported to syllable-based languages, and can be enhanced by employing improved feature front-ends. Yajie Miao, Mohammad Gowayyed, Xingyu Na, Tom Ko, Florian Metze, Alex Waibel |
ICASSP | 6 |
| 2016 | Language Adaptive DNNs for Improved Low Resource Speech Recognition
Markus Müller 0001, Sebastian Stüker, Alex Waibel |
INTERSPEECH | 3 |
| 2016 | Dynamic Transcription for Low-Latency Speech Translation
Jan Niehues, Thai Son Nguyen, Eunah Cho, Thanh-Le Ha, Kevin Kilgour, Markus Müller 0001, Matthias Sperber, Sebastian Stüker, Alex Waibel |
INTERSPEECH | 9 |
| 2016 | Unsupervised Phoneme Segmentation of Previously Unseen Languages
Marco Vetter, Markus Müller 0001, Fatima Hamlaoui, Graham Neubig, Satoshi Nakamura 0001, Sebastian Stüker, Alex Waibel |
INTERSPEECH | 7 |
| 2016 | Evaluation of the KIT Lecture Translation System
Markus Müller 0001, Sarah Fünfer, Sebastian Stüker, Alex Waibel |
LREC | 4 |
| 2016 | Optimizing Computer-Assisted Transcription Quality with Iterative User Interfaces
Matthias Sperber, Graham Neubig, Satoshi Nakamura 0001, Alex Waibel |
LREC | 4 |
| 2015 | Stripping Adjectives: Integration Techniques for Selective Stemming in SMT Systems
Isabel Slawik, Jan Niehues, Alex Waibel |
EAMT | 3 |
| 2015 | Combination of NN and CRF models for joint detection of punctuation and disfluenciesabstractInserting proper punctuation marks and deleting speech disfluencies are two of the most essential tasks in spoken language processing. This challenging task has prompted extensive research using various techniques, such as conditional random fields. Neural networks, however, are relatively under-explored for this task. Combining different modeling techniques with different advantages has the potential to lead to improvements. In this work, we first establish the performance of joint modeling of punctuation prediction and disfluency detection using neural networks. We then combine a conditional random fields based model and a neural networks based model log-linearly, and show that the combined approach outperforms both individual models, by 2.7% and 3.5% in F-score for speech disfluency and punctuation detection, respectively. When used as a preprocessing step to machine translation this also results in an improved translation quality of 2.5 BLEU points compared to the baseline and of 0.6 BLEU points compared to the non-combined model. Index Terms: speech disfluency detection, punctuation insertion, speech translation Eunah Cho, Kevin Kilgour, Jan Niehues, Alex Waibel |
INTERSPEECH | 4 |
| 2015 | Gaussian free cluster tree construction using deep neural networkabstractThis paper presents a Gaussian free approach to constructing the cluster tree (CT) that context dependent acoustic models (CDAM) depend on. Over the last few years deep neural networks (DNN) have supplanted Gaussian mixture models (GMM) as the default method for acoustic modeling (AM). DNN AMs have also been successfully used to flat start context independent (CI) AMs and generate alignments on which CTs can be trained. Those approaches however still required Gaussians to build their CTs. Our proposed Gaussian free CT algorithm eliminates this requirements and allows, for the first time, the flat start training of state of the art DNN AMs without the use of Gaussian. An evaluation on the IWSLT transcription task demonstrates the effectiveness of this approach. Linchen Zhu, Kevin Kilgour, Sebastian Stüker, Alex Waibel |
INTERSPEECH | 4 |
| 2015 | Evaluation of Crowdsourced User Input Data for Spoken Dialog SystemsabstractMaria Schmidt, Markus Müller, Martin Wagner, Sebastian Stüker, Alex Waibel, Hansjörg Hofmann, Steffen Werner. Proceedings of the 16th Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2015. Maria Schmidt, Markus Müller 0001, Sebastian Stüker, Alex Waibel, Hansjörg Hofmann, Steffen Werner |
SIGDIAL Conference | 5 |
| 2014 | Tight Integration of Speech Disfluency Removal into SMTabstractSpeech disfluencies are one of the main challenges of spoken language processing.Conventional disfluency detection systems deploy a hard decision, which can have a negative influence on subsequent applications such as machine translation.In this paper we suggest a novel approach in which disfluency detection is integrated into the translation process.We train a CRF model to obtain a disfluency probability for each word.The SMT decoder will then skip the potentially disfluent word based on its disfluency probability.Using the suggested scheme, the translation score of both the manual transcript and ASR output is improved by around 0.35 BLEU points compared to the CRF hard decision system. Eunah Cho, Jan Niehues, Alex Waibel |
EACL | 3 |
| 2014 | Optimization of Neural Network Language Models for keyword searchabstractRecent works have shown Neural Network based Language Models (NNLMs) to be an effective modeling technique for Automatic Speech Recognition. Prior works have shown that these models obtain lower perplexity and word error rate (WER) compared to both standard n-gram language models (LMs) and more advanced language models including maximum entropy and random forest LMs. While these results are compelling, prior works were limited to evaluating NNLMs on perplexity and word error rate. Our initial results showed that while NNLMs improved speech recognition accuracy, the improvement in keyword search was negligible. In this paper we propose alternate optimizations of NNLMs for the task of keyword search. We evaluate the performance of the proposed methods for keyword search on the Vietnamese dataset provided in phase one of the BABEL1project and demonstrate that by penalizing low frequency words during NNLM training, keyword search metrics such as actual term weighted value (ATWV) can be improved by up to 9.3% compared to the standard training methods. Ankur Gandhe, Florian Metze, Alex Waibel, Ian Lane |
ICASSP | 3 |
| 2014 | Multilingual shifting deep bottleneck features for low-resource ASRabstractIn this work, we propose a deep bottleneck feature architecture that is able to leverage data from multiple languages. We also show that tonal features are helpful for non-tonal languages. Evaluations are performed on a low-resource conversational telephone speech transcription task in Bengali, while additional data for DBNF training is provided in Assamese, Pashto, Tagalog, Turkish, and Vietnamese. We obtain relative reductions of up to 17.3% and 9.4% WER over mono-lingual GMMs and DBNFs, respectively. Quoc Bao Nguyen, Jonas Gehring, Markus Müller 0001, Sebastian Stüker, Alex Waibel |
ICASSP | 5 |
| 2014 | Training time reduction and performance improvements from multilingual techniques on the BABEL ASR taskabstractIn the IARPA sponsored program BABEL we are faced with the challenge of training automatic speech recognition systems in sparse data conditions in very little time. In this paper we show that by using multilingual bootstrapping techniques in combination with multilingual deep belief bottle neck features that are only fine tuned on the target language the training time of an LVCSR system can be essentially halved while the word error rate stays the same. We show this for recognition systems on Tagalog, making use of multilingual systems trained on the other four languages of the Babel base period: Cantonese, Pashto, Turkish, and Vietnamese. Sebastian Stüker, Markus Müller 0001, Quoc Bao Nguyen, Alex Waibel |
ICASSP | 4 |
| 2014 | A World without Barriers: Connecting the World across Languages, Distances and MediaabstractAs our world becomes increasingly interdependent and globalization brings people together more than ever, we quickly discover that it is no longer the absence of connectivity (the "digital divide") that separates us, but that new and different forms of alienation still keep us apart, including language, culture, distance and interfaces. Can technology provide solutions to bring us closer to our fellow humans? Alex Waibel |
ICMI | 1 |
| 2014 | A Corpus of Spontaneous Speech in Lectures: The KIT Lecture Corpus for Spoken Language Processing and Translation
Eunah Cho, Sarah Fünfer, Sebastian Stüker, Alex Waibel |
LREC | 4 |
| 2014 | Manual Analysis of Structurally Informed Reordering in German-English Machine Translation
Teresa Herrmann, Jan Niehues, Alex Waibel |
LREC | 3 |
| 2014 | On-the-fly user modeling for cost-sensitive correction of speech transcriptsabstractWe propose an on-the-fly updating framework for cost-sensitive manual correction of automatically recognized speech transcripts. This framework trains cost-models during the transcription process, and does not require the transcriber enrollment necessary in previous work. We use a baseline method that optimizes a segmentation into segments to supervise or not to supervise in a cost-sensitive fashion that minimizes human effort, and introduce a much faster algorithm for computing such a segmentation that can be used for on-the-fly updates. Besides removing the need to carry out enrollments, experiments show that our updating framework results in 28% higher human supervision efficiency than previous cost-sensitive approaches. Matthias Sperber, Graham Neubig, Satoshi Nakamura 0001, Alex Waibel |
SLT | 4 |
| 2014 | Segmentation for Efficient Supervised Language Annotation with an Explicit Cost-Utility TradeoffabstractIn this paper, we study the problem of manually correcting automatic annotations of natural language in as efficient a manner as possible. We introduce a method for automatically segmenting a corpus into chunks such that many uncertain labels are grouped into the same chunk, while human supervision can be omitted altogether for other segments. A tradeoff must be found for segment sizes. Choosing short segments allows us to reduce the number of highly confident labels that are supervised by the annotator, which is useful because these labels are often already correct and supervising correct labels is a waste of effort. In contrast, long segments reduce the cognitive effort due to context switches. Our method helps find the segmentation that optimizes supervision efficiency by defining user models to predict the cost and utility of supervising each segment and solving a constrained optimization problem balancing these contradictory objectives. A user study demonstrates noticeable gains over pre-segmented, confidence-ordered baselines on two natural language processing tasks: speech transcription and word segmentation. Matthias Sperber, Mirjam Simantzik, Graham Neubig, Satoshi Nakamura 0001, Alex Waibel |
Trans. Assoc. Comput. Linguistics | 5 |
| 2013 | DNN acoustic modeling with modular multi-lingual feature extraction networksabstractIn this work, we propose several deep neural network architectures that are able to leverage data from multiple languages. Modularity is achieved by training networks for extracting high-level features and for estimating phoneme state posteriors separately, and then combining them for decoding in a hybrid DNN/HMM setup. This approach has been shown to achieve superior performance for single-language systems, and here we demonstrate that feature extractors benefit significantly from being trained as multi-lingual networks with shared hidden representations. We also show that existing mono-lingual networks can be re-used in a modular fashion to achieve a similar level of performance without having to train new networks on multi-lingual data. Furthermore, we investigate in extending these architectures to make use of language-specific acoustic features. Evaluations are performed on a low-resource conversational telephone speech transcription task in Vietnamese, while additional data for acoustic model training is provided in Pashto, Tagalog, Turkish, and Cantonese. Improvements of up to 17.4% and 13.8% over mono-lingual GMMs and DNNs, respectively, are obtained. Jonas Gehring, Quoc Bao Nguyen, Florian Metze, Alex Waibel |
ASRU | 4 |
| 2013 | Models of tone for tonal and non-tonal languagesabstractConventional wisdom in automatic speech recognition asserts that pitch information is not helpful in building speech recognizers for non-tonal languages and contributes only modestly to performance in speech recognizers for tonal languages. To maintain consistency between different systems, pitch is therefore often ignored, trading the slight performance benefits for greater system uniformity/ simplicity. In this paper, we report results that challenge this conventional approach. We present new models of tone that deliver consistent performance improvements for tonal languages (Cantonese, Vietnamese) and even modest improvements for non-tonal languages. Using neural networks for feature integration and fusion, these models achieve significant gains throughout, and provide us with system uniformity and standardization across all languages, tonal and non-tonal. Florian Metze, Zaid Sheikh, Alex Waibel, Jonas Gehring, Kevin Kilgour, Quoc Bao Nguyen, Van Huy Nguyen |
ASRU | 3 |
| 2013 | Extracting deep bottleneck features using stacked auto-encodersabstractIn this work, a novel training scheme for generating bottleneck features from deep neural networks is proposed. A stack of denoising auto-encoders is first trained in a layer-wise, unsupervised manner. Afterwards, the bottleneck layer and an additional layer are added and the whole network is fine-tuned to predict target phoneme states. We perform experiments on a Cantonese conversational telephone speech corpus and find that increasing the number of auto-encoders in the network produces more useful features, but requires pre-training, especially when little training data is available. Using more unlabeled data for pre-training only yields additional gains. Evaluations on larger datasets and on different system setups demonstrate the general applicability of our approach. In terms of word error rate, relative improvements of 9.2% (Cantonese, ML training), 9.3% (Tagalog, BMMI-SAT training), 12% (Tagalog, confusion network combinations with MFCCs), and 8.7% (Switchboard) are achieved. Jonas Gehring, Yajie Miao, Florian Metze, Alex Waibel |
ICASSP | 4 |
| 2013 | Warped Minimum Variance Distortionless Response based bottle neck features for LVCSRabstractThis paper presents the results of our experiments on bottleneck feature applied to a wMVDR (Warped Minimum Variance Distortionless Response) frontend. We examine how to best optimize wMVDR-BNF features and wMVDR combined with MFCC bottleneck features (wMVDR+MFCC-BNF). Our wMVDR+MFCC-BNF frontend improves a single pass system from 18.7% (20.7%) to 18.1% compared to a MFCC-BNF (MFCC) system tested on the Quaero 2010 German evaluation set. When used in a system combination our wMVDR-BNF and wMVDR+MFCC-BNF systems reduced the overall WER from 14.3% to 13.3% on the IWSLT 2010 test set while at the same time reducing the number of systems needed from 9 to 5. Our result of 11.9% on the 2012 IWSLT testset is better than the best result submitted during the evaluation campaign. Kevin Kilgour, Igor Tseyzer, Quoc Bao Nguyen, Alex Waibel |
ICASSP | 4 |
| 2013 | Subspace mixture model for low-resource speech recognition in cross-lingual settingsabstractThe subspace Gaussian mixture model (SGMM) has been exploited for cross-lingual speech recognition. The general motivation is that the subspace parameters can be estimated on multiple source languages and then transferred to the target language. In this work, we investigate an extension to SGMM, referred to as subspace mixture model (SMM), in which subspace parameters on the target language are casted as a linear mixture of the subspaces derived from source languages. This approach reduces the number of SGMM model parameters, while retaining the flexibility of subspace learning on the target language. Experiments show that the proposed SMM method outperforms SGMM significantly when the target language has limited training data. Yajie Miao, Florian Metze, Alex Waibel |
ICASSP | 3 |
| 2013 | Learning discriminative basis coefficients for eigenspace MLLR unsupervised adaptationabstractEigenspace MLLR is effective for fast adaptation when the amount of adaptation data is limited, e.g., less than 5s. The general motivation is to represent the MLLR transform as a linear combination of basis matrices. In this paper, we present a framework to estimate a speaker-independent discriminative transform over the combination coefficients. This discriminative basis coefficients transform (DBCT) is learned by optimizing discriminative criteria over all the training speakers. During recognition, the ML basis coefficients for each testing speaker are firstly found, on which DBCT is applied to give the final MLLR transform discrimination ability. Experiments show that DBCT results in consistent WER reduction in unsupervised adaptation, compared with both standard ML and discriminatively trained transforms. Yajie Miao, Florian Metze, Alex Waibel |
ICASSP | 3 |
| 2013 | A real-world system for simultaneous translation of German lecturesabstractWe present a real-time automatic speech translation system for university lectures that can interpret several lectures in parallel. University lectures are characterized by a multitude of diverse topics and a large amount of technical terms. This poses specific challenges, e.g., a very specific vocabulary and language model are needed. In addition, in order to be able to translate simultaneously, i.e., to interpret the lectures, the components of the systems need special modifications. The output of the system is delivered in the form or realtime subtitles via a web site that can be accessed by the students attending the lecture through mobile phones, tablet computers or laptops. We evaluated the system on our German to English lecture translation task at the Karlsruhe Institute of Technology. The system is now being installed in several lecture halls at KIT and is able to provide the translation to the students in several parallel sessions. Eunah Cho, Christian Fügen, Teresa Herrmann, Kevin Kilgour, Mohammed Mediani, Christian Mohr, Jan Niehues, Kay Rottmann, Christian Saam, Sebastian Stüker, Alex Waibel |
INTERSPEECH | 11 |
| 2013 | Modular combination of deep neural networks for acoustic modelingabstractIn this work, we propose a modular combination of two pop-ular applications of neural networks to large-vocabulary con-tinuous speech recognition. First, a deep neural network is trained to extract bottleneck features from frames of mel scale filterbank coefficients. In a similar way as is usually done for GMM/HMM systems, this network is then applied as a non-linear discriminative feature-space transformation for a hybrid setup where acoustic modeling is performed by a deep belief network. This effectively results in a very large network, where the layers of the bottleneck network are fixed and applied to suc-cessive windows of feature frames in a time-delay fashion. We show that bottleneck features improve the recognition perfor-mance of DBN/HMM hybrids, and that the modular combina-tion enables the acoustic model to benefit from a larger tempo-ral context. Our architecture is evaluated on a recently released and challenging Tagalog corpus containing conversational tele-phone speech. Jonas Gehring, Wonkyum Lee, Kevin Kilgour, Ian Lane, Yajie Miao, Alex Waibel |
INTERSPEECH | 6 |
| 2013 | Efficient speech transcription through respeakingabstractWe propose a method for efficient off-line speech transcription through respeaking. Speech is segmented into smaller utterances using an initial automatic transcript. Respeaking is performed segment by segment, while confidence filtering helps save supervision effort. We conduct detailed experiments comparing speaking vs. typing, sequential vs. confidence-ordered supervision, and examine the effect of the respeaking word error rate on correction efficiency. Our results demonstrate that the proposed method can not only outperform typing in terms of correction efficiency, but is also much less demanding for the respeakers than traditional respeaking methods, consequently helping to keep costs down. Matthias Sperber, Graham Neubig, Christian Fügen, Satoshi Nakamura 0001, Alex Waibel |
INTERSPEECH | 5 |
| 2013 | Measuring the Structural Importance through Rhetorical Structure Index
Narine Kokhlikyan, Alex Waibel, Joy Ying Zhang |
HLT-NAACL | 2 |
| 2013 | Training speech translation from audio recordings of interpreter-mediated communication
Matthias Paulik, Alex Waibel |
Comput. Speech Lang. | 2 |
| 2012 | A hybrid phonotactic language identification system with an SVM back-end for simultaneous lecture translationabstractIn this paper we describe our work in constructing a language identification system for use in our simultaneous lecture translation system. We first built PPR and PPRLM baseline systems that produce score-fusing language cue feature vectors for language discrimination and utilize an SVM back-end classifier for the actual language identification. On our bi-lingual lecture tasks the PPRLM system clearly outperforms the PPR system in various segment length conditions, however at the cost of slower run-time. By using lexical information in the form of keyword spotting, and additional language models we show ways to improve the performance of both baseline systems. In order to combine the faster run-time of the PPR system with the better performance of the PPRLM system we finally built a hybrid of both approaches that clearly outperforms the PPR system while not adding any additional computing time. This hybrid system is therefore our choice for the use in the lecture translation system due to its faster run-time and good performance. Michael Heck, Sebastian Stüker, Alex Waibel |
ICASSP | 3 |
| 2012 | Blind dereverberation of sinusoid signals using PLL-based combined phase and amplitude analysisabstractIn this paper a new `blind' single-microphone method for the dereverberation of sinusoid signals is presented. A phase-locked-loop is utilized to precisely track the amplitude, frequency and phase offset of a reverberated recording. This information can then be combined to calculate the amplitude and phase offset of single reverberated wavefronts, which allows to subtract them from the original recording. Experimental results have shown that the direct-to-reverberant ratio of recordings can be improved to an extent equal to a delay-and-sum beamformer with 5 microphones. At the end, extensions are outlined which might make the method suitable for dereverberation of real speech. Ralf Huber, Florian Kraft, Alex Waibel |
ICASSP | 3 |
| 2012 | Unsupervised vocabulary selection for real-time speech recognition of lecturesabstractIn this work, we propose a novel method for vocabulary selection to automatically adapt automatic speech recognition systems to the diverse topics that occur in educational and scientific lectures. Utilizing materials that are available before the lecture begins, such as lecture slides, our proposed framework iteratively searches for related documents on the web and generates a lecture-specific vocabulary based on the resulting documents. In this paper, we propose a novel method for vocabulary selection where we first collect documents similar to an initial seed document and then rank the resulting vocabulary based on a score which is calculated using a combination of word features. This is a critical component for adaptation that has typically been overlooked in prior works. On the inter ACT German-English simultaneous lecture translation system our proposed approach significantly improved vocabulary coverage, reducing the out-of-vocabulary rate, on average by 57.0% and up to 84.9%, compared to a lecture-independent baseline. Furthermore, our approach reduced the word error rate, by 12.5% on average and up to 25.3%, compared to a lecture-independent baseline. Paul Maergner, Alex Waibel, Ian Lane |
ICASSP | 2 |
| 2012 | The KIT Lecture Corpus for Speech Translation
Sebastian Stüker, Florian Kraft, Christian Mohr, Teresa Herrmann, Eunah Cho, Alex Waibel |
LREC | 6 |
| 2011 | TriS: A Statistical Sentence Simplifier with Log-linear Models and Margin-based Discriminative Training
Nguyen Bach, Qin Gao, Stephan Vogel, Alex Waibel |
IJCNLP | 4 |
| 2011 | Unsupervised Vocabulary Selection for Domain-Independent Simultaneous Lecture Translation
Paul Maergner, Ian Lane, Alex Waibel |
MTSummit | 3 |
| 2010 | Domain Adaptation in Statistical Machine Translation using Factored Translation Models
Jan Niehues, Alex Waibel |
EAMT | 2 |
| 2010 | Spoken language translation from parallel speech audio: Simultaneous interpretation as SLT training dataabstractIn recent work, we proposed an alternative to parallel text as translation model (TM) training data: audio recordings of parallel speech (pSp), as it occurs in any communication scenario where interpreters are involved. Although interpretation compares poorly to translation, we reported surprisingly strong translation results for systems based on pSp trained TMs. This work extends the use of pSp as a data source for unsupervised training of all major models involved in statistical spoken language translation. We consider the scenario of speech translation between a resource rich and a resource-deficient language. Our seed models are based on 10h of transcribed audio and parallel text comprised of 100k translated words. With the help of 92h of untranscribed pSp audio, and by taking advantage of the redundancy inherent to pSp (the same information is given twice, in two languages), we report significant improvements for the resource-deficient acoustic, language and translation models. Matthias Paulik, Alex Waibel |
ICASSP | 2 |
| 2010 | Named-entity projection and data-driven morphological decomposition for field maintainable speech-to-speech translation systems
Ian Lane, Alex Waibel |
INTERSPEECH | 2 |
| 2010 | Rapid development of speech translation using consecutive interpretationabstractThe development of a speech translation (ST) system is costly, largely because it is expensive to collect parallel data. A new language pair is typically only considered in the aftermath of an international crisis that incurs a major need of crosslingual communication. Urgency justifies the deployment of interpreters while data is being collected. In recent work, we have shown that audio recordings of interpreter-mediated communication can present a low-cost data resource for the rapid development of automatic text and speech translation. However, our previous experiments remain limited to English/Spanish simultaneous interpretation. In this work, we examine our approaches for exploiting interpretation audio as translation model training data in the context of English/Pashto consecutive interpretation. We show that our previously made findings remain valid, despite the more complex language pair and the additional challenges introduced by the strong resource-limitations of Pashto. Index Terms: speech translation, machine translation, parallel speech Matthias Paulik, Alex Waibel |
INTERSPEECH | 2 |
| 2010 | Jibbigo: Speech-to-speech translation on mobile devicesabstractJibbigo is a speech-to-speech translation application for iPhone, iPod touch, and iPad devices. Jibbigo allows the user to simply speak a sentence, and it speaks the sentence aloud in the other language, much like a personal human interpreter would. The speech-to-speech translation is bidirectional for a two way dialog between participants. Matthias Eck 0001, Ian Lane, Ying Zhang 0048, Alex Waibel |
SLT | 4 |
| 2009 | Pronunciation modeling for dialectal arabic speech recognitionabstractShort vowels in Arabic are normally omitted in written text which leads to ambiguity in the pronunciation. This is even more pronounced for dialectal Arabic where a single word can be pronounced quite differently based on the speaker's nationality, level of education, social class and religion. In this paper we focus on pronunciation modeling for Iraqi-Arabic speech. We introduce multiple pronunciations into the Iraqi speech recognition lexicon, and compare the performance, when weights computed via forced alignment are assigned to the different pronunciations of a word. Incorporating multiple pronunciations improved recognition accuracy compared to a single pronunciation baseline and introducing pronunciation weights further improved performance. Using these techniques an absolute reduction in word-error-rate of 2.4% was obtained compared to the baseline system. Hassan Al-Haj, Roger Hsiao, Ian Lane, Alan W. Black, Alex Waibel |
ASRU | 5 |
| 2009 | Automatic translation from parallel speech: Simultaneous interpretation as MT training dataabstractState-of-the art statistical machine translation depends heavily on the availability of domain-specific bilingual parallel text. However, acquiring large amounts of bilingual parallel text is costly and, depending on the language pair, sometimes impossible. We propose an alternative to parallel text as machine translation (MT) training data; audio recordings of parallel speech (pSp) as it occurs in any scenario where interpreters are involved. Although interpretation (pSp) differs significantly from translation (parallel text), we achieve surprisingly strong translation results with our pSp-trained MT and speech translation systems.We argue that the presented approach is of special interest for developing speech translation in the context of resource-deficient languages where even monolingual resources are scarce. Matthias Paulik, Alex Waibel |
ASRU | 2 |
| 2009 | End-to-End Evaluation in Simultaneous Translation
Olivier Hamon, Christian Fügen, Djamel Mostefa, Victoria Arranz, Muntsin Kolss, Alex Waibel, Khalid Choukri |
EACL | 6 |
| 2009 | Human translations guided language discovery for ASR systemsabstractInternational audience Sebastian Stüker, Laurent Besacier, Alex Waibel |
INTERSPEECH | 3 |
| 2008 | Extracting clues from human interpreter speech for spoken language translationabstractIn previous work, we reported dramatic improvements in automatic speech recognition (ASR) and spoken language translation (SLT) gained by applying information extracted from spoken human interpretations. These interpretations were artificially created by collecting read sentences from a clean parallel text corpus. Real human interpretations are significantly different. They suffer from frequent synopses, omissions and self-corrections. Expressing these differences in BLEU score by evaluating human interpretations with carefully created human translations, we found that human interpretations perform two to three times worse than state-of-the art SLT. Facing these stark differences, we address the question if and how ASR and SLT can profit from human interpretations. In the following we describe initial experiments that apply knowledge derived from real human interpretations for improving English and Spanish ASR and SLT. Our experiments are conducted on a small European Parliamentary Plenary Sessions development set. Matthias Paulik, Alex Waibel |
ICASSP | 2 |
| 2008 | Stream decoding for simultaneous spoken language translationabstractIn the typical speech translation system, the first-best speech recognizer hypothesis is segmented into sentence-like units which are then fed to the downstream machine translation component. The need for a sufficiently large context in this intermediate step and for the MT introduces delays which are undesirable in many application scenarios, such as real-time subtitling of foreign language broadcasts or simultaneous translation of speeches and lectures. In this paper, we propose a statistical machine translation decoder which processes a continuous input stream, such as that produced by a run-on speech recognizer. By decoupling decisions about the timing of translation output generation from any fixed input segmentation, this design can guarantee a maximum output lag for each input word while allowing for full word reordering within this time window. Experimental results show that this system achieves competitive translation performance with a minimum of translationinduced latency. Muntsin Kolss, Stephan Vogel, Alex Waibel |
INTERSPEECH | 3 |
| 2008 | Class-based statistical machine translation for field maintainable speech-to-speech translationabstractCurrent speech-to-speech translation systems lack any mechanism to handle out-of-vocabulary words that did not appear in the training data. To improve the usability of these systems we have developed a field maintainable speech-to-speech translation framework that enables users to add new words to the system while it is being used in the field. To realize such a framework, a novel class-based statistical machine translation framework is proposed, that applies class-based translation models and class n-gram language models during translation. To obtain consistent labelling of the parallel training corpora, on which these models are trained, we introduce a bilingual tagger that jointly labels both sides of the parallel corpora. On a Japanese-English evaluation system, the proposed framework significantly improved translation quality, obtaining a relative improvement in BLEU-score of 15% for both translation directions. Ian Lane, Alex Waibel |
INTERSPEECH | 2 |
| 2008 | Lightly supervised acoustic model training on EPPS recordingsabstractDebates in the European Parliament are simultaneously translated into the official languages of the Union. These interpretations are broadcast live via satellite on separate audio channels. After several months, the parliamentary proceedings are published as final text editions (FTE). FTEs are formatted for an easy readability and can differ significantly from the original speeches and the live broadcast interpretations. We examine the impact on German word error rate (WER) when introducing supervision based on German FTEs and supervision based on German automatic translations extracted from the English and Spanish audio. We show that FTE based supervision and additional interpretation based supervision provide significant reductions in WER. We successfully apply FTE supervised acoustic model (AM) training using 143h of recordings. Combining the new AM with the mentioned supervision techniques, we achieve a significant WER reduction of 13.3 % relative. Index Terms: lightly supervised acoustic model training, EPPS, speech recognition 1. Matthias Paulik, Alex Waibel |
INTERSPEECH | 2 |
| 2008 | Communicating Unknown Words in Machine Translation
Matthias Eck 0001, Stephan Vogel, Alex Waibel |
LREC | 3 |
| 2008 | Probabilistic integration of sparse audio-visual cues for identity trackingabstractIn the context of smart environments, the ability to track and identify persons is a key factor, determining the scope and flexibility of analytical components or intelligent services that can be provided. While some amount of work has been done concerning the camera-based tracking of multiple users in a variety of scenarios, technologies for acoustic and visual identification, such as face or voice ID, are unfortunately still subjected to severe limitations when distantly placed sensors have to be used. Because of this, reliable cues for identification can be hard to obtain without user cooperation, especially when multiple users are involved. Keni Bernardin, Rainer Stiefelhagen, Alex Waibel |
ACM Multimedia | 3 |
| 2008 | Confidence based multimodal fusion for person identificationabstractPerson identification is of great interest for various kinds of applications and interactive systems. In our system we use face recognition and voice recognition from data recorded in an interactive dialogue system. In such a system, sequential images and sequential utterances can be used to improve recognition accuracy over single hypotheses. The presented approach uses confidence-based fusion for sequence hypotheses, for multimodal fusion, and to provide a reliability measure of the classification quality that can be used to decide when to trust and when to ignore classification results. Philipp Große, Hartwig Holzapfel, Alex Waibel |
ACM Multimedia | 3 |
| 2008 | Modelling multimodal user ID in dialogueabstractThis paper presents an approach to model user ID in dialogue. A belief network is used to integrate ID classifiers, such as face ID and voice ID, and person related information, such as the first name and last name of a person from speech recognition or spelling. Different network structures are analyzed and compared with each other and are compared with a rule-based user model. The approach is evaluated on dialogue data collected in a person identification scenario, which includes both, identification of known persons and interactive learning of names and ID of unknown persons. Hartwig Holzapfel, Alex Waibel |
SLT | 2 |
| 2008 | Simultaneous machine translation of german lectures into english: Investigating research challenges for the futureabstractAn increasingly globalized world fosters the exchange of students, researchers or employees. As a result, situations in which people of different native tongues are listening to the same lecture become more and more frequent. In many such situations, human interpreters are prohibitively expensive or simply not available. For this reason, and because first prototypes have already demonstrated the feasibility of such systems, automatic translation of lectures receives increasing attention. A large vocabulary and strong variations in speaking style make lecture translation a challenging, however not hopeless, task. The scope of this paper is to investigate a variety of challenges and to highlight possible solutions in building a system for simultaneous translation of lectures from German to English. While some of the investigated challenges are more general, e.g. environment robustness, other challenges are more specific for this particular task, e.g. pronunciation of foreign words or sentence segmentation. We also report our progress in building an end-to-end system and analyze its performance in terms of objective and subjective measures. Matthias Wölfel, Muntsin Kolss, Florian Kraft, Jan Niehues, Matthias Paulik, Alex Waibel |
SLT | 6 |
| 2007 | Consolidation based speech translationabstractTo alleviate the degradation of the performance of speech translation, this paper proposes a new approach to translate ASR results through consolidation which extracts meaningful phrases and remove redundant and irrelevant information caused by speaker’s disfluency and recognition errors. The speech translation results via consolidation are partial translation and can not be directly compared with gold standards in which all words are translated. We would like to propose a new evaluation framework for partial translation by comparing with the most similar set of words extracted from a word network created by merging gradual summarizations of the gold standard translation. Chinese broadcast news speech in RT04 were recognized, consolidated and then translated. The performance of MT results was evaluated using BLEU. We propose Information Preservation Accuracy (IPAccy) and Meaning Preservation Accuracy (MPAccy) for consolidation and consolidation-based MT. Chiori Hori, Bing Zhao 0005, Stephan Vogel, Alex Waibel |
ASRU | 4 |
| 2007 | Continuous Electromyographic Speech Recognition with a Multi-Stream Decoding ArchitectureabstractIn our previous work, we reported a surface electromyographic (EMG) continuous speech recognition system with a novel EMG feature extraction method, E4, which is more robust to EMG noise than traditional spectral features. In this paper, we show that articulatory feature (AF) classifiers can also benefit from the E4 feature, which improve the F-score of the AF classifiers from 0.492 to 0.686. We also show that the E4 feature is less correlated across EMG channels and thus channel combination gains larger improvement in F-score. With a stream architecture, the AF classifiers are then integrated into the decoding framework and improve the word error rate by 11.8% relative from 33.9% to 29.9%. Szu-Chen Stan Jou, Tanja Schultz, Alex Waibel |
ICASSP (4) | 3 |
| 2007 | Speech Translation Enhanced ASR for European Parliament Speeches - On the Influence of ASR Performance on Speech TranslationabstractIn this paper we describe our work in coupling automatic speech recognition (ASR) and machine translation (MT) in a speech translation enhanced automatic speech recognition (STE-ASR) framework for transcribing and translating European parliament speeches. We demonstrate the influence of the quality of the ASR component on the MT performance, by comparing a series of WERs with the corresponding automatic translation scores. By porting an STE-ASR framework to the task at hand, we show how the word errors for transcribing English and Spanish speeches can be lowered by 3.0% and 4.8% relative, respectively. Sebastian Stüker, Matthias Paulik, Muntsin Kolss, Christian Fügen, Alex Waibel |
ICASSP (4) | 5 |
| 2007 | Behavior models for learning and receptionist dialogs
Hartwig Holzapfel, Alex Waibel |
INTERSPEECH | 2 |
| 2007 | Computer-supported human-human multilingual communication
Alex Waibel, Keni Bernardin, Matthias Wölfel |
INTERSPEECH | 1 |
| 2007 | Estimating phrase pair relevance for translation model pruning
Matthias Eck 0001, Stephan Vogel, Alex Waibel |
MTSummit | 3 |
| 2007 | Simultaneous translation of lectures and speeches
Christian Fügen, Alex Waibel, Muntsin Kolss |
Mach. Transl. | 2 |
| 2007 | Far-Field Speaker RecognitionabstractIn this paper, we study robust speaker recognition in far-field microphone situations. Two approaches are investigated to improve the robustness of speaker recognition in such scenarios. The first approach applies traditional techniques based on acoustic features. We introduce reverberation compensation as well as feature warping and gain significant improvements, even under mismatched training-testing conditions. In addition, we performed multiple channel combination experiments to make use of information from multiple distant microphones. Overall, we achieved up to 87.1% relative improvements on our Distant Microphone database and found that the gains hold across different data conditions and microphone settings. The second approach makes use of higher-level linguistic features. To capture speaker idiosyncrasies, we apply n-gram models trained on multilingual phone strings and show that higher-level features are more robust under mismatching conditions. Furthermore, we compared the performances between multilingual and multiengine systems, and examined the impact of a number of involved languages on recognition results. Our findings confirm the usefulness of language variety and indicate a language independent nature of this approach, which suggests that speaker recognition using multilingual phone strings could be successfully applied to any given language. Qin Jin, Tanja Schultz, Alex Waibel |
IEEE Trans. Speech Audio Process. | 3 |
| 2007 | Enabling Multimodal Human-Robot Interaction for the Karlsruhe Humanoid RobotabstractIn this paper, we present our work in building technologies for natural multimodal human-robot interaction. We present our systems for spontaneous speech recognition, multimodal dialogue processing, and visual perception of a user, which includes localization, tracking, and identification of the user, recognition of pointing gestures, as well as the recognition of a person's head orientation. Each of the components is described in the paper and experimental results are presented. We also present several experiments on multimodal human-robot interaction, such as interaction using speech and gestures, the automatic determination of the addressee during human-human-robot interaction, as well on interactive learning of dialogue strategies. The work and the components presented here constitute the core building blocks for audiovisual perception of humans and multimodal human-robot interaction used for the humanoid robot developed within the German research project (Sonderforschungsbereich) on humanoid cooperative robots. Rainer Stiefelhagen, Hazim Kemal Ekenel, Christian Fügen, Petra Gieselmann, Hartwig Holzapfel, Florian Kraft, Kai Nickel, Michael Voit, Alex Waibel |
IEEE Trans. Robotics | 9 |
| 2006 | A Flexible Online Server for Machine Translation Evaluation
Matthias Eck 0001, Stephan Vogel, Alex Waibel |
EAMT | 3 |
| 2006 | Open Domain Speech Recognition & Translation: Lectures and SpeechesabstractFor years speech translation has focused on the recognition and translation of discourses in limited domains, such as hotel reservations or scheduling tasks. Only recently research projects have been started to tackle the problem of open domain speech recognition and translation of complex tasks such as lectures and speeches. In this paper we present the on-going work at our laboratory in open domain speech translation of lectures and parliamentary speeches. Starting from a translation system for European parliamentary plenary sessions and a lecture speech recognition system we show how both components perform in unison on speech translation of lectures Christian Fügen, Muntsin Kolss, Dietmar Bernreuther, Matthias Paulik, Sebastian Stüker, Stephan Vogel, Alex Waibel |
ICASSP (1) | 7 |
| 2006 | Articulatory Feature Classification using Surface ElectromyographyabstractIn this paper, we present an approach for articulatory feature classification based on surface electromyographic signals generated by the facial muscles. With parallel recorded audible speech and electromyographic signals, experiments are conducted to show the anticipatory behavior of electromyographic signals with respect to speech signals. On average, we found that the signals to be time delayed by 0.02 to 0.12 second. Furthermore, it is shown that different articulators have different anticipatory behavior. With offset-aligned signals, we improved the average F-score of the articulatory feature classifiers in our baseline system from 0.467 to 0.502. Szu-Chen Stan Jou, Lena Maier-Hein, Tanja Schultz, Alex Waibel |
ICASSP (1) | 4 |
| 2006 | Directing Attention in Online Aggregate Sensor Streams via Auditory Blind Value AssignmentabstractMultiparty collaborative applications in which groups of people act in concert to achieve some real-world goal abound. In these situations, it is useful for a central planning agent to receive online audio-visual information from all participants. However, as the size of the group grows, it becomes difficult to process all the sensory streams; cognitive overload prevents direct analysis of sensory streams for situational awareness. To avoid this situation, an automatic method is needed to assign value to each stream and direct the attention of the planning agent to those streams which are most valuable. We present an audio-based blind value assignment (BVA) method to address this problem, and experiments demonstrating the method's efficacy. We demonstrate that use of audio BVA techniques results in automatic value judgments which are broadly similar to human value judgments and superior to automatic judgments based on video information Robert G. Malkin, Datong Chen, Jie Yang 0001, Alex Waibel |
ICME | 4 |
| 2006 | Multimodal estimation of user interruptibility for smart mobile telephonesabstractContext-aware computer systems are characterized by the ability to consider user state information in their decision logic. One example application of context-aware computing is the smart mobile telephone. Ideally, a smart mobile telephone should be able to consider both social factors (i.e., known relationships between contactor and contactee) and environmental factors (i.e., the contactee's current locale and activity) when deciding how to handle an incoming request for communication.Toward providing this kind of user state information and improving the ability of the mobile phone to handle calls intelligently, we present work on inferring environmental factors from sensory data and using this information to predict user interruptibility. Specifically, we learn the structure and parameters of a user state model from continuous ambient audio and visual information from periodic still images, and attempt to associate the learned states with user-reported interruptibility levels. We report experimental results using this technique on real data, and show how such an approach can allow for adaptation to specific user preferences. Robert G. Malkin, Datong Chen, Jie Yang 0001, Alex Waibel |
ICMI | 4 |
| 2006 | Dynamic extension of a grammar-based dialogue system: constructing an all-recipes knowing robotabstractIn the upcoming field of humanoid and human-friendly robots, the ability of the robot for simple, unconstrained and natural communication with its users is of central importance. The basis for appropriate actions of the robot is the correct understanding of the user utterances. To be able to cover all the entities a user might talk about, we enhanced our dialogue manager with an ability for dynamic vocabulary generation out of information found across the internet. As a test case, we chose an internet recipe database integrated in the dialogue manager of our household robot so that it can understand several thousand recipes and ingredients now. Index Terms: dialogue management, human-robot interaction, vocabulary extension Petra Gieselmann, Alex Waibel |
INTERSPEECH | 2 |
| 2006 | A multilingual expectations model for contextual utterances in mixed-initiative spoken dialogue
Hartwig Holzapfel, Alex Waibel |
INTERSPEECH | 2 |
| 2006 | Optimizing components for handheld two-way speech translation for an English-iraqi Arabic systemabstractThis paper described our handheld two-way speech translation system for English and Iraqi. The focus is on developing a field usable handheld device for speech-to-speech translation. The computation and memory limitations on the handheld impose critical constraints on the ASR, SMT, and TTS components. In this paper we discuss our approaches to optimize these components for the handheld device and present performance numbers from the evaluations that were an integral part of the project. Since one major aspect of the TransTac program is to build fieldable systems, we spent significant effort on developing an intuitive interface that minimizes the training time for users but also provides useful information such as back translations for translation quality feedback. Roger Hsiao, Ashish Venugopal, Thilo Köhler, Ying Zhang 0048, Paisarn Charoenpornsawat, Andreas Zollmann, Stephan Vogel, Alan W. Black, Tanja Schultz, Alex Waibel |
INTERSPEECH | 10 |
| 2006 | Towards continuous speech recognition using surface electromyographyabstractWe present our research on continuous speech recognition of the surface electromyographic signals that are generated by the human articulatory muscles.Previous research on electromyographic speech recognition was limited to isolated word recognition because it was very difficult to train phoneme-based acoustic models for the electromyographic speech recognizer.In this paper, we demonstrate how to train the phoneme-based acoustic models with carefully designed electromyographic feature extraction methods.By decomposing the signal into different feature space, we successfully keep the useful information while reducing the noise.Additionally, we also model the anticipatory effect of the electromyographic signals compared to the speech signal.With a 108-word decoding vocabulary, the experimental results show that the word error rate improves from 86.8% to 32.0% by using our novel feature extraction methods. Szu-Chen Stan Jou, Tanja Schultz, Matthias Walliczek, Florian Kraft, Alex Waibel |
INTERSPEECH | 5 |
| 2006 | Rapid simulation-driven reinforcement learning of multimodal dialog strategies in human-robot interaction
Thomas Prommer, Hartwig Holzapfel, Alex Waibel |
INTERSPEECH | 3 |
| 2006 | Sub-word unit based non-audible speech recognition using surface electromyographyabstractIn this paper we present a novel approach for a surface electromyographic speech recognition system based on sub-word units. Rather than using full word models as integrated in our previous work we propose here smaller sub-word units as prerequisites for large vocabulary speech recognition. This allows the recognition of words not seen in the training set based on seen sub-word units. Therefore we report on experiments with syllables and phonemes as sub-word units. We also developed a new feature extraction method that gains significant improvement for words and sub-word units. Index Terms: silent speech, non-audible speech recognition, electromyography, sub-word unit comparison Matthias Walliczek, Florian Kraft, Szu-Chen Stan Jou, Tanja Schultz, Alex Waibel |
INTERSPEECH | 5 |
| 2005 | Clustering and Classifying Person Names by Origin
Fei Huang 0002, Stephan Vogel, Alex Waibel |
AAAI | 3 |
| 2005 | Augmenting a statistical translation system with a translation memory
Sanjika Hewavitharana, Stephan Vogel, Alex Waibel |
EAMT | 3 |
| 2005 | Adaptation of the translation model for statistical machine translation based on information retrieval
Almut Silja Hildebrand, Matthias Eck 0001, Stephan Vogel, Alex Waibel |
EAMT | 4 |
| 2005 | Whispery Speech Recognition using Adapted Articulatory FeaturesabstractThis paper describes our research on adaptation methods applied to articulatory feature detection on soft whispery speech recorded with a throat microphone. Since the amount of adaptation data is small and the testing data is very different from the training data, a series of adaptation methods is necessary. The adaptation methods include: maximum likelihood linear regression, feature-space adaptation, and re-training with downsampling, sigmoidal low-pass filter, and linear multivariate regression. Adapted articulatory feature detectors are used in parallel to standard senone-based HMM models in a stream architecture for decoding. With these adaptation methods, articulatory feature detection accuracy improves from 87.82% to 90.52% with corresponding F-measure from 0.504 to 0.617, while the final word error rate improves from 33.8% to 31.2%. Szu-Chen Stan Jou, Tanja Schultz, Alex Waibel |
ICASSP (1) | 3 |
| 2005 | Classifying user environment for mobile applications using linear autoencoding of ambient audioabstractMany mobile devices and applications can act in context-sensitive ways, but rely on explicit human action for context awareness. It would be preferable if our devices were able to attain context awareness without human intervention. One important aspect of user context is environment. We present a novel method for classifying environment types based on acoustic signals. This method makes use of linear autoencoding neural networks, and is motivated by the observation that biological coding systems seem to be heavily influenced by the statistics of their environments. We show that the autoencoder method achieved a lower error rate than a standard Gaussian mixture model on a representative sample task, and that a linear combination of autoencoders and GMMs yielded better performance than either alone. Robert G. Malkin, Alex Waibel |
ICASSP (5) | 2 |
| 2005 | Automatically Transcribing Meetings using Distant MicrophonesabstractIn this paper, we describe our efforts to develop acoustic models suitable for distant microphone automatic speech recognition. Our goal is to investigate how the performance of a system trained on a combination of close-talking and distant microphone data can be optimized, while assuming as little information about the configuration of (multiple) distant microphones as possible, to avoid guesstimates and lengthy calibration runs. We evaluated our system in NIST's RT-04S "Meeting" speech-to-text evaluation, where speech data was recorded at several sites with a varying number of different table-top microphones, but not with microphone arrays. Body-mounted microphones provide baseline numbers for distant ASR performance and allow for comparisons of meeting speech with other spontaneous speech data. Florian Metze, Christian Fügen, Alex Waibel |
ICASSP (1) | 4 |
| 2005 | The connector: facilitating context-aware communicationabstractWe present the Connector, a context-aware service that intelligently connects people. It maintains an awareness of its users' activities, preoccupations and social relationships to mediate a proper connection at the right time between them. In addition to providing users with important contextual cues about the availability of potential callees, the Connector adapts the behavior of the contactee's device automatically in order to avoid inappropriate interruptions.To acquire relevant context information, perceptual components analyze sensor input obtained from a smart mobile phone and --- if available --- from a variety of audio-visual sensors built into a smart meeting room environment. The Connector also uses any available multimodal interface (e.g. a speech interface to the smart phone, steerable camera-projector, targeted loudspeakers) in the smart meeting room, to deliver information to users in the most unobtrusive way possible. Maria Danninger, G. Flaherty, Keni Bernardin, Hazim Kemal Ekenel, Thilo Köhler, Robert G. Malkin, Rainer Stiefelhagen, Alex Waibel |
ICMI | 8 |
| 2005 | Spontaneous speech consolidation for spoken language applicationsabstractThis paper describes the work done as a part of the International Workshop on Speech Summarization for Information Extraction and Machine Translation (IWSpS) , on spoken language processing including summarization, machine translation and question answering on lecture speech in the Translanguage English Database (TED) corpus . The hypotheses of lecture speech obtained by automatic speech recognition (ASR) system are ill-formed due to the spontaneity of speakers and recognition errors. The overall performance of spoken language processing components is affected by the errors introduced by the ASR system. In order to get more reliable phrases which maintain the original meaning and contribute positively to the total performance of the spoken language system, this paper proposes a consolidation fram ework. The consolidation approach extracts words by excluding redundant and irrelevant information and concatenating words so as to maintain the original meaning. Automatic consolidation performance is evaluated by comparing with manual consolidation by humans using a word accuracy metric . Our approach gives 58% accuracy on ASR output with 70% word accuracy. Chiori Hori, Alex Waibel |
INTERSPEECH | 2 |
| 2005 | Rapid porting of ASR-systems to mobile devicesabstractPortable devices for the consumer market are becoming available in large quantities. Because of their design and use, human speech often is the input modality of choice, for example for car navigation systems or portable speech-to-speech translation devices. In this paper we describe our work in porting our existing desktop PC based speech recognition system to an off-the-shelf PDA running WindowsCE3.0. We do this in a way that our already well performing language and acoustic models can be taken over without the need of retraining them for the PDA. In order to achieve an acceptable run-time behavior we apply several optimization techniques to the preprocessing and decoding process. Among other things we introduce the newly developed early feature vector reduction. In that way the execution time of our recognition system can be reduced from initially 28x realtime to 2.6x real-time with a tolerable increase in word error rate. The size of the acoustic models is reduced to 25 % of its original size. 1. Thilo Köhler, Christian Fügen, Sebastian Stüker, Alex Waibel |
INTERSPEECH | 4 |
| 2005 | Temporal ICA for classification of acoustic events i a kitchen environmentabstractWe describe a feature extraction method for general audio modeling using a temporal extension of Independent Component Analysis (ICA) and demonstrate its utility in the context of a sound classification task in a kitchen environment. Our approach accounts for temporal dependencies over multiple analysis frames much like the standard audio modeling technique of adding first and second temporal derivatives to the feature set. Using a real-world dataset of kitchen sounds, we show that our approach outperforms a canonical version of this standard front end, the mel-frequency cepstral coefficients (MFCCs), which has found successful application in automatic speech recognition tasks. Florian Kraft, Robert G. Malkin, Thomas Schaaf, Alex Waibel |
INTERSPEECH | 4 |
| 2005 | Clarification questions to improve dialogue flow and speech recognition in spoken dialogue systems
Ulf Krum, Hartwig Holzapfel, Alex Waibel |
INTERSPEECH | 3 |
| 2005 | Document driven machine translation enhanced ASRabstractIn human-mediated translation scenarios a human interpreter translates between a source and a target language using either a spoken or a written representation of the source language. In this paper we improve the recognition performance on the speech of the human translator spoken in the target language by taking advantage of the source language representations. We use machine translation techniques to translate between the source and target language resources and then bias the target language speech recognizer towards the gained knowledge, hence the name Machine Translation Enhanced Automatic Speech Recognition. We investigate several different techniques among which are restricting the search vocabulary, selecting hypotheses from n-best lists, applying cache and interpolation schemes to language modeling, and combining the most successful techniques into our final, iterative system. Overall we outperform the baseline system by a relative word error rate reduction of 37.6%. Matthias Paulik, Christian Fügen, Sebastian Stüker, Tanja Schultz, Thomas Schaaf, Alex Waibel |
INTERSPEECH | 6 |
| 2005 | Low Cost Portability for Statistical Machine Translation based on N-gram CoverageabstractStatistical machine translation relies heavily on the available training data. However, in some cases, it is necessary to limit the amount of training data that can be created for or actually used by the systems. To solve that problem, we introduce a weighting scheme that tries to select more informative sentences first. This selection is based on the previously unseen n-grams the sentences contain, and it allows us to sort the sentences according to their estimated importance. After sorting, we can construct smaller training corpora, and we are able to demonstrate that systems trained on much less training data show a very competitive performance compared to baseline systems using all available training data. Matthias Eck 0001, Stephan Vogel, Alex Waibel |
MTSummit | 3 |
| 2004 | Improving Statistical Machine Translation in the Medical Domain using the Unified Medical Language system
Matthias Eck 0001, Stephan Vogel, Alex Waibel |
COLING | 3 |
| 2004 | Phrase Pair Rescoring with Term Weighting for Statistical Machine Translatio
Bing Zhao 0005, Stephan Vogel, Matthias Eck 0001, Alex Waibel |
EMNLP | 4 |
| 2004 | Performance comparisons of all-pass transform adaptation with maximum likelihood linear regressionabstractAll-pass transform (APT) adaptation transforms the cepstral means of a hidden Markov model so as to mimic the effect of warping the short-time frequency axis of a segment of speech, much like vocal tract length normalization (VTLN). However, APT adaptation can be implemented as a linear transformation in the cepstral domain, much like the better known maximum likelihood linear regression (MLLR). Recent work demonstrated the superior performance of APT adaptation to MLLR for a large vocabulary conversational speech recognition task. This work presents similar comparisons on the switchboard corpus. We found that without VTLN, the best MLLR and APT systems achieved word error rates (WERs) of 43.0% and 40.2% respectively. Similarly, with VTLN the respective error rates were 40.3%, and 39.2%, so that APT adaptation is significantly better in both cases. We also undertook a set of experiments to determine whether APT adaptation can be combined with a linear semi-tied covariance (STC) transform. With a single APT per speaker, the application of STC reduced the WER from 42.9% to 39.4%. John W. McDonough, Alex Waibel |
ICASSP (1) | 2 |
| 2004 | Minimum Kullback-Leibler distance based multivariate Gaussian feature adaptation for distant-talking speech recognitionabstractMultivariate Gaussian based speech compensation or mapping has been developed to reduce the mismatch between training and deployment conditions for robust speech recognition. The acoustic mapping procedure can be formulated as a feature space adaptation where a noisy input signal is transformed by a multivariate Gaussian network. We propose a novel algorithm to update the network parameters based on minimizing the Kullback-Leibler distance between the core recognizer's acoustic model and transformed features. It is designed to achieve optimal overall system performance rather than MMSE on a specific feature domain. An online stochastic gradient descent learning rule is derived. We evaluate the performance of the new algorithm using a JRTk broadcast news system on a distance-talking speech corpus and compare its performance with that of previous MMSE based approaches. The experiments show the KL based approach is more effective for a large vocabulary continuous speech recognition (LVCSR) system. Alex Waibel |
ICASSP (1) | 2 |
| 2004 | Towards language portability in statistical speech translationabstractSpeech translation has made significant advances over the last years. We believe that we can overcome today's limits of language and domain portable conversational speech translation systems by relying more radically on learning approaches and by the use of multiple layers of reduction and transformation to extract the desired content in another language. Therefore, we cascade stochastic source-channel models that extract an underlying message from a corrupt observed output. The three models effectively translate: (1) speech to word lattices (automatic speech recognition, ASR); (2) ill-formed fragments of word strings into a compact well-formed sentence (Clean); (3) sentences in one language to sentences in another (machine translation, MT). We present results of our research efforts towards rapid language portability of all these components. The results on translation suggest that MT systems can be successfully constructed for any language pair by cascading multiple MT systems via English. Moreover, end-to-end performance can be improved, if the interlingua language is enriched with additional linguistic information that can be derived automatically and monolingually in a data-driven fashion. Alex Waibel, Tanja Schultz, Stephan Vogel, Christian Fügen, Matthias Honal, Muntsin Kolss, Jürgen Reichert, Sebastian Stüker |
ICASSP (3) | 1 |
| 2004 | Integrating thumbnail features for speech recognition using conditional exponential modelsabstractWe describe a novel approach for modeling segmental information in speech recognition, through the use of thumbnail features. By taking into account dependencies at the segmental level, thumbnail features are more resistant to changes in speaking rates and other factors. While the traditional acoustic features are fixed for every utterance, one set of thumbnail features is computed for each hypothesis, which may violate the traditional scoring paradigm. To this end, we introduce a conditional exponential modeling framework. It allows better integration of various knowledge sources in a discriminative fashion. We present preliminary experiments on the Switchboard task. Hua Yu 0008, Alex Waibel |
ICASSP (1) | 2 |
| 2004 | Tight coupling of speech recognition and dialog management - dialog-context dependent grammar weighting for speech recognitionabstractIn this paper we present our current work on a tight coupling of a speech recognizer with a dialog manager and our results by restricting the search space of our grammar based speech recognizer through the information given by the dialog manager. As a result of the tight coupling the same lingustic knowledge sources can be used in both, speech recognizer and dialog manager. Furthermore, the flexible context-free grammar implementation of our speech decoder Ibis allows weighting of specific rules at run-time to restrict the search space of the recognizer for the next decoding step. These rules are given by the dialog manager depending on the current dialog context. With this approach we were able to reduce the word error rate of user responses to system questions by 3.3% relative for close talking and 16.0% relative, when using distant speech input. The sentence error rates were reduced by 2.2%, 9.2% respectively. Christian Fügen, Hartwig Holzapfel, Alex Waibel |
INTERSPEECH | 3 |
| 2004 | Adaptation for soft whisper recognition using a throat microphoneabstractThis paper describes various adaptation methods applied to recognizing soft whisper recorded with a throat microphone.Since the amount of adaptation data is small and the testing data is very different from the training data, a series of adaptation methods is necessary.The adaptation methods include: maximum likelihood linear regression, feature-space adaptation, and re-training with downsampling, sigmoidal low-pass filter, or linear multivariate regression.With these adaptation methods, the word error rate improves from 99.3% to 32.9%. Szu-Chen Stan Jou, Tanja Schultz, Alex Waibel |
INTERSPEECH | 3 |
| 2004 | Worldwide ongoing activities on multilingual speech to speech translationabstractThis paper presents an overview of worldwide going on activities on Speech-to-Speech Translation. After a short introduction of the field, including the major projects and milestones, activities and projects going on in Asia, Europe and US are presented and described. Gianni Lazzari, Alex Waibel, Chengqing Zong |
INTERSPEECH | 2 |
| 2004 | Speech translation: past, present and futureabstractA decade after its first beginnings, the grand challenge of building automatic systems that translate speech has grown into an active research area, a focus for speech and language researchers worldwide. Workshops, conferences, sessions, journal issues are devoted to it, and research is supported by governments in Asia, Europe and the US. The problem has attracted much attention, as practical needs in an increasingly globalized world converge with scientific advances that bring possible solutions within reach. While great progress has been made, the problem is certainly not solved and much remains to be done. At a midway point along the way, we present this paper as a review of the past, of the successes so far, and as an attempt to chart a course for the future. Alex Waibel |
INTERSPEECH | 1 |
| 2004 | Natural human-robot interaction using speech, head pose and gesturesabstractIn this paper we present our ongoing work in building technologies for natural multimodal human-robot interaction. We present our systems for spontaneous speech recognition, multimodal dialogue processing and visual perception of a user, which includes the recognition of pointing gestures as well as the recognition of a person's head orientation. Each of the components is described in the paper and experimental results are presented. In order to demonstrate and measure the usefulness of such technologies for human-robot interaction, all components have been integrated on a mobile robot platform and have been used for real-time human-robot interaction in a kitchen scenario. Rainer Stiefelhagen, Christian Fügen, Petra Gieselmann, Hartwig Holzapfel, Kai Nickel, Alex Waibel |
IROS | 6 |
| 2004 | Language Model Adaptation for Statistical Machine Translation Based on Information Retrieval
Matthias Eck 0001, Stephan Vogel, Alex Waibel |
LREC | 3 |
| 2004 | Interpreting BLEU/NIST Scores: How Much Improvement do We Need to Have a Better System?
Ying Zhang 0048, Stephan Vogel, Alex Waibel |
LREC | 3 |
| 2004 | Improving Named Entity Translation Combining Phonetic and Semantic Similarities
Fei Huang 0002, Stephan Vogel, Alex Waibel |
HLT-NAACL | 3 |
| 2004 | Speaker adaptation with all-pass transforms
John W. McDonough, Thomas Schaaf, Alex Waibel |
Speech Commun. | 3 |
| 2004 | Automatic detection and recognition of signs from natural scenesabstractIn this paper, we present an approach to automatic detection and recognition of signs from natural scenes, and its application to a sign translation task. The proposed approach embeds multiresolution and multiscale edge detection, adaptive searching, color analysis, and affine rectification in a hierarchical framework for sign detection, with different emphases at each phase to handle the text in different sizes, orientations, color distributions and backgrounds. We use affine rectification to recover deformation of the text regions caused by an inappropriate camera view angle. The procedure can significantly improve text detection rate and optical character recognition (OCR) accuracy. Instead of using binary information for OCR, we extract features from an intensity image directly. We propose a local intensity normalization method to effectively handle lighting variations, followed by a Gabor transform to obtain local features, and finally a linear discriminant analysis (LDA) method for feature selection. We have applied the approach in developing a Chinese sign translation system, which can automatically detect and recognize Chinese signs as input from a camera, and translate the recognized text into English. Xilin Chen 0001, Jie Yang 0001, Jing Zhang 0011, Alex Waibel |
IEEE Trans. Image Process. | 4 |
| 2003 | Effective Phrase Translation Extraction from Alignment ModelsabstractPhrase level translation models are effective in improving translation quality by addressing the problem of local re-ordering across language boundaries. Methods that attempt to fundamentally modify the traditional IBM translation model to incorporate phrases typically do so at a prohibitive computational cost. We present a technique that begins with improved IBM models to create phrase level knowledge sources that effectively represent local as well as global phrasal context. Our method is robust to noisy alignments at both the sentence and corpus level, delivering high quality phrase level translation pairs that contribute to significant improvements in translation quality (as measured by the BLEU metric) over word based lexica as well as a competing alignment based method. Ashish Venugopal, Stephan Vogel, Alex Waibel |
ACL | 3 |
| 2003 | Maximum mutual information speaker adapted training with semi-tied covariance matricesabstractWe present re-estimation formulae for semi-tied covariance (STC) transformation matrices based on a maximum mutual information (MMI) criterion. These re-estimation formulae are different from those that have appeared previously in the literature. Moreover, we present a positive definiteness criterion with which the regularization constant present in all NMI re-estimation formulae can be reliably set to provide both consistent improvements in the total mutual information of the training set, as well as fast convergence. We combine the STC re-estimation formulae with their like for speaker-independent means and variances, and update all parameters during NMI speaker adapted training (MMI-SAT). We present the results of two sets of speech recognition experiments conducted on the the 1998 Broadcast News evaluation set, as well as a corpus of meeting room data collected at the Interactive Systems Laboratories of the Carnegie Mellon University. John W. McDonough, Alex Waibel |
ICASSP (1) | 2 |
| 2003 | Multilingual articulatory featuresabstractSpeech recognition systems based on or aided by articulatory features, such as place and manner of articulation, have been shown to be useful under varying circumstances. Recognizers based on features better compensate channel and noise variability. We show that it is also possible to compensate for inter language variability using articulatory feature detectors. We come to the conclusion that articulatory features can be recognized across languages and that using detectors from many languages can improve the classification accuracy of the feature detectors on a single language. We further demonstrate how those multilingual and cross-lingual detectors can support an HMM based recognizer and thereby significantly reduce the word error rate by up to 12.3% relative. We expect that with the use of multilingual articulatory features it is possible to support the rapid deployment of recognition systems for new target languages. Sebastian Stüker, Tanja Schultz, Florian Metze, Alex Waibel |
ICASSP (1) | 4 |
| 2003 | SMaRT: the Smart Meeting Room Task at ISLabstractAs computational and communications systems become increasingly smaller, faster, more powerful, and more integrated, the goal of interactive, integrated meeting support rooms is slowly becoming reality. It is already possible, for instance, to rapidly locate task-related information during a meeting, filter it, and share it with remote users. Unfortunately, the technologies that provide such capabilities are as obstructive as they are useful - they force humans to focus on the tool rather than the task. Thus the veneer of utility often hides the true costs of use, which are longer, less focused human interactions. To address this issue, we present our current research efforts towards SMaRT: the Smart Meeting Room Task. The goal of SMaRT is to provide meeting support services that do not require explicit human-computer interaction. Instead, by monitoring the activities in the meeting room using both video and audio analysis, the room is able to react appropriately to users' needs and allow the users to focus on their own goals. Alex Waibel, Tanja Schultz, Michael Bett, Matthias Denecke, Robert G. Malkin, Ivica Rogina, Rainer Stiefelhagen, Jie Yang 0001 |
ICASSP (4) | 1 |
| 2003 | Comparison of acoustic model adaptation techniques on non-native speechabstractThe performance of speech recognition systems is consistently poor on non-native speech. The challenge for non-native speech recognition is to maximize the recognition performance with a small amount of available non-native data. We report on acoustic modeling adaptation for the recognition of non-native speech. Using non-native data from German speakers, we investigate how bilingual models, speaker adaptation, acoustic model interpolation and polyphone decision tree specialization methods can help to improve the recognizer performance. Results obtained from the experiments demonstrate the feasibility of these methods. Zhirong Wang, Tanja Schultz, Alex Waibel |
ICASSP (1) | 3 |
| 2003 | Calibration of a Hybrid Camera NetworkabstractVisual surveillance using a camera network has imposed new challenges to camera calibration. An essential problem is that a large number of cameras may not have a common field of view or even be synchronized well. We propose to use a hybrid camera network that consists of catadioptric and perspective cameras for a visual surveillance task. The relations between multiple views of a scene captured from different cameras can be then calibrated under the catadioptric camera's coordinate system. We address the important issue of how to calibrate the hybrid camera network. We calibrate the hybrid camera network in three steps. First, we calibrate the catadioptric camera using only the vanishing points. In order to reduce computational complexity, we calibrate the camera without the mirror first and then calibrate the catadioptric camera system. Second, we determine 3D positions of some points using as few as two spatial parallel lines and some equidistance points. Finally, we calibrate other perspective cameras based on these known spatial points. Xilin Chen 0001, Jie Yang 0001, Alex Waibel |
ICCV | 3 |
| 2003 | Integrating multilingual articulatory features into speech recognitionabstractThe use of articulatory features, such as place and manner of articulation, has been shown to reduce the word error rate of speech recognition systems under different conditions and in different settings.For example recognition systems based on features are more robust to noise and reverberation.In earlier work we showed that articulatory features can compensate for inter language variability and can be recognized across languages.In this paper we show that using cross-and multilingual detectors to support an HMM based speech recognition system significantly reduces the word error rate.By selecting and weighting the features in a discriminative way, we achieve an error rate reduction that lies in the same range as that seen when using language specific feature detectors.By combining feature detectors from many languages and training the weights discriminatively, we even outperform the case where only monolingual detectors are being used. Sebastian Stüker, Florian Metze, Tanja Schultz, Alex Waibel |
INTERSPEECH | 4 |
| 2003 | Speechalator: two-way speech-to-speech translation on a consumer PDAabstractThis paper describes a working two-way speech-to-speech translation system that runs in near real-time on a consumer handheld computer. It can translate from English to Arabic and Arabic to English in the domain of medical interviews. We describe the general architecture and frameworks within which we developed each of the components: HMM-based recognition, interlingua translation (both rule and statistically based), and unit selection synthesis. Alex Waibel, Ahmed Badran, Alan W. Black, Robert E. Frederking, Donna Gates, Alon Lavie, Lori S. Levin, Kevin A. Lenzo, Laura Mayfield Tomokiyo, Jürgen Reichert, Tanja Schultz, Dorcas Wallace, Monika Woszczyna, Jing Zhang 0011 |
INTERSPEECH | 1 |
| 2003 | Minimum variance distortionless response on a warped frequency scaleabstractIn this work we propose a time domain technique to estimate an all-pole model based on the minimum variance distortionless response (MVDR) using a warped short time frequency axis such as the Mel scale. The use of the MVDR eliminates the overemphasis of harmonic peaks typically seen in medium and high pitched voiced speech when spectral estimation is based on linear prediction (LP). Moreover, warping the frequency axis prior to MVDR spectral estimation ensures more parameters in the spectral model are allocated to the low, as opposed to high, frequency regions of the spectrum, thereby mimicking the human auditory system. In a series of speech recognition experiments on the Switchboard Corpus (spontaneous English telephone speech), the proposed approach achieved a word error rate (WER) of 32.1% for female speakers, which is clearly superior to the 33.2% WER obtained by the usual combination of Mel warping and linear prediction. Matthias Wölfel, John W. McDonough, Alex Waibel |
INTERSPEECH | 3 |
| 2003 | The CMU statistical machine translation systemabstractIn this paper we describe the components of our statistical machine translation system. This system combines phrase-to-phrase translations extracted from a bilingual corpus using different alignment approaches. Special methods to extract and align named entities are used. We show how a manual lexicon can be incorporated into the statistical system in an optimized way. Experiments on Chinese-to-English and Arabic-to-English translation tasks are presented. Stephan Vogel, Ying Zhang 0048, Fei Huang 0002, Alicia Tribble, Ashish Venugopal, Bing Zhao 0005, Alex Waibel |
MTSummit | 7 |
| 2003 | Speechalator: Two-Way Speech-to-Speech Translation in Your Hand
Alex Waibel, Ahmed Badran, Alan W. Black, Robert E. Frederking, Donna Gates, Alon Lavie, Lori S. Levin, Kevin A. Lenzo, Laura Mayfield Tomokiyo, Jürgen Reichert, Tanja Schultz, Dorcas Wallace, Monika Woszczyna, Jing Zhang 0011 |
HLT-NAACL | 1 |
| 2003 | Extracting named entity translingual equivalence with limited resourcesabstractIn this article we present an automatic approach to extracting Hindi-English (H-E) Named Entity (NE) translingual equivalences from bilingual parallel corpora. In the absence of a Hindi NE tagger or H-E translation dictionary, this approach adapts a Chinese-English (C-E) surface string transliteration model for H-E NE extraction. The model is initially trained using automatically extracted C-E NE pairs, then iteratively updated based on newly extracted H-E NE pairs. For each English person and location NE in each sentence pair, this approach searches for its Hindi correspondence with minimum transliteration cost and constructs an H-E NE list from the bilingual corpus. Experiments show that this approach extracted 1000 H-E NE pairs with a precision of 91.8%. Fei Huang 0002, Stephan Vogel, Alex Waibel |
ACM Trans. Asian Lang. Inf. Process. | 3 |
| 2002 | Automatic speech summarization applied to English broadcast news speechabstractThis paper reports an automatic speech summarization method and experimental results using English broadcast news speech. In our proposed method, a set of words maximizing a summarization score indicating an appropriateness of summarization is extracted from automatically transcribed speech. This extraction is performed using a Dynamic Programming (DP) technique according to a target compression ratio. We have previously tested the performance of our method using Japanese broadcast news speech. Since our method is based on a statistical approach, it could be applied to any language. In this paper, English broadcast news speech transcribed using a speech recognizer is automatically summarized. In order to apply our method to English, the model of estimating word concatenation probabilities based on a dependency structure in the original speech given by a Stochastic Dependency Context Free Grammar (SDCFG) is modified. A summarization method for multiple utterances using two-level DP technique is also proposed. Chiori Hori, Sadaoki Furui, Robert G. Malkin, Hua Yu 0008, Alex Waibel |
ICASSP | 5 |
| 2002 | Speaker identification using multilingual phone stringsabstractFar-field speaker identification is very challenging since varying recording conditions often result in un-matching training and testing situations. Although the widely used Gaussian Mixture Models (GMM) approach achieves reasonable good results when training and testing conditions match, its performance degrades dramatically under un-matching conditions. In this paper we propose a new approach for far-field speaker identification: the usage of multilingual phone strings derived from phone recognizers in eight different languages. The experiments are carried out on a database of 30 speakers recorded with eight different microphone distances. The results show that the multi-lingual phone string approach is robust against un-matching conditions and significantly outperforms the GMMs. On 10-second test chunks, the average closed-set identification performance achieves 96.7% on variable distance data. Qin Jin, Tanja Schultz, Alex Waibel |
ICASSP | 3 |
| 2002 | On maximum mutual information speaker-adapted trainingabstractIn this work, we combine maximum mutual information-based parameter estimation with speaker-adapted training (SAT). As will be shown, this can be achieved by performing unsupervised parameter estimation on the test data, a distinct advantage for many recognition tasks involving conversational speech. We also propose an approximation to the maximum likelihood and maximum mutual information SAT re-estimation formulae that greatly reduces the amount of disk space required to conduct training on corpora such as Broadcast News, which contains speech from thousands of speakers. We present the results of a set of speech recognition experiments on three test sets: the English Spontaneous Scheduling Task corpus, Broadcast News, and a new corpus of Meeting Room data collected at the Interactive Systems Laboratories of the Carnegie Mellon University. John W. McDonough, Thomas Schaaf, Alex Waibel |
ICASSP | 3 |
| 2002 | Experiments on distant-talking speech recognition in meeting room using extended MAMabstractMAM has been successfully used to improve the noise robustness of speech recognizers. We apply this method in the exploration of distant-talking speech recognition for meeting room system, and propose a new extended MAM algorithm combined with LDA and MLLR techniques to compensate for additive noise and channel variation. By sharing the LDA matrix with the core recognition system, the MAM secondary model is coupled better with the intended system and can take the advantage of context speech frames. Using MAM in conjunction with model adaptation method MLLR can result in further improved recognition accuracy. Our database includes simulated 10dB additive noisy speech and real distant-talking speech by 8-channel simultaneous recording. We started from a baseline LVCSR meeting transcription system, with word error rate range from 25 to 55%. By applying extended MAM, the improved system achieves an average relative error rate reduction of 27% for 10dB additive noisy speech and 15% for real distant-talking speech. Alex Waibel |
ICASSP | 2 |
| 2002 | Efficient language model lookahead through polymorphic linguistic context assignmentabstractIn this study, we examine how fast decoding of conversational speech with large vocabularies profits from efficient use of linguistic information, i.e. language models and grammars. Based on a re-entrant single pronunciation prefix tree, we use the concept of linguistic context polymorphism to achieve an early incorporation of language model information. This approach allows us to use all available language model information in a one-pass decoder, using the same engine to decode with statistical n-gram language models as well as context free grammars or re-scoring of lattices in an efficient way. We compare this approach to our previous decoder, which needed three passes to incorporate all available information. The results on a very large vocabulary task show that the search can be speeded up by almost a factor of three, without introducing additional search errors. On all examined tasks, we observed significant improvements by using an exact language model lookahead over usual bigram lookahead strategies, even for very hard tasks with unmatched conditions, without introducing extra memory overhead. Hagen Soltau, Florian Metze, Christian Fügen, Alex Waibel |
ICASSP | 4 |
| 2002 | Automatic detection and translation of text from natural scenesabstractLarge amounts of information are embedded in natural scenes. Signs are good examples of natural objects with high information content. In this paper, we discuss problems in automatic detection and translation of text from natural scenes. We describe the chal1enges of automatic text detection and propose methods to address these chal1enges. We extend example based machine translation technology for sign translation and present a prototype system for Chinese sign translation. This system is capable of capturing images, automatically detecting and recognizing text, and translating the text into English. The translation can be displayed on a palm size PDA, or synthesized as a voice output message over the earphones. Jie Yang 0001, Xilin Chen 0001, Jing Zhang 0011, Ying Zhang 0048, Alex Waibel |
ICASSP | 5 |
| 2002 | Integrating Emotional Cues into a Framework for Dialogue ManagementabstractEmotions are very important in human-human communication but are usually ignored in human-computer interaction. Recent work focuses on recognition and generation of emotions as well as emotion driven behavior. Our work focuses on the use of emotions in dialogue systems that can be used with speech input or as well in multi-modal environments. We describe a framework for using emotional cues in a dialogue system and their informational characterization. We describe emotion models that can be integrated into the dialogue system and can be used in different domains and tasks. Our application of the dialogue system is planned to model multi-modal human-computer-interaction with a humanoid robotic system. Hartwig Holzapfel, Christian Fügen, Matthias Denecke, Alex Waibel |
ICMI | 4 |
| 2002 | Flexi-Modal and Multi-Machine User InterfacesabstractWe describe our system which facilitates collaboration using multiple modalities, including speech, handwriting, gestures, gaze tracking, direct manipulation, large projected touch-sensitive displays, laser pointer tracking, regular monitors with a mouse and keyboard, and wireless networked handhelds. Our system allows multiple, geographically dispersed participants to simultaneously and flexibly mix different modalities using the right interface at the right time on one or more machines. We discuss each of the modalities provided, how they were integrated in the system architecture, and how the user interface enabled one or more people to flexibly use one or more devices. Brad A. Myers, Robert G. Malkin, Michael Bett, Alex Waibel, Ben Bostwick, Rob Miller 0001, Jie Yang 0001, Matthias Denecke, Edgar Seemann, Choon Hong Peck, Dave Kong, Jeffrey Nichols 0001, William L. Scherlis |
ICMI | 4 |
| 2002 | Towards Universal Speech RecognitionabstractThe increasing interest in multilingual applications like speech-to-speech translation systems is accompanied by the need for speech recognition front-ends in many languages that can also handle multiple input languages at the same time. We describe a universal speech recognition system that fulfills such needs. It is trained by sharing speech and text data across languages and thus reduces the number of parameters and overhead significantly at the cost of only slight accuracy loss. The final recognizer eases the burden of maintaining several monolingual engines, makes dedicated language identification obsolete and allows for code-switching within an utterance. To achieve these goals we developed new methods for constructing multilingual acoustic models and multilingual n-gram language models. Zhirong Wang, Umut Topkara, Tanja Schultz, Alex Waibel |
ICMI | 4 |
| 2002 | A PDA-Based Sign TranslatorabstractWe propose an effective approach for a PDA-based sign system and present the sign translator. Its main functions include three parts: detection, recognition and translation. Automatic detection and recognition of text in natural scenes is a prerequisite for the automatic sign translator. In order to make the system robust for text detection in various natural scenes, the detection approach efficiently embeds multi-resolution, adaptive search in a hierarchical framework with different emphases at each layer. We also introduce an intensity-based OCR method to recognize characters in various fonts and lighting conditions, where we employ the Gabor transform to obtain local features, and LDA for selection and classification of features. The recognition rate is 92.4% for the testing set obtained from the natural sign. A sign is different from the normal used sentence. It is brief with a lot of abbreviations and place nouns. We only briefly introduce a rule-based place name translation. We have integrated all these functions in a PDA, which can capture sign images, auto segment and recognize the Chinese sign, and translate it into English. Jing Zhang 0011, Xilin Chen 0001, Jie Yang 0001, Alex Waibel |
ICMI | 4 |
| 2002 | Phonetic speaker identification
Qin Jin, Tanja Schultz, Alex Waibel |
INTERSPEECH | 3 |
| 2002 | Interlingua based statistical machine translation
Manuel Kauers, Stephan Vogel, Christian Fügen, Alex Waibel |
INTERSPEECH | 4 |
| 2002 | A flexible stream architecture for ASR using articulatory features
Florian Metze, Alex Waibel |
INTERSPEECH | 2 |
| 2002 | Compensating for hyperarticulation by modeling articulatory properties
Hagen Soltau, Florian Metze, Alex Waibel |
INTERSPEECH | 3 |
| 2002 | Automatic sign translationabstractLarge amounts of information is embedded in the natural scenes. Signs are good examples of objects in natural environments which have rich information content. In this paper, we present our efforts in the automatic sign translation. We describe the challenges in the automatic sign translation and introduce the architecture of our current system for automatic detection and translation of Chinese signs. Two data-driven machine translation methods: Example Based Machine Translation (EBMT) and Statistical Machine Translation (SMT) are compared for the task of translating Chinese signs into English. We report the experimental results of both methods that are trained from a small bilingual sign corpus combined with a bilingual glossary. The experiment results indicate that EBMT generates more correct translations while SMT is better at inferring unseen patterns. We are currently working on developing a multi-engine machine translation system that can incrementally learn from the data and combine the results from EBMT and SMT. Ying Zhang 0048, Bing Zhao 0005, Jie Yang 0001, Alex Waibel |
INTERSPEECH | 4 |
| 2002 | Automatic Detection of Signs with Affine TransformationabstractIn this paper, we propose an approach for detecting signs from natural scenes. The approach efficiently embeds multiresolution, adaptive search, and affine rectification algorithms in a hierarchical framework, with different emphases at each layer. We combine in multi-resolution and multi-scale edge detection techniques to effectively detect text in different sizes. By using the cites from text inside the image, we introduce affine rectification transformation to recover deformation of the text region caused by air inappropriate camera view angle. This procedure can significantly improve text detection rate and OCR (Optical Character Recognition) accuracy. Experimental results have demonstrated feasibility of the proposed algorithms. We have applied the proposed approach to a Chinese sign translation system, which can automatically detect Chinese text input from a camera, recognize the text, and translate the recognized text into English or voice stream. Xilin Chen 0001, Jie Yang 0001, Jing Zhang 0011, Alex Waibel |
WACV | 4 |
| 2002 | Modeling focus of attention for meeting indexing based on multiple cuesabstractA user's focus of attention plays an important role in human-computer interaction applications, such as a ubiquitous computing environment and intelligent space, where the user's goal and intent have to be continuously monitored. We are interested in modeling people's focus of attention in a meeting situation. We propose to model participants' focus of attention from multiple cues. We have developed a system to estimate participants' focus of attention from gaze directions and sound sources. We employ an omnidirectional camera to simultaneously track participants' faces around a meeting table and use neural networks to estimate their head poses. In addition, we use microphones to detect who is speaking. The system predicts participants' focus of attention from acoustic and visual information separately. The system then combines the output of the audio- and video-based focus of attention predictors. We have evaluated the system using the data from three recorded meetings. The acoustic information has provided 8% relative error reduction on average compared to only using one modality. The focus of attention model can be used as an index for a multimedia meeting record. It can also be used for analyzing a meeting. Rainer Stiefelhagen, Jie Yang 0001, Alex Waibel |
IEEE Trans. Neural Networks | 3 |
| 2001 | Speaker compensation with sine-log all-pass transformsabstractIn previous work, we proposed the rational all-pass transform (RAPT) as the basis of a speaker adaptation scheme intended for use with a large vocabulary speech recognition system. It was shown that RAPT-based adaptation reduces to a linear transformation of cepstral means, much like the better known maximum likelihood linear regression (MLLR). In a set of speech recognition experiments conducted on the Switchboard Corpus, we obtained a word error rate (WER) of 37.9% using RAPT adaptation, a significant improvement over the 39.5% WER achieved with MLLR. In the present work, we propose the sine-log all-pass transform (SLAPT) as a replacement for the RAPT. Our findings indicate the SLAPT is just as effective as the RAPT at reducing WER when used as the basis for a variety of speaker compensation schemes, but in addition conduces to far more tractable computation of transformed cepstral sequences, and the estimation of optimal transform parameters. John W. McDonough, Florian Metze, Hagen Soltau, Alex Waibel |
ICASSP | 4 |
| 2001 | The ISL evaluation system for Verbmobil-IIabstractDescribes the 2000 ISL large vocabulary speech recognition system for fast decoding of conversational speech which was used in the German Verbmobil-II project. The challenge of this task is to build robust acoustic models to handle different dialects, spontaneous effects, and crosstalk as occur in conversational speech. We present speaker incremental normalization and adaptation experiments close to real-time constraints. To reduce the number of consequential errors caused by out-of-vocabulary words, we conducted filler-model experiments to handle unknown proper names. The overall improvements from 1998 to 2000 resulted in a word error reduction from 40% to 17% on our development test set. Hagen Soltau, Thomas Schaaf, Florian Metze, Alex Waibel |
ICASSP | 4 |
| 2001 | Advances in automatic meeting record creation and accessabstractOral communication is transient, but many important decisions, social contracts and fact findings are first carried out in an oral setup, documented in written form and later retrieved. At Carnegie Mellon University's Interactive Systems Laboratories we have been experimenting with the documentation of meetings. The paper summarizes part of the progress that we have made in this test bed, specifically on the question of automatic transcription using large vocabulary continuous speech recognition, information access using non-keyword based methods, summarization and user interfaces. The system is capable of automatically constructing a searchable and browsable audio-visual database of meetings and provide access to these records. Alex Waibel, Michael Bett, Florian Metze, Klaus Ries 0001, Thomas Schaaf, Tanja Schultz, Hagen Soltau, Hua Yu 0008, Klaus Zechner |
ICASSP | 1 |
| 2001 | Model-combination-based acoustic mappingabstractWe propose a method for compensating distortions in the speech signal caused by environment changes. The basic method concentrates on additive noise, but can be extended to address also channel and to some extent speaker changes. By combining compensation with adaptation techniques it leads to high error rate reductions for mobile speech applications. Thereby, it is more efficient than adapting the acoustic model of the recognizer and more powerful than simple noise reduction techniques. Martin Westphal, Alex Waibel |
ICASSP | 2 |
| 2001 | Experiments on cross-language acoustic modelingabstractWith the distribution of speech products all over the world, the portability to new target languages becomes a practical concern. As a consequence our research focuses on rapid transfer of LVCSR systems to other languages. In former studies we evaluated the performance if limited adaptation data is available. Particularly for very time constrained tasks and minority languages, it is even reasonable that no training data is available at all. In this paper we examine what performance can be expected in this scenario. All experiments are run in the framework of the GlobalPhone project which investigates LVCSR systems in 15 languages. Tanja Schultz, Alex Waibel |
INTERSPEECH | 2 |
| 2001 | Online handwriting recognition: the NPen++ recognizer
Stefan Jäger 0002, Stefan Manke, Jürgen Reichert, Alex Waibel |
Int. J. Document Anal. Recognit. | 4 |
| 2001 | Language-independent and language-adaptive acoustic modeling for speech recognition
Tanja Schultz, Alex Waibel |
Speech Commun. | 2 |
| 2001 | Multimodal error correction for speech user interfacesabstractAlthough commercial dictation systems and speech-enabled telephone voice user interfaces have become readily available, speech recognition errors remain a serious problem in the design and implementation of speech user interfaces. Previous work hypothesized that switching modality could speed up interactive correction of recognition errors. This article presents multimodal error correction methods that allow the user to correct recognition errors efficiently without keyboard input. Correction accuracy is maximized by novel recognition algorithms that use context information for recognizing correction input. Multimodal error correction is evaluated in the context of a prototype multimodal dictation system. The study shows that unimodal repair is less accurate than multimodal error correction. On a dictation task, multimodal correction is faster than unimodal correction by respeaking. The study also provides empirical evidence that system-initiated error correction (based on confidence measures) may not expedite error correction. Furthermore, the study suggests that recognition accuracy determines user choice between modalities: while users initially prefer speech, they learn to avoid ineffective correction modalities with experience. To extrapolate results from this user study, the article introduces a performance model of (recognition-based) multimodal interaction that predicts input speed including time needed for error correction. Applied to interactive error correction, the model predicts the impact of improvements in recognition technology on correction speeds, and the influence of recognition accuracy and correction method on the productivity of dictation systems. This model is a first step toward formalizing multimodal interaction. Bernhard Suhm, Brad A. Myers, Alex Waibel |
ACM Trans. Comput. Hum. Interact. | 3 |
| 2000 | DIASUMM: Flexible Summarization of Spontaneous Dialogues in Unrestricted Domains
Klaus Zechner, Alex Waibel |
COLING | 2 |
| 2000 | Face Recognition in a Meeting RoomabstractWe investigate the recognition of human faces in a meeting room. The major challenges of identifying human faces in this environment include low quality of input images, poor illumination, unrestricted head poses and continuously changing facial expressions and occlusion. In order to address these problems we propose a novel algorithm, dynamic space warping (DSW). The basic idea of the algorithm is to combine local features under certain spatial constraints. We compare DSW with the eigenface approach on data collected from various meetings. We have tested both front and profile face images and images with two stages of occlusion. The experimental results indicate that the DSW approach outperforms the eigenface approach in both cases. Ralph Gross, Jie Yang 0001, Alex Waibel |
FG | 3 |
| 2000 | Segmenting Hands of Arbitrary ColorabstractHand segmentation is a prerequisite for many gesture recognition tasks. Color has been widely used for hand segmentation. However, many approaches rely on predefined skin color models. It is very difficult to predefine a color model in a mobile application where the light condition may change dramatically over time. We propose a novel statistical approach to hand segmentation based on Bayes decision theory. The proposed method requires no predefined skin color model. Instead it generates a hand color model and a background color model for a given image, and uses these models to classify each pixel in the image as either a hand pixel or a background pixel. Models are generated using a Gaussian mixture model with the restricted EM algorithm. Our method is capable of segmenting hands of arbitrary color in a complex scene. It performs well even when there is a significant overlap between hand and background colors, or when the user wears gloves. We show that the Bayes decision method is superior to a commonly used method by comparing their upper bound performance. Experimental results demonstrate the feasibility of the proposed method. Xiaojin Zhu 0001, Jie Yang 0001, Alex Waibel |
FG | 3 |
| 2000 | Strategies for automatic segmentation of audio dataabstractIn many applications, like indexing of broadcast news or surveillance applications, the input data consists of a continuous, unsegmented audio stream. Speech recognition technology, however, usually requires segments of relatively short length as input. For such applications, effective methods to segment continuous audio streams into homogeneous segments are required. In this paper, three different segmenting strategies (model-based, metric-based and energy-based) are compared on the same broadcast news test data. It is shown that model-based and metric-based techniques outperform the simpler energy-based algorithms. While model based segmenters achieve very high level of segment boundary precision, the metric-based segmenter preforms better in terms of segment boundary recall (RCL). To combine the advantages of both strategies, a new hybrid algorithm is introduced. For this, the results of a preliminary metric-based segmentation are used to construct the models for the final model-based segmenter run. The new hybrid approach is shown to outperform the other segmenting strategies. Thomas Kemp, Martin Westphal, Alex Waibel |
ICASSP | 4 |
| 2000 | Polyphone decision tree specialization for language adaptationabstractWith the distribution of speech technology products all over the world, the fast and efficient portability to new target languages becomes a practical concern. The authors explore the relative effectiveness of adapting multilingual LVCSR systems to a new target language with limited adaptation data. For this purpose they introduce a polyphone decision tree specialization method. Several recognition results are presented based on mono- and multilingual recognizers. These recognizers are developed in the framework of the project GlobalPhone. In this project we investigate speech recognition in 15 languages: Arabic, Mandarin and Shanghai Chinese, Croatian, English, French, German, Japanese, Korean, Portuguese, Russian, Spanish, Swedish, Tamil, and Turkish. Tanja Schultz, Alex Waibel |
ICASSP | 2 |
| 2000 | Specialized acoustic models for hyperarticulated speechabstractThis study aims to improve the performance of automatic speech recognizers at hyperarticulated speech. Hyperarticulation often occur as a strategy to recover previous recognition errors in spoken dialogue systems. Contrary to this intention a significant performance degradation can be observed at hyperarticulation. In this paper we present an analysis of features that caused the performance loss. The average phone duration is nearby 20% longer. Pitch contour and fundamental frequency change significantly at hyperarticulation. We report on adapting acoustic and transition models to hyperarticulated speech. We achieved a word error reduction about 23% at hyperarticulation. Hagen Soltau, Alex Waibel |
ICASSP | 2 |
| 2000 | Growing Gaussian Mixture Models for Pose Invariant Face RecognitionabstractA major challenge for face recognition algorithms lies in the variance faces undergo while changing pose. This problem is typically addressed by building view dependent models based on face images taken from predefined head poses. However, it is impossible to determine all head poses beforehand in an unrestricted setting such as a meeting room, where people can move and interact freely. We present an approach to pose invariant face recognition. We employ Gaussian mixture models to characterize human faces and model pose variance with different numbers of mixture components. The optimal number of mixture components for each person is automatically learned from training data by growing the mixture models. The proposed algorithm is tested on real data recorded in a meeting room. The experimental results indicate that the new method outperforms standard eigenface and Gaussian mixture model approaches. Our algorithm achieved as much as 42% error reduction compared to the standard eigenface approach on the same test data. Ralph Gross, Jie Yang 0001, Alex Waibel |
ICPR | 3 |
| 2000 | Simultaneous Tracking of Head Poses in a Panoramic ViewabstractIn this paper we present an approach to simultaneously estimate gaze directions of multiple people in the view of a panoramic camera. Human faces are located and tracked using a probabilistic skin-color model and motion detection. Neural networks are used to estimate head poses of the detected faces. With this approach, it is possible to simultaneously track the locations of multiple people around a meeting table and estimate their gaze directions using only a panoramic camera. We have achieved an accuracy of 9 degrees for head pan estimation and 6 degrees for tilt estimation for a multi-user system. Rainer Stiefelhagen, Jie Yang 0001, Alex Waibel |
ICPR | 3 |
| 2000 | Dialogue management for multimodal user registrationabstract... information with a user. It is widely used in hospitals, hotels and conferences. In this paper, we propose an approach to interactive user registration by combining face recognition, speech recognition and speech synthesis technologies together through an efficient dialogue manager. In order to minimize a user's effort, we employ a new dialogue management model based on a finite state automaton (FSA), which uses a Baysian network to fuse the user's information from multiple channels (e.g., face image, speech, records stored in a pre-constructed database) to reliably estimate the confidence about user identity. Instead of fixing weights, the FSA adjusts its weights dynamically by integrating partial information from multiple information sources. This is achieved by maximizing an objective function to determine an optimal action at each succeeding state according to current confidence and information cues. Thus the transition between states can be done along the shortest path from the initial state to the goal state. We have developed a multimodal user registration system to demonstrate the feasibility of the proposed approach. Fei Huang 0002, Jie Yang 0001, Alex Waibel |
INTERSPEECH | 3 |
| 2000 | Application of LDA to speaker recognitionabstractThe speaker recognition task falls under the general problem of pattern classification. Speaker recognition as a pattern classification problem, its ultimate objective is design of a system that classifies the vector of features in different classes by partitioning the feature space into optimal speaker discriminative space. Linear Discriminant Analysis (LDA) is a feature extraction method that provides a linear transformation of n-dimensional feature vectors (or samples) into m- dimensional space (m < n), so that samples belonging to the same class are close together but samples from different classes are far apart from each other. In this paper we discuss the issue of the application of LDA to our Gaussian Mixture Model (GMM) based speaker identification task. Applying LDA improved the identification performance. Keywords: Speaker recognition, Linear Discriminant Analysis, Gaussian Mixture Model 1. INTRODUCTION Speaker recognition is the task of automatically recognizing who i... Qin Jin, Alex Waibel |
INTERSPEECH | 2 |
| 2000 | A na ve de-lambing method for speaker identificationabstractThis paper addresses the issue of close-set text-independent speaker identification from speech samples recorded over telephone. We have known that the speaker identification performance variability can be attributed to many factors. One major factor is the inherent differences in the recognizability of different speakers. In speaker recognition systems such differences are characterized by the use of animal names for different types of speakers. In this paper we use lambs to refer to those speakers who are particularly easy to imitate in our closeset text-independent speaker identification system. That is, other speakers are much more likely to be recognized as these lamb speakers when they cannot be correctly recognized. Lambs adversely affect our close-set text-independent speaker identification performance a lot. In this paper we describe a naive de-lambing method to deal with these lamb speakers so as to improve our system performance. The speech data of our close-set speaker ide... Qin Jin, Alex Waibel |
INTERSPEECH | 2 |
| 2000 | The effects of room acoustics on MFCC speech parameterabstractAutomatic speech recognition systems attain high performance for close-talking applications, but they deteriorate significantly in distant-talking environment. The reason is the mismatch between training and testing conditions. We have carried out a research work for a better understanding of the effects of room acoustics on speech feature by comparing simultaneous recordings of close talking and distant talking speech utterances. The characteristics of two degrading sources, background noise and room reverberation are discussed. Their impacts on the spectrum are different. The noise affects on the valley of the spectrum while the reverberation causes the distortion at the peaks at the pitch frequency and its multiples. In the situation of very few training data, we attempt to choose the efficient compensation approaches in the spectrum, spectrum subband or cepstrum domain. Vector Quantization based model is used to study the influence of the variation on feature vector distribution. ... Alex Waibel |
INTERSPEECH | 2 |
| 2000 | Phone dependent modeling of hyperarticulated effects#
Hagen Soltau, Alex Waibel |
INTERSPEECH | 2 |
| 2000 | New developments in automatic meeting transcriptionabstractIn this paper we report on new developments in the automatic meeting transcription task. Unlike other types of speech (such as those found in Broadcast News and Switchboard), meetings are unique in their richer dynamics of human-to-human interaction. An intuitive "fingernail" plot is proposed to visualize such turntaking behavior. We will also show how recognition of short turns can be improved by building a language model tailored specifically for short turns. Out-Of-Vocabulary (OOV) words become a more salient problem in the meeting transcription task, as they are mostly topic words and proper names, lack of which not only causes Word Error Rate (WER) increase, but also limits further use of recognition hypotheses. We describe a prototype system which uses the Web as a source for vocabulary expansion, and present preliminary OOV retrieval results. 1. INTRODUCTION As speech recognition research progresses from read speech(Wall Street Journal), to prepared speech (a major part of Bro... Hua Yu 0008, Takashi Tomokiyo, Zhirong Wang, Alex Waibel |
INTERSPEECH | 4 |
| 2000 | Streamlining the front end of a speech recognizerabstractIn this paper we seek to streamline various operations within the front end of a speech recognizer, both to reduce unnecessary computation and to simplify the conceptual framework. First, a novel view of the front end in terms of linear transformations is presented. Then we study the invariance property of recognition performance with respect to linear transformations (LT) at the front end. Analysis reveals that several LT steps can be consolidated into a single LT, which effectively eliminates the Discrete Cosine Transform (DCT) step, part of the traditional MFCC (Mel-Frequency Cepstral Coefficient) front end. Moreover, a highly simplified, data-driven front-end scheme is proposed as a direct generalization of this idea. The new setup has no Mel-scale filtering, another part of the MFCC front end. Experimental results show a 5% relative improvement on the Broadcast News task. 1. LINEAR TRANSFORMATIONS IN THE TRADITIONAL FRONT END The front end is a relatively independent component ... Hua Yu 0008, Alex Waibel |
INTERSPEECH | 2 |
| 2000 | Shallow Discourse Genre Annotation in CallHome Spanish
Klaus Ries 0001, Lori S. Levin, Liza Valle, Alon Lavie, Alex Waibel |
LREC | 5 |
| 2000 | Towards Unrestricted Lip ReadingabstractLip reading provides useful information in speech perception and language understanding, especially when the auditory speech is degraded. However, many current automatic lip reading systems impose some restrictions on users. In this paper, we present our research efforts in the Interactive System Laboratory, towards unrestricted lip reading. We first introduce a top–down approach to automatically track and extract lip regions. This technique makes it possible to acquire visual information in real-time without limiting the user's freedom of movement. We then discuss normalization algorithms to preprocess images for different lightning conditions (global illumination and side illumination). We also compare different visual preprocessing methods such as raw image, Linear Discriminant Analysis (LDA), and Principle Component Analysis (PCA). We demonstrate the feasibility of the proposed methods by the development of a modular system for flexible human–computer interaction via both visual and acoustic speech. The system is based on an extension of the existing state-of-the-art speech recognition system, a modular Multiple State–Time Delayed Neural Network (MS–TDNN) system. We have developed adaptive combination methods at several different levels of the recognition network. The system can automatically track a speaker and extract his/her lip region in real-time. The system has been evaluated under different noisy conditions such as white noise, music, and mechanical noise. The experimental results indicate that the system can achieve up to 55% error reduction using visual information in addition to the acoustic signal. Uwe Meier, Rainer Stiefelhagen, Jie Yang 0001, Alex Waibel |
Int. J. Pattern Recognit. Artif. Intell. | 4 |
| 2000 | The Janus-III Translation System: Speech-to-Speech Translation in Multiple Domains
Lori S. Levin, Alon Lavie, Monika Woszczyna, Donna Gates, Marsal Gavaldà, Detlef Koll, Alex Waibel |
Mach. Transl. | 7 |
| 2000 | Multilinguality in speech and spoken language systemsabstractBuilding modern speech and language systems currently requires large data resources such as texts, voice recordings, pronunciation lexicons, morphological decomposition information and parsing grammars. Based on a study of the most important differences between language groups, we introduce approaches to efficiently deal with the enormous task of covering even a small percentage of the world's languages. For speech recognition, we have reduced the resource requirements by applying acoustic model combination, bootstrapping and adaption techniques. Similar algorithms have been applied to improve the recognition of foreign accents. Segmenting language into appropriate units reduces the amount of data required to robustly estimate statistical models. The underlying morphological principles are also used to automatically adapt the coverage of our speech recognition dictionaries with the Hypothesis-Driven Lexical Adaptation (HDLA) algorithm. This reduces the out-of-vocabulary problems encountered in agglutinative languages. Speech recognition results are reported for the read GlobalPhone database and some broadcast news data. For speech translation, using a task-oriented Interlingua allows to build a system with N languages with linear, rather than quadratic effort. We have introduced a modular grammar design to maximize reusability and portability. End-to-end translation results are reported on a travel-domain task in the framework of C-STAR. Alex Waibel, Petra Geutner, Laura Mayfield Tomokiyo, Tanja Schultz, Monika Woszczyna |
Proc. IEEE | 1 |
| 1999 | Model-Based and Empirical Evaluation of Multimodal Interactive Error CorrectionabstractOur research addresses the problem of error correction in speech user interfaces. Previous work hypothesized that switching modality could speed up interactive correction of recognition errors (so-called multimodal error correction). We present a user study that compares, on a dictation task, multimodal error correction with conventional interactive correction, such as speaking again, choosing Tom a list, and keyboard input. Results show that multimodal correction is faster than conventional correction without keyboard input, but slower than correction by typing for users with good typing skills. Furthermore, while users initially prefer speech, they learn to avoid ineffective correction modalities with experience. To extrapolate results from this user study we developed a performance model of multimodal interaction that predicts input speed including time needed for error correction. We apply the model to estimate the impact of recognition technology improvements on correction speeds and the influence of recognition accuracy and correction method on the productivity of dictation systems. Our model is a first step towards formalizing multimodal (recognition-based) interaction. Bernhard Suhm, Brad A. Myers, Alex Waibel |
CHI | 3 |
| 1999 | Selection criteria for hypothesis driven lexical adaptationabstractAdapting the vocabulary of a speech recognizer to the utterance to be recognized has proven to be successful both in reducing high out-of-vocabulary as well as word error rates. This applies especially to languages that have a rapid vocabulary growth due to a large number of inflections and composita. This paper presents various adaptation methods within the hypothesis driven lexical adaptation (HDLA) framework which allow speech recognition on a virtually unlimited vocabulary. Selection criteria for the adaptation process are either based on morphological knowledge or distance measures at phoneme or grapheme level. Different methods are introduced for determining distances between phoneme pairs and for creating the large fallback lexicon the adapted vocabulary is chosen from. HDLA reduces the out-of-vocabulary-rate by 55% for Serbo-Croatian, 35% for German and 27% for Turkish. The reduced out-of-vocabulary rate also decreases the word error rate by an absolute 4.1% to 25.4% on Serbo-Croatian broadcast news data. Petra Geutner, Michael Finke, Alex Waibel |
ICASSP | 3 |
| 1999 | Modeling and efficient decoding of large vocabulary conversational speechabstractCapturing the large variability of conversational speech in the framework of purely phone based speech recognizers is virtually impossible. It has been shown earlier that suprasegmental features such as speaking rate, duration and syllabic, syntactic and semantic structure are important predictors of pronunciation variation. In order to allow for a tighter coupling of these predictors of pronunciation, duration and acoustic modeling a new recognition toolkit has been developed. The phonetic transcription of speech has been generalized to an attribute based representation, thus enabling the integration of suprasegmental, non-phonetic features. A pronunciation model is trained to augment the attribute transcription to mark possible pronunciation effects which are then taken into account by the acoustic model induction algorithm. A finite state machine single-prefix-tree, one-pass, time-synchronous decoder is presented that efficiently decodes highly spontaneous speech within this new representational framework. Michael Finke, Jürgen Fritsch, Detlef Koll, Alex Waibel |
EUROSPEECH | 4 |
| 1999 | Navigating German cities by spontaneous French queries
Harouna Kabré, Alex Waibel |
EUROSPEECH | 2 |
| 1999 | Unsupervised training of a speech recognizer: recent experimentsabstractCurrent speech recognition systems require large amounts of transcribed data for parameter estimation. The transcription, however, is tedious and expensive. In this work we describe our experiments which are aimed at training a speech recognizer with only a minimal amount (30 minutes) of transcriptions and a large portion (50 hours) of untranscribed data. A recognizer is bootstrapped on the transcribed part of the data and initial transcripts are generated with it for the remainder (the untranscribed part). Using a lattice-based confidence measure, the recognition errors are (partially) detected and the remainder of the hypotheses is used for training. Using this scheme, the word error rate on a broadcast news speech recognition task dropped from more than 32.0% to 21.4%. In a cheating experiment we show, that this performance cannot be significantly improved by improving the measure of confidence. By combining the unsupervisedly trained system with our currently best recognizer which ... Thomas Kemp, Alex Waibel |
EUROSPEECH | 2 |
| 1999 | Mandarin large vocabulary speech recognition using the globalphone database
Jürgen Reichert, Tanja Schultz, Alex Waibel |
EUROSPEECH | 3 |
| 1999 | Towards spontaneous speech recognition for on-board car navigation and information systemsabstractThis paper presents a hierarchical approach to the Large-Scale Speaker Recognition problem. In here the authors present a binary tree data-base approach for arranging the trained speaker models based on a distance measure designed for comparing two sets of distributions. The combination of this hierarchical structure and the distance measure [1] provide the means for conducting a large-scale verification task. In addition, two techniques are presented for creating a model of the complement-space to the cohort which is used for rejection purposes. Results are presented for the drastic improvements achieved mainly in reducing the false-acceptance of the speaker verification system without any significant false-rejection degradation. Martin Westphal, Alex Waibel |
EUROSPEECH | 2 |
| 1999 | Progress in automatic meeting transcriptionabstractIn this paper we report recent developments on the meeting transcription task, a large vocabulary conversational speech recognition task. Previous experiments showed this is a very challenging task, with about 50% word error rate (WER) using existing recognizers. The difficulty mostly comes from highly disfluent/conversational nature of meetings, and lack of domain specific training data. For the first problem, our SWB(Switchboard) system --- a conversational telephone speech recognizer --- was used to recognize wide-band meeting data; for the latter, we leveraged the large amount of Broadcast News (BN) data to build a robust system. This paper will especially focus on two experiments in the BN system development: model combination and HMM topology/duration modeling. Model combination can be done at various stages of recognition: post-processing schemes such as ROVER can lead to significant improvements; to reduce computation we tried model combination at acoustic score level. We will ... Hua Yu 0008, Michael Finke, Alex Waibel |
EUROSPEECH | 3 |
| 1999 | Modeling focus of attention for meeting indexingabstractVisual cues, such as gesturing, looking at each other or monitoring each others facial expressions, play an important role in meetings.Such information can be used for indexing of multimedia meeting recordings.In this paper, we present an approach to detect who is looking at whom during a meeting.Our proposal is to employ Hidden Markov Models to characterize participants' focus of attention by using gaze information as well as knowledge about the number and positions of people present in a meeting.The number and positions of the participants faces are detected in the field of view of a panoramic camera.We use neural networks to estimate the directions of participants' gaze from camera images.We discuss the implementation of the approach in detail including system architecture, data collection, and evaluation.The system has achieved an accuracy rate of up to 93 % in detecting focus of attention on test sequences taken from meetings.We have used focus of attention as an index in a multimedia meeting browser. Rainer Stiefelhagen, Jie Yang 0001, Alex Waibel |
ACM Multimedia (1) | 3 |
| 1999 | Multimodal people ID for a multimedia meeting browserabstractA meeting browser is a system that allows users to review a multimedia meeting record from a variety of indexing methods. Identification of meeting participants is essential for creating such a multimedia meeting record. Moreover, knowing who is speaking can enhance the performance of speech recognition and indexing meeting transcription. In this paper, we present an approach that identifies meeting participants by fusing multimodal inputs. We use face ID, speaker ID, color appearance ID, and sound source directional ID to identify and track meeting. After describing the different modules in detail, we will discuss a framework for combining the information sources. Integration of the multimodal people ID into the multimedia meeting browser is in its preliminary stage. Jie Yang 0001, Xiaojin Zhu 0001, Ralph Gross, John Kominek, Alex Waibel |
ACM Multimedia (1) | 6 |
| 1999 | Translation systems under the C-STAR frameworkabstractThis talk will review our work on Speech Translation under the recent worldwide C-STAR demonstration. C-STAR is the Consortium for Speech Translation Advanced Research and now includes 6 partners and 20 partner/affiliate laboratories around the world. The work demonstrated concludes the second phase of the consortium, which has focused on translating conversational spontaneous speech as opposed to well formed, well structured text. As such, much of the work has focused on exploiting semantic and pragmatic constraints derived from the task domain and dialog situation to produce an understandable translation. Six partners have connected their respective systems with each other and allowed travel related spoken dialogs to provide communication between each of them. A common Interlingua representation was developed and used between the partners to make this multilingual deployment possible. The systems were also complemented by the introduction of Web based shared workspaces that allow one user in one country to communicate pictures, documents, sounds, tables, etc. to the other over the Web while referring to these documents in the dialog. Some of the partners’ systems were also deployed in wearable situations, such as a traveler exploring a foreign city. In this case speech and language technology was installed on a wearable computer with a small hand-held display. It was used to provide language translation as well as human-machine information access for the purpose of navigation (using GPS localization) and tour guidance. This combination of human-machine and human-machine-human dialogs could allow a user explore a foreign environment more effectively by resorting to human-machine and human-human dialogs wherever most appropriate. Alex Waibel |
MTSummit | 1 |
| 1999 | Stochastically-based semantic analysis for machine translation
Wolfgang Minker, Marsal Gavaldà, Alex Waibel |
Comput. Speech Lang. | 3 |
| 1998 | Skin-Color Modeling and Adaptation
Jie Yang 0001, Weier Lu, Alex Waibel |
ACCV (2) | 3 |
| 1998 | Visual Tracking for Multimodal Human Computer InteractionabstractIn this paper, we present visual tracking techniques for multimodal human computer interaction.First, we discuss techniques for tracking human faces in which human skin-color is used as a major feature.An adaptive stochastic model has been developed to characterize the skin-color distributions.Based on the maximum likelihood method, the model parameters can be adapted for different people and different lighting conditions.The feasibility of the model has been demonstrated by the development of a real-time face tracker.The system has achieved a rate of 30-t-frames/second using a low-end workstation with a framegrabber and a camera.We also present a top-down approach for tracking facial features such as eyes, nostrils, and lip comers.These real-time visual tracking techniques have been successfully applied to many applications such as gaze tracking, and lipreading.The face tracker has been combined with a microphone array for extracting speech signal from a specific person.The gaze tracker has been combined with a speech recognizer in a multimodal interface for controlling a panoramic image viewer. Jie Yang 0001, Rainer Stiefelhagen, Uwe Meier, Alex Waibel |
CHI | 4 |
| 1998 | Hierarchies of neural networks for connectionist speech recognition
Jürgen Fritsch, Alex Waibel |
ESANN | 2 |
| 1998 | Serbo-Croatian LVCSR on the dictation and broadcast news domainabstractThis paper describes the development of a Serbo-Croatian dictation and broadcast news speech recognizer. The intention is to generate an automatic text transcription of a news show, which will be submitted to a multilingual informedia database. We outline the complete system development process using the JanusRTk, beginning with data collection, design and training of the parameters, tuning and evaluation. We report on general recognition techniques like segmentation, adaptation and language model interpolation, as well as language specific problems, e.g. high OOV rate due to inflected word forms. We show that even with a low amount of acoustic training data, combined with Web based interpolated language models, it is sufficient to build up a fairly reliable automatic news transcription system, which yields a performance of 36.0% word error (WE). Peter Scheytt, Petra Geutner, Alex Waibel |
ICASSP | 3 |
| 1998 | Recognition of music typesabstractThis paper describes a music type recognition system that can be used to index and search in multimedia databases. A new approach to temporal structure modeling is supposed. The so called ETM-NN (explicit time modelling with neural network) method uses abstraction of acoustical events to the hidden units of a neural network. This new set of abstract features representing temporal structures, can be then learned via a traditional neural networks to discriminate between different types of music. The experiments show that this method outperforms HMMs significantly. Hagen Soltau, Tanja Schultz, Martin Westphal, Alex Waibel |
ICASSP | 4 |
| 1998 | Experiments in automatic meeting transcription using JRTKabstractWe describe our early exploration of automatic recognition of conversational speech in meetings for use in automatic summarizers and browsers to produce meeting minutes effectively and rapidly. To achieve optimal performance we started from two different baseline English recognizers adapted to meeting conditions and tested the resulting performance. The data were found to be highly disfluent (conversational human to human speech), noisy (due to lapel microphones and environment), and overlapped with background noise, resulting in error rates comparable so far to those on the CallHome conversational database (40-50% WER). A meeting browser is presented that allows the user to search and skim through highlights from a meeting efficiently despite the recognition errors. Hua Yu 0008, Cortis Clark, Robert G. Malkin, Alex Waibel |
ICASSP | 4 |
| 1998 | Effective structural adaptation of LVCSR systems to unseen domains using hierarchical connectionist acoustic modelsabstractWe present an approach to efficiently and effectively downsize and adapt the structure of large vocabulary conversational speech recognition (LVCSR) systems to unseen domains, requiring only small amounts of transcribed adaptation data. Our approach aims at bringing todays mostly task dependent systems closer to the aspired goal of domain independence. To achieve this, we rely on the ACID/HNN framework [2, 3], a hierarchical connectionist modeling paradigm which allows to dynamically adapt a tree structured modeling hierarchy to differing specifity of phonetic context in new domains. Experimental validation of the proposed approach has been carried out by adapting size and structure of ACID/HNN based acoustic models trained on Switchboard to two quite different, unseen domains, Wall Street Journal and an English Spontaneous Scheduling Task. In both cases, our approach yields considerably downsized acoustic models with performance improvements of up to 18% over the unadapted baseline models. Jürgen Fritsch, Michael Finke, Alex Waibel |
ICSLP | 3 |
| 1998 | Probabilistic dialogue act extraction for concept based multilingual translation systemsabstractThis paper describes a probabilistic method for dialogue act #DA# extraction for concept-based multilingual translation systems. ADA is a unit of a semantic interlingua and it consists of speaker information, speech act, concept and argument. Probabilistic models for the extraction of speech acts or concepts are trained as speech act or concept dependent word n-gram models. The proposed method is evaluated on DA-annotated English and Japanese databases. The experimental results show that the proposed method gives a better performance compared to the conventioanl grammar -based approach. In addition, the proposed method is much more robust for erroneous inputs obtained as speech recognition outputs. 1. INTRODUCTION In the C-STAR #Consortium for Speech Translation Advanced Research# project, several sites of spoken language groups, i.e., at CMU, ATR , UKA, ETRI, IRST 1 , etc. are developing multilingual speech-to-speech translation systems #1##2##3#. To facilitate multilingual transl... Toshiaki Fukada, Detlef Koll, Alex Waibel, Koichi Tanigaki |
ICSLP | 3 |
| 1998 | Conversational speech systems for on-board car navigation and assistanceabstractThis paper describes our latest efforts in building a speech recognizer for operating a navigation system through speech instead of typed input. Compared to conventional speech recognition for navigation systems, where the input is usually restricted to a fixed set of keywords and keyword phrases, complete spontaneous sentences are allowed as speech input. We will present the interaction of speech input, parsing and the necessary reactions to the requested queries. Our system has been trained on German spontaneous speech data and has been adapted to navigation queries using MLLR. As the system is not restricted to command word input, a parser is necessary to further process the recognized utterance. We show that within a lab environment our system is able to handle arbitrary spontaneous sentences as input to a navigation system successfully. The performance of the recognizer measured in word error rate gives a result of 18%. The parser has also been evaluated and yields an error rate o... Petra Geutner, Matthias Denecke, Uwe Meier, Martin Westphal, Alex Waibel |
ICSLP | 5 |
| 1998 | Phonetic-distance-based hypothesis driven lexical adaptation for transcribing multlingual broadcast newsabstractHigh out-of-vocabulary (OOV) rates are one of the most prevailing problems for languages with a rapid vocabulary growth due to a large number of inflections. Especially when transcribing SerboCroatian and German broadcast news, the OOV-rate is between 8.7% and 4.5%. Hypothesis Driven Lexical Adaptation (HDLA) has already been shown to decrease high OOV-rates significantly by using morphology-based linguistic knowledge. This paper introduces another approach to dynamically adapt a recognition lexicon to the utterance to be recognized. Instead of morphological knowledge about word stems and inflection endings, distance measures based on Levenstein distance are used. Results based on phoneme and grapheme distances will be presented. Compared to the use of morphological knowledge, our distance-based approach offers the distinct advantage that no expert knowledge about a specific language is required, no definition of complex grammar rules is necessary. Instead, grapheme sequences or the ph... Petra Geutner, Michael Finke, Alex Waibel |
ICSLP | 3 |
| 1998 | The interactive systems labs view4you video indexing systemabstractThe recognition of broadcast news is a challenging problem in speech recognition. To achieve the long-term goal of robust, real-time news transcription, several problems have to be overcome, e.g. the variety of acoustic conditions and the unlimited vocabulary. Recently, a number of sites have been working on content-addressable multi-media information sources. In the presented paper, we focus on extending this work towards a multi-lingual environment, where queries and multimedia documents may appear in multiple languages. In cooperation with the Informedia project at CMU [4], we attempt to provide cross-lingual access to German and Serbo-Croatian newscasts. 1. THE VIEW4YOU SYSTEM In the View4You system, German and Serbocroatian public newscasts are recorded daily using standard consumer electronics equipment. The newscasts are automatically segmented and an index is created for each of the segments by means of automatic speech recognition. The user can query the system in natural lan... Thomas Kemp, Petra Geutner, Borislav Tomaz, Manfred Weber, Martin Westphal, Alex Waibel |
ICSLP | 7 |
| 1998 | Reducing the OOV rate in broadcast news speech recognitionabstractThe recognition of broadcast news is a challenging problem in speech recognition. To achieve the long-term goal of robust, real-time news transcription, several problems have to be overcome, e.g. the variety of acoustic conditions and the unlimited vocabulary. In this paper we address the problem of unlimited vocabulary. We show, that this problem is more serious for German than it is for English. Using a speech recognition system with a large vocabulary, we dynamically adapt the active vocabulary to the topic of the current news segment. This is done by using information retrieval (IR) techniques on a large collection of texts automatically gathered from the internet. The same technique is also used to adapt the language model of the recognition system. The process of vocabulary adaptation and language model retraining is completely unsupervised. We show, that dynamic vocabulary adaptation can significantly reduce the out-of-vocabulary (OOV) rate and improve the word error rate of our... Thomas Kemp, Alex Waibel |
ICSLP | 2 |
| 1998 | Unsupervised training of a speech recognizer using TV broadcastsabstractCurrent speech recognition systems require large amounts of transcribed data for parameter estimation. The transcription, however, is tedious and expensive. In this work we describe our experiments which are aimed at training a speech recognizer without transcriptions. The experiments were carried out with TV newscasts, that were recorded using a satellite receiver and a simple MPEG coding hardware. The newscasts were automatically segmented into segments of similar acoustic background condition. This material is inexpensive and can be made available in large quantities, but there are no transcriptions available. We develop a training scheme, where a recognizer is bootstrapped using very little transcribed data and is improved using new, untranscribed speech. We show that it is necessary to use a confidence measure to judge the initial transcriptions of the recognizer before using them. Higher improvements can be achieved if the number of parameters in the system is increased when more... Thomas Kemp, Alex Waibel |
ICSLP | 2 |
| 1998 | An interlingua based on domain actions for machine translation of task-oriented dialoguesabstractThis paper describes an interlingua for spoken language translation that is based on domain actions in the travel planning domain. Domain actions are composed of speech acts (e.g., requestinformation) , attributes (e.g., size, price), and objects (e.g., hotel, flight) and can take arguments. Development of the interlingua is guided by a database containing travel dialogues in English, Korean, Japanese, and Italian. There are currently 423 domain actions that cover hotel reservation and transportation. The interlingua will soon be extended to cover tours, tourist attractions, and events. The interlingua is used by the C-STAR speech translation consortium for translating travel planning dialogues in six languages: English, Japanese, German, Korean, Italian, and French. The paper also addresses the role of the interlingua in Carnegie Mellon's JANUS translation system. Lori S. Levin, Donna Gates, Alon Lavie, Alex Waibel |
ICSLP | 4 |
| 1998 | Language independent and language adaptive large vocabulary speech recognitionabstractThis paper describes the design of a multilingual speech recognizer using an LVCSR dictation database which has been collected under the project GlobalPhone. This project at the University of Karlsruhe investigates LVCSR systems in 15 languages of the world, namely Arabic, Chinese, Croatian, English, French, German, Italian, Japanese, Korean, Portuguese, Russian, Spanish, Swedish, Tamil, and Turkish. Based on a global phoneme set we built different multilingual speech recognition systems for five of the 15 languages. Context dependent phoneme models are created data-driven by introducing questions about language and language groups to our polyphone clustering procedure. We apply the resulting multilingual models to unseen languages and present several recognition results in language independent and language adaptive setups. 1. Introduction As the demand for speech recognition systems in multiple languages grows, the development of multilingual systems which combine the phonetic inven... Tanja Schultz, Alex Waibel |
ICSLP | 2 |
| 1998 | On the influence of hyperarticulated speech on recognition performanceabstractSince we cannot exclude that speech recognizers fail sometimes, it is important to examine how users react to recognition errors. In correction situations, speaking style becomes more accentuated to disambiguate the original mistake. We examine the effect of speaking style in such situations on speech recognition performance. Our results indicate that hyperarticulated effects occur in correction situations and decrease word accuracy significantly. 1. INTRODUCTION Considerable progress has been achieved in speech recognition over the last years through techniques such as vocal tract length normalization (VTLN), maximum likelihood linear regression (MLLR) or speaker adapted training. However, even in dictation applications with 95% word accuracy (WA) it is often necessary to correct word errors. Studies [5] show that a user can lose a lot of time through error correction that he won through dictating instead of typing. Since recognition systems will always exhibit some errors it is impo... Hagen Soltau, Alex Waibel |
ICSLP | 2 |
| 1998 | Fast decoding for statistical machine translationabstractWe investigated an efficient decoding algorithm for statistical machine translation. Compared to the other algorithms, this new algorithm is applicable to different translation models, and it is much faster. Experiments showed that the algorithm achieved an overall performance comparable to the state of the art decoding algorithms. 1. INTRODUCTION A statistical machine translation system consists of three sub-tasks: the modeling task describes machine translation processes with stochastic models; the learning task estimates the parameters in the models; and the decoding task searches for the translation that has the highest score according to the models. [1, 2, 3] described different translation models and their learning algorithms. [4, 2, 5, 6] introduced different decoding algorithms. However, those decoding algorithms have many limitations. Below is a brief review of these algorithms: 1.1. IBM Stack Decoder In the IBM Stack Decoder [4], a hypothesis is comprised of a source sent... Ye-Yi Wang, Alex Waibel |
ICSLP | 2 |
| 1998 | Linear discriminant - a new criterion for speaker normalizationabstractIn Vocal Tract Length Normalization (VTLN) a linear or nonlinear frequency transformation compensates for different vocal tract lengths. Finding good estimates for the speaker specific warp parameters is a critical issue. Despite good results using the Maximum Likelihood criterion to find parameters for a linear warping, there are concerns using this method. We searched for a new criterion that enhances the inter-class separability in addition to optimizing the distribution of each phonetic class. Using such a criterion, Linear Discriminant Analysis determines a linear transformation in a lower dimensional space. For VTLN, we keep the dimension constant and warp the training samples of each speaker such that the Linear Discriminant is optimized. Although that criterion depends on all training samples of all speakers it can iteratively provide speaker specific warp factors. We discuss how this approach can be applied in speech recognition and present first results on two different recog... Martin Westphal, Tanja Schultz, Alex Waibel |
ICSLP | 3 |
| 1997 | Decoding Algorithm in Statistical Machine TranslationabstractDecoding algorithm is a crucial part in statistical machine translation. We describe a stack decoding algorithm in this paper. We present the hypothesis scoring method and the heuristics used in our algorithm. We report several techniques deployed to improve the performance of the decoder. We also introduce a simplified model to moderate the sparse data problem and to speed up the decoding process. We evaluate and compare these techniques/models in our statistical machine translation system. Ye-Yi Wang, Alex Waibel |
ACL | 2 |
| 1997 | Verbmobil: the combination of deep and shallow processing for spontaneous speech translationabstractVerbmobil is a speech-to-speech translation system for spontaneously spoken negotiation dialogs. The actual system translates 74.2% of spontaneously spoken German input. We give an overview of the Verbmobil system. After the introduction of the Verbmobil scenario and the unique constraints of the project, we describe the underlying system architecture and its realization. The progress that was achieved on the end-to-end translation rate owes much to the increase of the word recognition rate from 45% in 1993 to 87% in 1996. In order to achieve the envisaged coverage on the uncertain speech recognizer output, deep and shallow approaches to the analysis and transfer problem had to be combined. Thomas Bub, Wolfgang Wahlster, Alex Waibel |
ICASSP | 3 |
| 1997 | Context-dependent hybrid HME/HMM speech recognition using polyphone clustering decision treesabstractThis paper presents a context-dependent hybrid connectionist speech recognition system that uses a set of generalized hierarchical mixtures of experts (HME) to estimate context-dependent posterior acoustic class probabilities. The connectionist part of the system is organized in a modular fashion, allowing the distributed training of such a system on regular workstations. Context classes are based on polyphonic contexts, clustered using decision trees which we adopt from our continuous density HMM recognizer JANUS (Waibel et al., 1996). The system is evaluated on ESST, an English speaker-independent spontaneous speech database. Context dependent modeling is shown to yield significant improvements over simple context-independent modeling, requiring only small additional overhead in terms of training and decoding time. Jürgen Fritsch, Michael Finke, Alex Waibel |
ICASSP | 3 |
| 1997 | Janus-III: speech-to-speech translation in multiple languagesabstractThis paper describes JANUS-III, our most recent version of the JANUS speech-to-speech translation system. We present an overview of the system and focus on how system design facilitates speech translation between multiple languages, and allows for easy adaptation to new source and target languages. We also describe our methodology for evaluation of end-to-end system performance with a variety of source and target languages. For system development and evaluation, we have experimented with both push-to-talk as well as cross-talk recording conditions. To date, our system has achieved performance levels of over 80% acceptable translations on transcribed input, and over 70% acceptable translations on speech input recognized with a 75-90% word accuracy. Our current major research is concentrated on enhancing the capabilities of the system to deal with input in broad and general domains. Alon Lavie, Alex Waibel, Lori S. Levin, Michael Finke, Donna Gates, Marsal Gavaldà, Torsten Zeppenfeld, Puming Zhan |
ICASSP | 2 |
| 1997 | Multimodal interfaces for multimedia information agentsabstractWhen humans communicate they take advantage of a rich spectrum of cues. Some are verbal and acoustic. Some are non-verbal and non-acoustic. Signal processing technology has devoted much attention to the recognition of speech, as a single human communication signal. Most other complementary communication cues, however, remain unexplored and unused in human-computer interaction. In this paper we show that the addition of non-acoustic or non-verbal cues can significantly enhance robustness, flexibility, naturalness and performance of human-computer interaction. We demonstrate computer agents that use speech, gesture, handwriting, pointing, spelling jointly for more robust, natural and flexible human-computer interaction in the various tasks of an information worker: information creation, access, manipulation or dissemination. Alex Waibel, Bernhard Suhm, Minh Tue Vo, Jie Yang 0001 |
ICASSP | 1 |
| 1997 | Recognition of conversational telephone speech using the JANUS speech engineabstractRecognition of conversational speech is one of the most challenging speech recognition tasks to-date. While recognition error rates of 10% or lower can now be reached on speech dictation tasks over vocabularies in excess of 60,000 words, recognition of conversational speech has persistently resisted most attempts at improvements by way of the proven techniques to date. Difficulties arise from shorter words, telephone channel degradation, and highly disfluent and coarticulated speech. In this paper, we describe the application, adaptation, and performance evaluation of our JANUS speech recognition engine to the Switchboard conversational speech recognition task. Through a number of algorithmic improvements, we have been able to reduce error rates from more than 50% word error to 38%, measured on the offical 1996 NIST evaluation test set. Improvements include vocal tract length normalization, polyphonic modeling, label boosting, speaker adaptation with and without confidence measures, and speaking mode dependent pronunciation modeling. Torsten Zeppenfeld, Michael Finke, Klaus Ries 0001, Martin Westphal, Alex Waibel |
ICASSP | 5 |
| 1997 | Dialogue strategies guiding users to their communicative goalsabstractMuch work has been done in dialogue modeling for Human - Computer Interaction. Problems arise in situations where disambiguation of highly ambiguous data base output is necessary. We propose to model the task rather than the dialogue itself. Furthermore, we propose underspecified representations to represent relevant data and to serve as a base for generating clarification questions that guide the user efficiently to arrive at his communicative goal. In this paper, we establish a connection between underspecified representations as representations of disjunctions and clarification questions. Our approach to clarifying dialogues differs from other approaches in that the form of the clarification dialogues is entirely determined by the domain modeling and by the underspecified representations. 1 Introduction In spoken dialogue systems, the need for clarification questions arises in situations in which information is missing (e.g. due to partial interpretation in the presence of recogni... Matthias Denecke, Alex Waibel |
EUROSPEECH | 2 |
| 1997 | Speaking mode dependent pronunciation modeling in large vocabulary conversational speech recognitionabstractIn spontaneous conversational speech there is a large amount ofvariability due to accents, speaking styles and speaking rates (also known as the speaking mode) [3]. Because current recognition systems usually use only a relatively small number of pronunciation variants for the words in their dictionaries, the amount ofvariability that can be modeled is limited. Increasing the number of variants per dictionary entry is the obvious solution. Unfortunately, this also means increasing the confusability between the dictionary entries, and thus often leads to an actual performance decrease. In this paper we present a framework for speaking mode dependent pronunciation modeling. The probability of encountering pronunciation variants is de ned to be a function of the speaking style. The probability function is learned through decision trees from rule based generated pronunciation variants as observed on the Switchboard corpus. The framework is successfully applied to increase the performance of our state-of-the-art Janus Recognition Toolkit Switchboard recognizer signi cantly. 1. Michael Finke, Alex Waibel |
EUROSPEECH | 2 |
| 1997 | Japanese LVCSR on the spontaneous scheduling task with JANUS-3abstractThis paper presents our findings during the development of the recognition engine for the Japanese part of the VERBMOBIL speech-to-speech translation project. We describe an efficient method to bootstrap a large vocabulary speech recognizer for spontaneously spoken Japanese speech from a German recognizer and show that the amount of effort in developing the system could be reduced by using this rapid cross language bootstrapping technique. The Japanese recognizer is integrated into the VERBMOBIL system and shows very promising results achieving 9.3% word error rate. 1. INTRODUCTION The overall goal of the first phase of the VERBMOBIL project is to build a speech-to-speech translation system from both German and Japanese spontaneously spoken input speech to English, German and Japanese output in an appointment scenario [1]. The Japanese recognizer described in this paper is beeing designed to be part of this translation system. Unlike Japanese dictation systems [2] there is no need fo... Tanja Schultz, Detlef Koll, Alex Waibel |
EUROSPEECH | 3 |
| 1997 | Fast bootstrapping of LVCSR systems with multilingual phoneme setsabstractIn this paper we described an efficient method to bootstrap continuously spoken, large vocabulary speech recognition systems by multilingual phoneme sets. To evaluate this techniques we collected the multilingual database GlobalPhone which currently consists of 9 different languages. A multilingual recognizer (MULTI) based on the four languages German, English, Japanese and Spanish was developed to serve as a source system. Likewise this system is very useful for language identification and achieves 100% language identification rate. Based on the MULTI system we evaluated our bootstrap technique on such completely different languages as Chinese, Croatian, and Turkish. 1. INTRODUCTION As the demand for speech recognition and translation systems in multiple languages grows, the development of multilingual systems is of increasing concern. On the one hand a multilingual system can be used as a language independent speech recognition and translation system with integrated automatic langu... Tanja Schultz, Alex Waibel |
EUROSPEECH | 2 |
| 1997 | Exploiting repair context in interactive error recovery
Bernhard Suhm, Alex Waibel |
EUROSPEECH | 2 |
| 1997 | Statistical analysis of dialogue structureabstractWe introduce a statistical model for dialogues. We describe a dynamic programming algorithm that can be used to bracket a dialogue into segments and label each segment with its speech act. We evaluate the performance of the model. We also use this model for language modelling and get perplexity reduction. 1 INTRODUCTION Dialogue structure provides important information for spoken language understanding. This structure comprises the current topic, discourse state, and speech act, etc. Many researchers used topic information to reduce the perplexity of a task [1, 2]. In our experiments, we also found that dialogue structure information also helps to reduce ambiguities and improve spoken language translation performance. While knowledge-based approaches are used widely and successfully in dialogue structure analysis[3, 4], they require intensive human effort in defining linguistic structures and developing grammars to detect the structures. We would like to build a model that is able to ... Ye-Yi Wang, Alex Waibel |
EUROSPEECH | 2 |
| 1997 | Speaker normalization and speaker adaptation - a combination for conversational speech recognitionabstractSpeaker normalization and speaker adaptation are two strategies to tackle the variations from speaker, channel, and environment. The vocal tract length normalization (VTLN) is an effective speaker normalization approach to compensate for the variations of vocal tract shapes. The Maximum Likelihood Linear Regression(MLLR) is a recent proposed method for speaker-adaptation. In this paper, we propose a speaker-specific Bark scale VTLN method, investigate the combination of the VTLN with MLLR, and present an iterative procedure for decoding the combined system of VTLN and MLLR. The results show that: (1) the new VTLN method is very effective with which the word error rate can be reduced up to 11%; (2) the combination of VTLN and MLLR can provide up to 15% word error reduction; (3) both VTLN and MLLR are more effective for the push-to-talk data than for the cross-talk data. 1 INTRODUCTION Almost all speech recognizers are, in some extent, sensitive to the variations of speakers and/or env... Puming Zhan, Martin Westphal, Michael Finke, Alex Waibel |
EUROSPEECH | 4 |
| 1996 | FeasPar - A Feature Structure Parser Learning to Parse Spoken Language
Finn Dag Buø, Alex Waibel |
COLING | 2 |
| 1996 | Multi-lingual Translation of Spontaneously Spoken Language in a Limited Domain
Alon Lavie, Donna Gates, Marsal Gavaldà, Laura Mayfield Tomokiyo, Alex Waibel, Lori S. Levin |
COLING | 5 |
| 1996 | Search in a Learnable Spoken Language Parser
Finn Dag Buø, Alex Waibel |
ECAI | 2 |
| 1996 | LVCSR-based language identificationabstractAutomatic language identification is an important problem in building multilingual speech recognition and understanding systems. Building a language identification module for four languages we studied the influence of applying different levels of knowledge sources on a large vocabulary continuous speech recognition (LVCSR) approach, i.e. phonetic, phonotactic, lexical, and syntactic-semantic knowledge. The resulting language identification (LID) module can identify spontaneous speech input and can be used as a front end for the multilingual speech-to-speech translation system JANUS-II. A comparison of five LID systems showed that the incorporation of lexical and linguistic knowledge reduces the language identification error for the 2-language tests up to 50%. Based on these results we build a LID module for German, English, Spanish, and Japanese which yields 84% identification rate on the spontaneous scheduling task (SST). Tanja Schultz, Ivica Rogina, Alex Waibel |
ICASSP | 3 |
| 1996 | JANUS-II-translation of spontaneous conversational speechabstractJANUS-II is a research system to design and test components of speech-to-speech translation systems as well as a research prototype for such a system. We focus on two aspects of the system: (1) the new features of the speech recognition component JANUS-SR, and (2) the end-to-end performance of JANUS-II, including a comparison of two machine translation strategies used for JANUS-MT (PHOENIX and GLR*). Alex Waibel, Michael Finke, Donna Gates, Marsal Gavaldà, Thomas Kemp, Alon Lavie, Lori S. Levin, Laura Mayfield Tomokiyo, Arthur E. McNair, Ivica Rogina, Kaori Shima, Tilo Sloboda, Monika Woszczyna, Torsten Zeppenfeld, Puming Zhan |
ICASSP | 1 |
| 1996 | Focus of attention: Towards low bitrate video tele-conferencingabstractLow bitrate video tele-conferencing requires adapting algorithms that may work perfectly well in a high-bitrate situation. When a slow transmission rate is unacceptable, compromise must be reached among the demands of speed, bandwidth limits and image quality. In this paper we present an approach to low bitrate video tele-conferencing by focusing attention on important information. We show that by selectively degrading the quality of less important regions, more important regions can be sent without loss of quality but with greatly reduced bandwidth requirements. A prototype system has been developed to demonstrate the concept. The experimental results show significant savings of required bandwidth for video subjected to the changes. Jie Yang 0001, Leejay Wu, Alex Waibel |
ICIP (2) | 3 |
| 1996 | Learning to parse spontaneous speechabstractWe describe and experimentally evaluate a system, FeasPar, that learns parsing spontaneous speech.To train and run FeasPar Feature Structure Parser, only limited handmodeled knowledge is required.The FeasPar architecture consists of neural networks and a search.The networks spilt the incoming sentence into chunks, which are labeled with feature values and chunk relations.Then, the search nds the most probable and consistent feature structure.FeasPar is trained, tested and evaluated with the Spontaneous Scheduling Task, and compared with two samples of a handmodeled GLR* parser, developed for 4 months and 2 y ears, respectively.The handmodeling e ort for FeasPar i s 2 w eeks.FeasPar performes better than the GLR* parser developed 4 months in all six comparisons that are made and has a similar performance as the GLR* parser developed for 2 y ears. Finn Dag Buø, Alex Waibel |
ICSLP | 2 |
| 1996 | Recognizing emotion in speechabstractThis paper explores several statistical pattern recognition techniques to classify utterances according to their emotional content.We have recorded a corpus containing emotional speech with over a 1000 utterances from different speakers.We present a new method of extracting prosodic features from speech, based on a smoothing spline approximation of the pitch contour.To make maximal use of the limited amount of training data available, we introduce a novel pattern recognition technique: majority voting of subspace specialists.Using this technique, we obtain classification performance that is close to human performance on the task. Frank Dellaert, Thomas Polzin, Alex Waibel |
ICSLP | 3 |
| 1996 | Recognition of spelled names over the telephoneabstractRecognition of spelled names over the telephone line is essential for applications such as telephone directory assistance, or automatic mail ordering.We present recognition results on the spelling section of the OGI Spelled and Spoken Word Telephone Corpus, using a Multi-State Time Delay Neural Network MS-TDNN.Many applications allow for strong language modeling constraints.In our experiments we examined the bene cial e ects of reducing the search space to a list of last names, ranging from about 1000 to 14 million entries.We compare tree search methods and show that signi cant improvements can be achieved by enriching the search trees with probabilities. Hermann Hild, Alex Waibel |
ICSLP | 2 |
| 1996 | Dialogue processing in a conversational speech translation systemabstractAttempts at discourse processing of spontaneously spoken dialogue face several difficulties: multiple hypotheses that result from the parser's attempts to make sense of the output from the speech recognizer, ambiguity that results from segmentation of multi-sentence utterances, and cumulative error -errors in the discourse context which cause further errors when subsequent sentences are processed.In this paper we will describe our robust parsers, our procedures for segmenting long utterances, and two approaches to discourse processing that attempt to deal with ambiguity and cumulative error. Alon Lavie, Lori S. Levin, Yan Qu, Alex Waibel, Donna Gates, Marsal Gavaldà, Laura Mayfield Tomokiyo, Maite Taboada |
ICSLP | 4 |
| 1996 | Translation of conversational speech with JANUS-IIabstractIn this paper we investigate the possibility of translating continuous spoken conversations in a cross-talk environment.This is a task known to be difficult for human translators due to several factors.It is characterized by rapid and even overlapping turn-taking, a high degree of co-articulation, and fragmentary language.We describe experiments using both push-to-talk as well as cross-talk recording conditions.Our results indicate that conversational speech recognition and translation is possible, even in a free crosstalk environment.To date, our system has achieved performances of over 80% acceptable translations on transcribed input, and over 70% acceptable translations on speech input recognized with a 70-80% word accuracy.The system's performance on spontaneous conversations recorded in a cross-talk environment is shown to be as good and even slightly superior to the simpler and easier push-to-talk scenario. Alon Lavie, Alex Waibel, Lori S. Levin, Donna Gates, Marsal Gavaldà, Torsten Zeppenfeld, Puming Zhan, Oren Glickman |
ICSLP | 2 |
| 1996 | Class phrase models for language modelling
Klaus Ries 0001, Finn Dag Buø, Alex Waibel |
ICSLP | 3 |
| 1996 | Dictionary learning for spontaneous speech recognitionabstractSpontaneous speech a d d s a v ariety of phenomena to a speech recognition task: false starts, human and nonhuman noises, new words, and alternative pronunciations.All of these phenomena have t o b e t a c kled when adapting a speech r e c o gnition system for spontaneous speech.In this paper we will focus on how to automatically expand and adapt phonetic dictionaries for spontaneous speech recognition.Especially for spontaneous speech it is important t o c hoose the pronunciations of a word according to the frequency in which t h e y appear in the database rather than the \correct" pronunciation as might be found in a lexicon.Therefore, we proposed a data-driven approach to add new pronunciations to a given phonetic dictionary 1] i n a w ay that they model the given occurrences of words in the database.We will show h o w t h i s algorithm can be extended to produce alternative p r o n unciations for word tuples and frequently misrecognized words.We will also discuss how further knowledge can be incorporated into the phoneme recognizer in a way t h a t i t l e a r n s t o generalize from pronunciations which w ere found previously.The experiments have been performed on the German Spontaneous Scheduling Task (GSST), using the speech recognition engine of JANUS 2, the spontaneous speech-to-speech translation system of the Interactive Tilo Sloboda, Alex Waibel |
ICSLP | 2 |
| 1996 | Interactive recovery from speech recognition errors in speech user interfacesabstractrepair are very accurate.N-best choice appears to be more effective in this setting, probably due to the small vocabularies. Bernhard Suhm, Brad A. Myers, Alex Waibel |
ICSLP | 3 |
| 1996 | Word clustering with parallel spoken language corporaabstractIn this paper we i n troduce a word clustering algorithm which uses a bilingual, parallel corpus to group together words in the source and target language.Our method generalizes previous mutual information clustering algorithms for monolingual data by incorporating a statistical translation model.Preliminary experiments have shown that the algorithm can e ectively employ the constraints implicit in bilingual data to extract classes which are well-suited to machine translation tasks. Ye-Yi Wang, John D. Lafferty, Alex Waibel |
ICSLP | 3 |
| 1996 | JANUS-II: towards spontaneous Spanish speech recognitionabstractJANUS-II is a research system for investigating various issues in speech-to-speech translations and has been implemented for speech-to-speech translations on many languages 1 .In this paper, we address the Spanish speech recognition part of JANUS-II.First, we report the bootstrap and optimization of the recognition system.Then we i n v estigate the di erence between push-to-talk and cross-talk dialogs, which are two di erent kinds of data in our database.We give a detail noise analysis for the push-to-talk and cross-talk dialogs and present some recognition results for the comparison.We h a v e observed that the cross-talk dialogs are harder than the pushto-talk dialogs for speech recognition, because they are more noisy than the latter.Currently, the error rate of our Spanish recognizer is 27 for push-to-talk test set and 32 for crosstalk test set. Puming Zhan, Klaus Ries 0001, Marsal Gavaldà, Donna Gates, Alon Lavie, Alex Waibel |
ICSLP | 6 |
| 1996 | Adaptively Growing Hierarchical Mixtures of Experts
Jürgen Fritsch, Michael Finke, Alex Waibel |
NIPS | 3 |
| 1996 | A real-time face trackerabstractThe authors present a real-time face tracker. The system has achieved a rate of 30+ frames/second using an HP-9000 workstation with a frame grabber and a Canon VC-Cl camera. It can track a person's face while the person moves freely (e.g., walks, jumps, sits down and stands up) in a room. Three types of models have been employed in developing the system. First, they present a stochastic model to characterize skin color distributions of human faces. The information provided by the model is sufficient for tracking a human face in various poses and views. This model is adaptable to different people and different lighting conditions in real-time. Second, a motion model is used to estimate image motion and to predict the search window. Third, a camera model is used to predict and compensate for camera motion. The system can be applied to teleconferencing and many HCI applications including lip reading and gaze tracking. The principle in developing this system can be extended to other tracking problems such as tracking the human hand. Jie Yang 0001, Alex Waibel |
WACV | 2 |
| 1995 | Knowing who to listen to in speech recognition: visually guided beamformingabstractWith speech recognition systems steadily improving in performance, freedom from head-sets and push-buttons to activate the recognizer is one of the most important issues to achieve user acceptance. Microphone arrays and beamforming can deliver signals that suppress undesired jamming signals but rely on knowledge where the signal is in space. This knowledge is usually derived by identifying the loudest signal source. Knowing who is speaking to whom and where should however not depend on loudness, but on the communication purpose. In this paper, we present acoustic and visual modules that use tracking of the face of a speaker of interest for sound source localization and beamforming for signal extraction. It is shown that in noisy environments a more accurate localization in space can be delivered visually than acoustically. Given a reliable location finder, beamforming substantially improves recognition accuracy. Udo Bub, Martin Hunke, Alex Waibel |
ICASSP | 3 |
| 1995 | Toward movement-invariant automatic lip-reading and speech recognitionabstractWe present the development of a modular system for flexible human-computer interaction via speech. The speech recognition component integrates acoustic and visual information (automatic lip-reading) improving overall recognition, especially in noisy environments. The image of the lips, constituting the visual input, is automatically extracted from the camera picture of the speaker's face by the lip locator module. Finally, the speaker's face is automatically acquired and followed by the face tracker sub-system. Integration of the three functions results in the first bi-modal speech recognizer allowing the speaker reasonable freedom of movement within a possibly noisy room while continuing to communicate with the computer via voice. Compared to audio-alone recognition, the combined system achieves a 20 to 50 percent error rate reduction for various signal/noise conditions. Paul Duchnowski, Martin Hunke, Dietrich Büsching, Uwe Meier, Alex Waibel |
ICASSP | 5 |
| 1995 | Concept-based speech translationabstractAs part of the JANUS speech-to-speech translation project, the authors have developed a robust translation system based on the information structures inherent to the task being performed. The basic premise is that the structure of the information to be transmitted is largely independent of the language used to encode it. The system performs no syntactic analysis; speaker utterances are parsed into semantic chunks, which can be strung together without grammatical rules, and passed through a simple template-based translation module. The authors have achieved encouraging coverage rates on English, German and Spanish input with English, German and Spanish output. Laura Mayfield Tomokiyo, Marsal Gavaldà, Wayne H. Ward, Alex Waibel |
ICASSP | 4 |
| 1995 | NPen++: a writer independent, large vocabulary on-line cursive handwriting recognition systemabstractIn this paper we describe the NPen/sup ++/ system for writer independent on-line handwriting recognition. This recognizer needs no training for a particular writer and can recognize any common writing style (cursive, hand-printed, or a mixture of both). The neural network architecture, which was originally proposed for continuous speech recognition tasks, and the preprocessing techniques of NPen/sup ++/ are designed to make heavy use of the dynamic writing information, i.e. the temporal sequence of data points recorded on an LCD tablet or digitizer. We present results for the writer independent recognition of isolated words. Tested on different dictionary sizes from 1,000 up to 100,000 words, recognition rates range from 98.0% for the 1,000 word dictionary to 91.4% on a 20,000 word dictionary and 82.9% for the 100,000 word dictionary. No language models are used to achieve these results. Stefan Manke, Michael Finke, Alex Waibel |
ICDAR | 3 |
| 1995 | Speeding up the score computation of HMM speech regognizers with the bucket voronoi intersection algorithmabstractWith increasing sizes of speech databases, speech recognizers with huge parameter spaces have become trainable. However, the time and memory requirements for high accuracy realtime speaker-independent continuous speech recognition will probably not be met by the available hardware for a reasonable price for the next few years. This paper describes the application of the Bucket Voronoi Intersection algorithm to the JANUS-2 speech recognizer, which reduces the time for the computation of HMM emission probabilities with large Gaussian mixtures by 50% to 80%. 1. INTRODUCTION Although the computation of Gaussians is only a part (for very large vocabularies, even a small part) of the overall run time, speeding it up does reduce the reaction time of the recognizer, and especially the time for training significantly. When computing the log probability of a Gaussian mixture, many speech recognizer do not use all Gaussians but only the top n. We have found that in our system using only the one ... Jürgen Fritsch, Ivica Rogina, Tilo Sloboda, Alex Waibel |
EUROSPEECH | 4 |
| 1995 | Integrating spelling into spoken dialogue recognitionabstractRecognition of spelled letter sequences is essential for many real-world applications which involve arbitrary names or addresses. Often the letter sequences carry the sentence's crucial information; therefore, it is important to correctly localize and recognize the spelled string. However, large vocabulary speech recognizers tend to perform poorly on spelled letters, especially if they have to deal with spontaneous speech. The research presented here aims at improving the recognition accuracy of spontaneous speech with embedded spelled-letter sequences. We propose methods to localize spelled-letter segments and reclassify them with a specialized letter recognizer. In: 4th European Conference on Speech Communication and Technology (EUROSPEECH '95), Madrid, Spain, 18 - 21 September 1995 1. INTRODUCTION Applications of spelling in spontaneous speech include the recognition of spelled names or addresses, as well as repair dialogues, where spelling can be used to disambiguate the inevitab... Hermann Hild, Alex Waibel |
EUROSPEECH | 2 |
| 1995 | Translation and interpretation of spontaneous speech
Alex Waibel |
MTSummit | 1 |
| 1995 | The challenge of spoken language systems: research directions for the ninetiesabstractA spoken language system combines speech recognition, natural language processing and human interface technology. It functions by recognizing the person's words, interpreting the sequence of words to obtain a meaning in terms of the application, and providing an appropriate response back to the user. Potential applications of spoken language systems range from simple tasks, such as retrieving information from an existing database (traffic reports, airline schedules), to interactive problem solving tasks involving complex planning and reasoning (travel planning, traffic routing), to support for multilingual interactions. We examine eight key areas in which basic research is needed to produce spoken language systems: (1) robust speech recognition; (2) automatic training and adaptation; (3) spontaneous speech; (4) dialogue models; (5) natural language response generation; (6) speech synthesis and speech generation; (7) multilingual systems; and (8) interactive multimodal systems. In each area, we identify key research challenges, the infrastructure needed to support research, and the expected benefits. We conclude by reviewing the need for multidisciplinary research, for development of shared corpora and related resources, for computational support and far rapid communication among researchers. The successful development of this technology will increase accessibility of computers to a wide range of users, will facilitate multinational communication and trade, and will create new research specialties and jobs in this rapidly expanding area.> Ronald A. Cole, Lynette Hirschman, Les E. Atlas, Mary E. Beckman, Alan Biermann, Marcia A. Bush, Mark A. Clements, Jordan Cohen, Oscar Garcia, Brian A. Hanson, Hynek Hermansky, Steve Levinson, Kathy McKeown, Nelson Morgan, David G. Novick, Mari Ostendorf, Sharon L. Oviatt, Patti Price, Harvey F. Silverman, Judy Spitz, Alex Waibel, Clifford J. Weinstein, Stephen A. Zahorian, Victor Zue |
IEEE Trans. Speech Audio Process. | 21 |
| 1994 | Learning complex output representations in connectionist parsing of spoken languageabstractDue to robustness, learnability and ease of integration of different information sources, connectionist parsing systems have proven to be applicable for parsing spoken language, However, most proposed connectionist parsers do not compute and represent complex structures. These parsers assign only a very limited structure to a given input string. For spoken language translation and data base access, more detailed syntactic and semantic representation is needed. In the present paper, the authors show that arbitrary linguistic features and arbitrary complex tree structures can indeed also be learned by a connectionist parsing system.> Finn Dag Buø, Thomas Polzin, Alex Waibel |
ICASSP (1) | 3 |
| 1994 | Learning state-dependent stream weights for multi-codebook HMM speech recognition systemsabstractMany speech recognition systems use multiple information streams to compute HMM output probabilities (e.g. systems based on semicontinuous or discrete HMMs use one codebook for cepstral coefficients, and another one for delta cepstral coefficients). The final score is a weighted sum of the contributions of every stream. These weights can be found empirically and usually the same set of weights is used for every acoustic model. There is reason to believe that there are features which are more important for some acoustic models than for others. Especially one would expect the beginning and ending segment of a phoneme to be more context dependent than the middle part, so in that case the probability estimator of the speech recognizer should put more emphasis on the delta-spectrum than on the spectrum. Experiments have shown that spectral or cepstral coefficients are more important than their derivatives and more important than power or delta-power coefficients. We propose an algorithm for learning individual stream weights for every HMM state. Since these individual weights are a superset of the stream-only dependent weights, they can reproduce the results of the stream-only dependent weights and, additionally, discriminate between HMM states. Thus, the recognition performance must improve.> Ivica Rogina, Alex Waibel |
ICASSP (1) | 2 |
| 1994 | JANUS 93: towards spontaneous speech translationabstractWe present first results from our efforts toward translation of spontaneously spoken speech. Improvements include increasing coverage, robustness, generality and speed of JANUS, the speech-to-speech translation system of Carnegie Mellon and Karlsruhe University. The recognition and machine translation engine have been upgraded to deal with requirements introduced by spontaneous human to human dialogs. To allow for development and evaluation of our system on adequate data, a large database with spontaneous scheduling dialogs is being gathered for English, German and Spanish.> Monika Woszczyna, Naomi Aoki-Waibel, Finn Dag Buø, Noah Coccaro, Keiko Horiguchi, Thomas Kemp, Alon Lavie, Arthur E. McNair, Thomas Polzin, Ivica Rogina, Carolyn P. Rosé, Tanja Schultz, Bernhard Suhm, Masaru Tomita, Alex Waibel |
ICASSP (1) | 15 |
| 1994 | Combining bitmaps with dynamic writing information for on-line handwriting recognitionabstractWriter independent, large vocabulary online handwriting recognition systems require robust input representations, which make optimal use of the dynamic writing information, i.e. the temporal ordering of the sampled data points. In this paper we describe an input representation for cursive handwriting, which combines this dynamic writing information with static bitmaps used in optical character recognition. This input representation is used with a connectionist recognizer, which is well suited for handling temporal sequences of patterns as provided by this kind of input representation. Our system has been tested on different cursive handwriting recognition tasks with vocabulary sizes up to 20000 words. We achieve recognition rates up to 99.5% on writer independent, single character recognition tasks and up to 98.1% on writer dependent, cursive handwriting tasks. Stefan Manke, Michael Finke, Alex Waibel |
ICPR (2) | 3 |
| 1994 | See me, hear me: integrating automatic speech recognition and lip-readingabstractWe present recent work on integration of visual information (automatic lip-reading) with acoustic speech for better overall speech recognition. A Multi-State Time Delay Neural Network performs the recognition of spelled letter sequences taking advantage of lip images from a standard camera. The problems addressed include efficient but effective representation of the visual information and optimum manner of combining the two modalities when rendering a decision. We show results for several alternatives to direct gray level image as the visual evidence. These are: Principal Components, Linear Discriminants, and DFT coefficients. Dimensionality of the input is decreased by a factor of 12 while maintaining recognition rates. Combination of the visual and acoustic information is performed at three different levels of abstraction. Results suggest that integration of higher order input features works best. On a continuous spelling task, visual-alone recognition of 45-55%, when combined with a... Paul Duchnowski, Uwe Meier, Alex Waibel |
ICSLP | 3 |
| 1994 | Improving recognizer acceptance through robust, natural speech repair
Arthur E. McNair, Alex Waibel |
ICSLP | 2 |
| 1994 | Towards better language models for spontaneous speechabstractIn our effort to build a speech--to--speech translation system for spontaneous spoken dialogs we have developed several methods to improve the language models of the speech decoder of the system. We attempt to take advantage of natural equivalence word classes, frequently occuring word phrases, and discourse structure. Each of these methods was tested on spontaneous English, German and Spanish human--human dialogs. 1. INTRODUCTION The goal of the JANUS project is multi-lingual machine translation of spontaneously spoken dialogs in a limited domain: two people scheduling a meeting with each other. We are currently working with English, German, and Spanish as source languages and German, English, and Japanese as target languages. Table 1 shows the size of training and test set for the English, German and Spanish Spontaneous Scheduling Task databases (ESST, GSST, SSST) used for all experiments reported in this paper, and the coverage of the dictionary over the test set. 1 ESST GSST SSST... Bernhard Suhm, Alex Waibel |
ICSLP | 2 |
| 1994 | Inferring linguistic structure in spoken languageabstractWe demonstrate the applications of Markov Chains and HMMs to modeling of the underlying structure in spontaneous spoken language. Experiments with supervised training cover the detection of the current dialog state and identification of the speech act as used by the speech translation component in our JANUS Speech-to-Speech Translation System. HMM training with hidden states is used to uncover other levels of structure in the task. The possible use of the model for perplexity reduction in a continuous speech recognition system is also demonstrated. To achieve improvement over a state independent bigram language model, great care must be taken to keep the number of model parameters small in the face of limited amounts of training data from transcribed spontaneous speech. 1. INTRODUCTION In spoken language understanding productive interpretation of an utterance has to incorporate the underlying linguistic structure in a dialog. This structure comprises the current topic, discourse state... Monika Woszczyna, Alex Waibel |
ICSLP | 2 |
| 1994 | The Use of Dynamic Writing Information in a Connectionist On-Line Cursive Handwriting Recognition SystemabstractIn this paper we present NPen ++, a connectionist system for writer independent, large vocabulary on-line cursive handwriting recognition. This system combines a robust input representation, which preserves the dynamic writing information, with a neural network architecture, a so called Multi-State Time Delay Neural Network (MS-TDNN), which integrates rec.ognition and segmen(cid:173) tation in a single framework. Our preprocessing transforms the original coordinate sequence into a (still temporal) sequence offea(cid:173) ture vectors, which combine strictly local features, like curvature or writing direction, with a bitmap-like representation of the co(cid:173) ordinate's proximity. The MS-TDNN architecture is well suited for handling temporal sequences as provided by this input rep(cid:173) resentation. Our system is tested both on writer dependent and writer independent tasks with vocabulary sizes ranging from 400 up to 20,000 words. For example, on a 20,000 word vocabulary we achieve word recognition rates up to 88.9% (writer dependent) and 84.1 % (writer independent) without using any language models. Stefan Manke, Michael Finke, Alex Waibel |
NIPS | 3 |
| 1994 | Introduction Structured Connectionist Systems
Alex Waibel |
Mach. Learn. | 1 |
| 1993 | Improving connected letter recognition by lipreading
Christoph Bregler, Hermann Hild, Stefan Manke, Alex Waibel |
ICASSP (1) | 4 |
| 1993 | Multi-speaker/speaker-independent architectures for the multi-state time delay neural network
Hermann Hild, Alex Waibel |
ICASSP (2) | 2 |
| 1993 | Improving the MS-TDNN for word spotting
Torsten Zeppenfeld, Rick Houghton, Alex Waibel |
ICASSP (2) | 3 |
| 1993 | Tuning by doing: flexibility through automatic structure optimization
Ulrich Bodenhausen, Alex Waibel |
EUROSPEECH | 2 |
| 1993 | Speaker-independent connected letter recognition with a multi-state time delay neural networkabstractThe Multi-State Time Delay Neural Network (MS-TDNN) inte-grates a nonlinear time alignment procedure (DTW) and the high-accuracy phoneme spotting capabilities of a TDNN into a connec-tionist speech recognition system with word-level classification and error backpropagation. We present an MS-TDNN for recognizing continuously spelled letters, a task characterized by a small but highly confusable vocabulary. Our MS-TDNN achieves 98.5/92.0% word accuracy on speaker dependent/independent tasks, outper-forming previously reported results on the same databases. We pro-pose training techniques aimed at improving sentence level perfor-mance, including free alignment across word boundaries, word du-ration modeling and error backpropagation on the sentence rather than the word level. Architectures integrating submodules special-ized on a subset of speakers achieved further improvements. 1 Hermann Hild, Alex Waibel |
EUROSPEECH | 2 |
| 1993 | Detection and transcription of new wordsabstractThis paper describes a model which enables a speech recognition system to automatically detect new words and to provide a rough phonetic transcription. In our approach to the new word problem the decision whether new words occurred in the speech input is not based exclusively on acoustic evidence but also on a language model designed to support the detection of new words. We describe preliminary experiments to create new word grammars on the Wall Street Journal task. Furthermore we present recognition results of our new word model using the recognition engine of the JANUS speech to speech translation system [1, 2], designed around the task of conference registration. 1. Bernhard Suhm, Monika Woszczyna, Alex Waibel |
EUROSPEECH | 3 |
| 1993 | Recent advances in JANUS: a speech translation systemabstractWe present recent advances from our efforts in increasing coverage, robustness, generality and speed of JANUS, CMU's speech-tospeech translation system. JANUS is a speaker-independent system translating spoken utterances in English and also in German into one of German, English or Japanese. The system has been designed around the task of conference registration (CR). It has initially been built based on a speech database of 12 read dialogs, encompassing a vocabulary of around 500 words. We have since been expanding the system along several dimensions to improve speed, robustness and coverage and to move toward spontaneous input. 1. INTRODUCTION In this paper we describe recent improvements of JANUS, a speech to speech translation system. Improvements have been made mainly along the following dimensions: 1.) better context-dependent modeling improves performance in the speech recognition module, 2.) improved language models, smoothing, and word equivalence classes improve coverage and ... Monika Woszczyna, Noah Coccaro, Andreas Eisele 0001, Alon Lavie, Arthur E. McNair, Thomas Polzin, Ivica Rogina, Carolyn P. Rosé, Tilo Sloboda, Masaru Tomita, J. Tsutsumi, Naomi Aoki-Waibel, Alex Waibel, Wayne H. Ward |
EUROSPEECH | 13 |
| 1992 | PARSEC: a structured connectionist parsing system for spoken languageabstractThe authors present PARSEC-a system for generating connectionist parsing networks from example parses. PARSEC is not based on formal grammar systems and has been geared towards spoken language tasks. PARSEC networks exhibit three strengths important for application to speech processing: they learn to parse, and generalize well compared to hand-coded grammars; they tolerate several types of noise; and they can learn to use multimodal input. The authors also present the PARSEC architecture, its training algorithms, and performance analyses along several dimensions that demonstrate PARSEC's features. They compare PARSEC's performance to that of traditional grammar-based parsing systems.> Ajay N. Jain, Alex Waibel, David S. Touretzky |
ICASSP | 2 |
| 1992 | Testing generality in JANUS: a multi-lingual speech translation systemabstractFor speech translation to be practical and useful, speech translation systems should be portable to multiple languages without substantial modification. The authors present results of expanding the English-based JANUS speech translation system to translate from spoken German sentences to English and Japanese utterances. The authors also report the results of implementing part of the linked predictive neural network (LPNN) speech recognition module on a massively parallel machine. The JANUS approach generalizes well, with overall system performance of 97%. This surpasses English-based JANUS performance.> Louise Osterholtz, Charles Augustine, Arthur E. McNair, Ivica Rogina, Hiroaki Saito 0001, Tilo Sloboda, Joe Tebelskis, Alex Waibel |
ICASSP | 8 |
| 1992 | A hybrid neural network, dynamic programming word spotterabstractA novel keyword-spotting system that combines both neural network and dynamic programming techniques is presented. This system makes use of the strengths of time delay neural networks (TDNNs), which include strong generalization ability, potential for parallel implementations, robustness to noise, and time shift invariant learning. Dynamic programming models are used by this system because they have the useful capability of time warping input speech patterns. This system was trained and tested on the Stonehenge Road Rally database, which is a 20-keyword-vocabulary, speaker-independent, continuous-speech corpus. Currently, this system performs at a figure of merit (FOM) rate of 82.5%. FOM is the detection rate averaged from 0 to 10 false alarms per keyword hour. This measure is explained in detail.> Torsten Zeppenfeld, Alex Waibel |
ICASSP | 2 |
| 1992 | Connected Letter Recognition with a Multi-State Time Delay Neural Network
Hermann Hild, Alex Waibel |
NIPS | 2 |
| 1992 | Performance Through Consistency: MS-TDNN's for Large Vocabulary Continuous Speech Recognition
Joe Tebelskis, Alex Waibel |
NIPS | 2 |
| 1992 | The Meta-Pi Network: Building Distributed Knowledge Representations for Robust Multisource Pattern RecognitionabstractThe authors present the Meta-Pi network, a multinetwork connectionist classifier that forms distributed low-level knowledge representations for robust pattern recognition, given random feature vectors generated by multiple statistically distinct sources. They illustrate how the Meta-Pi paradigm implements an adaptive Bayesian maximum a posteriori classifier. They also demonstrate its performance in the context of multispeaker phoneme recognition in which the Meta-Pi superstructure combines speaker-dependent time-delay neural network (TDNN) modules to perform multispeaker /b,d,g/ phoneme recognition with speaker-dependent error rates of 2%. Finally, the authors apply the Meta-Pi architecture to a limited source-independent recognition task, illustrating its discrimination of a novel source. They demonstrate that it can adapt to the novel source (speaker), given five adaptation examples of each of the three phonemes.> John B. Hampshire II, Alex Waibel |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1992 | Integrated phoneme and function word architecture of hidden control neural networks for continuous speech recognition
Bojan Petek, Alex Waibel, Joe Tebelskis |
Speech Commun. | 2 |
| 1991 | Learning the architecture of neural networks for speech recognitionabstractResults are presented that suggest that it is possible to learn the architecture of neural networks for speech recognition systems. The Tempo 2 algorithm is proposed. It is a training algorithm for neural networks that trains the temporal parameters of the network (delays and widths of the input windows) as well as the weights. A comparison of the performances with one adaptive parameter set (either weights, delays or widths) shows that the main parameters are the weights. Delays and widths seem to be of lesser importance, but in combination with the weights the temporal parameters can improve performance, especially generalization. A Tempo 2 network with trained delays and widths and random weights can classify 70% of the phonemes correctly. The application to phoneme classification, shows that this adaptive architecture can approach the performance of a carefully hand-tuned TDNN (time-delay neural network) and leads to more compact networks.> Ulrich Bodenhausen, Alex Waibel |
ICASSP | 2 |
| 1991 | Integrating time alignment and neural networks for high performance continuous speech recognitionabstractThe authors describe two systems in which neural network classifiers are merged with dynamic programming (DP) time alignment methods to produce high-performance continuous speech recognizers. One system uses the connectionist Viterbi-training (CVT) procedure, in which a neural network with frame-level outputs is trained using guidance from a time alignment procedure. The other system uses multi-state time-delay neural networks (MS-TDNNs), in which embedded DP time alignment allows network training with only word-level external supervision. The CVT results on the, TI Digits are 99.1% word accuracy and 98.0% string accuracy. The MS-TDNNs are described in detail, with attention focused on their architecture, the training procedure, and results of applying the MS-TDNNs to continuous speaker-dependent alphabet recognition: on two speakers, word accuracy is respectively 97.5% and 89.7%.> Patrick Haffner, Michael A. Franzini, Alex Waibel |
ICASSP | 3 |
| 1991 | Continuous speech recognition using linked predictive neural networksabstractThe authors present a large vocabulary, continuous speech recognition system based on linked predictive neural networks (LPNNs). The system uses neural networks as predictors of speech frames, yielding distortion measures which can be used by the one-stage DTW algorithm to perform continuous speech recognition. The system currently achieves 95%, 58%, and 39% word accuracy on tasks with perplexity 7, 111, and 402, respectively, outperforming several simple HMMs that have been tested. It was also found that the accuracy and speed of the LPNN can be slightly improved by the judicious use of hidden control inputs. The strengths and weaknesses of the predictive approach are discussed.> Joe Tebelskis, Alex Waibel, Bojan Petek, Otto Schmidbauer |
ICASSP | 2 |
| 1991 | JANUS: a speech-to-speech translation system using connectionist and symbolic processing strategiesabstractThe authors present JANUS, a speech-to-speech translation system that utilizes diverse processing strategies including dynamic programming, stochastic techniques, connectionist learning, and traditional AI knowledge representation approaches. JANUS translates continuously spoken English utterances into Japanese and German speech utterances. The overall system performance on a corpus of conference registration conversations is 87%. Two versions of JANUS are compared: one using a LR parser (JANUS 1) and one using a connectionist parser (JANUS 2). Performance results were mixed, with JANUS 1 deriving benefit from a tighter language model and JANUS 2 benefitting from greater flexibility.> Alex Waibel, Ajay N. Jain, Arthur E. McNair, Hiroaki Saito 0001, Alex Hauptmann 0001, Joe Tebelskis |
ICASSP | 1 |
| 1991 | A connectionist model for dialog processingabstractA novel connectionist system for dialog processing is described. Based on a script-like formalism, the system consists of several modular neural networks which can track the semantic flow of a dialog. The system can be extended to understand and translate dialogs in a certain domain.> Ye-Yi Wang, Alex Waibel |
ICASSP | 2 |
| 1991 | Recent work in continuous speech recognition using the connectionist viterbi training procedure
Michael A. Franzini, Alex Waibel, Kai-Fu Lee |
EUROSPEECH | 2 |
| 1991 | Time-delay neural networks embedding time alignment: a performance analysis
Patrick Haffner, Alex Waibel |
EUROSPEECH | 2 |
| 1991 | Evaluation of speaker-independent phoneme recognition on TIMIT database using TDNNs
Nobuo Hataoka, Alex Waibel |
EUROSPEECH | 2 |
| 1991 | Integrated phoneme-function word architecture of hidden control neural networks for continuous speech recognition
Bojan Petek, Alex Waibel, Joe Tebelskis |
EUROSPEECH | 2 |
| 1991 | Multi-State Time Delay Networks for Continuous Speech Recognition
Patrick Haffner, Alex Waibel |
NIPS | 2 |
| 1991 | JANUS: Speech-to-Speech Translation Using Connectionist and Non-Connectionist Techniques
Alex Waibel, Ajay N. Jain, Arthur E. McNair, Joe Tebelskis, Louise Osterholtz, Hiroaki Saito 0001, Otto Schmidbauer, Tilo Sloboda, Monika Woszczyna |
NIPS | 1 |
| 1990 | Connectionist Viterbi training: a new hybrid method for continuous speech recognitionabstractA hybrid method for continuous-speech recognition which combines hidden Markov models (HMMs) and a connectionist technique called connectionist Viterbi training (CVT) is presented. CVT can be run iteratively and can be applied to large-vocabulary recognition tasks. Successful completion of training the connectionist component of the system, despite the large network size and volume of training data, depends largely on several measures taken to reduce learning time. The system is trained and tested on the TI/NBS speaker-independent continuous-digits database. Performance on test data for unknown-length strings is 98.5% word accuracy and 95.0% string accuracy. Several improvements to the current system are expected to increase these accuracies significantly.> Michael A. Franzini, Kai-Fu Lee, Alex Waibel |
ICASSP | 3 |
| 1990 | The Meta-Pi network: connectionist rapid adaptation for high-performance multi-speaker phoneme recognitionabstractA multinetwork time-delay-neural-network (TDNN)-based connectionist architecture that allows multispeaker phoneme discrimination (/b,d,g/) to be performed at the speaker-dependent recognition rate of 98.4% is presented. The overall network gates the phonemic decisions of modules trained on individual speakers to form its overall classification decision. By dynamically adapting to the input speech and focusing on a combination of speaker-specific modules, the network outperforms a single TDNN trained on the speech of all six speakers (95.9%). To train this network a form of multiplicative connection called the Meta-Pi connection is developed. It is illustrated how the Mega-Pi paradigm implements a dynamically adaptive Bayesian MAP classifier. It learns-without supervision-to recognize the speech of one particular speaker (99.8%) using a dynamic combination of internal models of other speakers exclusively. The Meta-Pi model is a viable basis for a connectionist speech recognition system that can rapidly adapt to new speakers and varying speaker dialects.> John B. Hampshire II, Alex Waibel |
ICASSP | 2 |
| 1990 | Robust connectionist parsing of spoken languageabstractA modular, recurrent connectionist network architecture which learns to robustly perform incremental parsing of complex sentences is presented. From sequential input, one word at a time, the networks learn to do semantic role assignment, noun phrase attachment, and clause structure recognition for sentences with passive constructions and center embedded clauses. The networks make syntactic and semantic predictions at every point in time, and previous predictions are revised as expectations are affirmed or violated with the arrival of new information. The networks induce their own grammar rules for dynamically transforming an input sequence of words into a syntactic/semantic interpretation. These networks generalize and display tolerance to input which has been corrupted in ways common in spoken language.> Ajay N. Jain, Alex Waibel |
ICASSP | 2 |
| 1990 | Large vocabulary recognition using linked predictive neural networksabstractA large-vocabulary isolated word recognition system based on linked predictive neural networks (LPNNs) is presented. In this system, neural networks are used as predictors of speech frames, enabling a pool of such networks to serve as phoneme models. Higher-level algorithms are used to organize these networks, linking them into sequences corresponding to the phonetic spellings of words, and to train and evaluate the networks for word recognition. By virtue of linking phonemic networks, the LPNN is vocabulary independent and can be applied to large-vocabulary recognition. Recognition rates of 94% for a 234-word Japanese vocabulary of acoustically similar words and 90% for a larger vocabulary of 924 words are obtained.> Joe Tebelskis, Alex Waibel |
ICASSP | 2 |
| 1990 | Speaker-independent phoneme recognition on TIMIT database using integrated time-delay neural networks (TDNNs)abstractA structure of neural networks (NNs) is described for speaker-independent and context-independent phoneme recognition. This structure is based on the integration of time-delay neural networks (TDNN) which have several TDNNs separated according to the duration of phonemes. As a result, the proposed structure deals with phonemes of varying duration more effectively. In the experimental evaluation of the proposed structure, 16 English vowel recognition was performed using 5268 vowel tokens picked from 480 sentences spoken by 140 speakers (98 males and 42 females) on the TIMIT (TI-MIT) database. The number of training tokens and testing tokens was 4326 from 100 speakers (69 males and 31 females) and 942 from 40 speakers (29 males and 11 females), respectively. The result was a 60.5% recognition rate (around 70% for a collapsed 13-vowel case), which was improved from 56% in the single TDNN structure, showing the effectiveness of the proposed structure's use of temporal information Nobuo Hataoka, Alex Waibel |
IJCNN | 2 |
| 1990 | Phoneme-based word recognition by neural network - a step toward large vocabulary recognitionabstractA neural-network-based word recognition system extendible to large-vocabulary isolated word recognition is presented. The system consists of (1) time-delay neural networks (TDNNs) for phoneme spotting and (2) a higher-level network and a dynamic programming (DP) time alignment procedure for word recognition. TDNN-based phenome-spotting networks are used whose role is to fire when a particular phenome is input. A higher-level network then improves these phenome firing patterns in view of an idealized phoneme sequence. For training of the higher-level network, DP matching is used to determine idealized phoneme firing patterns which are nearest to the actual phoneme firings. During recognition, the system selects the most probable word by applying DP matching to the outputs of the higher-level network. Speaker-dependent and isolated word recognition experiments show that word recognition rates of around 92% can be achieved for medium-size vocabularies Akihiro Hirai, Alex Waibel |
IJCNN | 2 |
| 1990 | Speech recognition using sub-phoneme recognition neural network
Kiyoaki Aikawa, Alex Waibel |
ICSLP | 2 |
| 1990 | The Tempo 2 Algorithm: Adjusting Time-Delays By Supervised Learning
Ulrich Bodenhausen, Alex Waibel |
NIPS | 2 |
| 1990 | Continuous Speech Recognition by Linked Predictive Neural Networks
Joe Tebelskis, Alex Waibel, Bojan Petek, Otto Schmidbauer |
NIPS | 2 |
| 1990 | A time-delay neural network architecture for isolated word recognition
Kevin J. Lang, Alex Waibel, Geoffrey E. Hinton |
Neural Networks | 2 |
| 1990 | A novel objective function for improved phoneme recognition using time-delay neural networksabstractSingle-speaker and multispeaker recognition results are presented for the voice-stop consonants /b,d,g/ using time-delay neural networks (TDNNs) with a number of enhancements, including a new objective function for training these networks. The new objective function, called the classification figure of merit (CFM), differs markedly from the traditional mean-squared-error (MSE) objective function and the related cross entropy (CE) objective function. Where the MSE and CE objective functions seek to minimize the difference between each output node and its ideal activation, the CFM function seeks to maximize the difference between the output activation of the node representing incorrect classifications. A simple arbitration mechanism is used with all three objective functions to achieve a median 30% reduction in the number of misclassifications when compared to TDNNs trained with the traditional MSE back-propagation objective function alone. John B. Hampshire II, Alex Waibel |
IEEE Trans. Neural Networks | 2 |
| 1989 | Spotting Japanese CV-syllables and phonemes using time-delay neural networksabstractThe authors present techniques for spotting Japanese CV syllables/phonemes in input speech based on TDNNs. They constructed a TDNN which can discriminate a single CV syllable or phoneme group. In Japanese, there are only about one hundred syllables, or fewer than 30 phonemes, which makes it feasible to prepare and train the TDNN to spot all possible syllables or phonemes extracted as training tokens from training words. Syllable and phoneme spotting experiments show excellent results, including a syllable spotting rate of better than 96.7% correct. These spotting techniques are proved to be a significant step toward continuous speech recognition.> Hidefumi Sawai, Alex Waibel, Masanori Miyatake, Kiyohiro Shikano |
ICASSP | 2 |
| 1989 | Consonant recognition by modular construction of large phonemic time-delay neural networksabstractIt is shown that neural networks for speech recognition can be constructed in a modular fashion by exploiting the hidden structure of previously trained phonetic subcategory networks. The performance of resulting larger phonetic nets was found to be as good as the performance of the subcomponent nets by themselves. This approach avoids the excessive learning times that would be necessary to train larger networks and allows for incremental learning. Large time-delay neural networks constructed incrementally by applying these modular training techniques achieved a recognition performance of 96.0% for all consonants and 94.7% for all phonemes.> Alex Waibel, Hidefumi Sawai, Kiyohiro Shikano |
ICASSP | 1 |
| 1989 | Fast back-propagation learning methods for large phonemic neural networks
Patrick Haffner, Alex Waibel, Hidefumi Sawai, Kiyohiro Shikano |
EUROSPEECH | 2 |
| 1989 | Connectionist Architectures for Multi-Speaker Phoneme Recognition
John B. Hampshire II, Alex Waibel |
NIPS | 2 |
| 1989 | Incremental Parsing by Modular Recurrent Connectionist Networks
Ajay N. Jain, Alex Waibel |
NIPS | 2 |
| 1989 | Modular Construction of Time-Delay Neural Networks for Speech RecognitionabstractSeveral strategies are described that overcome limitations of basic network models as steps towards the design of large connectionist speech recognition systems. The two major areas of concern are the problem of time and the problem of scaling. Speech signals continuously vary over time and encode and transmit enormous amounts of human knowledge. To decode these signals, neural networks must be able to use appropriate representations of time and it must be possible to extend these nets to almost arbitrary sizes and complexity within finite resources. The problem of time is addressed by the development of a Time-Delay Neural Network; the problem of scaling by Modularity and Incremental Design of large nets based on smaller subcomponent nets. It is shown that small networks trained to perform limited tasks develop time invariant, hidden abstractions that can subsequently be exploited to train larger, more complex nets efficiently. Using these techniques, phoneme recognition networks of increasing complexity can be constructed that all achieve superior recognition performance. Alex Waibel |
Neural Comput. | 1 |
| 1988 | Noise reduction using connectionist modelsabstractUsing a back propagation network learning algorithm, a four-layered feed-forward network is trained on learning samples to realize a mapping from the set of noisy signals a set of noise-free signals. Computer experiments were carried out on 12 kHz sampled Japanese speech data, using stationary and nonstationary noise. The experiments showed that the network can indeed learn to perform noise reduction. Even for noisy speech signals that had not been part of the training data, the network successfully produced noise-suppressed output signals.> Shin'ichi Tamura, Alex Waibel |
ICASSP | 2 |
| 1988 | Phoneme recognition: neural networks vs. hidden Markov modelsabstractA time-delay neural network (TDNN) for phoneme recognition is discussed. By the use of two hidden layers in addition to an input and output layer it is capable of representing complex nonlinear decision surfaces. Three important properties of the TDNNs have been observed. First, it was able to invent without human interference meaningful linguistic abstractions in time and frequency such as formant tracking and segmentation. Second, it has learned to form alternate representations linking different acoustic events with the same higher level concept. In this fashion it can implement trading relations between lower level acoustic events leading to robust recognition performance despite considerable variability in the input speech. Third, the network is translation-invariant and does not rely on precise alignment or segmentation of the input. The TDNNs performance is compared with the best of hidden Markov models (HMMs) on a speaker-dependent phoneme-recognition task. The TDNN achieved a recognition of 98.5% compared to 93.7% for the HMM, i.e., a fourfold reduction in error.> Alex Waibel, Toshiyuki Hanazawa, Geoffrey E. Hinton, Kiyohiro Shikano, Kevin J. Lang |
ICASSP | 1 |
| 1988 | Consonant Recognition by Modular Construction of Large Phonemic Time-Delay Neural Networks
Alex Waibel |
NIPS | 1 |
| 1987 | Prosodic knowledge sources for word hypothesization in a continuous speech recognition systemabstractPreviously we have reported on the extraction of prosodic cues (such as stress, pitch, duration) from continuous speech [1] and have reported on possible uses of some prosodic information (e.g., temporal cues [2]) in large vocabulary word recognition systems. In this paper we extend these previous findings to a speaker-independent continuous speech recognition system. Speaker-independent knowledge sources (KS) were implemented that attempt to hypothesize words based on only prosodic cues found in the signal. The prosodic cues exploited were temporal cues (syllable durations, ratios of unvoiced segment durations to syllable durations, voiced segment durations), intensity profiles and likelihoods of stressedness. Each KS extracts the appropriate prosodic cue and searches its knowledge base for words whose prosodic patterns satisfy the constraints found in the signal. Usign a multispeaker continuous speechdatabase for evaluation, each prosodic KS is shown to hypothesize the correct word substantially better than chance. All prosodic KSs were then combined and compared with a speaker-independent acoustic-phonetic word hypothesizer. After applying the prosodic KSs, the correct word ranked on average 25th (out of 252 words). The acoustic-phonetic KS alone yielded an average rank of 40 (out of 252) without the addition of prosodic information. After prosodic and phonetic KSs were combined the average rank was reduced to 15 out of 252. The results indicate that prosodic information indeed adds complementary information that substantially improves word hypothesization in speaker-independent continuous speech recognition systems. Alex Waibel |
ICASSP | 1 |
| 1986 | Recognition of lexical stress in a continuous speech understanding system - A pattern recognition approachabstractStress is one of the key components in human speech perception. Its uses extend from the phonetic level over the lexical to the syntactic and semantic level. Several methods have been developed in the past to detect stress automatically from the signal. This paper takes a pattern recognition approach to the the problem of stress detection. The algorithm presented has three key features: (1) optimal combination of the evidence obtained from the acoustic correlates of stress is achieved by means of a Bayesian classifier assuming multivariate Gaussian distributions; (2) the algorithm detects lexical stress in continuously spoken English utterances; (3) rather than making hard decisions, the algorithm returns probabilities for each syllable, i.e., a measure of stressedness. The algorithm was tested over 4 databases of differing continuous speech data. When a forced decision is imposed by setting a threshold at stress probability 0.5, error rates of 7.79% to 14.85% missed stresses were obtained. Unlike in other languages (such as Japanese), amplitude integrals are the strongest predictor of English stress. Performance results and an analysis of errors are presented. Alex Waibel |
ICASSP | 1 |
| 1985 | A coarse phonetic knowledge source for template independent large vocabulary word recognitionabstractIn this paper we present a template independent knowledge source (KS), that uses coarse phonetic information to substantially constrain the candidate vocabulary for use in word hypothesization with very large vocabularies. It consists of three parts: the segmenter that breaks a test utterance up into a sequence of coarse phonetic classes, the knowledge compiler that generates a reference dictionary containing the appropriate coarse phonetic representations for each word candidate and finally, a matching engine. Coarse phonetic classification is performed using linear discriminant analysis, more specifically perceptron classification. The knowledge compiler first generates a phonemic representation and segmental durations by rule from a list of word candidates (i.e., from text), and then derives coarse phonetic class segments. Matching is performed by a nonlinear time alignment algorithm based on dissimilarity scores between detected and lexical coarse class segments. The coarse phonetic KS was tested by compiling a word list of approximately 1500 words. Using only the coarse classes Silence, Plosive, Fricative, Vocalic, Front Vowel, Back Vowel, Nasal and R, a vocabulary reduction to 5% of the original vocabulary is achieved at lower than 5% error rate for three different speakers. Helmut Lagger, Alex Waibel |
ICASSP | 2 |
| 1984 | Suprasegmentals in very large vocabulary isolated word recognitionabstractProsodic information is believed to be valuable informnation in human speech perception, but speech recognition systems to date have largely been based on segmental spectral analysis. In this paper I describe parts of a front end to a very-large-vocabulary isolated word recognition system using prosodic information. The present front end is template independent (speaker training for large vocabulary systems (> 20,000 words) is undesirable) and makes use of robust cues in the incoming speech to obtain a presorted vocabulary of candidates. It is shown that prosodic information, e.g., the rhythmic structure of an input word, its syllabic structure, voiced/unvoiced regions in the word and the temporal distribution of back/front vowels, nasals and liquids and glides, can be used effectively to select a substantially reduced subvocabulary of candidates, before any fine phonetic analysis is attempted to recognize the word. Alex Waibel |
ICASSP | 1 |
| 1982 | Performance trade-offs in search techniques for isolated word speech recognitionabstractThe Cost effectiveness of various search methods used in experimental and practical discrete utterance speech recognition systems is a very critical factor for the usefulness of such systems. The advantages of some cost effective search techniques, e.g. branch and bound search, branch and bound search with pruning and beam search, have been previoulsy reported. In this paper we analyze the properties that affect the practical usefulness of these algorithms when task characteristics and machine architecture are considered. Roberto Bisiani, Alex Waibel |
ICASSP | 2 |