Florian Metze

dblp:26/1652 · DBLP profile ↗
← Back
178ranked-venue papers
21as first author
31since 2021 · last 2025
0000-0002-6663-8600ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 150 · 20 first-author · 20 since 2021Artificial intelligence and machine learning · 105 · 13 first-author · 25 since 2021Databases, data management, data science and information retrieval · 3Applied, interdisciplinary, general and emerging computing · 2Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 Long-Form Fuzzy Speech-to-Text Alignment for 1000+ Languages
abstract
Conventional speech-to-text forced alignment typically operates at the utterance level. In practice, however, we do not usually have short segments (e.g., 10 seconds) of audio with exact, verbatim transcriptions (e.g., the LibriSpeech corpus) as in lab conditions. Instead, audio often comes in long-form (e.g., an hour-long lecture recording), and the available transcription may be non-verbatim or include unspoken annotations, making it misaligned with the actual speech. This motivates the need for long-form fuzzy speech-to-text alignment, which has practical applications - for example, preparing segmented supervised audio data for training machine learning models. We demonstrate the Torchaudio long-form aligner, which supports such use cases. Moreover, it can be equipped with any CTC model that predicts frame-wise labels, turning the model into a robust and powerful aligner.
Ruizhe Huang, Xiaohui Zhang 0007, Zhaoheng Ni, Moto Hira, Jeff Hwang, Vineel Pratap, Ju Lin, Ming Sun 0013, Florian Metze
ASRU9
2025 MMW: Side Talk Rejection Multi-Microphone Whisper On Smart Glasses
abstract
Smart glasses are increasingly positioned as the nextgeneration interface for ubiquitous access to large language models (LLMs). Nevertheless, achieving reliable interaction in real-world noisy environments remains a major challenge, particularly due to interference from side speech. In this work, we introduce a novel side-talk rejection multimicrophone Whisper (MMW) framework for smart glasses, incorporating three key innovations. First, we propose a Mix Block based on a Tri-Mamba architecture to effectively fuse multi-channel audio at the raw waveform level, while maintaining compatibility with streaming processing. Second, we design a Frame Diarization Mamba Layer to enhance frame-level side-talk suppression, facilitating more efficient fine-tuning of Whisper models. Third, we employ a Multi-Scale Group Relative Policy Optimization (GRPO) strategy to jointly optimize frame-level and utterance-level side speech suppression. Experimental evaluations demonstrate that the proposed MMW system can reduce the word error rate (WER) by 4.95% in noisy conditions.
Yiteng Huang, Yangyang Shi, Saurabh Adya, Ming Sun 0013, Florian Metze
ASRU8
2025 Directional Speech Recognition with Full-Duplex Capability
Ju Lin, Yiteng Huang, Ming Sun 0013, Frank Seide, Florian Metze
INTERSPEECH5
2025 MASV: Speaker Verification with Global and Local Context Mamba
Yiteng Huang, Ming Sun 0013, Xinhao Mei, Yangyang Shi, Florian Metze
INTERSPEECH8
2025 Thinking in Directivity: Speech Large Language Model for Multi-Talker Directional Speech Recognition
Jiamin Xie, Ju Lin, Yiteng Huang, Tyler Vuong, Zhaojiang Lin, Prashant Rawat, Sangeeta Srivastava, Ming Sun 0013, Florian Metze
INTERSPEECH11
2024 Audio-Journey: Open Domain Latent Diffusion Based Text-To-Audio Generation
abstract
Despite recent progress, machine learning (ML) models for open-domain audio generation need to catch up to generative models for image, text, speech, and music. The lack of massive open-domain audio datasets is the main reason for this performance gap; we overcome this challenge through a novel data augmentation approach. We leverage state-of-the-art (SOTA) Large Language Models (LLMs) to enrich captions in the weakly-labeled audio dataset. We then use a SOTA video-captioning model to generate captions for the videos from which the audio data originated, and we again use LLMs to merge the audio and video captions to form a rich, large-scale dataset. We experimentally evaluate the quality of our audio-visual captions, showing a 12.5% gain in semantic score over baselines. Using our augmented dataset, we train a Latent Diffusion Model to generate in an encodec encoding latent space. Our model is novel in the current SOTA audio generation landscape due to our generation space, text encoder, noise schedule, and attention mechanism. Together, these innovations provide competitive open-domain audio generation. The samples, models, and implementation will be at https://audiojourney.github.io.
Jackson Michaels, Juncheng Li 0001, Laura Yao, Lijun Yu, Zach Wood-Doughty, Florian Metze
ICASSP6
2023 CTC Alignments Improve Autoregressive Translation
abstract
Brian Yan, Siddharth Dalmia, Yosuke Higuchi, Graham Neubig, Florian Metze, Alan W Black, Shinji Watanabe. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023.
Brian Yan, Siddharth Dalmia, Yosuke Higuchi, Graham Neubig, Florian Metze, Alan W. Black, Shinji Watanabe 0001
EACL5
2023 LegoNN: Building Modular Encoder-Decoder Models
abstract
State-of-the-art encoder-decoder models (e.g. for machine translation (MT) or automatic speech recognition (ASR)) are constructed and trained end-to-end as an atomic unit. No component of the model can be (re-)used without the others, making it impossible to share parts, e.g. a high resourced decoder, across tasks. We describe LegoNN, a procedure for building encoder-decoder architectures in a way so that its parts can be applied to other tasks without the need for any fine-tuning. To achieve this reusability, the interface between encoder and decoder modules is grounded to a sequence of marginal distributions over a pre-defined discrete vocabulary. We present two approaches for ingesting these marginals; one is differentiable, allowing the flow of gradients across the entire network, and the other is gradient-isolating. To enable the portability of decoder modules between MT tasks for different source languages and across other tasks like ASR, we introduce a modality agnostic encoder which consists of a length control mechanism to dynamically adapt encoders' output lengths in order to match the expected input length range of pre-trained decoders. We present several experiments to demonstrate the effectiveness of LegoNN models: a trained language generation LegoNN decoder module from German-English (De-En) MT task can be reused without any fine-tuning for the Europarl English ASR and the Romanian-English (Ro-En) MT tasks, matching or beating the performance of baseline. After fine-tuning, LegoNN models improve the Ro-En MT task by 1.5 BLEU points and achieve 12.5% relative WER reduction on the Europarl ASR task. To show how the approach generalizes, we compose a LegoNN ASR model from three modules – each has been learned within different end-to-end trained models on three different datasets – achieving an overall WER reduction of 19.5%.
Siddharth Dalmia, Dmytro Okhonko, Mike Lewis, Sergey Edunov, Shinji Watanabe 0001, Florian Metze, Luke Zettlemoyer, Abdel-rahman Mohamed
IEEE ACM Trans. Audio Speech Lang. Process.6
2022 Self-supervised object detection from audio-visual correspondence
abstract
We tackle the problem of learning object detectors without supervision. Differently from weakly-supervised object detection, we do not assume image-level class labels. Instead, we extract a supervisory signal from audio-visual data, using the audio component to “teach” the object detector. While this problem is related to sound source localisation, it is considerably harder because the detector must classify the objects by type, enumerate each instance of the object, and do so even when the object is silent. We tackle this problem by first designing a self-supervised framework with a contrastive objective that jointly learns to classify and localise objects. Then, without using any supervision, we simply use these self-supervised labels and boxes to train an image-based object detector. With this, we outperform previous unsupervised and weakly-supervised detectors for the task of object detection and sound source localization. We also show that we can align this detector to ground-truth classes with as little as one label per pseudo-class, and show how our method can learn to detect generic objects that go beyond instruments, such as airplanes and cats.
Triantafyllos Afouras, Yuki Markus Asano, Francois Fagan, Andrea Vedaldi, Florian Metze
CVPR5
2022 Normalized Contrastive Learning for Text-Video Retrieval
abstract
Cross-modal contrastive learning has led the recent advances in multimodal retrieval with its simplicity and effectiveness.In this work, however, we reveal that cross-modal contrastive learning suffers from incorrect normalization of the sum retrieval probabilities of each text or video instance.Specifically, we show that many test instances are either overor under-represented during retrieval, significantly hurting the retrieval performance.To address this problem, we propose Normalized Contrastive Learning (NCL) which utilizes the Sinkhorn-Knopp algorithm to compute the instance-wise biases that properly normalize the sum retrieval probabilities of each instance so that every text and video instance is fairly represented during cross-modal retrieval.Empirical study shows that NCL brings consistent and significant gains in text-video retrieval on different model architectures, with new stateof-the-art multimodal retrieval metrics on the ActivityNet, MSVD, and MSR-VTT datasets without any architecture engineering.
Yookoon Park, Mahmoud Azab, Seungwhan Moon, Florian Metze, Gourab Kundu, Ahmed Kirmani
EMNLP5
2022 On Adversarial Robustness Of Large-Scale Audio Visual Learning
abstract
As audio-visual systems are being deployed for safety-critical tasks such as surveillance and malicious content filtering, their robustness remains an under-studied area. Existing published work on robustness either does not scale to large-scale dataset, or does not deal with multiple modalities. This work aims to study several key questions related to multi-modal learning through the lens of robustness: 1) Are multi-modal models necessarily more robust than uni-modal models? 2) How to efficiently measure the robustness of multi-modal learning? 3) How to fuse different modalities to achieve a more robust multi-modal model? To understand the robustness of the multi-modal model in a large-scale setting, we propose a density-based metric, and a convexity metric to efficiently measure the distribution of each modality in high-dimensional latent space. Our work provides a theoretical intuition together with empirical evidence showing how multi-modal fusion affects adversarial robustness through these metrics. We further devise a mix-up strategy based on our metrics to improve the robustness of the trained model. Our experiments on AudioSet [1] and Kinetics-Sounds [2] verify our hypothesis that multi-modal models are not necessarily more robust than their uni-modal counterparts in the face of adversarial examples. We also observe our mix-up trained method could achieve as much protection as traditional adversarial training, offering a computationally cheap alternative.
Juncheng Li 0001, Shuhui Qu, Po-Yao Huang 0001, Florian Metze
ICASSP5
2022 End-to-End Speech Summarization Using Restricted Self-Attention
abstract
Speech summarization is typically performed by using a cascade of speech recognition and text summarization models. End-to-end modeling of speech summarization models is challenging due to memory and compute constraints arising from long input audio sequences. Recent work in document summarization has inspired methods to reduce the complexity of self-attentions, which enables transformer models to handle long sequences. In this work, we introduce a single model optimized end-to-end for speech summarization. We apply the restricted self-attention technique from text-based models to speech models to address the memory and compute constraints. We demonstrate that the proposed model learns to directly summarize speech for the How-2 corpus of instructional videos. The proposed end-to-end model outperforms the previously proposed cascaded model by 3 points absolute on ROUGE. Further, we consider the spoken language understanding task of predicting concepts from speech inputs and show that the proposed end-to-end model outperforms the cascade model by 4 points absolute F-1.
Shruti Palaskar, Alan W. Black, Florian Metze
ICASSP4
2022 AudioTagging Done Right: 2nd comparison of deep learning methods for environmental sound classification
abstract
After its sweeping success in vision and language tasks, pure attention-based neural architectures (e.g.DeiT) [1] are emerging to the top of audio tagging (AT) leaderboards [2], which seemingly obsoletes traditional convolutional neural networks (CNNs), feed-forward networks or recurrent networks.However, taking a closer look, there is great variability in published research, for instance, performances of models initialized with pretrained weights differ drastically from without pretraining [2], training time for a model varies from hours to weeks, and often, essences are hidden in seemingly trivial details.This urgently calls for a comprehensive study since our 1st comparison [3] is half-decade old.In this work, we perform extensive experiments on AudioSet [4] which is the largest weakly-labeled sound event dataset available, we also did analysis based on the data quality and efficiency.We compare a few state-of-the-art baselines on the AT task, and study the performance and efficiency of 2 major categories of neural architectures: CNN variants and attention-based variants.We also closely examine their optimization procedures.Our opensourced experimental results 1 provide insights to trade off between performance, efficiency, optimization process, for both practitioners and researchers. 2
Juncheng Li 0001, Shuhui Qu, Po-Yao Huang 0001, Florian Metze
INTERSPEECH4
2022 ASR2K: Speech Recognition for Around 2000 Languages without Audio
Florian Metze, David R. Mortensen, Alan W. Black, Shinji Watanabe 0001
INTERSPEECH2
2022 Phone Inventories and Recognition for Every Language
abstract
Identifying phone inventories is a crucial component in language documentation and the preservation of endangered languages. However, even the largest collection of phone inventory only covers about 2000 languages, which is only 1/4 of the total number of languages in the world. A majority of the remaining languages are endangered. In this work, we attempt to solve this problem by estimating the phone inventory for any language listed in Glottolog, which contains phylogenetic information regarding 8000 languages. In particular, we propose one probabilistic model and one non-probabilistic model, both using phylogenetic trees (“language family trees”) to measure the distance between languages. We show that our best model outperforms baseline models by 6.5 F1. Furthermore, we demonstrate that, with the proposed inventories, the phone recognition model can be customized for every language in the set, which improved the PER (phone error rate) in phone recognition by 25%.
Florian Metze, David R. Mortensen, Alan W. Black, Shinji Watanabe 0001
LREC2
2022 Masked Autoencoders that Listen
abstract
This paper studies a simple extension of image-based Masked Autoencoders (MAE) to self-supervised representation learning from audio spectrograms. Following the Transformer encoder-decoder design in MAE, our Audio-MAE first encodes audio spectrogram patches with a high masking ratio, feeding only the non-masked tokens through encoder layers. The decoder then re-orders and decodes the encoded context padded with mask tokens, in order to reconstruct the input spectrogram. We find it beneficial to incorporate local window attention in the decoder, as audio spectrograms are highly correlated in local time and frequency bands. We then fine-tune the encoder with a lower masking ratio on target datasets. Empirically, Audio-MAE sets new state-of-the-art performance on six audio and speech classification tasks, outperforming other recent models that use external supervised pre-training. Our code and models is available at https://github.com/facebookresearch/AudioMAE.
Po-Yao Huang 0001, Hu Xu 0001, Juncheng Li 0001, Alexei Baevski, Michael Auli, Wojciech Galuba, Florian Metze, Christoph Feichtenhofer
NeurIPS7
2021 How2Sign: A Large-Scale Multimodal Dataset for Continuous American Sign Language
abstract
One of the factors that have hindered progress in the areas of sign language recognition, translation, and production is the absence of large annotated datasets. Towards this end, we introduce How2Sign, a multimodal and multiview continuous American Sign Language (ASL) dataset, consisting of a parallel corpus of more than 80 hours of sign language videos and a set of corresponding modalities including speech, English transcripts, and depth. A three-hour subset was further recorded in the Panoptic studio enabling detailed 3D pose estimation. To evaluate the potential of How2Sign for real-world impact, we conduct a study with ASL signers and show that synthesized videos using our dataset can indeed be understood. The study further gives insights on challenges that computer vision should address in order to make progress in this field.
Amanda Cardoso Duarte, Shruti Palaskar, Lucas Ventura, Deepti Ghadiyaram, Kenneth DeHaan, Florian Metze, Jordi Torres, Xavier Giró-i-Nieto
CVPR6
2021 NoiseQA: Challenge Set Evaluation for User-Centric Question Answering
abstract
Abhilasha Ravichander, Siddharth Dalmia, Maria Ryskina, Florian Metze, Eduard Hovy, Alan W Black. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021.
Abhilasha Ravichander, Siddharth Dalmia, Maria Ryskina, Florian Metze, Eduard H. Hovy, Alan W. Black
EACL4
2021 VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding
abstract
Hu Xu, Gargi Ghosh, Po-Yao Huang, Dmytro Okhonko, Armen Aghajanyan, Florian Metze, Luke Zettlemoyer, Christoph Feichtenhofer. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021.
Hu Xu 0001, Gargi Ghosh, Po-Yao Huang 0001, Dmytro Okhonko, Armen Aghajanyan, Florian Metze, Luke Zettlemoyer, Christoph Feichtenhofer
EMNLP (1)6
2021 Phone Distribution Estimation for Low Resource Languages
Juncheng Li 0001, Jiali Yao, Alan W. Black, Florian Metze
ICASSP5
2021 Multilingual Phonetic Dataset for Low Resource Speech Recognition
abstract
Phone Recognition is one of the most important tasks in the field of multilingual speech recognition, especially for low-resource languages whose orthographies are not available. However, most speech recognition datasets so far only focus on high-resource languages, there are very few datasets available for low-resource languages, especially datasets with detailed phone annotation. In this work, we present a large multilingual phonetic dataset, which is preprocessed and aligned from the UCLA phonetic dataset. The dataset contains around 100 low-resource languages and 7000 utterances in total. This dataset would provide an ideal training/evaluation set for universal phone recognition.
David R. Mortensen, Florian Metze, Alan W. Black
ICASSP3
2021 Audio-Visual Event Recognition Through the Lens of Adversary
abstract
As audio/visual classification models are widely deployed for sensitive tasks like content filtering at scale, it is critical to understand their robustness along with improving the accuracy. This work aims to study several key questions related to multimodal learning through the lens of adversarial noises: 1) The trade-off between early/middle/late fusion affecting its robustness and accuracy 2) How does different frequency/time domain features contribute to the robustness? 3) How does different neural modules contribute against the adversarial noise? In our experiment, we construct adversarial examples to attack state-of-the-art neural models trained on Google AudioSet.[1]1We compare how much attack potency in terms of adversarial perturbation of size using different Lpnorms we would need to "deactivate" the victim model. Using adversarial noise to ablate multimodal models, we are able to provide insights into what is the best potential fusion strategy to balance the model parameters/accuracy and robustness trade-off, and distinguish the robust features versus the non-robust features that various neural networks model tend to learn.
Juncheng Li 0001, Kaixin Ma, Shuhui Qu, Po-Yao Huang 0001, Florian Metze
ICASSP5
2021 Space-Time Crop & Attend: Improving Cross-modal Video Representation Learning
abstract
The quality of the image representations obtained from self-supervised learning depends strongly on the type of data augmentations used in the learning formulation. Recent papers have ported these methods from still images to videos and found that leveraging both audio and video signals yields strong gains; however, they did not find that spatial augmentations such as cropping, which are very important for still images, work as well for videos. In this paper, we improve these formulations in two ways unique to the spatio-temporal aspect of videos. First, for space, we show that spatial augmentations such as cropping do work well for videos too, but that previous implementations, due to the high processing and memory cost, could not do this at a scale sufficient for it to work well. To address this issue, we first introduce Feature Crop, a method to simulate such augmentations much more efficiently directly in feature space. Second, we show that as opposed to naïve average pooling, the use of transformer-based attention improves performance significantly, and is well suited for processing feature crops. Combining both of our discoveries into a new method, Space-Time Crop & Attend (STiCA) we achieve state-of-the-art performance across multiple video-representation learning benchmarks. In particular, we achieve new state-of-the-art accuracies of 67.0% on HMDB-51 and 93.1% on UCF-101 when pre-training on Kinetics-400. Code and pretrained models are available1.
Mandela Patrick, Po-Yao Huang 0001, Ishan Misra, Florian Metze, Andrea Vedaldi, Yuki Markus Asano, João F. Henriques
ICCV4
2021 Support-set bottlenecks for video-text representation learning
Mandela Patrick, Po-Yao Huang 0001, Yuki Markus Asano, Florian Metze, Alex Hauptmann 0001, João F. Henriques, Andrea Vedaldi
ICLR4
2021 Rethinking End-to-End Evaluation of Decomposable Tasks: A Case Study on Spoken Language Understanding
abstract
Decomposable tasks are complex and comprise of a hierarchy of sub-tasks. Spoken intent prediction, for example, combines automatic speech recognition and natural language understanding. Existing benchmarks, however, typically hold out examples for only the surface-level sub-task. As a result, models with similar performance on these benchmarks may have unobserved performance differences on the other sub-tasks. To allow insightful comparisons between competitive end-to-end architectures, we propose a framework to construct robust test sets using coordinate ascent over sub-task specific utility functions. Given a dataset for a decomposable task, our method optimally creates a test set for each sub-task to individually assess sub-components of the end-to-end model. Using spoken language understanding as a case study, we generate new splits for the Fluent Speech Commands and Snips SmartLights datasets. Each split has two test sets: one with held-out utterances assessing natural language understanding abilities, and one with held-out speakers to test speech processing skills. Our splits identify performance gaps up to 10% between end-to-end systems that were within 1% of each other on the original test sets. These performance gaps allow more realistic and actionable comparisons between different architectures, driving future model development. We release our splits and tools for the community.
Siddhant Arora, Alissa Ostapenko, Vijay Viswanathan 0002, Siddharth Dalmia, Florian Metze, Shinji Watanabe 0001, Alan W. Black
Interspeech5
2021 Hierarchical Phone Recognition with Compositional Phonetics
Juncheng Li 0001, Florian Metze, Alan W. Black
Interspeech3
2021 Multimodal Speech Summarization Through Semantic Concept Learning
Shruti Palaskar, Ruslan Salakhutdinov, Alan W. Black, Florian Metze
Interspeech4
2021 Differentiable Allophone Graphs for Language-Universal Speech Recognition
abstract
Building language-universal speech recognition systems entails producing phonological units of spoken sound that can be shared across languages.While speech annotations at the language-specific phoneme or surface levels are readily available, annotations at a universal phone level are relatively rare and difficult to produce.In this work, we present a general framework to derive phone-level supervision from only phonemic transcriptions and phone-to-phoneme mappings with learnable weights represented using weighted finite-state transducers, which we call differentiable allophone graphs.By training multilingually, we build a universal phone-based speech recognition model with interpretable probabilistic phone-to-phoneme mappings for each language.These phone-based systems with learned allophone graphs can be used by linguists to document new languages, build phone-based lexicons that capture rich pronunciation variations, and re-evaluate the allophone mappings of seen language.We demonstrate the aforementioned benefits of our proposed framework with a system trained on 7 diverse languages.
Brian Yan, Siddharth Dalmia, David R. Mortensen, Florian Metze, Shinji Watanabe 0001
Interspeech4
2021 Searchable Hidden Intermediates for End-to-End Models of Decomposable Sequence Tasks
abstract
Siddharth Dalmia, Brian Yan, Vikas Raunak, Florian Metze, Shinji Watanabe. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Siddharth Dalmia, Brian Yan, Vikas Raunak, Florian Metze, Shinji Watanabe 0001
NAACL-HLT4
2021 Multilingual Multimodal Pre-training for Zero-Shot Cross-Lingual Transfer of Vision-Language Models
abstract
Po-Yao Huang, Mandela Patrick, Junjie Hu, Graham Neubig, Florian Metze, Alexander Hauptmann. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Po-Yao Huang 0001, Mandela Patrick, Junjie Hu 0001, Graham Neubig, Florian Metze, Alex Hauptmann 0001
NAACL-HLT5
2021 Keeping Your Eye on the Ball: Trajectory Attention in Video Transformers
abstract
In video transformers, the time dimension is often treated in the same way as the two spatial dimensions. However, in a scene where objects or the camera may move, a physical point imaged at one location in frame $t$ may be entirely unrelated to what is found at that location in frame $t+k$. These temporal correspondences should be modeled to facilitate learning about dynamic scenes. To this end, we propose a new drop-in block for video transformers - trajectory attention - that aggregates information along implicitly determined motion paths. We additionally propose a new method to address the quadratic dependence of computation and memory on the input size, which is particularly important for high resolution or long videos. While these ideas are useful in a range of settings, we apply them to the specific task of video action recognition with a transformer model and obtain state-of-the-art results on the Kinetics, Something-Something V2, and Epic-Kitchens datasets.
Mandela Patrick, Dylan Campbell, Yuki Markus Asano, Ishan Misra, Florian Metze, Christoph Feichtenhofer, Andrea Vedaldi, João F. Henriques
NeurIPS5
2020 Towards Zero-Shot Learning for Automatic Phonemic Transcription
abstract
Automatic phonemic transcription tools are useful for low-resource language documentation. However, due to the lack of training sets, only a tiny fraction of languages have phonemic transcription tools. Fortunately, multilingual acoustic modeling provides a solution given limited audio training data. A more challenging problem is to build phonemic transcribers for languages with zero training data. The difficulty of this task is that phoneme inventories often differ between the training languages and the target language, making it infeasible to recognize unseen phonemes. In this work, we address this problem by adopting the idea of zero-shot learning. Our model is able to recognize unseen phonemes in the target language without any training data. In our model, we decompose phonemes into corresponding articulatory attributes such as vowel and consonant. Instead of predicting phonemes directly, we first predict distributions over articulatory attributes, and then compute phoneme distributions with a customized acoustic model. We evaluate our model by training it using 13 languages and testing it using 7 unseen languages. We find that it achieves 7.7% better phoneme error rate on average over a standard multilingual model.
Siddharth Dalmia, David R. Mortensen, Juncheng Li 0001, Alan W. Black, Florian Metze
AAAI6
2020 Universal Phone Recognition with a Multilingual Allophone System
abstract
Multilingual models can improve language processing, particularly for low resource situations, by sharing parameters across languages. Multilingual acoustic models, however, generally ignore the difference between phonemes (sounds that can support lexical contrasts in a particular language) and their corresponding phones (the sounds that are actually spoken, which are language independent). This can lead to performance degradation when combining a variety of training languages, as identically annotated phonemes can actually correspond to several different underlying phonetic realizations. In this work, we propose a joint model of both language-independent phone and language-dependent phoneme distributions. In multilingual ASR experiments over 11 languages, we find that this model improves testing performance by 2% phoneme error rate absolute in low-resource conditions. Additionally, because we are explicitly modeling language-independent phones, we can build a (nearly-)universal phone recognizer that, when combined with the PHOIBLE [1] large, manually curated database of phone inventories, can be customized into 2,000 language dependent recognizers. Experiments on two low-resourced indigenous languages, Inuktitut and Tusom, show that our recognizer achieves phone accuracy improvements of more than 17%, moving a step closer to speech recognition for all languages in the world.1
Siddharth Dalmia, Juncheng Li 0001, Matthew Lee 0012, Patrick Littell, Jiali Yao, Antonios Anastasopoulos, David R. Mortensen, Graham Neubig, Alan W. Black, Florian Metze
ICASSP11
2020 ASR Error Correction and Domain Adaptation Using Machine Translation
abstract
Off-the-shelf pre-trained Automatic Speech Recognition (ASR) systems are an increasingly viable service for companies of any size building speech-based products. While these ASR systems are trained on large amounts of data, domain mismatch is still an issue for many such parties that want to use this service as-is leading to not so optimal results for their task. We propose a simple technique to perform domain adaptation for ASR error correction via machine translation. The machine translation model is a strong candidate to learn a mapping from out-of-domain ASR errors to in-domain terms in the corresponding reference files. We use two off-the-shelf ASR systems in this work: Google ASR (commercial) and the ASPIRE model (open-source). We observe 7% absolute improvement in word error rate and 4 point absolute improvement in BLEU score in Google ASR output via our proposed method. We also evaluate ASR error correction via a downstream task of Speaker Diarization that captures speaker style, syntax, structure and semantic improvements we obtain via ASR correction.
Anirudh Mani, Shruti Palaskar, Nimshi Venkat Meripo, Sandeep Konam, Florian Metze
ICASSP5
2020 Looking Enhances Listening: Recovering Missing Speech Using Images
abstract
Speech is understood better by using visual context; for this reason, there have been many attempts to use images to adapt automatic speech recognition (ASR) systems. Current work, however, has shown that visually adapted ASR models only use images as a regularization signal, while completely ignoring their semantic content. In this paper, we present a set of experiments where we show the utility of the visual modality under noisy conditions. Our results show that multimodal ASR models can recover words which are masked in the input acoustic signal, by grounding its transcriptions using the visual representations. We observe that integrating visual context can result in up to 35% relative improvement in masked word recovery. These results demonstrate that end-to-end multimodal ASR systems can become more robust to noise by leveraging the visual context.
Tejas Srinivasan, Ramon Sanabria, Florian Metze
ICASSP3
2020 Contextual RNN-T for Open Domain ASR
abstract
End-to-end (E2E) systems for automatic speech recognition (ASR), such as RNN Transducer (RNN-T) and Listen-Attend-Spell (LAS) blend the individual components of a traditional hybrid ASR system - acoustic model, language model, pronunciation model - into a single neural network. While this has some nice advantages, it limits the system to be trained using only paired audio and text. Because of this, E2E models tend to have difficulties with correctly recognizing rare words that are not frequently seen during training, such as entity names. In this paper, we propose modifications to the RNN-T model that allow the model to utilize additional metadata text with the objective of improving performance on these named entity words. We evaluate our approach on an in-house dataset sampled from de-identified public social media videos, which represent an open domain ASR task. By using an attention model and a biasing model to leverage the contextual metadata that accompanies a video, we observe a relative improvement of about 16% in Word Error Rate on Named Entities (WER-NE) for videos with related metadata.
Mahaveer Jain, Gil Keren, Jay Mahadeokar, Geoffrey Zweig, Florian Metze, Yatharth Saraf
INTERSPEECH5
2020 Towards Context-Aware End-to-End Code-Switching Speech Recognition
Zimeng Qiu, Yiyuan Li, Florian Metze, William M. Campbell
INTERSPEECH4
2020 AlloVera: A Multilingual Allophone Database
abstract
We introduce a new resource, AlloVera, which provides mappings from 218 allophones to phonemes for 14 languages. Phonemes are contrastive phonological units, and allophones are their various concrete realizations, which are predictable from phonological context. While phonemic representations are language specific, phonetic representations (stated in terms of (allo)phones) are much closer to a universal (language-independent) transcription. AlloVera allows the training of speech recognition models that output phonetic transcriptions in the International Phonetic Alphabet (IPA), regardless of the input language. We show that a “universal” allophone model, Allosaurus, built with AlloVera, outperforms “universal” phonemic models and language-specific models on a speech-transcription task. We explore the implications of this technology (and related technologies) for the documentation of endangered and minority languages. We further explore other applications for which AlloVera will be suitable as it grows, including phonological typology.
David R. Mortensen, Patrick Littell, Alexis Michaud, Shruti Rijhwani, Antonios Anastasopoulos, Alan W. Black, Florian Metze, Graham Neubig
LREC8
2020 Transfer learning for multimodal dialog
abstract
Audio-Visual Scene-Aware Dialog (AVSD) is best understood as an extension of Visual Question Answering, the task of generating a textual answer in response to a textual question on multi-media content. In AVSD, the answer-relevant “context” is expanded to include past dialog turns, which we view as a specialized form of extra textual knowledge (in addition to the standard video features). We have developed a framework that uses hierarchical attention to fuse contributions from different modalities, and had shown how it can be used to generate textual summaries from multi-modal sources, specifically videos with accompanying commentary. In this paper, we transfer the algorithmic approach, models, and data from this background corpus of 2000 h of how-to videos to the AVSD task, and report our findings. Our approach uses dialog context, but makes no assumption about the ordering of the history. Our system achieves the best performance in both automatic and human evaluations in the 7th Dialog State Tracking Challenge (AVSD).
Shruti Palaskar, Ramon Sanabria, Florian Metze
Comput. Speech Lang.3
2020 Speech Technology for Unwritten Languages
abstract
Speech technology plays an important role in our everyday life. Among others, speech is used for human-computer interaction, for instance for information retrieval and on-line shopping. In the case of an unwritten language, however, speech technology is unfortunately difficult to create, because it cannot be created by the standard combination of pre-trained speech-to-text and text-to-speech subsystems. The research presented in this article takes the first steps towards speech technology for unwritten languages. Specifically, the aim of this work was 1) to learn speech-to-meaning representations without using text as an intermediate representation, and 2) to test the sufficiency of the learned representations to regenerate speech or translated text, or to retrieve images that depict the meaning of an utterance in an unwritten language. The results suggest that building systems that go directly from speech-to-meaning and from meaning-to-speech, bypassing the need for text, is possible.
Odette Scharenborg, Lucas Ondel Yang, Shruti Palaskar, Philip Arthur, Francesco Ciannella, Mingxing Du, Elin Larsen, Danny Merkx, Rachid Riad, Liming Wang 0003, Emmanuel Dupoux, Laurent Besacier, Alan W. Black, Mark Hasegawa-Johnson, Florian Metze, Graham Neubig, Sebastian Stüker, Pierre Godard, Markus Müller 0001
IEEE ACM Trans. Audio Speech Lang. Process.15
2020 Machine Listening for Heart Status Monitoring: Introducing and Benchmarking HSS - The Heart Sounds Shenzhen Corpus
abstract
Auscultation of the heart is a widely studied technique, which requires precise hearing from practitioners as a means of distinguishing subtle differences in heart-beat rhythm. This technique is popular due to its non-invasive nature, and can be an early diagnosis aid for a range of cardiac conditions. Machine listening approaches can support this process, monitoring continuously and allowing for a representation of both mild and chronic heart conditions. Despite this potential, relevant databases and benchmark studies are scarce. In this paper, we introduce our publicly accessible database, the Heart Sounds Shenzhen Corpus (HSS), which was first released during the recent INTERSPEECH 2018 ComParE Heart Sound sub-challenge. Additionally, we provide a survey of machine learning work in the area of heart sound recognition, as well as a benchmark for HSS utilising standard acoustic features and machine learning models. At best our support vector machine with Log Mel features achieves 49.7% unweighted average recall on a three category task (normal, mild, moderate/severe).
Fengquan Dong, Kun Qian 0003, Zhao Ren, Alice Baird, Zhenyu Dai, Florian Metze, Yoshiharu Yamamoto, Björn W. Schuller
IEEE J. Biomed. Health Informatics8
2019 Gated Embeddings in End-to-End Speech Recognition for Conversational-Context Fusion
abstract
We present a novel conversational-context aware end-to-end speech recognizer based on a gated neural network that incorporates conversational-context/word/speech embeddings.Unlike conventional speech recognition models, our model learns longer conversational-context information that spans across sentences and is consequently better at recognizing long conversations.Specifically, we propose to use text-based external word and/or sentence embeddings (i.e., fast-Text, BERT) within an end-to-end framework, yielding significant improvement in word error rate with better conversational-context representation.We evaluated the models on the Switchboard conversational speech corpus and show that our model outperforms standard end-to-end speech recognition models.
Suyoun Kim, Siddharth Dalmia, Florian Metze
ACL (1)3
2019 Multimodal Abstractive Summarization for How2 Videos
abstract
In this paper, we study abstractive summarization for open-domain videos.Unlike the traditional text news summarization, the goal is less to "compress" text information but rather to provide a fluent textual summary of information that has been collected and fused from different source modalities, in our case video and audio transcripts (or text).We show how a multi-source sequence-to-sequence model with hierarchical attention can integrate information from different modalities into a coherent output, compare various models trained with different modalities and present pilot experiments on the How2 corpus of instructional videos.We also propose a new evaluation metric (Content F1) for abstractive summarization task that measures semantic adequacy rather than fluency of the summaries, which is covered by metrics like ROUGE and BLEU.
Shruti Palaskar, Jindrich Libovický, Spandana Gella, Florian Metze
ACL (1)4
2019 A Comparison of Five Multiple Instance Learning Pooling Functions for Sound Event Detection with Weak Labeling
abstract
Sound event detection (SED) entails two subtasks: recognizing what types of sound events are present in an audio stream (audio tagging), and pinpointing their onset and offset times (localization). In the popular multiple instance learning (MIL) framework for SED with weak labeling, an important component is the pooling function. This paper compares five types of pooling functions both theoretically and experimentally, with special focus on their performance of localization. Although the attention pooling function is currently receiving the most attention, we find the linear softmax pooling function to perform the best among the five. Using this pooling function, we build a neural network called TALNet. It is the first system to reach state-of-the-art audio tagging performance on Audio Set, while exhibiting strong localization performance on the DCASE 2017 challenge at the same time.
Yun Wang 0005, Juncheng Li 0001, Florian Metze
ICASSP3
2019 Connectionist Temporal Localization for Sound Event Detection with Sequential Labeling
abstract
Research on sound event detection (SED) with weak labeling has mostly focused on presence/absence labeling, which provides no temporal information at all about the event occurrences. In this paper, we consider SED with sequential labeling, which specifies the temporal order of the event boundaries. The conventional connectionist temporal classification (CTC) framework, when applied to SED with sequential labeling, does not localize long events well due to a "peak clustering" problem. We adapt the CTC framework and propose connectionist temporal localization (CTL), which successfully solves the problem. Evaluation on a subset of Audio Set shows that CTL closes a third of the gap between presence/ absence labeling and strong labeling, demonstrating the usefulness of the extra temporal information in sequential labeling. CTL also makes it easy to combine sequential labeling with presence/absence labeling and strong labeling.
Yun Wang 0005, Florian Metze
ICASSP2
2019 Multimodal Grounding for Sequence-to-sequence Speech Recognition
abstract
Humans are capable of processing speech by making use of multiple sensory modalities. For example, the environment where a conversation takes place generally provides semantic and/or acoustic context that helps us to resolve ambiguities or to recall named entities. Motivated by this, there have been many works studying the integration of visual information into the speech recognition pipeline. Specifically, in our previous work, we propose a multistep visual adaptive training approach which improves the accuracy of an audio-based Automatic Speech Recognition (ASR) system. This approach, however, is not end-to-end as it requires fine-tuning the whole model with an adaptation layer. In this paper, we propose novel end-to-end multimodal ASR systems and compare them to the adaptive approach by using a range of visual representations obtained from state-of-the-art convolutional neural networks. We show that adaptive training is effective for S2S models leading to an absolute improvement of 1.4% in word error rate. As for the end-to-end systems, although they perform better than baseline, the improvements are slightly less than adaptive training, 0.8 absolute WER reduction in single-best models. Using ensemble decoding, end-to-end models reach a WER of 15% which is the lowest score among all systems.
Ozan Caglayan, Ramon Sanabria, Shruti Palaskar, Loïc Barrault, Florian Metze
ICASSP5
2019 Phoneme Level Language Models for Sequence Based Low Resource ASR
abstract
Building multilingual and crosslingual models help bring different languages together in a language universal space. It allows models to share parameters and transfer knowledge across languages, enabling faster and better adaptation to a new language. These approaches are particularly useful for low resource languages. In this paper, we propose a phoneme-level language model that can be used multilingually and for crosslingual adaptation to a target language. We show that our model performs almost as well as the monolingual models by using six times fewer parameters, and is capable of better adaptation to languages not seen during training in a low resource scenario. We show that these phoneme-level language models can be used to decode sequence based Connectionist Temporal Classification (CTC) acoustic model outputs to obtain comparable word error rates with Weighted Finite State Transducer (WFST) based decoding in Babel languages. We also show that these phoneme-level language models outperform WFST decoding in various low-resource conditions like adapting to a new language and domain mismatch between training and testing data.
Siddharth Dalmia, Alan W. Black, Florian Metze
ICASSP4
2019 Learning from Multiview Correlations in Open-domain Videos
abstract
An increasing number of datasets contain multiple views, such as video, sound and automatic captions. A basic challenge in representation learning is how to leverage multiple views to learn better representations. This is further complicated by the existence of a latent alignment between views, such as between speech and its transcription, and by the multitude of choices for the learning objective. We explore an advanced, correlation-based representation learning method on a 4-way parallel, multimodal dataset, and assess the quality of the learned representations on retrieval-based tasks. We show that the proposed approach produces rich representations that capture most of the information shared across views. Our best models for speech and textual modalities achieve retrieval rates from 70.7% to 96.9% on open-domain, user-generated instructional videos. This shows it is possible to learn reliable representations across disparate, unaligned and noisy modalities, and encourages using the proposed approach on larger datasets.
Nils Holzenberger, Shruti Palaskar, Pranava Swaroop Madhyastha, Florian Metze, Raman Arora
ICASSP4
2019 Learned in Speech Recognition: Contextual Acoustic Word Embeddings
abstract
End-to-end acoustic-to-word speech recognition models have recently gained popularity because they are easy to train, scale well to large amounts of training data, and do not require a lexicon. In addition, word models may also be easier to integrate with downstream tasks such as spoken language understanding, because inference (search) is much simplified compared to phoneme, character or any other sort of sub-word units. In this paper, we describe methods to construct contextual acoustic word embeddings directly from a supervised sequence-to-sequence acoustic-to-word speech recognition model using the learned attention distribution. On a suite of 16 standard sentence evaluation tasks, our embeddings show competitive performance against a word2vec model trained on the speech transcriptions. In addition, we evaluate these embeddings on a spoken language understanding task, and observe that our embeddings match the performance of text-based embeddings in a pipeline of first performing speech recognition and then constructing word embeddings from transcriptions.
Shruti Palaskar, Vikas Raunak, Florian Metze
ICASSP3
2019 On Leveraging the Visual Modality for Neural Machine Translation
abstract
Leveraging the visual modality effectively for Neural Machine Translation (NMT) remains an open problem in computational linguistics.Recently, Caglayan et al. posit that the observed gains are limited mainly due to the very simple, short, repetitive sentences of the Multi30k dataset (the only multimodal MT dataset available at the time), which renders the source text sufficient for context.In this work, we further investigate this hypothesis on a new large scale multimodal Machine Translation (MMT) dataset, How2, which has 1.57 times longer mean sentence length than Multi30k and no repetition.We propose and evaluate three novel fusion techniques, each of which is designed to ensure the utilization of visual context at different stages of the Sequence-to-Sequence transduction pipeline, even under full linguistic context.However, we still obtain only marginal gains under full linguistic context and posit that visual embeddings extracted from deep vision models (ResNet for Multi30k, ResNext for How2) do not lend themselves to increasing the discriminativeness between the vocabulary elements at token level prediction in NMT.We demonstrate this qualitatively by analyzing attention distribution and quantitatively through Principal Component Analysis, arriving at the conclusion that it is the quality of the visual embeddings rather than the length of sentences, which need to be improved in existing MMT datasets.
Vikas Raunak, Sang Keun Choe, Quanyang Lu, Florian Metze
INLG5
2019 Cross-Attention End-to-End ASR for Two-Party Conversations
abstract
We present an end-to-end speech recognition model that learns interaction between two speakers based on the turn-changing information. Unlike conventional speech recognition models, our model exploits two speakers' history of conversational-context information that spans across multiple turns within an end-to-end framework. Specifically, we propose a speaker-specific cross-attention mechanism that can look at the output of the other speaker side as well as the one of the current speaker for better at recognizing long conversations. We evaluated the models on the Switchboard conversational speech corpus and show that our model outperforms standard end-to-end speech recognition models.
Suyoun Kim, Siddharth Dalmia, Florian Metze
INTERSPEECH3
2019 Multilingual Speech Recognition with Corpus Relatedness Sampling
abstract
Multilingual acoustic models have been successfully applied to low-resource speech recognition.Most existing works have combined many small corpora together, and pretrained a multilingual model by sampling from each corpus uniformly.The model is eventually fine-tuned on each target corpus.This approach, however, fails to exploit the relatedness and similarity among corpora in the training set.For example, the target corpus might benefit more from a corpus in the same domain or a corpus from a close language.In this work, we propose a simple but useful sampling strategy to take advantage of this relatedness.We first compute the corpus-level embeddings and estimate the similarity between each corpus.Next we start training the multilingual model with uniform-sampling from each corpus at first, then we gradually increase the probability to sample from related corpora based on its similarity with the target corpus.Finally the model would be fine-tuned automatically on the target corpus.Our sampling strategy outperforms the baseline multilingual model on 16 low-resource tasks.Additionally, we demonstrate that our corpus embeddings capture the language and domain information of each corpus.
Siddharth Dalmia, Alan W. Black, Florian Metze
INTERSPEECH4
2019 SANTLR: Speech Annotation Toolkit for Low Resource Languages
Zhong Zhou, Siddharth Dalmia, Alan W. Black, Florian Metze
INTERSPEECH5
2019 Survey Talk: Multimodal Processing of Speech and Language
Florian Metze
INTERSPEECH1
2019 Adversarial Music: Real world Audio Adversary against Wake-word Detection System
abstract
Voice Assistants (VAs) such as Amazon Alexa or Google Assistant rely on wake-word detection to respond to people's commands, which could potentially be vulnerable to audio adversarial examples. In this work, we target our attack on the wake-word detection system. Our goal is to jam the model with some inconspicuous background music to deactivate the VAs while our audio adversary is present. We implemented an emulated wake-word detection system of Amazon Alexa based on recent publications. We validated our models against the real Alexa in terms of wake-word detection accuracy. Then we computed our audio adversaries with consideration of expectation over transform and we implemented our audio adversary with a differentiable synthesizer. Next we verified our audio adversaries digitally on hundreds of samples of utterances collected from the real world. Our experiments show that we can effectively reduce the recognition F1 score of our emulated model from 93.4% to 11.0%. Finally, we tested our audio adversary over the air, and verified it works effectively against Alexa, reducing its F1 score from 92.5% to 11.0%. To the best of our knowledge, this is the first real-world adversarial attack against a commercial grade VA wake-word detection system. Our demo video is included in the supplementary material.
Juncheng Li 0001, Shuhui Qu, Joseph Szurley, J. Zico Kolter, Florian Metze
NeurIPS6
2019 Automatic word count estimation from daylong child-centered recordings in various language environments using language-independent syllabification of speech
abstract
Automatic word count estimation (WCE) from audio recordings can be used to quantify the amount of verbal communication in a recording environment. One key application of WCE is to measure language input heard by infants and toddlers in their natural environments, as captured by daylong recordings from microphones worn by the infants. Although WCE is nearly trivial for high-quality signals in high-resource languages, daylong recordings are substantially more challenging due to the unconstrained acoustic environments and the presence of near- and far-field speech. Moreover, many use cases of interest involve languages for which reliable ASR systems or even well-defined lexicons are not available. A good WCE system should also perform similarly for low- and high-resource languages in order to enable unbiased comparisons across different cultures and environments. Unfortunately, the current state-of-the-art solution, the LENA system, is based on proprietary software and has only been optimized for American English, limiting its applicability. In this paper, we build on existing work on WCE and present the steps we have taken towards a freely available system for WCE that can be adapted to different languages or dialects with a limited amount of orthographically transcribed speech data. Our system is based on language-independent syllabification of speech, followed by a language-dependent mapping from syllable counts (and a number of other acoustic features) to the corresponding word count estimates. We evaluate our system on samples from daylong infant recordings from six different corpora consisting of several languages and socioeconomic environments, all manually annotated with the same protocol to allow direct comparison. We compare a number of alternative techniques for the two key components in our system: speech activity detection and automatic syllabification of speech. As a result, we show that our system can reach relatively consistent WCE accuracy across multiple corpora and languages (with some limitations). In addition, the system outperforms LENA on three of the four corpora consisting of different varieties of English. We also demonstrate how an automatic neural network-based syllabifier, when trained on multiple languages, generalizes well to novel languages beyond the training data, outperforming two previously proposed unsupervised syllabifiers as a feature extractor for WCE.
Okko Johannes Räsänen, Shreyas Seshadri, Julien Karadayi, Eric Riebling, John P. Bunce, Alejandrina Cristià, Florian Metze, Marisa Casillas, Celia Rosemberg, Elika Bergelson, Melanie Soderstrom
Speech Commun.7
2018 Sequence-Based Multi-Lingual Low Resource Speech Recognition
abstract
Techniques for multi-lingual and cross-lingual speech recognition can help in low resource scenarios, to bootstrap systems and enable analysis of new languages and domains. End-to-end approaches, in particular sequence-based techniques, are attractive because of their simplicity and elegance. While it is possible to integrate traditional multi-lingual bottleneck feature extractors as front-ends, we show that end-to-end multi-lingual training of sequence models is effective on context independent models trained using Connectionist Temporal Classification (CTC) loss. We show that our model improves performance on Babel languages by over 6% absolute in terms of word/phoneme error rate when compared to mono-lingual systems built in the same setting for these languages. We also show that the trained model can be adapted cross-lingually to an unseen language using just 25% of the target data. We show that training on multiple languages is important for very low resource cross-lingual target scenarios, but not for multi-lingual testing scenarios. Here, it appears beneficial to include large well prepared datasets.
Siddharth Dalmia, Ramon Sanabria, Florian Metze, Alan W. Black
ICASSP3
2018 A Light-Weight Multimodal Framework for Improved Environmental Audio Tagging
abstract
The lack of strong labels has severely limited the state-of-the-art fully supervised audio tagging systems to be scaled to larger dataset. Meanwhile, audio-visual learning models based on unlabeled videos have been successfully applied to audio tagging, but they are inevitably resource hungry and require a long time to train. In this work, we propose a light-weight, multimodal framework for environmental audio tagging. The audio branch of the framework is a convolutional and recurrent neural network (CRNN) based on multiple instance learning (MIL). It is trained with the audio tracks of a large collection of weakly labeled YouTube video excerpts; the video branch uses pretrained state-of-the-art image recognition networks and word embeddings to extract information from the video track and to map visual objects to sound events. Experiments on the audio tagging task of the DCASE 2017 challenge show that the incorporation of video information improves a strong baseline audio tagging system by 5.3% in terms of F1score. The entire system can be trained within 6 hours on a single GPU, and can be easily carried over to other audio tasks such as speech sentimental analysis.
Juncheng Li 0001, Yun Wang 0005, Joseph Szurley, Florian Metze, Samarjit Das
ICASSP4
2018 End-to-end Multimodal Speech Recognition
abstract
Transcription or sub-titling of open-domain videos is still a challenging domain for Automatic Speech Recognition (ASR) due to the data's challenging acoustics, variable signal processing and the essentially unrestricted domain of the data. In previous work, we have shown that the visual channel - specifically object and scene features - can help to adapt the acoustic model (AM) and language model (LM) of a recognizer, and we are now expanding this work to end-to-end approaches. In the case of a Connectionist Temporal Classification (CTC)-based approach, we retain the separation of AM and LM, while for a sequence-to-sequence (S2S) approach, both information sources are adapted together, in a single model. This paper also analyzes the behavior of CTC and S2S models on noisy video data (How-To corpus), and compares it to results on the clean Wall Street Journal (WSJ) corpus, providing insight into the robustness of both approaches.
Shruti Palaskar, Ramon Sanabria, Florian Metze
ICASSP3
2018 Enhancement and Analysis of Conversational Speech: JSALT 2017
abstract
Automatic speech recognition is more and more widely and effectively used. Nevertheless, in some automatic speech analysis tasks the state of the art is surprisingly poor. One of these is “diarization”, the task of determining who spoke when. Diarization is key to processing meeting audio and clinical interviews, extended recordings such as police body cam or child language acquisition data, and any other speech data involving multiple speakers whose voices are not cleanly separated into individual channels. Overlapping speech, environmental noise and suboptimal recording techniques make the problem harder. During the JSALT Summer Workshop at CMU in 2017, an international team of researchers worked on several aspects of this problem, including calibration of the state of the art, detection of overlaps, enhancement of noisy recordings, and classification of shorter speech segments. This paper sketches the workshop's results, and announces plans for a “Diarization Challenge” to encourage further progress.
Neville Ryant, Elika Bergelson, Kenneth Church 0001, Alejandrina Cristià, Jun Du 0002, Sriram Ganapathy, Sanjeev Khudanpur, Diana Kowalski, Mahesh Krishnamoorthy, Rajat Kulshreshta, Mark Y. Liberman, Yu-Ding Lu, Matthew Maciejewski, Florian Metze, Ján Profant, Lei Sun 0010, Yu Tsao 0001
ICASSP14
2018 Linguistic Unit Discovery from Multi-Modal Inputs in Unwritten Languages: Summary of the "Speaking Rosetta" JSALT 2017 Workshop
abstract
We summarize the accomplishments of a multi-disciplinary workshop exploring the computational and scientific issues surrounding the discovery of linguistic units (subwords and words) in a language without orthography. We study the replacement of orthographic transcriptions by images and/or translated text in a well-resourced language to help unsupervised discovery from raw speech.
Odette Scharenborg, Laurent Besacier, Alan W. Black, Mark Hasegawa-Johnson, Florian Metze, Graham Neubig, Sebastian Stüker, Pierre Godard, Markus Müller 0001, Lucas Ondel Yang, Shruti Palaskar, Philip Arthur, Francesco Ciannella, Mingxing Du, Elin Larsen, Danny Merkx, Rachid Riad, Liming Wang 0003, Emmanuel Dupoux
ICASSP5
2018 The ACLEW DiViMe: An Easy-to-use Diarization Tool
Adrien Le Franc, Eric Riebling, Julien Karadayi, Yun Wang 0005, Camila Scaff, Florian Metze, Alejandrina Cristià
INTERSPEECH6
2018 Multiple Instance Deep Learning for Weakly Supervised Small-Footprint Audio Event Detection
abstract
State-of-the-art audio event detection (AED) systems rely on supervised learning using strongly labeled data. However, this dependence severely limits scalability to large-scale datasets where fine resolution annotations are too expensive to obtain. In this paper, we propose a small-footprint multiple instance learning (MIL) framework for multi-class AED using weakly annotated labels. The proposed MIL framework uses audio embeddings extracted from a pre-trained convolutional neural network as input features. We show that by using audio embeddings the MIL framework can be implemented using a simple DNN with performance comparable to recurrent neural networks. We evaluate our approach by training an audio tagging system using a subset of AudioSet, which is a large collection of weakly labeled YouTube video excerpts. Combined with a late-fusion approach, we improve the F1 score of a baseline audio tagging system by 17%. We show that audio embeddings extracted by the convolutional neural networks significantly boost the performance of all MIL models. This framework reduces the model complexity of the AED system and is suitable for applications where computational resources are limited.
Shao-Yen Tseng, Juncheng Li 0001, Yun Wang 0005, Florian Metze, Joseph Szurley, Samarjit Das
INTERSPEECH4
2018 Comparing the Max and Noisy-Or Pooling Functions in Multiple Instance Learning for Weakly Supervised Sequence Learning Tasks
abstract
Many sequence learning tasks require the localization of certain events in sequences. Because it can be expensive to obtain strong labeling that specifies the starting and ending times of the events, modern systems are often trained with weak labeling without explicit timing information. Multiple instance learning (MIL) is a popular framework for learning from weak labeling. In a common scenario of MIL, it is necessary to choose a pooling function to aggregate the predictions for the individual steps of the sequences. In this paper, we compare the and pooling functions on a speech recognition task and a sound event detection task. We find that max pooling is able to localize phonemes and sound events, while noisy-or pooling fails. We provide a theoretical explanation of the different behavior of the two pooling functions on sequence learning tasks.
Yun Wang 0005, Juncheng Li 0001, Florian Metze
INTERSPEECH3
2018 Subword and Crossword Units for CTC Acoustic Models
abstract
This paper proposes a novel approach to create an unit set for CTC based speech recognition systems. By using Byte Pair Encoding we learn an unit set of an arbitrary size on a given training text. In contrast to using characters or words as units this allows us to find a good trade-off between the size of our unit set and the available training data. We evaluate both Crossword units, that may span multiple word, and Subword units. By combining this approach with decoding methods using a separate language model we are able to achieve state of the art results for grapheme based CTC systems.
Thomas Zenkel, Ramon Sanabria, Florian Metze, Alex Waibel
INTERSPEECH3
2018 Annotating High-Level Structures of Short Stories and Personal Anecdotes
Boyang Li 0001, Beth Cardier, Tong Wang 0007, Florian Metze
LREC4
2018 Learning Joint Embedding with Multimodal Cues for Cross-Modal Video-Text Retrieval
abstract
Constructing a joint representation invariant across different modalities (e.g., video, language) is of significant importance in many multimedia applications. While there are a number of recent successes in developing effective image-text retrieval methods by learning joint representations, the video-text retrieval task, however, has not been explored to its fullest extent. In this paper, we study how to effectively utilize available multimodal cues from videos for the cross-modal video-text retrieval task. Based on our analysis, we propose a novel framework that simultaneously utilizes multi-modal features (different visual characteristics, audio inputs, and text) by a fusion strategy for efficient retrieval. Furthermore, we explore several loss functions in training the embedding and propose a modified pairwise ranking loss for the task. Experiments on MSVD and MSR-VTT datasets demonstrate that our method achieves significant performance gain compared to the state-of-the-art approaches.
Niluthpol Chowdhury Mithun, Juncheng Li 0001, Florian Metze, Amit K. Roy-Chowdhury
ICMR3
2018 Domain Robust Feature Extraction for Rapid Low Resource ASR Development
abstract
Developing a practical speech recognizer for a low resource language is challenging, not only because of the (potentially unknown) properties of the language, but also because test data may not be from the same domain as the available training data.In this paper, we focus on the latter challenge, i.e. domain mismatch, for systems trained using a sequence-based criterion. We demonstrate the effectiveness of using a pre-trained English recognizer, which is robust to such mismatched conditions, as a domain normalizing feature extractor on a low resource language. In our example, we use Turkish Conversational Speech and Broadcast News data.This enables rapid development of speech recognizers for new languages which can easily adapt to any domain. Testing in various cross-domain scenarios, we achieve relative improvements of around 25% in phoneme error rate, with improvements being around 50% for some domains.
Siddharth Dalmia, Florian Metze, Alan W. Black
SLT3
2018 Dialog-Context Aware end-to-end Speech Recognition
abstract
Existing speech recognition systems are typically built at the sentence level, although it is known that dialog context, e.g. higher-level knowledge that spans across sentences or speakers, can help the processing of long conversations. The recent progress in end-to-end speech recognition systems promises to integrate all available information (e.g. acoustic, language resources) into a single model, which is then jointly optimized. It seems natural that such dialog context information should thus also be integrated into the end-to-end models to improve recognition accuracy further. In this work, we present a dialog-context aware speech recognition model, which explicitly uses context information beyond sentence-level information, in an end-to-end fashion. Our dialog-context model captures a history of sentence-level contexts, so that the whole system can be trained with dialog-context information in an end-to-end manner. We evaluate our proposed approach on the Switchboard conversational speech corpus, and show that our system outperforms a comparable sentence-level end-to-end speech recognition system.
Suyoun Kim, Florian Metze
SLT2
2018 Acoustic-to-Word Recognition with Sequence-to-Sequence Models
abstract
Acoustic-to-Word recognition provides a straightforward solution to end-to-end speech recognition without needing external decoding, language model re-scoring or lexicon. While character-based models offer a natural solution to the-of-vocabulary problem, word models can be simpler to decode and may also be able to directly recognize semantically meaningful units. We present effective methods to train Sequence-to-Sequence models for direct word-level recognition (and character-level recognition) and show an absolute improvement of 4.4-5.0% in Word Error Rate on the Switchboard corpus compared to prior work. In addition to these promising results, word-based models are more interpretable than character models, which have to be composed into words using a separate decoding step. We analyze the encoder hidden states and the attention behavior, and show that location-aware attention naturally represents words as a single speech-word-vector, despite spanning multiple frames in the input. We finally show that the Acoustic-to-Word model also learns to segment speech into words with a mean standard deviation of 3 frames as compared with human annotated forced-alignments for the Switchboard corpus.
Shruti Palaskar, Florian Metze
SLT2
2018 Hierarchical Multitask Learning With CTC
abstract
In Automatic Speech Recognition, it is still challenging to learn useful intermediate representations when using high-level (or abstract) target units such as words. For that reason, when only a few hundreds of hours of training data are available, character or phoneme-based systems tend to outperform word-based systems. In this paper, we show how Hierarchical Multitask Learning can encourage the formation of useful intermediate representations. We achieve this by performing Connectionist Temporal Classification at different levels of the network with targets of different granularity. Our model thus performs predictions in multiple scales for the same input. On the standard 300h Switchboard training setup, our hierarchical multitask architecture demonstrates improvements over singletask architectures with the same number of parameters. Our model obtains 14.0% Word Error Rate on the Switchboard subset of the Eval2000 test set without any decoder or language model, outperforming the current state-of-the-art on non-autoregressive Acoustic-to-Word models.
Ramon Sanabria, Florian Metze
SLT2
2017 Visual features for context-aware speech recognition
abstract
Automatic transcriptions of consumer generated multi-media content such as “Youtube” videos still exhibit high word error rates. Such data typically occupies a very broad domain, has been recorded in challenging conditions, with cheap hardware and a focus on the visual modality, and may have been post-processed or edited.
Abhinav Gupta 0002, Yajie Miao, Leonardo Neves, Florian Metze
ICASSP4
2017 A comparison of Deep Learning methods for environmental sound detection
abstract
Environmental sound detection is a challenging application of machine learning because of the noisy nature of the signal, and the small amount of (labeled) data that is typically available. This work thus presents a comparison of several state-of-the-art Deep Learning models on the IEEE challenge on Detection and Classification of Acoustic Scenes and Events (DCASE) 2016 challenge task and data, classifying sounds into one of fifteen common indoor and outdoor acoustic scenes, such as bus, cafe, car, city center, forest path, library, train, etc. In total, 13 hours of stereo audio recordings are available, making this one of the largest datasets available.
Juncheng Li 0001, Florian Metze, Shuhui Qu, Samarjit Das
ICASSP3
2017 A first attempt at polyphonic sound event detection using connectionist temporal classification
abstract
Sound event detection is the task of detecting the type, starting time, and ending time of sound events in audio streams. Recently, recurrent neural networks (RNNs) have become the mainstream solution for sound event detection. Because RNNs make a prediction at every frame, it is necessary to provide exact starting and ending times of the sound events in the training data, making data annotation an extremely time-consuming process. Connectionist temporal classification (CTC), as a sequence-to-sequence model, can relax this constraint, because it suffices to provide ordered sequences of sound events without exact starting and ending times. This paper presents a first attempt at using CTC for sound event detection. In the polyphonic situation, sound events may overlap with each other, making it hard to define ordered sequences of sound events. We propose to use the boundaries (i.e. starts and ends) of the sound events as tokens for CTC. We show that CTC is able to locate the boundaries of sound events on a very noisy corpus of consumer generated content with rough hints about their positions. The CTC approach seems to be particularly suited to detecting short and transient sounds, which have traditionally been hardest to detect.
Yun Wang 0005, Florian Metze
ICASSP2
2017 A Transfer Learning Based Feature Extractor for Polyphonic Sound Event Detection Using Connectionist Temporal Classification
Yun Wang 0005, Florian Metze
INTERSPEECH2
2017 Comparison of Decoding Strategies for CTC Acoustic Models
abstract
Connectionist Temporal Classification has recently attracted a lot of interest as it offers an elegant approach to building acoustic models (AMs) for speech recognition. The CTC loss function maps an input sequence of observable feature vectors to an output sequence of symbols. Output symbols are conditionally independent of each other under CTC loss, so a language model (LM) can be incorporated conveniently during decoding, retaining the traditional separation of acoustic and linguistic components in ASR. For fixed vocabularies, Weighted Finite State Transducers provide a strong baseline for efficient integration of CTC AMs with n-gram LMs. Character-based neural LMs provide a straight forward solution for open vocabulary speech recognition and all-neural models, and can be decoded with beam search. Finally, sequence-to-sequence models can be used to translate a sequence of individual sounds into a word string. We compare the performance of these three approaches, and analyze their error patterns, which provides insightful guidance for future research and development in this important area.
Thomas Zenkel, Ramon Sanabria, Florian Metze, Jan Niehues, Matthias Sperber, Sebastian Stüker, Alex Waibel
INTERSPEECH3
2016 An empirical exploration of CTC acoustic models
abstract
The connectionist temporal classification (CTC) loss function has several interesting properties relevant for automatic speech recognition (ASR): applied on top of deep recurrent neural networks (RNNs), CTC learns the alignments between speech frames and label sequences automatically, which removes the need for pre-generated frame-level labels. CTC systems also do not require context decision trees for good performance, using context-independent (CI) phonemes or characters as targets. This paper presents an extensive exploration of CTC-based acoustic models applied to a variety of ASR tasks, including an empirical study of the optimal configuration and architectural variants for CTC. We observe that on large amounts of training data, CTC models tend to outperform state-of-the-art hybrid approach. Further experiments reveal that CTC can be readily ported to syllable-based languages, and can be enhanced by employing improved feature front-ends.
Yajie Miao, Mohammad Gowayyed, Xingyu Na, Tom Ko, Florian Metze, Alex Waibel
ICASSP5
2016 Audio-based multimedia event detection using deep recurrent neural networks
abstract
Multimedia event detection (MED) is the task of detecting given events (e.g. birthday party, making a sandwich) in a large collection of video clips. While visual features and automatic speech recognition typically provide the best features for this task, nonspeech audio can also contribute useful information, such as crowds cheering, engine noises, or animal sounds. MED is typically formulated as a two-stage process: the first stage generates clip-level feature representations, often by aggregating frame-level features; the second stage performs binary or multi-class classification to decide whether a given event occurs in a video clip. Both stages are usually performed "statically", i.e. using only local temporal information, or bag-of-words models. In this paper, we introduce longer-range temporal information with deep recurrent neural networks (RNNs) for both stages. We classify each audio frame among a set of semantic units called "noisemes" the sequence of frame-level confidence distributions is used as a variable-length clip-level representation. Such confidence vector sequences are then fed into long short-term memory (LSTM) networks for clip-level classification. We observe improvements in both frame-level and clip-level performance compared to SVM and feed-forward neural network baselines.
Yun Wang 0005, Leonardo Neves, Florian Metze
ICASSP3
2016 Experiences with Shared Resources for Research and Education in Speech and Language Processing
abstract
\n Contains fulltext :\n 161889.pdf (Publisher’s version ) (Open Access)\n
Rebecca Bates 0001, Eric Fosler-Lussier, Florian Metze, Martha A. Larson, Gina-Anne Levow, Emily Mower Provost
INTERSPEECH3
2016 Manipulating Word Lattices to Incorporate Human Corrections
Yashesh Gaur, Florian Metze, Jeffrey P. Bigham
INTERSPEECH2
2016 Virtual Machines and Containers as a Platform for Experimentation
abstract
Copyright © 2016 ISCA. Research on computational speech processing has traditionally relied on the availability of a relatively large and complex infrastructure, which encompasses data (text and audio), tools (feature extraction, model training, scoring, possibly on-line and off-line, etc.), glue code, and computing. Traditionally, it has been very hard to move experiments from one site to another, and to replicate experiments. With the increasing availability of shared platforms such as commercial cloud computing platforms or publicly funded super-computing centers, there is a need and an opportunity to abstract the experimental environment from the hardware, and distribute complete setups as a virtual machine, a container, or some other shareable resource, that can be deployed and worked with anywhere. In this paper, we discuss our experience with this concept and present some tools that the community might find useful. We outline, as a case study, how such tools can be applied to a naturalistic language acquisition audio corpus.
Florian Metze, Eric Riebling, Anne S. Warlaumont, Elika Bergelson
INTERSPEECH1
2016 Open-Domain Audio-Visual Speech Recognition: A Deep Learning Approach
Yajie Miao, Florian Metze
INTERSPEECH2
2016 Recurrent Support Vector Machines for Audio-Based Multimedia Event Detection
abstract
Multimedia event detection (MED) is the task of detecting given events (e.g. parade, birthday party) in a large collection of video clips. While the most useful information comes from visual features and speech recognition, a lot can also be inferred from the non-speech audio content, either alone or in conjunction with visual and speech cues. This paper studies MED with non-speech audio information only. MED is usually performed in two stages. The first stage generates a representation for each clip in the form of either a single vector or a sequence of vectors, often by aggregating frame-level features; the second stage performs binary or multi-class classification to decide whether each target event occurs in each clip. Common classifiers used for the second stage include support vector machines (SVMs), feed-forward deep neural networks (DNNs), and recurrent neural networks (RNNs).
Yun Wang 0005, Florian Metze
ICMR2
2015 EESEN: End-to-end speech recognition using deep RNN models and WFST-based decoding
abstract
The performance of automatic speech recognition (ASR) has improved tremendously due to the application of deep neural networks (DNNs). Despite this progress, building a new ASR system remains a challenging task, requiring various resources, multiple training stages and significant expertise. This paper presents our Eesen framework which drastically simplifies the existing pipeline to build state-of-the-art ASR systems. Acoustic modeling in Eesen involves learning a single recurrent neural network (RNN) predicting context-independent targets (phonemes or characters). To remove the need for pre-generated frame labels, we adopt the connectionist temporal classification (CTC) objective function to infer the alignments between speech and label sequences. A distinctive feature of Eesen is a generalized decoding approach based on weighted finite-state transducers (WFSTs), which enables the efficient incorporation of lexicons and language models into CTC decoding. Experiments show that compared with the standard hybrid DNN systems, Eesen achieves comparable word error rates (WERs), while at the same time speeding up decoding significantly.
Yajie Miao, Mohammad Gowayyed, Florian Metze
ASRU3
2015 QUESST2014: Evaluating Query-by-Example Speech Search in a zero-resource setting with real-life queries
abstract
In this paper, we present the task and describe the main findings of the 2014 “Query-by-Example Speech Search Task” (QUESST) evaluation. The purpose of QUESST was to perform language independent search of spoken queries on spoken documents, while targeting languages or acoustic conditions for which very few speech resources are available. This evaluation investigated for the first time the performance of query-by-example search against morphological and morpho-syntactic variability, requiring participants to match variants of a spoken query in several languages of different morphological complexity. Another novelty is the use of the normalized cross entropy cost (Cnxe) as the primary performance metric, keeping Term-Weighted Value (TWV) as a secondary metric for comparison with previous evaluations. After analyzing the most competitive submissions (by five teams), we find that, although low-level “pattern matching” approaches provide the best performance for “exact” matches, “symbolic” approaches working on higher-level representations seem to perform better in more complex settings, such as matching morphological variants. Finally, optimizing the output scores for Cnxe seems to generate systems that are more robust to differences in the operating point and that also perform well in terms of TWV, whereas the opposite might not be always true.
Xavier Anguera Miró, Luis Javier Rodríguez-Fuentes, Andi Buzo, Florian Metze, Igor Szöke, Mikel Peñagarikano
ICASSP4
2015 Semi-supervised training in low-resource ASR and KWS
abstract
In particular for “low resource” Keyword Search (KWS) and Speech-to-Text (STT) tasks, more untranscribed test data may be available than training data. Several approaches have been proposed to make this data useful during system development, even when initial systems have Word Error Rates (WER) above 70%. In this paper, we present a set of experiments on low-resource languages in telephony speech quality in Assamese, Bengali, Lao, Haitian, Zulu, and Tamil, demonstrating the impact that such techniques can have, in particular learning robust bottle-neck features on the test data. In the case of Tamil, when significantly more test data than training data is available, we integrated semi-supervised training and speaker adaptation on the test data, and achieved significant additional improvements in STT and KWS.
Florian Metze, Ankur Gandhe, Yajie Miao, Zaid Sheikh, Yun Wang 0005, Hao Zhang 0025, Jungsuk Kim, Ian Lane, Wonkyum Lee, Sebastian Stüker, Markus Müller 0001
ICASSP1
2015 Regularizing DNN acoustic models with Gaussian stochastic neurons
abstract
Dropout and DropConnect can be viewed as regularization methods for deep neural network (DNN) training. In DNN acoustic modeling, the huge number of speech samples makes it expensive to sample the neuron mask (Dropout) or the weight mask (DropConnect) repetitively from a high dimensional distribution. In this paper we investigate the effect of Gaussian stochastic neurons on DNN acoustic modeling. The pre-Gaussian stochastic term can be viewed as a variant of Dropout/DropConnect and the post-Gaussian stochastic term generalizes the idea of data augmentation into hidden layers. Gaussian stochastic neurons can give improvement on large data sets where Dropout tends to be less useful. Under the low resource condition, its performance is comparable with Dropout, but with a lower time complexity during fine-tuning.
Hao Zhang 0025, Yajie Miao, Florian Metze
ICASSP3
2015 Using keyword spotting to help humans correct captioning faster
abstract
Automatic real-time captioning provides immediate and on de-mand access to spoken content in lectures or talks, and is a cru-cial accommodation for deaf and hard of hearing (DHH) people. However, in the presence of specialized content, like in techni-cal talks, automatic speech recognition (ASR) still makes mis-takes which may render the output incomprehensible. In this paper, we introduce a new approach, which allows audience or crowd workers, to quickly correct errors that they spot in ASR output. Prior approaches required the crowd worker to manu-ally “edit ” the ASR hypothesis by selecting and replacing the text, which is not suitable for real-time scenarios. Our approach is faster and allows the worker to simply type corrections for misrecognized words as soon as he or she spots them. The sys-tem then finds the most likely position for the correction in the ASR output using keyword search (KWS) and stitches the word into the ASR output. Our work demonstrates the potential of computation to incorporate human input quickly enough to be usable in real-time scenarios, and may be a better method for providing this vital accommodation to DHH people. Index Terms: speech recognition, human-computer interac-tion, spoken term detection, real-time crowd sourcing.
Yashesh Gaur, Florian Metze, Yajie Miao, Jeffrey P. Bigham
INTERSPEECH2
2015 The speech recognition virtual kitchen turns one
Florian Metze, Eric Riebling, Eric Fosler-Lussier, Andrew R. Plummer, Rebecca Bates 0001
INTERSPEECH1
2015 Distance-aware DNNs for robust speech recognition
abstract
Distant speech recognition (DSR) remains to be an open chal-lenge, even for the state-of-the-art deep neural network (DNN) models. Previous work has attempted to improve DNNs un-der constantly distant speech. However, in real applications, the speaker-microphone distance (SMD) can be quite dynamic, varying even within a single utterance. This paper investigates how to alleviate the impact of dynamic SMD on DNN models. Our solution is to incorporate the frame-level SMD information into DNN training. Generation of the SMD information relies on a universal extractor that is learned on a meeting corpus. We study the utility of different architectures in instantiating the SMD extractor. On our target acoustic modeling task, two approaches are proposed to build distance-aware DNN models using the SMD information: simple concatenation and distance adaptive training (DAT). Our experiments show that in the sim-plest case, incorporating the SMD descriptors improves word error rates of DNNs by 5.6 % relative. Further optimizing SMD extraction and integration results in more gains. Index Terms: Deep neural networks, speaker-microphone dis-tance, robust acoustic modeling
Yajie Miao, Florian Metze
INTERSPEECH2
2015 On speaker adaptation of long short-term memory recurrent neural networks
abstract
Long Short-Term Memory (LSTM) is a recurrent neural net-work (RNN) architecture specializing in modeling long-range temporal dynamics. On acoustic modeling tasks, LSTM-RNNs have shown better performance than DNNs and conventional RNNs. In this paper, we conduct an extensive study on speaker adaptation of LSTM-RNNs. Speaker adaptation helps to reduce the mismatch between acoustic models and testing speakers. We have two main goals for this study. First, on a benchmark dataset, the existing DNN adaptation techniques are evaluated on the adaptation of LSTM-RNNs. We observe that LSTM-RNNs can be effectively adapted by using speaker-adaptive (SA) front-end, or by inserting speaker-dependent (SD) layers. Second, we propose two adaptation approaches that implement the SD-layer-insertion idea specifically for LSTM-RNNs. Us-ing these approaches, speaker adaptation improves word error rates by 3-4 % relative over a strong LSTM-RNN baseline. This improvement is enlarged to 6-7 % if we exploit SA features for further adaptation. Index Terms: Long Short-Term Memory, recurrent neural net-work, acoustic modeling, speaker adaptation
Yajie Miao, Florian Metze
INTERSPEECH2
2015 Speaker Adaptive Training of Deep Neural Network Acoustic Models Using I-Vectors
abstract
In acoustic modeling, speaker adaptive training (SAT) has been a long-standing technique for the traditional Gaussian mixture models (GMMs). Acoustic models trained with SAT become independent of training speakers and generalize better to unseen testing speakers. This paper ports the idea of SAT to deep neural networks (DNNs), and proposes a framework to perform feature-space SAT for DNNs. Using i-vectors as speaker representations, our framework learns an adaptation neural network to derive speaker-normalized features. Speaker adaptive models are obtained by fine-tuning DNNs in such a feature space. This framework can be applied to various feature types and network structures, posing a very general SAT solution. In this paper, we fully investigate how to build SAT-DNN models effectively and efficiently. First, we study the optimal configurations of SAT-DNNs for large-scale acoustic modeling tasks. Then, after presenting detailed comparisons between SAT-DNNs and the existing DNN adaptation methods, we propose to combine SAT-DNNs and model-space DNN adaptation during decoding. Finally, to accelerate learning of SAT-DNNs, a simple yet effective strategy, frame skipping, is employed to reduce the size of training data. Our experiments show that compared with a strong DNN baseline, the SAT-DNN model achieves 13.5% and 17.5% relative improvement on word error rates (WERs), without and with model-space adaptation applied respectively. Data reduction based on frame skipping results in 2 × speed-up for SAT-DNN training, while causing negligible WER loss on the testing data.
Yajie Miao, Hao Zhang 0025, Florian Metze
IEEE ACM Trans. Audio Speech Lang. Process.3
2014 Augmenting Translation Models with Simulated Acoustic Confusions for Improved Spoken Language Translation
abstract
We propose a novel technique for adapting text-based statistical machine translation to deal with input from automatic speech recognition in spoken language translation tasks.We simulate likely misrecognition errors using only a source language pronunciation dictionary and language model (i.e., without an acoustic model), and use these to augment the phrase table of a standard MT system.The augmented system can thus recover from recognition errors during decoding using synthesized phrases.Using the outputs of five different English ASR systems as input, we find consistent and significant improvements in translation quality.Our proposed technique can also be used in conjunction with lattices as ASR output, leading to further improvements.
Yulia Tsvetkov, Florian Metze, Chris Dyer
EACL2
2014 Optimization of Neural Network Language Models for keyword search
abstract
Recent works have shown Neural Network based Language Models (NNLMs) to be an effective modeling technique for Automatic Speech Recognition. Prior works have shown that these models obtain lower perplexity and word error rate (WER) compared to both standard n-gram language models (LMs) and more advanced language models including maximum entropy and random forest LMs. While these results are compelling, prior works were limited to evaluating NNLMs on perplexity and word error rate. Our initial results showed that while NNLMs improved speech recognition accuracy, the improvement in keyword search was negligible. In this paper we propose alternate optimizations of NNLMs for the task of keyword search. We evaluate the performance of the proposed methods for keyword search on the Vietnamese dataset provided in phase one of the BABEL1project and demonstrate that by penalizing low frequency words during NNLM training, keyword search metrics such as actual term weighted value (ATWV) can be improved by up to 9.3% compared to the standard training methods.
Ankur Gandhe, Florian Metze, Alex Waibel, Ian Lane
ICASSP2
2014 Exploring audio semantic concepts for event-based video retrieval
abstract
The audio semantic concepts (sound events) play important roles in audio-based content analysis. How to capture the semantic information effectively from the complex occurrence pattern of sound events in YouTube quality videos is a challenging problem. This paper presents a novel framework to handle the complex situation for semantic information extraction in real-world videos and evaluate through the NIST multimedia event detection task (MED). We calculate the occurrence confidence matrix of sound events and explore multiple strategies to generate clip-level semantic features from the matrix. We evaluate the performance using TRECVID2011 MED dataset. The proposed method outperforms previous HMM-based system. The late fusion experiment with the low-level features and text feature (ASR) shows that audio semantic concepts capture complementary information in the soundtrack.
Shourabh Rawat, Florian Metze
ICASSP3
2014 Semi-automatic audio semantic concept discovery for multimedia retrieval
abstract
Huge amount of videos on the Internet have rare textual information, which makes video retrieval challenging given a text query. Previous work explored semantic concepts for content analysis to assist retrieval. However, the human-defined concepts might fail to cover the data and there is a potential gap between these concepts and the semantics expected from user's query. Also, building a corpus is expensive and time-consuming. To address these issues, we propose a semi-automatic framework to discover the semantic concepts. We limit ourselves in audio modality here. In the paper, we also discuss how to select meaningful vocabulary from the discovered hierarchical sub-categories and provide an approach to detect all the concepts without further annotation. We evaluate the method on NIST 2011 multimedia event detection (MED) dataset.
Shourabh Rawat, Florian Metze
ICASSP3
2014 Improved audio features for large-scale multimedia event detection
abstract
In this paper, we present recent experiments on using Artificial Neural Networks (ANNs), a new “delayed” approach to speech vs. non-speech segmentation, and extraction of large-scale pooling feature (LSPF) for detecting “events” within consumer videos, using the audio channel only. A “event” is defined to be a sequence of observations in a video, that can be directly observed or inferred. Ground truth is given by a semantic description of the event, and by a number of example videos. We describe and compare several algorithmic approaches, and report results on the 2013 TRECVID Multimedia Event Detection (MED) task, using arguably the largest such research set currently available. The presented system achieved the best results in most audio-only conditions. While the overall finding is that MFCC features perform best, we find that ANN as well as LSP features provide complementary information at various levels of temporal resolution. This paper provides analysis of both low-level and high-level features, investigating their relative contributions to overall system performance.
Florian Metze, Shourabh Rawat
ICME1
2014 Query-by-example spoken term detection on multilingual unconstrained speech
abstract
As part of the MediaEval 2013 benchmark evaluation campaign, the objective of the Spoken Web Search (SWS) task was to perform Query-by-Example Spoken Term Detection (QbESTD) using audio queries in a low-resource setting. After two successful editions and a continuously growing interest in the scientific community, a special effort was made in SWS 2013 to prepare a challenging database, including speech in 9 different languages with diverse environment and channel conditions. In this paper, first we describe the database and the performance metrics. Then, we briefly review the algorithmic approaches followed by participants and present and discuss the obtained performances, which demonstrate the feasibility of the proposed task, even under such challenging conditions (multiple languages and unconstrained acoustic conditions). Finally, we analyze the fusion of the top-performing systems, which achieved a 30% relative improvement over the best single system in the evaluation, proving that a variety of approaches can be effectively combined to bring complementary information in the search for queries.
Xavier Anguera Miró, Luis Javier Rodríguez-Fuentes, Igor Szöke, Andi Buzo, Florian Metze, Mikel Peñagarikano
INTERSPEECH5
2014 Neural network language models for low resource languages
abstract
For resource rich languages, recent works have shown Neural Network based Language Models (NNLMs) to be an effective modeling technique for Automatic Speech Recognition, out performing standard n-gram language models (LMs). For low resource languages, however, the performance of NNLMs has not been well explored. In this paper, we evaluate the effectiveness of NNLMs for low resource languages and show that NNLMs learn better word probabilities than state-of-theart n-gram models even when the amount of training data is severely limited. We show that interpolated NNLMs obtain a lower WER than standard n-gram models, no mater the amount of training data. Additionally, we observe that with small amounts of data (approx. 100k training tokens), feed-forward NNLMs obtain lower perplexity than recurrent NNLMs, while for the larger data condition (500k-1M training tokens), recurrent NNLMs can obtain lower perplexity than feed-forward models.
Ankur Gandhe, Florian Metze, Ian Lane
INTERSPEECH2
2014 Improving language-universal feature extraction with deep maxout and convolutional neural networks
abstract
When deployed in automated speech recognition (ASR), deep neural networks (DNNs) can be treated as a complex feature extractor plus a simple linear classifier. Previous work has investigated the utility of multilingual DNNs acting as language-universal feature extractors (LUFEs). In this paper, we explore different strategies to further improve LUFEs. First, we replace the standard sigmoid nonlinearity with the recently proposed maxout units. The resulting maxout LUFEs have the nice property of generating sparse feature representations. Second, the convolutional neural network (CNN) architecture is applied to obtain more invariant feature space. We evaluate the performance of LUFEs on a cross-language ASR task. Each of the proposed techniques results in word error rate reduction compared with the existing DNN-based LUFEs. Combining the two methods together brings additional improvement on the target language.
Yajie Miao, Florian Metze
INTERSPEECH2
2014 Distributed learning of multilingual DNN feature extractors using GPUs
abstract
Multilingual deep neural networks (DNNs) can act as deep feature extractors and have been applied successfully to crosslanguage acoustic modeling. Learning these feature extractors becomes an expensive task, because of the enlarged multilingual training data and the sequential nature of stochastic gradient descent (SGD). This paper investigates strategies to accelerate the learning process over multiple GPU cards. We propose the DistModel and DistLang frameworks which distribute feature extractor learning by models and languages respectively. The time-synchronous DistModel has the nice property of tolerating infrequent model averaging. With 3 GPUs, DistModel achieves 2.6× speed-up and causes no loss on word error rates. When using DistLang, we observe better acceleration but worse recognition performance. Further evaluations are conducted to scale DistModel to more languages and GPU cards.
Yajie Miao, Hao Zhang 0025, Florian Metze
INTERSPEECH3
2014 Towards speaker adaptive training of deep neural network acoustic models
abstract
We investigate the concept of speaker adaptive training (SAT) in the context of deep neural network (DNN) acoustic models. Previous studies have shown success of performing speaker adaptation for DNNs in speech recognition. In this paper, we apply SAT to DNNs by learning two types of feature mapping neural networks. Given an initial DNN model, these networks take speaker i-vectors as additional information and project DNN inputs into a speaker-normalized space. The final SAT model is obtained by updating the canonical DNN in the normalized feature space. Experiments on a Switchboard 110- hour setup show that compared with the baseline DNN, the SAT-DNN model brings 7.5% and 6.0% relative improvement when DNN inputs are speaker-independent and speakeradapted features respectively. Further evaluations on the more challenging BABEL datasets reveal significant word error rate reduction achieved by SAT-DNN.
Yajie Miao, Hao Zhang 0025, Florian Metze
INTERSPEECH3
2014 The speech recognition virtual kitchen: launch party
Andrew R. Plummer, Eric Riebling, Florian Metze, Eric Fosler-Lussier, Rebecca Bates 0001
INTERSPEECH4
2014 An in-depth comparison of keyword specific thresholding and sum-to-one score normalization
abstract
The quality of a spoken term detection (STD) system critically depends on the choice of a “thresholding” function, which is used to determine whether to output a candidate detection or not based on its score. In the context of the IARPA Babel program and the NIST OpenKWS evaluation series, the penalty for missing an occurrence depends on the frequency of the keyword, so it is desirable either to apply different thresholds to different keywords, or to normalize the scores before applying a global threshold. This paper compares two widely used thresholding algorithms: keyword specific thresholding (KST) and sum-to-one score normalization (STO), analyzes the difference in their performance in detail, and recommends the use of the “estimated KST” algorithm
Yun Wang 0005, Florian Metze
INTERSPEECH2
2014 Word-based probabilistic phonetic retrieval for low-resource spoken term detection
abstract
Two problems make Spoken Term Detection (STD) particularly challenging under low-resource conditions: the low quality of speech recognition hypotheses, and a high number of out-ofvocabulary (OOV) words. In this paper, we propose an intuitive way to handle OOV terms for STD on word-based Confusion Networks using phonetic similarities, and generalize it into a probabilistic and vocabulary-independent retrieval framework. We then reflect on how several heuristics and Machine Learning based methods can be incorporated into this framework to improve retrieval performance. We present experimental results on several low-resource languages from IARPA’s Babel program, such as Assamese, Bengali, Haitian, and Lao.
Florian Metze
INTERSPEECH2
2014 A methodology for using crowdsourced data to measure uncertainty in natural speech
abstract
People sometimes express uncertainty unconsciously in order to add layers of meaning on top of their speech, conveying doubts about the accuracy of the information they are trying to communicate. In this paper, we propose a methodology for annotating uncertainty, which is usually a subjective and expensive process, by using crowdsourcing. In our experiment, we used an online database which consists of colors that more than 200,000 users have named. Based on the amount of unique names that users have given each color, an entropy value was calculated to represent the uncertainty level of the color. A model, which performed better than chance, was created to predict whether or not the color that the participant was describing was ambiguous or borderline, given certain prosodic cues of their speech when asked to name the color verbally. Using crowdsourced data can greatly streamline the process of annotating uncertainty, but our methods have yet to be tested in other domains besides color. By using methods such as ours to measure prosodic attributes of uncertainty, it should be possible to increase the accuracy of voice search.
Lara J. Martin, Matthew Stone, Florian Metze, Jack Mostow
SLT3
2014 Improvements to speaker adaptive training of deep neural networks
abstract
Speaker adaptive training (SAT) is a well studied technique for Gaussian mixture acoustic models (GMMs). Recently we proposed to perform SAT for deep neural networks (DNNs), with speaker i-vectors applied in feature learning. The resulting SAT-DNN models significantly outperform DNNs on word error rates (WERs). In this paper, we present different methods to further improve and extend SAT-DNN. First, we conduct detailed analysis to investigate i-vector extractor training and flexible feature fusion. Second, the SAT-DNN approach is extended to improve tasks including bottleneck feature (BNF) generation, convolutional neural network (CNN) acoustic modeling and multilingual DNN-based feature extraction. Third, for transcribing multimedia data, we enrich the i-vector representation with global speaker attributes (age, gender, etc.) obtained automatically from video signals. On a collection of instructional videos, incorporation of the additional visual features is observed to boost the recognition accuracy of SAT-DNN.
Yajie Miao, Lu Jiang 0004, Hao Zhang 0025, Florian Metze
SLT4
2014 A keyword search system using open source software
abstract
Provides an overview of a speech-to-text (STT) and keyword search (KWS) system architecture build primarily on the top of the Kaldi toolkit and expands on a few highlights. The system was developed as a part of the research efforts of the Radical team while participating in the IARPA Babel program. Our aim was to develop a general system pipeline which could be easily and rapidly deployed in any language, independently on the language script and phonological and linguistic features of the language.
Jan Trmal, Guoguo Chen, Daniel Povey, Sanjeev Khudanpur, Pegah Ghahremani, Xiaohui Zhang 0007, Vimal Manohar, Chunxi Liu, Aren Jansen, Dietrich Klakow, David Yarowsky, Florian Metze
SLT12
2014 EM-based phoneme confusion matrix generation for low-resource spoken term detection
abstract
The idea of using a data-driven phoneme confusion matrix (PCM) to enhance speech recognition and retrieval performance is not new to the speech community. Although empirical results show various degrees of improvements brought by introducing a PCM, the underlying data-driven processes introduced in most papers are rather ad-hoc and lack rigorous statistical justifications. In this paper we will focus on the statistical aspects of PCM generation, propose and justify a novel expectation-maximization based algorithm for data-driven PCM generation. We will evaluate the performance of the generated PCMs under the context of low-resource spoken term detection, with primary focus on out-of-vocabulary keywords.
Yun Wang 0005, Florian Metze
SLT3
2014 Language independent search in MediaEval's Spoken Web Search task
Florian Metze, Xavier Anguera Miró, Etienne Barnard, Marelie H. Davel, Guillaume Gravier
Comput. Speech Lang.1
2013 Using web text to improve keyword spotting in speech
abstract
For low resource languages, collecting sufficient training data to build acoustic and language models is time consuming and often expensive. But large amounts of text data, such as online newspapers, web forums or online encyclopedias, usually exist for languages that have a large population of native speakers. This text data can be easily collected from the web and then used to both expand the recognizer's vocabulary and improve the language model. One challenge, however, is normalizing and filtering the web data for a specific task. In this paper, we investigate the use of online text resources to improve the performance of speech recognition specifically for the task of keyword spotting. For the five languages provided in the base period of the IARPA BABEL project, we automatically collected text data from the web using only Limited LP resources. We then compared two methods for filtering the web data, one based on perplexity ranking and the other based on out-of-vocabulary (OOV) word detection. By integrating the web text into our systems, we observed significant improvements in keyword spotting accuracy for four out of the five languages. The best approach obtained an improvement in actual term weighted value (ATWV) of 0.0424 compared to a baseline system trained only on LimitedLP resources. On average, ATWV was improved by 0.0243 across five languages.
Ankur Gandhe, Florian Metze, Alexander I. Rudnicky, Ian Lane, Matthias Eck 0001
ASRU3
2013 DNN acoustic modeling with modular multi-lingual feature extraction networks
abstract
In this work, we propose several deep neural network architectures that are able to leverage data from multiple languages. Modularity is achieved by training networks for extracting high-level features and for estimating phoneme state posteriors separately, and then combining them for decoding in a hybrid DNN/HMM setup. This approach has been shown to achieve superior performance for single-language systems, and here we demonstrate that feature extractors benefit significantly from being trained as multi-lingual networks with shared hidden representations. We also show that existing mono-lingual networks can be re-used in a modular fashion to achieve a similar level of performance without having to train new networks on multi-lingual data. Furthermore, we investigate in extending these architectures to make use of language-specific acoustic features. Evaluations are performed on a low-resource conversational telephone speech transcription task in Vietnamese, while additional data for acoustic model training is provided in Pashto, Tagalog, Turkish, and Cantonese. Improvements of up to 17.4% and 13.8% over mono-lingual GMMs and DNNs, respectively, are obtained.
Jonas Gehring, Quoc Bao Nguyen, Florian Metze, Alex Waibel
ASRU3
2013 Models of tone for tonal and non-tonal languages
abstract
Conventional wisdom in automatic speech recognition asserts that pitch information is not helpful in building speech recognizers for non-tonal languages and contributes only modestly to performance in speech recognizers for tonal languages. To maintain consistency between different systems, pitch is therefore often ignored, trading the slight performance benefits for greater system uniformity/ simplicity. In this paper, we report results that challenge this conventional approach. We present new models of tone that deliver consistent performance improvements for tonal languages (Cantonese, Vietnamese) and even modest improvements for non-tonal languages. Using neural networks for feature integration and fusion, these models achieve significant gains throughout, and provide us with system uniformity and standardization across all languages, tonal and non-tonal.
Florian Metze, Zaid Sheikh, Alex Waibel, Jonas Gehring, Kevin Kilgour, Quoc Bao Nguyen, Van Huy Nguyen
ASRU1
2013 Deep maxout networks for low-resource speech recognition
abstract
As a feed-forward architecture, the recently proposed maxout networks integrate dropout naturally and show state-of-the-art results on various computer vision datasets. This paper investigates the application of deep maxout networks (DMNs) to large vocabulary continuous speech recognition (LVCSR) tasks. Our focus is on the particular advantage of DMNs under low-resource conditions with limited transcribed speech. We extend DMNs to hybrid and bottleneck feature systems, and explore optimal network structures (number of maxout layers, pooling strategy, etc) for both setups. On the newly released Babel corpus, behaviors of DMNs are extensively studied under different levels of data availability. Experiments show that DMNs improve low-resource speech recognition significantly. Moreover, DMNs introduce sparsity to their hidden activations and thus can act as sparse feature extractors.
Yajie Miao, Florian Metze, Shourabh Rawat
ASRU2
2013 Neighbour selection and adaptation for rapid speaker-dependent ASR
abstract
Speaker dependent (SD) ASR systems have significantly lower word error rates (WER) compared to speaker independent (SI) systems. However, SD systems require sufficient training data from the target speaker, which is impractical to collect in a short time. We present a technique for training SD models using just few minutes of speaker's data. We compensate for the lack of adequate speaker-specific data by selecting neighbours from a database of existing speakers who are acoustically close to the target speaker. These neighbours provide ample training data, which is used to adapt the SI model to obtain an initial SD model for the new speaker with significantly lower WER. We evaluate various neighbour selection algorithms on a large-scale medical transcription task and report significant reduction in WER using only 5 mins of speaker-specific data. We conduct a detailed analysis of various factors such as gender and accent in the neighbour selection. Finally, we study neighbour selection and adaptation in the context of discriminative objective functions.
Udhyakumar Nallasamy, Mark C. Fuhs, Monika Woszczyna, Florian Metze, Tanja Schultz
ASRU4
2013 Extracting deep bottleneck features using stacked auto-encoders
abstract
In this work, a novel training scheme for generating bottleneck features from deep neural networks is proposed. A stack of denoising auto-encoders is first trained in a layer-wise, unsupervised manner. Afterwards, the bottleneck layer and an additional layer are added and the whole network is fine-tuned to predict target phoneme states. We perform experiments on a Cantonese conversational telephone speech corpus and find that increasing the number of auto-encoders in the network produces more useful features, but requires pre-training, especially when little training data is available. Using more unlabeled data for pre-training only yields additional gains. Evaluations on larger datasets and on different system setups demonstrate the general applicability of our approach. In terms of word error rate, relative improvements of 9.2% (Cantonese, ML training), 9.3% (Tagalog, BMMI-SAT training), 12% (Tagalog, confusion network combinations with MFCCs), and 8.7% (Switchboard) are achieved.
Jonas Gehring, Yajie Miao, Florian Metze, Alex Waibel
ICASSP3
2013 A summary of the 2012 JHU CLSP workshop on zero resource speech technologies and models of early language acquisition
abstract
We summarize the accomplishments of a multi-disciplinary workshop exploring the computational and scientific issues surrounding zero resource (unsupervised) speech technologies and related models of early language acquisition. Centered around the tasks of phonetic and lexical discovery, we consider unified evaluation metrics, present two new approaches for improving speaker independence in the absence of supervision, and evaluate the application of Bayesian word segmentation algorithms to automatic subword unit tokenizations. Finally, we present two strategies for integrating zero resource techniques into supervised settings, demonstrating the potential of unsupervised methods to improve mainstream technologies.
Aren Jansen, Emmanuel Dupoux, Sharon Goldwater, Mark Johnson 0001, Sanjeev Khudanpur, Kenneth Church 0001, Naomi Feldman, Hynek Hermansky, Florian Metze, Richard C. Rose, Mike Seltzer, Pascal Clark, Ian McGraw, Balakrishnan Varadarajan, Erin D. Bennett, Benjamin Börschinger, Justin T. Chiu, Ewan Dunbar, Abdellah Fourtassi, David F. Harwath, Chia-ying Lee, Keith D. Levin, Atta Norouzian, Vijayaditya Peddinti, Rachael Richardson, Thomas Schatz, Samuel Thomas 0001
ICASSP9
2013 The spoken web search task at MediaEval 2012
abstract
In this paper, we describe the “Spoken Web Search” Task, which was held as part of the 2012 MediaEval benchmark evaluation campaign. The purpose of this task was to perform audio search with audio input in four languages, with very few resources being available. Continuing in the spirit of the 2011 SpokenWeb Search Task, which used speech from four Indian languages, the 2012 data was taken from the LWAZI corpus, to provide even more diversity and allow for a task that will allow both zero resource “pattern matching” approaches and “speech recognition” based approaches to participate. In this paper, we summarize the results from several independent systems, developed by nine teams, analyze their performance, and provide directions for future research.
Florian Metze, Xavier Anguera Miró, Etienne Barnard, Marelie H. Davel, Guillaume Gravier
ICASSP1
2013 Subspace mixture model for low-resource speech recognition in cross-lingual settings
abstract
The subspace Gaussian mixture model (SGMM) has been exploited for cross-lingual speech recognition. The general motivation is that the subspace parameters can be estimated on multiple source languages and then transferred to the target language. In this work, we investigate an extension to SGMM, referred to as subspace mixture model (SMM), in which subspace parameters on the target language are casted as a linear mixture of the subspaces derived from source languages. This approach reduces the number of SGMM model parameters, while retaining the flexibility of subspace learning on the target language. Experiments show that the proposed SMM method outperforms SGMM significantly when the target language has limited training data.
Yajie Miao, Florian Metze, Alex Waibel
ICASSP2
2013 Learning discriminative basis coefficients for eigenspace MLLR unsupervised adaptation
abstract
Eigenspace MLLR is effective for fast adaptation when the amount of adaptation data is limited, e.g., less than 5s. The general motivation is to represent the MLLR transform as a linear combination of basis matrices. In this paper, we present a framework to estimate a speaker-independent discriminative transform over the combination coefficients. This discriminative basis coefficients transform (DBCT) is learned by optimizing discriminative criteria over all the training speakers. During recognition, the ML basis coefficients for each testing speaker are firstly found, on which DBCT is applied to give the final MLLR transform discrimination ability. Experiments show that DBCT results in consistent WER reduction in unsupervised adaptation, compared with both standard ML and discriminatively trained transforms.
Yajie Miao, Florian Metze, Alex Waibel
ICASSP2
2013 Identification and modeling of word fragments in spontaneous speech
abstract
This paper presents a novel approach to handling disfluencies, word fragments and self-interruption points in Cantonese conversational speech. We train a classifier that exploits lexical and acoustic information to automatically identify disfluencies during training of a speech recognition system on conversational speech, and then use this classifier to augment reference annotations used for acoustic model training. We experiment with approaches to modeling disfluencies in the pronunciation dictionary, and their effect on the polyphonic decision tree clustering. We achieve automatic detection of disfluencies with 88% accuracy, which leads to a reduction in character error rate of 1.9% absolute. While the high baseline error rates are due to the task we are currently working on, we demonstrate that this approach works well on the Switchboard corpus, for which the conversational nature of speech is also a major problem.
Yulia Tsvetkov, Zaid Sheikh, Florian Metze
ICASSP3
2013 Prosody-Based Unsupervised Speech Summarization with Two-Layer Mutually Reinforced Random Walk
Sujay Kumar Jauhar, Yun-Nung Chen, Florian Metze
IJCNLP3
2013 Multi-layer mutually reinforced random walk with hidden parameters for improved multi-party meeting summarization
abstract
This paper proposes an improved approach of summarization for spoken multi-party interaction, in which a multi-layer graph with hidden parameters is constructed. The graph includes utterance-to-utterance relation, utterance-to-parameter weight, and speaker-to-parameter weight. Each utterance and each speaker are represented as a node in the utterance-layer and speaker-layer of the graph respectively. We use terms/ topics as hidden parameters for estimating utterance-to-parameter and speaker-to-parameter weight, and compute topical similarity between utterances as the utterance-to-utterance relation. By within- and between-layer propagation in the graph, the scores from different layers can be mutually reinforced so that utterances can automatically share the scores with the utterances from the speakers who focus on similar terms/ topics. For both ASR output and manual transcripts, experiments confirmed the efficacy of including hidden parameters and involving speaker information in the multi-layer graph for summarization. We find that choosing latent topics as hidden parameters significantly reduces computational complexity and does not hurt the performance.
Yun-Nung Chen, Florian Metze
INTERSPEECH2
2013 Formalizing expert knowledge for developing accurate speech recognizers
abstract
The expertise required to develop a speech recognition system with reasonable accuracy for a given task is quite significant, and precludes most non-speech experts from integrating speech recognition into their own research. While an initial baseline recognizer may readily be available or relatively simple to acquire, identifying the necessary accuracy optimizations require an expert understanding of the application domain as well as significant experience in building speech recognition systems. This paper describes our efforts and experiments in formalizing knowledge from speech experts that would help novices by automatically analyzing an acoustic context and recommending appropriate techniques for accuracy gains. Through two recognition experiments, we show that it is possible to model experts' understanding of developing accurate speech recognition systems in a rule-based knowledge base, and that this knowledge base can accurately predict successful optimization techniques for previously seen acoustic situations, both in seen and unseen datasets. We argue that such a knowledge base, once fully developed, will be of tremendous value for boosting the use of speech recognition in research and development on non-mainstream languages and acoustic conditions.
Florian Metze, Matthew Kam
INTERSPEECH2
2013 The speech recognition virtual kitchen
Florian Metze, Eric Fosler-Lussier, Rebecca Bates 0001
INTERSPEECH1
2013 Improving low-resource CD-DNN-HMM using dropout and multilingual DNN training
abstract
We investigate two strategies to improve the context-dependent deep neural network hidden Markov model (CD-DNN-HMM) in low-resource speech recognition. Although outperforming the conventional Gaussian mixture model (GMM) HMM on various tasks, CD-DNN-HMM acoustic modeling becomes challenging with limited transcribed speech, e.g., less than 10 hours. To resolve this issue, we firstly exploit dropout which prevents overfitting in DNN finetuning and improves model robustness under data sparseness. Then, the effectiveness of multilingual DNN training is evaluated when additional auxiliary languages are available. The hidden layer parameters of the target language are shared and learned over multiple languages. Experiments show that both strategies boost the recognition performance significantly. Combining them results in further reduction in word error rate, achieving 11.6% and 6.2% relative improvement on two limited data conditions.
Yajie Miao, Florian Metze
INTERSPEECH2
2013 Robust audio-codebooks for large-scale event detection in consumer videos
abstract
In this paper we present our audio based system for detecting "events" within consumer videos (e.g. You Tube) and report our experiments on the TRECVID Multimedia Event Detection (MED) task and development data. Codebook or bag-of-words models have been widely used in text, visual and audio domains and form the state-of-the-art in MED tasks. The overall effectiveness of these models on such datasets depends critically on the choice of lowlevel features, clustering approach, sampling method, codebook size, weighting schemes and choice of classifier. In this work we empirically evaluate several approaches to model expressive and robust audio codebooks for the task of MED while ensuring compactness. First, we introduce the Large Scale Pooling Features (LSPF) and Stacked Cepstral Features for encoding local temporal information in audio codebooks. Second, we discuss several design decisions for generating and representing expressive audio codebooks and show how they scale to large datasets. Third, we apply text based techniques like Latent Dirichlet Allocation (LDA) to learn acoustic-topics as a means of providing compact representation while maintaining performance. By aggregating these decisions into our model, we obtained 11% relative improvement over our baseline audio systems.
Shourabh Rawat, Peter Schulam 0001, Susanne Burger, Duo Ding, Florian Metze
INTERSPEECH6
2012 Articulatory features for expressive speech synthesis
abstract
This paper describes some of the results from the project entitled “New Parameterization for Emotional Speech Synthesis” held at the Summer 2011 JHU CLSP workshop. We describe experiments on how to use articulatory features as a meaningful intermediate representation for speech synthesis. This parameterization not only allows us to reproduce natural sounding speech but also allows us to generate stylistically varying speech.
Alan W. Black, H. Timothy Bunnell, Ying Dou, Prasanna Kumar Muthukumar, Florian Metze, Daniel Perry 0002, Tim Polzehl, Kishore Prahallad, Stefan Steidl, Callie Vaughn
ICASSP5
2012 The Spoken Web Search Task at MediaEval 2011
abstract
In this paper, we describe the “Spoken Web Search” Task, which was held as part of the 2011 MediaEval benchmark campaign. The purpose of this task was to perform audio search with audio input in four languages, with very few resources being available in each language. The data was taken from “spoken web” material collected over mobile phone connections by IBM India. We present results from several independent systems, developed by five teams and using different approaches, compare them, and provide analysis and directions for future research.
Florian Metze, Nitendra Rajput, Xavier Anguera Miró, Marelie H. Davel, Guillaume Gravier, Charl Johannes van Heerden, Gautam Varma Mantena, Armando Muscariello, Kishore Prahallad, Igor Szöke, Javier Tejedor
ICASSP1
2012 Generating Natural Language Summaries for Multimedia
Duo Ding, Florian Metze, Shourabh Rawat, Peter Schulam 0001, Susanne Burger
INLG2
2012 Integrating Intra-Speaker Topic Modeling and Temporal-Based Inter-Speaker Topic Modeling in Random Walk for Improved Multi-Party Meeting Summarization
abstract
This paper proposes an improved approach of summarization for spoken multi-party interaction, in which intra-speaker and inter-speaker topics are modeled in a graph constructed with topical relations. Each utterance is represented as a node of the graph and the edge between two nodes is weighted by the similarity between the two utterances, which is topical similarity evaluated by probabilistic latent semantic analysis (PLSA). We model intra-speaker topics by sharing the topics from the same speaker and inter-speaker topics by partially sharing the topics from the adjacent utterances based on temporal information. We did experiments for ASR and manual transcripts. For both transcripts, experiments showed combining intra-speaker and inter-speaker topic modeling can help include the important utterances to offer the improvement for summarization.
Yun-Nung Chen, Florian Metze
INTERSPEECH2
2012 Event-based Video Retrieval Using Audio
abstract
Multimedia Event Detection (MED) is an annual task in the NIST TRECVID evaluation, and requires participants to build indexing and retrieval systems for locating videos in which certain predefined events are shown. Typical systems focus heavily on the use of visual data. Audio data, however, also contains rich information that can be effectively used for video retrieval, and MED could benefit from the attention of researchers in audio analysis. We present several systems for performing MED using only audio data, report the results of each system on the TRECVID MED 2011 development dataset, and compare the strengths and weaknesses of each approach.
Qin Jin, Peter Schulam 0001, Shourabh Rawat, Susanne Burger, Duo Ding, Florian Metze
INTERSPEECH6
2012 The Speech Recognition Virtual Kitchen: An Initial Prototype
Florian Metze, Eric Fosler-Lussier
INTERSPEECH1
2012 Enhanced Polyphone Decision Tree Adaptation for Accented Speech Recognition
abstract
State-of-the-art Automatic Speech Recognition (ASR) models struggle to handle accented speech, particularly if the target accent is under-represented in the training data. The acoustic variations presented by an unfamiliar accent, render the ASR polyphone decision tree (PDT) and its associated Gaussian mixture models (GMM) misfit to the test data. In this paper, we improve on the previous work of adapting the polyphone decision tree, using a semi-continuous model based approach to address the problem of data sparsity. We extend the existing PDT to introduce additional states with shared parameters, corresponding to the new contextual variations identified in the adaptation data, while still robustly estimating the state based parameters on a small adaptation set. We conduct ASR experiments on Arabic and English accents and show that our technique performs better than Maximum A-Posteriori (MAP) adaptation and a previous implementation of polyphone decision tree specialization (PDTS). Compared to MAP adaptation, we obtain 7% relative improvement for Dialectal Arabic and 13.8% relative improvement for Accented English.
Udhyakumar Nallasamy, Florian Metze, Tanja Schultz
INTERSPEECH2
2012 On Speaker-Independent Personality Perception and Prediction from Speech
abstract
In this paper, we present ongoing experiments and insights regarding automatic assessment of perceived personality. While within the INTERSPEECH Speaker Trait Challenge participants will train systems in order to recognize binary targets along the Big 5 personality trait, we will analyze and discuss properties of the data, the labeling scheme and the predictive quality. Conducting factor analyses, estimating reliability, and building regression models capturing dimensions of personality we compare all results to our former and current work and introduce a new extension of our personality database. Eventually, this paper contributes in methodology and understanding on how to asses the perceived personality from an unknown speaker by humans and machines.
Tim Polzehl, Katrin Schoenenberg, Sebastian Möller 0001, Florian Metze, Gelareh Mohammadi, Alessandro Vinciarelli
INTERSPEECH4
2012 Initialization Schemes for Multilayer Perceptron Training and their Impact on ASR Performance using Multilingual Data
Ngoc Thang Vu, Wojtek Breiter, Florian Metze, Tanja Schultz
INTERSPEECH3
2012 Beyond audio and video retrieval: towards multimedia summarization
abstract
Given the deluge of multimedia content that is becoming available over the Internet, it is increasingly important to be able to effectively examine and organize these large stores of information in ways that go beyond browsing or collaborative filtering. In this paper we review previous work on audio and video processing, and define the task of Topic-Oriented Multimedia Summarization (TOMS) using natural language generation: given a set of automatically extracted features from a video (such as visual concepts and ASR transcripts) a TOMS system will automatically generate a paragraph of natural language ("a recounting"), which summarizes the important information in a video belonging to a certain topic area, and provides explanations for why a video was matched and retrieved. We see this as a first step towards systems that will be able to discriminate visually similar, but semantically different videos, compare two videos and provide textual output or summarize a large number of videos at once. In this paper, we introduce our approach of solving the TOMS problem. We extract visual concept features and ASR transcription features from a given video, and develop a template-based natural language generation system to produce a textual recounting based on the extracted features. We also propose possible experimental designs for continuously evaluating and improving TOMS systems, and present results of a pilot evaluation of our initial system.
Duo Ding, Florian Metze, Shourabh Rawat, Peter Schulam 0001, Susanne Burger, Ehsan Younessian, Michael G. Christel, Alex Hauptmann 0001
ICMR2
2012 AMVA'12: ACM international workshop on audio and multimedia methods for large-scale video analysis
abstract
Media sharing sites on the Internet and the one-click upload capability of smartphones have led to a deluge of online multimedia content. Everyday, thousands of videos are uploaded into the web creating an ever-growing demand for methods to make them easier to retrieve, search, and index. While visual information is a very important part of a video, acoustic information often complements it. This is especially true for the analysis of consumer-produced, "unconstrained" videos from social media networks, such as YouTube uploads or Flickr content. The goal of the 1st ACM International Workshop on Audio and Multimedia Methods for Large-Scale Video Analysis (AMVA) is to bring together researchers and practitioners in this newly emerging field, and to foster discussion on future directions of the topic by providing a forum for focused exchanges on new ideas, developments, and results. The aim is to build a strong community and a venue that at some point can become its own conference.
Gerald Friedland, Daniel P. W. Ellis, Florian Metze
ACM Multimedia3
2012 Intra-Speaker Topic Modeling for Improved Multi-Party Meeting Summarization with Integrated Random Walk
Yun-Nung Chen, Florian Metze
HLT-NAACL2
2012 Two-layer mutually reinforced random walk for improved multi-party meeting summarization
abstract
This paper proposes an improved approach of summarization for spoken multi-party interaction, in which a two-layer graph with utterance-to-utterance, speaker-to-speaker, and speaker-to-utterance relations is constructed. Each utterance and each speaker are represented as a node in the utterance-layer and speaker-layer of the graph respectively, and the edge between two nodes is weighted by the similarity between the two utterances, the two speakers, or the utterance and the speaker. The relation between utterances is evaluated by lexical similarity via word overlap or topical similarity via probabilistic latent semantic analysis (PLSA). By within- and between-layer propagation in the graph, the scores from different layers can be mutually reinforced so that utterances can automatically share the scores with the utterances from the same speaker and similar utterances. For both ASR output and manual transcripts, experiments confirmed the efficacy of involving speaker information in the two-layer graph for summarization.
Yun-Nung Chen, Florian Metze
SLT2
2012 Active learning for accent adaptation in Automatic Speech Recognition
abstract
We experiment with active learning for speech recognition in the context of accent adaptation. We adapt a source recognizer on the target accent by selecting a relatively small, matched subset of utterances from a large, untranscribed and multi-accented corpus for human transcription. Traditionally, active learning in speech recognition has relied on uncertainty based sampling to choose the most informative data for manual labeling. Such an approach doesn't include explicit relevance criterion during data selection, which is crucial for choosing utterances to match the target accent, from datasets with wide-ranging speakers of different accents. We formulate a cross-entropy based relevance measure to complement uncertainty based sampling for active learning to aid accent adaptation. We evaluate the algorithm on two different setups for Arabic and English accents and show that our approach performs favorably to conventional data selection. We analyze the results to show the effectiveness of our approach in finding the most relevant subset of utterances for improving the speech recognizer on the target accent.
Udhyakumar Nallasamy, Florian Metze, Tanja Schultz
SLT2
2011 Analysis of Dialectal Influence in Pan-Arabic ASR
abstract
In this paper, we analyze the impact of five Arabic dialects on the front-end and pronunciation dictionary component of an Automatic Speech Recognition (ASR) system.We use ASR"s phonetic decision tree as a diagnostic tool to compare the robustness of MFCC to MLP front-ends to dialectal variations in the speech data and found that MLP Bottle-Neck features are less robust to dialectal variation.We also perform a rulebased analysis of the pronunciation dictionary, which enables us to identify dialectal words in the vocabulary and automatically generate pronunciations for unseen words.We show that our technique produces pronunciations with an average phone error rate 9.2%.
Udhyakumar Nallasamy, Michael Garbus, Florian Metze, Qin Jin, Thomas Schaaf, Tanja Schultz
INTERSPEECH3
2011 Modeling Speaker Personality Using Voice
abstract
In this paper, we validate the application of an established personality assessment and modeling paradigm to speech input, and extend earlier work towards text independent speech input. We show that human labelers can consistently label acted speech data generated across multiple recording sessions, and investigate further which of the 5 scales in the NEO-FFI scheme can be assessed from speech, and how a manipulation of one scale influences the perception of another. Finally, we present a clustering of human labels of perceived personality traits, which will be useful in future experiments on automatic classification and generation of personality traits from speech.
Tim Polzehl, Sebastian Möller 0001, Florian Metze
INTERSPEECH3
2011 Anger recognition in speech using acoustic and linguistic cues
Tim Polzehl, Alexander Schmitt, Florian Metze, Michael Wagner 0004
Speech Commun.3
2010 Late fusion of individual engines for improved recognition of negative emotion in speech - learning vs. democratic vote
abstract
The fusion of multiple recognition engines is known to be able to outperform individual ones, given sufficient independence of methods, models, and knowledge sources. We therefore investigate late fusion of different speech-based recognizers of emotion. Two generally different streams of information are considered: acoustics and linguistics fed by state-of-the-art automatic speech recognition. A total of five emotion recognition engines from different sites that provide heterogeneous output information are integrated by either simple democratic vote or learning `which predictor to trust when'. We are able to significantly outperform the best individual engine by fusion, and the so far best reported result on the recently introduced Emotion Challenge task.
Björn W. Schuller, Florian Metze, Stefan Steidl, Anton Batliner, Florian Eyben, Tim Polzehl
ICASSP2
2010 Improvements to generalized discriminative feature transformation for speech recognition
abstract
Generalized Discriminative Feature Transformation (GDFT) is a feature space discriminative training algorithm for automatic speech recognition (ASR). GDFT uses Lagrange relaxation to transform the constrained maximum likelihood linear regression (CMLLR) algorithm for feature space discriminative training. This paper presents recent improvements on GDFT, which are achieved by regularization to the optimization problem. The resulting algorithm is called regularized GDFT (rGDFT) and we show that many regularization and smoothing techniques developed for model space discriminative training are also applicable to feature space training. We evaluated rGDFT on a real-time Iraqi ASR system and also on a large scale Arabic ASR task.
Roger Hsiao, Florian Metze, Tanja Schultz
INTERSPEECH2
2010 Emotion recognition using imperfect speech recognition
abstract
This paper investigates the use of speech-to-text methods for assigning an emotion class to a given speech utterance. Previous work shows that an emotion extracted from text can convey complementary evidence to the information extracted by classifiers based on spectral, or other non-linguistic features. As speech-to-text usually presents significantly more computational effort, in this study we investigate the degree of speech-to-text accuracy needed for reliable detection of emotions from an automatically generated transcription of an utterance. We evaluate the use of hypotheses in both training and testing, and compare several classification approaches on the same task. Our results show that emotion recognition performance stays roughly constant as long as word accuracy doesn't fall below a reasonable value, making the use of speech-to-text viable for training of emotion classifiers based on linguistics.
Florian Metze, Anton Batliner, Florian Eyben, Tim Polzehl, Björn W. Schuller, Stefan Steidl
INTERSPEECH1
2010 The 2010 CMU GALE speech-to-text system
abstract
This paper describes the latest Speech-to-Text system developed for the Global Autonomous Language Exploitation ("GALE") domain by Carnegie Mellon University (CMU). This systems uses discriminative training, bottle-neck features and other techniques that were not used in previous versions of our system, and is trained on 1150 hours of data from a variety of Arabic speech sources. In this paper, we show how different lexica, pre-processing, and system combination techniques can be used to improve the final output, and provide analysis of the improvements achieved by the individual techniques.
Florian Metze, Roger Hsiao, Qin Jin, Udhyakumar Nallasamy, Tanja Schultz
INTERSPEECH1
2010 Analysis of gender normalization using MLP and VTLN features
abstract
This paper analyzes the capability of multilayer perceptron frontends to perform speaker normalization. We find the context decision tree to be a very useful tool to assess the speaker normalization power of different frontends. We introduce a gender question into the training of the phonetic context decision tree. After the context clustering the gender specific models are counted. We compare this for the following frontends: (1) Bottle-Neck (BN) with and without vocal tract length normalization (VTLN), (2) standard MFCC, (3) stacking of multiple MFCC frames with linear discriminant analysis (LDA). We find the BN-frontend to be even more effective in reducing the number of gender questions than VTLN. From this we conclude that a Bottle-Neck frontend is more effective for gender normalization. Combining VTLN and BN-features reduces the number of gender specific models further.
Thomas Schaaf, Florian Metze
INTERSPEECH2
2010 Multimedia content with a speech track: ACM multimedia 2010 workshop on searching spontaneous conversational speech
abstract
No abstract available.
Martha A. Larson, Roeland Ordelman, Florian Metze, Wessel Kraaij, Franciska de Jong
ACM Multimedia3
2010 Automatically assessing acoustic manifestations of personality in speech
abstract
In this paper, we present first results on applying a personality assessment paradigm to speech input, and comparing human and automatic performance on this task. We cue a professional speaker to produce speech using different personality profiles and encode the resulting vocal personality impressions in terms of the Big Five NEO-FFI personality traits. We then have human raters, who do not know the speaker, estimate the five factors. We analyze the recordings using signal-based acoustic and prosodic methods and observe high consistency between the acted personalities, the raters' assessments, and initial automatic classification results. This presents a first step towards being able to handle personality traits in speech, which we envision will be used in future voice-based communication between humans and machines.
Tim Polzehl, Sebastian Möller 0001, Florian Metze
SLT3
2009 Detecting real life anger
abstract
Acoustic anger detection in voice portals can help to enhance human computer interaction. A comprehensive voice portal data collection has been carried out and gives new insight on the nature of real life data. Manual labeling revealed a high percentage of non-classifiable data. Experiments with a statistical classifier indicate that, in contrast to pitch and energy related features, duration measures do not play an important role for this data while cepstral information does. Also in a direct comparison between Gaussian Mixture Models and Support Vector Machines the latter gave better results.
Felix Burkhardt, Tim Polzehl, Joachim Stegmann, Florian Metze, Richard Huber
ICASSP4
2009 Emotion classification in children's speech using fusion of acoustic and linguistic features
abstract
This paper describes a system to detect angry vs. non-angry utterances of children who are engaged in dialog with an Aibo robot dog. The system was submitted to the Interspeech2009 Emotion Challenge evaluation. The speech data consist of short utterances of the children’s speech, and the proposed system is designed to detect anger in each given chunk. Frame-based cepstral features, prosodic and acoustic features as well as glottal excitation features are extracted automatically, reduced in dimensionality and classified by means of an artificial neural network and a support vector machine. An automatic speech recognizer transcribes the words in an utterance and yields a separate classification based on the degree of emotional salience of the words. Late fusion is applied to make a final decision on anger vs. non-anger of the utterance. Preliminary results show 75.9% unweighted average recall on the training data and 67.6 % on the test set. Index Terms: speech processing, meta-data extraction, emotion recognition, evaluation
Tim Polzehl, Shiva Sundaram, Hamed Ketabdar, Michael Wagner 0004, Florian Metze
INTERSPEECH5
2009 Influence of training on direct and indirect measures for the evaluation of multimodal systems
Julia Seebode, Stefan Schaffer, Ina Wechsung, Florian Metze
INTERSPEECH4
2009 Predicting the quality of multimodal systems based on judgments of single modalities
Ina Wechsung, Klaus-Peter Engelbrecht, Anja Naumann, Stefan Schaffer, Julia Seebode, Florian Metze, Sebastian Möller 0001
INTERSPEECH6
2008 Detecting trends in social bookmarking systems using a probabilistic generative model and smoothing
abstract
We propose a method for the detection of trends in social bookmarking systems. Compared to other work in this emerging field, our approach has a more sound statistical basis. In order to cope with the problem of vanishing probabilities due to data sparsity, we apply smoothing and show that it allows for an easy calibration of our trend detector resulting in better generalization and scalability. We test our approach on a collection of 105, 000, 000 bookmarks collected from the del.icio.us bookmarking service. To our knowledge, this is the largest corpus of a real world bookmarking service analyzed in this context. The results show that our method outperforms previously proposed methods and successfully detects trends in the data.
Robert Wetzker, Till Plumbaum, Alexander Korth, Christian Bauckhage, Tansu Alpcan, Florian Metze
ICPR6
2008 User perception of multi-modal interfaces for mobile applications
abstract
This paper presents a comparative study on the usability of a service presented in telephone, PC-based web interface, and mobile/ multi-modal variants. The goal is not to analyze individual strengths and weaknesses of the different modalities, but to understand the user’s perception of the SUMI criteria (efficiency, affect/ likability, helpfulness, control, learnability), and the overall impression of a service with respect to the access variant tested. As multi-modality is often framed as a technology to make usage more “intuitive”, we were particularly interested in the differences between experienced and novice users. To this end, we conducted a study with 80 participants and conclude that, while multi-modality is accepted by experienced users, it seems to be asking too much from novice users, particularly with respect to learnability and efficiency.
Florian Metze, Roman Englert, Udo Bub, Ingmar Kliche, Thomas Scheerbarth
INTERSPEECH1
2007 Spotting using Durational Entropy
abstract
This paper deals with the task of detection of a given keyword in continuous speech. We build upon a previously proposed algorithm where a modified Viterbi search algorithm is used to detect keywords, without requiring any explicit garbage or filler models. In this work, the concept of durational entropy is used to further discard a large fraction of false alarm errors. Durational entropy is defined as the entropy of the distribution of state occupancies. A method to recursively compute it for all Viterbi paths is also presented in this paper. Experimental results on one hour of broadcast news data suggest that durational entropy constraints can indeed be used to avoid a large number of false alarms errors at a minimal cost of degradation in keyword detection accuracy.
Jitendra Ajmera, Florian Metze
ICASSP (4)2
2007 Comparison of Four Approaches to Age and Gender Recognition for Telephone Applications
abstract
This paper presents a comparative study of four different approaches to automatic age and gender classification using seven classes on a telephony speech task and also compares the results with human performance on the same data. The automatic approaches compared are based on (1) a parallel phone recognizer, derived from an automatic language identification system; (2) a system using dynamic Bayesian networks to combine several prosodic features; (3) a system based solely on linear prediction analysis; and (4) Gaussian mixture models based on MFCCs for separate recognition of age and gender. On average, the parallel phone recognizer performs as well as Human listeners do, while loosing performance on short utterances. The system based on prosodic features however shows very little dependence on the length of the utterance.
Florian Metze, Jitendra Ajmera, Roman Englert, Udo Bub, Felix Burkhardt, Joachim Stegmann, Christian Müller 0014, Richard Huber, Bernt Andrassy, Josef G. Bauer, Bernhard Littel
ICASSP (4)1
2007 An intelligent knowledge sharing system for web communities
abstract
This paper presents an expert peering system for information exchange in the knowledge society. Our system realizes an intelligent, real-time search engine for enterprise Intranets or online communities that automatically relays user queries to knowledgable specialists. According to its very nature, the system requires a sound integration of concepts drawn from various areas of Computer Science. In addition to our solutions to problems in scalable processing, data transfer, and networking, we also address issues of interface design and usability, as well as aspects of machine intelligence. Results obtained from extensive experiments demonstrate the efficiency and robustness of our system and convey its potential for next generation web services.
Christian Bauckhage, Tansu Alpcan, Sachin Agarwal 0001, Florian Metze, Robert Wetzker, Milena Ilic, Sahin Albayrak
SMC4
2007 Discriminative speaker adaptation using articulatory features
Florian Metze
Speech Commun.1
2006 Articulatory features for "meeting" speech recognition
abstract
“Meeting ” speech, for example from the RT-04S task, contains a mixture of different speaking styles that leads to word error rates higher than 25 % even when close-talking microphones are being used. The problem is even more serious, as word error rates are particularly high when speakers use a clear speaking mode, for example because they want to stress an important point. Previ-ous work showed that an approach that combines standard phone-based acoustic models with models detecting the presence or ab-sence of “Articulatory Features ” such as “Rounded ” or “Voiced” can improve ASR performance particularly for these cases. This paper presents a discriminative approach to automatically comput-ing from training or adaptation data the feature stream weights needed for the above approach, therefore presenting a framework for integrating articulatory features into existing automatic speech recognition systems. We find a 7 % relative improvements on top of our best RT-04S system using discriminative adaptation. 1.
Florian Metze
INTERSPEECH1
2005 Automatically Transcribing Meetings using Distant Microphones
abstract
In this paper, we describe our efforts to develop acoustic models suitable for distant microphone automatic speech recognition. Our goal is to investigate how the performance of a system trained on a combination of close-talking and distant microphone data can be optimized, while assuming as little information about the configuration of (multiple) distant microphones as possible, to avoid guesstimates and lengthy calibration runs. We evaluated our system in NIST's RT-04S "Meeting" speech-to-text evaluation, where speech data was recorded at several sites with a varying number of different table-top microphones, but not with microphone arrays. Body-mounted microphones provide baseline numbers for distant ASR performance and allow for comparisons of meeting speech with other spontaneous speech data.
Florian Metze, Christian Fügen, Alex Waibel
ICASSP (1)1
2004 The 2003 ISL rich transcription system for conversational telephony speech
abstract
This paper describes the ISL large vocabulary conversational telephony speech recognition system, which was tested in NIST's RT-03S ("Switchboard") evaluation. We present our experiments on improving preprocessing, acoustic modelling, and language modelling. The system features phone-dependent semi-tied full covariances, semi-tied clustering of septa-phones, clustering across phones, feature adaptive training, robust estimation of VTLN and MLLR, as well as context-dependent interpolation of language models. We present detailed results for each stage of our multi-pass transcription scheme. System development started with a 1997 SWB system, yielding a word error rate of 35.1% on our internal 1h development set. The final system performed at 21.8%, a 38% relative improvement. The error rate on the RT-03 CTS evaluation set is 23.4%.
Hagen Soltau, Hua Yu 0008, Florian Metze, Christian Fügen, Qin Jin, Szu-Chen Stan Jou
ICASSP (1)3
2004 Issues in meeting transcription - the ISL meeting transcription system
abstract
This paper describes the Interactive Systems Lab’s Meeting transcription system, which performs segmentation, speaker clustering as well as transcriptions of conversational meeting speech. The system described here was evaluated in NIST’s RT-04S “Meeting” speech evaluation. This paper compares the performance of our Broadcast News and the most recent Switchboard system on the Meeting data and compares both with a newly-trained meeting recognizer. Furthermore we investigate the effects of automatic segmentation on adaptation. Our best meeting system achieves 44.5% on the MDM condition in NIST’s RT-04S evaluation.
Tanja Schultz, Qin Jin, Kornel Laskowski, Florian Metze, Christian Fügen
INTERSPEECH5
2003 Multilingual articulatory features
abstract
Speech recognition systems based on or aided by articulatory features, such as place and manner of articulation, have been shown to be useful under varying circumstances. Recognizers based on features better compensate channel and noise variability. We show that it is also possible to compensate for inter language variability using articulatory feature detectors. We come to the conclusion that articulatory features can be recognized across languages and that using detectors from many languages can improve the classification accuracy of the feature detectors on a single language. We further demonstrate how those multilingual and cross-lingual detectors can support an HMM based recognizer and thereby significantly reduce the word error rate by up to 12.3% relative. We expect that with the use of multilingual articulatory features it is possible to support the rapid deployment of recognition systems for new target languages.
Sebastian Stüker, Tanja Schultz, Florian Metze, Alex Waibel
ICASSP (1)3
2003 The NESPOLE! voIP multilingual corpora in tourism and medical domains
abstract
In this paper we present the multilingual VoIP (Voice over Internet Protocol networks) corpora collected for the second showcase of the Nespole! project in the tourism and medical domains. The corpora comprise over 20 hours of human-to-human monolingual dialogues in English, French, German and Italian: 66 dialogues in the tourism domain and 49 in the medical domain. We describe in detail the data collection (technical set-up, scenarios for each domain, recording procedure and data transcription), as well as statistically illustrated corpora and a preliminary data analysis. 1.
Nadia Mana, Susanne Burger, Roldano Cattoni, Laurent Besacier, Victoria MacLaren, John W. McDonough, Florian Metze
INTERSPEECH7
2003 Integrating multilingual articulatory features into speech recognition
abstract
The use of articulatory features, such as place and manner of articulation, has been shown to reduce the word error rate of speech recognition systems under different conditions and in different settings.For example recognition systems based on features are more robust to noise and reverberation.In earlier work we showed that articulatory features can compensate for inter language variability and can be recognized across languages.In this paper we show that using cross-and multilingual detectors to support an HMM based speech recognition system significantly reduces the word error rate.By selecting and weighting the features in a discriminative way, we achieve an error rate reduction that lies in the same range as that seen when using language specific feature detectors.By combining feature detectors from many languages and training the weights discriminatively, we even outperform the case where only monolingual detectors are being used.
Sebastian Stüker, Florian Metze, Tanja Schultz, Alex Waibel
INTERSPEECH2
2002 Efficient language model lookahead through polymorphic linguistic context assignment
abstract
In this study, we examine how fast decoding of conversational speech with large vocabularies profits from efficient use of linguistic information, i.e. language models and grammars. Based on a re-entrant single pronunciation prefix tree, we use the concept of linguistic context polymorphism to achieve an early incorporation of language model information. This approach allows us to use all available language model information in a one-pass decoder, using the same engine to decode with statistical n-gram language models as well as context free grammars or re-scoring of lattices in an efficient way. We compare this approach to our previous decoder, which needed three passes to incorporate all available information. The results on a very large vocabulary task show that the search can be speeded up by almost a factor of three, without introducing additional search errors. On all examined tasks, we observed significant improvements by using an exact language model lookahead over usual bigram lookahead strategies, even for very hard tasks with unmatched conditions, without introducing extra memory overhead.
Hagen Soltau, Florian Metze, Christian Fügen, Alex Waibel
ICASSP2
2002 A flexible stream architecture for ASR using articulatory features
Florian Metze, Alex Waibel
INTERSPEECH1
2002 Compensating for hyperarticulation by modeling articulatory properties
Hagen Soltau, Florian Metze, Alex Waibel
INTERSPEECH2
2001 Speaker compensation with sine-log all-pass transforms
abstract
In previous work, we proposed the rational all-pass transform (RAPT) as the basis of a speaker adaptation scheme intended for use with a large vocabulary speech recognition system. It was shown that RAPT-based adaptation reduces to a linear transformation of cepstral means, much like the better known maximum likelihood linear regression (MLLR). In a set of speech recognition experiments conducted on the Switchboard Corpus, we obtained a word error rate (WER) of 37.9% using RAPT adaptation, a significant improvement over the 39.5% WER achieved with MLLR. In the present work, we propose the sine-log all-pass transform (SLAPT) as a replacement for the RAPT. Our findings indicate the SLAPT is just as effective as the RAPT at reducing WER when used as the basis for a variety of speaker compensation schemes, but in addition conduces to far more tractable computation of transformed cepstral sequences, and the estimation of optimal transform parameters.
John W. McDonough, Florian Metze, Hagen Soltau, Alex Waibel
ICASSP2
2001 The ISL evaluation system for Verbmobil-II
abstract
Describes the 2000 ISL large vocabulary speech recognition system for fast decoding of conversational speech which was used in the German Verbmobil-II project. The challenge of this task is to build robust acoustic models to handle different dialects, spontaneous effects, and crosstalk as occur in conversational speech. We present speaker incremental normalization and adaptation experiments close to real-time constraints. To reduce the number of consequential errors caused by out-of-vocabulary words, we conducted filler-model experiments to handle unknown proper names. The overall improvements from 1998 to 2000 resulted in a word error reduction from 40% to 17% on our development test set.
Hagen Soltau, Thomas Schaaf, Florian Metze, Alex Waibel
ICASSP3
2001 Advances in automatic meeting record creation and access
abstract
Oral communication is transient, but many important decisions, social contracts and fact findings are first carried out in an oral setup, documented in written form and later retrieved. At Carnegie Mellon University's Interactive Systems Laboratories we have been experimenting with the documentation of meetings. The paper summarizes part of the progress that we have made in this test bed, specifically on the question of automatic transcription using large vocabulary continuous speech recognition, information access using non-keyword based methods, summarization and user interfaces. The system is capable of automatically constructing a searchable and browsable audio-visual database of meetings and provide access to these records.
Alex Waibel, Michael Bett, Florian Metze, Klaus Ries 0001, Thomas Schaaf, Tanja Schultz, Hagen Soltau, Hua Yu 0008, Klaus Zechner
ICASSP3
2001 The nespole! voIP dialogue database
abstract
This paper presents the status of the NESPOLE! data collection as of end of February, 2001. A multilingual VoIP (Voice over Internet Protocol networks) database consisting of 200 dialogues in 4 languages (English, German, Italian and French) was recorded and transcribed. Dialogue speakers were connected via a H323 video-conferencing terminal. We describe the task, the technical architecture, the recording procedure and the transcription process of the NESPOLE! data collection. We provide some statistics concerning the data and, finally, we address problems that arose during the collection and annotation process. 1.
Susanne Burger, Laurent Besacier, Paolo Coletti, Florian Metze, Céline Morel
INTERSPEECH4
2001 Speech recognition over netmeeting connections
abstract
In this paper we evaluate the performance of the ISL's German Verbmobil spontaneous speech recognizer on the Nespole! database. In this task, people talk to an agent in a tourist office to plan their holidays via a NetMeeting connection, also sharing screen contents (web-pages). Stereo recordings were made both before and after speech transmission over an IP connection using the G.711 codec, so that we are able to directly measure the loss in LVCSR performance due to NetMeeting's segmentation and compression. The aim of this work is to quantify this loss, which is a consequence of using protocols which were not designed for speech recognition purposes. We report on techniques employed to port our existing clean-speech recognizer to this new data quality, using about 1.5h of labeled adaptation data, but avoiding a complete retraining of the system.
Florian Metze, John W. McDonough, Hagen Soltau
INTERSPEECH1
2000 Confidence measure based language identification
abstract
In this paper we present a new application for confidence measures in spoken language processing. In today's computerized dialogue systems, language identification (LID) is typically achieved via dedicated modules. In our approach, LID is integrated into the speech recognizer, therefore profiting from high-level linguistic knowledge at very little extra cost. Our new approach is based on a word lattice based confidence measure (Kemp and Schaaf, 1997), which was originally devised for unsupervised training. In this work, we show that the confidence based language identification algorithm outperforms conventional score based methods. Also, this method is less dependent on the acoustic characteristics of the transmission channel than score based methods. By introducing additional parameters, unknown languages can be rejected. The proposed method is compared to a score based approach on the Verbmobil database, a three language task.
Florian Metze, Thomas Kemp, Thomas Schaaf, Tanja Schultz, Hagen Soltau
ICASSP1
2000 Generalized radial basis function networks for classification and novelty detection: self-organization of optimal Bayesian decision
Sebastian Albrecht 0002, Jan Busch, Martin Kloppenburg, Florian Metze, Paul Tavan
Neural Networks4