Holger Schwenk

dblp:92/6322 · DBLP profile ↗
← Back
66ranked-venue papers
23as first author
12since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 59 · 21 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 10 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2023 BLASER: A Text-Free Speech-to-Speech Translation Evaluation Metric
abstract
Mingda Chen, Paul-Ambroise Duquenne, Pierre Andrews, Justine Kao, Alexandre Mourachko, Holger Schwenk, Marta R. Costa-jussà. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Mingda Chen, Paul-Ambroise Duquenne, Pierre Andrews, Justine Kao, Alexandre Mourachko, Holger Schwenk, Marta R. Costa-jussà
ACL (1)6
2023 SpeechMatrix: A Large-Scale Mined Corpus of Multilingual Speech-to-Speech Translations
abstract
Paul-Ambroise Duquenne, Hongyu Gong, Ning Dong, Jingfei Du, Ann Lee, Vedanuj Goswami, Changhan Wang, Juan Pino, Benoît Sagot, Holger Schwenk. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Paul-Ambroise Duquenne, Hongyu Gong, Jingfei Du, Ann Lee 0001, Vedanuj Goswami, Changhan Wang, Juan Pino 0001, Benoît Sagot, Holger Schwenk
ACL (1)10
2023 Multilingual Representation Distillation with Contrastive Learning
abstract
Multilingual sentence representations from large models encode semantic information from two or more languages and can be used for different cross-lingual information retrieval and matching tasks.In this paper, we integrate contrastive learning into multilingual representation distillation and use it for quality estimation of parallel sentences (i.e., find semantically similar sentences that can be used as translations of each other).We validate our approach with multilingual similarity search and corpus filtering tasks.Experiments across different low-resource languages show that our method greatly outperforms previous sentence encoders such as LASER, LASER3, and LaBSE.
Weiting Tan, Kevin Heffernan, Holger Schwenk, Philipp Koehn
EACL3
2023 DiffEdit: Diffusion-based semantic image editing with mask guidance
Guillaume Couairon, Jakob Verbeek, Holger Schwenk, Matthieu Cord
ICLR3
2023 Modular Speech-to-Text Translation for Zero-Shot Cross-Modal Transfer
abstract
International audience
Paul-Ambroise Duquenne, Holger Schwenk, Benoît Sagot
INTERSPEECH2
2022 FlexIT: Towards Flexible Semantic Image Translation
abstract
Deep generative models, like GANs, have considerably improved the state of the art in image synthesis, and are able to generate near photo-realistic images in structured domains such as human faces. Based on this success, recent work on image editing proceeds by projecting images to the GAN latent space and manipulating the latent vector. However, these approaches are limited in that only images from a narrow domain can be transformed, and with only a limited number of editing operations. We propose FlexIT, a novel method which can take any input image and a user-defined text instruction for editing. Our method achieves flexible and natural editing, pushing the limits of semantic image translation. First, FlexIT combines the input image and text into a single target point in the CLIP multimodal embedding space. Via the latent space of an autoencoder, we iteratively transform the input image toward the target point, ensuring coherence and quality with a variety of novel regularization terms. We propose an evaluation protocol for semantic image translation, and thoroughly evaluate our method on ImageNet. Code will be available at https://github.com/facebookresearch/SemanticImageTranslation/.
Guillaume Couairon, Asya Grechka, Jakob Verbeek, Holger Schwenk, Matthieu Cord
CVPR4
2022 T-Modules: Translation Modules for Zero-Shot Cross-Modal Machine Translation
abstract
We present a new approach to perform zeroshot cross-modal transfer between speech and text for translation tasks.Multilingual speech and text are encoded in a joint fixed-size representation space.Then, we compare different approaches to decode these multimodal and multilingual fixed-size representations, enabling zero-shot translation between languages and modalities.All our models are trained without the need of cross-modal labeled translation data.Despite a fixed-size representation, we achieve very competitive results on several text and speech translation tasks.In particular, we outperform the state of the art for zero-shot speech translation on Must-C.We also introduce the first results for zero-shot direct speechto-speech and text-to-speech translation.
Paul-Ambroise Duquenne, Hongyu Gong, Benoît Sagot, Holger Schwenk
EMNLP4
2022 Textless Speech-to-Speech Translation on Real Data
abstract
Ann Lee, Hongyu Gong, Paul-Ambroise Duquenne, Holger Schwenk, Peng-Jen Chen, Changhan Wang, Sravya Popuri, Yossi Adi, Juan Pino, Jiatao Gu, Wei-Ning Hsu. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Ann Lee 0001, Hongyu Gong, Paul-Ambroise Duquenne, Holger Schwenk, Peng-Jen Chen, Changhan Wang, Sravya Popuri, Yossi Adi, Juan Pino 0001, Jiatao Gu, Wei-Ning Hsu
NAACL-HLT4
2021 CCMatrix: Mining Billions of High-Quality Parallel Sentences on the Web
abstract
Holger Schwenk, Guillaume Wenzek, Sergey Edunov, Edouard Grave, Armand Joulin, Angela Fan. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Holger Schwenk, Guillaume Wenzek, Sergey Edunov, Edouard Grave, Armand Joulin, Angela Fan
ACL/IJCNLP (1)1
2021 WikiMatrix: Mining 135M Parallel Sentences in 1620 Language Pairs from Wikipedia
abstract
Holger Schwenk, Vishrav Chaudhary, Shuo Sun, Hongyu Gong, Francisco Guzmán. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021.
Holger Schwenk, Vishrav Chaudhary, Hongyu Gong, Francisco Guzmán
EACL1
2021 Multimodal and Multilingual Embeddings for Large-Scale Speech Mining
abstract
We present an approach to encode a speech signal into a fixed-size representation which minimizes the cosine loss with the existing massively multilingual LASER text embedding space. Sentences are close in this embedding space, independently of their language and modality, either text or audio. Using a similarity metric in that multimodal embedding space, we perform mining of audio in German, French, Spanish and English from Librivox against billions of sentences from Common Crawl. This yielded more than twenty thousand hours of aligned speech translations. To evaluate the automatically mined speech/text corpora, we train neural speech translation systems for several languages pairs. Adding the mined data, achieves significant improvements in the BLEU score on the CoVoST2 and the MUST-C test sets with respect to a very competitive baseline. Our approach can also be used to directly perform speech-to-speech mining, without the need to first transcribe or translate the data. We obtain more than one thousand three hundred hours of aligned speech in French, German, Spanish and English. This speech corpus has the potential to boost research in speech-to-speech translation which suffers from scarcity of natural end-to-end training data. All the mined multimodal corpora will be made freely available.
Paul-Ambroise Duquenne, Hongyu Gong, Holger Schwenk
NeurIPS3
2021 Beyond English-Centric Multilingual Machine Translation
abstract
Existing work in translation demonstrated the potential of massively multilingual machine translation by training a single model able to translate between any pair of languages. However, much of this work is English-Centric, training only on data which was translated from or to English.While this is supported by large sources of training data, it does not reflect translation needs worldwide. In this work, we create a true Many-to-Many multilingual translation model that can translate directly between any pair of 100 languages. We build and open-source a training data set that covers thousands of language directions with parallel data, created through large-scale mining. Then, we explore how to effectively increase model capacity through a combination of dense scaling and language-specific sparse parameters to create high quality models. Our focus on non-English-Centric models brings gains of more than 10 BLEU when directly translating between non-English directions while performing competitively to the best single systems from the Workshop on Machine Translation (WMT). We open-source our scripts so that others may reproduce the data, evaluation, and final M2M-100 model.
Angela Fan, Shruti Bhosale, Holger Schwenk, Zhiyi Ma, Ahmed El-Kishky, Siddharth Goyal, Mandeep Baines, Onur Celebi, Guillaume Wenzek, Vishrav Chaudhary, Naman Goyal 0001, Tom Birch, Vitaliy Liptchinsky, Sergey Edunov, Michael Auli, Armand Joulin
J. Mach. Learn. Res.3
2020 MLQA: Evaluating Cross-lingual Extractive Question Answering
abstract
Question answering (QA) models have shown rapid progress enabled by the availability of large, high-quality benchmark datasets.Such annotated datasets are difficult and costly to collect, and rarely exist in languages other than English, making building QA systems that work well in other languages challenging.In order to develop such systems, it is crucial to invest in high quality multilingual evaluation benchmarks to measure progress.We present MLQA, a multi-way aligned extractive QA evaluation benchmark intended to spur research in this area.1 MLQA contains QA instances in 7 languages, English, Arabic, German, Spanish, Hindi, Vietnamese and Simplified Chinese.MLQA has over 12K instances in English and 5K in each other language, with each instance parallel between 4 languages on average.We evaluate stateof-the-art cross-lingual models and machinetranslation-based baselines on MLQA.In all cases, transfer results are significantly behind training-language performance.
Patrick S. H. Lewis, Barlas Oguz, Ruty Rinott, Sebastian Riedel 0001, Holger Schwenk
ACL5
2020 Searching the Web for Cross-lingual Parallel Data
abstract
While the World Wide Web provides a large amount of text in many languages, cross-lingual parallel data is more difficult to obtain. Despite its scarcity, this parallel cross-lingual data plays a crucial role in a variety of tasks in natural language processing with applications in machine translation, cross-lingual information retrieval, and document classification, as well as learning cross-lingual representations. Here, we describe the end-to-end process of searching the web for parallel cross-lingual texts. We motivate obtaining parallel text as a retrieval problem whereby the goal is to retrieve cross-lingual parallel text from a large, multilingual web-crawled corpus. We introduce techniques for searching for cross-lingual parallel data based on language, content, and other metadata. We motivate and introduce multilingual sentence embeddings as a core tool and demonstrate techniques and models that leverage them for identifying parallel documents and sentences as well as techniques for retrieving and filtering this data. We describe several large-scale datasets curated using these techniques and show how training on sentences extracted from parallel or comparable documents mined from the Web can improve machine translation models and facilitate cross-lingual NLP.
Ahmed El-Kishky, Philipp Koehn, Holger Schwenk
SIGIR3
2019 Analysis of Joint Multilingual Sentence Representations and Semantic K-Nearest Neighbor Graphs
Holger Schwenk, Douwe Kiela, Matthijs Douze
AAAI1
2019 Margin-based Parallel Corpus Mining with Multilingual Sentence Embeddings
abstract
Machine translation is highly sensitive to the size and quality of the training data, which has led to an increasing interest in collecting and filtering large parallel corpora. In this paper, we propose a new method for this task based on multilingual sentence embeddings. In contrast to previous approaches, which rely on nearest neighbor retrieval with a hard threshold over cosine similarity, our proposed method accounts for the scale inconsistencies of this measure, considering the margin between a given sentence pair and its closest candidates instead. Our experiments show large improvements over existing methods. We outperform the best published results on the BUCC mining task and the UN reconstruction task by more than 10 F1 and 30 precision points, respectively. Filtering the English-German ParaCrawl corpus with our approach, we obtain 31.2 BLEU points on newstest2014, an improvement of more than one point over the best official filtered version.
Mikel Artetxe, Holger Schwenk
ACL (1)2
2019 Massively Multilingual Sentence Embeddings for Zero-Shot Cross-Lingual Transfer and Beyond
abstract
We introduce an architecture to learn joint multilingual sentence representations for 93 languages, belonging to more than 30 different families and written in 28 different scripts. Our system uses a single BiLSTM encoder with a shared BPE vocabulary for all languages, which is coupled with an auxiliary decoder and trained on publicly available parallel corpora. This enables us to learn a classifier on top of the resulting embeddings using English annotated data only, and transfer it to any of the 93 languages without any modification. Our experiments in cross-lingual natural language inference (XNLI dataset), cross-lingual document classification (MLDoc dataset) and parallel corpus mining (BUCC dataset) show the effectiveness of our approach. We also introduce a new test set of aligned sentences in 112 languages, and show that our sentence embeddings obtain strong results in multilingual similarity search even for low-resource languages. Our implementation, the pre-trained encoder and the multilingual test set are available at https://github.com/facebookresearch/LASER
Mikel Artetxe, Holger Schwenk
Trans. Assoc. Comput. Linguistics2
2018 XNLI: Evaluating Cross-lingual Sentence Representations
abstract
Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel Bowman, Holger Schwenk, Veselin Stoyanov. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. 2018.
Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel R. Bowman, Holger Schwenk, Veselin Stoyanov
EMNLP6
2018 A Corpus for Multilingual Document Classification in Eight Languages
Holger Schwenk, Xian Li 0003
LREC1
2017 Very Deep Convolutional Networks for Text Classification
abstract
The dominant approach for many NLP tasks are recurrent neural networks, in particular LSTMs, and convolutional neural networks. However, these architectures are rather shallow in comparison to the deep convolutional networks which have pushed the state-of-the-art in computer vision. We present a new architecture (VDCNN) for text processing which operates directly at the character level and uses only small convolutions and pooling operations. We are able to show that the performance of this model increases with depth: using up to 29 convolutional layers, we report improvements over the state-of-the-art on several public text classification tasks. To the best of our knowledge, this is the first time that very deep convolutional nets have been applied to text processing.
Alexis Conneau, Holger Schwenk, Loïc Barrault, Yann LeCun
EACL (1)2
2017 Supervised Learning of Universal Sentence Representations from Natural Language Inference Data
abstract
Many modern NLP systems rely on word embeddings, previously trained in an unsupervised manner on large corpora, as base features.Efforts to obtain embeddings for larger chunks of text, such as sentences, have however not been so successful.Several attempts at learning unsupervised representations of sentences have not reached satisfactory enough performance to be widely adopted.In this paper, we show how universal sentence representations trained using the supervised data of the Stanford Natural Language Inference datasets can consistently outperform unsupervised methods like SkipThought vectors (Kiros et al., 2015) on a wide range of transfer tasks.Much like how computer vision uses ImageNet to obtain features, which can then be transferred to other tasks, our work tends to indicate the suitability of natural language inference for transfer learning to other NLP tasks.Our encoder is publicly available 1 .
Alexis Conneau, Douwe Kiela, Holger Schwenk, Loïc Barrault, Antoine Bordes
EMNLP3
2017 Parallel fragments : Measuring their impact on translation performance
Sadaf Abdul-Rauf, Holger Schwenk, Mohammad Nawaz
Comput. Speech Lang.2
2017 Introduction to the special issue on deep learning approaches for machine translation
Marta R. Costa-jussà, Alexandre Allauzen, Loïc Barrault, Kyunghyun Cho, Holger Schwenk
Comput. Speech Lang.5
2016 Building and using multimodal comparable corpora for machine translation
abstract
Abstract In recent decades, statistical approaches have significantly advanced the development of machine translation systems. However, the applicability of these methods directly depends on the availability of very large quantities of parallel data. Recent works have demonstrated that a comparable corpus can compensate for the shortage of parallel corpora. In this paper, we propose an alternative to comparable corpora containing text documents as resources for extracting parallel data: a multimodal comparable corpus with audio documents in source language and text document in target language, built fromEuronewsandTEDweb sites. The audio is transcribed by an automatic speech recognition system, and translated with a baseline statistical machine translation system. We then use information retrieval in a large text corpus in the target language in order to extract parallel sentences/phrases. We evaluate the quality of the extracted data on an English to French translation task and show significant improvements over a state-of-the-art baseline.
Haithem Afli, Loïc Barrault, Holger Schwenk
Nat. Lang. Eng.3
2016 Empirical Use of Information Retrieval to Build Synthetic Data for SMT Domain Adaptation
abstract
In this paper, we present information retrieval as a powerful tool for addressing an imperative problem in the field of statistical machine translation, i.e., improving translation quality when not enough parallel corpora are available. We devise a framework, which uses information retrieval to create a synthetic corpus from the easily available monolingual corpora. We propose an improved unsupervised training approach with a data selection mechanism, which selects only the most appropriate sentences, thus reducing the amount of data, which is less related to the domain in the additional bitext. We also introduce a new method to choose sentences based on their relative similarity/difference from the query sentence. Using the synthetic corpus created by our method, we are able to improve state-of-the-art statistical machine translation systems.
Sadaf Abdul-Rauf, Holger Schwenk, Patrik Lambert, Mohammad Nawaz
IEEE ACM Trans. Audio Speech Lang. Process.2
2015 Continuous Adaptation to User Feedback for Statistical Machine Translation
abstract
Frédéric Blain, Fethi Bougares, Amir Hazem, Loïc Barrault, Holger Schwenk. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015.
Frédéric Blain, Fethi Bougares, Amir Hazem, Loïc Barrault, Holger Schwenk
HLT-NAACL5
2014 Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation
abstract
Kyunghyun Cho, Bart van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, Yoshua Bengio. Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2014.
Kyunghyun Cho, Bart van Merrienboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, Yoshua Bengio
EMNLP6
2014 Translation project adaptation for MT-enhanced computer assisted translation
Mauro Cettolo, Nicola Bertoldi, Marcello Federico, Holger Schwenk, Loïc Barrault, Christophe Servan
Mach. Transl.4
2013 A Multi-Domain Translation Model Framework for Statistical Machine Translation
Rico Sennrich, Holger Schwenk, Walid Aransa
ACL (1)2
2013 Multimodal Comparable Corpora as Resources for Extracting Parallel Data: Parallel Phrases Extraction
Haithem Afli, Loïc Barrault, Holger Schwenk
IJCNLP3
2013 CSLM - a modular open-source continuous space language modeling toolkit
abstract
Language models play a very important role in many natural language processing applications, in particular large vocabulary speech recognition and statistical machine translation. For a long time, back-off n-gram language models were considered to be the state-of-art when large amounts of training data are available. Recently, so called continuous space methods or neural network language models have shown to systematically outperform these models and they are getting increasingly popular. This article describes an open-source toolkit that implements these models in a very efficient way, including support for GPU cards. The modular architecture makes it very easy to work with different data formats and to support various alternative models. Using data selection, resampling techniques and a highly optimized code, training on more than five billions words takes less than 24 hours. The resulting models achieve reductions in the perplexity of almost 20%. This toolkit has been very successfully applied to various languages for large vocabulary speech recognition and statistical machine translation. By making available this toolkit we hope that many more researchers will be able to work on this very promising technique, and by these means, quickly advance the field.
Holger Schwenk
INTERSPEECH1
2012 Collaborative Machine Translation Service for Scientific texts
Patrik Lambert, Jean Senellart, Laurent Romary, Holger Schwenk, Florian Zipser, Patrice Lopez, Frédéric Blain
EACL4
2012 Automatic Translation of Scientific Documents in the HAL Archive
Patrik Lambert, Holger Schwenk, Frédéric Blain
LREC2
2011 Parametric Weighting of Parallel Data for Statistical Machine Translation
Kashif Shah, Loïc Barrault, Holger Schwenk
IJCNLP3
2011 Qualitative Analysis of Post-Editing for High Quality Machine Translation
Frédéric Blain, Jean Senellart, Holger Schwenk, Mirko Plitt, Johann Roturier
MTSummit3
2011 Parallel sentence generation from comparable corpora for improved SMT
Sadaf Abdul-Rauf, Holger Schwenk
Mach. Transl.2
2009 Trends and challenges in language modeling for speech recognition and machine translation
abstract
Summary form only given. Language models play an important role in large vocabulary continuous speech recognition (LVCSR) systems and statistical approaches to machine translation (SMT), in particular when modeling morphologically rich languages. Despite intensive research over more than 20 years, state-of-the-art LVCSR and SMT systems seem to use only one dominant approach: n-gram back-off language models. This talk first reviews the most important approaches to language modeling. I then discuss some of the recent trends and challenges for the future. An interesting alternative to the back-off n-gram approach are the so-called continuous space methods. The basic idea is to perform the probability estimation in a continuous space. By these means better probability estimations of unseen word sequences can be expected. There is also a relative large body of works on adaptive language models. The adaptation can aim to tailor a language model to a particular task or domain, or it can be performed over time. Another very active research area are discriminative language models. Finally, I will review the challenges and benefits of language models trained an very large amounts of training material.
Holger Schwenk
ASRU1
2009 On the Use of Comparable Corpora to Improve SMT performance
Sadaf Abdul-Rauf, Holger Schwenk
EACL2
2008 Large and Diverse Language Models for Statistical Machine Translation
Holger Schwenk, Philipp Koehn
IJCNLP1
2008 Data selection and smoothing in an open-source system for the 2008 NIST machine translation evaluation
abstract
International audience
Holger Schwenk, Yannick Estève
INTERSPEECH1
2008 System Combination for Machine Translation of Spoken and Written Language
abstract
This paper describes an approach for computing a consensus translation from the outputs of multiple machine translation (MT) systems. The consensus translation is computed by weighted majority voting on a confusion network, similarly to the well-established ROVER approach of Fiscus for combining speech recognition hypotheses. To create the confusion network, pairwise word alignments of the original MT hypotheses are learned using an enhanced statistical alignment algorithm that explicitly models word reordering. The context of a whole corpus of automatic translations rather than a single sentence is taken into account in order to achieve high alignment quality. The confusion network is rescored with a special language model, and the consensus translation is extracted as the best path. The proposed system combination approach was evaluated in the framework of the TC-STAR speech translation project. Up to six state-of-the-art statistical phrase-based translation systems from different project partners were combined in the experiments. Significant improvements in translation quality from Spanish to English and from English to Spanish in comparison with the best of the individual MT systems were achieved under official evaluation conditions.
Evgeny Matusov, Gregor Leusch, Rafael E. Banchs, Nicola Bertoldi, Daniel Déchelotte, Marcello Federico, Muntsin Kolss, Young-Suk Lee 0001, José B. Mariño, Matthias Paulik, Salim Roukos, Holger Schwenk, Hermann Ney
IEEE Trans. Speech Audio Process.12
2007 Smooth Bilingual N-Gram Translation
Holger Schwenk, Marta R. Costa-jussà, José A. R. Fonollosa
EMNLP-CoNLL1
2007 The LIMSI 2006 TC-STAR EPPS Transcription Systems
abstract
This paper describes the speech recognizers developed to transcribe European Parliament Plenary Sessions (EPPS) in English and Spanish in the 2nd TC-STAR Evaluation Campaign. The speech recognizers are state-of-the-art systems using multiple decoding passes with models (lexicon, acoustic models, language models) trained for the different transcription tasks. Compared to the LIMSI TC-STAR 2005 EPPS systems, relative word error rate reductions of about 30% have been achieved on the 2006 development data. The word error rates with the LIMSI systems on the 2006 EPPS evaluation data are 8.2% for English and 7.8% for Spanish. Experiments with cross-site adaptation and system combination are also described.
Lori Lamel, Jean-Luc Gauvain, Gilles Adda, Claude Barras, Eric Bilinski, Olivier Galibert, Agusti Pujol, Holger Schwenk, Xuan Zhu 0001
ICASSP (4)8
2007 Improved machine translation of speech-to-text outputs
abstract
International audience
Daniel Déchelotte, Holger Schwenk, Gilles Adda, Jean-Luc Gauvain
INTERSPEECH2
2007 A state-of-the-art statistical machine translation system based on Moses
Daniel Déchelotte, Holger Schwenk, Hélène Bonneau-Maynard, Alexandre Allauzen, Gilles Adda
MTSummit2
2007 Continuous space language models
Holger Schwenk
Comput. Speech Lang.1
2006 Continuous Space Language Models for Statistical Machine Translation
Holger Schwenk, Daniel Déchelotte, Jean-Luc Gauvain
ACL1
2006 Advances in transcription of broadcast news and conversational telephone speech within the combined EARS BBN/LIMSI system
abstract
This paper describes the progress made in the transcription of broadcast news (BN) and conversational telephone speech (CTS) within the combined BBN/LIMSI system from May 2002 to September 2004. During that period, BBN and LIMSI collaborated in an effort to produce significant reductions in the word error rate (WER), as directed by the aggressive goals of the Effective, Affordable, Reusable, Speech-to-text [Defense Advanced Research Projects Agency (DARPA) EARS] program. The paper focuses on general modeling techniques that led to recognition accuracy improvements, as well as engineering approaches that enabled efficient use of large amounts of training data and fast decoding architectures. Special attention is given on efforts to integrate components of the BBN and LIMSI systems, discussing the tradeoff between speed and accuracy for various system combination strategies. Results on the EARS progress test sets show that the combined BBN/LIMSI system achieved relative reductions of 47% and 51% on the BN and CTS domains, respectively.
Spyridon Matsoukas, Jean-Luc Gauvain, Gilles Adda, Thomas Colthurst, Chia-Lin Kao, Owen Kimball, Lori Lamel, Fabrice Lefèvre, Jeff Z. Ma, John Makhoul, Long Nguyen 0001, Rohit Prasad, Richard M. Schwartz, Holger Schwenk, Bing Xiang
IEEE Trans. Speech Audio Process.14
2005 Where are we in transcribing French broadcast news?
abstract
International audience
Jean-Luc Gauvain, Gilles Adda, Martine Adda-Decker, Alexandre Allauzen, Véronique Gendner, Lori Lamel, Holger Schwenk
INTERSPEECH7
2005 The 2004 BBN/LIMSI 20xRT English conversational telephone speech recognition system
abstract
In this paper we describe the English Conversational Telephone Speech (CTS) recognition system jointly developed by BBN and LIMSI under the DARPA EARS program for the 2004 evalua-tion conducted by NIST. The 2004 BBN/LIMSI system achieved a word error rate (WER) of 13.5 % at 18.3xRT (real-time as mea-sured on Pentium 4 Xeon 3.4 GHz Processor) on the EARS progress test set. This translates into a 22.8 % relative improvement in WER over the 2003 BBN/LIMSI EARS evaluation system, which was run without any time constraints. In addition to reporting on the system architecture and the evaluation results, we also highlight the significant improvements made at both sites. 1.
Rohit Prasad, Spyridon Matsoukas, Chia-Lin Kao, Jeff Z. Ma, Dongxin Xu, Thomas Colthurst, Owen Kimball, Richard M. Schwartz, Jean-Luc Gauvain, Lori Lamel, Holger Schwenk, Gilles Adda, Fabrice Lefèvre
INTERSPEECH11
2005 Building continuous space language models for transcribing european languages
abstract
International audience
Holger Schwenk, Jean-Luc Gauvain
INTERSPEECH1
2004 Speech transcription in multiple languages
abstract
The paper summarizes recent work underway at LIMSI on speech-to-text transcription in multiple languages. The research has been oriented towards the processing of broadcast audio and conversational speech for information access. Broadcast news transcription systems have been developed for seven languages, and it is planned to address several other languages in the near term. Research on conversational speech has mainly focused on the English language, with some initial work on French, Arabic and Spanish. Automatic processing must take into account the characteristics of the audio data, such as needing to deal with the continuous data stream, specificities of the language and the use of an imperfect word transcription for accessing the information content. Our experience thus far indicates that at today's word error rates, the techniques used in one language can be successfully ported to other languages, and most of the language specificities concern lexical and pronunciation modeling.
Lori Lamel, Jean-Luc Gauvain, Gilles Adda, Martine Adda-Decker, Leonardo Canseco-Rodriguez, Langzhou Chen, Olivier Galibert, Abdelkhalek Messaoudi, Holger Schwenk
ICASSP (3)9
2004 Speech recognition in multiple languages and domains: the 2003 BBN/LIMSI EARS system
abstract
We report on the results of the first evaluations for the BBN/LIMSI system under the new DARPA EARS program. The evaluations were carried out for conversational telephone speech (CTS) and broadcast news (BN) for three languages: English, Mandarin, and Arabic. In addition to providing system descriptions and evaluation results, the paper highlights methods that worked well across the two domains and those few that worked well on one domain but not the other. For the BN evaluations, which had to be run under 10 times real-time, we demonstrated that a joint BBN/LIMSI system with a time constraint achieved better results than either system alone.
Richard M. Schwartz, Thomas Colthurst, Nicolae Duta, Herbert Gish, Rukmini Iyer, Chia-Lin Kao, Daben Liu, Owen Kimball, Jeff Z. Ma, John Makhoul, Spyridon Matsoukas, Long Nguyen 0001, Mohammed Noamany, Rohit Prasad, Bing Xiang, Dongxin Xu, Jean-Luc Gauvain, Lori Lamel, Holger Schwenk, Gilles Adda, Langzhou Chen
ICASSP (3)19
2004 Efficient training of large neural networks for language modeling
abstract
Recently there has been increasing interest in using neural networks for language modeling. In contrast to the well-known backoff n-gram language models, the neural network approach tries to limit the data sparseness problem by performing the estimation in a continuous space, allowing by this means smooth interpolations. The complexity to train such a model and to calculate one n-gram probability is however several orders of magnitude higher than for the backoff models, making the new approach difficult to use in real applications. In this paper several techniques are presented that allow the use of a neural network language model in a large vocabulary speech recognition system, in particular very, fast lattice rescoring and efficient training of large neural networks on training corpora of over 10 million words. The described approach achieves significant word error reductions with respect to a carefully tuned 4-gram backoff language model in a state of the art conversational speech recognizer for the DARPA rich transcriptions evaluations.
Holger Schwenk
IJCNN1
2004 Language recognition using phone latices
abstract
This paper proposes a new phone lattice based method for automatic language recognition from speech data. By using phone lattices some approximations usually made by language identification (LID) systems relying on phonotactic constraints to simplify the training and decoding processes can be avoided. We demonstrate the use of phone lattices both in training and testing significantly improves the accuracy of a phonotactically based LID system. Performance is further enhanced by using a neural network to combine the results of multiple phone recognizers. Using three phone recognizers with context independent phone models, the system achieves an equal error rate of 2.7% on the Eval03 NIST detection test (30s segment, primary condition) with an overall decoding process that runs faster than real-time (0.5xRT).
Jean-Luc Gauvain, Abdelkhalek Messaoudi, Holger Schwenk
INTERSPEECH3
2004 Neural network language models for conversational speech recognition
abstract
International audience
Holger Schwenk, Jean-Luc Gauvain
INTERSPEECH1
2003 Conversational telephone speech recognition
abstract
This paper describes the development of a speech recognition system for the processing of telephone conversations, starting with a state-of-the-art broadcast news transcription system. We identify major changes and improvements in acoustic and language modeling, as well as decoding, which are required to achieve state-of-the-art performance on conversational speech. Some major changes on the acoustic side include the use of speaker normalization (VTLN), the need to cope with channel variability, and the need for efficient speaker adaptation and better pronunciation modeling. On the linguistic side the primary challenge is to cope with the limited amount of language model training data. To address this issue we make use of a data selection technique, and a smoothing technique based on a neural network language model. At the decoding level lattice rescoring and minimum word error decoding are applied. On the development data, the improvements yield an overall word error rate of 24.9% whereas the original BN transcription system had a word error rate of about 50% on the same data.
Jean-Luc Gauvain, Lori Lamel, Holger Schwenk, Gilles Adda, Langzhou Chen, Fabrice Lefèvre
ICASSP (1)3
2002 Connectionist language modeling for large vocabulary continuous speech recognition
abstract
This paper describes ongoing work on a new approach for language modeling for large vocabulary continuous speech recognition. Almost all state.. o. f-the-art systems use statistical n-gram language models estimated on text corpora. One principle problem with such language models is the fact that many of the n-grams are never observed even in very large training corpora, and therefore it is common to back-off to a lower-order model. In this paper we propose to address this problem by carrying out the estimation task in a continuous space, enabling a smooth interpolation of the probabilities. A neural network is used to learn the projection of the words onto a continuous space and to estimate the n-gram probabilities. The connectionist language model is being evaluated on the DARPA HUB5 conversational telephone speech recognition task and preliminary results show consistent improvements in both perplexity and word error rate.
Holger Schwenk, Jean-Luc Gauvain
ICASSP1
2000 Combining multiple speech recognizers using voting and language model information
Holger Schwenk, Jean-Luc Gauvain
INTERSPEECH1
2000 Boosting Neural Networks
abstract
Boosting is a general method for improving the performance of learning algorithms. A recently proposed boosting algorithm, AdaBoost, has been applied with great success to several benchmark machine learning problems using mainly decision trees as base classifiers. In this article we investigate whether AdaBoost also works as well with neural networks, and we discuss the advantages and drawbacks of different versions of the AdaBoost algorithm. In particular, we compare training methods based on sampling the training set and weighting the cost function. The results suggest that random resampling of the training data is not the main explanation of the success of the improvements brought by AdaBoost. This is in contrast to bagging, which directly aims at reducing variance and for which random resampling is essential to obtain the reduction in generalization error. Our system achieves about 1.4% error on a data set of on-line handwritten digits from more than 200 writers. A boosted multilayer network achieved 1.5% error on the UCI letters and 8.1% error on the UCI satellite data set, which is significantly better than boosted decision trees.
Holger Schwenk, Yoshua Bengio
Neural Comput.1
1999 Using boosting to improve a hybrid HMM/neural network speech recognizer
abstract
"Boosting" is a general method for improving the performance of almost any learning algorithm. A previously proposed and very promising boosting algorithm is AdaBoost. In this paper we investigate if AdaBoost can be used to improve a hybrid HMM/neural network continuous speech recognizer. Boosting significantly improves the word error rate from 6.3% to 5.3% on a test set of the OGI Numbers 95 corpus, a medium size continuous numbers recognition task. These results compare favorably with other combining techniques using several different feature representations or additional information from longer time spans. In summary, we can say that the reasons for the impressive success of AdaBoost are still not completely understood. To the best of our knowledge, an application of AdaBoost to a real world problem has not yet been reported in the literature either. In this paper we investigate if AdaBoost can be applied to boost the performance of a continuous speech recognition system. In this domain we have to deal with large amounts of data (often more than 1 million training examples) and inherently noisy phoneme labels. The paper is organized as follows. We summarize the AdaBoost algorithm and our baseline speech recognizer. We show how AdaBoost can be applied to this task and we report results on the Numbers 95 corpus and compare them with other classifier combination techniques. The paper finishes with a conclusion and perspectives for future work.
Holger Schwenk
ICASSP1
1998 The Diabolo Classifier
abstract
We present a new classification architecture based on autoassociative neural networks that are used to learn discriminant models of each class. The proposed architecture has several interesting properties with respect to other model-based classifiers like nearest-neighbors or radial basis functions: it has a low computational complexity and uses a compact distributed representation of the models. The classifier is also well suited for the incorporation of a priori knowledge by means of a problem-specific distance measure. In particular, we will show that tangent distance (Simard, Le Cun, & Denker, 1993) can be used to achieve transformation invariance during learning and recognition. We demonstrate the application of this classifier to optical character recognition, where it has achieved state-of-the-art results on several reference databases. Relations to other models, in particular those based on principal component analysis, are also discussed.
Holger Schwenk
Neural Comput.1
1997 AdaBoosting Neural Networks: Application to on-line Character Recognition
Holger Schwenk, Yoshua Bengio
ICANN1
1997 Training Methods for Adaptive Boosting of Neural Networks
Holger Schwenk, Yoshua Bengio
NIPS1
1996 Constraint tangent distance for on-line character recognition
abstract
In online character recognition we can observe two kinds of intra-class variations: small geometric deformations and completely different writing styles. We propose a new approach to deal with these problems by defining an extension of tangent distance, well known in off-line character recognition. The system has been implemented with a k-nearest neighbor classifier and a so called diabolo classifier respectively. Both classifiers are invariant under transformations like rotation, scale or slope and can deal with variations in stroke order and writing direction. Results are presented for our digit database with more than 200 writers.
Holger Schwenk, Maurice Milgram
ICPR1
1994 Transformation Invariant Autoassociation with Application to Handwritten Character Recognition
abstract
When training neural networks by the classical backpropagation algo(cid:173) rithm the whole problem to learn must be expressed by a set of inputs and desired outputs. However, we often have high-level knowledge about the learning problem. In optical character recognition (OCR), for in(cid:173) stance, we know that the classification should be invariant under a set of transformations like rotation or translation. We propose a new modular classification system based on several autoassociative multilayer percep(cid:173) trons which allows the efficient incorporation of such knowledge. Results are reported on the NIST database of upper case handwritten letters and compared to other approaches to the invariance problem. 1 INCORPORATION OF EXPLICIT KNOWLEDGE The aim of supervised learning is to learn a mapping between the input and the output space from a set of example pairs (input, desired output). The classical implementation in the domain of neural networks is the backpropagation algorithm. If this learning set is sufficiently representative of the underlying data distributions, one hopes that after learning, the system is able to generalize correctly to other inputs of the same distribution. 992 Holger Schwenk, Maurice Milgram It would be better to have more powerful techniques to incorporate knowledge into the learning process than the choice of a set of examples. The use of additional knowledge is often limited to the feature extraction module. Besides simple operations like (size) normalization, we can find more sophisticated approaches like zernike moments in the domain of optical character recognition (OCR). In this paper we will not investigate this possibility, all discussed classifiers work directly on almost non preprocessed data (pixels). In the context of OCR interest focuses on invariance of the classifier under a number of given transformations (translation, rotation, ... ) of the data to classify. In general a neural network could extract those properties of a large enough learning set, but it is very hard to learn and will probably take a lot of time. In the last years two main approaches for this invariance problem have been proposed: tangent-prop and tangent-distance. An indirect incorporation can be achieved by boosting (Drucker, Schapire and Simard, 1993). In this paper we briefly discuss these approaches and will present a new classification system which allows the efficient incorporation of transformation invariances. 1.1 TANGENT PROPAGATION The principle of tangent-prop is to specify besides desired outputs also desired changes jJJ. of the output vector when transforming the net input x by the transformations tJJ. (Simard, Victorri, LeCun and Denker, 1992). For this, let us define a transformation of pattern p as t(p, a) : P --t P where P is the space of all patterns and a a parameter. Such transformations are in general highly nonlinear operations in the pixel space P and their analytical expressions are seldom known. It is therefore favorable to use a first order approximation:
Holger Schwenk, Maurice Milgram
NIPS1