EDBT 2026 Demo / reviewers in the wild / expert
Alexandre Allauzen
dblp:54/8163 · also Alex Allauzen, Alexander Allauzen
· DBLP profile ↗
49ranked-venue papers
5as first author
20since 2021 · last 2026
0000-0002-8627-1965ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 42 · 4 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 3 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | D-LIM: A neural network for interpretable gene-gene interactionsabstractRecent advances in gene editing can produce large genotype-fitness maps for targeted genes, yet predicting the effects of mutations between genes remains challenging. Indeed, biochemical models require knowledge of underlying parameters and interactions, whereas machine learning methods typically lack interpretability, as they do not link model parameters to biological quantities. We introduce D-LIM, a neural network that infers low-dimensional fitness landscapes directly from mutation-fitness data. The distinctive feature of D-LIM is that it assumes genes act through independent gene-specific molecular phenotypes whose nonlinear interactions determine fitness. When this assumption holds, the model yields accurate predictions and interpretable effective phenotypes. Conversely, failure reveals that a low-dimensional model is insufficient. Applied to deep mutational scanning of metabolic pathways, protein-protein interactions, and yeast environmental adaptation, D-LIM achieves state-of-the-art predictive accuracy. The inferred phenotype-fitness landscapes reveal whether epistatic interactions can be captured by a low-dimensional continuous model and identify potential trade-offs. Moreover, D-LIM estimates mutational effects on the effective phenotypes, enabling weak extrapolation beyond the training domain. D-LIM demonstrates how simple structure constraints in a neural network can help inference and hypothesis generation in biology. Shuhui Wang, Alexandre Allauzen, Philippe Nghe, Vaitea Opuu |
PLoS Comput. Biol. | 2 |
| 2025 | Bridging the Theoretical Gap in Randomized SmoothingabstractRandomized smoothing has become a leading approach for certifying adversarial robustness in machine learning models. However, a persistent gap remains between theoretical certified robustness and empirical robustness accuracy. This paper introduces a new framework that bridges this gap by leveraging Lipschitz continuity for certification and proposing a novel, less conservative method for computing confidence intervals in randomized smoothing. Our approach tightens the bounds of certified robustness, offering a more accurate reflection of model robustness in practice. Through rigorous experimentation we show that our method improves the robust accuracy, compressing the gap between empirical findings and previous theoretical results. We argue that investigating local Lipschitz constants and designing ad-hoc confidence intervals can further enhance the performance of randomized smoothing. These results pave the way for a deeper understanding of the relationship between Lipschitz continuity and certified robustness. Blaise Delattre, Paul Caillon, Quentin Barthélemy, Erwan Fagnou, Alexandre Allauzen |
AISTATS | 5 |
| 2025 | Forward Only Learning for Orthogonal Neural Networks of Any DepthabstractBackpropagation is still the de facto algorithm used today to train neural networks. With the exponential growth of recent architectures, the computational cost of this algorithm also becomes a burden. The recent PEPITA and forward-only frameworks have proposed promising alternatives, but they failed to scale up to a handful of hidden layers, yet limiting their use. In this paper, we first analyze theoretically the main limitations of these approaches. It allows us the design of a forward-only algorithm, which is equivalent to backpropagation under the linear and orthogonal assumptions. By relaxing the linear assumption, we then introduce FOTON (Forward-Only Training of Orthogonal Networks) that bridges the gap with the backpropagation algorithm. Experimental results show that it outperforms PEPITA, enabling us to train neural networks of any depth, without the need for a backward pass. Moreover its performance on convolutional networks clearly opens up avenues for its application to more advanced architectures. The code is open-sourced on https://github.com/p0lcAi/FOTON. Paul Caillon, Alex Colagrande, Erwan Fagnou, Blaise Delattre, Alexandre Allauzen |
ECAI | 5 |
| 2025 | Training compute-optimal transformer encoder modelsabstractTransformer encoders are critical for a wide range of Natural Language Processing (NLP) tasks, yet their compute-efficiency remains poorly understood.We present the first comprehensive empirical investigation of compute-optimal pretraining for encoder transformers using the Masked Language Modeling (MLM) objective.Across hundreds of carefully controlled runs we vary model size, data size, batch size, learning rate, and masking ratio, with increasing compute budget.The compute-optimal data-to-model ratio of Transformer encoder models is 10 to 100 times larger than the ratio of auto-regressive models.Using these recipes, we train OptiBERT, a family of compute-optimal BERT-style models that matches or surpasses leading baselines-including ModernBERT and NeoBERT-on GLUE and MTEB while training with dramatically less FLOPS. Megi Dervishi, Alexandre Allauzen, Gabriel Synnaeve, Yann LeCun |
EMNLP | 2 |
| 2025 | SCOPE: A Self-supervised Framework for Improving Faithfulness in Conditional Text GenerationabstractLarge Language Models (LLMs), when used for conditional text generation, often produce hallucinations, i.e., information that is unfaithful or not grounded in the input context. This issue arises in typical conditional text generation tasks, such as text summarization and data-to-text generation, where the goal is to produce fluent text based on contextual input. When fine-tuned on specific domains, LLMs struggle to provide faithful answers to a given context, often adding information or generating errors. One underlying cause of this issue is that LLMs rely on statistical patterns learned from their training data. This reliance can interfere with the model's ability to stay faithful to a provided context, leading to the generation of ungrounded information. We build upon this observation and introduce a novel self-supervised method for generating a training set of unfaithful samples. We then refine the model using a training process that encourages the generation of grounded outputs over unfaithful ones, drawing on preference-based training. Our approach leads to significantly more grounded text generation, outperforming existing self-supervised techniques in faithfulness, as evaluated through automatic metrics, LLM-based assessments, and human evaluations. Song Duong, Florian Le Bronnec, Alexandre Allauzen, Vincent Guigue, Alberto Lumbreras, Laure Soulier, Patrick Gallinari |
ICLR | 3 |
| 2025 | Accelerated training through iterative gradient propagation along the residual pathabstractDespite being the cornerstone of deep learning, backpropagation is criticized for its inherent sequentiality, which can limit the scalability of very deep models.
Such models faced convergence issues due to vanishing gradient, later resolved using residual connections. Variants of these are now widely used in modern architectures.
However, the computational cost of backpropagation remains a major burden, accounting for most of the training time.
Taking advantage of residual-like architectural designs, we introduce Highway backpropagation, a parallelizable iterative algorithm that approximates backpropagation, by alternatively i) accumulating the gradient estimates along the residual path, and ii) backpropagating them through every layer in parallel. This algorithm is naturally derived from a decomposition of the gradient as the sum of gradients flowing through all paths, and is adaptable to a diverse set of common architectures, ranging from ResNets and Transformers to recurrent neural networks.
Through an extensive empirical study on a large selection of tasks and models, we evaluate Highway-BP and show that major speedups can be achieved with minimal performance degradation. Erwan Fagnou, Paul Caillon, Blaise Delattre, Alexandre Allauzen |
ICLR | 4 |
| 2025 | Improving Diversity in Language Models: When Temperature Fails, Change the LossabstractIncreasing diversity in language models is a challenging yet essential objective. A common approach is to raise the decoding temperature. In this work, we investigate this approach through a simplistic yet common case to provide insights into why decreasing temperature can improve quality (Precision), while increasing it often fails to boost coverage (Recall). Our analysis reveals that for a model to be effectively tunable through temperature adjustments, it must be trained toward coverage. To address this, we propose rethinking loss functions in language models by leveraging the Precision-Recall framework. Our results demonstrate that this approach achieves a substantially better trade-off between Precision and Recall than merely combining negative log-likelihood training with temperature scaling. These findings offer a pathway toward more versatile and robust language modeling techniques. Alexandre Verine, Florian Le Bronnec, Kunhao Zheng, Alexandre Allauzen, Yann Chevaleyre, Benjamin Négrevergne |
ICML | 4 |
| 2024 | Exploring Precision and Recall to assess the quality and diversity of LLMsabstractFlorian Le Bronnec, Alexandre Verine, Benjamin Negrevergne, Yann Chevaleyre, Alexandre Allauzen. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Florian Le Bronnec, Alexandre Verine, Benjamin Négrevergne, Yann Chevaleyre, Alexandre Allauzen |
ACL (1) | 5 |
| 2024 | LOCOST: State-Space Models for Long Document Abstractive SummarizationabstractFlorian Le Bronnec, Song Duong, Mathieu Ravaut, Alexandre Allauzen, Nancy Chen, Vincent Guigue, Alberto Lumbreras, Laure Soulier, Patrick Gallinari. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Florian Le Bronnec, Song Duong, Mathieu Ravaut, Alexandre Allauzen, Nancy F. Chen, Vincent Guigue, Alberto Lumbreras, Laure Soulier, Patrick Gallinari |
EACL (1) | 4 |
| 2024 | Chain and Causal Attention for Efficient Entity TrackingabstractThis paper investigates the limitations of transformers for entity-tracking tasks in large language models.We identify a theoretical constraint, showing that transformers require at least log 2 (n + 1) layers to handle entity tracking with n state changes.To address this issue, we propose an efficient and frugal enhancement to the standard attention mechanism, enabling it to manage long-term dependencies more efficiently.By considering attention as an adjacency matrix, our model can track entity states with a single layer.Empirical results demonstrate significant improvements in entity tracking datasets while keeping competitive performance on standard natural language modeling.Our modified attention allows us to achieve the same performance with drastically fewer layers.Additionally, our enhanced mechanism reveals structured internal representations of attention.Extensive experiments on both toy and complex datasets validate our approach.Our contributions include theoretical insights, an improved attention mechanism, and empirical validation. Erwan Fagnou, Paul Caillon, Blaise Delattre, Alexandre Allauzen |
EMNLP | 4 |
| 2024 | The Lipschitz-Variance-Margin Tradeoff for Enhanced Randomized SmoothingabstractReal-life applications of deep neural networks are hindered by their unsteady predictions when faced with noisy inputs and adversarial attacks. The certified radius in this context is a crucial indicator of the robustness of models. However how to design an efficient classifier with an associated certified radius? Randomized smoothing provides a promising framework by relying on noise injection into the inputs to obtain a smoothed and robust classifier. In this paper, we first show that the variance introduced by the Monte-Carlo sampling in the randomized smoothing procedure estimate closely interacts with two other important properties of the classifier, \textit{i.e.} its Lipschitz constant and margin. More precisely, our work emphasizes the dual impact of the Lipschitz constant of the base classifier, on both the smoothed classifier and the empirical variance. To increase the certified robust radius, we introduce a different way to convert logits to probability vectors for the base classifier to leverage the variance-margin trade-off. We leverage the use of Bernstein's concentration inequality along with enhanced Lipschitz bounds for randomized smoothing. Experimental results show a significant improvement in certified accuracy compared to current state-of-the-art methods. Our novel certification procedure allows us to use pre-trained models with randomized smoothing, effectively improving the current certification radius in a zero-shot manner. Blaise Delattre, Alexandre Araujo, Quentin Barthélemy, Alexandre Allauzen |
ICLR | 4 |
| 2024 | LeBenchmark 2.0: A standardized, replicable and enhanced framework for self-supervised representations of French speech
Titouan Parcollet, Solène Evain, Marcely Zanon Boito, Adrien Pupier, Salima Mdhaffar, Hang Le 0001, Sina Alisamir, Natalia A. Tomashenko, Marco Dinarelli, Shucong Zhang, Alexandre Allauzen, Maximin Coavoux, Yannick Estève, Mickael Rouvier, Jérôme Goulian, Benjamin Lecouteux, François Portet, Solange Rossato, Fabien Ringeval, Didier Schwab, Laurent Besacier |
Comput. Speech Lang. | 12 |
| 2023 | What does KnowBert-UMLS forget?abstractIntegrating a source of structured prior knowledge, such as a knowledge graph, into transformer-based language models is an increasingly popular method for increasing data efficiency and adapting them to a target domain. However, most methods for integrating structured knowledge into language models require additional training in order to adapt the model to the non-textual modality. This process typically leads to some amount of catastrophic forgetting on the general domain. KnowBert is one such knowledge integration method which can incorporate information from a variety of knowledge graphs to enhance the capabilities of transformer-based language models such as BERT. We conduct a qualitative analysis of the results of KnowBert-UMLS, a biomedically specialized KnowBert model, on a variety of linguistic tasks. Our results reveal that its increased understanding of biomedical concepts comes at the cost, specifically, of general common-sense knowledge and understanding of casual speech. Guilhem Piat, Nasredine Semmar, Julien Tourille, Alexandre Allauzen, Hassane Essafi |
AICCSA | 4 |
| 2023 | A Unified Algebraic Perspective on Lipschitz Neural Networks
Alexandre Araujo, Aaron J. Havens, Blaise Delattre, Alexandre Allauzen, Bin Hu 0002 |
ICLR | 4 |
| 2023 | Efficient Bound of Lipschitz Constant for Convolutional Layers by Gram IterationabstractSince the control of the Lipschitz constant has a great impact on the training stability, generalization, and robustness of neural networks, the estimation of this value is nowadays a real scientific challenge. In this paper we introduce a precise, fast, and differentiable upper bound for the spectral norm of convolutional layers using circulant matrix theory and a new alternative to the Power iteration. Called the Gram iteration, our approach exhibits a superlinear convergence. First, we show through a comprehensive set of experiments that our approach outperforms other state-of-the-art methods in terms of precision, computational cost, and scalability. Then, it proves highly effective for the Lipschitz regularization of convolutional neural networks, with competitive results against concurrent approaches. Blaise Delattre, Quentin Barthélemy, Alexandre Araujo, Alexandre Allauzen |
ICML | 4 |
| 2023 | Relational data embeddings for feature enrichment with background information
Alexis Cvetkov-Iliev, Alexandre Allauzen, Gaël Varoquaux |
Mach. Learn. | 2 |
| 2022 | A Dynamical System Perspective for Lipschitz Neural NetworksabstractThe Lipschitz constant of neural networks has been established as a key quantity to enforce the robustness to adversarial examples. In this paper, we tackle the problem of building $1$-Lipschitz Neural Networks. By studying Residual Networks from a continuous time dynamical system perspective, we provide a generic method to build $1$-Lipschitz Neural Networks and show that some previous approaches are special cases of this framework. Then, we extend this reasoning and show that ResNet flows derived from convex potentials define $1$-Lipschitz transformations, that lead us to define the Convex Potential Layer (CPL). A comprehensive set of experiments on several datasets demonstrates the scalability of our architecture and the benefits as an $\ell_2$-provable defense against adversarial examples. Our code is available at \url{https://github.com/MILES-PSL/Convex-Potential-Layer} Laurent Meunier, Blaise Delattre, Alexandre Araujo, Alexandre Allauzen |
ICML | 4 |
| 2022 | Adapting without forgetting: KnowBert-UMLSabstractDomain adaptation in pretrained language models usually comes at some cost, most notably out-of-domain performance. This type of specialization typically relies on pre-training over a large in-domain corpus, which has the side effect of causing catastrophic forgetting on general text. We seek to specialize a language model by incorporating information from a knowledge base into its contextualized representations, thus reducing its reliance on specialized text. We achieve this by following the KnowBert method, applied to the UMLS biomedical knowledge base. We evaluate our model on in-domain and out-of-domain tasks, comparing against BERT and other specialized models. We find that our performance on biomedical tasks is competitive with the state-of-the-art with virtually no loss of generality. Our results demonstrate the applicability of this knowledge integration technique to the biomedical domain as well as its shortcomings. The reduced risk of catastrophic forgetting displayed by this approach to domain adaptation broadens the scope of applicability of specialized language models. Guilhem Piat, Nasredine Semmar, Alexandre Allauzen, Hassane Essafi, Gaël Bernard, Julien Tourille |
WiMob | 3 |
| 2021 | Measure and Evaluation of Semantic Divergence across Two LanguagesabstractSyrielle Montariol, Alexandre Allauzen. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Syrielle Montariol, Alexandre Allauzen |
ACL/IJCNLP (1) | 2 |
| 2021 | LeBenchmark: A Reproducible Framework for Assessing Self-Supervised Representation Learning from SpeechabstractSelf-Supervised Learning (SSL) using huge unlabeled data has been successfully explored for image and natural language processing. Recent works also investigated SSL from speech. They were notably successful to improve performance on downstream tasks such as automatic speech recognition (ASR). While these works suggest it is possible to reduce dependence on labeled data for building efficient speech systems, their evaluation was mostly made on ASR and using multiple and heterogeneous experimental settings (most of them for English). This questions the objective comparison of SSL approaches and the evaluation of their impact on building speech systems. In this paper, we propose LeBenchmark: a reproducible framework for assessing SSL from speech. It not only includes ASR (high and low resource) tasks but also spoken language understanding, speech translation and emotion recognition. We also focus on speech technologies in a language different than English: French. SSL models of different sizes are trained from carefully sourced and documented datasets. Experiments show that SSL is beneficial for most but not all tasks which confirms the need for exhaustive and reliable benchmarks to evaluate its real impact. LeBenchmark is shared with the scientific community for reproducible research in SSL from speech. Solène Evain, Hang Le 0001, Marcely Zanon Boito, Salima Mdhaffar, Sina Alisamir, Ziyi Tong, Natalia A. Tomashenko, Marco Dinarelli, Titouan Parcollet, Alexandre Allauzen, Yannick Estève, Benjamin Lecouteux, François Portet, Solange Rossato, Fabien Ringeval, Didier Schwab, Laurent Besacier |
Interspeech | 11 |
| 2020 | FlauBERT: Unsupervised Language Model Pre-training for FrenchabstractLanguage models have become a key step to achieve state-of-the art results in many different Natural Language Processing (NLP) tasks. Leveraging the huge amount of unlabeled texts nowadays available, they provide an efficient way to pre-train continuous word representations that can be fine-tuned for a downstream task, along with their contextualization at the sentence level. This has been widely demonstrated for English using contextualized representations (Dai and Le, 2015; Peters et al., 2018; Howard and Ruder, 2018; Radford et al., 2018; Devlin et al., 2019; Yang et al., 2019b). In this paper, we introduce and share FlauBERT, a model learned on a very large and heterogeneous French corpus. Models of different sizes are trained using the new CNRS (French National Centre for Scientific Research) Jean Zay supercomputer. We apply our French language models to diverse NLP tasks (text classification, paraphrasing, natural language inference, parsing, word sense disambiguation) and show that most of the time they outperform other pre-training approaches. Different versions of FlauBERT as well as a unified evaluation protocol for the downstream tasks, called FLUE (French Language Understanding Evaluation), are shared to the research community for further reproducible experiments in French NLP. Hang Le 0001, Loïc Vial, Jibril Frej, Vincent Segonne, Maximin Coavoux, Benjamin Lecouteux, Alexandre Allauzen, Benoît Crabbé, Laurent Besacier, Didier Schwab |
LREC | 7 |
| 2018 | Unsupervised Learning of Word Segmentation: Does Tone Matter?
Pierre Godard, Kevin Löser, Alexandre Allauzen, Laurent Besacier, François Yvon |
CICLing (1) | 3 |
| 2018 | Learning with Noise-Contrastive Estimation: Easing training by learning to scaleabstractNoise-Contrastive Estimation (NCE) is a learning criterion that is regularly used to train neural language models in place of Maximum Likelihood Estimation, since it avoids the computational bottleneck caused by the output softmax. In this paper, we analyse and explain some of the weaknesses of this objective function, linked to the mechanism of self-normalization, by closely monitoring comparative experiments. We then explore several remedies and modifications to propose tractable and efficient NCE training strategies. In particular, we propose to make the scaling factor a trainable parameter of the model, and to use the noise distribution to initialize the output bias. These solutions, yet simple, yield stable and competitive performances in either small and large scale language modelling tasks. Matthieu Labeau, Alexandre Allauzen |
COLING | 2 |
| 2017 | Introduction to the special issue on deep learning approaches for machine translation
Marta R. Costa-jussà, Alexandre Allauzen, Loïc Barrault, Kyunghyun Cho, Holger Schwenk |
Comput. Speech Lang. | 2 |
| 2017 | Document Neural Autoregressive Distribution EstimationabstractWe present an approach based on feed-forward neural networks for learning the distribution over textual documents. This approach is inspired by the Neural Autoregressive Distribution Estimator (NADE) model which has been shown to be a good estimator of the distribution over discrete-valued high-dimensional vectors. In this paper, we present how NADE can successfully be adapted to textual data, retaining the property that sampling or computing the probability of an observation can be done exactly and efficiently. The approach can also be used to learn deep representations of documents that are competitive to those learned by alternative topic modeling approaches. Finally, we describe how the approach can be combined with a regular neural network N-gram model and substantially improve its performance, by making its learned representation sensitive to the larger, document-level context. Stanislas Lauly, Alexandre Allauzen, Hugo Larochelle |
J. Mach. Learn. Res. | 3 |
| 2017 | A comparison of discriminative training criteria for continuous space translation models
Alexandre Allauzen, Quoc-Khanh Do, François Yvon |
Mach. Transl. | 1 |
| 2016 | Preliminary Experiments on Unsupervised Word Discovery in MboshiabstractInternational audience Pierre Godard, Gilles Adda, Martine Adda-Decker, Alexandre Allauzen, Laurent Besacier, Hélène Bonneau-Maynard, Guy-Noël Kouarata, Kevin Löser, Annie Rialland, François Yvon |
INTERSPEECH | 4 |
| 2015 | A Discriminative Training Procedure for Continuous Translation ModelsabstractContinuous-space translation models have recently emerged as extremely powerful ways to boost the performance of existing translation systems.A simple, yet effective way to integrate such models in inference is to use them in an N -best rescoring step.In this paper, we focus on this scenario and show that the performance gains in rescoring can be greatly increased when the neural network is trained jointly with all the other model parameters, using an appropriate objective function.Our approach is validated on two domains, where it outperforms strong baselines. Quoc-Khanh Do, Alexandre Allauzen, François Yvon |
EMNLP | 2 |
| 2015 | Non-lexical neural architecture for fine-grained POS TaggingabstractIn this paper we explore a POS tagging application of neural architectures that can infer word representations from the raw character stream.It relies on two modelling stages that are jointly learnt: a convolutional network that infers a word representation directly from the character stream, followed by a prediction stage.Models are evaluated on a POS and morphological tagging task for German.Experimental results show that the convolutional network can infer meaningful word representations, while for the prediction stage, a well designed and structured strategy allows the model to outperform stateof-the-art results, without any feature engineering. Matthieu Labeau, Kevin Löser, Alexandre Allauzen |
EMNLP | 3 |
| 2014 | Rule-based Reordering Space in Statistical Machine Translation
Nicolas Pécheux, Alexandre Allauzen, François Yvon |
LREC | 2 |
| 2014 | "Sheldon speaking, Bonjour!": Leveraging Multilingual Tracks for (Weakly) Supervised Speaker IdentificationabstractWe address the problem of speaker identification in multimedia data, and TV series in particular. While speaker identification is traditionally a supervised machine-learning task, our first contribution is to significantly reduce the need for costly preliminary manual annotations through the use of automatically aligned (and potentially noisy) fan-generated transcripts and subtitles. Hervé Bredin, Anindya Roy, Nicolas Pécheux, Alexandre Allauzen |
ACM Multimedia | 4 |
| 2014 | Maximum-entropy word alignment and posterior-based phrase extraction for machine translation
Nadi Tomeh, Alexandre Allauzen, François Yvon |
Mach. Transl. | 2 |
| 2013 | Structure learning in hidden conditional random fields for grapheme-to-phoneme conversionabstractAccurate grapheme-to-phoneme (g2p) conversion is needed for several speech processing applications, such as automatic speech synthesis and recognition.For some languages, notably English, improvements of g2p systems are very slow, due to the intricacy of the associations between letter and sounds.In recent years, several improvements have been obtained either by using variable-length associations in generative models (jointn-grams), or by recasting the problem as a conventional sequence labeling task, enabling to integrate rich dependencies in discriminative models.In this paper, we consider several ways to reconciliate these two approaches.Introducing hidden variable-length alignments through latent variables, our Hidden Conditional Random Field (HCRF) models are able to produce comparative performance compared to strong generative and discriminative models on the CELEX database. Patrick Lehnen, Alexandre Allauzen, Thomas Lavergne, François Yvon, Stefan Hahn, Hermann Ney |
INTERSPEECH | 2 |
| 2013 | Structured Output Layer Neural Network Language Models for Speech RecognitionabstractThis paper extends a novel neural network language model (NNLM) which relies on word clustering to structure the output vocabulary: Structured OUtput Layer (SOUL) NNLM. This model is able to handle arbitrarily-sized vocabularies, hence dispensing with the need for shortlists that are commonly used in NNLMs. Several softmax layers replace the standard output layer in this model. The output structure depends on the word clustering which is based on the continuous word representation determined by the NNLM. Mandarin and Arabic data are used to evaluate the SOUL NNLM accuracy via speech-to-text experiments. Well tuned speech-to-text systems (with error rates around 10%) serve as the baselines. The SOUL model achieves consistent improvements over a classical shortlist NNLM both in terms of perplexity and recognition accuracy for these two languages that are quite different in terms of their internal structure and recognition vocabulary size. An enhanced training scheme is proposed that allows more data to be used at each training iteration of the neural network. Hai Son Le, Ilya Oparin, Alexandre Allauzen, Jean-Luc Gauvain, François Yvon |
IEEE Trans. Speech Audio Process. | 3 |
| 2012 | Continuous Space Translation Models with Neural Networks
Hai Son Le, Alexandre Allauzen, François Yvon |
HLT-NAACL | 2 |
| 2011 | Discriminative Weighted Alignment Matrices For Statistical Machine Translation
Nadi Tomeh, Alexandre Allauzen, François Yvon |
EAMT | 2 |
| 2011 | Structured Output Layer neural network language modelabstractThis paper introduces a new neural network language model (NNLM) based on word clustering to structure the output vocabulary: Structured Output Layer NNLM. This model is able to handle vocabularies of arbitrary size, hence dispensing with the design of short-lists that are commonly used in NNLMs. Several softmax layers replace the standard output layer in this model. The output structure depends on the word clustering which uses the continuous word representation induced by a NNLM. The GALE Mandarin data was used to carry out the speech-to-text experiments and evaluate the NNLMs. On this data the well tuned baseline system has a character error rate under 10%. Our model achieves consistent improvements over the combination of an n-gram model and classical short-list NNLMs both in terms of perplexity and recognition accuracy. Hai Son Le, Ilya Oparin, Alexandre Allauzen, Jean-Luc Gauvain, François Yvon |
ICASSP | 3 |
| 2011 | Large Vocabulary SOUL Neural Network Language ModelsabstractInternational audience Hai Son Le, Ilya Oparin, Abdelkhalek Messaoudi, Alexandre Allauzen, Jean-Luc Gauvain, François Yvon |
INTERSPEECH | 4 |
| 2011 | Using Dynamic Time Warping to Compute Prosodic Similarity MeasuresabstractInternational audience Albert Rilliard, Alexandre Allauzen, Philippe Boula de Mareüil |
INTERSPEECH | 2 |
| 2010 | Training Continuous Space Language Models: Some Practical Issues
Hai Son Le, Alexandre Allauzen, Guillaume Wisniewski, François Yvon |
EMNLP | 2 |
| 2010 | Assessing Phrase-Based Translation Models with Oracle Decoding
Guillaume Wisniewski, Alexandre Allauzen, François Yvon |
EMNLP | 2 |
| 2009 | Perception of the evolution of prosody in the French broadcast news styleabstractInternational audience Philippe Boula de Mareüil, Albert Rilliard, Alexandre Allauzen |
INTERSPEECH | 3 |
| 2008 | Training and Evaluation of POS Taggers on the French MULTITAG Corpus
Alexandre Allauzen, Hélène Bonneau-Maynard |
LREC | 1 |
| 2007 | Error detection in confusion network
Alexandre Allauzen |
INTERSPEECH | 1 |
| 2007 | A state-of-the-art statistical machine translation system based on Moses
Daniel Déchelotte, Holger Schwenk, Hélène Bonneau-Maynard, Alexandre Allauzen, Gilles Adda |
MTSummit | 4 |
| 2005 | Open Vocabulary ASR for Audiovisual Document IndexationabstractThe paper reports on an investigation of an open vocabulary recognizer that allows new words to be introduced in the recognition vocabulary, without the need to retrain or adapt the language model. This method uses special word classes, whose n-gram probabilities are estimated during the training process by discounting a mass of probability from the out of vocabulary words. A part-of-speech tagger is used to determine the word classes during language model training and for vocabulary adaptation. Metadata information provided by a French audiovisual archive institute are used to identify important document-specific missing words which are added to appropriate word classes in the system vocabulary. Pronunciations for the new words are derived by grapheme-to-phoneme conversion. On over 3 hours of broadcast news data, this approach leads to a reduction of 0.35% in the OOV rate, of 0.6% of the word error rate, with 80% of the occurrences of the newly introduced words being correctly recognized. Alexandre Allauzen, Jean-Luc Gauvain |
ICASSP (1) | 1 |
| 2005 | Diachronic vocabulary adaptation for broadcast news transcription
Alexandre Allauzen, Jean-Luc Gauvain |
INTERSPEECH | 1 |
| 2005 | Where are we in transcribing French broadcast news?abstractInternational audience Jean-Luc Gauvain, Gilles Adda, Martine Adda-Decker, Alexandre Allauzen, Véronique Gendner, Lori Lamel, Holger Schwenk |
INTERSPEECH | 4 |
| 2002 | Transcribing audio-video archivesabstractThis paper addresses the automatic transcription of audiovideo archives using a state-of-the-art broadcast news speech transcription system. A 9-hour corpus spanning the latter half of the 20th century (1945-1995) has been transcribed and an analysis of the transcription quality carried out. In addition to the challenges of transcribing heterogenous broadcast news data, we are faced with changing properties of the archive over time, such as the audio quality, the speaking style, vocabulary items and manner of expression. After assessing the performance of the transcription system, several paths are explored in an attempt to reduce the mismatch between the acoustic and language models and the archived data. Claude Barras, Alexandre Allauzen, Lori Lamel, Jean-Luc Gauvain |
ICASSP | 2 |