David Grangier

dblp:57/1192 · DBLP profile ↗
← Back
52ranked-venue papers
10as first author
20since 2021 · last 2025
0000-0002-8847-9532ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 44 · 8 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 6 · 2 first-author
YearPublicationVenuePosition
2025 No Need to Talk: Asynchronous Mixture of Language Models
abstract
We introduce SMALLTALK LM, an innovative method for training a mixture of language models in an almost asynchronous manner. Each model of the mixture specializes in distinct parts of the data distribution, without the need of high-bandwidth communication between the nodes training each model. At inference, a lightweight router directs a given sequence to a single expert, according to a short prefix. This inference scheme naturally uses a fraction of the parameters from the overall mixture model. Unlike prior works on asynchronous LLM training, our routing method does not rely on full corpus clustering or access to metadata, making it more suitable for real-world applications. Our experiments on language modeling demonstrate that SMALLTALK LM achieves significantly lower perplexity than dense model baselines for the same total training FLOPs and an almost identical inference cost. Finally, in our downstream evaluations we outperform the dense baseline on 75% of the tasks.
Anastasiia Filippova, Angelos Katharopoulos, David Grangier, Ronan Collobert
ICLR3
2025 Task-Adaptive Pretrained Language Models via Clustered-Importance Sampling
abstract
Specialist language models (LMs) focus on a specific task or domain on which they often outperform generalist LMs of the same size. However, the specialist data needed to pretrain these models is only available in limited amount for most tasks. In this work, we build specialist models from large generalist training sets instead. We adjust the training distribution of the generalist data with guidance from the limited domain-specific data. We explore several approaches, with clustered importance sampling standing out. This method clusters the generalist dataset and samples from these clusters based on their frequencies in the smaller specialist dataset. It is scalable, suitable for pretraining and continued pretraining, it works well in multi-task settings. Our findings demonstrate improvements across different domains in terms of language modeling perplexity and accuracy on multiple-choice question tasks. We also present ablation studies that examine the impact of dataset sizes, clustering configurations, and model sizes.
David Grangier, Simin Fan, Skyler Seto, Pierre Ablin
ICLR1
2025 The AdEMAMix Optimizer: Better, Faster, Older
abstract
Momentum based optimizers are central to a wide range of machine learning applications. These typically rely on an Exponential Moving Average (EMA) of gradients, which decays exponentially the present contribution of older gradients. This accounts for gradients being local linear approximations which lose their relevance as the iterate moves along the loss landscape. This work questions the use of a single EMA to accumulate past gradients and empirically demonstrates how this choice can be sub-optimal: a single EMA cannot simultaneously give a high weight to the immediate past, and a non-negligible weight to older gradients. Building on this observation, we propose AdEMAMix, a simple modification of the Adam optimizer with a mixture of two EMAs to better take advantage of past gradients. Our experiments on language modeling and image classification show---quite surprisingly---that gradients can stay relevant for tens of thousands of steps. They help to converge faster, and often to lower minima: e.g., a $1.3$B parameter AdEMAMix LLM trained on $101$B tokens performs comparably to an AdamW model trained on $197$B tokens ($+95\%$). Moreover, our method significantly slows-down model forgetting during training. Our work motivates further exploration of different types of functions to leverage past gradients, beyond EMAs.
Matteo Pagliardini, Pierre Ablin, David Grangier
ICLR3
2025 Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging
abstract
Machine learning models are routinely trained on a mixture of different data domains. Different domain weights yield very different downstream performances. We propose the Soup-of-Experts, a novel architecture that can instantiate a model at test time for any domain weights with minimal computational cost and without re-training the model. Our architecture consists of a bank of expert parameters, which are linearly combined to instantiate one model. We learn the linear combination coefficients as a function of the input domain weights. To train this architecture, we sample random domain weights, instantiate the corresponding model, and backprop through one batch of data sampled with these domain weights. We demonstrate how our approach obtains small specialized models on several language modeling tasks quickly. Soup-of-Experts are particularly appealing when one needs to ship many different specialist models quickly under a size constraint.
Pierre Ablin, Angelos Katharopoulos, Skyler Seto, David Grangier
ICML4
2025 Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection
abstract
A widespread strategy to obtain a language model that performs well on a target domain is to finetune a pretrained model to perform unsupervised next-token prediction on data from that target domain. Finetuning presents two challenges: (i) if the amount of target data is limited, as in most practical applications, the model will quickly overfit, and (ii) the model will drift away from the original model, forgetting the pretraining data and the generic knowledge that comes with it. Our goal is to derive scaling laws that quantify these two phenomena for various target domains, amounts of available target data, and model scales. We measure the efficiency of injecting pretraining data into the finetuning data mixture to avoid forgetting and mitigate overfitting. A key practical takeaway from our study is that injecting as little as $1%$ of pretraining data in the finetuning data mixture prevents the model from forgetting the pretraining set.
Louis Béthune, David Grangier, Dan Busbridge, Eleonora Gualdoni, Marco Cuturi, Pierre Ablin
ICML2
2025 Scaling Laws for Optimal Data Mixtures
abstract
Large foundation models are typically trained on data from multiple domains, with the data mixture--the proportion of each domain used--playing a critical role in model performance. The standard approach to selecting this mixture relies on trial and error, which becomes impractical for large-scale pretraining. We propose a systematic method to determine the optimal data mixture for any target domain using scaling laws. Our approach accurately predicts the loss of a model of size $N$ trained with $D$ tokens and a specific domain weight vector $h$. We validate the universality of these scaling laws by demonstrating their predictive power in three distinct and large-scale settings: large language model (LLM), native multimodal model (NMM), and large vision models (LVM) pretraining. We further show that these scaling laws can extrapolate to new data mixtures and across scales: their parameters can be accurately estimated using a few small-scale training runs, and used to estimate the performance at larger scales and unseen domain weights. The scaling laws allow to derive the optimal domain weights for any target domain under a given training budget ($N$,$D$), providing a principled alternative to costly trial-and-error methods.
Mustafa Shukor, Louis Béthune, Dan Busbridge, David Grangier, Enrico Fini, Alaaeldin El-Nouby, Pierre Ablin
NeurIPS4
2024 Rephrasing the Web: A Recipe for Compute and Data-Efficient Language Modeling
abstract
Pratyush Maini, Skyler Seto, Richard Bai, David Grangier, Yizhe Zhang, Navdeep Jaitly. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Pratyush Maini, Skyler Seto, He Bai 0002, David Grangier, Yizhe Zhang 0002, Navdeep Jaitly
ACL (1)4
2024 Aggregate-and-Adapt Natural Language Prompts for Downstream Generalization of CLIP
abstract
Large pretrained vision-language models like CLIP have shown promising generalization capability, but may struggle in specialized domains (e.g., satellite imagery) or fine-grained classification (e.g., car models) where the visual concepts are unseen or under-represented during pretraining. Prompt learning offers a parameter-efficient finetuning framework that can adapt CLIP to downstream tasks even when limited annotation data are available. In this paper, we improve prompt learning by distilling the textual knowledge from natural language prompts (either human- or LLM-generated) to provide rich priors for those under-represented concepts. We first obtain a prompt ``summary'' aligned to each input image via a learned prompt aggregator. Then we jointly train a prompt generator, optimized to produce a prompt embedding that stays close to the aggregated summary while minimizing task loss at the same time. We dub such prompt embedding as Aggregate-and-Adapted Prompt Embedding (AAPE). AAPE is shown to be able to generalize to different downstream data distributions and tasks, including vision-language understanding tasks (e.g., few-shot classification, VQA) and generation tasks (image captioning) where AAPE achieves competitive performance. We also show AAPE is particularly helpful to handle non-canonical and OOD examples. Furthermore, AAPE learning eliminates LLM-based inference cost as required by baselines, and scales better with data and LLM model size.
Chen Huang 0001, Skyler Seto, Samira Abnar, David Grangier, Navdeep Jaitly, Joshua M. Susskind
NeurIPS4
2023 AudioLM: A Language Modeling Approach to Audio Generation
abstract
We introduce AudioLM, a framework for high-quality audio generation with long-term consistency. AudioLM maps the input audio to a sequence of discrete tokens and casts audio generation as a language modeling task in this representation space. We show how existing audio tokenizers provide different trade-offs between reconstruction quality and long-term structure, and we propose a hybrid tokenization scheme to achieve both objectives. Namely, we leverage the discretized activations of a masked language model pre-trained on audio to capture long-term structure and the discrete codes produced by a neural audio codec to achieve high-quality synthesis. By training on large corpora of raw audio waveforms, AudioLM learns to generate natural and coherent continuations given short prompts. When trained on speech, and without any transcript or annotation, AudioLM generates syntactically and semantically plausible speech continuations while also maintaining speaker identity and prosody for unseen speakers. Furthermore, we demonstrate how our approach extends beyond speech by generating coherent piano music continuations, despite being trained without any symbolic representation of music.
Zalan Borsos, Raphaël Marinier, Damien Vincent, Eugene Kharitonov, Olivier Pietquin, Matthew Sharifi, Dominik Roblek, Olivier Teboul, David Grangier, Marco Tagliasacchi, Neil Zeghidour
IEEE ACM Trans. Audio Speech Lang. Process.9
2022 The Trade-offs of Domain Adaptation for Neural Language Models
abstract
This work connects language model adaptation with concepts of machine learning theory.We consider a training setup with a large outof-domain set and a small in-domain set.We derive how the benefit of training a model on either set depends on the size of the sets and the distance between their underlying distributions.We analyze how out-of-domain pretraining before in-domain fine-tuning achieves better generalization than either solution independently.Finally, we present how adaptation techniques based on data selection, such as importance sampling, intelligent data selection and influence functions, can be presented in a common framework which highlights their similarity and also their subtle differences.
David Grangier, Dan Iter
ACL (1)1
2022 Learning Strides in Convolutional Neural Networks
Rachid Riad, Olivier Teboul, David Grangier, Neil Zeghidour
ICLR3
2022 On Systematic Style Differences between Unsupervised and Supervised MT and an Application for High-Resource Machine Translation
abstract
Modern unsupervised machine translation (MT) systems reach reasonable translation quality under clean and controlled data conditions.As the performance gap between supervised and unsupervised MT narrows, it is interesting to ask whether the different training methods result in systematically different output beyond what is visible via quality metrics like adequacy or BLEU.We compare translations from supervised and unsupervised MT systems of similar quality, finding that unsupervised output is more fluent and more structurally different in comparison to human translation than is supervised MT.We then demonstrate a way to combine the benefits of both methods into a single system which results in improved adequacy and fluency as rated by human evaluators.Our results open the door to interesting discussions about how supervised and unsupervised MT might be different yet mutually-beneficial.
Kelly Marchisio, Markus Freitag, David Grangier
NAACL-HLT3
2022 High Quality Rather than High Model Probability: Minimum Bayes Risk Decoding with Neural Metrics
abstract
Abstract In Neural Machine Translation, it is typically assumed that the sentence with the highest estimated probability should also be the translation with the highest quality as measured by humans. In this work, we question this assumption and show that model estimates and translation quality only vaguely correlate. We apply Minimum Bayes Risk (MBR) decoding on unbiased samples to optimize diverse automated metrics of translation quality as an alternative inference strategy to beam search. Instead of targeting the hypotheses with the highest model probability, MBR decoding extracts the hypotheses with the highest estimated quality. Our experiments show that the combination of a neural translation model with a neural reference-based metric, Bleurt, results in significant improvement in human evaluations. This improvement is obtained with translations different from classical beam-search output: These translations have much lower model likelihood and are less favored by surface metrics like Bleu.
Markus Freitag, David Grangier, Qijun Tan
Trans. Assoc. Comput. Linguistics2
2021 Dive: End-to-End Speech Diarization Via Iterative Speaker Embedding
abstract
We introduce DIVE, an end-to-end speaker diarization sys-tem. DIVE presents the diarization task as an iterative pro-cess: it repeatedly builds a representation for each speaker before predicting their voice activity conditioned on the ex-tracted representations. This strategy intrinsically resolves the speaker ordering ambiguity without requiring the classi-cal permutation invariant training loss. In contrast with prior work, our model does not rely on pretrained speaker represen-tations and jointly optimizes all parameters of the system with a multi-speaker voice activity loss. DIVE does not require the training speaker identities and allows efficient window-based training. Importantly, our loss explicitly excludes unreliable speaker turn boundaries from training, which is adapted to the standard collar-based Diarization Error Rate (DER) eval-uation. Overall, these contributions yield a system redefining the state-of-the-art on the CALLHOME benchmark, with 6.7% DER compared to 7.8% for the best alternative.
Neil Zeghidour, Olivier Teboul, David Grangier
ASRU3
2021 Learning From Heterogeneous Eeg Signals with Differentiable Channel Reordering
abstract
We propose CHARM, a method for training a single neural network across inconsistent input channels. Our work is motivated by Electroencephalography (EEG), where data collection protocols from different headsets result in varying channel ordering and number, which limits the feasibility of transferring trained systems across datasets. Our approach builds upon attention mechanisms to estimate a latent reordering matrix from each input signal and map input channels to a canonical order. CHARM is differentiable and can be composed further with architectures expecting a consistent channel ordering to build end-to-end trainable classifiers. We perform experiments on four EEG classification datasets and demonstrate the efficacy of CHARM via simulated shuffling and masking of input channels. Moreover, our method improves the transfer of pre-trained representations between datasets collected with different protocols.
Aaqib Saeed, David Grangier, Olivier Pietquin, Neil Zeghidour
ICASSP2
2021 Contrastive Learning of General-Purpose Audio Representations
abstract
We introduce COLA, a self-supervised pre-training approach for learning a general-purpose representation of audio. Our approach is based on contrastive learning: it learns a representation which assigns high similarity to audio segments extracted from the same recording while assigning lower similarity to segments from different recordings. We build on top of recent advances in contrastive learning for computer vision and reinforcement learning to design a lightweight, easy-to-implement self-supervised model of audio. We pre-train embeddings on the large-scale Audioset database and transfer these representations to 9 diverse classification tasks, including speech, music, animal sounds, and acoustic scenes. We show that despite its simplicity, our method significantly outperforms previous self-supervised systems. We furthermore conduct ablation studies to identify key design choices and release a library1to pre-train and fine-tune COLA models.
Aaqib Saeed, David Grangier, Neil Zeghidour
ICASSP2
2021 Auxiliary Task Update Decomposition: the Good, the Bad and the neutral
Lucio M. Dery, Yann N. Dauphin, David Grangier
ICLR3
2021 Experts, Errors, and Context: A Large-Scale Study of Human Evaluation for Machine Translation
abstract
Abstract Human evaluation of modern high-quality machine translation systems is a difficult problem, and there is increasing evidence that inadequate evaluation procedures can lead to erroneous conclusions. While there has been considerable research on human evaluation, the field still lacks a commonly accepted standard procedure. As a step toward this goal, we propose an evaluation methodology grounded in explicit error analysis, based on the Multidimensional Quality Metrics (MQM) framework. We carry out the largest MQM research study to date, scoring the outputs of top systems from the WMT 2020 shared task in two language pairs using annotations provided by professional translators with access to full document context. We analyze the resulting data extensively, finding among other results a substantially different ranking of evaluated systems from the one established by the WMT crowd workers, exhibiting a clear preference for human over machine output. Surprisingly, we also find that automatic metrics based on pre-trained embeddings can outperform human crowd workers. We make our corpus publicly available for further research.
Markus Freitag, George F. Foster, David Grangier, Viresh Ratnakar, Qijun Tan, Wolfgang Macherey
Trans. Assoc. Comput. Linguistics3
2021 Efficient Content-Based Sparse Attention with Routing Transformers
abstract
Self-attention has recently been adopted for a wide range of sequence modeling problems. Despite its effectiveness, self-attention suffers from quadratic computation and memory requirements with respect to sequence length. Successful approaches to reduce this complexity focused on attending to local sliding windows or a small set of locations independent of content. Our work proposes to learn dynamic sparse attention patterns that avoid allocating computation and memory to attend to content unrelated to the query of interest. This work builds upon two lines of research: It combines the modeling flexibility of prior work on content-based sparse attention with the efficiency gains from approaches based on local, temporal sparse attention. Our model, the Routing Transformer, endows self-attention with a sparse routing module based on online k-means while reducing the overall complexity of attention to O( n 1.5 d) from O( n 2 d) for sequence length n and hidden dimension d. We show that our model outperforms comparable sparse attention models on language modeling on Wikitext-103 (15.8 vs 18.3 perplexity), as well as on image generation on ImageNet-64 (3.43 vs 3.44 bits/dim) while using fewer self-attention layers. Additionally, we set a new state-of-the-art on the newly released PG-19 data-set, obtaining a test perplexity of 33.2 with a 22 layer Routing Transformer model trained on sequences of length 8192. We open-source the code for Routing Transformer in Tensorflow. 1
Aurko Roy, Mohammad Saffar, Ashish Vaswani, David Grangier
Trans. Assoc. Comput. Linguistics4
2021 Wavesplit: End-to-End Speech Separation by Speaker Clustering
abstract
We introduce Wavesplit, an end-to-end source separation system. From a single mixture, the model infers a representation for each source and then estimates each source signal given the inferred representations. The model is trained to jointly perform both tasks from the raw waveform. Wavesplit infers a set of source representations via clustering, which addresses the fundamental permutation problem of separation. For speech separation, our sequence-wide speaker representations provide a more robust separation of long, challenging recordings compared to prior work. Wavesplit redefines the state-of-the-art on clean mixtures of 2 or 3 speakers (WSJ0-2/3mix), as well as in noisy and reverberated settings (WHAM/WHAMR). We also set a new benchmark on the recent LibriMix dataset. Finally, we show that Wavesplit is also applicable to other domains, by separating fetal and maternal heart rates from a single abdominal electrocardiogram.
Neil Zeghidour, David Grangier
IEEE ACM Trans. Audio Speech Lang. Process.2
2020 Toward Better Storylines with Sentence-Level Language Models
abstract
We propose a sentence-level language model which selects the next sentence in a story from a finite set of fluent alternatives.Since it does not need to model fluency, the sentence-level language model can focus on longer range dependencies, which are crucial for multisentence coherence.Rather than dealing with individual words, our method treats the story so far as a list of pre-trained sentence embeddings and predicts an embedding for the next sentence, which is more efficient than predicting word embeddings.Notably this allows us to consider a large number of candidates for the next sentence during training.We demonstrate the effectiveness of our approach with state-of-the-art accuracy on the unsupervised Story Cloze task and with promising results on larger-scale next sentence prediction tasks.
Daphne Ippolito, David Grangier, Douglas Eck, Chris Callison-Burch
ACL2
2020 Translationese as a Language in "Multilingual" NMT
abstract
Machine translation has an undesirable propensity to produce "translationese" artifacts, which can lead to higher BLEU scores while being liked less by human raters.Motivated by this, we model translationese and original (i.e.natural) text as separate languages in a multilingual model, and pose the question: can we perform zero-shot translation between original source text and original target text?There is no data with original source and original target, so we train a sentence-level classifier to distinguish translationese from original target text, and use this classifier to tag the training data for an NMT model.Using this technique we bias the model to produce more natural outputs at test time, yielding gains in human evaluation scores on both adequacy and fluency.Additionally, we demonstrate that it is possible to bias the model to produce translationese and game the BLEU score, increasing it while decreasing human-rated quality.We analyze these outputs using metrics measuring the degree of translationese, and present an analysis of the volatility of heuristic-based train-data tagging.
Parker Riley, Isaac Caswell, Markus Freitag, David Grangier
ACL4
2020 BLEU might be Guilty but References are not Innocent
abstract
The quality of automatic metrics for machine translation has been increasingly called into question, especially for high-quality systems.This paper demonstrates that, while choice of metric is important, the nature of the references is also critical.We study different methods to collect references and compare their value in automated evaluation by reporting correlation with human evaluation for a variety of systems and metrics.Motivated by the finding that typical references exhibit poor diversity, concentrating around translationese language, we develop a paraphrasing task for linguists to perform on existing reference translations, which counteracts this bias.Our method yields higher correlation with human judgment not only for the submissions of WMT 2019 English→German, but also for Back-translation and APE augmented MT output, which have been shown to have low correlation with automatic metrics using standard references.We demonstrate that our methodology improves correlation with all modern evaluation metrics we look at, including embedding-based methods.To complete this picture, we reveal that multireference BLEU does not improve the correlation for high quality output, and present an alternative multi-reference formulation that is more effective.
Markus Freitag, David Grangier, Isaac Caswell
EMNLP (1)2
2020 Modeling Human Motion with Quaternion-Based Neural Networks
Dario Pavllo, Christoph Feichtenhofer, Michael Auli, David Grangier
Int. J. Comput. Vis.4
2019 ELI5: Long Form Question Answering
abstract
We introduce the first large-scale corpus for long-form question answering, a task requiring elaborate and in-depth answers to openended questions.The dataset comprises 270K threads from the Reddit forum "Explain Like I'm Five" (ELI5) where an online community provides answers to questions which are comprehensible by five year olds.Compared to existing datasets, ELI5 comprises diverse questions requiring multi-sentence answers.We provide a large set of web documents to help answer the question.Automatic and human evaluations show that an abstractive model trained with a multi-task objective outperforms conventional Seq2Seq, language modeling, as well as a strong extractive baseline.However, our best model is still far from human performance since raters prefer gold responses in over 86% of cases, leaving ample opportunity for future improvement.1
Angela Fan, Yacine Jernite, Ethan Perez, David Grangier, Jason Weston, Michael Auli
ACL (1)4
2019 Unsupervised Paraphrasing without Translation
abstract
Paraphrasing exemplifies the ability to abstract semantic content from surface forms.Recent work on automatic paraphrasing is dominated by methods leveraging Machine Translation (MT) as an intermediate step.This contrasts with humans, who can paraphrase without being bilingual.This work proposes to learn paraphrasing models from an unlabeled monolingual corpus only.To that end, we propose a residual variant of vector-quantized variational auto-encoder.We compare with MT-based approaches on paraphrase identification, generation, and training augmentation.Monolingual paraphrasing outperforms unsupervised translation in all settings.Comparisons with supervised translation are more mixed: monolingual paraphrasing is interesting for identification and augmentation; supervised translation is superior for generation.
Aurko Roy, David Grangier
ACL (1)2
2019 3D Human Pose Estimation in Video With Temporal Convolutions and Semi-Supervised Training
abstract
In this work, we demonstrate that 3D poses in video can be effectively estimated with a fully convolutional model based on dilated temporal convolutions over 2D keypoints. We also introduce back-projection, a simple and effective semi-supervised training method that leverages unlabeled video data. We start with predicted 2D keypoints for unlabeled video, then estimate 3D poses and finally back-project to the input 2D keypoints. In the supervised setting, our fully-convolutional model outperforms the previous best result from the literature by 6 mm mean per-joint position error on Human3.6M, corresponding to an error reduction of 11%, and the model also shows significant improvements on HumanEva-I. Moreover, experiments with back-projection show that it comfortably outperforms previous state-of-the-art results in semi-supervised settings where labeled data is scarce. Code and models are available at https://github.com/facebookresearch/VideoPose3D.
Dario Pavllo, Christoph Feichtenhofer, David Grangier, Michael Auli
CVPR3
2018 QuaterNet: A Quaternion-based Recurrent Model for Human Motion
Dario Pavllo, David Grangier, Michael Auli
BMVC2
2018 Understanding Back-Translation at Scale
abstract
An effective method to improve neural machine translation with monolingual data is to augment the parallel training corpus with back-translations of target language sentences.This work broadens the understanding of back-translation and investigates a number of methods to generate synthetic source sentences.We find that in all but resource poor settings back-translations obtained via sampling or noised beam outputs are most effective.Our analysis shows that sampling or noisy synthetic data gives a much stronger training signal than data generated by beam or greedy search.We also compare how synthetic data compares to genuine bitext and study various domain effects.Finally, we scale to hundreds of millions of monolingual sentences and achieve a new state of the art of 35 BLEU on the WMT'14 English-German test set.
Sergey Edunov, Myle Ott, Michael Auli, David Grangier
EMNLP4
2018 Analyzing Uncertainty in Neural Machine Translation
abstract
Machine translation is a popular test bed for research in neural sequence-to-sequence models but despite much recent research, there is still a lack of understanding of these models. Practitioners report performance degradation with large beams, the under-estimation of rare words and a lack of diversity in the final translations. Our study relates some of these issues to the inherent uncertainty of the task, due to the existence of multiple valid translations for a single source sentence, and to the extrinsic uncertainty caused by noisy training data. We propose tools and metrics to assess how uncertainty in the data is captured by the model distribution and how it affects search strategies that generate translations. Our results show that search works remarkably well but that the models tend to spread too much probability mass over the hypothesis space. Next, we propose tools to assess model calibration and show how to easily fix some shortcomings of current models. We release both code and multiple human reference translations for two popular benchmarks.
Myle Ott, Michael Auli, David Grangier, Marc'Aurelio Ranzato
ICML3
2018 Classical Structured Prediction Losses for Sequence to Sequence Learning
abstract
Sergey Edunov, Myle Ott, Michael Auli, David Grangier, Marc’Aurelio Ranzato. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Sergey Edunov, Myle Ott, Michael Auli, David Grangier, Marc'Aurelio Ranzato
NAACL-HLT4
2018 QuickEdit: Editing Text & Translations by Crossing Words Out
abstract
David Grangier, Michael Auli. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
David Grangier, Michael Auli
NAACL-HLT1
2017 A Convolutional Encoder Model for Neural Machine Translation
abstract
The prevalent approach to neural machine translation relies on bi-directional LSTMs to encode the source sentence.We present a faster and simpler architecture based on a succession of convolutional layers.This allows to encode the source sentence simultaneously compared to recurrent networks for which computation is constrained by temporal dependencies.On WMT'16 English-Romanian translation we achieve competitive accuracy to the state-of-the-art and on WMT'15 English-German we outperform several recently published results.Our models obtain almost the same accuracy as a very deep LSTM setup on WMT'14 English-French translation.We speed up CPU decoding by more than two times at the same or higher accuracy as a strong bidirectional LSTM. 1
Jonas Gehring, Michael Auli, David Grangier, Yann N. Dauphin
ACL (1)3
2017 Language Modeling with Gated Convolutional Networks
abstract
The pre-dominant approach to language modeling to date is based on recurrent neural networks. Their success on this task is often linked to their ability to capture unbounded context. In this paper we develop a finite context approach through stacked convolutions, which can be more efficient since they allow parallelization over sequential tokens. We propose a novel simplified gating mechanism that outperforms Oord et al. (2016) and investigate the impact of key architectural decisions. The proposed approach achieves state-of-the-art on the WikiText-103 benchmark, even though it features long-term dependencies, as well as competitive results on the Google Billion Words benchmark. Our model reduces the latency to score a sentence by an order of magnitude compared to a recurrent baseline. To our knowledge, this is the first time a non-recurrent approach is competitive with strong recurrent models on these large scale language tasks.
Yann N. Dauphin, Angela Fan, Michael Auli, David Grangier
ICML4
2017 Convolutional Sequence to Sequence Learning
abstract
The prevalent approach to sequence to sequence learning maps an input sequence to a variable length output sequence via recurrent neural networks. We introduce an architecture based entirely on convolutional neural networks. Compared to recurrent models, computations over all elements can be fully parallelized during training to better exploit the GPU hardware and optimization is easier since the number of non-linearities is fixed and independent of the input length. Our use of gated linear units eases gradient propagation and we equip each decoder layer with a separate attention module. We outperform the accuracy of the deep LSTM setup of Wu et al. (2016) on both WMT’14 English-German and WMT’14 English-French translation at an order of magnitude faster speed, both on GPU and CPU.
Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, Yann N. Dauphin
ICML3
2017 Efficient softmax approximation for GPUs
abstract
We propose an approximate strategy to efficiently train neural network based language models over very large vocabularies. Our approach, called adaptive softmax, circumvents the linear dependency on the vocabulary size by exploiting the unbalanced word distribution to form clusters that explicitly minimize the expectation of computation time. Our approach further reduces the computational cost by exploiting the specificities of modern architectures and matrix-matrix vector operations, making it particularly suited for graphical processing units. Our experiments carried out on standard benchmarks, such as EuroParl and One Billion Word, show that our approach brings a large gain in efficiency over standard approximations while achieving an accuracy close to that of the full softmax. The code of our method is available at https://github.com/facebookresearch/adaptive-softmax.
Edouard Grave, Armand Joulin, Moustapha Cissé, David Grangier, Hervé Jégou
ICML4
2016 Strategies for Training Large Vocabulary Neural Language Models
abstract
Training neural network language models over large vocabularies is computationally costly compared to count-based models such as Kneser-Ney.We present a systematic comparison of neural strategies to represent and train large vocabularies, including softmax, hierarchical softmax, target sampling, noise contrastive estimation and self normalization.We extend self normalization to be a proper estimator of likelihood and introduce an efficient variant of softmax.We evaluate each method on three popular benchmarks, examining performance on rare words, the speed/accuracy trade-off and complementarity to Kneser-Ney.
David Grangier, Michael Auli
ACL (1)2
2016 Neural Text Generation from Structured Data with Application to the Biography Domain
abstract
This paper introduces a neural model for concept-to-text generation that scales to large, rich domains.It generates biographical sentences from fact tables on a new dataset of biographies from Wikipedia.This set is an order of magnitude larger than existing resources with over 700k samples and a 400k vocabulary.Our model builds on conditional neural language models for text generation.To deal with the large vocabulary, we extend these models to mix a fixed vocabulary with copy actions that transfer sample-specific words from the input database to the generated output sentence.To deal with structured data, we allow the model to embed words differently depending on the data fields in which they occur.Our neural model significantly outperforms a Templated Kneser-Ney language model by nearly 15 BLEU.
Rémi Lebret, David Grangier, Michael Auli
EMNLP2
2012 Learning from Heterogeneous Sources via Gradient Boosting Consensus
abstract
Multiple data sources containing different types of features may be available for a given task. For instance, users' profiles can be used to build recommendation systems. In addition, a model can also use users' historical behaviors and social networks to infer users' interests on related products. We argue that it is desirable to collectively use any available multiple heterogeneous data sources in order to build effective learning models. We call this framework heterogeneous learning. In our proposed setting, data sources can include (i) non-overlapping features, (ii) non-overlapping instances, and (iii) multiple networks (i.e. graphs) that connect instances. In this paper, we propose a general optimization framework for heterogeneous learning, and devise a corresponding learning model from gradient boosting. The idea is to minimize the empirical loss with two constraints: (1) There should be consensus among the predictions of overlapping instances (if any) from different data sources; (2) Connected instances in graph datasets may have similar predictions. The objective function is solved by stochastic gradient boosting trees. Furthermore, a weighting strategy is designed to emphasize informative data sources, and deemphasize the noisy ones. We formally prove that the proposed strategy leads to a tighter error bound. This approach consistently outperforms a standard concatenation of data sources on movie rating prediction, number recognition and terrorist attack detection tasks. We observe that the proposed model can improve out-of-sample error rate by as much as 80%.
Xiaoxiao Shi, Jean-François Paiement, David Grangier, Philip S. Yu
SDM3
2010 Label Embedding Trees for Large Multi-Class Tasks
abstract
Multi-class classification becomes challenging at test time when the number of classes is very large and testing against every possible class can become computationally infeasible. This problem can be alleviated by imposing (or learning) a structure over the set of classes. We propose an algorithm for learning a tree-structure of classifiers which, by optimizing the overall tree loss, provides superior accuracy to existing tree labeling methods. We also propose a method that learns to embed labels in a low dimensional space that is faster than non-embedding approaches and has superior accuracy to existing embedding approaches. Finally we combine the two ideas resulting in the label embedding tree that outperforms alternative methods including One-vs-Rest while being orders of magnitude faster.
Samy Bengio, Jason Weston, David Grangier
NIPS3
2010 Feature Set Embedding for Incomplete Data
abstract
We present a new learning strategy for classification problems in which train and/or test data suffer from missing features. In previous work, instances are represented as vectors from some feature space and one is forced to impute missing values or to consider an instance-specific subspace. In contrast, our method considers instances as sets of (feature,value) pairs which naturally handle the missing value case. Building onto this framework, we propose a classification strategy for sets. Our proposal maps (feature,value) pairs into an embedding space and then non-linearly combines the set of embedded vectors. The embedding and the combination parameters are learned jointly on the final classification objective. This simple strategy allows great flexibility in encoding prior knowledge about the features in the embedding step and yields advantageous results compared to alternative solutions over several datasets.
David Grangier, Iain Melvin
NIPS1
2010 Learning to rank with (a lot of) word features
Jason Weston, David Grangier, Ronan Collobert, Kunihiko Sadamasa, Yanjun Qi, Olivier Chapelle, Kilian Q. Weinberger
Inf. Retr.3
2009 Supervised semantic indexing
abstract
In this article we propose Supervised Semantic Indexing (SSI), an algorithm that is trained on (query, document) pairs of text documents to predict the quality of their match. Like Latent Semantic Indexing (LSI), our models take account of correlations between words (synonymy, polysemy). However, unlike LSI our models are trained with a supervised signal directly on the ranking task of interest, which we argue is the reason for our superior results. As the query and target texts are modeled separately, our approach is easily generalized to different retrieval tasks, such as online advertising placement. Dealing with models on all pairs of words features is computationally challenging. We propose several improvements to our basic model for addressing this issue, including low rank (but diagonal preserving) representations, and correlated feature hashing (CFH). We provide an empirical study of all these methods on retrieval tasks based on Wikipedia documents as well as an Internet advertisement task. We obtain state-of-the-art performance while providing realistically scalable methods.
Jason Weston, David Grangier, Ronan Collobert, Kunihiko Sadamasa, Yanjun Qi, Olivier Chapelle, Kilian Q. Weinberger
CIKM3
2009 Supervised Semantic Indexing
Jason Weston, Ronan Collobert, David Grangier
ECIR4
2009 Polynomial Semantic Indexing
abstract
We present a class of nonlinear (polynomial) models that are discriminatively trained to directly map from the word content in a query-document or document-document pair to a ranking score. Dealing with polynomial models on word features is computationally challenging. We propose a low rank (but diagonal preserving) representation of our polynomial models to induce feasible memory and computation requirements. We provide an empirical study on retrieval tasks based on Wikipedia documents, where we obtain state-of-the-art performance while providing realistically scalable methods.
Jason Weston, David Grangier, Ronan Collobert, Kunihiko Sadamasa, Yanjun Qi, Corinna Cortes, Mehryar Mohri
NIPS3
2009 Discriminative keyword spotting
Joseph Keshet, David Grangier, Samy Bengio
Speech Commun.2
2008 A Discriminative Kernel-Based Approach to Rank Images from Text Queries
abstract
This paper introduces a discriminative model for the retrieval of images from text queries. Our approach formalizes the retrieval task as a ranking problem, and introduces a learning procedure optimizing a criterion related to the ranking performance. The proposed model hence addresses the retrieval problem directly and does not rely on an intermediate image annotation task, which contrasts with previous research. Moreover, our learning procedure builds upon recent work on the online learning of kernel-based classifiers. This yields an efficient, scalable algorithm, which can benefit from recent kernels developed for image comparison. The experiments performed over stock photography data show the advantage of our discriminative ranking approach over state-of-the-art alternatives (e.g. our model yields 26.3% average precision over the Corel dataset, which should be compared to 22.0%, for the best alternative model evaluated). Further analysis of the results shows that our model is especially advantageous over difficult queries such as queries with few relevant pictures or multiple-word queries.
David Grangier, Samy Bengio
IEEE Trans. Pattern Anal. Mach. Intell.1
2007 Learning the inter-frame distance for discriminative template-based keyword detection
abstract
This paper proposes a discriminative approach to template-based keyword detection. We introduce a method to learn the distance used to compare acoustic frames, a crucial element for template matching approaches. The proposed algorithm estimates the distance from data, with the objective to produce a detector maximizing the Area Under the receiver operating Curve (AUC), i.e. the standard evaluation measure for the keyword detection problem. The experiments performed over a large corpus, SpeechDatII, suggest that our model is effective compared to an HMM system, e.g. the proposed approach reaches 93.8\% of averaged AUC compared to 87.9\% for the HMM.
David Grangier, Samy Bengio
INTERSPEECH1
2006 A Discriminative Approach for the Retrieval of Images from Text Queries
David Grangier, Florent Monay, Samy Bengio
ECML1
2006 A Neural Network to Retrieve Images from Text Queries
David Grangier, Samy Bengio
ICANN (2)1
2005 Inferring document similarity from hyperlinks
abstract
Assessing semantic similarity between text documents is a crucial aspect in Information Retrieval systems. In this work, we propose to use hyperlink information to derive a similarity measure that can then be applied to compare any text documents, with or without hyperlinks. As linked documents are generally semantically closer than unlinked documents, we use a training corpus with hyperlinks to infer a function a,b → sim(a,b) that assigns a higher value to linked documents than to unlinked ones. Two sets of experiments on different corpora show that this function compares favorably with OKAPI matching on document retrieval tasks.
David Grangier, Samy Bengio
CIKM1
2005 Effect of segmentation method on video retrieval performance
abstract
This paper presents experiments that evaluate the effect of different video segmentation methods on text-based video retrieval. Segmentations relying on modalities like speech, video and text or their combination are compared with a baseline sliding window segmentation. The results suggest that even with the sliding window segmentation, acceptable performance can be obtained on a broadcast news retrieval task. Moreover, in the case where manually segmented data are available for training, the approach combining the different modalities can lead to IR results close to those obtained with a manual segmentation.
David Grangier, Alessandro Vinciarelli
ICME1