EDBT 2026 Demo / reviewers in the wild / expert
Pierre Colombo
dblp:229/3167 · also Pierre Jean A. Colombo
· DBLP profile ↗
26ranked-venue papers
11as first author
22since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 26 · 11 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Context is Gold to find the Gold Passage: Evaluating and Training Contextual Document EmbeddingsabstractA limitation of modern document retrieval embedding methods is that they typically encode passages (chunks) from the same documents independently, often overlooking crucial contextual information from the rest of the document that could greatly improve individual chunk representations.In this work, we introduce ConTEB (Contextaware Text Embedding Benchmark), a benchmark designed to evaluate retrieval models on their ability to leverage document-wide context.Our results show that state-of-the-art embedding models struggle in retrieval scenarios where context is required.To address this limitation, we propose InSeNT (In-sequence Negative Training), a novel contrastive posttraining approach which combined with late chunking pooling enhances contextual representation learning while preserving computational efficiency.Our method significantly improves retrieval quality on ConTEB without sacrificing base model performance.We further find chunks embedded with our method are more robust to suboptimal chunking strategies and larger retrieval corpus sizes.We opensource all artifacts at https://github.com/ illuin-tech/contextual-embeddings. Max Conti, Manuel Faysse, Gautier Viaud, Antoine Bosselut, Céline Hudelot, Pierre Colombo |
EMNLP | 6 |
| 2025 | ColPali: Efficient Document Retrieval with Vision Language ModelsabstractDocuments are visually rich structures that convey information through text, but also figures, page layouts, tables, or even fonts. Since modern retrieval systems mainly rely on the textual information they extract from document pages to index documents -often through lengthy and brittle processes-, they struggle to exploit key visual cues efficiently. This limits their capabilities in many practical document retrieval applications such as Retrieval Augmented Generation (RAG).
To benchmark current systems on visually rich document retrieval, we introduce the Visual Document Retrieval Benchmark $\textit{ViDoRe}$, composed of various page-level retrieval tasks spanning multiple domains, languages, and practical settings.
The inherent complexity and performance shortcomings of modern systems motivate a new concept; doing document retrieval by directly embedding the images of the document pages. We release $\textit{ColPali}$, a Vision Language Model trained to produce high-quality multi-vector embeddings from images of document pages. Combined with a late interaction matching mechanism, $\textit{ColPali}$ largely outperforms modern document retrieval pipelines while being drastically simpler, faster and end-to-end trainable.
We release models, data, code and benchmarks under open licenses at https://hf.co/vidore. Manuel Faysse, Hugues Sibille, Bilel Omrani, Gautier Viaud, Céline Hudelot, Pierre Colombo |
ICLR | 7 |
| 2024 | Unsupervised Layer-Wise Score Aggregation for Textual OOD DetectionabstractOut-of-distribution (OOD) detection is a rapidly growing field due to new robustness and security requirements driven by an increased number of AI-based systems. Existing OOD textual detectors often rely on anomaly scores (\textit{e.g.}, Mahalanobis distance) computed on the embedding output of the last layer of the encoder. In this work, we observe that OOD detection performance varies greatly depending on the task and layer output. More importantly, we show that the usual choice (the last layer) is rarely the best one for OOD detection and that far better results can be achieved, provided that an oracle selects the best layer. We propose a data-driven, unsupervised method to leverage this observation to combine layer-wise anomaly scores. In addition, we extend classical textual OOD benchmarks by including classification tasks with a more significant number of classes (up to 150), which reflects more realistic settings. On this augmented benchmark, we show that the proposed post-aggregation methods achieve robust and consistent results comparable to using the best layer according to an oracle while removing manual feature selection altogether. Maxime Darrin, Guillaume Staerman, Eduardo Dadalto Câmara Gomes, Jackie Chi Kit Cheung, Pablo Piantanida, Pierre Colombo |
AAAI | 6 |
| 2024 | Enhanced Hallucination Detection in Neural Machine Translation through Simple Detector AggregationabstractHallucinated translations pose significant threats and safety concerns when it comes to practical deployment of machine translation systems.Previous research works have identified that detectors exhibit complementary performance -different detectors excel at detecting different types of hallucinations.In this paper, we propose to address the limitations of individual detectors by combining them and introducing a straightforward method for aggregating multiple detectors.Our results demonstrate the efficacy of our aggregated detector, providing a promising step towards evermore reliable machine translation systems. Anas Himmi, Guillaume Staerman, Marine Picot, Pierre Colombo, Nuno Miguel Guerreiro |
EMNLP | 4 |
| 2024 | SaulLM-54B & SaulLM-141B: Scaling Up Domain Adaptation for the Legal DomainabstractIn this paper, we introduce SaulLM-medium and SaulLM-large, two large language models (LLMs) families tailored for the legal sector. These models, which feature architectures of 54 billion and 140 billion parameters, respectively, are based on the Mixtral architecture. The development of SaulLM-54B and SaulLM-140B is guided by large-scale domain adaptation, divided into strategies: (1) the exploitation of continued pretaining involving a legal corpus that includes over $400$ billion tokens, (2) the implementation of a specialized legal instruction-following protocol, and (3) the alignment of model outputs with human preferences in legal interpretations. The integration of synthetically generated data in the second and third steps enhances the models' capabilities in interpreting and processing legal texts, effectively reaching state-of-the-art performance and outperforming all previous open-source models on LegalBench Instruct. This research thoroughly explores the trade-offs involved in domain-specific adaptation at this scale, offering insights that may inform future studies on domain adaptation using strong decoder models. Building upon SaulLM-7B, this study refines the approach to produce an LLM better equipped for legal tasks and domains. Additionally, we release base, instruct and aligned versions on top of SaulLM-medium and SaulLM-large under the MIT License to facilitate reuse and collaborative research. Pierre Colombo, Telmo Pires, Malik Boudiaf, Rui Melo, Gabriel Hautreux, Etienne Malaboeuf, Johanne Charpentier, Dominic Culver, Michael Desa |
NeurIPS | 1 |
| 2024 | xcomet : Transparent Machine Translation Evaluation through Fine-grained Error DetectionabstractAbstract Widely used learned metrics for machine translation evaluation, such as Comet and Bleurt, estimate the quality of a translation hypothesis by providing a single sentence-level score. As such, they offer little insight into translation errors (e.g., what are the errors and what is their severity). On the other hand, generative large language models (LLMs) are amplifying the adoption of more granular strategies to evaluation, attempting to detail and categorize translation errors. In this work, we introduce xcomet, an open-source learned metric designed to bridge the gap between these approaches. xcomet integrates both sentence-level evaluation and error span detection capabilities, exhibiting state-of-the-art performance across all types of evaluation (sentence-level, system-level, and error span detection). Moreover, it does so while highlighting and categorizing error spans, thus enriching the quality assessment. We also provide a robustness analysis with stress tests, and show that xcomet is largely capable of identifying localized critical errors and hallucinations. Nuno Miguel Guerreiro, Ricardo Rei, Daan van Stigt, Luísa Coheur, Pierre Colombo, André F. T. Martins |
Trans. Assoc. Comput. Linguistics | 5 |
| 2023 | Optimal Transport for Unsupervised Hallucination Detection in Neural Machine TranslationabstractNeural machine translation (NMT) has become the de-facto standard in real-world machine translation applications.However, NMT models can unpredictably produce severely pathological translations, known as hallucinations, that seriously undermine user trust.It becomes thus crucial to implement effective preventive strategies to guarantee their proper functioning.In this paper, we address the problem of hallucination detection in NMT by following a simple intuition: as hallucinations are detached from the source content, they exhibit cross-attention patterns that are statistically different from those of good quality translations.We frame this problem with an optimal transport formulation and propose a fully unsupervised, plug-in detector that can be used with any attention-based NMT model.Experimental results show that our detector not only outperforms all previous model-based detectors, but is also competitive with detectors that employ external models trained on millions of samples for related tasks such as quality estimation and cross-lingual sentence similarity. Nuno Miguel Guerreiro, Pierre Colombo, Pablo Piantanida, André F. T. Martins |
ACL (1) | 2 |
| 2023 | Transductive Learning for Textual Few-Shot Classification in API-based Embedding ModelsabstractPierre Colombo, Victor Pellegrain, Malik Boudiaf, Myriam Tami, Victor Storchan, Ismail Ayed, Pablo Piantanida. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Pierre Colombo, Victor Pellegrain, Malik Boudiaf, Myriam Tami, Victor Storchan, Ismail Ben Ayed, Pablo Piantanida |
EMNLP | 1 |
| 2023 | RainProof: An Umbrella to Shield Text Generator from Out-Of-Distribution DataabstractImplementing effective control mechanisms to ensure the proper functioning and security of deployed NLP models, from translation to chatbots, is essential.A key ingredient to ensure safe system behaviour is Out-Of-Distribution (OOD) detection, which aims to detect whether an input sample is statistically far from the training distribution.Although OOD detection is a widely covered topic in classification tasks, most methods rely on hidden features output by the encoder.In this work, we focus on leveraging soft-probabilities in a black-box framework, i.e. we can access the soft-predictions but not the internal states of the model.Our contributions include: (i) RAINPROOF a Relative informAItioN Projection OOD detection framework; and (ii) a more operational evaluation setting for OOD detection.Surprisingly, we find that OOD detection is not necessarily aligned with task-specific measures.The OOD detector may filter out samples well processed by the model and keep samples that are not, leading to weaker performance.Our results show that RAINPROOF provides OOD detection methods more aligned with task-specific performance metrics than traditional OOD detectors. Maxime Darrin, Pablo Piantanida, Pierre Colombo |
EMNLP | 3 |
| 2023 | Revisiting Instruction Fine-tuned Model Evaluation to Guide Industrial ApplicationsabstractInstruction Fine-Tuning (IFT) is a powerful paradigm that strengthens the zero-shot capabilities of Large Language Models (LLMs), but in doing so induces new evaluation metric requirements.We show LLM-based metrics to be well adapted to these requirements, and leverage them to conduct an investigation of taskspecialization strategies, quantifying the tradeoffs that emerge in practical industrial settings.Our findings offer practitioners actionable insights for real-world IFT model deployment. Manuel Faysse, Gautier Viaud, Céline Hudelot, Pierre Colombo |
EMNLP | 4 |
| 2023 | Hallucinations in Large Multilingual Translation ModelsabstractAbstract Hallucinated translations can severely undermine and raise safety issues when machine translation systems are deployed in the wild. Previous research on the topic focused on small bilingual models trained on high-resource languages, leaving a gap in our understanding of hallucinations in multilingual models across diverse translation scenarios. In this work, we fill this gap by conducting a comprehensive analysis—over 100 language pairs across various resource levels and going beyond English-centric directions—on both the M2M neural machine translation (NMT) models and GPT large language models (LLMs). Among several insights, we highlight that models struggle with hallucinations primarily in low-resource directions and when translating out of English, where, critically, they may reveal toxic patterns that can be traced back to the training data. We also find that LLMs produce qualitatively different hallucinations to those of NMT models. Finally, we show that hallucinations are hard to reverse by merely scaling models trained with the same data. However, employing more diverse models, trained on different data or with different procedures, as fallback systems can improve translation quality and virtually eliminate certain pathologies. Nuno Miguel Guerreiro, Duarte M. Alves, Jonas Waldendorf, Barry Haddow, Alexandra Birch, Pierre Colombo, André F. T. Martins |
Trans. Assoc. Comput. Linguistics | 6 |
| 2022 | InfoLM: A New Metric to Evaluate Summarization & Data2Text GenerationabstractAssessing the quality of natural language generation (NLG) systems through human annotation is very expensive. Additionally, human annotation campaigns are time-consuming and include non-reusable human labour. In practice, researchers rely on automatic metrics as a proxy of quality. In the last decade, many string-based metrics (e.g., BLEU or ROUGE) have been introduced. However, such metrics usually rely on exact matches and thus, do not robustly handle synonyms. In this paper, we introduce InfoLM a family of untrained metrics that can be viewed as a string-based metric that addresses the aforementioned flaws thanks to a pre-trained masked language model. This family of metrics also makes use of information measures allowing the possibility to adapt InfoLM to different evaluation criteria. Using direct assessment, we demonstrate that InfoLM achieves statistically significant improvement and two figure correlation gains in many configurations compared to existing metrics on both summarization and data2text generation tasks. Pierre Colombo, Chloé Clavel, Pablo Piantanida |
AAAI | 1 |
| 2022 | Learning Disentangled Textual Representations via Statistical Measures of SimilarityabstractWhen working with textual data, a natural application of disentangled representations is fair classification where the goal is to make predictions without being biased (or influenced) by sensitive attributes that may be present in the data (e.g., age, gender or race).Dominant approaches to disentangle a sensitive attribute from textual representations rely on learning simultaneously a penalization term that involves either an adversarial loss (e.g., a discriminator) or an information measure (e.g., mutual information).However, these methods require the training of a deep neural network with several parameter updates for each update of the representation model.As a matter of fact, the resulting nested optimization loop is both time consuming, adding complexity to the optimization dynamic, and requires a fine hyperparameter selection (e.g., learning rates, architecture).In this work, we introduce a family of regularizers for learning disentangled representations that do not require training.These regularizers are based on statistical measures of similarity between the conditional probability distributions with respect to the sensitive attributes.Our novel regularizers do not require additional training, are faster and do not involve additional tuning while achieving better results both when combined with pretrained and randomly initialized text encoders. Pierre Colombo, Guillaume Staerman, Nathan Noiry, Pablo Piantanida |
ACL (1) | 1 |
| 2022 | Of Human Criteria and Automatic Metrics: A Benchmark of the Evaluation of Story GenerationabstractResearch on Automatic Story Generation (ASG) relies heavily on human and automatic evaluation. However, there is no consensus on which human evaluation criteria to use, and no analysis of how well automatic criteria correlate with them. In this paper, we propose to re-evaluate ASG evaluation. We introduce a set of 6 orthogonal and comprehensive human criteria, carefully motivated by the social sciences literature. We also present HANNA, an annotated dataset of 1,056 stories produced by 10 different ASG systems. HANNA allows us to quantitatively evaluate the correlations of 72 automatic metrics with human criteria. Our analysis highlights the weaknesses of current metrics for ASG and allows us to formulate practical recommendations for ASG evaluation. Cyril Chhun, Pierre Colombo, Fabian M. Suchanek, Chloé Clavel |
COLING | 2 |
| 2022 | A Differential Entropy Estimator for Training Neural NetworksabstractMutual Information (MI) has been widely used as a loss regularizer for training neural networks. This has been particularly effective when learn disentangled or compressed representations of high dimensional data. However, differential entropy (DE), another fundamental measure of information, has not found widespread use in neural network training. Although DE offers a potentially wider range of applications than MI, off-the-shelf DE estimators are either non differentiable, computationally intractable or fail to adapt to changes in the underlying distribution. These drawbacks prevent them from being used as regularizers in neural networks training. To address shortcomings in previously proposed estimators for DE, here we introduce KNIFE, a fully parameterized, differentiable kernel-based estimator of DE. The flexibility of our approach also allows us to construct KNIFE-based estimators for conditional (on either discrete or continuous variables) DE, as well as MI. We empirically validate our method on high-dimensional synthetic data and further apply it to guide the training of neural networks for real-world tasks. Our experiments on a large variety of tasks, including visual domain adaptation, textual fair classification, and textual fine-tuning demonstrate the effectiveness of KNIFE-based estimation. Code can be found at https://github.com/g-pichler/knife. Georg Pichler, Pierre Colombo, Malik Boudiaf, Günther Koliander, Pablo Piantanida |
ICML | 2 |
| 2022 | Beyond Mahalanobis Distance for Textual OOD DetectionabstractAs the number of AI systems keeps growing, it is fundamental to implement and develop efficient control mechanisms to ensure the safe and proper functioning of machine learning (ML) systems. Reliable out-of-distribution (OOD) detection aims to detect test samples that are statistically far from the training distribution, as they might cause failures of in-production systems. In this paper, we propose a new detector called TRUSTED. Different from previous works, TRUSTED key components (i) include a novel OOD score relying on the concept of statistical data depth, (ii) rely on the idea’s full potential that all hidden layers of the network carry information regarding OOD. Our extensive experiments, comparing over 51k model configurations including different checkpoints, seed and various datasets, demonstrate that TRUSTED achieve state-of-the-art performances by producing an improvement of over 3 AUROC points. Pierre Colombo, Eduardo Dadalto Câmara Gomes, Guillaume Staerman, Nathan Noiry, Pablo Piantanida |
NeurIPS | 1 |
| 2022 | What are the best Systems? New Perspectives on NLP BenchmarkingabstractIn Machine Learning, a benchmark refers to an ensemble of datasets associated with one or multiple metrics together with a way to aggregate different systems performances. They are instrumental in {\it (i)} assessing the progress of new methods along different axes and {\it (ii)} selecting the best systems for practical use. This is particularly the case for NLP with the development of large pre-trained models (\textit{e.g.} GPT, BERT) that are expected to generalize well on a variety of tasks. While the community mainly focused on developing new datasets and metrics, there has been little interest in the aggregation procedure, which is often reduced to a simple average over various performance measures. However, this procedure can be problematic when the metrics are on a different scale, which may lead to spurious conclusions. This paper proposes a new procedure to rank systems based on their performance across different tasks. Motivated by the social choice theory, the final system ordering is obtained through aggregating the rankings induced by each task and is theoretically grounded. We conduct extensive numerical experiments (on over 270k scores) to assess the soundness of our approach both on synthetic and real scores (\textit{e.g.} GLUE, EXTREM, SEVAL, TAC, FLICKR). In particular, we show that our method yields different conclusions on state-of-the-art systems than the mean-aggregation procedure while being both more reliable and robust. Pierre Colombo, Nathan Noiry, Ekhine Irurozki, Stéphan Clémençon |
NeurIPS | 1 |
| 2022 | The BigScience ROOTS Corpus: A 1.6TB Composite Multilingual DatasetabstractAs language models grow ever larger, the need for large-scale high-quality text datasets has never been more pressing, especially in multilingual settings. The BigScience workshop, a 1-year international and multidisciplinary initiative, was formed with the goal of researching and training large language models as a values-driven undertaking, putting issues of ethics, harm, and governance in the foreground. This paper documents the data creation and curation efforts undertaken by BigScience to assemble the Responsible Open-science Open-collaboration Text Sources (ROOTS) corpus, a 1.6TB dataset spanning 59 languages that was used to train the 176-billion-parameter BigScience Large Open-science Open-access Multilingual (BLOOM) language model. We further release a large initial subset of the corpus and analyses thereof, and hope to empower large-scale monolingual and multilingual modeling projects with both the data and the processing tools, as well as stimulate research around this large multilingual corpus. Hugo Laurençon, Lucile Saulnier, Thomas Wang, Christopher Akiki, Albert Villanova del Moral, Teven Le Scao, Leandro von Werra, Chenghao Mou, Eduardo G. Ponferrada, Huu Nguyen, Jörg Frohberg, Mario Sasko, Quentin Lhoest, Angelina McMillan-Major, Gérard Dupont, Stella Biderman, Anna Rogers, Loubna Ben Allal, Francesco De Toni, Giada Pistilli, Olivier Nguyen, Somaieh Nikpoor, Maraim Masoud, Pierre Colombo, Javier de la Rosa 0001, Paulo Villegas, Tristan Thrush, Shayne Longpre, Sebastian Nagel 0005, Leon Weber-Genzel, Manuel Muñoz, Daniel van Strien, Zaid Alyafeai, Khalid Almubarak, Minh Chien Vu, Itziar Gonzalez-Dios, Aitor Soroa, Kyle Lo, Manan Dey, Pedro Ortiz Suarez, Aaron Gokaslan, Shamik Bose, David Ifeoluwa Adelani, Long Phan, Hieu Tran, Ian Yu, Suhas Pai, Jenny Chim, Violette Lepercq, Suzana Ilic, Margaret Mitchell, Sasha Luccioni, Yacine Jernite |
NeurIPS | 24 |
| 2021 | A Novel Estimator of Mutual Information for Learning to Disentangle Textual RepresentationsabstractPierre Colombo, Pablo Piantanida, Chloé Clavel. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Pierre Colombo, Pablo Piantanida, Chloé Clavel |
ACL/IJCNLP (1) | 1 |
| 2021 | Improving Multimodal fusion via Mutual Dependency MaximisationabstractMultimodal sentiment analysis is a trending area of research, and the multimodal fusion is one of its most active topic.Acknowledging humans communicate through a variety of channels (i.e visual, acoustic, linguistic), multimodal systems aim at integrating different unimodal representations into a synthetic one.So far, a consequent effort has been made on developing complex architectures allowing the fusion of these modalities.However, such systems are mainly trained by minimising simple losses such as L 1 or cross-entropy.In this work, we investigate unexplored penalties and propose a set of new objectives that measure the dependency between modalities.We demonstrate that our new penalties lead to a consistent improvement (up to 4.3 on accuracy) across a large variety of state-of-the-art models on two well-known sentiment analysis datasets: CMU-MOSI and CMU-MOSEI.Our method not only achieves a new SOTA on both datasets but also produces representations that are more robust to modality drops.Finally, a by-product of our methods includes a statistical network which can be used to interpret the high dimensional representations learnt by the model. Pierre Colombo, Emile Chapuis, Matthieu Labeau, Chloé Clavel |
EMNLP (1) | 1 |
| 2021 | Code-switched inspired losses for spoken dialog representationsabstractSpoken dialog systems need to be able to handle both multiple languages and multilinguality inside a conversation (e.g in case of codeswitching).In this work, we introduce new pretraining losses tailored to learn multilingual spoken dialog representations.The goal of these losses is to expose the model to codeswitched language.To scale up training, we automatically build a pretraining corpus composed of multilingual conversations in five different languages (French, Italian, English, German and Spanish) from OpenSubtitles, a huge multilingual corpus composed of 24.3G tokens.We test the generic representations on MIAM, a new benchmark composed of five dialog act corpora on the same aforementioned languages as well as on two novel multilingual downstream tasks (i.e multilingual mask utterance retrieval and multilingual inconsistency identification).Our experiments show that our new code switched-inspired losses achieve a better performance in both monolingual and multilingual settings. Pierre Colombo, Emile Chapuis, Matthieu Labeau, Chloé Clavel |
EMNLP (1) | 1 |
| 2021 | Automatic Text Evaluation through the Lens of Wasserstein BarycentersabstractA new metric BaryScore to evaluate text generation based on deep contextualized embeddings (e.g., BERT, Roberta, ELMo) is introduced.This metric is motivated by a new framework relying on optimal transport tools, i.e., Wasserstein distance and barycenter.By modelling the layer output of deep contextualized embeddings as a probability distribution rather than by a vector embedding; this framework provides a natural way to aggregate the different outputs through the Wasserstein space topology.In addition, it provides theoretical grounds to our metric and offers an alternative to available solutions (e.g., Mover-Score and BertScore).Numerical evaluation is performed on four different tasks: machine translation, summarization, data2text generation and image captioning.Our results show that BaryScore outperforms other BERT based metrics and exhibits more consistent behaviour in particular for text summarization. Pierre Colombo, Guillaume Staerman, Chloé Clavel, Pablo Piantanida |
EMNLP (1) | 1 |
| 2020 | Guiding Attention in Sequence-to-Sequence Models for Dialogue Act PredictionabstractThe task of predicting dialog acts (DA) based on conversational dialog is a key component in the development of conversational agents. Accurately predicting DAs requires a precise modeling of both the conversation and the global tag dependencies. We leverage seq2seq approaches widely adopted in Neural Machine Translation (NMT) to improve the modelling of tag sequentiality. Seq2seq models are known to learn complex global dependencies while currently proposed approaches using linear conditional random fields (CRF) only model local tag dependencies. In this work, we introduce a seq2seq model tailored for DA classification using: a hierarchical encoder, a novel guided attention mechanism and beam search applied to both training and inference. Compared to the state of the art our model does not require handcrafted features and is trained end-to-end. Furthermore, the proposed approach achieves an unmatched accuracy score of 85% on SwDA, and state-of-the-art accuracy score of 91.6% on MRDA. Pierre Colombo, Emile Chapuis, Matteo Manica, Emmanuel Vignon, Giovanna Varni, Chloé Clavel |
AAAI | 1 |
| 2020 | The importance of fillers for text representations of speech transcriptsabstractWhile being an essential component of spoken language, fillers (e.g."um" or "uh") often remain overlooked in Spoken Language Understanding (SLU) tasks. We explore the possibility of representing them with deep contextualised embeddings, showing improvements on modelling spoken language and two downstream tasks - predicting a speaker's stance and expressed confidence. Tanvi Dinkar, Pierre Colombo, Matthieu Labeau, Chloé Clavel |
EMNLP (1) | 2 |
| 2020 | Heavy-tailed Representations, Text Polarity Classification & Data AugmentationabstractThe dominant approaches to text representation in natural language rely on learning embeddings on massive corpora which have convenient properties such as compositionality and distance preservation. In this paper, we develop a novel method to learn a heavy-tailed embedding with desirable regularity properties regarding the distributional tails, which allows to analyze the points far away from the distribution bulk using the framework of multivariate extreme value theory. In particular, a classifier dedicated to the tails of the proposed embedding is obtained which exhibits a scale invariance property exploited in a novel text generation method for label preserving dataset augmentation. Experiments on synthetic and real text data show the relevance of the proposed framework and confirm that this method generates meaningful sentences with controllable attribute, e.g. positive or negative sentiments. Hamid Jalalzai, Pierre Colombo, Chloé Clavel, Éric Gaussier, Giovanna Varni, Emmanuel Vignon, Anne Sabourin |
NeurIPS | 2 |
| 2019 | From the Token to the Review: A Hierarchical Multimodal approach to Opinion MiningabstractAlexandre Garcia, Pierre Colombo, Florence d’Alché-Buc, Slim Essid, Chloé Clavel. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Alexandre Garcia 0001, Pierre Colombo, Florence d'Alché-Buc, Slim Essid, Chloé Clavel |
EMNLP/IJCNLP (1) | 2 |