Corentin Dancette

dblp:218/5644 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
7since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Vision and language · 34% Trustworthy machine learning · 30% Language models and text generation · 11%

Topics — the 19 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language
visual question answering
1.742023
Improving Selective Visual Question Answering by Learning from Your Peers · CVPR 2023
Beyond Question-Based Biases: Assessing Multimodal Shortcut Learning in Visual Question Answering · ICCV 2021
RUBi: Reducing Unimodal Biases for Visual Question Answering · NeurIPS 2019
Computer vision › Vision and language
compositionality
0.812024
Beyond task performance: evaluating and reducing the flaws of large multimodal models with in-context-learning · ICLR 2024
Machine learning › Trustworthy machine learning
hallucination
0.812024
Beyond task performance: evaluating and reducing the flaws of large multimodal models with in-context-learning · ICLR 2024
Natural language and speech › Language models and text generation
in-context learning
0.812024
Beyond task performance: evaluating and reducing the flaws of large multimodal models with in-context-learning · ICLR 2024
Computer vision › Vision and language
multimodal in-context learning
0.812024
Beyond task performance: evaluating and reducing the flaws of large multimodal models with in-context-learning · ICLR 2024
Computer vision › Vision and language › vision-language model › multimodal large language model
multimodal large language model evaluation
0.812024
Beyond task performance: evaluating and reducing the flaws of large multimodal models with in-context-learning · ICLR 2024
Machine learning › Trustworthy machine learning
out-of-distribution generalization
0.722022
Fishr: Invariant Gradient Variances for Out-of-Distribution Generalization · ICML 2022
RUBi: Reducing Unimodal Biases for Visual Question Answering · NeurIPS 2019
Machine learning › Transfer learning and domain adaptation
fine-tuning
0.712023
Rewarded soups: towards Pareto-optimal alignment by interpolating weights fine-tuned on diverse rewards · NeurIPS 2023
Machine learning › Reinforcement learning
reinforcement learning from human feedback
0.712023
Rewarded soups: towards Pareto-optimal alignment by interpolating weights fine-tuned on diverse rewards · NeurIPS 2023
Machine learning › Reinforcement learning › reinforcement learning from human feedback
reward alignment
0.712023
Rewarded soups: towards Pareto-optimal alignment by interpolating weights fine-tuned on diverse rewards · NeurIPS 2023
Machine learning › Trustworthy machine learning › uncertainty estimation
selective classification
0.712023
Improving Selective Visual Question Answering by Learning from Your Peers · CVPR 2023
Machine learning › Trustworthy machine learning
uncertainty and abstention
0.712023
Improving Selective Visual Question Answering by Learning from Your Peers · CVPR 2023
Machine learning › Efficient and distributed learning › model composition
weight interpolation
0.712023
Rewarded soups: towards Pareto-optimal alignment by interpolating weights fine-tuned on diverse rewards · NeurIPS 2023
Machine learning › Transfer learning and domain adaptation
domain invariance
0.612022
Fishr: Invariant Gradient Variances for Out-of-Distribution Generalization · ICML 2022
Machine learning › Trustworthy machine learning › robustness
shortcut learning
0.512021
Beyond Question-Based Biases: Assessing Multimodal Shortcut Learning in Visual Question Answering · ICCV 2021
Machine learning › Trustworthy machine learning › fairness › social bias
social bias in vision-language models
0.512021
Beyond Question-Based Biases: Assessing Multimodal Shortcut Learning in Visual Question Answering · ICCV 2021
Computer vision › Vision and language › vision-language model
multimodal large language model
0.212024
Beyond task performance: evaluating and reducing the flaws of large multimodal models with in-context-learning · ICLR 2024
Computer vision › Vision and language
image captioning
0.212023
eP-ALM: Efficient Perceptual Augmentation of Language Models · ICCV 2023
Machine learning › Trustworthy machine learning
robustness
0.112019
RUBi: Reducing Unimodal Biases for Visual Question Answering · NeurIPS 2019

Methods — techniques the papers use, named apart from their topics

instruction tuning · 0.8in-context learning · 0.8RLHF · 0.8weight interpolation · 0.7parameter-efficient fine-tuning · 0.7multimodal selection function · 0.7multi-policy training · 0.7linear projection · 0.7learning from peers · 0.7fisher information · 0.6
YearPublicationVenuePosition
2025 RadSAM: Segmenting 3D Radiological Images with a 2D Promptable Model
Julien Khlaut, Elodie Ferreres, Daniel Tordjman, Hélène Philippe, Tom Boeken, Pierre Manceron, Corentin Dancette
MICCAI (4)7
2024 Beyond task performance: evaluating and reducing the flaws of large multimodal models with in-context-learning
abstract
Following the success of Large Language Models (LLMs), Large Multimodal Models (LMMs), such as the Flamingo model and its subsequent competitors, have started to emerge as natural steps towards generalist agents. However, interacting with recent LMMs reveals major limitations that are hardly captured by the current evaluation benchmarks. Indeed, task performances (e.g., VQA accuracy) alone do not provide enough clues to understand their real capabilities, limitations, and to which extent such models are aligned to human expectations. To refine our understanding of those flaws, we deviate from the current evaluation paradigm, and (1) evaluate 10 recent open-source LMMs from 3B up to 80B parameter scale, on 5 different axes; hallucinations, abstention, compositionality, explainability and instruction following. Our evaluation on these axes reveals major flaws in LMMs. While the current go-to solution to align these models is based on training, such as instruction tuning or RLHF, we rather (2) explore the training-free in-context learning (ICL) as a solution, and study how it affects these limitations. Based on our ICL study, (3) we push ICL further and propose new multimodal ICL variants such as; Multitask-ICL, Chain-of-Hindsight-ICL, and Self-Correcting-ICL. Our findings are as follows; (1) Despite their success, LMMs have flaws that remain unsolved with scaling alone. (2) The effect of ICL on LMMs flaws is nuanced; despite its effectiveness for improved explainability, answer abstention, ICL only slightly improves instruction following, does not improve compositional abilities, and actually even amplifies hallucinations. (3) The proposed ICL variants are promising as post-hoc approaches to efficiently tackle some of those flaws. The code is available here: https://github.com/mshukor/EvALign-ICL.
Mustafa Shukor, Alexandre Ramé, Corentin Dancette, Matthieu Cord
ICLR3
2023 Improving Selective Visual Question Answering by Learning from Your Peers
abstract
Despite advances in Visual Question Answering (VQA), the ability of models to assess their own correctness remains under-explored. Recent work has shown that VQA models, out-of-the-box, can have difficulties abstaining from answering when they are wrong. The option to abstain, also called Selective Prediction, is highly relevant when deploying systems to users who must trust the system's output (e.g., VQA assistants for users with visual impairments). For such scenarios, abstention can be especially important as users may provide out-of-distribution (OOD) or adversarial inputs that make incorrect answers more likely. In this work, we explore Selective VQA in both in-distribution (ID) and OOD scenarios, where models are presented with mixtures of ID and OOD data. The goal is to maximize the number of questions answered while minimizing the risk of error on those questions. We propose a simple yet effective Learning from Your Peers (LYP) approach for training multimodal selection functions for making abstention decisions. Our approach uses predictions from models trained on distinct subsets of the training data as targets for optimizing a Selective VQA model. It does not require additional manual labels or held-out data and provides a signal for identifying examples that are easy/difficult to generalize to. In our extensive evaluations, we show this benefits a number of models across different architectures and scales. Overall, for ID, we reach 32.92% in the selective prediction metric coverage at 1 % risk of error$(\mathcal{C} {@} 1\%)$which doubles the previous best coverage of 15.79% on this task. For mixed ID/OOD, using models' softmax confidences for abstention decisions performs very poorly, answering$\mathcal{C}$@1%.
Corentin Dancette, Spencer Whitehead, Rishabh Maheshwary, Ramakrishna Vedantam, Stefan Scherer, Xinlei Chen, Matthieu Cord, Marcus Rohrbach
CVPR1
2023 eP-ALM: Efficient Perceptual Augmentation of Language Models
abstract
Large Language Models (LLMs) have so far impressed the world, with unprecedented capabilities that emerge in models at large scales. On the vision side, transformer models (i.e., ViT) are following the same trend, achieving the best performance on challenging benchmarks. With the abundance of such unimodal models, a natural question arises; do we need also to follow this trend to tackle multimodal tasks? In this work, we propose to rather direct effort to efficient adaptations of existing models, and propose to augment Language Models with perception. Existing approaches for adapting pretrained models for vision-language tasks still rely on several key components that hinder their efficiency. In particular, they still train a large number of parameters, rely on large multimodal pretraining, use encoders (e.g., CLIP) trained on huge image-text datasets, and add significant inference overhead. In addition, most of these approaches have focused on Zero-Shot and In Context Learning, with little to no effort on direct finetuning. We investigate the minimal computational effort needed to adapt unimodal models for multimodal tasks and propose a new challenging setup, alongside different approaches, that efficiently adapts unimodal pretrained models. We show that by freezing more than 99% of total parameters, training only one linear projection layer, and prepending only one trainable token, our approach (dubbed eP-ALM) significantly outperforms other baselines on VQA and Captioning across Image, Video, and Audio modalities, following the proposed setup. The code is available here: https://github.com/mshukor/eP-ALM.
Mustafa Shukor, Corentin Dancette, Matthieu Cord
ICCV2
2023 Rewarded soups: towards Pareto-optimal alignment by interpolating weights fine-tuned on diverse rewards
abstract
Foundation models are first pre-trained on vast unsupervised datasets and then fine-tuned on labeled data. Reinforcement learning, notably from human feedback (RLHF), can further align the network with the intended usage. Yet the imperfections in the proxy reward may hinder the training and lead to suboptimal results; the diversity of objectives in real-world tasks and human opinions exacerbate the issue. This paper proposes embracing the heterogeneity of diverse rewards by following a multi-policy strategy. Rather than focusing on a single a priori reward, we aim for Pareto-optimal generalization across the entire space of preferences. To this end, we propose rewarded soup, first specializing multiple networks independently (one for each proxy reward) and then interpolating their weights linearly. This succeeds empirically because we show that the weights remain linearly connected when fine-tuned on diverse rewards from a shared pre-trained initialization. We demonstrate the effectiveness of our approach for text-to-text (summarization, Q&A, helpful assistant, review), text-image (image captioning, text-to-image generation, visual grounding), and control (locomotion) tasks. We hope to enhance the alignment of deep models, and how they interact with the world in all its diversity.
Alexandre Ramé, Guillaume Couairon, Corentin Dancette, Jean-Baptiste Gaya, Mustafa Shukor, Laure Soulier, Matthieu Cord
NeurIPS3
2022 Fishr: Invariant Gradient Variances for Out-of-Distribution Generalization
abstract
Learning robust models that generalize well under changes in the data distribution is critical for real-world applications. To this end, there has been a growing surge of interest to learn simultaneously from multiple training domains - while enforcing different types of invariance across those domains. Yet, all existing approaches fail to show systematic benefits under controlled evaluation protocols. In this paper, we introduce a new regularization - named Fishr - that enforces domain invariance in the space of the gradients of the loss: specifically, the domain-level variances of gradients are matched across training domains. Our approach is based on the close relations between the gradient covariance, the Fisher Information and the Hessian of the loss: in particular, we show that Fishr eventually aligns the domain-level loss landscapes locally around the final weights. Extensive experiments demonstrate the effectiveness of Fishr for out-of-distribution generalization. Notably, Fishr improves the state of the art on the DomainBed benchmark and performs consistently better than Empirical Risk Minimization. Our code is available at https://github.com/alexrame/fishr.
Alexandre Ramé, Corentin Dancette, Matthieu Cord
ICML2
2021 Beyond Question-Based Biases: Assessing Multimodal Shortcut Learning in Visual Question Answering
abstract
We introduce an evaluation methodology for visual question answering (VQA) to better diagnose cases of shortcut learning. These cases happen when a model exploits spurious statistical regularities to produce correct answers but does not actually deploy the desired behavior. There is a need to identify possible shortcuts in a dataset and assess their use before deploying a model in the real world. The research community in VQA has focused exclusively on question-based shortcuts, where a model might, for example, answer "What is the color of the sky" with "blue" by relying mostly on the question-conditional training prior and give little weight to visual evidence. We go a step further and consider multimodal shortcuts that involve both questions and images. We first identify potential shortcuts in the popular VQA v2 training set by mining trivial predictive rules such as co-occurrences of words and visual elements. We then introduce VQA-CounterExamples (VQACE), an evaluation protocol based on our subset of CounterExamples i.e. image-question-answer triplets where our rules lead to incorrect answers. We use this new evaluation in a large-scale study of existing approaches for VQA. We demonstrate that even state-of-the-art models perform poorly and that existing techniques to reduce biases are largely ineffective in this context. Our findings suggest that past work on question-based biases in VQA has only addressed one facet of a complex issue. The code for our method is available at https://github.com/cdancette/detect-shortcuts
Corentin Dancette, Rémi Cadène, Damien Teney, Matthieu Cord
ICCV1
2019 RUBi: Reducing Unimodal Biases for Visual Question Answering
abstract
Visual Question Answering (VQA) is the task of answering questions about an image. Some VQA models often exploit unimodal biases to provide the correct answer without using the image information. As a result, they suffer from a huge drop in performance when evaluated on data outside their training set distribution. This critical issue makes them unsuitable for real-world settings. We propose RUBi, a new learning strategy to reduce biases in any VQA model. It reduces the importance of the most biased examples, i.e. examples that can be correctly classified without looking at the image. It implicitly forces the VQA model to use the two input modalities instead of relying on statistical regularities between the question and the answer. We leverage a question-only model that captures the language biases by identifying when these unwanted regularities are used. It prevents the base VQA model from learning them by influencing its predictions. This leads to dynamically adjusting the loss in order to compensate for biases. We validate our contributions by surpassing the current state-of-the-art results on VQA-CP v2. This dataset is specifically designed to assess the robustness of VQA models when exposed to different question biases at test time than what was seen during training.
Rémi Cadène, Corentin Dancette, Hédi Ben-Younes, Matthieu Cord, Devi Parikh
NeurIPS2
2018 Sampling Strategies in Siamese Networks for Unsupervised Speech Representation Learning
abstract
Recent studies have investigated siamese network architectures for learning invariant speech representations using same-different side information at the word level. Here we investigate systematically an often ignored component of siamese networks: the sampling procedure (how pairs of same vs. different tokens are selected). We show that sampling strategies taking into account Zipf's Law, the distribution of speakers and the proportions of same and different pairs of words significantly impact the performance of the network. In particular, we show that word frequency compression improves learning across a large range of variations in number of training pairs. This effect does not apply to the same extent to the fully unsupervised setting, where the pairs of same-different words are obtained by spoken term discovery. We apply these results to pairs of words discovered using an unsupervised algorithm and show an improvement on state-of-the-art in unsupervised representation learning using siamese networks.
Rachid Riad, Corentin Dancette, Julien Karadayi, Neil Zeghidour, Thomas Schatz, Emmanuel Dupoux
INTERSPEECH2
2018 A K-Nearest Neighbours Approach To Unsupervised Spoken Term Discovery
abstract
The following topics are dealt with: speech recognition; neural nets; speech processing; learning (artificial intelligence); natural language processing; recurrent neural nets; speaker recognition; speech synthesis; feature extraction; and text analysis.
Alexis Thual, Corentin Dancette, Julien Karadayi, Juan Benjumea, Emmanuel Dupoux
SLT2