EDBT 2026 Demo / reviewers in the wild / expert
Corentin Dancette
dblp:218/5644
· DBLP profile ↗
10ranked-venue papers
2as first author
7since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Vision and language · 34% Trustworthy machine learning · 30% Language models and text generation · 11% |
Topics — the 19 heaviest of 20, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language
visual question answering |
1.7 | 4 | 2023 | Improving Selective Visual Question Answering by Learning from Your Peers · CVPR 2023 Beyond Question-Based Biases: Assessing Multimodal Shortcut Learning in Visual Question Answering · ICCV 2021 RUBi: Reducing Unimodal Biases for Visual Question Answering · NeurIPS 2019 |
Computer vision › Vision and language
compositionality |
0.8 | 1 | 2024 | Beyond task performance: evaluating and reducing the flaws of large multimodal models with in-context-learning · ICLR 2024 |
Machine learning › Trustworthy machine learning
hallucination |
0.8 | 1 | 2024 | Beyond task performance: evaluating and reducing the flaws of large multimodal models with in-context-learning · ICLR 2024 |
Natural language and speech › Language models and text generation
in-context learning |
0.8 | 1 | 2024 | Beyond task performance: evaluating and reducing the flaws of large multimodal models with in-context-learning · ICLR 2024 |
Computer vision › Vision and language
multimodal in-context learning |
0.8 | 1 | 2024 | Beyond task performance: evaluating and reducing the flaws of large multimodal models with in-context-learning · ICLR 2024 |
Computer vision › Vision and language › vision-language model › multimodal large language model
multimodal large language model evaluation |
0.8 | 1 | 2024 | Beyond task performance: evaluating and reducing the flaws of large multimodal models with in-context-learning · ICLR 2024 |
Machine learning › Trustworthy machine learning
out-of-distribution generalization |
0.7 | 2 | 2022 | Fishr: Invariant Gradient Variances for Out-of-Distribution Generalization · ICML 2022 RUBi: Reducing Unimodal Biases for Visual Question Answering · NeurIPS 2019 |
Machine learning › Transfer learning and domain adaptation
fine-tuning |
0.7 | 1 | 2023 | Rewarded soups: towards Pareto-optimal alignment by interpolating weights fine-tuned on diverse rewards · NeurIPS 2023 |
Machine learning › Reinforcement learning
reinforcement learning from human feedback |
0.7 | 1 | 2023 | Rewarded soups: towards Pareto-optimal alignment by interpolating weights fine-tuned on diverse rewards · NeurIPS 2023 |
Machine learning › Reinforcement learning › reinforcement learning from human feedback
reward alignment |
0.7 | 1 | 2023 | Rewarded soups: towards Pareto-optimal alignment by interpolating weights fine-tuned on diverse rewards · NeurIPS 2023 |
Machine learning › Trustworthy machine learning › uncertainty estimation
selective classification |
0.7 | 1 | 2023 | Improving Selective Visual Question Answering by Learning from Your Peers · CVPR 2023 |
Machine learning › Trustworthy machine learning
uncertainty and abstention |
0.7 | 1 | 2023 | Improving Selective Visual Question Answering by Learning from Your Peers · CVPR 2023 |
Machine learning › Efficient and distributed learning › model composition
weight interpolation |
0.7 | 1 | 2023 | Rewarded soups: towards Pareto-optimal alignment by interpolating weights fine-tuned on diverse rewards · NeurIPS 2023 |
Machine learning › Transfer learning and domain adaptation
domain invariance |
0.6 | 1 | 2022 | Fishr: Invariant Gradient Variances for Out-of-Distribution Generalization · ICML 2022 |
Machine learning › Trustworthy machine learning › robustness
shortcut learning |
0.5 | 1 | 2021 | Beyond Question-Based Biases: Assessing Multimodal Shortcut Learning in Visual Question Answering · ICCV 2021 |
Machine learning › Trustworthy machine learning › fairness › social bias
social bias in vision-language models |
0.5 | 1 | 2021 | Beyond Question-Based Biases: Assessing Multimodal Shortcut Learning in Visual Question Answering · ICCV 2021 |
Computer vision › Vision and language › vision-language model
multimodal large language model |
0.2 | 1 | 2024 | Beyond task performance: evaluating and reducing the flaws of large multimodal models with in-context-learning · ICLR 2024 |
Computer vision › Vision and language
image captioning |
0.2 | 1 | 2023 | eP-ALM: Efficient Perceptual Augmentation of Language Models · ICCV 2023 |
Machine learning › Trustworthy machine learning
robustness |
0.1 | 1 | 2019 | RUBi: Reducing Unimodal Biases for Visual Question Answering · NeurIPS 2019 |
Methods — techniques the papers use, named apart from their topics
instruction tuning · 0.8in-context learning · 0.8RLHF · 0.8weight interpolation · 0.7parameter-efficient fine-tuning · 0.7multimodal selection function · 0.7multi-policy training · 0.7linear projection · 0.7learning from peers · 0.7fisher information · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | RadSAM: Segmenting 3D Radiological Images with a 2D Promptable Model
Julien Khlaut, Elodie Ferreres, Daniel Tordjman, Hélène Philippe, Tom Boeken, Pierre Manceron, Corentin Dancette |
MICCAI (4) | 7 |
| 2024 | Beyond task performance: evaluating and reducing the flaws of large multimodal models with in-context-learningabstractFollowing the success of Large Language Models (LLMs), Large Multimodal Models (LMMs), such as the Flamingo model and its subsequent competitors, have started to emerge as natural steps towards generalist agents. However, interacting with recent LMMs reveals major limitations that are hardly captured by the current evaluation benchmarks. Indeed, task performances (e.g., VQA accuracy) alone do not provide enough clues to understand their real capabilities, limitations, and to which extent such models are aligned to human expectations. To refine our understanding of those flaws, we deviate from the current evaluation paradigm, and (1) evaluate 10 recent open-source LMMs from 3B up to 80B parameter scale, on 5 different axes; hallucinations, abstention, compositionality, explainability and instruction following. Our evaluation on these axes reveals major flaws in LMMs. While the current go-to solution to align these models is based on training, such as instruction tuning or RLHF, we rather (2) explore the training-free in-context learning (ICL) as a solution, and study how it affects these limitations. Based on our ICL study, (3) we push ICL further and propose new multimodal ICL variants such as; Multitask-ICL, Chain-of-Hindsight-ICL, and Self-Correcting-ICL. Our findings are as follows; (1) Despite their success, LMMs have flaws that remain unsolved with scaling alone. (2) The effect of ICL on LMMs flaws is nuanced; despite its effectiveness for improved explainability, answer abstention, ICL only slightly improves instruction following, does not improve compositional abilities, and actually even amplifies hallucinations. (3) The proposed ICL variants are promising as post-hoc approaches to efficiently tackle some of those flaws. The code is available here: https://github.com/mshukor/EvALign-ICL. Mustafa Shukor, Alexandre Ramé, Corentin Dancette, Matthieu Cord |
ICLR | 3 |
| 2023 | Improving Selective Visual Question Answering by Learning from Your PeersabstractDespite advances in Visual Question Answering (VQA), the ability of models to assess their own correctness remains under-explored. Recent work has shown that VQA models, out-of-the-box, can have difficulties abstaining from answering when they are wrong. The option to abstain, also called Selective Prediction, is highly relevant when deploying systems to users who must trust the system's output (e.g., VQA assistants for users with visual impairments). For such scenarios, abstention can be especially important as users may provide out-of-distribution (OOD) or adversarial inputs that make incorrect answers more likely. In this work, we explore Selective VQA in both in-distribution (ID) and OOD scenarios, where models are presented with mixtures of ID and OOD data. The goal is to maximize the number of questions answered while minimizing the risk of error on those questions. We propose a simple yet effective Learning from Your Peers (LYP) approach for training multimodal selection functions for making abstention decisions. Our approach uses predictions from models trained on distinct subsets of the training data as targets for optimizing a Selective VQA model. It does not require additional manual labels or held-out data and provides a signal for identifying examples that are easy/difficult to generalize to. In our extensive evaluations, we show this benefits a number of models across different architectures and scales. Overall, for ID, we reach 32.92% in the selective prediction metric coverage at 1 % risk of error$(\mathcal{C} {@} 1\%)$which doubles the previous best coverage of 15.79% on this task. For mixed ID/OOD, using models' softmax confidences for abstention decisions performs very poorly, answering$\mathcal{C}$@1%. Corentin Dancette, Spencer Whitehead, Rishabh Maheshwary, Ramakrishna Vedantam, Stefan Scherer, Xinlei Chen, Matthieu Cord, Marcus Rohrbach |
CVPR | 1 |
| 2023 | eP-ALM: Efficient Perceptual Augmentation of Language ModelsabstractLarge Language Models (LLMs) have so far impressed the world, with unprecedented capabilities that emerge in models at large scales. On the vision side, transformer models (i.e., ViT) are following the same trend, achieving the best performance on challenging benchmarks. With the abundance of such unimodal models, a natural question arises; do we need also to follow this trend to tackle multimodal tasks? In this work, we propose to rather direct effort to efficient adaptations of existing models, and propose to augment Language Models with perception. Existing approaches for adapting pretrained models for vision-language tasks still rely on several key components that hinder their efficiency. In particular, they still train a large number of parameters, rely on large multimodal pretraining, use encoders (e.g., CLIP) trained on huge image-text datasets, and add significant inference overhead. In addition, most of these approaches have focused on Zero-Shot and In Context Learning, with little to no effort on direct finetuning. We investigate the minimal computational effort needed to adapt unimodal models for multimodal tasks and propose a new challenging setup, alongside different approaches, that efficiently adapts unimodal pretrained models. We show that by freezing more than 99% of total parameters, training only one linear projection layer, and prepending only one trainable token, our approach (dubbed eP-ALM) significantly outperforms other baselines on VQA and Captioning across Image, Video, and Audio modalities, following the proposed setup. The code is available here: https://github.com/mshukor/eP-ALM. Mustafa Shukor, Corentin Dancette, Matthieu Cord |
ICCV | 2 |
| 2023 | Rewarded soups: towards Pareto-optimal alignment by interpolating weights fine-tuned on diverse rewardsabstractFoundation models are first pre-trained on vast unsupervised datasets and then fine-tuned on labeled data. Reinforcement learning, notably from human feedback (RLHF), can further align the network with the intended usage. Yet the imperfections in the proxy reward may hinder the training and lead to suboptimal results; the diversity of objectives in real-world tasks and human opinions exacerbate the issue. This paper proposes embracing the heterogeneity of diverse rewards by following a multi-policy strategy. Rather than focusing on a single a priori reward, we aim for Pareto-optimal generalization across the entire space of preferences. To this end, we propose rewarded soup, first specializing multiple networks independently (one for each proxy reward) and then interpolating their weights linearly. This succeeds empirically because we show that the weights remain linearly connected when fine-tuned on diverse rewards from a shared pre-trained initialization. We demonstrate the effectiveness of our approach for text-to-text (summarization, Q&A, helpful assistant, review), text-image (image captioning, text-to-image generation, visual grounding), and control (locomotion) tasks. We hope to enhance the alignment of deep models, and how they interact with the world in all its diversity. Alexandre Ramé, Guillaume Couairon, Corentin Dancette, Jean-Baptiste Gaya, Mustafa Shukor, Laure Soulier, Matthieu Cord |
NeurIPS | 3 |
| 2022 | Fishr: Invariant Gradient Variances for Out-of-Distribution GeneralizationabstractLearning robust models that generalize well under changes in the data distribution is critical for real-world applications. To this end, there has been a growing surge of interest to learn simultaneously from multiple training domains - while enforcing different types of invariance across those domains. Yet, all existing approaches fail to show systematic benefits under controlled evaluation protocols. In this paper, we introduce a new regularization - named Fishr - that enforces domain invariance in the space of the gradients of the loss: specifically, the domain-level variances of gradients are matched across training domains. Our approach is based on the close relations between the gradient covariance, the Fisher Information and the Hessian of the loss: in particular, we show that Fishr eventually aligns the domain-level loss landscapes locally around the final weights. Extensive experiments demonstrate the effectiveness of Fishr for out-of-distribution generalization. Notably, Fishr improves the state of the art on the DomainBed benchmark and performs consistently better than Empirical Risk Minimization. Our code is available at https://github.com/alexrame/fishr. Alexandre Ramé, Corentin Dancette, Matthieu Cord |
ICML | 2 |
| 2021 | Beyond Question-Based Biases: Assessing Multimodal Shortcut Learning in Visual Question AnsweringabstractWe introduce an evaluation methodology for visual question answering (VQA) to better diagnose cases of shortcut learning. These cases happen when a model exploits spurious statistical regularities to produce correct answers but does not actually deploy the desired behavior. There is a need to identify possible shortcuts in a dataset and assess their use before deploying a model in the real world. The research community in VQA has focused exclusively on question-based shortcuts, where a model might, for example, answer "What is the color of the sky" with "blue" by relying mostly on the question-conditional training prior and give little weight to visual evidence. We go a step further and consider multimodal shortcuts that involve both questions and images. We first identify potential shortcuts in the popular VQA v2 training set by mining trivial predictive rules such as co-occurrences of words and visual elements. We then introduce VQA-CounterExamples (VQACE), an evaluation protocol based on our subset of CounterExamples i.e. image-question-answer triplets where our rules lead to incorrect answers. We use this new evaluation in a large-scale study of existing approaches for VQA. We demonstrate that even state-of-the-art models perform poorly and that existing techniques to reduce biases are largely ineffective in this context. Our findings suggest that past work on question-based biases in VQA has only addressed one facet of a complex issue. The code for our method is available at https://github.com/cdancette/detect-shortcuts Corentin Dancette, Rémi Cadène, Damien Teney, Matthieu Cord |
ICCV | 1 |
| 2019 | RUBi: Reducing Unimodal Biases for Visual Question AnsweringabstractVisual Question Answering (VQA) is the task of answering questions about an image. Some VQA models often exploit unimodal biases to provide the correct answer without using the image information. As a result, they suffer from a huge drop in performance when evaluated on data outside their training set distribution. This critical issue makes them unsuitable for real-world settings. We propose RUBi, a new learning strategy to reduce biases in any VQA model. It reduces the importance of the most biased examples, i.e. examples that can be correctly classified without looking at the image. It implicitly forces the VQA model to use the two input modalities instead of relying on statistical regularities between the question and the answer. We leverage a question-only model that captures the language biases by identifying when these unwanted regularities are used. It prevents the base VQA model from learning them by influencing its predictions. This leads to dynamically adjusting the loss in order to compensate for biases. We validate our contributions by surpassing the current state-of-the-art results on VQA-CP v2. This dataset is specifically designed to assess the robustness of VQA models when exposed to different question biases at test time than what was seen during training. Rémi Cadène, Corentin Dancette, Hédi Ben-Younes, Matthieu Cord, Devi Parikh |
NeurIPS | 2 |
| 2018 | Sampling Strategies in Siamese Networks for Unsupervised Speech Representation LearningabstractRecent studies have investigated siamese network architectures for learning invariant speech representations using same-different side information at the word level. Here we investigate systematically an often ignored component of siamese networks: the sampling procedure (how pairs of same vs. different tokens are selected). We show that sampling strategies taking into account Zipf's Law, the distribution of speakers and the proportions of same and different pairs of words significantly impact the performance of the network. In particular, we show that word frequency compression improves learning across a large range of variations in number of training pairs. This effect does not apply to the same extent to the fully unsupervised setting, where the pairs of same-different words are obtained by spoken term discovery. We apply these results to pairs of words discovered using an unsupervised algorithm and show an improvement on state-of-the-art in unsupervised representation learning using siamese networks. Rachid Riad, Corentin Dancette, Julien Karadayi, Neil Zeghidour, Thomas Schatz, Emmanuel Dupoux |
INTERSPEECH | 2 |
| 2018 | A K-Nearest Neighbours Approach To Unsupervised Spoken Term DiscoveryabstractThe following topics are dealt with: speech recognition; neural nets; speech processing; learning (artificial intelligence); natural language processing; recurrent neural nets; speaker recognition; speech synthesis; feature extraction; and text analysis. Alexis Thual, Corentin Dancette, Julien Karadayi, Juan Benjumea, Emmanuel Dupoux |
SLT | 2 |