Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Benoît Gérin

dblp:376/2632 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2025
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Transfer learning and domain adaptation · 46% Vision and language · 23% Learning paradigms · 23%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Transfer learning and domain adaptation
few-shot learning
0.812024
Boosting Vision-Language Models with Transduction · NeurIPS 2024
Machine learning › Transfer learning and domain adaptation › few-shot learning
transductive few-shot learning
0.812024
Boosting Vision-Language Models with Transduction · NeurIPS 2024
Machine learning › Learning paradigms › semi-supervised learning
transductive learning
0.812024
Boosting Vision-Language Models with Transduction · NeurIPS 2024
Computer vision › Vision and language
vision-language model
0.812024
Boosting Vision-Language Models with Transduction · NeurIPS 2024
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › parameter estimation
maximum likelihood estimation
0.212024
Boosting Vision-Language Models with Transduction · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

block majorize-minimize · 0.8KL-divergence penalty · 0.8
YearPublicationVenuePosition
2025 Enhancing Remote Sensing Vision-Language Models for Zero-Shot Scene Classification
abstract
peer reviewed
Karim El Khoury, Maxime Zanella, Benoît Gérin, Tiffanie Godelaine, Benoît Macq, Saïd Mahmoudi, Christophe De Vleeschouwer, Ismail Ben Ayed
ICASSP3
2024 Boosting Vision-Language Models with Transduction
abstract
Transduction is a powerful paradigm that leverages the structure of unlabeled data to boost predictive accuracy. We present TransCLIP, a novel and computationally efficient transductive approach designed for Vision-Language Models (VLMs). TransCLIP is applicable as a plug-and-play module on top of popular inductive zero- and few-shot models, consistently improving their performances. Our new objective function can be viewed as a regularized maximum-likelihood estimation, constrained by a KL divergence penalty that integrates the text-encoder knowledge and guides the transductive learning process. We further derive an iterative Block Majorize-Minimize (BMM) procedure for optimizing our objective, with guaranteed convergence and decoupled sample-assignment updates, yielding computationally efficient transduction for large-scale datasets. We report comprehensive evaluations, comparisons, and ablation studies that demonstrate: (i) Transduction can greatly enhance the generalization capabilities of inductive pretrained zero- and few-shot VLMs; (ii) TransCLIP substantially outperforms standard transductive few-shot learning methods relying solely on vision features, notably due to the KL-based language constraint.
Maxime Zanella, Benoît Gérin, Ismail Ben Ayed
NeurIPS2