Yohann Benchetrit

dblp:42/9992 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Theory of computation · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Generative modeling · 52% Deep learning architectures and training · 33% 3D vision · 15%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling › diffusion model
latent diffusion model
1.622025
Boosting Latent Diffusion with Perceptual Objectives · ICLR 2025
On improved Conditioning Mechanisms and Pre-training Strategies for Diffusion Models · NeurIPS 2024
Machine learning › Deep learning architectures and training › loss function design
perceptual loss
0.912025
Boosting Latent Diffusion with Perceptual Objectives · ICLR 2025
Machine learning › Generative modeling
diffusion model
0.812024
On improved Conditioning Mechanisms and Pre-training Strategies for Diffusion Models · NeurIPS 2024
Machine learning › Deep learning architectures and training
pre-training strategy
0.812024
On improved Conditioning Mechanisms and Pre-training Strategies for Diffusion Models · NeurIPS 2024
Bioinformatics and computational biology › neuroscience › neuroinformatics
brain decoding
0.812024
Brain decoding: toward real-time reconstruction of visual perception · ICLR 2024
Machine learning › Generative modeling
image generation
0.212024
Brain decoding: toward real-time reconstruction of visual perception · ICLR 2024

Methods — techniques the papers use, named apart from their topics

regression · 1.5pretrained image embeddings · 1.5magnetoencephalography · 1.5contrastive learning · 1.5latent perceptual loss · 0.9flow matching · 0.9DDPM · 0.9text-to-image generation · 0.8latent diffusion model · 0.8
YearPublicationVenuePosition
2025 Boosting Latent Diffusion with Perceptual Objectives
abstract
Latent diffusion models (LDMs) power state-of-the-art high-resolution generative image models. LDMs learn the data distribution in the latent space of an autoencoder (AE) and produce images by mapping the generated latents into RGB image space using the AE decoder. While this approach allows for efficient model training and sampling, it induces a disconnect between the training of the diffusion model and the decoder, resulting in a loss of detail in the generated images. To remediate this disconnect, we propose to leverage the internal features of the decoder to define a latent perceptual loss (LPL). This loss encourages the models to create sharper and more realistic images. Our loss can be seamlessly integrated with common autoencoders used in latent diffusion models, and can be applied to different generative modeling paradigms such as DDPM with epsilon and velocity prediction, as well as flow matching. Extensive experiments with models trained on three datasets at 256 and 512 resolution show improved quantitative -- with boosts between 6% and 20% in FID -- and qualitative results when using our perceptual loss.
Tariq Berrada, Pietro Astolfi, Melissa Hall, Marton Havasi, Yohann Benchetrit, Adriana Romero-Soriano, Karteek Alahari, Michal Drozdzal, Jakob Verbeek
ICLR5
2024 Brain decoding: toward real-time reconstruction of visual perception
abstract
In the past five years, the use of generative and foundational AI systems has greatly improved the decoding of brain activity. Visual perception, in particular, can now be decoded from functional Magnetic Resonance Imaging (fMRI) with remarkable fidelity. This neuroimaging technique, however, suffers from a limited temporal resolution ($\approx$0.5\,Hz) and thus fundamentally constrains its real-time usage. Here, we propose an alternative approach based on magnetoencephalography (MEG), a neuroimaging device capable of measuring brain activity with high temporal resolution ($\approx$5,000 Hz). For this, we develop an MEG decoding model trained with both contrastive and regression objectives and consisting of three modules: i) pretrained embeddings obtained from the image, ii) an MEG module trained end-to-end and iii) a pretrained image generator. Our results are threefold: Firstly, our MEG decoder shows a 7X improvement of image-retrieval over classic linear decoders. Second, late brain responses to images are best decoded with DINOv2, a recent foundational image model. Third, image retrievals and generations both suggest that high-level visual features can be decoded from MEG signals, although the same approach applied to 7T fMRI also recovers better low-level features. Overall, these results, while preliminary, provide an important step towards the decoding - in real-time - of the visual processes continuously unfolding within the human brain.
Yohann Benchetrit, Hubert J. Banville, Jean-Rémi King
ICLR1
2024 On improved Conditioning Mechanisms and Pre-training Strategies for Diffusion Models
abstract
Large-scale training of latent diffusion models (LDMs) has enabled unprecedented quality in image generation. However, large-scale end-to-end training of these models is computationally costly, and hence most research focuses either on finetuning pretrained models or experiments at smaller scales. In this work we aim to improve the training efficiency and performance of LDMs with the goal of scaling to larger datasets and higher resolutions. We focus our study on two points that are critical for good performance and efficient training: (i) the mechanisms used for semantic level (\eg a text prompt, or class name) and low-level (crop size, random flip, \etc) conditioning of the model, and (ii) pre-training strategies to transfer representations learned on smaller and lower-resolution datasets to larger ones. The main contributions of our work are the following: we present systematic experimental study of these points, we propose a novel conditioning mechanism that disentangles semantic and low-level conditioning, we obtain state-of-the-art performance on CC12M for text-to-image at 512 resolution.
Tariq Berrada, Pietro Astolfi, Melissa Hall, Reyhane Askari Hemmat, Yohann Benchetrit, Marton Havasi, Matthew J. Muckley, Karteek Alahari, Adriana Romero-Soriano, Jakob Verbeek, Michal Drozdzal
NeurIPS5
2011 An Excluded Minor Characterization of Seymour Graphs
Alexander A. Ageev, Yohann Benchetrit, András Sebö, Zoltán Szigeti
IPCO2