Marlène Careil

dblp:301/9344 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2024
0000-0003-4542-5371ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Generative modeling · 78% Transfer learning and domain adaptation · 17% Segmentation and scene understanding · 5%
Computer graphics and multimedia
2 papers
Image and video coding · 54% Visual content generation and editing · 46%

Topics — the 11 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
0.812024
Towards image compression with perfect realism at ultra-low bitrates · ICLR 2024
Image and video coding
image compression
0.812024
Towards image compression with perfect realism at ultra-low bitrates · ICLR 2024
Image and video coding › image compression
perceptual image compression
0.812024
Towards image compression with perfect realism at ultra-low bitrates · ICLR 2024
Machine learning › Transfer learning and domain adaptation › few-shot learning
few-shot transfer
0.712023
Few-shot Semantic Image Synthesis with Class Affinity Transfer · CVPR 2023
Machine learning › Generative modeling
image generation
0.712023
Few-shot Semantic Image Synthesis with Class Affinity Transfer · CVPR 2023
Machine learning › Generative modeling › image generation › conditional image synthesis
semantic image synthesis
0.712023
Few-shot Semantic Image Synthesis with Class Affinity Transfer · CVPR 2023
Visual content generation and editing
image generation
0.712023
Zero-shot spatial layout conditioning for text-to-image diffusion models · ICCV 2023
Visual content generation and editing › image generation › text-to-image generation
text-to-image diffusion
0.712023
Zero-shot spatial layout conditioning for text-to-image diffusion models · ICCV 2023
Machine learning › Generative modeling › diffusion model
conditional generation
0.512021
Instance-Conditioned GAN · NeurIPS 2021
Machine learning › Generative modeling
generative adversarial network
0.512021
Instance-Conditioned GAN · NeurIPS 2021
Computer vision › Segmentation and scene understanding
semantic segmentation
0.212023
Few-shot Semantic Image Synthesis with Class Affinity Transfer · CVPR 2023

Methods — techniques the papers use, named apart from their topics

vector quantization · 1.5adversarial loss · 1.5diffusion model · 1.3textual label embeddings · 0.7self-supervised vision features · 0.7segmentation guidance · 0.7cross-attention · 0.7class affinity matrix · 0.7GAN · 0.7nearest neighbor · 0.5kernel density estimation · 0.5
YearPublicationVenuePosition
2024 Towards image compression with perfect realism at ultra-low bitrates
abstract
Image codecs are typically optimized to trade-off bitrate vs. distortion metrics. At low bitrates, this leads to compression artefacts which are easily perceptible, even when training with perceptual or adversarial losses. To improve image quality and remove dependency on the bitrate we propose to decode with iterative diffusion models. We condition the decoding process on a vector-quantized image representation, as well as a global image description to provide additional context. We dub our model `PerCo'' for ``perceptual compression'', and compare it to state-of-the-art codecs at rates from 0.1 down to 0.003 bits per pixel. The latter rate is more than an order of magnitude smaller than those considered in most prior work, compressing a 512x768 Kodak image with less than 153 bytes. Despite this ultra-low bitrate, our approach maintains the ability to reconstruct realistic images. We find that our model leads to reconstructions with state-of-the-art visual quality as measured by FID and KID. As predicted by rate-distortion-perception theory, visual quality is less dependent on the bitrate than previous methods.
Marlène Careil, Matthew J. Muckley, Jakob Verbeek, Stéphane Lathuilière
ICLR1
2023 Few-shot Semantic Image Synthesis with Class Affinity Transfer
abstract
Semantic image synthesis aims to generate photo realistic images given a semantic segmentation map. Despite much recent progress, training them still requires large datasets of images annotated with per-pixel label maps that are extremely tedious to obtain. To alleviate the high annotation cost, we propose a transfer method that leverages a model trained on a large source dataset to improve the learning ability on small target datasets via estimated pairwise relations between source and target classes. The class affinity matrix is introduced as a first layer to the source model to make it compatible with the target label maps, and the source model is then further finetuned for the target domain. To estimate the class affinities we consider different approaches to leverage prior knowledge: semantic segmentation on the source domain, textual label embeddings, and self-supervised vision features. We apply our approach to GAN-based and diffusion-based architectures for semantic synthesis. Our experiments show that the different ways to estimate class affinity can be effectively combined, and that our approach significantly improves over existing state-of-the-art transfer approaches for generative image models.
Marlène Careil, Jakob Verbeek, Stéphane Lathuilière
CVPR1
2023 Zero-shot spatial layout conditioning for text-to-image diffusion models
abstract
Large-scale text-to-image diffusion models have significantly improved the state of the art in generative image modeling and allow for an intuitive and powerful user interface to drive the image generation process. Expressing spatial constraints, e.g. to position specific objects in particular locations, is cumbersome using text; and current text-based image generation models are not able to accurately follow such instructions. In this paper we consider image generation from text associated with segments on the image canvas, which combines an intuitive natural language interface with precise spatial control over the generated content. We propose ZestGuide, a "zero-shot" segmentation guidance approach that can be plugged into pre-trained text-to-image diffusion models, and does not require any additional training. It leverages implicit segmentation maps that can be extracted from cross-attention layers, and uses them to align the generation with input masks. Our experimental results combine high image quality with accurate alignment of generated content with input segmentations, and improve over prior work both quantitatively and qualitatively, including methods that require training on images with corresponding segmentations. Compared to Paint with Words, the previous state-of-the art in image generation with zero-shot segmentation conditioning, we improve by 5 to 10 mIoU points on the COCO dataset with similar FID scores.
Guillaume Couairon, Marlène Careil, Matthieu Cord, Stéphane Lathuilière, Jakob Verbeek
ICCV2
2021 Instance-Conditioned GAN
abstract
Generative Adversarial Networks (GANs) can generate near photo realistic images in narrow domains such as human faces. Yet, modeling complex distributions of datasets such as ImageNet and COCO-Stuff remains challenging in unconditional settings. In this paper, we take inspiration from kernel density estimation techniques and introduce a non-parametric approach to modeling distributions of complex datasets. We partition the data manifold into a mixture of overlapping neighborhoods described by a datapoint and its nearest neighbors, and introduce a model, called instance-conditioned GAN (IC-GAN), which learns the distribution around each datapoint. Experimental results on ImageNet and COCO-Stuff show that IC-GAN significantly improves over unconditional models and unsupervised data partitioning baselines. Moreover, we show that IC-GAN can effortlessly transfer to datasets not seen during training by simply changing the conditioning instances, and still generate realistic images. Finally, we extend IC-GAN to the class-conditional case and show semantically controllable generation and competitive quantitative results on ImageNet; while improving over BigGAN on ImageNet-LT. Code and trained models to reproduce the reported results are available at https://github.com/facebookresearch/ic_gan.
Arantxa Casanova, Marlène Careil, Jakob Verbeek, Michal Drozdzal, Adriana Romero-Soriano
NeurIPS2