Qianli Feng

dblp:198/0940 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
9since 2021 · last 2025
0000-0002-7550-2019ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021
YearPublicationVenuePosition
2025 Be More Specific: Evaluating Object-centric Realism in Synthetic Images
abstract
Evaluation of synthetic images is important for both model development and selection. An ideal evaluation should be specific, accurate and aligned with human perception. This paper addresses the problem of evaluating realism of objects in synthetic images. Although methods have been proposed to evaluate holistic realism, there are no methods tailored towards object-centric realism evaluation. In this work, we define a new standard for assessing object-centric realism that follows a shape-texture breakdown and proposes the first object-centric realism evaluation dataset for synthetic images. The dataset contains images generated from state-of-the-art image generative models and is richly annotated at object level across a diverse set of object categories. We then design and train the OLIP model, an architecture that considerably outperforms any existing baseline on object-centric realism evaluation.
Anqi Liang, Ciprian A. Corneanu, Qianli Feng, Giorgio Giannone, Aleix Martinez
CVPR3
2025 Structured Human Assessment of Text-to-Image Generative Models
abstract
Following the great progress in text-conditioned image generation there is a dire need for establishing clear com-parison benchmarks. Unfortunately, assessing performance of such models is highly subjective and notoriously difficult. Current automatic assessment of generated images quality and their alignment to text are approximate at best while human assessment is subjective, poorly calibrated and not very well defined. To address these concerns, we propose GenomeBench, a new framework for assessing quality of text-to-image generative models. It consists of a prompt dataset richly annotated with semantic components based on a formalized grounding of language and images. On top of it, we define a procedure to collect human assessment through a carefully guided question answering process. Fi-nally, these assessments are summarized into a novel score built around quality and alignment to text. We show the proposal achieves higher inter-annotator agreement with respect to the baseline human assessment and better cor-relation between quality and alignment compared to automatic assessment. Finally, we use this framework to dissect the performance of recent text-to-image models, providing insights on strength and weakness of each.
Ciprian A. Corneanu, Qianli Feng, Aleix Martinez
WACV2
2025 DreamBlend: Advancing Personalized Fine-Tuning of Text-to-Image Diffusion Models
abstract
Given a small number of images of a subject, personalized image generation techniques can fine-tune large pretrained text-to-image diffusion models to generate images of the subject in novel contexts, conditioned on text prompts. In doing so, a tradeoff is made between prompt fidelity, subject fidelity and diversity. As the pretrained model is fine-tuned, earlier checkpoints synthesize images with low subject fidelity but high prompt fidelity and diversity. In contrast, later checkpoints generate images with low prompt fidelity and diversity but high subject fidelity. This inherent tradeoff limits the prompt fidelity, subject fidelity and diversity of generated images. In this work, we propose DreamBlend to combine the prompt fidelity from earlier checkpoints and the subject fidelity from later checkpoints during inference. We perform a cross attention guided image synthesis from a later checkpoint, guided by an image
Shwetha Ram, Tal Neiman, Qianli Feng, Son Tran, Trishul Chilimbi
WACV3
2025 Why does Knowledge Distillation work? Rethink its attention and fidelity mechanism
Chenqi Guo, Shiwei Zhong, Qianli Feng, Yinglong Ma 0001
Expert Syst. Appl.4
2025 Recommendation feedback-based dynamic adaptive training for efficient social item recommendation
Chenqi Guo, Yinglong Ma 0001, Qianli Feng
Expert Syst. Appl.4
2023 Network-Free, Unsupervised Semantic Segmentation with Synthetic Images
abstract
We derive a method that yields highly accurate semantic segmentation maps without the use of any additional neural network, layers, manually annotated training data, or supervised training. Our method is based on the observation that the correlation of a set of pixels belonging to the same semantic segment do not change when generating synthetic variants of an image using the style mixing approach in GANs. We show how we can use GAN inversion to accurately semantically segment synthetic and real photos as well as generate large training image-semantic segmentation mask pairs for downstream tasks.
Qianli Feng, Raghudeep Gadde, Wentong Liao, Eduard Ramon, Aleix Martinez
CVPR1
2021 When do GANs replicate? On the choice of dataset size
abstract
Do GANs replicate training images? Previous studies have shown that GANs do not seem to replicate training data without significant change in the training procedure. This leads to a series of research on the exact condition needed for GANs to overfit to the training data. Although a number of factors has been theoretically or empirically identified, the effect of dataset size and complexity on GANs replication is still unknown. With empirical evidence from BigGAN and StyleGAN2, on datasets CelebA, Flower and LSUN-bedroom, we show that dataset size and its complexity play an important role in GANs replication and perceptual quality of the generated images. We further quantify this relationship, discovering that replication percentage decays exponentially with respect to dataset size and complexity, with a shared decaying factor across GAN-dataset combinations. Meanwhile, the perceptual image quality follows a U-shape trend w.r.t dataset size. This finding leads to a practical tool for one-shot estimation on minimal dataset size to prevent GAN replication which can be used to guide datasets construction and selection.
Qianli Feng, Chenqi Guo, Fabian Benitez-Quiroz, Aleix Martinez
ICCV1
2021 Detail Me More: Improving GAN's photo-realism of complex scenes
abstract
Generative models can synthesize photo-realistic images of a single object. For example, for human faces, algorithms learn to model the local shape and shading of the face components, i.e., changes in the brows, eyes, nose, mouth, jaw line, etc. This is possible because all faces have two brows, two eyes, a nose and a mouth, approximately in the same location. The modeling of complex scenes is however much more challenging because the scene components and their location vary from image to image. For example, living rooms contain a varying number of products belonging to many possible categories and locations, e.g., a lamp may or may not be present in an endless number of possible locations. In the present work, we propose to add a "broker" module in Generative Adversarial Networks (GAN) to solve this problem. The broker is tasked to mediate the use of multiple discriminators in the appropriate image locales. For example, if a lamp is detected or wanted in a specific area of the scene, the broker assigns a fine-grained lamp discriminator to that image patch. This allows the generator to learn the shape and shading models of the lamp. The resulting multi-fine-grained optimization problem is able to synthesize complex scenes with almost the same level of photo-realism as single object images. We demonstrate the generability of the proposed approach on several GAN algorithms (BigGAN, ProGAN, StyleGAN, StyleGAN2), image resolutions (2562to 10242), and datasets. Our approach yields significant improvements over state-of-the-art GAN algorithms.
Raghudeep Gadde, Qianli Feng, Aleix Martinez
ICCV2
2021 Adding Knowledge to Unsupervised Algorithms for the Recognition of Intent
Stuart Synakowski, Qianli Feng, Aleix Martinez
Int. J. Comput. Vis.2