Yuzhen Du

dblp:363/9844 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
3 papers
Image and video processing · 78% Visual content generation and editing · 22%
Artificial intelligence
3 papers
Generative modeling · 77% Representation and self-supervised learning · 18% Time series and sequential data · 5%

Topics — the 11 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Image and video processing › image restoration
image inpainting
1.722025
Pinco: Position-Induced Consistent Adapter for Diffusion Transformer in Foreground-Conditioned Inpainting · ICCV 2025
ATA: Adaptive Transformation Agent for Text-Guided Subject-Position Variable Background Inpainting · CVPR 2025
Machine learning › Generative modeling
diffusion model
1.522024
LD-BFR: Vector-Quantization-Based Face Restoration Model with Latent Diffusion Enhancement · ACM Multimedia 2024
AnomalyDiffusion: Few-Shot Anomaly Image Generation with Diffusion Model · AAAI 2024
Visual content generation and editing › image editing
text-guided image editing
0.912025
ATA: Adaptive Transformation Agent for Text-Guided Subject-Position Variable Background Inpainting · CVPR 2025
Machine learning › Generative modeling › synthetic data generation
anomaly image generation
0.812024
AnomalyDiffusion: Few-Shot Anomaly Image Generation with Diffusion Model · AAAI 2024
Machine learning › Generative modeling › diffusion model
latent diffusion model
0.812024
LD-BFR: Vector-Quantization-Based Face Restoration Model with Latent Diffusion Enhancement · ACM Multimedia 2024
Machine learning › Representation and self-supervised learning
vector quantization
0.812024
LD-BFR: Vector-Quantization-Based Face Restoration Model with Latent Diffusion Enhancement · ACM Multimedia 2024
Image and video processing › image restoration › face restoration
blind face restoration
0.812024
LD-BFR: Vector-Quantization-Based Face Restoration Model with Latent Diffusion Enhancement · ACM Multimedia 2024
Image and video processing › image restoration
face restoration
0.812024
LD-BFR: Vector-Quantization-Based Face Restoration Model with Latent Diffusion Enhancement · ACM Multimedia 2024
Image and video processing
image restoration
0.812024
LD-BFR: Vector-Quantization-Based Face Restoration Model with Latent Diffusion Enhancement · ACM Multimedia 2024
Machine learning › Generative modeling › cross-modal generation
text-conditioned generation
0.312025
ATA: Adaptive Transformation Agent for Text-Guided Subject-Position Variable Background Inpainting · CVPR 2025
Machine learning › Time series and sequential data
anomaly detection
0.212024
AnomalyDiffusion: Few-Shot Anomaly Image Generation with Diffusion Model · AAAI 2024

Methods — techniques the papers use, named apart from their topics

position embedding · 1.7adaptive transformation · 1.7vector quantization · 1.5latent diffusion · 1.5cross-attention · 1.5diffusion transformer · 0.9adapter · 0.9spatial anomaly embedding · 0.8latent diffusion model · 0.8adaptive attention re-weighting · 0.8
YearPublicationVenuePosition
2025 ATA: Adaptive Transformation Agent for Text-Guided Subject-Position Variable Background Inpainting
abstract
Image inpainting aims to fill the missing region of an image. Recently, there has been a surge of interest in foreground-conditioned background inpainting, a sub-task that fills the background of an image while the foreground subject and associated text prompt are provided. Existing background inpainting methods typically strictly preserve the subject’s original position from the source image, resulting in inconsistencies between the subject and the generated background. To address this challenge, we propose a new task, the "Text-Guided Subject-Position Variable Background Inpainting", which aims to dynamically adjust the subject position to achieve a harmonious relationship between the subject and the inpainted background, and propose the Adaptive Transformation Agent (ATA) for this task. Firstly, we design a PosAgent Block that adaptively predicts an appropriate displacement based on given features to achieve variable subject-position. Secondly, we design the Reverse Displacement Transform (RDT) module, which arranges multiple PosAgent blocks in a reverse structure, to transform hierarchical feature maps from deep to shallow based on semantic information. Thirdly, we equip ATA with a Position Switch Embedding to control whether the subject’s position in the generated image is adaptively predicted or fixed. Extensive comparative experiments validate the effectiveness of our ATA approach, which not only demonstrates superior inpainting capabilities in subject-position variable inpainting, but also ensures good performance on subjectposition fixed inpainting.
Yizhe Tang, Zhimin Sun, Yuzhen Du, Ran Yi 0002, Guangben Lu, Lizhuang Ma, Fangyuan Zou
CVPR3
2025 Pinco: Position-Induced Consistent Adapter for Diffusion Transformer in Foreground-Conditioned Inpainting
abstract
Foreground-conditioned inpainting aims to seamlessly fill the background region of an image by utilizing the provided foreground subject and a text description. While existing T2I-based image inpainting methods can be applied to this task, they suffer from issues of subject shape expansion, distortion, or impaired ability to align with the text description, resulting in inconsistencies between the visual elements and the text description. To address these challenges, we propose Pinco, a plug-and-play foreground-conditioned inpainting adapter that generates high-quality backgrounds with good text alignment while effectively preserving the shape of the foreground subject. Firstly, we design a Self-Consistent Adapter that integrates the foreground subject features into the layout-related self-attention layer, which helps to alleviate conflicts between the text and subject features by ensuring that the model can effectively consider the foreground subject's characteristics while processing the overall image layout. Secondly, we design a Decoupled Image Feature Extraction method that employs distinct architectures to extract semantic and spatial features separately, significantly improving subject feature extraction and ensuring high-quality preservation of the subject's shape. Thirdly, to ensure precise utilization of the extracted features and to focus attention on the subject region, we introduce a Shared Positional Embedding Anchor, greatly improving the model's understanding of subject features and boosting training efficiency. Extensive experiments demonstrate that our method achieves superior performance and efficiency in foreground-conditioned inpainting.
Guangben Lu, Yuzhen Du, Yizhe Tang, Zhimin Sun, Ran Yi 0002, Yifan Qi, Lizhuang Ma, Fangyuan Zou
ICCV2
2024 AnomalyDiffusion: Few-Shot Anomaly Image Generation with Diffusion Model
abstract
Anomaly inspection plays an important role in industrial manufacture. Existing anomaly inspection methods are limited in their performance due to insufficient anomaly data. Although anomaly generation methods have been proposed to augment the anomaly data, they either suffer from poor generation authenticity or inaccurate alignment between the generated anomalies and masks. To address the above problems, we propose AnomalyDiffusion, a novel diffusion-based few-shot anomaly generation model, which utilizes the strong prior information of latent diffusion model learned from large-scale dataset to enhance the generation authenticity under few-shot training data. Firstly, we propose Spatial Anomaly Embedding, which consists of a learnable anomaly embedding and a spatial embedding encoded from an anomaly mask, disentangling the anomaly information into anomaly appearance and location information. Moreover, to improve the alignment between the generated anomalies and the anomaly masks, we introduce a novel Adaptive Attention Re-weighting Mechanism. Based on the disparities between the generated anomaly image and normal sample, it dynamically guides the model to focus more on the areas with less noticeable generated anomalies, enabling generation of accurately-matched anomalous image-mask pairs. Extensive experiments demonstrate that our model significantly outperforms the state-of-the-art methods in generation authenticity and diversity, and effectively improves the performance of downstream anomaly inspection tasks. The code and data are available in https://github.com/sjtuplayer/anomalydiffusion.
Jiangning Zhang, Ran Yi 0002, Yuzhen Du, Xu Chen 0024, Liang Liu 0007, Yabiao Wang, Chengjie Wang 0001
AAAI4
2024 LD-BFR: Vector-Quantization-Based Face Restoration Model with Latent Diffusion Enhancement
abstract
Blind Face Restoration (BFR) aims to restore high-quality face images from low-quality images with unknown degradation. Previous GAN-based or ViT-based methods have shown promising results, but have identity details loss once degradation is severe; while recent diffusion-based methods work on image level and take a lot of time to infer. To restore images in any degradation types with high quality and spend less time compared to the classic diffusion-based method, we propose LD-BFR, a novel BFR framework that integrates both the strengths of vector quantization and latent diffusion. First, we employ a Dual Cross-Attention vector quantization to restore the degraded image in a global manner. Then we utilize the restored high-quality quantized feature as the guidance in our latent diffusion model to generate high-quality restored images with rich details. With the help of the proposed high-quality feature injection module, our LD-BFR effectively injects the high-quality feature as a condition to guide the generation of our latent diffusion model. Extensive experiments demonstrate the superior performance of our model over the SOTA BFR methods. The code is available at: https://github.com/YuzhenD/LD-BFR.git
Yuzhen Du, Ran Yi 0002, Lizhuang Ma
ACM Multimedia1