Jiafeng Mao

dblp:274/7104 · DBLP profile ↗
← Back
12ranked-venue papers
6as first author
11since 2021 · last 2026
0009-0003-0907-7522ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 6 first-author · 10 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021
YearPublicationVenuePosition
2026 Difficulty Controlled Diffusion Model for Synthesizing Effective Training Data
abstract
Generative models have become a powerful tool for synthesizing training data in computer vision tasks. Current approaches solely focus on aligning generated images with the target dataset distribution. As a result, they capture only the common features in the real dataset and mostly generate "easy samples", which are already well learned by models trained on real data. In contrast, those rare "hard samples", with atypical features but crucial for enhancing performance, cannot be effectively generated. Consequently, these approaches must synthesize large volumes of data to yield appreciable performance gains, yet the improvement remains limited. To overcome this limitation, we present a novel method that can learn to control the learning difficulty of samples during generation while also achieving domain alignment. Thus, it can efficiently generate valuable "hard samples" that yield significant performance improvements for target tasks. This is achieved by incorporating learning difficulty as an additional conditioning signal in generative models, together with a designed encoder structure and training–generation strategy. Experimental results across multiple datasets show that our method can achieve higher performance with lower generation cost. Specifically, we obtain the best performance with only 10% additional synthetic data, saving 63.4 GPU hours of generation time compared to the previous SOTA on ImageNet. Moreover, our method provides insightful visualizations of category-specific hard factors, serving as a tool for analyzing datasets.
Zerun Wang, Jiafeng Mao, Toshihiko Yamasaki
AAAI2
2025 Diversity-Driven Generative Dataset Distillation Based on Diffusion Model with Self-Adaptive Memory
abstract
Dataset distillation enables the training of deep neural networks with comparable performance in significantly reduced time by compressing large datasets into small and representative ones. Although the introduction of generative models has made great achievements in this field, the distributions of their distilled datasets are not diverse enough to represent the original ones, leading to a decrease in downstream validation accuracy. In this paper, we present a diversity-driven generative dataset distillation method based on a diffusion model to solve this problem. We introduce self-adaptive memory to align the distribution between distilled and real datasets, assessing the representativeness. The degree of alignment leads the diffusion model to generate more diverse datasets during the distillation process. Extensive experiments show that our method outperforms existing state-of-the-art methods in most situations, proving its ability to tackle dataset distillation tasks.
Mingzhuo Li, Guang Li 0008, Jiafeng Mao, Takahiro Ogawa 0001, Miki Haseyama
ICIP3
2025 Noisy Label Refinement with Semantically Reliable Synthetic Images
abstract
Semantic noise in image classification datasets, where visually similar categories are frequently mislabeled, poses a significant challenge to conventional supervised learning approaches. In this paper, we explore the potential of using synthetic images generated by advanced text-to-image models to address this issue. Although these high-quality synthetic images come with reliable labels, their direct application in training is limited by domain gaps and diversity constraints. Unlike conventional approaches, we propose a novel method that leverages synthetic images as reliable reference points to identify and correct mislabeled samples in noisy datasets. Extensive experiments across multiple benchmark datasets show that our approach significantly improves classification accuracy under various noise conditions, especially in challenging scenarios with semantic label noise. Additionally, since our method is orthogonal to existing noise-robust learning techniques, when combined with state-of-the-art noise-robust training methods, it achieves superior performance, improving accuracy by 30% on CIFAR-10 and by 11% on CIFAR-100 under 70% semantic noise, and by 24% on ImageNet-100 under real-world noise conditions.
Yingxuan Li, Jiafeng Mao, Yusuke Matsui 0001
ICIP2
2025 Exploring Palette based Color Guidance in Diffusion Models
Qianru Qiu, Jiafeng Mao
ACM Multimedia2
2024 The Lottery Ticket Hypothesis in Denoising: Towards Semantic-Driven Initialization
Jiafeng Mao, Kiyoharu Aizawa
ECCV (74)1
2024 SCOMatch: Alleviating Overtrusting in Open-Set Semi-supervised Learning
Zerun Wang, Liuyu Xiang, Lang Huang 0001, Jiafeng Mao, Ling Xiao 0001, Toshihiko Yamasaki
ECCV (51)4
2024 Dealing with Synthetic Data Contamination in Online Continual Learning
abstract
Image generation has shown remarkable results in generating high-fidelity realistic images, in particular with the advancement of diffusion-based models. However, the prevalence of AI-generated images may have side effects for the machine learning community that are not clearly identified. Meanwhile, the success of deep learning in computer vision is driven by the massive dataset collected on the Internet. The extensive quantity of synthetic data being added to the Internet would become an obstacle for future researchers to collect "clean" datasets without AI-generated content. Prior research has shown that using datasets contaminated by synthetic images may result in performance degradation when used for training. In this paper, we investigate the potential impact of contaminated datasets on Online Continual Learning (CL) research. We experimentally show that contaminated datasets might hinder the training of existing online CL methods. Also, we propose Entropy Selection with Real-synthetic similarity Maximization (ESRM), a method to alleviate the performance deterioration caused by synthetic images when training online CL models. Experiments show that our method can significantly alleviate performance deterioration, especially when the contamination is severe. For reproducibility, the source code of our work is available at https://github.com/maorong-wang/ESRM.
Maorong Wang, Nicolas Michel, Jiafeng Mao, Toshihiko Yamasaki
NeurIPS3
2023 Training-Free Location-Aware Text-to-Image Synthesis
abstract
Current large-scale generative models have impressive efficiency in generating high-quality images based on text prompts. However, they lack the ability to precisely control the size and position of objects in the generated image. In this study1, we analyze the generative mechanism of the stable diffusion model and propose a new interactive generation paradigm that allows users to specify the position of generated objects without additional training. Moreover, we propose an object detection-based evaluation metric to assess the control capability of location aware generation task. Our experimental results show that our method outperforms state-of-the-art methods on both control capacity and image quality.
Jiafeng Mao
ICIP1
2023 Noise-Avoidance Sampling for Annotation Missing Object Detection
abstract
Excellent results can be achieved using object detection with fully supervised training on large well-annotated datasets. However, the problem of missing annotations in real-world datasets can considerably reduce the performance of object detectors. In this study, we thoroughly analyze the effect of missing annotations on both positive and negative samples in object detector training. To mitigate the negative impact caused by annotation missing problem, we propose a simple yet effective method, noise-avoidance sampling, to distinguish noisy training samples and subsequently reduce their negative impact. Experiments are conducted on the PASCAL VOC 07+12 dataset with varying levels of missing annotations. The results reveal that the proposed method achieves comparable or superior performance with state-of-the-art methods.
Jiafeng Mao, Qing Yu 0013, Go Irie, Kiyoharu Aizawa
ICIP1
2023 Guided Image Synthesis via Initial Image Editing in Diffusion Model
abstract
Diffusion models have the ability to generate high quality images by denoising pure Gaussian noise images. While previous research has primarily focused on improving the control of image generation through adjusting the denoising process, we propose a novel direction of manipulating the initial noise to control the generated image. Through experiments on stable diffusion, we show that blocks of pixels in the initial latent images have a preference for generating specific content, and that modifying these blocks can significantly influence the generated image. In particular, we show that modifying a part of the initial image affects the corresponding region of the generated image while leaving other regions unaffected, which is useful for repainting tasks. Furthermore, we find that the generation preferences of pixel blocks are primarily determined by their values, rather than their position. By moving pixel blocks with a tendency to generate user-desired content to user-specified regions, our approach achieves state-of-the-art performance in layout-to-image generation. Our results highlight the flexibility and power of initial image manipulation in controlling the generated image.
Jiafeng Mao, Kiyoharu Aizawa
ACM Multimedia1
2021 Noisy Annotation Refinement for Object Detection
Jiafeng Mao, Qing Yu 0013, Yoko Yamakata, Kiyoharu Aizawa
BMVC1
2020 Noisy Localization Annotation Refinement For Object Detection
abstract
The production of finely annotated datasets for object detection tasks is labor-intensive, therefore, cloud sourcing is often used to create datasets, which leads to these datasets tending to contain incorrect annotations such as inaccurate localization bounding boxes. In this study, we highlight a problem of object detection with noisy bounding box annotations and show that these noisy annotations are harmful to the performance of deep neural networks. To solve this problem, we further propose a framework to allow the network to modify the noisy datasets by alternating refinement. The experimental results demonstrate that our proposed framework can significantly alleviate the influences of noise on model performance.
Jiafeng Mao, Qing Yu 0013, Kiyoharu Aizawa
ICIP1