EDBT 2026 Demo / reviewers in the wild / expert
Mengping Yang
dblp:324/0385
· DBLP profile ↗
19ranked-venue papers
9as first author
19since 2021 · last 2025
0000-0003-1503-9621ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 7 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 8 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | FreeCus: Free Lunch Subject-Driven Customization in Diffusion TransformersabstractIn light of recent breakthroughs in text-to-image (T2I) generation, particularly with diffusion transformers (DiT), subject-driven technologies are increasingly being employed for high-fidelity customized production that preserves subject identity from reference inputs, enabling thrilling design workflows and engaging entertainment. Existing alternatives typically require either per-subject optimization via trainable text embeddings or training specialized encoders for subject feature extraction on large-scale datasets. Such dependencies on training procedures fundamentally constrain their practical applications. More importantly, current methodologies fail to fully leverage the inherent zero-shot potential of modern diffusion transformers (e.g., the Flux series) for authentic subject-driven synthesis. To bridge this gap, we propose FreeCus, a genuinely training-free framework that activates DiT's capabilities through three key innovations: 1) We introduce a pivotal attention sharing mechanism that captures the subject's layout integrity while preserving crucial editing flexibility. 2) Through a straightforward analysis of DiT's dynamic shifting, we propose an upgraded variant that significantly improves fine-grained feature extraction. 3) We further integrate advanced Multimodal Large Language Models (MLLMs) to enrich cross-modal semantic representations. Extensive experiments reflect that our method successfully unlocks DiT's zero-shot ability for consistent subject synthesis across diverse contexts, achieving state-of-the-art or comparable results compared to approaches that require additional training. Notably, our framework demonstrates seamless compatibility with existing inpainting pipelines and control modules, facilitating more compelling experiences. Our code is available at: https://github.com/Monalissaa/FreeCus. Yanbing Zhang, Mengping Yang |
ICCV | 4 |
| 2025 | Image Synthesis Under Limited Data: A Survey and Taxonomy
Mengping Yang, Zhe Wang 0002 |
Int. J. Comput. Vis. | 1 |
| 2025 | Transductive Parameter-Free Propagation Framework for Few-Shot Distribution RectificationabstractFew-shot learning (FSL) is challenging due to the scarce labeled novel-class data. Researchers have to train the embedding function with auxiliary base-class data to obtain the novel-class embeddings. However, the domain gap makes the novel-class embedding unsatisfactory, as the novel class and the base class are disjoint. Recent studies prove that embedding rectification shows great potential, introduces miscellaneous variants, and achieves similar performances. Nonetheless, while each method demonstrates unique strengths, they often address distinct challenges in isolation, limiting their applicability in more complex or diverse scenarios. In this article, we take a closer look at these methods and hypothesize that a general embedding rectification framework is more essential to the model's performance. To verify our observation, we propose: 1) a distribution propagation (DisP) layer distinguishes the inter-class margin and increases intra-class aggregation, performing the task-level rectification; and 2) a prototype propagation (ProtoP) layer moves the prototype toward the ideal class center, applying the prototype-query level rectification. Our framework aims to maximize the actual data distribution. Although pseudo-labeling proves effective in achieving this goal, a significant challenge is ensuring the reliable retention of only high-confidence predictions. To overcome this, we introduce a distribution-based pseudo-labeling method pseudo-query upgrade (PseQUp) that provides more reliable pseudo-labeling samples without relying on confidence scores. We evaluate the proposed method in both transfer learning and meta-learning scenarios. Empirical experiments show the applicable and plug-and-play ability of the proposed methods. Heng Tian, Ziqiu Chi, Zhe Wang 0002, Wei Guo 0023, Mengping Yang, Xinlei Xu |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Attention Calibration for Disentangled Text-to-Image PersonalizationabstractRecent thrilling progress in large-scale text-to-image (T2I) models has unlocked unprecedented synthesis quality of AI-generated content (AIGC) including image generation, 3D and video composition. Further, personalized techniques enable appealing customized production of a novel concept given only several images as reference. However, an intriguing problem persists: Is it possible to capture multiple, novel concepts from one single reference image? In this paper, we identify that existing approaches fail to preserve visual consistency with the reference image and eliminate cross-influence from concepts. To alleviate this, we propose an attention calibration mechanism to improve the concept-level understanding of the T2I model. Specifically, we first introduce new learnable modifiers bound with classes to capture attributes of multiple concepts. Then, the classes are separated and strengthened following the activation of the cross-attention operation, ensuring comprehensive and self-contained concepts. Additionally, we suppress the attention activation of different classes to mitigate mutual influence among concepts. Together, our proposed method, dubbed DisenDiff, can learn disentangled multiple concepts from one single image and produce novel customized images with learned concepts. We demonstrate that our method outperforms the current state of the art in both qualitative and quantitative evaluations. More importantly, our proposed techniques are compatible with LoRA and inpainting pipelines, enabling more interactive experiences. Yanbing Zhang, Mengping Yang, Qin Zhou 0002, Zhe Wang 0002 |
CVPR | 2 |
| 2024 | An Empirical Study and Analysis of Text-to-Image Generation Using Large Language Model-Powered Textual Representation
Zhiyu Tan, Mengping Yang, Luozheng Qin, Ye Qian, Qiang Zhou 0001, Cheng Zhang 0014, Hao Li 0030 |
ECCV (80) | 2 |
| 2024 | Freezing partial source representations matters for image inpainting under limited data
Yanbing Zhang, Mengping Yang, Ting Xiao 0002, Zhe Wang 0002, Ziqiu Chi |
Eng. Appl. Artif. Intell. | 2 |
| 2024 | Adaptive weighted dictionary representation using anchor graph for subspace clustering
Wenyi Feng, Zhe Wang 0002, Ting Xiao 0002, Mengping Yang |
Pattern Recognit. | 4 |
| 2023 | Semantic-Aware Generator and Low-level Feature Augmentation for Few-shot Image GenerationabstractFew-shot image generation aims to generate novel images for an unseen category with only a few samples. Prior studies fail to produce novel images with desirable diversity and fidelity. To ameliorate the generation quality, we in this paper propose a Semantic-Aware Generator (SAG) to provide explicit semantic guidance to the discriminator, and a Low-level Feature Augmentation (LFA) technique to provide fine-grained information, facilitating the diversity. Specifically, we observe that the generator feature layers contain different levels of semantic information. Such observation motivates us to employ intermediate feature maps of the generator as semantic labels to guide the discriminator, improving the semantic awareness of the generator. Moreover, spatially informative and diverse features obtained via LFA contribute to better generation quality. Together with the aforementioned module, we conduct extensive experiments on three representative benchmarks and the results demonstrate the effectiveness and advancement of our method. Zhe Wang 0002, Jiaoyan Guan, Mengping Yang, Ting Xiao 0002, Ziqiu Chi |
ACM Multimedia | 3 |
| 2023 | Improving Few-shot Image Generation by Structural Discrimination and Textural ModulationabstractFew-shot image generation, which aims to produce plausible and diverse images for one category given a few images from this category, has drawn extensive attention. Existing approaches either globally interpolate different images or fuse local representations with pre-defined coefficients. However, such an intuitive combination of images/features only exploits the most relevant information for generation, leading to poor diversity and coarse-grained semantic fusion. To remedy this, this paper proposes a novel textural modulation (TexMod) mechanism to inject external semantic signals into internal local representations. Parameterized by the feedback from the discriminator, our TexMod enables more fined-grained semantic injection while maintaining the synthesis fidelity. Moreover, a global structural discriminator (StructD) is developed to explicitly guide the model to generate images with reasonable layout and outline. Furthermore, the frequency awareness of the model is reinforced by encouraging the model to distinguish frequency signals. Together with these techniques, we build a novel and effective model for few-shot image generation. The effectiveness of our model is identified by extensive experiments on three popular datasets and various settings. Besides achieving state-of-the-art synthesis performance on these datasets, our proposed techniques could be seamlessly integrated into existing models for a further performance boost. Our code and models are available at \hrefhttps://github.com/kobeshegu/SDTM-GAN-ACMMM-2023 here. Mengping Yang, Zhe Wang 0002, Wenyi Feng, Qian Zhang 0068, Ting Xiao 0002 |
ACM Multimedia | 1 |
| 2023 | Revisiting the Evaluation of Image Synthesis with GANsabstractA good metric, which promises a reliable comparison between solutions, is essential for any well-defined task. Unlike most vision tasks that have per-sample ground-truth, image synthesis tasks target generating unseen data and hence are usually evaluated through a distributional distance between one set of real samples and another set of generated samples. This study presents an empirical investigation into the evaluation of synthesis performance, with generative adversarial networks (GANs) as a representative of generative models. In particular, we make in-depth analyses of various factors, including how to represent a data point in the representation space, how to calculate a fair distance using selected samples, and how many instances to use from each set. Extensive experiments conducted on multiple datasets and settings reveal several important findings. Firstly, a group of models that include both CNN-based and ViT-based architectures serve as reliable and robust feature extractors for measurement evaluation. Secondly, Centered Kernel Alignment (CKA) provides a better comparison across various extractors and hierarchical layers in one model. Finally, CKA is more sample-efficient and enjoys better agreement with human judgment in characterizing the similarity between two internal data correlations. These findings contribute to the development of a new measurement system, which enables a consistent and reliable re-evaluation of current state-of-the-art generative models. Mengping Yang, Ceyuan Yang, Yichi Zhang 0013, Qingyan Bai, Yujun Shen, Bo Dai 0002 |
NeurIPS | 1 |
| 2023 | Adaptive federated few-shot feature learning with prototype rectification
Mengping Yang, Yonghui Xi, Saisai Niu, Zhe Wang 0002 |
Eng. Appl. Artif. Intell. | 1 |
| 2023 | DFSGAN: Introducing editable and representative attributes for few-shot image generation
Mengping Yang, Saisai Niu, Zhe Wang 0002, Dongdong Li 0003, Wenli Du |
Eng. Appl. Artif. Intell. | 1 |
| 2023 | ProtoGAN: Towards high diversity and fidelity image synthesis under limited data
Mengping Yang, Zhe Wang 0002, Ziqiu Chi, Wenli Du |
Inf. Sci. | 1 |
| 2023 | Semantic alignment with self-supervision for class incremental learning
Zhiling Fu, Zhe Wang 0002, Xinlei Xu, Mengping Yang, Ziqiu Chi, Weichao Ding |
Knowl. Based Syst. | 4 |
| 2022 | WaveGAN: Frequency-Aware GAN for High-Fidelity Few-Shot Image Generation
Mengping Yang, Zhe Wang 0002, Ziqiu Chi, Wenyi Feng |
ECCV (15) | 1 |
| 2022 | Better Embedding and More Shots for Few-shot LearningabstractIn few-shot learning, methods are enslaved to the scarce labeled data, resulting in suboptimal embedding. Recent studies learn the embedding network by other large-scale labeled data. However, the trained network may give rise to the distorted embedding of target data. We argue two respects are required for an unprecedented and promising solution. We call them Better Embedding and More Shots (BEMS). Suppose we propose to extract embedding from the embedding network. BE maximizes the extraction of general representation and prevents over-fitting information. For this purpose, we introduce the topological relation for global reconstruction, avoiding excessive memorizing. MS maximizes the relevance between the reconstructed embedding and the target class space. In this respect, increasing the number of shots is a pivotal but intractable strategy. As a creative method, we derive the bound of information-theory-based loss function and implicitly achieve infinite shots with negligible cost. A substantial experimental analysis is carried out to demonstrate the state-of-the-art performance. Compared to the baseline, our method improves by up to 10%+. We also prove that BEMS is suitable for both standard pre-trained and meta-learning embedded networks. Ziqiu Chi, Zhe Wang 0002, Mengping Yang, Wei Guo 0023, Xinlei Xu |
IJCAI | 3 |
| 2022 | FreGAN: Exploiting Frequency Components for Training GANs under Limited DataabstractTraining GANs under limited data often leads to discriminator overfitting and memorization issues, causing divergent training. Existing approaches mitigate the overfitting by employing data augmentations, model regularization, or attention mechanisms. However, they ignore the frequency bias of GANs and take poor consideration towards frequency information, especially high-frequency signals that contain rich details. To fully utilize the frequency information of limited data, this paper proposes FreGAN, which raises the model's frequency awareness and draws more attention to synthesising high-frequency signals, facilitating high-quality generation. In addition to exploiting both real and generated images' frequency information, we also involve the frequency signals of real images as a self-supervised constraint, which alleviates the GAN disequilibrium and encourages the generator to synthesis adequate rather than arbitrary frequency signals. Extensive results demonstrate the superiority and effectiveness of our FreGAN in ameliorating generation quality in the low-data regime (especially when training data is less than 100). Besides, FreGAN can be seamlessly applied to existing regularization and attention mechanism models to further boost the performance. Mengping Yang, Zhe Wang 0002, Ziqiu Chi, Yanbing Zhang |
NeurIPS | 1 |
| 2022 | Gravitation balanced multiple kernel learning for imbalanced classification
Mengping Yang, Zhe Wang 0002, Yanqiong Li, Yangming Zhou, Dongdong Li 0003, Wenli Du |
Neural Comput. Appl. | 1 |
| 2022 | Learning to Capture the Query Distribution for Few-Shot LearningabstractIn the Few-Shot Learning (FSL), much of the related efforts only rely on the few available labeled samples (support set) building approach. However, the challenge is that the support set is easy-to-be-biased, so that they cannot be competent prototypes and are hard to represent the class distribution, leading to performance bottlenecks. In this paper, we propose to solve this obstacle by capturing the distribution of the unlabeled samples (query set). We propose two sampling methods: DeepSearch ($\cal DS$) and WideSearch ($\cal WS$). Both approaches are simple to implement and have no trainable parameters. They search the query samples near to the support set in different manners. Afterward, the statistic information is calculated, and we generate the latent samples according to it. The generated latent set is promising. First, it brings the query set distribution information to the classifier, which significantly improves the performance of the cross-entropy-based classifier. Second, it helps the support set become the better prototypes, which boosts the performance of the prototype-based classifier. Third, we find few latent samples are enough to boost the performance. Abundant experiments prove the proposed method achieves state-of-the-art performance on the few-shot tasks. Finally, rich ablation studies explain the compelling details of our approach. Ziqiu Chi, Zhe Wang 0002, Mengping Yang, Dongdong Li 0003, Wenli Du |
IEEE Trans. Circuits Syst. Video Technol. | 3 |