EDBT 2026 Demo / reviewers in the wild / expert
Mingfu Yan
dblp:361/6443
· DBLP profile ↗
11ranked-venue papers
1as first author
11since 2021 · last 2026
0009-0000-4820-0704ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-author · 10 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TARA: Token-Aware LoRA for Composable Personalization in Diffusion ModelsabstractPersonalized text-to-image generation aims to synthesize novel images of a specific subject or style using only a few reference images. Recent methods based on Low-Rank Adaptation (LoRA) enable efficient single-concept customization by injecting lightweight, concept-specific adapters into pre-trained diffusion models. However, combining multiple LoRA modules for multi-concept generation often leads to identity missing and visual feature leakage. In this work, we identify two key issues behind these failures: (1) token-wise interference among different LoRA modules, and (2) spatial misalignment between the attention map of a rare token and its corresponding concept-specific region. To address these issues, we propose Token-Aware LoRA (TARA), which introduces a token mask to explicitly constrain each module to focus on its associated rare token to avoid interference, and a training objective that encourages the spatial attention of a rare token to align with its concept region. Our method enables training-free multi-concept composition by directly injecting multiple independently trained TARA modules at inference time. Experimental results demonstrate that TARA enables efficient multi-concept inference and effectively preserving the visual identity of each concept by avoiding mutual interference between LoRA modules. Yuqi Peng, Lingtao Zheng, Yi Huang 0035, Mingfu Yan, Jianzhuang Liu, Shifeng Chen |
AAAI | 5 |
| 2025 | Efficient Document Shadow Removal with Contrast-Aware Guidance
Yifan Liu 0001, Jiyu Wu, Jiancheng Huang, Mingfu Yan, Yi Huang 0035, Shifeng Chen |
CGI (3) | 5 |
| 2025 | DIVE: Taming DINO for Subject-Driven Video EditingabstractBuilding on the success of diffusion models in image generation and editing, video editing has recently gained substantial attention. However, maintaining temporal consistency and motion alignment still remains challenging. To address these issues, this paper proposes DINO-guided Video Editing (DIVE), a framework designed to facilitate subject-driven editing in source videos conditioned on either target text prompts or reference images with specific identities. The core of DIVE lies in leveraging the powerful semantic features extracted from a pretrained DINOv2 model as implicit correspondences to guide the editing process. Specifically, to ensure temporal motion consistency, DIVE employs DINO features to align with the motion trajectory of the source video. For precise subject editing, DIVE incorporates the DINO features of reference images into a pretrained text-to-image model to learn Low-Rank Adaptations (LoRAs), effectively registering the target subject's identity. Extensive experiments on diverse real-world videos demonstrate that our framework can achieve high-quality editing results with robust motion consistency, highlighting the potential of DINO to contribute to video editing. Project page: https://dino-video-editing.github.io Yi Huang 0035, Wei Xiong 0008, He Zhang 0004, Chaoqi Chen, Jianzhuang Liu, Mingfu Yan, Shifeng Chen |
ICCV | 6 |
| 2025 | Component Adaptive Clustering for Generalized Category DiscoveryabstractGeneralized Category Discovery (GCD) tackles the challenging problem of categorizing unlabeled images into both known and novel classes within a partially labeled dataset, without prior knowledge of the number of unknown categories. Traditional methods often rely on rigid assumptions, such as predefining the number of classes, which limits their ability to handle the inherent variability and complexity of real-world data. To address these shortcomings, we propose AdaGCD, a cluster-centric contrastive learning framework that incorporates Adaptive Slot Attention (AdaSlot) into the GCD framework. AdaSlot dynamically determines the optimal number of slots based on data complexity, removing the need for predefined slot counts. This adaptive mechanism facilitates the flexible clustering of unlabeled data into known and novel categories by dynamically allocating representational capacity. By integrating adaptive representation with dynamic slot allocation, our method captures both instance-specific and spatially clustered features, improving class discovery in open-world scenarios. Extensive experiments on public and fine-grained datasets validate the effectiveness of our framework, emphasizing the advantages of leveraging spatial local information for category discovery in unlabeled image datasets. Mingfu Yan, Jiancheng Huang, Yifan Liu 0001, Shifeng Chen |
ICME | 1 |
| 2025 | Diffusion Model-Based Image Editing: A SurveyabstractDenoising diffusion models have emerged as a powerful tool for various image generation and editing tasks, facilitating the synthesis of visual content in an unconditional or input-conditional manner. The core idea behind them is learning to reverse the process of gradually adding noise to images, allowing them to generate high-quality samples from a complex distribution. In this survey, we provide an exhaustive overview of existing methods using diffusion models for image editing, covering both theoretical and practical aspects in the field. We delve into a thorough analysis and categorization of these works from multiple perspectives, including learning strategies, user-input conditions, and the array of specific editing tasks that can be accomplished. In addition, we pay special attention to image inpainting and outpainting, and explore both earlier traditional context-driven and current multimodal conditional methods, offering a comprehensive analysis of their methodologies. To further evaluate the performance of text-guided image editing algorithms, we propose a systematic benchmark, EditEval, featuring an innovative metric, LMM Score. Finally, we address current limitations and envision some potential directions for future research. Yi Huang 0035, Jiancheng Huang, Yifan Liu 0001, Mingfu Yan, Jiaxi Lv, Jianzhuang Liu, Wei Xiong 0008, He Zhang 0004, Liangliang Cao, Shifeng Chen |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | DCD-Net: Weakly supervised decomposition learning for real-world image dehazing
Yi Huang 0035, Jiancheng Huang, Mingfu Yan, Shifeng Chen |
Signal Process. | 4 |
| 2024 | MambaDW: Semantic-Aware Mamba for Document Watermark Removal
Yifan Liu 0001, Mingfu Yan, He Hua, Jiancheng Huang, Shifeng Chen |
CGI (1) | 2 |
| 2024 | Color-SD: Stable Diffusion Model Already has a Color Style Noisy Latent SpaceabstractWe present Color-SD, a comprehensive color style transfer framework that utilizes either image or text references. Built on the pretrained Stable Diffusion Model, Color-SD exploits an existing color style space, enabling a training-free and tuning-free zero-shot color style transfer method without introducing new parameters. For image references, we first invert the source and reference images to the noisy latent space, followed by parallel sampling. During this process, we execute distribution transformation in the noisy latent space, effectively completing the color style transfer and generating the stylized result. For text references, we capitalize on the Stable Diffusion model’s inherent text-to-image capability. We only invert the source image to the noisy latent, and the given text reference prompt is utilized during the parallel sampling. This approach eliminates the need for training or tuning, yet produces impressive open-set transfer results. Comprehensive experiments validate the effectiveness of our method, demonstrating significant superiority over existing methods in both qualitative and quantitative evaluations. Jiancheng Huang, Mingfu Yan, Shifeng Chen |
ICME | 2 |
| 2024 | SBCR: Stochasticity Beats Content Restriction Problem in Training and Tuning Free Image EditingabstractText-conditional image editing is a practical AIGC task that has recently emerged with great commercial and academic value. For real image editing, most diffusion model-based methods use DDIM Inversion as a first stage before editing. However, DDIM Inversion often results in reconstruction failure, leading to unsatisfactory performance for downstream editing. Many inversion-based works modify the formula to address this problem but this leads to another content restriction problem. To solve the content restriction problem, we first analyze why the reconstruction via DDIM Inversion fails and then propose Reconstruction-and-Generation Balancing Noises (R&G-B noises) that can achieve superior reconstruction and editing performance with the following advantages: 1) It can perfectly reconstruct real images without fine-tuning. 2) It can overcome the content restriction problem and generate diverse content. Jiancheng Huang, Mingfu Yan, Shifeng Chen |
ICMR | 2 |
| 2024 | MagicFight: Personalized Martial Arts Combat Video GenerationabstractAmid the surge in generic text-to-video generation, the field of personalized human video generation has witnessed notable advancements, primarily concentrated on single-person scenarios. However, to our knowledge, the domain of two-person interactions, particularly in the context of martial arts combat, remains uncharted. We identify a significant gap: existing models for single-person dancing generation prove insufficient for capturing the subtleties and complexities of two engaged fighters, resulting in challenges such as identity confusion, anomalous limbs, and action mismatches. To address this, we introduce a pioneering new task, Personalized Martial Arts Combat Video Generation. Our approach, MagicFight, is specifically crafted to overcome these hurdles. Given this pioneering task, we face a lack of appropriate datasets. Thus, we generate a bespoke dataset using the game physics engine Unity, meticulously crafting a multitude of 3D characters, martial arts moves, and scenes designed to represent the diversity of combat. MagicFight refines and adapts existing models and strategies to generate high-fidelity two-person combat videos that maintain individual identities and ensure seamless, coherent action sequences, thereby laying the groundwork for future innovations in the realm of interactive video content creation. Jiancheng Huang, Mingfu Yan, Songyan Chen, Yi Huang 0035, Shifeng Chen |
ACM Multimedia | 2 |
| 2024 | WaveDM: Wavelet-Based Diffusion Models for Image RestorationabstractLatest diffusion-based methods for many image restoration tasks outperform traditional models, but they encounter the long-time inference problem. To tackle it, this paper proposes a Wavelet-Based Diffusion Model (WaveDM). WaveDM learns the distribution of clean images in the wavelet domain conditioned on the wavelet spectrum of degraded images after wavelet transform, which is more time-saving in each step of sampling than modeling in the spatial domain. To ensure restoration performance, a unique training strategy is proposed where the low-frequency and high-frequency spectrums are learned using distinct modules. In addition, an Efficient Conditional Sampling (ECS) strategy is developed from experiments, which reduces the number of total sampling steps to around 5. Evaluations on twelve benchmark datasets including image raindrop removal, rain steaks removal, dehazing, defocus deblurring, demoiréing, and denoising demonstrate that WaveDM achieves state-of-the-art performance with the efficiency that is comparable to traditional one-pass methods and over 100× faster than existing image restoration methods using vanilla diffusion models. The code is available athttps://github.com/stayalive16/WaveDM Yi Huang 0035, Jiancheng Huang, Jianzhuang Liu, Mingfu Yan, Jiaxi Lv, Chaoqi Chen, Shifeng Chen |
IEEE Trans. Multim. | 4 |