EDBT 2026 Demo / reviewers in the wild / expert
Siming Fu
dblp:324/0710
· DBLP profile ↗
17ranked-venue papers
6as first author
17since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 3 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 10 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FocusDPO: Dynamic Preference Optimization for Multi-Subject Personalized Image Generation via Adaptive FocusabstractMulti-subject personalized image generation aims to synthesize customized images containing multiple specified subjects without requiring test-time optimization. However, achieving fine-grained independent control over multiple subjects remains challenging due to difficulties in preserving subject fidelity and preventing cross-subject attribute leakage. We present FocusDPO, a framework that adaptively identifies focus regions based on dynamic semantic correspondence and supervision image complexity. During training, our method progressively adjusts these focal areas across noise timesteps, implementing a weighted strategy that rewards information-rich patches while penalizing regions with low prediction confidence. The framework dynamically adjusts focus allocation during the DPO process according to the semantic complexity of reference images and establishes robust correspondence mappings between generated and reference subjects. Extensive experiments demonstrate that our method substantially enhances the performance of existing pre-trained personalized generation models, achieving state-of-the-art results on both single-subject and multi-subject personalized image synthesis benchmarks. Our method effectively mitigates attribute leakage while preserving superior subject fidelity across diverse generation scenarios, advancing the frontier of controllable multi-subject image synthesis. Qiaoqiao Jin, Siming Fu, Dong She, Weinan Jia, Hualiang Wang, Mu Liu, Jidong Jiang |
AAAI | 2 |
| 2025 | MARS: Mixture of Auto-Regressive Models for Fine-grained Text-to-image SynthesisabstractAuto-regressive models have made significant progress in the realm of text-to-image synthesis, yet devising an appropriate model architecture and training strategy to achieve a satisfactory level remains an important avenue of exploration. In this work, we introduce MARS, a novel framework for T2I generation that incorporates a specially designed Semantic Vision-Language Integration Expert (SemVIE). This innovative component integrates pre-trained LLMs by independently processing linguistic and visual information—freezing the textual component while fine-tuning the visual component. This methodology preserves the NLP capabilities of LLMs while imbuing them with exceptional visual understanding. Building upon the powerful base of the pre-trained Qwen-7B, MARS stands out with its bilingual generative capabilities corresponding to both English and Chinese language prompts and the capacity for joint image and text generation. The flexibility of this framework lends itself to migration towards any-to-any task adaptability. Furthermore, MARS employs a multi-stage training strategy that first establishes robust image-text alignment through complementary bidirectional tasks and subsequently concentrates on refining the T2I generation process, significantly augmenting text-image synchrony and the granularity of image details. Notably, MARS requires only 9% of the GPU days needed by SD1.5, yet it achieves remarkable results across a variety of benchmarks, illustrating the training efficiency and the potential for swift deployment in various applications. Wanggui He, Siming Fu, Mushui Liu, Xierui Wang, Wenyi Xiao, Fangxun Shu, Yi Wang 0068, Lei Zhang 0006, Zhelun Yu, Haoyuan Li 0002, Ziwei Huang 0005, Leilei Gan, Hao Jiang 0014 |
AAAI | 2 |
| 2025 | Resolving Multi-Condition Confusion for Finetuning-Free Personalized Image GenerationabstractPersonalized text-to-image generation methods can generate customized images based on the reference images, which have garnered wide research interest. Recent methods propose a finetuning-free approach with a decoupled cross-attention mechanism to generate personalized images requiring no test-time finetuning. However, when multiple reference images are provided, the current decoupled cross-attention mechanism encounters the object confusion problem and fails to map each reference image to its corresponding object, thereby seriously limiting its scope of application. To address the object confusion problem, in this work we investigate the relevance of different positions of the latent image features to the target object in diffusion model, and accordingly propose a weighted-merge method to merge multiple reference image features into the corresponding objects. Next, we integrate this weighted-merge method into existing pre-trained models and continue to train the model on a multi-object dataset constructed from the open-sourced SA-1B dataset. To mitigate object confusion and reduce training costs, we propose an object quality score to estimate the image quality for the selection of high-quality training samples. Furthermore, our weighted-merge training framework can be employed on single-object generation when a single object has multiple reference images. The experiments verify that our method achieves superior performance to the state-of-the-arts on the Concept101 dataset and DreamBooth dataset of multi-object personalized image generation, and remarkably improves the performance on single-object personalized image generation. Qihan Huang, Siming Fu, Hao Jiang 0014, Yipeng Yu, Jie Song 0011 |
AAAI | 2 |
| 2025 | TFCustom: Customized Image Generation with Time-Aware Frequency Feature GuidanceabstractSubject-driven image personalization has seen notable advancements, especially with the ReferenceNet paradigm, which excels in integrating reference image features for creative and commercial applications. However, current ReferenceNet implementations mainly function as latent-level feature extractors, limiting their potential. This restricts the delivery of suitable features to the denoising backbone across timesteps, resulting in suboptimal image consistency. In this paper, we revisit reference feature extraction and propose TFCustom, a framework that focuses on reference image features at different temporal and frequency levels. We introduce synchronized ReferenceNet to extract reference features while optimizing noise injection and denoising. We also propose a time-aware frequency refinement module that uses high- and low-frequency filters with time embeddings to adaptively select reference feature injection. Additionally, we introduce a reward-based loss to improve the similarity between reference objects and generated images. Experimental results show that TFCustom outperforms existing methods in single-object and multi-object reference generation, with significant improvements in textual details. Mushui Liu, Dong She, Jingxuan Pang, Qihan Huang, Jiacheng Ying, Wanggui He, Yuanlei Hou, Siming Fu |
CVPR | 8 |
| 2025 | LLaVA-MoD: Making LLaVA Tiny via MoE-Knowledge DistillationabstractWe introduce LLaVA-MoD, a novel framework designed to enable the efficient training of small-scale Multimodal Language Models ($s$-MLLM) distilling knowledge from large-scale MLLM ($l$-MLLM). Our approach tackles two fundamental challenges in MLLM distillation. First, we optimize the network structure of $s$-MLLM by integrating a sparse Mixture of Experts (MoE) architecture into the language model, striking a balance between computational efficiency and model expressiveness. Second, we propose a progressive knowledge transfer strategy for comprehensive knowledge transfer. This strategy begins with mimic distillation, where we minimize the Kullback-Leibler (KL) divergence between output distributions to enable $s$-MLLM to emulate $s$-MLLM's understanding. Following this, we introduce preference distillation via Preference Optimization (PO), where the key lies in treating $l$-MLLM as the reference model. During this phase, the $s$-MLLM's ability to discriminate between superior and inferior examples is significantly enhanced beyond $l$-MLLM, leading to a better $s$-MLLM that surpasses $l$-MLLM, particularly in hallucination benchmarks.
Extensive experiments demonstrate that LLaVA-MoD surpasses existing works across various benchmarks while maintaining a minimal activated parameters and low computational costs. Remarkably, LLaVA-MoD-2B surpasses Qwen-VL-Chat-7B with an average gain of 8.8\%, using merely $0.3\%$ of the training data and 23\% trainable parameters. The results underscore LLaVA-MoD's ability to effectively distill comprehensive knowledge from its teacher model, paving the way for developing efficient MLLMs. Fangxun Shu, Yue Liao, Lei Zhang 0006, Le Zhuo, Chenning Xu, Long Chan, Zhelun Yu, Wanggui He, Siming Fu, Haoyuan Li 0002, Si Liu 0001, Hongsheng Li 0001, Hao Jiang 0062 |
ICLR | 12 |
| 2025 | MS-Diffusion: Multi-subject Zero-shot Image Personalization with Layout GuidanceabstractRecent advancements in text-to-image generation models have dramatically enhanced the generation of photorealistic images from textual prompts, leading to an increased interest in personalized text-to-image applications, particularly in multi-subject scenarios. However, these advances are hindered by two main challenges: firstly, the need to accurately maintain the details of each referenced subject in accordance with the textual descriptions; and secondly, the difficulty in achieving a cohesive representation of multiple subjects in a single image without introducing inconsistencies. To address these concerns, our research introduces the MS-Diffusion framework for layout-guided zero-shot image personalization with multi-subjects. This innovative approach integrates grounding tokens with the feature resampler to maintain detail fidelity among subjects. With the layout guidance, MS-Diffusion further improves the cross-attention to adapt to the multi-subject inputs, ensuring that each subject condition acts on specific areas. The proposed multi-subject cross-attention orchestrates harmonious inter-subject compositions while preserving the control of texts. Comprehensive quantitative and qualitative experiments affirm that this method surpasses existing models in both image and text fidelity, promoting the development of personalized text-to-image generation. Xierui Wang, Siming Fu, Qihan Huang, Wanggui He, Hao Jiang 0062 |
ICLR | 2 |
| 2025 | LTB-Solver: Long-tailed Bias Solver for image synthesis of diffusion modelsabstractThough diffusion models have shown the merits of generating high-quality visual data while preserving better diversity in recent studies, they do not generalize well on long-tailed datasets due to the minority classes lacking of diversity and semantic information . To overcome the aforementioned challenges, we first take a closer look at the collapse of tail category patterns under long-tail distributed data and propose an alternative but easy-to-use and effective solution, a L ong- T ailed B ias Solver in diffusion model image synthesis ( LTB-Solver ), which thereby enhances the overall diversity and quality of synthetic samples building upon the properties of the long-tailed distribution training data. Especially, we extract rich generative distribution knowledge of ‘head’ categories within proxy model and transfer the head-tail consistency distance to ‘tail’ categories, enabling the target diffusion model to learn diverse generation preserving inter-sample variation during the diffusion training process. Moreover, we incorporate the minority guidance loss function that better aligns training objectives with sampling behaviors and adjust the loss values for different classes by multiplying them with different weights. Extensive experiments are conducted on various datasets and several state-of-the-art diffusion model frameworks to verify the effectiveness of the proposed method. The results show that our method significantly improves the performance of diffusion models on long-tailed datasets by a large margin. Siming Fu, Xiaoxuan He, Haoji Hu |
Neurocomputing | 1 |
| 2025 | SemiGMMPoint: Semi-supervised point cloud segmentation based on Gaussian mixture models
Xianwei Zhuang, Hualiang Wang, Xiaoxuan He, Siming Fu, Haoji Hu |
Pattern Recognit. | 4 |
| 2025 | Beyond Cross-Temporal Difference: Style-Aligned and Fusion-Difference Learning for Change DetectionabstractAccurate change detection in bi-temporal remote sensing images remains challenging due to style differences and multi-scale target differences. Existing methods relying on cross-temporal feature comparison often suffer from style-induced false alarms. Our research proposes a Style-Aligned and Fusion-Difference Network (SAFDNet) which can distinguish style differences from genuine changes. In the encoding stage, we designed a style alignment module based on the idea of histogram equalization to normalize feature distributions of bi-temporal images after feature extraction. When calculating the difference between bi-temporal images, we optimized the existing methods from two perspectives: the object and the scale. For objects, we make full use of layer-exchange fusion mechanism: features are exchanged between temporal streams, fused within each stream, and then compared to their original versions to capture changes while mitigating style interference. For scale, we use a multi-scale Structural Similarity (SSIM) metric replaces conventional distance measures to better quantify structural differences across varying object sizes. With this design, we effectively avoid the noise effect caused by directly using the cross-temporal feature maps to calculate the changes. Our model was tested on five remote sensing datasets: WHUCD, LEVIR-CD, SYSU-CD, CDD and PX-CLCD, and all showed excellent performance. The code and pre-trained model are available at https://github.com/Vgrant0/SAFDNet.git. Siming Fu, Sijun Dong, Xiaoliang Meng |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Video Surveillance on Mobile Edge Networks: Exploiting Multi-Exit NetworkabstractVideo surveillance systems are playing increasingly important roles in our everyday lives. To get meaningful surveillance information in a timely and accurate manner, it is vital to optimally allocate computation and communication resources for image classification tasks. In this paper, taking face recognition as an example, we propose a novel end-to-edge collaborative computing system based on a multi-exit network to dynamically allocate computation at the front end (the camera sensor) and back end (the mobile edge computing server). With the ∊-greedy algorithm for reinforcement learning, the decision module decides whether to obtain recognition results from earlier exits at the front end or transmit the feature maps to the back end to obtain more accurate results. The module balances recognition accuracy and time overhead under different channel conditions. Experimental results show that the proposed system can significantly save inference time and maintain competitive accuracy in various communication channel conditions. Yuchen Cao 0005, Siming Fu, Xiaoxuan He, Haoji Hu, Hangguan Shan, Lu Yu 0003 |
ICC | 2 |
| 2023 | Uniformly Distributed Category Prototype-Guided Vision-Language Framework for Long-Tail RecognitionabstractRecently, large-scale pre-trained vision-language models have presented benefits for alleviating class imbalance in long-tailed recognition. However, the long-tailed data distribution can corrupt the representation space, where the distance between head and tail categories is much larger than the distance between two tail categories. This uneven feature space distribution causes the model to exhibit unclear and inseparable decision boundaries on the uniformly distributed test set, which lowers its performance. To address these challenges, we propose the uniformly category prototype-guided vision-language framework to effectively mitigate feature space bias caused by data imbalance. Especially, we generate a set of category prototypes uniformly distributed on a hypersphere. Category prototype-guided mechanism for image-text matching makes the features of different classes converge to these distinct and uniformly distributed category prototypes, which maintain a uniform distribution in the feature space, and improve class boundaries. Additionally, our proposed irrelevant text filtering and attribute enhancement module allows the model to ignore irrelevant noisy text and focus more on key attribute information, thereby enhancing the robustness of our framework. In the image recognition fine-tuning stage, to address the positive bias problem of the learnable classifier, we design the class feature prototype-guided classifier, which compensates for the performance of tail classes while maintaining the performance of head classes. Our method outperforms previous vision-language methods for long-tailed learning work by a large margin and achieves state-of-the-art performance. Xiaoxuan He, Siming Fu, Xinpeng Ding, Yuchen Cao 0005, Hualiang Wang |
ACM Multimedia | 2 |
| 2023 | Class semantic enhancement network for semantic segmentation
Siming Fu, Hualiang Wang, Haoji Hu, Xiaoxuan He, Yongwen Long, Jianhong Bai, Yangtao Ou, Yuanjia Huang, Mengqiu Zhou |
J. Vis. Commun. Image Represent. | 1 |
| 2023 | AuxBranch: Binarization residual-aware network design via auxiliary branch search
Siming Fu, Huanpeng Chu, Lu Yu 0003, Zheyang Li, Wenming Tan, Haoji Hu |
Pattern Recognit. | 1 |
| 2022 | Renovate Yourself: Calibrating Feature Representation of Misclassified Pixels for Semantic SegmentationabstractExisting image semantic segmentation methods favor learning consistent representations by extracting long-range contextual features with the attention, multi-scale, or graph aggregation strategies. These methods usually treat the misclassified and correctly classified pixels equally, hence misleading the optimization process and causing inconsistent intra-class pixel feature representations in the embedding space during learning. In this paper, we propose the auxiliary representation calibration head (RCH), which consists of the image decoupling, prototype clustering, error calibration modules and a metric loss function, to calibrate these error-prone feature representations for better intra-class consistency and segmentation performance. RCH could be incorporated into the hidden layers, trained together with the segmentation networks, and decoupled in the inference stage without additional parameters. Experimental results show that our method could significantly boost the performance of current segmentation methods on multiple datasets (e.g., we outperform the original HRNet and OCRNet by 1.1% and 0.9% mIoU on the Cityscapes test set). Codes are available at https://github.com/VipaiLab/RCH. Hualiang Wang, Huanpeng Chu, Siming Fu, Zuozhu Liu, Haoji Hu |
AAAI | 3 |
| 2022 | Meta-prototype Decoupled Training for Long-Tailed Learning
Siming Fu, Huanpeng Chu, Xiaoxuan He, Hualiang Wang, Haoji Hu |
ACCV (6) | 1 |
| 2022 | Towards Calibrated Hyper-Sphere Representation via Distribution Overlap Coefficient for Long-Tailed Learning
Hualiang Wang, Siming Fu, Xiaoxuan He, Hangxiang Fang, Zuozhu Liu, Haoji Hu |
ECCV (24) | 2 |
| 2022 | Meta-BNS FOR Adversarial Data-Free QuantizationabstractData-free quantization has recently been a promising method to perform quantization without access to the original data. However, the drawback of such approaches is the homogenization of synthetic data due to low efficiency for diverse data generation and the performance collapse of the generator. To alleviate the above issue, we propose a novel Meta-BNS for adversarial data-free quantization scheme which consists of Meta-BNS module and adversarial exploration module. Meta-BNS module automatically learns an enhancement coefficient matrix function for BN loss module to provide a suitable constrain on the generator. Adversarial exploration module leverages minimax game between the generator and quantized model via input gradient to encourage the generator to learn high-dimensional and complex real data distribution. The experimental results show that our method achieves state-of-the-art performance for various settings on data-free quantization. Siming Fu, Hualiang Wang, Yuchen Cao 0005, Haoji Hu, Wenming Tan, Tingqun Ye |
ICIP | 1 |