Mushui Liu

dblp:334/2912 · DBLP profile ↗
← Back
17ranked-venue papers
8as first author
17since 2021 · last 2026
0000-0002-2909-7702ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 8 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 11 since 2021
YearPublicationVenuePosition
2026 FUSE: Fine-Grained and Semantic-Aware Learning for Unified Image Understanding and Generation
Wanggui He, Mushui Liu, Wenyi Xiao, Siyu Zou, Yanpeng Liu, Weilong Dai, Shuyi Ying, Ruikai Zhou, Yubo Tao, Hao Jiang 0062
AAAI3
2026 CMMCoT: Enhancing Complex Multi-Image Comprehension via Multi-Modal Chain-of-Thought and Memory Augmentation
abstract
While previous multimodal slow-thinking methods have demonstrated remarkable success in single-image understanding scenarios, their effectiveness becomes fundamentally constrained when extended to more complex multi-image comprehension tasks. This limitation stems from their predominant reliance on text-based intermediate reasoning processes. While for human, when engaging in sophisticated multi-image analysis, they typically perform two complementary cognitive operations: (1) continuous cross-image visual comparison through region-of-interest matching, and (2) dynamic memorization of critical visual concepts throughout the reasoning chain. Motivated by these observations, we propose the Complex Multi-Modal Chain-of-Thought (CMMCoT) framework, a multi-step reasoning framework that mimics human-like "slow thinking" for multi-image understanding. Our approach incorporates two key innovations: (1) The construction of interleaved multimodal multi-step reasoning chains, which utilize critical visual region tokens, extracted from intermediate reasoning steps, as supervisory signals. This mechanism not only facilitates comprehensive cross-modal understanding but also enhances model interpretability. (2) The introduction of a test-time memory augmentation module that expands the model’s reasoning capacity during inference while preserving parameter efficiency. Furthermore, to facilitate research in this direction, we have curated a novel multi-image slow-thinking dataset. Extensive experiments demonstrate the effectiveness of our model.
Yan Xia 0006, Mushui Liu, Zhelun Yu, Haoyuan Li 0002, Wanggui He, Dong She, Yi Wang 0068, Hao Jiang 0014
AAAI4
2025 MARS: Mixture of Auto-Regressive Models for Fine-grained Text-to-image Synthesis
abstract
Auto-regressive models have made significant progress in the realm of text-to-image synthesis, yet devising an appropriate model architecture and training strategy to achieve a satisfactory level remains an important avenue of exploration. In this work, we introduce MARS, a novel framework for T2I generation that incorporates a specially designed Semantic Vision-Language Integration Expert (SemVIE). This innovative component integrates pre-trained LLMs by independently processing linguistic and visual information—freezing the textual component while fine-tuning the visual component. This methodology preserves the NLP capabilities of LLMs while imbuing them with exceptional visual understanding. Building upon the powerful base of the pre-trained Qwen-7B, MARS stands out with its bilingual generative capabilities corresponding to both English and Chinese language prompts and the capacity for joint image and text generation. The flexibility of this framework lends itself to migration towards any-to-any task adaptability. Furthermore, MARS employs a multi-stage training strategy that first establishes robust image-text alignment through complementary bidirectional tasks and subsequently concentrates on refining the T2I generation process, significantly augmenting text-image synchrony and the granularity of image details. Notably, MARS requires only 9% of the GPU days needed by SD1.5, yet it achieves remarkable results across a variety of benchmarks, illustrating the training efficiency and the potential for swift deployment in various applications.
Wanggui He, Siming Fu, Mushui Liu, Xierui Wang, Wenyi Xiao, Fangxun Shu, Yi Wang 0068, Lei Zhang 0006, Zhelun Yu, Haoyuan Li 0002, Ziwei Huang 0005, Leilei Gan, Hao Jiang 0014
AAAI3
2025 Frame Order Matters: A Temporal Sequence-Aware Model for Few-Shot Action Recognition
abstract
In this paper, we propose a novel Temporal Sequence-Aware-Model (TSAM) for few-shot action recognition (FSAR), which incorporates a sequential perceiver adapter into the pre-training framework, to integrate both the spatial information and the sequential temporal dynamics into the feature embeddings. Different from the existing fine-tuning approaches that capture temporal information by exploring the relationships among all the frames, our perceiver-based adapter recurrently captures the sequential dynamics alongside the timeline, which could perceive the frame order change. To obtain the discriminative representations for each class, we extend a textual corpus for each class derived from the large language models (LLMs) and enrich the visual prototypes by integrating the contextual semantic information. Besides, We introduce an unbalanced optimal transport strategy for feature matching that mitigates the impact of class-unrelated features, thereby facilitating more effective decision-making. Experimental results on five FSAR datasets demonstrate that our method establishes a new benchmark, outperforming the second-best competitors.
Bozheng Li, Mushui Liu, Gaoang Wang
AAAI2
2025 LLM4GEN: Leveraging Semantic Representation of LLMs for Text-to-Image Generation
abstract
Diffusion models have exhibited substantial success in text-to-image generation. However, they often encounter challenges when dealing with complex and dense prompts involving multiple objects, attribute binding, and long descriptions. In this paper, we propose a novel framework called LLM4GEN, which enhances the semantic understanding of text-to-image diffusion models by leveraging the representation of Large Language Models (LLMs). It can be seamlessly incorporated into various diffusion models as a plug-and-play component. A specially designed Cross-Adapter Module (CAM) integrates the original text features of text-to-image models with LLM features, thereby enhancing text-to-image generation. Additionally, to facilitate and correct entity-attribute relationships in text prompts, we develop an entity-guided regularization loss to further improve generation performance. We also introduce DensePrompts, which contains 7,000 dense prompts to provide a comprehensive evaluation for the text-to-image generation task. Experiments indicate that LLM4GEN significantly improves the semantic alignment of SD1.5 and SDXL, demonstrating increases of 9.69% and 12.90% in color on T2I-CompBench, respectively. Moreover, it surpasses existing models in terms of sample quality, image-text alignment, and human evaluation.
Mushui Liu, Jun Dan, Zeng Zhao, Zhipeng Hu, Bai Liu 0002, Changjie Fan
AAAI1
2025 Envisioning Class Entity Reasoning by Large Language Models for Few-shot Learning
abstract
Few-shot learning (FSL) aims to recognize new concepts using a limited number of visual samples. Existing methods attempt to incorporate semantic information into the limited visual data for category understanding. However, these methods often enrich class-level feature representations with abstract category names, failing to capture nuanced features essential for effective generalization. To address this issue, we propose a novel framework for FSL, which incorporates both the abstract class semantics and the concrete class entities extracted from Large Language Models (LLMs), to enhance the representation of the class prototypes. Specifically, our framework composes a Semantic-guided Visual Pattern Extraction (SVPE) module and a Prototype-Calibration (PC) module, where the SVPE meticulously extracts semantic-aware visual patterns across diverse scales, while the PC module seamlessly integrates these patterns to refine the visual prototype, enhancing its representativeness. Extensive experiments on four few-shot classification benchmarks and the BSCD-FSL cross-domain benchmark showcase remarkable advancements over the current state-of-the-art methods. Notably, for the challenging one-shot setting, our approach, utilizing the ResNet-12 backbone, achieves an impressive average improvement of 1.95% over the second-best competitor.
Mushui Liu, Fangtai Wu, Bozheng Li, Ziqian Lu
AAAI1
2025 TFCustom: Customized Image Generation with Time-Aware Frequency Feature Guidance
abstract
Subject-driven image personalization has seen notable advancements, especially with the ReferenceNet paradigm, which excels in integrating reference image features for creative and commercial applications. However, current ReferenceNet implementations mainly function as latent-level feature extractors, limiting their potential. This restricts the delivery of suitable features to the denoising backbone across timesteps, resulting in suboptimal image consistency. In this paper, we revisit reference feature extraction and propose TFCustom, a framework that focuses on reference image features at different temporal and frequency levels. We introduce synchronized ReferenceNet to extract reference features while optimizing noise injection and denoising. We also propose a time-aware frequency refinement module that uses high- and low-frequency filters with time embeddings to adaptively select reference feature injection. Additionally, we introduce a reward-based loss to improve the similarity between reference objects and generated images. Experimental results show that TFCustom outperforms existing methods in single-object and multi-object reference generation, with significant improvements in textual details.
Mushui Liu, Dong She, Jingxuan Pang, Qihan Huang, Jiacheng Ying, Wanggui He, Yuanlei Hou, Siming Fu
CVPR1
2025 Hybrid mask generation for infrared small target detection with single-point supervision
Mushui Liu, Yunlong Yu 0001
Neurocomputing2
2025 Fully fine-tuned CLIP models are efficient few-shot learners
Mushui Liu, Bozheng Li, Jun Dan, Ziqian Lu
Knowl. Based Syst.1
2025 Synth-CLIP: Synthetic data make CLIP generalize better in data-limited scenarios
Mushui Liu, Ziqian Lu, Jun Dan, Yunlong Yu 0001, Yingming Li, Xi Li 0001, Jungong Han
Neural Networks1
2025 Variational Adapter: Improving CLIP in Data-Imbalanced Scenarios
abstract
In this paper, we propose the Prompt-based Variational Adapter (PVA), a novel approach designed to fine-tune the pre-trained Vision-Language Models (VLMs) in data-imbalanced scenarios. Unlike existing methods that focus primarily on pairwise alignment of visual-text relationships during fine-tuning, PVA relaxes pairwise explicit constrains and emphasizes the harmonization of visual and text modality distributions, enhancing generalization and cross-modal understanding. To realize this harmonization, we develop two variational adapters, which are appended separately to the visual and text encoders. These adapters transform the feature embeddings into latent spaces that implicitly align with the corresponding modality distributions. We then adopt a divide-and-conquer strategy, dividing classes into data-abundant and data-limited sets to reduce prediction bias. Within each set, we independently fine-tune the models by incorporating both the model’s original general knowledge and specialized knowledge gained from training samples. Extensive experiments across two data-imbalanced scenarios validate the superiority of our approach, establishing a new state-of-the-art on popular benchmarks.
Ziqian Lu, Mushui Liu, Yunlong Yu 0001, Xi Li 0001, Jungong Han
IEEE Trans. Circuits Syst. Video Technol.2
2024 OmniCLIP: Adapting CLIP for Video Recognition with Spatial-Temporal Omni-Scale Feature Learning
abstract
Recent Vision-Language Models (VLMs) e.g. CLIP have made great progress in video recognition. Despite the improvement brought by the strong visual backbone in extracting spatial features, CLIP still falls short in capturing and integrating spatial-temporal features which is essential for video recognition. In this paper, we propose OmniCLIP, a framework that adapts CLIP for video recognition by focusing on learning comprehensive features encompassing spatial, temporal, and dynamic spatial-temporal scales, which we refer to as omni-scale features. This is achieved through the design of spatial-temporal blocks that include parallel temporal adapters (PTA), enabling efficient temporal modeling. Additionally, we introduce a self-prompt generator (SPG) module to capture dynamic object spatial features. The synergy between PTA and SPG allows OmniCLIP to discern varying spatial information across frames and assess object scales over time. We have conducted extensive experiments in supervised video recognition, few-shot video recognition, and zero-shot recognition tasks. The results demonstrate the effectiveness of our method, especially with OmniCLIP achieving a top-1 accuracy of 74.30% on HMDB51 in a 16-shot setting, surpassing the recent MotionPrompt approach even with full training data. The code is available at https://github.com/XiaoBuL/OmniCLIP.
Mushui Liu, Bozheng Li
ECAI1
2024 Improving Zero-Shot Generalization for CLIP with Variational Adapter
Ziqian Lu, Mushui Liu, Yunlong Yu 0001, Xi Li 0001
ECCV (20)3
2024 HOGDA: Boosting Semi-supervised Graph Domain Adaptation via High-Order Structure-Guided Adaptive Feature Alignment
abstract
Semi-supervised graph domain adaptation, as a subfield of graph transfer learning, seeks to precisely annotate unlabeled target graph nodes by leveraging transferable features acquired from the limited labeled source nodes. However, most existing studies often directly utilize graph convolutional networks (GCNs)-based feature extractors to capture domain-invariant node features, while neglecting the issue that GCNs are insufficient in collecting complex structure information in graph. Considering the importance of graph structure information in encoding the complex relationship among nodes and edges, this paper aims to utilize such powerful information to assist graph transfer learning. To achieve this goal, we develop a novel framework called HOGDA. Concretely, HOGDA introduces a high-order structure information mixing (HSIM) module to effectively capture abundant structure information in graph, greatly enhancing the feature extractor's ability to adapt across different domains. Moreover, to achieve fine-grained feature distributions alignment, a novel strategy called adaptive weighted domain alignment (AWDA) is proposed to dynamically adjust the node weight during adversarial domain adaptation process, effectively boosting the model's transfer ability. Furthermore, to mitigate the overfitting phenomenon caused by limited source labeled nodes, we also design a trust-aware node clustering (TNC) strategy to guide the unlabeled nodes to achieve discriminative clustering. Extensive experimental results show that our HOGDA outperforms the state-of-the-art methods on various transfer tasks.
Jun Dan, Weiming Liu 0005, Mushui Liu, Chunfeng Xie, Shunjie Dong, Guofang Ma, Yanchao Tan, Jiazheng Xing
ACM Multimedia3
2024 Similar norm more transferable: Rethinking feature norms discrepancy in adversarial domain adaptation
Jun Dan, Mushui Liu, Chunfeng Xie, Jiawang Yu, Haoran Xie 0004, Ruokun Li, Shunjie Dong
Knowl. Based Syst.2
2024 Tolerant Self-Distillation for image classification
Mushui Liu, Yunlong Yu 0001, Zhong Ji, Jungong Han, Zhongfei Zhang
Neural Networks1
2023 Lightweight MIMO-WNet for single image deblurring
Mushui Liu, Yingming Li, Zhong Ji
Neurocomputing1