Yunzhi Zhuge

dblp:304/1289 · also Yun-Zhi Zhuge · DBLP profile ↗
← Back
41ranked-venue papers
9as first author
36since 2021 · last 2026
0000-0002-4288-4516ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 29 · 6 first-author · 25 since 2021Artificial intelligence and machine learning · 24 · 4 first-author · 21 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Spatial-Frequency Spiking Neural Network for Underwater Object Detection
abstract
Underwater object detection presents significant challenges due to the unique visual degradations in underwater environments, such as low contrast, poor visibility, and blurry object boundaries. While ANNs have achieved impressive detection accuracy, their high computational cost and power consumption limit their deployment in resource-constrained underwater platforms. In this work, we propose a Spatial-Frequency Spiking Neural Network (SFSNN) that combines the energy-efficient and event-driven nature of Spiking Neural Networks (SNNs) with the discriminative power of spatial-frequency analysis. SFSNN introduces a novel spatial-frequency spiking module that integrates spatial and frequency-domain representations, enhancing edge and texture features crucial for object detection in murky waters. Furthermore, we adapt the YOLOX architecture into a spike-based detector via ANN-to-SNN conversion using signed spiking neurons. Extensive experiments on the RUOD dataset demonstrate that SFSNN achieves superior performance over both SNN- and ANN-based detection models, offering a compelling solution for low-power underwater object detection.
Long Chen 0019, Wei Miao 0006, Yunzhi Zhuge, Hongming Xu 0002, Qi Xu 0008
AAAI4
2026 One Aligned LLM to Serve Them All: A Transfer Recipe for Training VLMs without Visual-Language Re-Alignment
Jiazuo Yu 0001, Yunzhi Zhuge, Lu Zhang 0053, Wei Zhou 0021, Dong Wang 0004, Huchuan Lu, You He 0002
Int. J. Comput. Vis.2
2026 Aggregating global-scale pixel-wise forgery cues within a graph
Hengrun Zhao, Yifan Wang 0004, Yunzhi Zhuge, Lijun Wang 0001, Huchuan Lu
Neural Networks3
2026 SELongVLM: Empowering Long Video Language Models With Self-Corrective Clip Selection
abstract
Recent advances in multimodal large language models (MLLMs) have enabled impressive progress in visual-language reasoning, yet long-video understanding remains a formidable challenge due to the need for coherent reasoning over ultra-long spatiotemporal dependencies. Existing methods struggle with the vast candidate space for relevant information in long videos, often failing to distinguish meaningful events from redundant content. We identify two critical and previously under-explored issues: absolute redundancy, where static visual content inflates token counts without adding narrative value, and relative redundancy, where task-irrelevant segments introduce noise that impairs reasoning. Compounding these issues is the weak spatiotemporal modeling in current MLLMs, which limits their ability to capture complex event dynamics. To address these multifaceted challenges, we introduce SELongVLM, a dynamically lenient-to-stringent selection long video language model. SELongVLM integrates two coordinated branches: a Residual Token Pruner (RTP) that removes repetitive background tokens via inter-frame residual modeling thus mitigating absolute redundancy while preserving motion cues, and a Semantic-aware Self-Correction Selector (SCSelector) that progressively refines query-relevant clip selection without frame-level annotations to reduce relative redundancy, guided by a stringent-to-lenient self-correcting mechanism during optimization. To ensure causal continuity and bolster spatiotemporal reasoning across disjoint clips, the framework further incorporates an action-aware operation for intra-clip dynamics and a temporal memory for cross-clip context, enabling robust spatiotemporal inference on long videos. Extensive experiments across eight benchmarks demonstrate that SELongVLM markedly outperforms existing models on both general and specialized long-video tasks. Specifically, it achieves 65.5% on VideoMME and 69.8% on MLVU for general benchmarks, and delivers strong performance on four specialized benchmarks - for example, 39.2% on TOMATO for fine-grained temporal reasoning and 69.2% on EventBench for event-level understanding.
Kecheng Zhang, Zongxin Yang, Mingfei Han 0002, Yunzhi Zhuge, Haihong Hao, Zhihui Li 0001, Xiaojun Chang
IEEE Trans. Pattern Anal. Mach. Intell.4
2026 BeatDance: Generating beat-consistent 3D dance with hierarchical spatial-temporal modeling
Xiaojian Shen, Dahu Shi, Jianrong Zhang, Yunzhi Zhuge, Zhiliang Wu, Guanghui Yue 0001, Wei Zhou 0021
Pattern Recognit.7
2026 Bridging CLIP and CLAP for open-vocabulary audio-visual segmentation with semantic coherence
Yunzhi Zhuge, Mengyuan Zhu, Yizhuang Peng, Lu Zhang 0053, Jin Zhan, Huchuan Lu
Pattern Recognit.1
2026 Parameter-Aware Mamba Model for Multitask Dense Prediction
abstract
Understanding the inter-relations and interactions between tasks is crucial for multitask dense prediction. Existing methods predominantly utilize convolutional layers and attention mechanisms to explore task-level interactions. In this work, we introduce a novel decoder-based framework, parameter-aware Mamba model (PAMM), specifically designed for dense prediction in multitask learning (MTL) setting. Distinct from approaches that employ Transformers to model holistic task relationships, PAMM leverages the rich, scalable parameters of state-space models (SSMs) to enhance task interconnectivity. It features dual state-space parameter experts (PEs) that integrate and set task-specific parameter priors (PPs), capturing the intrinsic properties of each task. This approach not only facilitates precise multitask interactions but also allows for the global integration of task priors through the structured state-space sequence (S4) model. Furthermore, we employ the multidirectional Hilbert scanning (MDHS) method to construct multiangle feature sequences, thereby enhancing the sequence model's perceptual capabilities for 2-D data. Extensive experiments on the NYUD-v2 and PASCAL-Context benchmarks demonstrate the effectiveness of our proposed method. Our code is available at https://github.com/CQC-gogopro/PAMM.
Xinzhuo Yu, Yunzhi Zhuge, Sitong Gong, Lu Zhang 0053, Huchuan Lu
IEEE Trans. Cybern.2
2026 DiMuS: Disentangled Multi-Signal Learning for Weakly Supervised Point-Based 3D Object Detection
abstract
Weakly supervised 3D object detection has emerged as a promising paradigm to reduce the reliance on costly 3D annotations. Existing methods often rely on 2D projection constraints or heuristic priors to supervise 3D box regression with inexpensive 2D labels. However, they still suffer from projection ambiguity and geometry inconsistency due to the entangled optimization of 3D parameters. In this paper, we propose DiMuS, a Disentangled Multi- $\boldsymbol {S}$ ignal learning framework that integrates complementary supervision from 2D boxes, LLM-derived semantic prior, and 3D geometric alignment to enhance distinct 3D properties of position, dimension, and orientation, respectively. Specifically, DiMuS incorporates three key components: (i) a Centerness-enhanced Projection Constraint (CPC) that improves position estimation through a centerness weighting strategy, (ii) a Semantic Prior Anchoring (SPA) module that leverages LLM-derived category-specific priors for robust dimension decoding, and (iii) a Rotation-aware Consistency Regularization (RCR) mechanism that enforces orientation consistency through synthetic rotations and self-supervised invariance learning. Additionally, an Adversarial Geometric Alignment (AGA) module is proposed to build attraction/repulsion forces between LiDAR points and box edges for dynamic boundary refinement. Extensive experiments on the KITTI dataset demonstrate that DiMuS outperforms previous weakly supervised methods, achieving 96.82% of fully supervised performance on car detection while maintaining robustness across different categories.
Yunzhi Zhuge, Lu Zhang 0053, Ping Hu 0001, Dong Wang 0004, Huchuan Lu
IEEE Trans. Image Process.2
2026 Context-Infused Trajectories: Enhancing Context and Frame Consistency in Reasoning Video Object Segmentation
abstract
Reasoning video object segmentation (ReaVOS) aims to segment referred objects in video sequences based on implicit and complex linguistic queries. Existing methods typically compress limited video frames into pooled representations and prompt multimodal large language models (MLLMs) to generate a single global segmentation token. However, this strategy lacks explicit contextual guidance and causes substantial loss of spatial details, limiting capability and segmentation consistency. To overcome these limitations, we introduce Context-infused Consistent Video Segmentor (CiCVS), a novel framework leveraging contextual information to guide generation of temporally coherent and accurate mask trajectories. CiCVS incorporates a Hierarchical Frame Sampling (HFS) module, which globally samples support frames across the entire video to ensure broad temporal coverage, and then uniformly selects target frames within the support set. It also employs a Contextual Token Prompting (CTP) module, which utilizes contextual cues from support frames to guide the MLLM in generating specialized tokens for various target frames, enabling the model to capture intricate temporal patterns and ensure consistency across long-range sequences. At the core of CTP is the Multimodal Injection Compressor (MIC) block, which efficiently integrates support frame features and textual semantic information into a compact set of latent queries, enhancing temporal-level object perception. To further advance the ReaVOS field, we introduce the CoCoRVOS benchmark, which features more temporally intricate reasoning instructions and a diverse set of video scenarios. Extensive experiments demonstrate that CiCVS establishes a new state-of-the-art on multiple benchmarks, achieving significant improvements in $\mathcal {J}\& \mathcal {F}$ scores, including +2.7 on CoCoRVOS, +1.4 on ReVOS, and +7.0 on ReasonVOS, underscoring its superior contextual reasoning and segmentation capabilities.
Yunzhi Zhuge, Sitong Gong, Lu Zhang 0053, Qi Xu 0008, Wenda Zhao 0003, Jin Zhan, Huchuan Lu
IEEE Trans. Image Process.1
2026 Exploiting Cross-Task Synergy via Frequency-Driven Hierarchical Learning for Multi-Task Dense Prediction
abstract
Multi-task dense prediction improves pixel-level performance by leveraging shared representations and inter-task collaboration. However, existing approaches either rely on implicit task relationships or neglect frequency-domain cues that are essential for preserving fine-grained details and enhancing cross-task feature learning at multiple scales. As a result, they face persistent challenges in multi-scale feature fusion, effective task interaction, and accurate decoding. To address these issues, we propose a hierarchical frequency-driven framework, termed Hierarchical Frequency-Adaptive Network (HiFAN), that facilitates cross-task collaborative optimization via frequency-domain analysis. Specifically, we first design a task-adaptive fusion module that exploits multi-scale frequency-domain information to enhance spatial details. This module generates dynamic convolutional kernels with task-specific parameters and positional biases to adaptively accommodate diverse task requirements. Next, we introduce an efficient cross-task interaction module that leverages compact low-frequency representations to enable global context exchange across tasks. Finally, we present a high-frequency-aware decoder that mitigates feature smoothing and detail loss commonly introduced by Transformer-based decoders. We demonstrate the effectiveness of HiFAN on two standard multi-task learning benchmarks, PASCAL-Context and NYUD-v2, achieving strong and competitive performance across multiple tasks. The code and model weights are available in HiFAN.
Yunzhi Zhuge, Xinzhuo Yu, Lu Zhang 0053, Xu Jia 0012, Jin Zhan, Huchuan Lu
IEEE Trans. Image Process.1
2026 3D-SceneQ: Empowering 3D LLM With Query-Guided Adaptive Pruning and Multi-Modal Feature Enhancement
abstract
Current 3D scene understanding pipelines typically concatenate vast numbers of 3D object tokens with text tokens and feed the resulting sequence to a Large Language Model (LLM). However, existing 3D-LLMs for scene understanding encounter two major limitations: (i) excessive inclusion of task-irrelevant object data, introducing noise that reduces reasoning accuracy and increases hallucinations; and (ii) reliance solely on point cloud data, which inherently lacks the rich semantic information available in complementary 2D modalities, such as color, material properties, texture, and high-level contextual relationships. To address these challenges, we introduce 3D-SceneQ, a 3D LLM that unifies adaptive token pruning with multimodal semantic enrichment, markedly advancing scene understanding, reasoning, and grounding. Specifically, we propose a Query-Guided Adaptive Pruning (QGAP) module that adaptively selects task-relevant objects based on the intended meaning of user instructions, leveraging a learnable latent query to retain critical information while effectively suppressing irrelevant noise. In addition, we introduce a Multi-modal Object-level Feature Enhancement (MOFE) module, which integrates 2D feature embeddings from pre-trained models into 3D representation to enrich object-level semantic information. Evaluations show that 3D-SceneQ outperforms current state-of-the-art methods across multiple benchmarks (e.g., +5.7 CIDEr on Scan2Cap and +17.0 EM@1 on SQA3D), demonstrating its remarkable capabilities in various 3D language reasoning tasks.
Yongqi Shan, Lu Zhang 0053, Jiazuo Yu 0001, Yunzhi Zhuge, Huchuan Lu
IEEE Trans. Multim.4
2025 Towards Open-Vocabulary Remote Sensing Image Semantic Segmentation
abstract
Recently, deep learning based methods have revolutionized remote sensing image segmentation. However, these methods usually rely on a predefined semantic class set, thus needing additional image annotation and model training when adapting to new classes. More importantly, they are unable to segment arbitrary semantic classes. In this work, we introduce Open-Vocabulary Remote Sensing Image Semantic Segmentation (OVRSISS), which aims to segment arbitrary semantic classes in remote sensing images. To address the lack of OVRSISS datasets, we develop LandDiscover50K, a comprehensive dataset of 51,846 images covering 40 diverse semantic classes. In addition, we propose a novel framework named GSNet that integrates domain priors from special remote sensing models and versatile capabilities of general vision-language models. Technically, GSNet consists of a Dual-Stream Image Encoder (DSIE), a Query-Guided Feature Fusion (QGFF), and a Residual Information Preservation Decoder (RIPD). DSIE first captures comprehensive features from both special models and general models in dual streams. Then, with the guidance of variable vocabularies, QGFF integrates specialist and generalist features, enabling them to complement each other. Finally, RIPD is proposed to aggregate multi-source features for more accurate mask predictions. Experiments show that our method outperforms other methods by a large margin, and our proposed LandDiscover50K improves the performance of OVRSISS methods. The dataset and method will be publicly available.
Chengyang Ye, Yunzhi Zhuge
AAAI2
2025 Bootstraping Clustering of Gaussians for View-consistent 3D Scene Understanding
abstract
Injecting semantics into 3D Gaussian Splatting (3DGS) has recently garnered significant attention. While current approaches typically distill 3D semantic features from 2D foundational models (e.g., CLIP and SAM) to facilitate novel view segmentation and semantic understanding, their heavy reliance on 2D supervision can undermine cross-view semantic consistency and necessitate complex data preparation processes, therefore hindering view-consistent scene understanding. In this work, we present FreeGS, an unsupervised semantic-embedded 3DGS framework that achieves view-consistent 3D scene understanding without the need for 2D labels. Instead of directly learning semantic features, we introduce the IDentity-coupled Semantic Field (IDSF) into 3DGS, which captures both semantic representations and view-consistent instance indices for each Gaussian. We optimize IDSF with a two-step alternating strategy: semantics help to extract coherent instances in 3D space, while the resulting instances regularize the injection of stable semantics from 2D space. Additionally, we adopt a 2D-3D joint contrastive loss to enhance the complementarity between view-consistent 3D geometry and rich semantics during the bootstrapping process, enabling FreeGS to uniformly perform tasks such as novel-view semantic segmentation, object selection, and 3D object detection. Extensive experiments on LERF-Mask, 3D-OVS, and ScanNet datasets demonstrate that FreeGS performs comparably to state-of-the-art methods while avoiding the complex data preprocessing workload.
Lu Zhang 0053, Ping Hu 0001, Liqian Ma, Yunzhi Zhuge, Huchuan Lu
AAAI5
2025 The Devil is in Temporal Token: High Quality Video Reasoning Segmentation
abstract
Existing methods for Video Reasoning Segmentation rely heavily on a single special token to represent the object in the keyframe or the entire video, inadequately capturing spatial complexity and inter-frame motion. To overcome these challenges, we propose VRS-HQ, an end-to-end video reasoning segmentation approach that leverages Multimodal Large Language Models (MLLMs) to inject rich spatiotemporal features into hierarchical tokens. Our key innovations include a Temporal Dynamic Aggregation (TDA) and a Token-driven Keyframe Selection (TKS). Specifically, we design frame-leveland temporal-leveltokens that utilize MLLM’s autoregressive learning to effectively capture both local and global information. Subsequently, we apply a similarity-based weighted fusion and frame selection strategy, then utilize SAM2 to perform keyframe segmentation and propagation. To enhance keyframe localization accuracy, the TKS filters keyframes based on SAM2’s occlusion scores during inference. VRSHQ achieves state-of-the-art performance on ReVOS, surpassing VISA by 5.9%/12.5%/9.1% in ${\mathcal{J}}{{ \& }}{\mathcal{F}}$ scores across the three subsets. These results highlight the strong temporal reasoning and segmentation capabilities of our method. Code and model weights are available at VRS-HQ.
Sitong Gong, Yunzhi Zhuge, Lu Zhang 0053, Zongxin Yang, Huchuan Lu
CVPR2
2025 Streaming Video Understanding and Multi-round Interaction with Memory-enhanced Knowledge
abstract
Recent advances in Large Language Models (LLMs) have enabled the development of Video-LLMs, advancing multimodal learning by bridging video data with language tasks. However, current video understanding models struggle with processing long video sequences, supporting multi-turn dialogues, and adapting to real-world dynamic scenarios. To address these issues, we propose StreamChat, a training-free framework for streaming video reasoning and conversational interaction. StreamChat leverages a novel hierarchical memory system to efficiently process and compress video features over extended sequences, enabling real-time, multi-turn dialogue. Our framework incorporates a parallel system scheduling strategy that enhances processing speed and reduces latency, ensuring robust performance in real-world applications. Furthermore, we introduce StreamBench, a versatile benchmark that evaluates streaming video understanding across diverse media types and interactive scenarios, including multi-turn interactions and complex reasoning tasks. Extensive evaluations on StreamBench and other public benchmarks demonstrate that StreamChat significantly outperforms existing state-of-the-art models in terms of accuracy and response times, confirming its effectiveness for streaming video understanding. Code is available at StreamChat.
Haomiao Xiong, Zongxin Yang, Jiazuo Yu 0001, Yunzhi Zhuge, Lu Zhang 0053, Jiawen Zhu 0003, Huchuan Lu
ICLR4
2025 FDAVS: Exploring Frequency-Driven Modality Enhancement in Audio-Visual Segmentation
abstract
Audio-Visual Segmentation (AVS) aims to identify and delineate sounding objects within a video stream guided by auditory cues. Existing research focuses on audio-visual interactions in the spatial domain while overlooking the intrinsic frequency-domain information, leading to insufficient multimodal feature integration and alignment. In response, we introduce FDAVS, which incorporates frequency-driven designs to achieve synergistic utilization of spatial and frequency information of multimodal features. To begin with, we propose Frequency-Oriented Audio Integration (FOAI), which decouples and optimizes the high- and low-frequency components of visual features within the image encoder while integrating audio signals. Subsequently, we employ Frequency-Based Cross-modal Fusion (FBCF) between the pixel decoder and mask decoder to enhance the guidance of multimodal frequency information on the visual modality, increasing the consistency of audiovisual features. FDAVS achieves state-of-the-art segmentation performance across three AVS benchmarks, highlighting the effectiveness of frequency-driven modality enhancement.
Mengyuan Zhu, Yunzhi Zhuge, Sitong Gong, Lu Zhang 0053, Huchuan Lu
ICME2
2025 Regularizing Subspace Redundancy of Low-Rank Adaptation
abstract
Low-Rank Adaptation (LoRA) and its variants have delivered strong capability in Parameter-Efficient Transfer Learning (PETL) by minimizing trainable parameters and benefiting from reparameterization. However, their projection matrices remain unrestricted during training, causing high representation redundancy and diminishing the effectiveness of feature adaptation in the resulting subspaces. While existing methods mitigate this by manually adjusting the rank or implicitly applying channel-wise masks, they lack flexibility and generalize poorly across various datasets and architectures. Hence, we propose ReSoRA, a method that explicitly models redundancy between mapping subspaces and adaptively Regularizes Subspace redundancy of Low-Rank Adaptation. Specifically, it theoretically decomposes the low-rank submatrices into multiple equivalent subspaces and systematically applies de-redundancy constraints to the feature distributions across different projections. Extensive experiments validate that our proposed method consistently facilitates existing state-of-the-art PETL methods across various backbones and datasets in vision-language retrieval and standard visual classification benchmarks. Besides, as a training supervision, ReSoRA can be seamlessly integrated into existing approaches in a plug-and-play manner, with no additional inference costs. Code is publicly available at: https://github.com/Lucenova/ReSoRA.
Yue Zhu 0012, Haiwen Diao, Shang Gao 0012, Jiazuo Yu 0001, Jiawen Zhu 0003, Yunzhi Zhuge, Shuai Hao 0007, Xu Jia 0012, Lu Zhang 0053, Ying Zhang 0021, Huchuan Lu
ACM Multimedia6
2025 FineRS: Fine-grained Reasoning and Segmentation of Small Objects with Reinforcement Learning
abstract
Multi-modal Large Language Models (MLLMs) have shown remarkable capabilities across a wide range of vision-language tasks. However, due to the restricted input resolutions, MLLMs face significant challenges in precisely understanding and localizing visual details in high-resolution images---particularly when dealing with extra-small objects embedded in cluttered contexts. To address this issue, we propose FineRS, a two-stage MLLM-based reinforcement learning framework for jointly reasoning and segmenting extremely small objects within high-resolution scenes. FineRS adopts a coarse-to-fine pipeline comprising Global Semantic Exploration (GSE) and Localized Perceptual Refinement (LPR). Specifically, GSE performs instruction-guided reasoning to generate a textural response and a coarse target region, while LPR refines this region to produce an accurate bounding box and segmentation mask. To couple the two stages, we introduce a locate-informed retrospective reward, where LPR's outputs are used to optimize GSE for more robust coarse region exploration. Additionally, we present FineRS-4k, a new dataset for evaluating MLLMs on attribute-level reasoning and pixel-level segmentation on subtle, small-scale targets in complex high-resolution scenes. Experimental results on FineRS-4k and public datasets demonstrate that our method consistently outperforms state-of-the-art MLLM-based approaches on both instruction-guided segmentation and visual reasoning tasks.
Lu Zhang 0053, Jiazuo Yu 0001, Haomiao Xiong, Ping Hu 0001, Yunzhi Zhuge, Huchuan Lu, You He 0002
NeurIPS5
2025 MoE-Adapters++: Toward More Efficient Continual Learning of Vision-Language Models Via Dynamic Mixture-of-Experts Adapters
abstract
In this paper, we first propose MoE-Adapters, a parameter-efficient training framework to alleviate long-term forgetting issues in incremental learning with Vision-Language Models (VLM). Our MoE-Adapters leverages incrementally added routers to activate and integrate exclusive expert adapters from a pre-defined static expert set, enabling the pre-trained CLIP to efficiently adapt to new tasks. To preserve the zero-shot capability of VLM, a Distribution Discriminative Auto-Selector (DDAS) is introduced that automatically routes in-distribution and out-of-distribution inputs to the MoE-Adapters and the original CLIP, respectively. However, relying on a static expert set and a separate distribution selector can lead to parameter redundancy and increased training complexity. In response, we further extend an MoE-Adapters++ framework by introducing dynamic MoE-adapters, which allows experts to be adaptively involved during the continual learning process. Additionally, a Latent Embedding Auto-Selector (LEAS) is proposed that incorporates distribution selection within CLIP to create a more unified architecture. Extensive experiments across diverse settings demonstrate that the proposed method consistently surpasses previous state-of-the-art approaches while concurrently improving training efficiency.
Jiazuo Yu 0001, Zichen Huang 0004, Yunzhi Zhuge, Lu Zhang 0053, Ping Hu 0001, Dong Wang 0004, Huchuan Lu, You He 0002
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 RVMamba: Selective Text-Vision Mamba for Referring Video Object Segmentation
abstract
Existing RVOS methods typically employ Transformers to model global cross-modal, temporal-spatial correspondences, but their quadratic complexity limits deployment on resource-constrained devices. To overcome this limitation, Mamba offers a sequence modeling framework with linear computational complexity. We introduceRVMamba, which utilizes weight modulation to selectively update hidden states across text-frame sequences, enabling effective linguistic context propagation, and a learning-based scanning strategy to efficiently capture spatio-temporal dependencies with linear memory consumption. Extensive experiments demonstrate thatRVMambaachieves state-of-the-art performance on public benchmarks, with significantly reduced memory growth, offering an efficient and scalable solution for long video processing.
Zhenyu Chen 0001, Jiawen Zhu 0003, Lu Zhang 0053, Ping Hu 0001, Yunzhi Zhuge, Huchuan Lu, You He 0002
IEEE Signal Process. Lett.5
2025 AVS-Mamba: Exploring Temporal and Multi-Modal Mamba for Audio-Visual Segmentation
abstract
The essence of audio-visual segmentation (AVS) lies in locating and delineating sound-emitting objects within a video stream. While Transformer-based methods have shown promise, their handling of long-range dependencies struggles due to quadratic computational costs, presenting a bottleneck in complex scenarios. To overcome this limitation and facilitate complex multi-modal comprehension with linear complexity, we introduce AVS-Mamba, a selective state space model to address the AVS task. Our framework incorporates two key components for video understanding and cross-modal learning: Temporal Mamba Block for sequential video processing and Vision-to-Audio Fusion Block for advanced audio-vision integration. Building on this, we develop the Multi-scale Temporal Encoder, aimed at enhancing the learning of visual features across scales, facilitating the perception of intra- and inter-frame information. To perform multi-modal fusion, we propose the Modality Aggregation Decoder, leveraging the Vision-to-Audio Fusion Block to integrate visual features into audio features across both frame and temporal levels. Further, we adopt the Contextual Integration Pyramid to perform audio-to-vision spatial-temporal context collaboration. Through these innovative contributions, our approach achieves new state-of-the-art results on the AVSBench-object and AVSBench-semantic datasets.
Sitong Gong, Yunzhi Zhuge, Lu Zhang 0053, Yifan Wang 0004, Lijun Wang 0001, Huchuan Lu
IEEE Trans. Multim.2
2025 Complementary and Contrastive Learning for Audio-Visual Segmentation
abstract
Audio-Visual Segmentation (AVS) aims to generate pixel-wise segmentation maps that correlate with the auditory signals of objects. This field has seen significant progress with numerous CNN and Transformer-based methods enhancing the segmentation accuracy and robustness. Traditional CNN approaches manage audio-visual interactions through basic operations like padding and multiplications but are restricted by CNNs' limited local receptive field. More recently, Transformer-based methods treat auditory cues as queries, utilizing attention mechanisms to enhance audio-visual cooperation within frames. Nevertheless, they typically struggle to extract multimodal coefficients and temporal dynamics adequately. To overcome these limitations, we present the Complementary and Contrastive Transformer (CCFormer), a novel framework adept at processing both local and global information and capturing spatial-temporal context comprehensively. Our CCFormer initiates with the Early Integration Module (EIM) that employs a parallel bilateral architecture, merging multi-scale visual features with audio data to boost cross-modal complementarity. To extract the intra-frame spatial features and facilitate the perception of temporal coherence, we introduce the Multi-query Transformer Module (MTM), which dynamically endows audio queries with learning capabilities and models the frame and video-level relations simultaneously. Furthermore, we propose the Bi-modal Contrastive Learning (BCL) to promote the alignment across both modalities in the unified feature space. Through the effective combination of those designs, our method sets new state-of-the-art benchmarks across the S4, MS3 and AVSS datasets.
Sitong Gong, Yunzhi Zhuge, Lu Zhang 0053, Huchuan Lu
IEEE Trans. Multim.2
2025 StableIdentity: Inserting Anybody Into Anywhere at First Sight
abstract
Recent advances in large pretrained text-to-image generation models have shown unprecedented capabilities for high-quality human-centric generation, however, customizing face identity is still an intractable problem. Existing methods cannot ensure stable identity preservation and flexible editability, even with several images for each subject during training. In this work, we propose StableIdentity, which allows identity-consistent recontextualization with just one face image from a person seen for the first time. More specifically, we employ a face encoder with the identity prior to encode the input face, and then calibrate the face representation to align the distribution of a space with the editability prior, which is constructed from celeb names. By incorporating identity prior and editability prior, the learned identity can be injected anywhere with various contexts. In addition, we design a masked two-phase diffusion loss to boost the pixel-level perception of the input face and maintain the diversity of generation. Extensive experiments demonstrate our method outperforms previous customization methods. In addition, the learned identity can be flexibly combined with the off-theshelf modules such as ControlNet. Notably, to the best of our knowledge, we are the first to directly inject the identity learned from a single image into video/3D generation without finetuning. We believe that the proposed StableIdentity is an important step to unify image, video, and 3D customized generation models. The code is available: https://github.com/qinghew/StableIdentity.
Xu Jia 0012, Xiaomin Li 0001, Taiqing Li, Liqian Ma, Yunzhi Zhuge, Huchuan Lu
IEEE Trans. Multim.6
2025 3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding
abstract
Multi-modal Large Language Models (MLLMs) exhibit impressive capabilities in 2D tasks, yet encounter challenges in discerning the spatial positions, interrelations, and causal logic in scenes when transitioning from 2D to 3D representations. We find that the limitations mainly lie in: i) the high annotation cost restricting the scale-up of volumes of 3D scene data, and ii) the lack of a straightforward and effective way to perceive 3D information which results in prolonged training durations and complicates the streamlined framework. To this end, we develop a pipeline based on open-source 2D MLLMs and LLMs to generate high-quality 3D-text pairs and construct 3DS-160 K, to enhance the pre-training process. Leveraging this high-quality pre-training data, we introduce the 3UR-LLM model, an end-to-end 3D MLLM designed for precise interpretation of 3D scenes, showcasing exceptional capability in navigating the complexities of the physical world. 3UR-LLM directly receives 3D point cloud as input and project 3D features fused with text instructions into a manageable set of tokens. Considering the computation burden derived from these hybrid tokens, we design a 3D compressor module to cohesively compress the 3D spatial cues and textual narrative. 3UR-LLM achieves promising performance with respect to the previous SOTAs, for instance, 3UR-LLM exceeds its counterparts by 7.1% CIDEr on ScanQA, while utilizing fewer training resources. The code and model weights for 3UR-LLM and the 3DS-160 K benchmark are available at 3UR-LLM.
Haomiao Xiong, Yunzhi Zhuge, Jiawen Zhu 0003, Lu Zhang 0053, Huchuan Lu
IEEE Trans. Multim.2
2025 Learning Motion and Temporal Cues for Unsupervised Video Object Segmentation
abstract
In this article, we address the challenges in unsupervised video object segmentation (UVOS) by proposing an efficient algorithm, termed MTNet, which concurrently exploits motion and temporal cues. Unlike previous methods that focus solely on integrating appearance with motion or on modeling temporal relations, our method combines both aspects by integrating them within a unified framework. MTNet is devised by effectively merging appearance and motion features during the feature extraction process within encoders, promoting a more complementary representation. To capture the intricate long-range contextual dynamics and information embedded within videos, a temporal transformer module is introduced, facilitating efficacious interframe interactions throughout a video clip. Furthermore, we employ a cascade of decoders all feature levels across all feature levels to optimally exploit the derived features, aiming to generate increasingly precise segmentation masks. As a result, MTNet provides a strong and compact framework that explores both temporal and cross-modality knowledge to robustly localize and track the primary object accurately in various challenging scenarios efficiently. Extensive experiments across diverse benchmarks conclusively show that our method not only attains state-of-the-art performance in UVOS but also delivers competitive results in video salient object detection (VSOD). These findings highlight the method's robust versatility and its adeptness in adapting to a range of segmentation tasks. The source code is available at https://github.com/hy0523/MTNet.
Yunzhi Zhuge, Hongyu Gu, Lu Zhang 0053, Jinqing Qi, Huchuan Lu
IEEE Trans. Neural Networks Learn. Syst.1
2025 SAMControl: Controlling Pose and Object for Image Editing with Soft Attention Mask
abstract
To achieve content-consistent results in text-conditioned image editing, existing methods typically employ a reconstruction branch to capture the source image details via diffusion inversion and a generation branch to synthesize the target image based on the given textual prompt and the masked source image details. However, accurately segmenting source details is challenging with the current fixed-threshold mask strategy. Additionally, the inadequacies in the inversion process can lead to insufficient retention of source details. In this article, we propose a method called SAMControl (Soft Attention Mask) to adaptively control the pose and object details for image editing. SAMControl dynamically learns flexible attention masks for different images at various diffusion steps. Furthermore, in the reconstruction branch, we utilize a direct inversion technique to ensure the fidelity of source details within SAM. Extensive qualitative and quantitative results demonstrate the effectiveness of the proposed method.
Yue Zhang 0004, Chao Wang 0102, Yunzhi Zhuge, Hehe Fan, Xiaojun Chang, Cheng Deng 0002, Yi Yang 0001
ACM Trans. Multim. Comput. Commun. Appl.4
2024 DME: Unveiling the Bias for Better Generalized Monocular Depth Estimation
abstract
This paper aims to design monocular depth estimation models with better generalization abilities. To this end, we have conducted quantitative analysis and discovered two important insights. First, the Simulation Correlation phenomenon, commonly seen in long-tailed classification problems, also exists in monocular depth estimation, indicating that the imbalanced depth distribution in training data may be the cause of limited generalization ability. Second, the imbalanced and long-tail distribution of depth values extends beyond the dataset scale, and also manifests within each individual image, further exacerbating the challenge of monocular depth estimation. Motivated by the above findings, we propose the Distance-aware Multi-Expert (DME) depth estimation model. Unlike prior methods that handle different depth range indiscriminately, DME adopts a divide-and-conquer philosophy where each expert is responsible for depth estimation of regions within a specific depth range. As such, the depth distribution seen by each expert is more uniform and can be more easily predicted. A pixel-level routing module is further designed and learned to stitch the prediction of all experts into the final depth map. Experiments show that DME achieves state-of-the-art performance on both NYU-Depth v2 and KITTI, and also delivers favorable zero-shot generalization capability on unseen datasets.
Songsong Yu, Yifan Wang 0004, Yunzhi Zhuge, Lijun Wang 0001, Huchuan Lu
AAAI3
2024 3D Prompt Learning for RGB-D Tracking
Bocen Li, Yunzhi Zhuge, Lijun Wang 0001, Yifan Wang 0004, Huchuan Lu
ACCV (2)2
2024 Boosting Continual Learning of Vision-Language Models via Mixture-of-Experts Adapters
abstract
Continual learning can empower vision-language models to continuously acquire new knowledge, without the need for access to the entire historical dataset. However, mitigating the performance degradation in large-scale models is non-trivial due to (i) parameter shifts throughout life-long learning and (ii) significant computational burdens associated with full-model tuning. In this work, we present a parameter-efficient continual learning framework to alleviate long-term forgetting in incremental learning with vision-language models. Our approach involves the dynamic expansion of a pre-trained CLIP model, through the integration of Mixture-of-Experts (MoE) adapters in response to new tasks. To preserve the zero-shot recognition capability of vision-language models, we further introduce a Distribution Discriminative Auto-Selector (DDAS) that automatically routes in-distribution and out-of-distribution inputs to the MoE Adapter and the original CLIP, respectively. Through extensive experiments across various settings, our proposed method consistently outperforms previous state-of-the-art approaches while concurrently reducing parameter training burdens by 60%. Our code locates at https://github.com/JiazuoYu/MoE-Adapters4CL
Jiazuo Yu 0001, Yunzhi Zhuge, Lu Zhang 0053, Ping Hu 0001, Dong Wang 0004, Huchuan Lu, You He 0002
CVPR2
2024 SHERL: Synthesizing High Accuracy and Efficient Memory for Resource-Limited Transfer Learning
Haiwen Diao, Xu Jia 0012, Yunzhi Zhuge, Ying Zhang 0021, Huchuan Lu, Long Chen 0016
ECCV (44)4
2024 LLMs Can Evolve Continually on Modality for X-Modal Reasoning
Jiazuo Yu 0001, Haomiao Xiong, Lu Zhang 0053, Haiwen Diao, Yunzhi Zhuge, Lanqing Hong, Dong Wang 0004, Huchuan Lu, You He 0002, Long Chen 0016
NeurIPS5
2024 Learning Local-Global Representation for Scribble-Based RGB-D Salient Object Detection via Transformer
abstract
Manual scribbles have been introduced to RGB-D Salient Object Detection (SOD) as a credible indicator for salient regions and backgrounds, helping to strike a balance between detection accuracy and labeling efficiency. Previous works address this task by constructing loss functions on semantics, edges, and structures to distinguish salient pixels from the background. However, using local representations extracted by CNNs or Transformers and the incomplete scribble annotations are ineffective in capturing the global contexts of salient objects, and thus cause inaccurate predictions in cluttered regions. In this paper, we propose a local-global representation learning framework by incorporating multi-perception information to boost scribble-based RGB-D SOD. Our system is composed of three sub-modules: Local Representation Aggregation (LRA), Global Representation Initialization (GRI) and Dual Transformer Decoder (DTD). The LRA module first conducts integration of multi-scale, multi-modal local representations extracted from RGB images and depth maps. The GRI module then learns inter- and intra-image representations to capture the global contexts of salient regions from different aspects. Finally, the DTD module alternately updates local-global representations through a dual Transformer architecture. Experimental results on six benchmarks demonstrate that the proposed method performs favorably against state-of-the-art scribble-based RGB-D SOD approaches and is competitive with the fully-supervised approaches.
Yue Wang 0038, Lu Zhang 0053, Yunzhi Zhuge, Huchuan Lu
IEEE Trans. Circuits Syst. Video Technol.4
2023 CTVIS: Consistent Training for Online Video Instance Segmentation
abstract
The discrimination of instance embeddings plays a vital role in associating instances across time for online video instance segmentation (VIS). Instance embedding learning is directly supervised by the contrastive loss computed upon the contrastive items (CIs), which are sets of anchor/positive/negative embeddings. Recent online VIS methods leverage CIs sourced from one reference frame only, which we argue is insufficient for learning highly discriminative embeddings. Intuitively, a possible strategy to enhance CIs is replicating the inference phase during training. To this end, we propose a simple yet effective training strategy, called Consistent Training for Online VIS (CTVIS), which devotes to aligning the training and inference pipelines in terms of building CIs. Specifically, CTVIS constructs CIs by referring inference the momentum-averaged embedding and the memory bank storage mechanisms, and adding noise to the relevant embeddings. Such an extension allows a reliable comparison between embeddings of current instances and the stable representations of historical instances, thereby conferring an advantage in modeling VIS challenges such as occlusion, re-identification, and deformation. Empirically, CTVIS outstrips the SOTA VIS models by up to +5.0 points on three VIS benchmarks, including YTVIS19 (55.1% AP), YTVIS21 (50.1% AP) and OVIS (35.5% AP). Furthermore, we find that pseudo-videos transformed from images can train robust models surpassing fully-supervised ones.
Kaining Ying, Weian Mao, Zhenhua Wang 0003, Hao Chen 0041, Lin Wu 0001, Yifan Liu 0001, Chengxiang Fan, Yunzhi Zhuge, Chunhua Shen
ICCV9
2023 Few-shot Semantic Segmentation by Exploiting Dynamic and Regional Contexts
abstract
Few-shot Semantic Segmentation (FSS) has received increasing interests recently. Modeling effective interaction be-tween support and query images is a crucial challenge in existing prototype based methods. In this paper, we propose a Dynamic and Regional Context Network (DRCNet) to achieve sufficient support-query interaction for accurate FSS. A Dynamic Context Module (DCM) is first proposed to capture the spatial details in query images by building dynamic convolutions in local views. To further alleviate the undesirable noises, a Regional Context Module (RCM) is proposed to mine and exclude the background and ambiguous objects in query images by modeling the prototypes for ambiguous regions. Experimental results on Pascal-5iand COCO-20idatasets demonstrate that our proposed DRCNet performs significantly superior against state-of-the-art methods.
Hongyu Gu, Yunzhi Zhuge, Lu Zhang 0053, Jinqing Qi, Huchuan Lu
ICME2
2022 Multi-granularity Transformer for Image Super-Resolution
Yunzhi Zhuge, Xu Jia 0012
ACCV (3)1
2021 Deep Reasoning Network for Few-shot Semantic Segmentation
abstract
Few-shot Semantic Segmentation (FSS) is a challenging problem in computer vision. It aims at segmenting objects of the unseen categories given only one or several annotated samples. The essence of FSS is to disseminate information from support images to query images for segmenting the mutual object categories. In this paper, we propose a Dynamic Reasoning Network (DRNet) to adaptively generate the parameters of predicting layers and infer the segmentation mask for each unseen category. More specifically, an Attentional Feature Integration Sub-network (AFIS) is first proposed to extract consistent features from support im-ages and query images. With shared weights, it stimulates the category consistency of different data streams. Then a Pooling-based Guidance Module (PGM) is used to cor-relate support features with query features progressively. To disseminate information from support images to various query images, we further propose a Dynamic PredictionModule (DPM) for generating the parameters of predicting layers. The proposed modules are unified for the dynamic reasoning of each query image segmentation. Experiments on two public benchmarks have demonstrated that our approach achieves superior performance and outperforms thevery recent state-of-the-art methods.
Yunzhi Zhuge, Chunhua Shen
ACM Multimedia1
2019 Deep Embedding Features for Salient Object Detection
abstract
Benefiting from the rapid development of Convolutional Neural Networks (CNNs), some salient object detection methods have achieved remarkable results by utilizing multi-level convolutional features. However, the saliency training datasets is of limited scale due to the high cost of pixel-level labeling, which leads to a limited generalization of the trained model on new scenarios during testing. Besides, some FCN-based methods directly integrate multi-level features, ignoring the fact that the noise in some features are harmful to saliency detection. In this paper, we propose a novel approach that transforms prior information into an embedding space to select attentive features and filter out outliers for salient object detection. Our network firstly generates a coarse prediction map through an encorder-decorder structure. Then a Feature Embedding Network (FEN) is trained to embed each pixel of the coarse map into a metric space, which incorporates much attentive features that highlight salient regions and suppress the response of non-salient regions. Further, the embedded features are refined through a deep-to-shallow Recursive Feature Integration Network (RFIN) to improve the details of prediction maps. Moreover, to alleviate the blurred boundaries, we propose a Guided Filter Refinement Network (GFRN) to jointly optimize the predicted results and the learnable guidance maps. Extensive experiments on five benchmark datasets demonstrate that our method outperforms state-of-the-art results. Our proposed method is end-to-end and achieves a realtime speed of 38 FPS.
Yunzhi Zhuge, Yu Zeng 0001, Huchuan Lu
AAAI1
2019 Multi-Source Weak Supervision for Saliency Detection
abstract
The high cost of pixel-level annotations makes it appealing to train saliency detection models with weak supervision. However, a single weak supervision source usually does not contain enough information to train a well-performing model. To this end, we propose a unified framework to train saliency detection models with diverse weak supervision sources. In this paper, we use category labels, captions, and unlabelled data for training, yet other supervision sources can also be plugged into this flexible framework. We design a classification network (CNet) and a caption generation network (PNet), which learn to predict object categories and generate captions, respectively, meanwhile highlight the most important regions for corresponding tasks. An attention transfer loss is designed to transmit supervision signal between networks, such that the network designed to be trained with one supervision source can benefit from another. An attention coherence loss is defined on unlabelled data to encourage the networks to detect generally salient regions instead of task-specific regions. We use CNet and PNet to generate pixel-level pseudo labels to train a saliency prediction network (SNet). During the testing phases, we only need SNet to predict saliency maps. Experiments demonstrate the performance of our method compares favourably against unsupervised and weakly supervised methods and even some supervised methods.
Yu Zeng 0001, Yunzhi Zhuge, Huchuan Lu, Lihe Zhang, Mingyang Qian, Yizhou Yu
CVPR2
2019 Joint Learning of Saliency Detection and Weakly Supervised Semantic Segmentation
abstract
Existing weakly supervised semantic segmentation (WSSS) methods usually utilize the results of pre-trained saliency detection (SD) models without explicitly modelling the connections between the two tasks, which is not the most efficient configuration. Here we propose a unified multi-task learning framework to jointly solve WSSS and SD using a single network, i.e. saliency and segmentation network (SSNet). SSNet consists of a segmentation network (SN) and a saliency aggregation module (SAM). For an input image, SN generates the segmentation result and, SAM predicts the saliency of each category and aggregating the segmentation masks of all categories into a saliency map. The proposed network is trained end-to-end with image-level category labels and class-agnostic pixel-level saliency labels. Experiments on PASCAL VOC 2012 segmentation dataset and four saliency benchmark datasets show the performance of our method compares favorably against state-of-the-art weakly supervised segmentation methods and fully supervised saliency detection methods.
Yu Zeng 0001, Yunzhi Zhuge, Huchuan Lu, Lihe Zhang
ICCV2
2018 Boundary-Guided Feature Aggregation Network for Salient Object Detection
abstract
Fully convolutional networks (FCN) has significantly improved the performance of many pixel-labeling tasks, such as semantic segmentation and depth estimation. However, it still remains nontrivial to thoroughly utilize the multilevel convolutional feature maps and boundary information for salient object detection. In this letter, we propose a novel FCN framework to integrate multilevel convolutional features recurrently with the guidance of object boundary information. First, a deep convolutional network is used to extract multilevel feature maps and separately aggregate them into multiple resolutions, which can be used to generate coarse saliency maps. Meanwhile, another boundary information extraction branch is proposed to generate boundary features. Finally, an attention-based feature fusion module is designed to fuse boundary information into salient regions to achieve accurate boundary inference and semantic enhancement. The final saliency maps are the combination of the predicted boundary maps and integrated saliency maps, which are more closer to the ground truths. Experiments and analysis on four large-scale benchmarks verify that our framework achieves new state-of-the-art results.
Yunzhi Zhuge, Gang Yang 0002, Huchuan Lu
IEEE Signal Process. Lett.1
2015 Robust Video Text Detection with Morphological Filtering Enhanced MSER
Yunzhi Zhuge, Huchuan Lu
J. Comput. Sci. Technol.1