VLDB 2026 Research / reviewers in the wild / expert
Lu Zhang 0053
dblp:82/10609-53
· DBLP profile ↗
46ranked-venue papers
6as first author
40since 2021 · last 2026
0000-0003-4648-4437ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 29 · 6 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 29 · 4 first-author · 24 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Explicit Geometry-Reflectance Domain Shift Modeling for Robust LiDAR Segmentation in Adverse Weather
Longyu Yang, Shangbo Yuan, Lu Zhang 0053, Jun Liu 0036, Heng Tao Shen, Xiaofeng Zhu 0001, Ping Hu 0001 |
Int. J. Comput. Vis. | 3 |
| 2026 | One Aligned LLM to Serve Them All: A Transfer Recipe for Training VLMs without Visual-Language Re-Alignment
Jiazuo Yu 0001, Yunzhi Zhuge, Lu Zhang 0053, Wei Zhou 0021, Dong Wang 0004, Huchuan Lu, You He 0002 |
Int. J. Comput. Vis. | 3 |
| 2026 | Bridging CLIP and CLAP for open-vocabulary audio-visual segmentation with semantic coherence
Yunzhi Zhuge, Mengyuan Zhu, Yizhuang Peng, Lu Zhang 0053, Jin Zhan, Huchuan Lu |
Pattern Recognit. | 4 |
| 2026 | Parameter-Aware Mamba Model for Multitask Dense PredictionabstractUnderstanding the inter-relations and interactions between tasks is crucial for multitask dense prediction. Existing methods predominantly utilize convolutional layers and attention mechanisms to explore task-level interactions. In this work, we introduce a novel decoder-based framework, parameter-aware Mamba model (PAMM), specifically designed for dense prediction in multitask learning (MTL) setting. Distinct from approaches that employ Transformers to model holistic task relationships, PAMM leverages the rich, scalable parameters of state-space models (SSMs) to enhance task interconnectivity. It features dual state-space parameter experts (PEs) that integrate and set task-specific parameter priors (PPs), capturing the intrinsic properties of each task. This approach not only facilitates precise multitask interactions but also allows for the global integration of task priors through the structured state-space sequence (S4) model. Furthermore, we employ the multidirectional Hilbert scanning (MDHS) method to construct multiangle feature sequences, thereby enhancing the sequence model's perceptual capabilities for 2-D data. Extensive experiments on the NYUD-v2 and PASCAL-Context benchmarks demonstrate the effectiveness of our proposed method. Our code is available at https://github.com/CQC-gogopro/PAMM. Xinzhuo Yu, Yunzhi Zhuge, Sitong Gong, Lu Zhang 0053, Huchuan Lu |
IEEE Trans. Cybern. | 4 |
| 2026 | DiMuS: Disentangled Multi-Signal Learning for Weakly Supervised Point-Based 3D Object DetectionabstractWeakly supervised 3D object detection has emerged as a promising paradigm to reduce the reliance on costly 3D annotations. Existing methods often rely on 2D projection constraints or heuristic priors to supervise 3D box regression with inexpensive 2D labels. However, they still suffer from projection ambiguity and geometry inconsistency due to the entangled optimization of 3D parameters. In this paper, we propose DiMuS, a Disentangled Multi- $\boldsymbol {S}$ ignal learning framework that integrates complementary supervision from 2D boxes, LLM-derived semantic prior, and 3D geometric alignment to enhance distinct 3D properties of position, dimension, and orientation, respectively. Specifically, DiMuS incorporates three key components: (i) a Centerness-enhanced Projection Constraint (CPC) that improves position estimation through a centerness weighting strategy, (ii) a Semantic Prior Anchoring (SPA) module that leverages LLM-derived category-specific priors for robust dimension decoding, and (iii) a Rotation-aware Consistency Regularization (RCR) mechanism that enforces orientation consistency through synthetic rotations and self-supervised invariance learning. Additionally, an Adversarial Geometric Alignment (AGA) module is proposed to build attraction/repulsion forces between LiDAR points and box edges for dynamic boundary refinement. Extensive experiments on the KITTI dataset demonstrate that DiMuS outperforms previous weakly supervised methods, achieving 96.82% of fully supervised performance on car detection while maintaining robustness across different categories. Yunzhi Zhuge, Lu Zhang 0053, Ping Hu 0001, Dong Wang 0004, Huchuan Lu |
IEEE Trans. Image Process. | 3 |
| 2026 | Context-Infused Trajectories: Enhancing Context and Frame Consistency in Reasoning Video Object SegmentationabstractReasoning video object segmentation (ReaVOS) aims to segment referred objects in video sequences based on implicit and complex linguistic queries. Existing methods typically compress limited video frames into pooled representations and prompt multimodal large language models (MLLMs) to generate a single global segmentation token. However, this strategy lacks explicit contextual guidance and causes substantial loss of spatial details, limiting capability and segmentation consistency. To overcome these limitations, we introduce Context-infused Consistent Video Segmentor (CiCVS), a novel framework leveraging contextual information to guide generation of temporally coherent and accurate mask trajectories. CiCVS incorporates a Hierarchical Frame Sampling (HFS) module, which globally samples support frames across the entire video to ensure broad temporal coverage, and then uniformly selects target frames within the support set. It also employs a Contextual Token Prompting (CTP) module, which utilizes contextual cues from support frames to guide the MLLM in generating specialized tokens for various target frames, enabling the model to capture intricate temporal patterns and ensure consistency across long-range sequences. At the core of CTP is the Multimodal Injection Compressor (MIC) block, which efficiently integrates support frame features and textual semantic information into a compact set of latent queries, enhancing temporal-level object perception. To further advance the ReaVOS field, we introduce the CoCoRVOS benchmark, which features more temporally intricate reasoning instructions and a diverse set of video scenarios. Extensive experiments demonstrate that CiCVS establishes a new state-of-the-art on multiple benchmarks, achieving significant improvements in $\mathcal {J}\& \mathcal {F}$ scores, including +2.7 on CoCoRVOS, +1.4 on ReVOS, and +7.0 on ReasonVOS, underscoring its superior contextual reasoning and segmentation capabilities. Yunzhi Zhuge, Sitong Gong, Lu Zhang 0053, Qi Xu 0008, Wenda Zhao 0003, Jin Zhan, Huchuan Lu |
IEEE Trans. Image Process. | 3 |
| 2026 | Exploiting Cross-Task Synergy via Frequency-Driven Hierarchical Learning for Multi-Task Dense PredictionabstractMulti-task dense prediction improves pixel-level performance by leveraging shared representations and inter-task collaboration. However, existing approaches either rely on implicit task relationships or neglect frequency-domain cues that are essential for preserving fine-grained details and enhancing cross-task feature learning at multiple scales. As a result, they face persistent challenges in multi-scale feature fusion, effective task interaction, and accurate decoding. To address these issues, we propose a hierarchical frequency-driven framework, termed Hierarchical Frequency-Adaptive Network (HiFAN), that facilitates cross-task collaborative optimization via frequency-domain analysis. Specifically, we first design a task-adaptive fusion module that exploits multi-scale frequency-domain information to enhance spatial details. This module generates dynamic convolutional kernels with task-specific parameters and positional biases to adaptively accommodate diverse task requirements. Next, we introduce an efficient cross-task interaction module that leverages compact low-frequency representations to enable global context exchange across tasks. Finally, we present a high-frequency-aware decoder that mitigates feature smoothing and detail loss commonly introduced by Transformer-based decoders. We demonstrate the effectiveness of HiFAN on two standard multi-task learning benchmarks, PASCAL-Context and NYUD-v2, achieving strong and competitive performance across multiple tasks. The code and model weights are available in HiFAN. Yunzhi Zhuge, Xinzhuo Yu, Lu Zhang 0053, Xu Jia 0012, Jin Zhan, Huchuan Lu |
IEEE Trans. Image Process. | 3 |
| 2026 | Generic-to-Personalised Learning for Multimodal Image Synthesis With Bidirectional Variational GANabstractMultimodal image synthesis, which predicts target-modality images from source-modality images, has garnered considerable attention in the field of clinical diagnosis. Both unidirectional and bidirectional multimodal image synthesis methods have been explored in the medical domain, however, unidirectional models heavily rely on paired images, while current bidirectional models typically overlook local image details due to their unsupervised training patterns. In this work, we propose a Bidirectional Variational Generative Adversarial Network (BVGAN) for multimodal image synthesis, which achieves high-quality bidirectional translations between any two modalities using only a limited number paired images. Firstly, BVGAN's generator incorporates a variational structure (VAS) to regularise the latent space for noise reduction. This regularisation imposes smoothness to the latent space, enabling BVGAN to produce high-quality, noise-free images. Secondly, a novel generic-to-personalised (GTP) learning strategy is introduced to train BVGAN and reduce its reliance on a large sets of paired images. GTP initially leverages an unsupervised learning model to capture the global mapping between two modalities using unpaired images from generic patients. It then applies a supervised learning model to refine the mapping for individual patient, enhancing image details. Finally, the GTP learning strategy along with VAS enables BVGAN to achieve state-of-the-art performance on two multi-modality medical datasets: Brain CTMRI and BRATS. Long Chen 0019, Xirui Dong, Jiangrong Shen, Lu Zhang 0053, Qi Xu 0008, Gang Pan 0001, Qiang Zhang 0008 |
IEEE Trans. Multim. | 4 |
| 2026 | 3D-SceneQ: Empowering 3D LLM With Query-Guided Adaptive Pruning and Multi-Modal Feature EnhancementabstractCurrent 3D scene understanding pipelines typically concatenate vast numbers of 3D object tokens with text tokens and feed the resulting sequence to a Large Language Model (LLM). However, existing 3D-LLMs for scene understanding encounter two major limitations: (i) excessive inclusion of task-irrelevant object data, introducing noise that reduces reasoning accuracy and increases hallucinations; and (ii) reliance solely on point cloud data, which inherently lacks the rich semantic information available in complementary 2D modalities, such as color, material properties, texture, and high-level contextual relationships. To address these challenges, we introduce 3D-SceneQ, a 3D LLM that unifies adaptive token pruning with multimodal semantic enrichment, markedly advancing scene understanding, reasoning, and grounding. Specifically, we propose a Query-Guided Adaptive Pruning (QGAP) module that adaptively selects task-relevant objects based on the intended meaning of user instructions, leveraging a learnable latent query to retain critical information while effectively suppressing irrelevant noise. In addition, we introduce a Multi-modal Object-level Feature Enhancement (MOFE) module, which integrates 2D feature embeddings from pre-trained models into 3D representation to enrich object-level semantic information. Evaluations show that 3D-SceneQ outperforms current state-of-the-art methods across multiple benchmarks (e.g., +5.7 CIDEr on Scan2Cap and +17.0 EM@1 on SQA3D), demonstrating its remarkable capabilities in various 3D language reasoning tasks. Yongqi Shan, Lu Zhang 0053, Jiazuo Yu 0001, Yunzhi Zhuge, Huchuan Lu |
IEEE Trans. Multim. | 2 |
| 2025 | Bootstraping Clustering of Gaussians for View-consistent 3D Scene UnderstandingabstractInjecting semantics into 3D Gaussian Splatting (3DGS) has recently garnered significant attention. While current approaches typically distill 3D semantic features from 2D foundational models (e.g., CLIP and SAM) to facilitate novel view segmentation and semantic understanding, their heavy reliance on 2D supervision can undermine cross-view semantic consistency and necessitate complex data preparation processes, therefore hindering view-consistent scene understanding. In this work, we present FreeGS, an unsupervised semantic-embedded 3DGS framework that achieves view-consistent 3D scene understanding without the need for 2D labels. Instead of directly learning semantic features, we introduce the IDentity-coupled Semantic Field (IDSF) into 3DGS, which captures both semantic representations and view-consistent instance indices for each Gaussian. We optimize IDSF with a two-step alternating strategy: semantics help to extract coherent instances in 3D space, while the resulting instances regularize the injection of stable semantics from 2D space. Additionally, we adopt a 2D-3D joint contrastive loss to enhance the complementarity between view-consistent 3D geometry and rich semantics during the bootstrapping process, enabling FreeGS to uniformly perform tasks such as novel-view semantic segmentation, object selection, and 3D object detection. Extensive experiments on LERF-Mask, 3D-OVS, and ScanNet datasets demonstrate that FreeGS performs comparably to state-of-the-art methods while avoiding the complex data preprocessing workload. Lu Zhang 0053, Ping Hu 0001, Liqian Ma, Yunzhi Zhuge, Huchuan Lu |
AAAI | 2 |
| 2025 | The Devil is in Temporal Token: High Quality Video Reasoning SegmentationabstractExisting methods for Video Reasoning Segmentation rely heavily on a single special token to represent the object in the keyframe or the entire video, inadequately capturing spatial complexity and inter-frame motion. To overcome these challenges, we propose VRS-HQ, an end-to-end video reasoning segmentation approach that leverages Multimodal Large Language Models (MLLMs) to inject rich spatiotemporal features into hierarchical tokens. Our key innovations include a Temporal Dynamic Aggregation (TDA) and a Token-driven Keyframe Selection (TKS). Specifically, we design frame-leveland temporal-leveltokens that utilize MLLM’s autoregressive learning to effectively capture both local and global information. Subsequently, we apply a similarity-based weighted fusion and frame selection strategy, then utilize SAM2 to perform keyframe segmentation and propagation. To enhance keyframe localization accuracy, the TKS filters keyframes based on SAM2’s occlusion scores during inference. VRSHQ achieves state-of-the-art performance on ReVOS, surpassing VISA by 5.9%/12.5%/9.1% in ${\mathcal{J}}{{ \& }}{\mathcal{F}}$ scores across the three subsets. These results highlight the strong temporal reasoning and segmentation capabilities of our method. Code and model weights are available at VRS-HQ. Sitong Gong, Yunzhi Zhuge, Lu Zhang 0053, Zongxin Yang, Huchuan Lu |
CVPR | 3 |
| 2025 | Towards Explicit Geometry-Reflectance Collaboration for Generalized LiDAR Segmentation in Adverse WeatherabstractExisting LiDAR semantic segmentation models often suffer from decreased accuracy when exposed to adverse weather conditions. Recent methods addressing this issue focus on enhancing training data through weather simulation or universal augmentation techniques. However, few works have studied the negative impacts caused by the heterogeneous domain shifts in the geometric structure and reflectance intensity of point clouds. In this paper, we delve into this challenge and address it with a novel Geometry-Reflectance Collaboration (GRC) framework that explicitly separates feature extraction for geometry and reflectance. Specifically, GRC employs a dual-branch architecture designed to independently process geometric and reflectance features initially, thereby capitalizing on their distinct characteristic. Then, GRC adopts a robust multi-level feature collaboration module to suppress redundant and unreliable information from both branches. Consequently, without complex simulation or augmentation, our method effectively extracts intrinsic information about the scene while suppressing interference, thus achieving better robustness and generalization in adverse weather conditions. We demonstrate the effectiveness of GRC through comprehensive experiments on challenging benchmarks, showing that our method out-performs previous approaches and establishes new state-of-the-art results. Longyu Yang, Ping Hu 0001, Shangbo Yuan, Lu Zhang 0053, Jun Liu 0036, Heng Tao Shen, Xiaofeng Zhu 0001 |
CVPR | 4 |
| 2025 | Streaming Video Understanding and Multi-round Interaction with Memory-enhanced KnowledgeabstractRecent advances in Large Language Models (LLMs) have enabled the development of Video-LLMs, advancing multimodal learning by bridging video data with language tasks. However, current video understanding models struggle with processing long video sequences, supporting multi-turn dialogues, and adapting to real-world dynamic scenarios. To address these issues, we propose StreamChat, a training-free framework for streaming video reasoning and conversational interaction. StreamChat leverages a novel hierarchical memory system to efficiently process and compress video features over extended sequences, enabling real-time, multi-turn dialogue. Our framework incorporates a parallel system scheduling strategy that enhances processing speed and reduces latency, ensuring robust performance in real-world applications. Furthermore, we introduce StreamBench, a versatile benchmark that evaluates streaming video understanding across diverse media types and interactive scenarios, including multi-turn interactions and complex reasoning tasks. Extensive evaluations on StreamBench and other public benchmarks demonstrate that StreamChat significantly outperforms existing
state-of-the-art models in terms of accuracy and response times, confirming its effectiveness for streaming video understanding. Code is available at StreamChat. Haomiao Xiong, Zongxin Yang, Jiazuo Yu 0001, Yunzhi Zhuge, Lu Zhang 0053, Jiawen Zhu 0003, Huchuan Lu |
ICLR | 5 |
| 2025 | FDAVS: Exploring Frequency-Driven Modality Enhancement in Audio-Visual SegmentationabstractAudio-Visual Segmentation (AVS) aims to identify and delineate sounding objects within a video stream guided by auditory cues. Existing research focuses on audio-visual interactions in the spatial domain while overlooking the intrinsic frequency-domain information, leading to insufficient multimodal feature integration and alignment. In response, we introduce FDAVS, which incorporates frequency-driven designs to achieve synergistic utilization of spatial and frequency information of multimodal features. To begin with, we propose Frequency-Oriented Audio Integration (FOAI), which decouples and optimizes the high- and low-frequency components of visual features within the image encoder while integrating audio signals. Subsequently, we employ Frequency-Based Cross-modal Fusion (FBCF) between the pixel decoder and mask decoder to enhance the guidance of multimodal frequency information on the visual modality, increasing the consistency of audiovisual features. FDAVS achieves state-of-the-art segmentation performance across three AVS benchmarks, highlighting the effectiveness of frequency-driven modality enhancement. Mengyuan Zhu, Yunzhi Zhuge, Sitong Gong, Lu Zhang 0053, Huchuan Lu |
ICME | 4 |
| 2025 | Regularizing Subspace Redundancy of Low-Rank AdaptationabstractLow-Rank Adaptation (LoRA) and its variants have delivered strong capability in Parameter-Efficient Transfer Learning (PETL) by minimizing trainable parameters and benefiting from reparameterization. However, their projection matrices remain unrestricted during training, causing high representation redundancy and diminishing the effectiveness of feature adaptation in the resulting subspaces. While existing methods mitigate this by manually adjusting the rank or implicitly applying channel-wise masks, they lack flexibility and generalize poorly across various datasets and architectures. Hence, we propose ReSoRA, a method that explicitly models redundancy between mapping subspaces and adaptively Regularizes Subspace redundancy of Low-Rank Adaptation. Specifically, it theoretically decomposes the low-rank submatrices into multiple equivalent subspaces and systematically applies de-redundancy constraints to the feature distributions across different projections. Extensive experiments validate that our proposed method consistently facilitates existing state-of-the-art PETL methods across various backbones and datasets in vision-language retrieval and standard visual classification benchmarks. Besides, as a training supervision, ReSoRA can be seamlessly integrated into existing approaches in a plug-and-play manner, with no additional inference costs. Code is publicly available at: https://github.com/Lucenova/ReSoRA. Yue Zhu 0012, Haiwen Diao, Shang Gao 0012, Jiazuo Yu 0001, Jiawen Zhu 0003, Yunzhi Zhuge, Shuai Hao 0007, Xu Jia 0012, Lu Zhang 0053, Ying Zhang 0021, Huchuan Lu |
ACM Multimedia | 9 |
| 2025 | FineRS: Fine-grained Reasoning and Segmentation of Small Objects with Reinforcement LearningabstractMulti-modal Large Language Models (MLLMs) have shown remarkable capabilities across a wide range of vision-language tasks. However, due to the restricted input resolutions, MLLMs face significant challenges in precisely understanding and localizing visual details in high-resolution images---particularly when dealing with extra-small objects embedded in cluttered contexts.
To address this issue, we propose FineRS, a two-stage MLLM-based reinforcement learning framework for jointly reasoning and segmenting extremely small objects within high-resolution scenes. FineRS adopts a coarse-to-fine pipeline comprising Global Semantic Exploration (GSE) and Localized Perceptual Refinement (LPR). Specifically, GSE performs instruction-guided reasoning to generate a textural response and a coarse target region, while LPR refines this region to produce an accurate bounding box and segmentation mask. To couple the two stages, we introduce a locate-informed retrospective reward, where LPR's outputs are used to optimize GSE for more robust coarse region exploration. Additionally, we present FineRS-4k, a new dataset for evaluating MLLMs on attribute-level reasoning and pixel-level segmentation on subtle, small-scale targets in complex high-resolution scenes. Experimental results on FineRS-4k and public datasets demonstrate that our method consistently outperforms state-of-the-art MLLM-based approaches on both instruction-guided segmentation and visual reasoning tasks. Lu Zhang 0053, Jiazuo Yu 0001, Haomiao Xiong, Ping Hu 0001, Yunzhi Zhuge, Huchuan Lu, You He 0002 |
NeurIPS | 1 |
| 2025 | MoE-Adapters++: Toward More Efficient Continual Learning of Vision-Language Models Via Dynamic Mixture-of-Experts AdaptersabstractIn this paper, we first propose MoE-Adapters, a parameter-efficient training framework to alleviate long-term forgetting issues in incremental learning with Vision-Language Models (VLM). Our MoE-Adapters leverages incrementally added routers to activate and integrate exclusive expert adapters from a pre-defined static expert set, enabling the pre-trained CLIP to efficiently adapt to new tasks. To preserve the zero-shot capability of VLM, a Distribution Discriminative Auto-Selector (DDAS) is introduced that automatically routes in-distribution and out-of-distribution inputs to the MoE-Adapters and the original CLIP, respectively. However, relying on a static expert set and a separate distribution selector can lead to parameter redundancy and increased training complexity. In response, we further extend an MoE-Adapters++ framework by introducing dynamic MoE-adapters, which allows experts to be adaptively involved during the continual learning process. Additionally, a Latent Embedding Auto-Selector (LEAS) is proposed that incorporates distribution selection within CLIP to create a more unified architecture. Extensive experiments across diverse settings demonstrate that the proposed method consistently surpasses previous state-of-the-art approaches while concurrently improving training efficiency. Jiazuo Yu 0001, Zichen Huang 0004, Yunzhi Zhuge, Lu Zhang 0053, Ping Hu 0001, Dong Wang 0004, Huchuan Lu, You He 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | RVMamba: Selective Text-Vision Mamba for Referring Video Object SegmentationabstractExisting RVOS methods typically employ Transformers to model global cross-modal, temporal-spatial correspondences, but their quadratic complexity limits deployment on resource-constrained devices. To overcome this limitation, Mamba offers a sequence modeling framework with linear computational complexity. We introduceRVMamba, which utilizes weight modulation to selectively update hidden states across text-frame sequences, enabling effective linguistic context propagation, and a learning-based scanning strategy to efficiently capture spatio-temporal dependencies with linear memory consumption. Extensive experiments demonstrate thatRVMambaachieves state-of-the-art performance on public benchmarks, with significantly reduced memory growth, offering an efficient and scalable solution for long video processing. Zhenyu Chen 0001, Jiawen Zhu 0003, Lu Zhang 0053, Ping Hu 0001, Yunzhi Zhuge, Huchuan Lu, You He 0002 |
IEEE Signal Process. Lett. | 3 |
| 2025 | AVS-Mamba: Exploring Temporal and Multi-Modal Mamba for Audio-Visual SegmentationabstractThe essence of audio-visual segmentation (AVS) lies in locating and delineating sound-emitting objects within a video stream. While Transformer-based methods have shown promise, their handling of long-range dependencies struggles due to quadratic computational costs, presenting a bottleneck in complex scenarios. To overcome this limitation and facilitate complex multi-modal comprehension with linear complexity, we introduce AVS-Mamba, a selective state space model to address the AVS task. Our framework incorporates two key components for video understanding and cross-modal learning: Temporal Mamba Block for sequential video processing and Vision-to-Audio Fusion Block for advanced audio-vision integration. Building on this, we develop the Multi-scale Temporal Encoder, aimed at enhancing the learning of visual features across scales, facilitating the perception of intra- and inter-frame information. To perform multi-modal fusion, we propose the Modality Aggregation Decoder, leveraging the Vision-to-Audio Fusion Block to integrate visual features into audio features across both frame and temporal levels. Further, we adopt the Contextual Integration Pyramid to perform audio-to-vision spatial-temporal context collaboration. Through these innovative contributions, our approach achieves new state-of-the-art results on the AVSBench-object and AVSBench-semantic datasets. Sitong Gong, Yunzhi Zhuge, Lu Zhang 0053, Yifan Wang 0004, Lijun Wang 0001, Huchuan Lu |
IEEE Trans. Multim. | 3 |
| 2025 | Complementary and Contrastive Learning for Audio-Visual SegmentationabstractAudio-Visual Segmentation (AVS) aims to generate pixel-wise segmentation maps that correlate with the auditory signals of objects. This field has seen significant progress with numerous CNN and Transformer-based methods enhancing the segmentation accuracy and robustness. Traditional CNN approaches manage audio-visual interactions through basic operations like padding and multiplications but are restricted by CNNs' limited local receptive field. More recently, Transformer-based methods treat auditory cues as queries, utilizing attention mechanisms to enhance audio-visual cooperation within frames. Nevertheless, they typically struggle to extract multimodal coefficients and temporal dynamics adequately. To overcome these limitations, we present the Complementary and Contrastive Transformer (CCFormer), a novel framework adept at processing both local and global information and capturing spatial-temporal context comprehensively. Our CCFormer initiates with the Early Integration Module (EIM) that employs a parallel bilateral architecture, merging multi-scale visual features with audio data to boost cross-modal complementarity. To extract the intra-frame spatial features and facilitate the perception of temporal coherence, we introduce the Multi-query Transformer Module (MTM), which dynamically endows audio queries with learning capabilities and models the frame and video-level relations simultaneously. Furthermore, we propose the Bi-modal Contrastive Learning (BCL) to promote the alignment across both modalities in the unified feature space. Through the effective combination of those designs, our method sets new state-of-the-art benchmarks across the S4, MS3 and AVSS datasets. Sitong Gong, Yunzhi Zhuge, Lu Zhang 0053, Huchuan Lu |
IEEE Trans. Multim. | 3 |
| 2025 | 3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene UnderstandingabstractMulti-modal Large Language Models (MLLMs) exhibit impressive capabilities in 2D tasks, yet encounter challenges in discerning the spatial positions, interrelations, and causal logic in scenes when transitioning from 2D to 3D representations. We find that the limitations mainly lie in: i) the high annotation cost restricting the scale-up of volumes of 3D scene data, and ii) the lack of a straightforward and effective way to perceive 3D information which results in prolonged training durations and complicates the streamlined framework. To this end, we develop a pipeline based on open-source 2D MLLMs and LLMs to generate high-quality 3D-text pairs and construct 3DS-160 K, to enhance the pre-training process. Leveraging this high-quality pre-training data, we introduce the 3UR-LLM model, an end-to-end 3D MLLM designed for precise interpretation of 3D scenes, showcasing exceptional capability in navigating the complexities of the physical world. 3UR-LLM directly receives 3D point cloud as input and project 3D features fused with text instructions into a manageable set of tokens. Considering the computation burden derived from these hybrid tokens, we design a 3D compressor module to cohesively compress the 3D spatial cues and textual narrative. 3UR-LLM achieves promising performance with respect to the previous SOTAs, for instance, 3UR-LLM exceeds its counterparts by 7.1% CIDEr on ScanQA, while utilizing fewer training resources. The code and model weights for 3UR-LLM and the 3DS-160 K benchmark are available at 3UR-LLM. Haomiao Xiong, Yunzhi Zhuge, Jiawen Zhu 0003, Lu Zhang 0053, Huchuan Lu |
IEEE Trans. Multim. | 4 |
| 2025 | MaskTrack: Auto-Labeling and Stable Tracking for Video Object SegmentationabstractVideo object segmentation (VOS) has witnessed notable progress due to the establishment of video training datasets and the introduction of diverse, innovative network architectures. However, video mask annotation is a highly intricate and labor-intensive task, as meticulous frame-by-frame comparisons are needed to ascertain the positions and identities of targets in the subsequent frames. Current VOS benchmarks often annotate only a few instances in each video to save costs, which, however, hinders the model's understanding of the complete context of the video scenes. To simplify video annotation and achieve efficient dense labeling, we introduce a zero-shot auto-labeling strategy based on the segment anything model (SAM), enabling it to densely annotate video instances without access to any manual annotations. Moreover, although existing VOS methods demonstrate improving performance, segmenting long-term and complex video scenes remains challenging due to the difficulties in stably discriminating and tracking instance identities. To this end, we further introduce a new framework, MaskTrack, which excels in long-term VOS and also exhibits significant performance advantages in distinguishing instances in complex videos with densely packed similar objects. We conduct extensive experiments to demonstrate the effectiveness of the proposed method and show that without introducing image datasets for pretraining, it achieves excellent performance on both short-term (86.2% in YouTube-VOS val) and long-term (68.2% in LVOS val) VOS benchmarks. Our method also surprisingly demonstrates strong generalization ability and performs well in visual object tracking (VOT) (65.6% in VOTS2023) and referring VOS (RVOS) (65.2% in Ref YouTube VOS) challenges. Zhenyu Chen 0001, Lu Zhang 0053, Ping Hu 0001, Huchuan Lu, You He 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | Real-Time Semantic Segmentation via a Densely Aggregated Bilateral NetworkabstractWith the growing demands of applications on online devices, the speed-accuracy trade-off is critical in the semantic segmentation system. Recently, the bilateral segmentation network has shown promising capacity to achieve the balance between favorable accuracy and fast speed, and has become the mainstream backbone in real-time semantic segmentation. Segmentation of target objects relies on high-level semantics, whereas it requires detailed low-level features to model specific local patterns for accurate location. However, the lightweight backbone of bilateral architecture limits the extraction of semantic context and spatial details. And the late fusion of the bilateral streams incurs the insufficient aggregation of semantic context and spatial details. In this article, we propose a densely aggregated bilateral network (DAB-Net) for real-time semantic segmentation. In the context path, a patchwise context enhancement (PCE) module is proposed to efficiently capture the local semantic contextual information from spatialwise and channelwise, respectively. Meanwhile, a context-guided spatial path (CGSP) is designed to exploit more spatial information by encoding finer details from the raw image and the transition from the context path. Finally, with multiple interactions between bilateral branches, the intertwined outputs from bilateral streams are combined in a unified decoder for a final interaction to further enhance the feature representation, which generates the final segmentation prediction. Experimental results on three public benchmarks demonstrate that our proposed method achieves higher accuracy with a limited decay in speed, which performs favorably against state-of-the-art real-time approaches and runs at 31.1 frames/s (FPS) on the high resolution of . The source code is released at https://github.com/isyangshu/DABNet. Shu Yang 0004, Lu Zhang 0053, Shuai Liu 0009, Huchuan Lu, Hao Chen 0011 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | Learning Motion and Temporal Cues for Unsupervised Video Object SegmentationabstractIn this article, we address the challenges in unsupervised video object segmentation (UVOS) by proposing an efficient algorithm, termed MTNet, which concurrently exploits motion and temporal cues. Unlike previous methods that focus solely on integrating appearance with motion or on modeling temporal relations, our method combines both aspects by integrating them within a unified framework. MTNet is devised by effectively merging appearance and motion features during the feature extraction process within encoders, promoting a more complementary representation. To capture the intricate long-range contextual dynamics and information embedded within videos, a temporal transformer module is introduced, facilitating efficacious interframe interactions throughout a video clip. Furthermore, we employ a cascade of decoders all feature levels across all feature levels to optimally exploit the derived features, aiming to generate increasingly precise segmentation masks. As a result, MTNet provides a strong and compact framework that explores both temporal and cross-modality knowledge to robustly localize and track the primary object accurately in various challenging scenarios efficiently. Extensive experiments across diverse benchmarks conclusively show that our method not only attains state-of-the-art performance in UVOS but also delivers competitive results in video salient object detection (VSOD). These findings highlight the method's robust versatility and its adeptness in adapting to a range of segmentation tasks. The source code is available at https://github.com/hy0523/MTNet. Yunzhi Zhuge, Hongyu Gu, Lu Zhang 0053, Jinqing Qi, Huchuan Lu |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Boosting Continual Learning of Vision-Language Models via Mixture-of-Experts AdaptersabstractContinual learning can empower vision-language models to continuously acquire new knowledge, without the need for access to the entire historical dataset. However, mitigating the performance degradation in large-scale models is non-trivial due to (i) parameter shifts throughout life-long learning and (ii) significant computational burdens associated with full-model tuning. In this work, we present a parameter-efficient continual learning framework to alleviate long-term forgetting in incremental learning with vision-language models. Our approach involves the dynamic expansion of a pre-trained CLIP model, through the integration of Mixture-of-Experts (MoE) adapters in response to new tasks. To preserve the zero-shot recognition capability of vision-language models, we further introduce a Distribution Discriminative Auto-Selector (DDAS) that automatically routes in-distribution and out-of-distribution inputs to the MoE Adapter and the original CLIP, respectively. Through extensive experiments across various settings, our proposed method consistently outperforms previous state-of-the-art approaches while concurrently reducing parameter training burdens by 60%. Our code locates at https://github.com/JiazuoYu/MoE-Adapters4CL Jiazuo Yu 0001, Yunzhi Zhuge, Lu Zhang 0053, Ping Hu 0001, Dong Wang 0004, Huchuan Lu, You He 0002 |
CVPR | 3 |
| 2024 | LLMs Can Evolve Continually on Modality for X-Modal Reasoning
Jiazuo Yu 0001, Haomiao Xiong, Lu Zhang 0053, Haiwen Diao, Yunzhi Zhuge, Lanqing Hong, Dong Wang 0004, Huchuan Lu, You He 0002, Long Chen 0016 |
NeurIPS | 3 |
| 2024 | Video Frame Interpolation for Large Motion with Generative Prior
Xu Jia 0012, Lu Zhang 0053, Xiaomin Li 0001, Huchuan Lu |
PRCV (10) | 4 |
| 2024 | Learning depth-aware decomposition for single image dehazing
Yumeng Kang, Lu Zhang 0053, Ping Hu 0001, Yu Liu 0005, Huchuan Lu, You He 0002 |
Comput. Vis. Image Underst. | 2 |
| 2024 | Video Frame Interpolation With Many-to-Many Splatting and Spatial Selective RefinementabstractIn this work, we first propose a fully differentiable Many-to-Many (M2M) splatting framework to interpolate frames efficiently. Given a frame pair, we estimate multiple bidirectional flows to directly forward warp the pixels to the desired time step before fusing any overlapping pixels. In doing so, each source pixel renders multiple target pixels and each target pixel can be synthesized from a larger area of visual context, establishing a many-to-many splatting scheme with robustness to undesirable artifacts. For each input frame pair, M2M has a minuscule computational overhead when interpolating an arbitrary number of in-between frames, hence achieving fast multi-frame interpolation. However, directly warping and fusing pixels in the intensity domain is sensitive to the quality of motion estimation and may suffer from less effective representation capacity. To improve interpolation accuracy, we further extend an M2M++ framework by introducing a flexible Spatial Selective Refinement (SSR) component, which allows for trading computational efficiency for interpolation quality and vice versa. Instead of refining the entire interpolated frame, SSR only processes difficult regions selected under the guidance of an estimated error map, thereby avoiding redundant computation. Evaluation on multiple benchmark datasets shows that our method is able to improve the efficiency while maintaining competitive video interpolation quality, and it can be adjusted to use more or less compute as needed. Ping Hu 0001, Simon Niklaus, Lu Zhang 0053, Stan Sclaroff, Kate Saenko |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Learning Local-Global Representation for Scribble-Based RGB-D Salient Object Detection via TransformerabstractManual scribbles have been introduced to RGB-D Salient Object Detection (SOD) as a credible indicator for salient regions and backgrounds, helping to strike a balance between detection accuracy and labeling efficiency. Previous works address this task by constructing loss functions on semantics, edges, and structures to distinguish salient pixels from the background. However, using local representations extracted by CNNs or Transformers and the incomplete scribble annotations are ineffective in capturing the global contexts of salient objects, and thus cause inaccurate predictions in cluttered regions. In this paper, we propose a local-global representation learning framework by incorporating multi-perception information to boost scribble-based RGB-D SOD. Our system is composed of three sub-modules: Local Representation Aggregation (LRA), Global Representation Initialization (GRI) and Dual Transformer Decoder (DTD). The LRA module first conducts integration of multi-scale, multi-modal local representations extracted from RGB images and depth maps. The GRI module then learns inter- and intra-image representations to capture the global contexts of salient regions from different aspects. Finally, the DTD module alternately updates local-global representations through a dual Transformer architecture. Experimental results on six benchmarks demonstrate that the proposed method performs favorably against state-of-the-art scribble-based RGB-D SOD approaches and is competitive with the fully-supervised approaches. Yue Wang 0038, Lu Zhang 0053, Yunzhi Zhuge, Huchuan Lu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Few-shot Semantic Segmentation by Exploiting Dynamic and Regional ContextsabstractFew-shot Semantic Segmentation (FSS) has received increasing interests recently. Modeling effective interaction be-tween support and query images is a crucial challenge in existing prototype based methods. In this paper, we propose a Dynamic and Regional Context Network (DRCNet) to achieve sufficient support-query interaction for accurate FSS. A Dynamic Context Module (DCM) is first proposed to capture the spatial details in query images by building dynamic convolutions in local views. To further alleviate the undesirable noises, a Regional Context Module (RCM) is proposed to mine and exclude the background and ambiguous objects in query images by modeling the prototypes for ambiguous regions. Experimental results on Pascal-5iand COCO-20idatasets demonstrate that our proposed DRCNet performs significantly superior against state-of-the-art methods. Hongyu Gu, Yunzhi Zhuge, Lu Zhang 0053, Jinqing Qi, Huchuan Lu |
ICME | 3 |
| 2023 | Video Diffusion Models with Local-Global Context GuidanceabstractDiffusion models have emerged as a powerful paradigm in video synthesis tasks including prediction, generation, and interpolation. Due to the limitation of the computational budget, existing methods usually implement conditional diffusion models with an autoregressive inference pipeline, in which the future fragment is predicted based on the distribution of adjacent past frames. However, only the conditions from a few previous frames can't capture the global temporal coherence, leading to inconsistent or even outrageous results in long-term video prediction. In this paper, we propose a Local-Global Context guided Video Diffusion model (LGC-VD) to capture multi-perception conditions for producing high-quality videos in both conditional/unconditional settings. In LGC-VD, the UNet is implemented with stacked residual blocks with self-attention units, avoiding the undesirable computational cost in 3D Conv. We construct a local-global context guidance strategy to capture the multi-perceptual embedding of the past fragment to boost the consistency of future prediction. Furthermore, we propose a two-stage training strategy to alleviate the effect of noisy frames for more stable predictions. Our experiments demonstrate that the proposed method achieves favorable performance on video prediction, interpolation, and unconditional video generation. We release code at https://github.com/exisas/LGC-VD. Lu Zhang 0053, Yu Liu 0005, Zhizhuo Jiang, You He 0002 |
IJCAI | 2 |
| 2023 | A uniform transformer-based structure for feature fusion and enhancement for RGB-D saliency detection
Yue Wang 0038, Xu Jia 0012, Lu Zhang 0053, James H. Elder, Huchuan Lu |
Pattern Recognit. | 3 |
| 2022 | You Only Infer Once: Cross-Modal Meta-Transfer for Referring Video Object SegmentationabstractWe present YOFO (You Only inFer Once), a new paradigm for referring video object segmentation (RVOS) that operates in an one-stage manner. Our key insight is that the language descriptor should serve as target-specific guidance to identify the target object, while a direct feature fusion of image and language can increase feature complexity and thus may be sub-optimal for RVOS. To this end, we propose a meta-transfer module, which is trained in a learning-to-learn fashion and aims to transfer the target-specific information from the language domain to the image domain, while discarding the uncorrelated complex variations of language description. To bridge the gap between the image and language domains, we develop a multi-scale cross-modal feature mining block that aggregates all the essential features required by RVOS from both domains and generates regression labels for the meta-transfer module. The whole system can be trained in an end-to-end manner and shows competitive performance against state-of-the-art two-stage approaches. Dezhuang Li, Ruoqi Li, Lijun Wang 0001, Yifan Wang 0004, Jinqing Qi, Lu Zhang 0053, Ting Liu 0018, Qingquan Xu, Huchuan Lu |
AAAI | 6 |
| 2022 | Video Object Segmentation via Structural Feature Reconfiguration
Zhenyu Chen 0001, Ping Hu 0001, Lu Zhang 0053, Huchuan Lu, You He 0002, Maodi Hu |
ACCV (7) | 3 |
| 2022 | Depth-inspired Label Mining for Unsupervised RGB-D Salient Object DetectionabstractExisting deep learning-based unsupervised Salient Object Detection (SOD) methods heavily rely on the pseudo labels predicted from handcrafted features. However, the pseudo ground truth obtained only from RGB space would easily bring undesirable noises, especially in some complex scenarios. This naturally leads to the incorporation of extra depth modality with RGB images for more robust object identification, namely RGB-D SOD. Compared with the well-studied unsupervised SOD in the RGB domain, deep unsupervised RGB-D SOD is a less explored direction in the literature. In this paper, we propose to tackle this task by introducing a novel systemic design for high-quality pseudo-label mining. Our framework consists of two key components, Depth-inspired Label Generation (DLG) and Multi-source Uncertainty-aware Label Optimization (MULO). In DLG, a lightweight deep network is designed for automatically producing pseudo labels from depth maps in a self-supervised manner. Then, MULO introduces an effective pseudo label optimization strategy by learning the uncertainty of the pseudo labels from the depth domain and heuristic features. Extensive experiments demonstrate that the proposed method significantly outperforms the state-of-the-art unsupervised methods on mainstream benchmarks. Yue Wang 0038, Lu Zhang 0053, Jinqing Qi, Huchuan Lu |
ACM Multimedia | 3 |
| 2022 | Center-Boundary Dual Attention for Oriented Object Detection in Remote Sensing ImagesabstractRecently, anchor-free object detectors have shown promising performance in oriented object detection on remote sensing images. However, the objects in remote sensing images always have large variations in arbitrary orientations, sizes, and aspect ratios, which makes the existing anchor-free methods hard to obtain satisfactory results. In this article, we propose a novel anchor-free detector, center-boundary dual attention (CBDA) network (CBDA-Net), for fast and accurate oriented object detection on remote sensing images. In CBDA-Net, we construct a CBDA module, which utilizes a dual attention mechanism to extract attention features on the center and boundary regions of objects. The CBDA module can learn more essential features for rotating objects and reduce the interference from complex background. Besides, to resolve the influence of object aspect ratio on angle errors, we propose an aspect ratio weighted angle loss (arwLoss), where diffident penalties are assigned on the angle loss based on the aspect ratios of objects. This loss construction is effective in improving the detection accuracy of oriented objects, especially for slender objects. We conduct extensive experiments on two publish benchmarks, i.e., DOTA and HRSC2016. The experimental results demonstrate that our CBDA-Net achieves favorable performance against other anchor-free state of the arts with a real-time speed of 50 FPS. Shuai Liu 0009, Lu Zhang 0053, Huchuan Lu, You He 0002 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | Learning Motion-Appearance Co-Attention for Zero-Shot Video Object SegmentationabstractHow to make the appearance and motion information interact effectively to accommodate complex scenarios is a fundamental issue in flow-based zero-shot video object segmentation. In this paper, we propose an Attentive Multi-Modality Collaboration Network (AMC-Net) to utilize appearance and motion information uniformly. Specifically, AMC-Net fuses robust information from multi-modality features and promotes their collaboration in two stages. First, we propose a Multi-Modality Co-Attention Gate (MCG) on the bilateral encoder branches, in which a gate function is used to formulate co-attention scores for balancing the contributions of multi-modality features and suppressing the redundant and misleading information. Then, we propose a Motion Correction Module (MCM) with a visualmotion attention mechanism, which is constructed to emphasize the features of foreground objects by incorporating the spatio-temporal correspondence between appearance and motion cues. Extensive experiments on three public challenging benchmark datasets verify that our proposed network performs favorably against existing state-of-the-art methods via training with fewer data. The code is released at https://github.com/isyangshu/AMC-Net. Shu Yang 0004, Lu Zhang 0053, Jinqing Qi, Huchuan Lu |
ICCV | 2 |
| 2021 | Polar Ray: A Single-stage Angle-free Detector for Oriented Object Detection in Aerial ImagesabstractOriented bounding boxes are widely used for object detection in aerial images. Existing oriented object detection methods typically follow the general object detection paradigm by adding an extra rotation angle on the horizontal bounding boxes. However, the angular periodicity incurs the difficulty in angle regression and rotation sensitivity on bounding boxes. In this paper, we propose a new anchor-free oriented object detector, Polar Ray Network (PRNet), where object keypoints are represented by polar coordinates without angle regression. Our PRNet learns a set of polar rays from the object center to boundary with predefined equal-distributed angles. We introduce a dynamic PointConv module to optimize the regression of polar ray by incorporating object corner features. Furthermore, a classification feature guidance module is presented to improve the classification accuracy by incorporating more spatial contents from polar rays. Experimental results on two public datasets, i.e., DOTA and HRSC2016, demonstrate that the proposed PRNet significantly outperforms existing anchor-free detectors, and shows highly competitiveness with the state-of-the-art two-stage anchor-based methods. Shuai Liu 0009, Lu Zhang 0053, Shuai Hao 0007, Huchuan Lu, You He 0002 |
ACM Multimedia | 2 |
| 2021 | Deeply supervised group recursive saliency prediction
Lingwei Kong, Lu Zhang 0053, Huchuan Lu |
Neurocomputing | 3 |
| 2020 | Synergistic Saliency and Depth Prediction for RGB-D Saliency Detection
Yue Wang 0038, James H. Elder, Runmin Wu, Huchuan Lu, Lu Zhang 0053 |
ACCV (2) | 6 |
| 2020 | Unsupervised Video Object Segmentation with Joint Hotspot Tracking
Lu Zhang 0053, Jianming Zhang 0001, Zhe Lin 0001, Radomír Mech, Huchuan Lu, You He 0002 |
ECCV (14) | 1 |
| 2019 | CapSal: Leveraging Captioning to Boost Semantics for Salient Object DetectionabstractDetecting salient objects in cluttered scenes is a big challenge. To address this problem, we argue that the model needs to learn discriminative semantic features for salient objects. To this end, we propose to leverage captioning as an auxiliary semantic task to boost salient object detection in complex scenarios. Specifically, we develop a CapSal model which consists of two sub-networks, the Image Captioning Network (ICN) and the Local-Global Perception Network (LGPN). ICN encodes the embedding of a generated caption to capture the semantic information of major objects in the scene, while LGPN incorporates the captioning embedding with local-global visual contexts for predicting the saliency map. ICN and LGPN are jointly trained to model high-level semantics as well as visual saliency. Extensive experiments demonstrate the effectiveness of image captioning in boosting the performance of salient object detection. In particular, our model performs significantly better than the state-of-the-art methods on several challenging datasets of complex scenarios. Lu Zhang 0053, Jianming Zhang 0001, Zhe Lin 0001, Huchuan Lu, You He 0002 |
CVPR | 1 |
| 2019 | Fast Video Object Segmentation via Dynamic Targeting NetworkabstractWe propose a new model for fast and accurate video object segmentation. It consists of two convolutional neural networks, a Dynamic Targeting Network (DTN) and a Mask Refinement Network (MRN). DTN locates the object by dynamically focusing on regions of interest surrounding the target object. The target region is predicted by DTN via two sub-streams, Box Propagation (BP) and Box Re-identification (BR). The BP stream is faster but less effective at objects with large deformation or occlusion. The BR stream performs better in difficult scenarios at a higher computation cost. We propose a Decision Module (DM) to adaptively determine which sub-stream to use for each frame. Finally, MRN is exploited to predict segmentation within the target region. Experimental results on two public datasets demonstrate that the proposed model significantly outperforms existing methods without online training in both accuracy and efficiency, and is comparable to online training-based methods in accuracy with an order of magnitude faster speed. Lu Zhang 0053, Zhe Lin 0001, Jianming Zhang 0001, Huchuan Lu, You He 0002 |
ICCV | 1 |
| 2018 | A Bi-Directional Message Passing Model for Salient Object DetectionabstractRecent progress on salient object detection is beneficial from Fully Convolutional Neural Network (FCN). The saliency cues contained in multi-level convolutional features are complementary for detecting salient objects. How to integrate multi-level features becomes an open problem in saliency detection. In this paper, we propose a novel bi-directional message passing model to integrate multi-level features for salient object detection. At first, we adopt a Multi-scale Context-aware Feature Extraction Module (MCFEM) for multi-level feature maps to capture rich context information. Then a bi-directional structure is designed to pass messages between multi-level features, and a gate function is exploited to control the message passing rate. We use the features after message passing, which simultaneously encode semantic information and spatial details, to predict saliency maps. Finally, the predicted results are efficiently combined to generate the final saliency map. Quantitative and qualitative experiments on five benchmark datasets demonstrate that our proposed model performs favorably against the state-of-the-art methods under different evaluation metrics. Lu Zhang 0053, Ju Dai, Huchuan Lu, You He 0002, Gang Wang 0012 |
CVPR | 1 |
| 2016 | Saliency detection via extreme learning machine
Lu Zhang 0053, Huchuan Lu |
Neurocomputing | 1 |