Huchuan Lu

dblp:64/6896 · DBLP profile ↗
← Back
512ranked-venue papers
18as first author
283since 2021 · last 2026
0000-0002-6668-9758ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 355 · 13 first-author · 184 since 2021Artificial intelligence and machine learning · 308 · 10 first-author · 168 since 2021Applied, interdisciplinary, general and emerging computing · 21 · 18 since 2021Systems, architecture and hardware · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 CADTrack: Learning Contextual Aggregation with Deformable Alignment for Robust RGBT Tracking
abstract
RGB-Thermal (RGBT) tracking aims to exploit visible and thermal infrared modalities for robust all-weather object tracking. However, existing RGBT trackers struggle to resolve modality discrepancies, which poses great challenges for robust feature representation. This limitation hinders effective cross-modal information propagation and fusion, which significantly reduces the tracking accuracy. To address this limitation, we propose a novel Contextual Aggregation with Deformable Alignment framework called CADTrack for RGBT Tracking. To be specific, we first deploy the Mamba-based Feature Interaction (MFI) that establishes efficient feature interaction via state space models. This interaction module can operate with linear complexity, reducing computational cost and improving feature discrimination. Then, we propose the Contextual Aggregation Module (CAM) that dynamically activates backbone layers through sparse gating based on the Mixture-of-Experts (MoE). This module can encode complementary contextual information from cross-layer features. Finally, we propose the Deformable Alignment Module (DAM) to integrate deformable sampling and temporal propagation, mitigating spatial misalignment and localization drift. With the above components, our CADTrack achieves robust and accurate tracking in complex scenarios. Extensive experiments on five RGBT tracking benchmarks verify the effectiveness of our proposed method.
Hao Li 0101, Xiantao Hu, Wenning Hao, Dong Wang 0004, Huchuan Lu
AAAI7
2026 X-ReID: Multi-granularity Information Interaction for Video-Based Visible-Infrared Person Re-Identification
abstract
Large-scale vision-language models (e.g., CLIP) have recently achieved remarkable performance in retrieval tasks, yet their potential for Video-based Visible-Infrared Person Re-Identification (VVI-ReID) remains largely unexplored. The primary challenges are narrowing the modality gap and leveraging spatiotemporal information in video sequences. To address the above issues, in this paper, we propose a novel cross-modality feature learning framework named X-ReID for VVI-ReID. Specifically, we first propose a Cross-modality Prototype Collaboration (CPC) to align and integrate features from different modalities, guiding the network to reduce the modality discrepancy. Then, a Multi-granularity Information Interaction (MII) is designed, incorporating short-term interactions from adjacent frames, long-term cross-frame information fusion, and cross-modality feature alignment to enhance temporal modeling and further reduce modality gaps. Finally, by integrating multi-granularity information, a robust sequence-level representation is achieved. Extensive experiments on two large-scale VVI-ReID benchmarks (i.e., HITSZ-VCM and BUPTCampus) demonstrate the superiority of our method over state-of-the-art methods.
Chenyang Yu, Xuehu Liu, Huchuan Lu
AAAI4
2026 SAM3-I: Segment Anything with Instructions
abstract
Jingjing Li, Yue Feng, Yuchen Guo, Jincai Huang, Wei Ji, Qi Bi, Yongri Piao, Miao Zhang, Xiaoqi Zhao, Qiang Chen, Shihao Zou, Huchuan Lu, Li Cheng. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Jincai Huang 0003, Wei Ji 0011, Qi Bi, Yongri Piao, Miao Zhang 0004, Xiaoqi Zhao 0003, Qiang Chen 0007, Shihao Zou, Huchuan Lu, Li Cheng 0001
ACL (1)12
2026 MedSegFM: A Generative Perspective for Lesion Segmentation via Flow Matching
Shijie Chang, Yilong Hu, Lihe Zhang, Zechen Liu, Huchuan Lu
Int. J. Comput. Vis.5
2026 One Aligned LLM to Serve Them All: A Transfer Recipe for Training VLMs without Visual-Language Re-Alignment
Jiazuo Yu 0001, Yunzhi Zhuge, Lu Zhang 0053, Wei Zhou 0021, Dong Wang 0004, Huchuan Lu, You He 0002
Int. J. Comput. Vis.7
2026 Power Battery Detection
Xiaoqi Zhao 0003, Peiqian Cao, Chenyang Yu, Zonglei Feng, Lihe Zhang, Hanqi Liu, Jiaming Zuo, Youwei Pang, Jinsong Ouyang, Weisi Lin, Georges El Fakhri, Huchuan Lu, Xiaofeng Liu 0001
Int. J. Comput. Vis.12
2026 Aggregating global-scale pixel-wise forgery cues within a graph
Hengrun Zhao, Yifan Wang 0004, Yunzhi Zhuge, Lijun Wang 0001, Huchuan Lu
Neural Networks5
2026 MRCNet: Motion Reasoning Chain for Cross Modal Video Camouflaged Object Detection
abstract
Video camouflaged object detection (VCOD) aims to identify objects that seamlessly blend into their surroundings in video sequences. Traditional methods merely rely on visual cues to capture inter-frame motion that reveals camouflaged objects. However, the high similarity between camouflaged objects and their environments often renders pure reliance on visual cues unreliable. Additionally, random motions including camera shaking and abrupt scene transitions also inevitably bring noise into the identification process. To overcome these challenges, we propose a Motion Reasoning Chain Network (MRCNet), a novel cross-modal VCOD framework that emulates the human thought process when observing camouflaged objects, i.e., motion reasoning. Specifically, we introduce a generative sampling strategy grounded in multimodal large language models (MLLMs) to bridge the implicit knowledge space of MLLMs and the explicit representation space regarding the attributes of camouflaged objects, thereby enabling the effective establishment of the motion reasoning chain tailored for VCOD. This process provides semantic guidance for visual comprehension of camouflaged objects through motion and concept attribute reasoning. To improve the identification capability of camouflaged objects, we develop motion representation learning driven by the motion reasoning chain. It introduces hierarchical de-biased motion prototype learning to mitigate hallucinations of MLLMs, boosting the motion perception. To learn precise prompts for the visual foundation model, cross-modal prompt learning further incorporates the de-biased concept prototype into visual representations to enhance the visual comprehension of camouflaged objects. Extensive experiments across three datasets demonstrate that MRCNet achieves state-of-the-art results on both general metrics and spatiotemporal consistency metrics.
Wenjun Hui, Zhenfeng Zhu, Shuai Zheng 0005, Ming-Ming Cheng, Huchuan Lu, Yao Zhao 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2026 First-Order Cross-Domain Meta Learning for Few-Shot Remote Sensing Object Classification
abstract
Remote sensing images exhibit intrinsic domain complexity arising from multi-source sensor variances, which heterogeneity fundamentally challenges conventional cross-domain few-shot methods that assume simple distribution shifts. Addressing this, we propose a first-order Cross-Domain Meta Learning (CDML) for few-shot remote sensing object classification. CDML implements a dual-stage domain adaptation task as the fundamental meta-learning unit, and includes a cross-domain meta-train phase (CDMTrain) and a cross-domain meta-test phase (CDMTest). In CDMTrain, we propose an inner-loop multi-domain few-shot task sampling, which enables a teacher model encapsulate both cross-category discriminative features and authentic inter-domain distributional divergence. This alternating cyclic learning paradigm captures genuine domain shifts, with each update direction progressively guiding the model toward parameters that balance multi-domain performance. In CDMTest, we evaluate a domain diversity enhancement by transferring teacher parameters to the student model for cross-domain capability assessment on the reserved pseudo-unseen domain. The task-level design progressively improves domain generalization through iterative domain adaptive task learning. Meanwhile, to mitigate the conflicts and inadequacies caused by multi-domain scenarios, we propose a learnable affine transformation model. It adaptively learns affine transformation parameters through intermediate layer features to fine-tune the update direction. Extensive experiments on five remote sensing classification benchmarks demonstrate a superior performance of the proposed method compared with the state-of-the-art methods.
Wenda Zhao 0003, Huchuan Lu
IEEE Trans. Pattern Anal. Mach. Intell.4
2026 CDTFusion: Crossing Domain and Task for Infrared and Visible Image Fusion
abstract
Infrared and visible images present different domains that hinder the fusion process, thereby losing texture details. Besides, the low-level fusion and subsequent high-level segmentation appear cross-task feature gap that impedes their mutual promotion, causing blurred object edges. Addressing the above issues, this paper proposes a novel infrared and visible image fusion method that simultaneously crosses domain and task. First, a swap image translation strategy is built to transfer the features of visible and infrared images into an adaptive domain. Meanwhile, a global-local constraint is introduced to achieve overall domain space transfer, and shorten their feature distance. Second, a task interaction & query module is designed to explore the cross-task feature interactive relationship, which is then used as a bridge to realize the gradient backpropagation. Thus, a fine-grained mapping from the segmentation feature to fusion feature is obtained. Extensive experiments demonstrate that the proposed method exhibits superior fusion and segmentation performance than the state-of-the-art methods.
Wenda Zhao 0003, You He 0002, Huchuan Lu
IEEE Trans. Pattern Anal. Mach. Intell.5
2026 DecoupleMAD: Boosting sensitivity in multimodal anomaly detection via representation decoupling
Yuan Zhao 0006, Bocen Li, Lihe Zhang, Huchuan Lu
Pattern Recognit.5
2026 Bridging CLIP and CLAP for open-vocabulary audio-visual segmentation with semantic coherence
Yunzhi Zhuge, Mengyuan Zhu, Yizhuang Peng, Lu Zhang 0053, Jin Zhan, Huchuan Lu
Pattern Recognit.7
2026 LiveMatte: Dynamic Scene Background Restoration and Selective Portrait Patch Enhancement
abstract
Real-time and accurate portrait matting in videos is a challenging problem in computer vision research. Recent approaches have explored incorporating prior conditions for accurate inference. Notably, some methods ask the user to provide the background image, which requires extra effort from the user to capture the background image and is limited to videos with static backgrounds only. We note that real-time video motion segmentation methods often train a background model to detect the foreground. Our insight of this work is that if we can directly restore the background content from the input video as a prior, we may be able to achieve more precise portrait matting. In addition, this approach could potentially work even with dynamic backgrounds, without requiring additional user input. However, automatically restoring the background content is not straightforward due to the difficulty in distinguishing between foreground and background. While it may seem that stationary pixel values represent the background, these values can vary across frames. To this end, we propose a novel dynamic scene background restoration (DSBR) module that learns a background model by accumulating background content from each input video frame. It restores the current background content, which serves as a matting prior for alpha prediction of the subsequent frame. DSBR is extremely lightweight and can be easily integrated into existing matting models. Based on it, we present a real-time portrait video matting framework,LiveMatte. To more efficiently process high-resolution videos, we also introduce a selective portrait patch enhancement (SPPE) module. Extensive experiments and user studies demonstrate that our method is better and faster than existing methods.
Zhanghan Ke, Lihe Zhang, Huchuan Lu, Rynson W. H. Lau
IEEE Trans. Circuits Syst. Video Technol.4
2026 What Makes You Unique? Attribute Prompt Composition for Object Re-Identification
Yingquan Wang, Dong Wang 0004, Huchuan Lu
IEEE Trans. Circuits Syst. Video Technol.5
2026 Classification and Calibration: Dual-Guidance Diffusion Model for Mitigating Reconstruction Hallucinations in Multi-Class Anomaly Detection
Yuan Zhao 0006, Xiaoqi Zhao 0003, Lihe Zhang, Huchuan Lu
IEEE Trans. Circuits Syst. Video Technol.5
2026 Parameter-Aware Mamba Model for Multitask Dense Prediction
abstract
Understanding the inter-relations and interactions between tasks is crucial for multitask dense prediction. Existing methods predominantly utilize convolutional layers and attention mechanisms to explore task-level interactions. In this work, we introduce a novel decoder-based framework, parameter-aware Mamba model (PAMM), specifically designed for dense prediction in multitask learning (MTL) setting. Distinct from approaches that employ Transformers to model holistic task relationships, PAMM leverages the rich, scalable parameters of state-space models (SSMs) to enhance task interconnectivity. It features dual state-space parameter experts (PEs) that integrate and set task-specific parameter priors (PPs), capturing the intrinsic properties of each task. This approach not only facilitates precise multitask interactions but also allows for the global integration of task priors through the structured state-space sequence (S4) model. Furthermore, we employ the multidirectional Hilbert scanning (MDHS) method to construct multiangle feature sequences, thereby enhancing the sequence model's perceptual capabilities for 2-D data. Extensive experiments on the NYUD-v2 and PASCAL-Context benchmarks demonstrate the effectiveness of our proposed method. Our code is available at https://github.com/CQC-gogopro/PAMM.
Xinzhuo Yu, Yunzhi Zhuge, Sitong Gong, Lu Zhang 0053, Huchuan Lu
IEEE Trans. Cybern.6
2026 Visual-Textual Information-Driven Tactile Data Generation Method
abstract
Tactile data can enhance the environmental perception and interaction capabilities of intelligent agents, serving as a foundational component for the development of embodied intelligence. Despite its critical role, tactile data acquisition remains cost-prohibitive and labor-intensive, resulting in severe data scarcity. Cross-modal generation offers a promising solution by leveraging abundant visual and textual data. However, effectively aligning heterogeneous visual-textual modalities under data-scarce and sparsely-annotated conditions remains a significant challenge. To address these challenges, a visual-textual information-driven tactile data generation (VTTac) framework is proposed, which features three key innovations. First, a multi-granularity text enhancement strategy is introduced to mitigate annotation sparsity through hierarchical semantic enrichment. Second, a cascaded dual cross-attention mechanism is designed to ensure cross-modal alignment. Third, a condition adapter injects a low-frequency background prior, enabling the generative backbone to focus on high-frequency texture synthesis. Subsequently, a wavelet transform seamlessly fuses these synthesized details with the real background. Extensive evaluations across three datasets demonstrate that VTTac consistently outperforms representative baselines. Furthermore, downstream tasks validate the physical faithfulness of the synthesized data for material classification and semantic reasoning, and zero-shot experiments confirm generalization to unseen objects.
Zhangzheng Tu, Hongchen Tan, Huchuan Lu
IEEE Trans. Image Process.6
2026 SD-ReID: View-Aware Stable Diffusion for Aerial-Ground Person Re-Identification
abstract
Aerial-Ground Person Re-IDentification (AG-ReID) aims to retrieve specific persons across cameras with different viewpoints. Previous works focus on designing discriminative models to maintain the identity consistency despite drastic changes in camera viewpoints. The core idea behind these methods is quite natural, but designing a view-robust model is a very challenging task. Moreover, they overlook the contribution of view-specific features in enhancing the model's ability to represent persons. To address these issues, we propose a novel generative framework named SD-ReID for AG-ReID, which leverages generative models to mimic the feature distribution of different views while extracting robust identity representations. More specifically, we first train a ViT-based model to extract person representations along with controllable conditions, including identity and view conditions. We then fine-tune the Stable Diffusion (SD) model to enhance person representations guided by these controllable conditions. Furthermore, we introduce the View-Refined Decoder (VRD) to bridge the gap between instance-level and global-level features. Finally, both person representations and all-view features are employed to retrieve target persons. Extensive experiments on five AG-ReID benchmarks (i.e., CARGO, AG-ReIDv1, AG-ReIDv2, LAGPeR and G2APS-ReID) demonstrate the effectiveness of our proposed method. The source code and pre-trained models are available at https://github.com/924973292/SD-ReID.
Huchuan Lu
IEEE Trans. Image Process.5
2026 HFP-SAM: Hierarchical Frequency Prompted SAM for Efficient Marine Animal Segmentation
abstract
Marine Animal Segmentation (MAS) aims at identifying and segmenting marine animals from complex marine environments. Most of previous deep learning-based MAS methods struggle with the long-distance modeling issue. Recently, Segment Anything Model (SAM) has gained popularity in general image segmentation. However, it lacks of perceiving fine-grained details and frequency information. To this end, we propose a novel learning framework, named Hierarchical Frequency Prompted SAM (HFP-SAM) for high-performance MAS. First, we design a Frequency Guided Adapter (FGA) to efficiently inject marine scene information into the frozen SAM backbone through frequency domain prior masks. Additionally, we introduce a Frequency-aware Point Selection (FPS) to generate highlighted regions through frequency analysis. These regions are combined with the coarse predictions of SAM to generate point prompts and integrate into SAM's decoder for fine predictions. Finally, to obtain comprehensive segmentation masks, we introduce a Full-View Mamba (FVM) to efficiently extract spatial and channel contextual information with linear computational complexity. Extensive experiments on four public datasets demonstrate the superior performance of our approach. We will make our code publicly available upon the acceptance.
Tianyu Yan, Yang Liu 0066, Tongdan Tang, Yili Ma, Long Lv, Feng Tian 0001, Weibing Sun, Huchuan Lu
IEEE Trans. Image Process.10
2026 DiMuS: Disentangled Multi-Signal Learning for Weakly Supervised Point-Based 3D Object Detection
abstract
Weakly supervised 3D object detection has emerged as a promising paradigm to reduce the reliance on costly 3D annotations. Existing methods often rely on 2D projection constraints or heuristic priors to supervise 3D box regression with inexpensive 2D labels. However, they still suffer from projection ambiguity and geometry inconsistency due to the entangled optimization of 3D parameters. In this paper, we propose DiMuS, a Disentangled Multi- $\boldsymbol {S}$ ignal learning framework that integrates complementary supervision from 2D boxes, LLM-derived semantic prior, and 3D geometric alignment to enhance distinct 3D properties of position, dimension, and orientation, respectively. Specifically, DiMuS incorporates three key components: (i) a Centerness-enhanced Projection Constraint (CPC) that improves position estimation through a centerness weighting strategy, (ii) a Semantic Prior Anchoring (SPA) module that leverages LLM-derived category-specific priors for robust dimension decoding, and (iii) a Rotation-aware Consistency Regularization (RCR) mechanism that enforces orientation consistency through synthetic rotations and self-supervised invariance learning. Additionally, an Adversarial Geometric Alignment (AGA) module is proposed to build attraction/repulsion forces between LiDAR points and box edges for dynamic boundary refinement. Extensive experiments on the KITTI dataset demonstrate that DiMuS outperforms previous weakly supervised methods, achieving 96.82% of fully supervised performance on car detection while maintaining robustness across different categories.
Yunzhi Zhuge, Lu Zhang 0053, Ping Hu 0001, Dong Wang 0004, Huchuan Lu
IEEE Trans. Image Process.6
2026 Multi-Modal Object Re-Identification With Prompt-S6 and Semantic-Aware Knowledge Guidance
abstract
Multi-modal object Re-Identification (ReID) aims to retrieve specific objects by integrating complementary information from multiple modalities. However, existing multi-modal ReID methods do not effectively address background interference suppression or achieve tri-modal alignment, instead focusing on pairwise feature fusion. Moreover, many current aggregation approaches suffer from high computational complexity. To address these limitations, we propose PRISM, a novel multi-modal ReID framework built upon Prompt-S6 (PS6) and semantic-aware knowledge guidance. PS6 maintains the linear complexity and strong sequence modeling capability of Mamba while enabling efficient cross-modal interaction. Leveraging these advantages, we design two key components: Semantic-Driven Token Pruning (SDTP) and Progressive Fusion Network (PFN). Parsing semantic priors from the segmentation foundation models, the SDTP then leverages these priors and applies dynamic token pruning to suppress background noise and refine feature representations. The PFN progressively aggregates multi-modal features to achieve tri-modal alignment and fully exploit modality complementarity. With the proposed modules, PRISM generates more robust multi-modal representations under complex scenarios. Extensive experiments on four multi-modal object ReID benchmarks demonstrate the effectiveness and efficiency of our approach. The source code is available at https://github.com/zw-absin/PRISM.
Weixiang Zhou, Jiabei Zuo, Cong Wang 0018, Huchuan Lu, Zhixun Su
IEEE Trans. Image Process.5
2026 Interactive Spatial-Frequency Fusion Mamba for Multi-Modal Image Fusion
abstract
Multi-Modal Image Fusion (MMIF) aims to combine images from different modalities to produce fused images, retaining texture details and preserving significant information. Recently, some MMIF methods incorporate frequency domain information to enhance spatial features. However, these methods typically rely on simple serial or parallel spatial-frequency fusion without interaction. In this paper, we propose a novel Interactive Spatial-Frequency Fusion Mamba (ISFM) framework for MMIF. Specifically, we begin with a Modality-Specific Extractor (MSE) to extract features from different modalities. It models long-range dependencies across the image with linear computational complexity. To effectively leverage frequency information, we then propose a Multi-scale Frequency Fusion (MFF). It adaptively integrates low-frequency and high-frequency components across multiple scales, enabling robust representations of frequency features. More importantly, we further propose an Interactive Spatial-Frequency Fusion (ISF). It incorporates frequency features to guide spatial features across modalities, enhancing complementary representations. Extensive experiments are conducted on six MMIF datasets. The experimental results demonstrate that our ISFM can achieve better performances than other state-of-the-art methods. The source code is available at https://github.com/Namn23/ISFM.
Long Lv, Xuehu Liu, Tongdan Tang, Feng Tian 0001, Weibing Sun, Huchuan Lu
IEEE Trans. Image Process.8
2026 Context-Infused Trajectories: Enhancing Context and Frame Consistency in Reasoning Video Object Segmentation
abstract
Reasoning video object segmentation (ReaVOS) aims to segment referred objects in video sequences based on implicit and complex linguistic queries. Existing methods typically compress limited video frames into pooled representations and prompt multimodal large language models (MLLMs) to generate a single global segmentation token. However, this strategy lacks explicit contextual guidance and causes substantial loss of spatial details, limiting capability and segmentation consistency. To overcome these limitations, we introduce Context-infused Consistent Video Segmentor (CiCVS), a novel framework leveraging contextual information to guide generation of temporally coherent and accurate mask trajectories. CiCVS incorporates a Hierarchical Frame Sampling (HFS) module, which globally samples support frames across the entire video to ensure broad temporal coverage, and then uniformly selects target frames within the support set. It also employs a Contextual Token Prompting (CTP) module, which utilizes contextual cues from support frames to guide the MLLM in generating specialized tokens for various target frames, enabling the model to capture intricate temporal patterns and ensure consistency across long-range sequences. At the core of CTP is the Multimodal Injection Compressor (MIC) block, which efficiently integrates support frame features and textual semantic information into a compact set of latent queries, enhancing temporal-level object perception. To further advance the ReaVOS field, we introduce the CoCoRVOS benchmark, which features more temporally intricate reasoning instructions and a diverse set of video scenarios. Extensive experiments demonstrate that CiCVS establishes a new state-of-the-art on multiple benchmarks, achieving significant improvements in $\mathcal {J}\& \mathcal {F}$ scores, including +2.7 on CoCoRVOS, +1.4 on ReVOS, and +7.0 on ReasonVOS, underscoring its superior contextual reasoning and segmentation capabilities.
Yunzhi Zhuge, Sitong Gong, Lu Zhang 0053, Qi Xu 0008, Wenda Zhao 0003, Jin Zhan, Huchuan Lu
IEEE Trans. Image Process.7
2026 Exploiting Cross-Task Synergy via Frequency-Driven Hierarchical Learning for Multi-Task Dense Prediction
abstract
Multi-task dense prediction improves pixel-level performance by leveraging shared representations and inter-task collaboration. However, existing approaches either rely on implicit task relationships or neglect frequency-domain cues that are essential for preserving fine-grained details and enhancing cross-task feature learning at multiple scales. As a result, they face persistent challenges in multi-scale feature fusion, effective task interaction, and accurate decoding. To address these issues, we propose a hierarchical frequency-driven framework, termed Hierarchical Frequency-Adaptive Network (HiFAN), that facilitates cross-task collaborative optimization via frequency-domain analysis. Specifically, we first design a task-adaptive fusion module that exploits multi-scale frequency-domain information to enhance spatial details. This module generates dynamic convolutional kernels with task-specific parameters and positional biases to adaptively accommodate diverse task requirements. Next, we introduce an efficient cross-task interaction module that leverages compact low-frequency representations to enable global context exchange across tasks. Finally, we present a high-frequency-aware decoder that mitigates feature smoothing and detail loss commonly introduced by Transformer-based decoders. We demonstrate the effectiveness of HiFAN on two standard multi-task learning benchmarks, PASCAL-Context and NYUD-v2, achieving strong and competitive performance across multiple tasks. The code and model weights are available in HiFAN.
Yunzhi Zhuge, Xinzhuo Yu, Lu Zhang 0053, Xu Jia 0012, Jin Zhan, Huchuan Lu
IEEE Trans. Image Process.6
2026 3D-SceneQ: Empowering 3D LLM With Query-Guided Adaptive Pruning and Multi-Modal Feature Enhancement
abstract
Current 3D scene understanding pipelines typically concatenate vast numbers of 3D object tokens with text tokens and feed the resulting sequence to a Large Language Model (LLM). However, existing 3D-LLMs for scene understanding encounter two major limitations: (i) excessive inclusion of task-irrelevant object data, introducing noise that reduces reasoning accuracy and increases hallucinations; and (ii) reliance solely on point cloud data, which inherently lacks the rich semantic information available in complementary 2D modalities, such as color, material properties, texture, and high-level contextual relationships. To address these challenges, we introduce 3D-SceneQ, a 3D LLM that unifies adaptive token pruning with multimodal semantic enrichment, markedly advancing scene understanding, reasoning, and grounding. Specifically, we propose a Query-Guided Adaptive Pruning (QGAP) module that adaptively selects task-relevant objects based on the intended meaning of user instructions, leveraging a learnable latent query to retain critical information while effectively suppressing irrelevant noise. In addition, we introduce a Multi-modal Object-level Feature Enhancement (MOFE) module, which integrates 2D feature embeddings from pre-trained models into 3D representation to enrich object-level semantic information. Evaluations show that 3D-SceneQ outperforms current state-of-the-art methods across multiple benchmarks (e.g., +5.7 CIDEr on Scan2Cap and +17.0 EM@1 on SQA3D), demonstrating its remarkable capabilities in various 3D language reasoning tasks.
Yongqi Shan, Lu Zhang 0053, Jiazuo Yu 0001, Yunzhi Zhuge, Huchuan Lu
IEEE Trans. Multim.5
2025 SUTrack: Towards Simple and Unified Single Object Tracking
abstract
In this paper, we propose a simple yet unified single object tracking (SOT) framework, dubbed SUTrack. It consolidates five SOT tasks (RGB-based, RGB-Depth, RGB-Thermal, RGB-Event, RGB-Language Tracking) into a unified model trained in a single session. Due to the distinct nature of the data, current methods typically design individual architectures and train separate models for each task. This fragmentation results in redundant training processes, repetitive technological innovations, and limited cross-modal knowledge sharing. In contrast, SUTrack demonstrates that a single model with a unified input representation can effectively handle various SOT tasks, eliminating the need for task-specific designs and separate training sessions. Additionally, we introduce a task-recognition training strategy and a soft token type embedding to further enhance SUTrack's performance with minimal overhead. Experiments show that SUTrack outperforms previous task-specific counterparts across 11 datasets spanning five SOT tasks. Moreover, we provide a range of models catering edge devices as well as high-performance GPUs, striking a good trade-off between speed and accuracy. We hope SUTrack could serve as a strong foundation for further compelling research into unified tracking models.
Xin Chen 0032, Ben Kang, Wanting Geng, Jiawen Zhu 0003, Dong Wang 0004, Huchuan Lu
AAAI7
2025 MambaPro: Multi-Modal Object Re-identification with Mamba Aggregation and Synergistic Prompt
abstract
Multi-modal object Re-IDentification (ReID) aims to retrieve specific objects by utilizing complementary image information from different modalities. Recently, large-scale pre-trained models like CLIP have demonstrated impressive performance in traditional single-modal ReID tasks. However, they remain unexplored for multi-modal object ReID. Furthermore, current multi-modal aggregation methods have obvious limitations in dealing with long sequences from different modalities. To address above issues, we introduce a novel framework called MambaPro for multi-modal object ReID. To be specific, we first employ a Parallel Feed-Forward Adapter (PFA) for adapting CLIP to multi-modal object ReID. Then, we propose the Synergistic Residual Prompt (SRP) to guide the joint learning of multi-modal features. Finally, leveraging Mamba's superior scalability for long sequences, we introduce Mamba Aggregation (MA) to efficiently model interactions between different modalities. As a result, MambaPro could extract more robust features with lower complexity. Extensive experiments on three multi-modal object ReID benchmarks (i.e., RGBNT201, RGBNT100 and MSVR310) validate the effectiveness of our proposed methods.
Xuehu Liu, Tianyu Yan, Aihua Zheng, Huchuan Lu
AAAI7
2025 CLIMB-ReID: A Hybrid CLIP-Mamba Framework for Person Re-Identification
abstract
Person Re-IDentification (ReID) aims to identify specific persons from non-overlapping cameras. Recently, some works have suggested using large-scale pre-trained vision-language models like CLIP to boost ReID performance. Unfortunately, existing methods still struggle to address two key issues simultaneously: efficiently transferring the knowledge learned from CLIP and comprehensively extracting the context information from images or videos. To address these issues, we introduce CLIMB-ReID, a pioneering hybrid framework that synergizes the impressive power of CLIP with the remarkable computational efficiency of Mamba. Specifically, we first propose a novel Multi-Memory Collaboration (MMC) strategy to transfer CLIP's knowledge in a parameter-free and prompt-free form. Then, we design a Multi-Temporal Mamba (MTM) to capture multi-granular spatiotemporal information in videos. Finally, with Importance-aware Reorder Mamba (IRM), information from various scales is combined to produce robust sequence features. Extensive experiments show that our proposed method outperforms other state-of-the-art methods on both image and video person ReID benchmarks.
Chenyang Yu, Xuehu Liu, Jiawen Zhu 0003, Huchuan Lu
AAAI6
2025 Bootstraping Clustering of Gaussians for View-consistent 3D Scene Understanding
abstract
Injecting semantics into 3D Gaussian Splatting (3DGS) has recently garnered significant attention. While current approaches typically distill 3D semantic features from 2D foundational models (e.g., CLIP and SAM) to facilitate novel view segmentation and semantic understanding, their heavy reliance on 2D supervision can undermine cross-view semantic consistency and necessitate complex data preparation processes, therefore hindering view-consistent scene understanding. In this work, we present FreeGS, an unsupervised semantic-embedded 3DGS framework that achieves view-consistent 3D scene understanding without the need for 2D labels. Instead of directly learning semantic features, we introduce the IDentity-coupled Semantic Field (IDSF) into 3DGS, which captures both semantic representations and view-consistent instance indices for each Gaussian. We optimize IDSF with a two-step alternating strategy: semantics help to extract coherent instances in 3D space, while the resulting instances regularize the injection of stable semantics from 2D space. Additionally, we adopt a 2D-3D joint contrastive loss to enhance the complementarity between view-consistent 3D geometry and rich semantics during the bootstrapping process, enabling FreeGS to uniformly perform tasks such as novel-view semantic segmentation, object selection, and 3D object detection. Extensive experiments on LERF-Mask, 3D-OVS, and ScanNet datasets demonstrate that FreeGS performs comparably to state-of-the-art methods while avoiding the complex data preprocessing workload.
Lu Zhang 0053, Ping Hu 0001, Liqian Ma, Yunzhi Zhuge, Huchuan Lu
AAAI6
2025 Two-stream Beats One-stream: Asymmetric Siamese Network for Efficient Visual Tracking
abstract
Efficient tracking has garnered attention for its ability to operate on resource-constrained platforms for real-world deployment beyond desktop GPUs. Current efficient trackers mainly follow precision-oriented trackers, adopting a one-stream framework with lightweight modules. However, blindly adhering to the one-stream paradigm may not be optimal, as incorporating template computation in every frame leads to redundancy, and pervasive semantic interaction between template and search region places stress on edge devices. In this work, we propose a novel asymmetric Siamese tracker named AsymTrack for efficient tracking. AsymTrack disentangles template and search streams into separate branches, with template computing only once during initialization to generate modulation signals. Building on this architecture, we devise an efficient template modulation mechanism to unidirectional inject crucial cues into the search features, and design an object perception enhancement module that integrates abstract semantics and local details to overcome the limited representation in lightweight tracker. Extensive experiments demonstrate that AsymTrack offers superior speed-precision trade-offs across different platforms compared to the current state-of-the-arts. For instance, AsymTrack-T achieves 60.8% AUC on LaSOT and 224/81/84 FPS on GPU/CPU/AGX, surpassing HiT-Tiny by 6.0% AUC with higher speeds.
Jiawen Zhu 0003, Huayi Tang, Xin Chen 0032, Xinying Wang 0005, Dong Wang 0004, Huchuan Lu
AAAI6
2025 Feature Selection and Dual Perturbation: Synergistic Approaches for Semi-Supervised Polyp Segmentation
abstract
Semi-supervised polyp segmentation (SSPS) is crucial for the computer-aided diagnosis of colorectal cancer, as it reduces the reliance on extensive labeled data. Although previous SSPS methods have achieved notable success, further investigation into critical issues remains necessary. In this paper, we focus on two key challenges in polyp segmentation: First, how can models extract discriminative features effectively? Second, how can we design a mechanism to suppress cognitive bias in SSPS? To address these, we propose a Feature Selection and Dual Perturbation Network integrating a feature selection and interaction module (FSIM) and a dual perturbation mechanism. The FSIM selects texture-rich information and aggregates discriminative features. Inspired by the immune response, the dual perturbation mechanism employs non-shared parameters to independently handle input and network perturbations. Moreover, a constrained loss function encourages effective collaboration among network components, enhancing robustness and reducing cognitive bias. Extensive experiments on multiple polyp datasets demonstrate our method consistently outperforms state-of-the-art SSPS approaches. Remarkably, even when using only 50 % of the labeled data, our approach surpasses several advanced fully supervised models.
Miao Zhang 0004, Yidi Tang, Zhuangze Hou, Weibing Sun, Yongri Piao, Huchuan Lu
BIBM7
2025 The Devil is in Temporal Token: High Quality Video Reasoning Segmentation
abstract
Existing methods for Video Reasoning Segmentation rely heavily on a single special token to represent the object in the keyframe or the entire video, inadequately capturing spatial complexity and inter-frame motion. To overcome these challenges, we propose VRS-HQ, an end-to-end video reasoning segmentation approach that leverages Multimodal Large Language Models (MLLMs) to inject rich spatiotemporal features into hierarchical tokens. Our key innovations include a Temporal Dynamic Aggregation (TDA) and a Token-driven Keyframe Selection (TKS). Specifically, we design frame-leveland temporal-leveltokens that utilize MLLM’s autoregressive learning to effectively capture both local and global information. Subsequently, we apply a similarity-based weighted fusion and frame selection strategy, then utilize SAM2 to perform keyframe segmentation and propagation. To enhance keyframe localization accuracy, the TKS filters keyframes based on SAM2’s occlusion scores during inference. VRSHQ achieves state-of-the-art performance on ReVOS, surpassing VISA by 5.9%/12.5%/9.1% in ${\mathcal{J}}{{ \& }}{\mathcal{F}}$ scores across the three subsets. These results highlight the strong temporal reasoning and segmentation capabilities of our method. Code and model weights are available at VRS-HQ.
Sitong Gong, Yunzhi Zhuge, Lu Zhang 0053, Zongxin Yang, Huchuan Lu
CVPR6
2025 ReNeg: Learning Negative Embedding with Reward Guidance
abstract
In text-to-image (T2I) generation applications, negative embeddings have proven to be a simple yet effective approach for enhancing generation quality. Typically, these negative embeddings are derived from user-defined negative prompts, which, while being functional, are not necessarily optimal. In this paper, we introduce ReNeg, an end-to-end method designed to learn improved Negative embeddings guided by a Reward model. We employ a reward feedback learning framework and integrate classifier-free guidance (CFG) into the training process, which was previously utilized only during inference, thus enabling the effec tive learning of negative embeddings. We also propose two strategies for learning both global and per-sample negative embeddings. Extensive experiments show that the learned negative embedding significantly outperforms null-text and handcrafted counterparts, achieving substantial improvements in human preference alignment. Additionally, the negative embedding learned within the same text embedding space exhibits strong generalization capabilities. For example, using the same CLIP text encoder, the negative embedding learned on SD1.5 can be seamlessly transferred to text-to-image or even text-to-video models such as ControlNet, ZeroScope, and VideoCrafter2, resulting in consistent performance improvements across the board. Code is available at https://github.com/AMD-AIG-AIMA/ReNeg.
Xiaomin Li 0001, Yixuan Liu 0004, Takashi Isobe, Xu Jia 0012, Qinpeng Cui, Dong Zhou 0003, Dong Li 0025, You He 0002, Huchuan Lu, Zhongdao Wang, Emad Barsoum
CVPR9
2025 DefMamba: Deformable Visual State Space Model
abstract
Recently, state space models (SSM), particularly Mamba, have attracted significant attention from scholars due to their ability to effectively balance computational efficiency and performance. However, most existing visual Mamba methods flatten images into 1D sequences using predefined scan orders, which results the model being less capable of utilizing the spatial structural information of the image during the feature extraction process. To address this issue, we proposed a novel visual foundation model called Def-Mamba. This model includes a multi-scale backbone structure and deformable mamba (DM) blocks, which dynamically adjust the scanning path to prioritize important information, thus enhancing the capture and processing of relevant input features. By combining a deformable scanning (DS) strategy, this model significantly improves its ability to learn image structures and detects changes in object details. Numerous experiments have shown that Def-Mamba achieves state-of-the-art performance in various visual tasks, including image classification, object detection, instance segmentation, and semantic segmentation. The code is open source on DefMamba .
Leiye Liu, Miao Zhang 0004, Jihao Yin, Tingwei Liu, Wei Ji 0011, Yongri Piao, Huchuan Lu
CVPR7
2025 IDEA: Inverted Text with Cooperative Deformable Aggregation for Multi-modal Object Re-Identification
abstract
Multi-modal object Re-IDentification (ReID) aims to retrieve specific objects by utilizing complementary information from various modalities. However, existing methods focus on fusing heterogeneous visual features, neglecting the potential benefits of text-based semantic information. To address this issue, we first construct three text-enhanced multi-modal object ReID benchmarks. To be specific, we propose a standardized multi-modal caption generation pipeline for structured and concise text annotations with Multi-modal Large Language Models (MLLMs). Besides, current methods often directly aggregate multi-modal information without selecting representative local features, leading to redundancy and high complexity. To address the above issues, we introduce IDEA, a novel feature learning framework comprising the Inverted Multi-modal Feature Extractor (IMFE) and Cooperative Deformable Aggregation (CDA). The IMFE utilizes Modal Prefixes and an InverseNet to integrate multi-modal information with semantic guidance from inverted text. The CDA adaptively generates sampling positions, enabling the model to focus on the interplay between global features and discriminative local features. With the constructed benchmarks and the proposed modules, our framework can generate more robust multi-modal features under complex scenarios. Extensive experiments on three multi-modal object ReID benchmarks demonstrate the effectiveness of our proposed method.
Yongfeng Lv, Huchuan Lu
CVPR4
2025 Mono2Stereo: A Benchmark and Empirical Study for Stereo Conversion
abstract
With the rapid proliferation of 3D devices and the shortage of 3D content, stereo conversion is attracting increasing attention. Recent works introduce pretrained Diffusion Models (DMs) into this task. However, due to the scarcity of large-scale training data and comprehensive benchmarks, the optimal methodologies for employing DMs in stereo conversion and the accurate evaluation of stereo effects remain largely unexplored. In this work, we introduce the Mono2Stereo dataset, providing high-quality training data and benchmark to support in-depth exploration of stereo conversion. With this dataset, we conduct an empirical study that yields two primary findings. 1) The differences between the left and right views are subtle, yet existing metrics consider overall pixels, failing to concentrate on regions critical to stereo effects. 2) Mainstream methods adopt either one-stage left-to-right generation or warp-and-inpaint pipeline, facing challenges of degraded stereo effect and image distortion respectively. Based on these findings, we introduce a new evaluation metric, Stereo Intersection-over-Union, which prioritizes disparity and achieves a high correlation with human judgments on stereo effect. Moreover, we propose a strong baseline model, harmonizing the stereo effect and image quality simultaneously, and notably surpassing current mainstream methods. Our code and data will be open-sourced to promote further research in stereo conversion. Our models are available at mono2stereo-bench.github.io.
Songsong Yu, Zhongang Qi, Zeke Xie, Yifan Wang 0004, Lijun Wang 0001, Ying Shan, Huchuan Lu
CVPR8
2025 KARST: Multi-Kernel Kronecker Adaptation with Re-Scaling Transmission for Visual Classification
abstract
Fine-tuning pre-trained vision models for specific tasks is a common practice in computer vision. However, this process becomes more expensive and resource-intensive as models grow larger. Recently, parameter-efficient fine-tuning (PEFT) methods have emerged as a popular solution to improve training efficiency and reduce storage needs by tuning additional low-rank modules within pre-trained backbones. Despite their advantages, they struggle with limited representation capabilities and misalignment with pre-trained intermediate features. To address these issues, we introduce an innovative Multi-Kernel Kronecker Adaptation with Re-Scaling Transmission (KARST) for various recognition tasks. Specifically, KARST’s multi-kernel design extends Kronecker projections horizontally and separates adaptation matrices into multiple complementary spaces, reducing parameter dependency and creating more compact subspaces. Besides, it incorporates extra learnable re-scaling factors to better align with pre-trained feature distributions, allowing for more flexible and balanced feature aggregation. Extensive experiments on diverse downstream datasets validate that our KARST not only outperforms other PEFT counterparts across model types and data domains, but also surpasses full fine-tuning with a negligible inference cost due to its re-parameterization characteristics.
Yue Zhu 0012, Haiwen Diao, Shang Gao 0012, Long Chen 0016, Huchuan Lu
ICASSP5
2025 Hierarchical Proxy Learning for Cloth-Changing Person Re-Identification
abstract
Cloth-Changing person Re-Identification (CC-ReID) depends significantly on learning discriminative features under the cloth-changing scenario. It is quite challenging due to the large intra-person variance and small inter-person variance caused by clothes changing. To address these issues, in this work we propose a Hierarchical Proxy Learning (HPL) framework to extract clothes-irrelevant and person-invariant features. Specifically, we employ person labels as the main proxy. Instead of leveraging clothing labels as sub proxy, we further propose a clustering-based automatic sub-proxy mining scheme. More specifically, we first construct a person-aware Main Proxy Learning (MPL) to improve the separability of different persons. Then, a Sub Proxy Learning (SPL) is constructed to enhance the intra-person compactness. Finally, a Sub-to-Main Proxy Learning (S2MPL) is proposed to promote the cooperation between the main proxies and sub proxies. In addition, to weed out the negative effect of clothes, we propose a Sample Balance and Diversity (SBD) module, which balances the number of sub proxies in a mini-batch and utilizes semantic guidance to enrich the diversity of clothes, simultaneously. Extensive experiments on two public CC-ReID datasets demonstrate the superiority of our proposed method over most state-of-the-art methods.
Chenyang Yu, Xuehu Liu, Ju Dai, Huchuan Lu
ICASSP5
2025 TrackFusion: Enhancing Multi-Object Tracking With Temporal Trajectory Modeling and Frame-Integrated Detection
abstract
Although MOTIP is the SOTA multi-object tracking method, there are still some issues that limit its performance. First, MOTIP still has defects in temporal information modeling, which leads to the failure to fully utilize the historical information of the tracked target and affects the correlation performance of the model. Second, in MOT, objects in consecutive video frames usually have temporal continuity and spatial consistency. Therefore, the object information of the previous frame can effectively assist the detection of the current frame. However, MOTIP performs independent detection between each frame, which does not fully utilize the correlation information between frames, resulting in suboptimal model performance. To address the above problems, we propose TrackFusion, which optimizes model performance from the perspective of trajectory modeling and inter-frame joint detection. First, we extract embeddings in video sequences through a Transformer-based detector, then combine the embeddings of the same object in different frames into sequences and input them into the trajectory modeling module for sequence association. This strategy effectively enhances the association ability. Thanks to these improvements, TrackFusion’s HOTA on the DanceTrack test set reached 68.6%, an increase of 1.1% compared to MOTIP’s 67.5%.
Shuai Liu 0009, Bingyang Wang, Jiaojiao Dai, Jinqing Qi, Huchuan Lu, You He 0002
ICASSP6
2025 EVEv2: Improved Baselines for Encoder-Free Vision-Language Models
abstract
Existing encoder-free vision-language models (VLMs) are rapidly narrowing the performance gap with their encoder-based counterparts, highlighting the promising potential for unified multimodal systems with structural simplicity and efficient deployment. We systematically clarify the performance gap between VLMs using pre-trained vision encoders, discrete tokenizers, and minimalist visual layers from scratch, deeply excavating the under-examined characteristics of encoder-free VLMs. We develop efficient strategies for encoder-free VLMs that rival mainstream encoder-based ones. After an in-depth investigation, we launch EVEv2.0, a new and improved family of encoder-free VLMs. We show that: (i) Properly decomposing and hierarchically associating vision and language within a unified model reduces interference between modalities. (ii) A well-designed training strategy enables effective optimization for encoder-free VLMs. Through extensive evaluation, our EVEv2.0 represents a thorough study for developing a decoder-only architecture across modalities, demonstrating superior data efficiency and strong vision-reasoning capability. Code is publicly available at: https://github.com/baaivision/EVE.
Haiwen Diao, Yufeng Cui, Yueze Wang, Haoge Deng, Wenxuan Wang 0002, Huchuan Lu
ICCV8
2025 CCL-LGS: Contrastive Codebook Learning for 3D Language Gaussian Splatting
abstract
Recent advances in 3D reconstruction techniques and vision-language models have fueled significant progress in 3D semantic understanding, a capability critical to robotics, autonomous driving, and virtual/augmented reality. However, methods that rely on 2D priors are prone to a critical challenge: cross-view semantic inconsistencies induced by occlusion, image blur, and view-dependent variations. These inconsistencies, when propagated via projection supervision, deteriorate the quality of 3D Gaussian semantic fields and introduce artifacts in the rendered outputs. To mitigate this limitation, we propose CCL-LGS, a novel framework that enforces view-consistent semantic supervision by integrating multi-view semantic cues. Specifically, our approach first employs a zero-shot tracker to align a set of SAM-generated 2D masks and reliably identify their corresponding categories. Next, we utilize CLIP to extract robust semantic encodings across views. Finally, our Contrastive Codebook Learning (CCL) module distills discriminative semantic features by enforcing intra-class compactness and inter-class distinctiveness. In contrast to previous methods that directly apply CLIP to imperfect masks, our framework explicitly resolves semantic conflicts while preserving category discriminability. Extensive experiments demonstrate that CCL-LGS outperforms previous state-of-the-art methods. Our project page is available at https://epsilontl.github.io/CCL-LGS/.
Xiaomin Li 0001, Liqian Ma, Zirui Zheng, Hefei Huang, Taiqing Li, Huchuan Lu, Xu Jia 0012
ICCV8
2025 VLIPP: Towards Physically Plausible Video Generation with Vision and Language Informed Physical Prior
abstract
Video diffusion models (VDMs) have advanced significantly in recent years, enabling the generation of highly realistic videos and drawing the attention of the community in their potential as world simulators. However, despite their capabilities, VDMs often fail to produce physically plausible videos due to an inherent lack of understanding of physics, resulting in incorrect dynamics and event sequences. To address this limitation, we propose a novel two-stage image-to-video generation framework that explicitly incorporates physics with vision and language informed physical prior. In the first stage, we employ a Vision Language Model (VLM) as a coarse-grained motion planner, integrating chain-of-thought and physics-aware reasoning to predict a rough motion trajectories/changes that approximate real-world physical dynamics while ensuring the inter-frame consistency. In the second stage, we use the predicted motion trajectories/changes to guide the video generation of a VDM. As the predicted motion trajectories/changes are rough, noise is added during inference to provide freedom to the VDM in generating motion with more fine details. Extensive experimental results demonstrate that our framework can produce physically plausible motion, and comparative evaluations highlight the notable superiority of our approach over existing methods. More video results are available on our Project Page: https://madaoer.github.io/projects/physically_plausible_video_generation.
Xindi Yang, Baolu Li 0001, Zhenfei Yin, Lei Bai 0001, Liqian Ma, Zhiyong Wang 0001, Jianfei Cai 0001, Tien-Tsin Wong, Huchuan Lu, Xu Jia 0012
ICCV10
2025 CAT: A Unified Click-and-Track Framework for Realistic Tracking
Yongsheng Yuan, Jie Zhao 0014, Dong Wang 0004, Huchuan Lu
ICCV4
2025 Learning Spatial-Semantic Features for Robust Video Object Segmentation
abstract
Tracking and segmenting multiple similar objects with distinct or complex parts in long-term videos is particularly challenging due to the ambiguity in identifying target components and the confusion caused by occlusion, background clutter, and changes in appearance or environment over time. In this paper, we propose a robust video object segmentation framework that learns spatial-semantic features and discriminative object queries to address the above issues. Specifically, we construct a spatial-semantic block comprising a semantic embedding component and a spatial dependency modeling part for associating global semantic features and local spatial features, providing a comprehensive target representation. In addition, we develop a masked cross-attention module to generate object queries that focus on the most discriminative parts of target objects during query propagation, alleviating noise accumulation to ensure effective long-term query propagation. The experimental results show that the proposed method sets new state-of-the-art performance on multiple data sets, including the DAVIS2017 test (\textbf{87.8\%}), YoutubeVOS 2019 (\textbf{88.1\%}), MOSE val (\textbf{74.0\%}), and LVOS test (\textbf{73.0\%}), which demonstrate the effectiveness and generalization capacity of the proposed method. We will make all the source code and trained models publicly available.
Xin Li 0034, Deshui Miao, Zhenyu He 0001, Yaowei Wang 0001, Huchuan Lu, Ming-Hsuan Yang 0001
ICLR5
2025 Autoregressive Video Generation without Vector Quantization
abstract
This paper presents a novel approach that enables autoregressive video generation with high efficiency. We propose to reformulate the video generation problem as a non-quantized autoregressive modeling of temporal frame-by-frame prediction and spatial set-by-set prediction. Unlike raster-scan prediction in prior autoregressive models or joint distribution modeling of fixed-length tokens in diffusion models, our approach maintains the causal property of GPT-style models for flexible in-context capabilities, while leveraging bidirectional modeling within individual frames for efficiency. With the proposed approach, we train a novel video autoregressive model without vector quantization, termed NOVA. Our results demonstrate that NOVA surpasses prior autoregressive video models in data efficiency, inference speed, visual fidelity, and video fluency, even with a much smaller model capacity, i.e., 0.6B parameters. NOVA also outperforms state-of-the-art image diffusion models in text-to-image generation tasks, with a significantly lower training cost. Additionally, NOVA generalizes well across extended video durations and enables diverse zero-shot applications in one unified model. Code and models are publicly available at https://github.com/baaivision/NOVA.
Haoge Deng, Haiwen Diao, Yufeng Cui, Huchuan Lu, Shiguang Shan, Yonggang Qi
ICLR6
2025 Streaming Video Understanding and Multi-round Interaction with Memory-enhanced Knowledge
abstract
Recent advances in Large Language Models (LLMs) have enabled the development of Video-LLMs, advancing multimodal learning by bridging video data with language tasks. However, current video understanding models struggle with processing long video sequences, supporting multi-turn dialogues, and adapting to real-world dynamic scenarios. To address these issues, we propose StreamChat, a training-free framework for streaming video reasoning and conversational interaction. StreamChat leverages a novel hierarchical memory system to efficiently process and compress video features over extended sequences, enabling real-time, multi-turn dialogue. Our framework incorporates a parallel system scheduling strategy that enhances processing speed and reduces latency, ensuring robust performance in real-world applications. Furthermore, we introduce StreamBench, a versatile benchmark that evaluates streaming video understanding across diverse media types and interactive scenarios, including multi-turn interactions and complex reasoning tasks. Extensive evaluations on StreamBench and other public benchmarks demonstrate that StreamChat significantly outperforms existing state-of-the-art models in terms of accuracy and response times, confirming its effectiveness for streaming video understanding. Code is available at StreamChat.
Haomiao Xiong, Zongxin Yang, Jiazuo Yu 0001, Yunzhi Zhuge, Lu Zhang 0053, Jiawen Zhu 0003, Huchuan Lu
ICLR7
2025 High-Precision Dichotomous Image Segmentation via Probing Diffusion Capacity
abstract
In the realm of high-resolution (HR), fine-grained image segmentation, the primary challenge is balancing broad contextual awareness with the precision required for detailed object delineation, capturing intricate details and the finest edges of objects. Diffusion models, trained on vast datasets comprising billions of image-text pairs, such as SD V2.1, have revolutionized text-to-image synthesis by delivering exceptional quality, fine detail resolution, and strong contextual awareness, making them an attractive solution for high-resolution image segmentation. To this end, we propose DiffDIS, a diffusion-driven segmentation model that taps into the potential of the pre-trained U-Net within diffusion models, specifically designed for high-resolution, fine-grained object segmentation. By leveraging the robust generalization capabilities and rich, versatile image representation prior of the SD models, coupled with a task-specific stable one-step denoising approach, we significantly reduce the inference time while preserving high-fidelity, detailed generation. Additionally, we introduce an auxiliary edge generation task to not only enhance the preservation of fine details of the object boundaries, but reconcile the probabilistic nature of diffusion with the deterministic demands of segmentation. With these refined strategies in place, DiffDIS serves as a rapid object mask generation model, specifically optimized for generating detailed binary maps at high resolutions, while demonstrating impressive accuracy and swift processing. Experiments on the DIS5K dataset demonstrate the superiority of DiffDIS, achieving state-of-the-art results through a streamlined inference process. The source code will be publicly available at \href{https://github.com/qianyu-dlut/DiffDIS}{DiffDIS}.
Qian Yu 0015, Peng-Tao Jiang, Hao Zhang 0063, Jinwei Chen 0003, Bo Li 0115, Lihe Zhang, Huchuan Lu
ICLR7
2025 FDAVS: Exploring Frequency-Driven Modality Enhancement in Audio-Visual Segmentation
abstract
Audio-Visual Segmentation (AVS) aims to identify and delineate sounding objects within a video stream guided by auditory cues. Existing research focuses on audio-visual interactions in the spatial domain while overlooking the intrinsic frequency-domain information, leading to insufficient multimodal feature integration and alignment. In response, we introduce FDAVS, which incorporates frequency-driven designs to achieve synergistic utilization of spatial and frequency information of multimodal features. To begin with, we propose Frequency-Oriented Audio Integration (FOAI), which decouples and optimizes the high- and low-frequency components of visual features within the image encoder while integrating audio signals. Subsequently, we employ Frequency-Based Cross-modal Fusion (FBCF) between the pixel decoder and mask decoder to enhance the guidance of multimodal frequency information on the visual modality, increasing the consistency of audiovisual features. FDAVS achieves state-of-the-art segmentation performance across three AVS benchmarks, highlighting the effectiveness of frequency-driven modality enhancement.
Mengyuan Zhu, Yunzhi Zhuge, Sitong Gong, Lu Zhang 0053, Huchuan Lu
ICME5
2025 Efficient Motion Prompt Learning for Robust Visual Tracking
abstract
Due to the challenges of processing temporal information, most trackers depend solely on visual discriminability and overlook the unique temporal coherence of video data. In this paper, we propose a lightweight and plug-and-play motion prompt tracking method. It can be easily integrated into existing vision-based trackers to build a joint tracking framework leveraging both motion and vision cues, thereby achieving robust tracking through efficient prompt learning. A motion encoder with three different positional encodings is proposed to encode the long-term motion trajectory into the visual embedding space, while a fusion decoder and an adaptive weight mechanism are designed to dynamically fuse visual and motion features. We integrate our motion module into three different trackers with five models in total. Experiments on seven challenging tracking benchmarks demonstrate that the proposed motion module significantly improves the robustness of vision-based trackers, with minimal training costs and negligible speed sacrifice. Code is available at https://github.com/zj5559/Motion-Prompt-Tracking.
Jie Zhao 0014, Xin Chen 0032, Yongsheng Yuan, Michael Felsberg, Dong Wang 0004, Huchuan Lu
ICML6
2025 GFM-Planner: Perception-Aware Trajectory Planning with Geometric Feature Metric
abstract
Like humans who rely on landmarks for orientation, autonomous robots depend on feature-rich environments for accurate localization. In this paper, we propose the GFM-Planner, a perception-aware trajectory planning framework based on the geometric feature metric, which enhances LiDAR localization accuracy by guiding the robot to avoid degraded areas. First, we derive the Geometric Feature Metric (GFM) from the fundamental LiDAR localization problem. Next, we design a 2D grid-based Metric Encoding Map (MEM) to efficiently store GFM values across the environment. A constant-time decoding algorithm is further proposed to retrieve GFM values for arbitrary poses from the MEM. Finally, we develop a perception-aware trajectory planning algorithm that improves LiDAR localization capabilities by guiding the robot in selecting trajectories through feature-rich areas. Both simulation and real-world experiments demonstrate that our approach enables the robot to actively select trajectories that significantly enhance LiDAR localization accuracy.
Dong Wang 0004, Huchuan Lu
IROS5
2025 UniSegDiff: Boosting Unified Lesion Segmentation via a Staged Diffusion Model
Yilong Hu, Shijie Chang, Lihe Zhang, Feng Tian 0001, Weibing Sun, Huchuan Lu
MICCAI (2)6
2025 GLDesigner: Leveraging Multi-Modal LLMs as Designer for Enhanced Aesthetic Text Glyph Layouts
abstract
Text logo design heavily relies on the creativity and expertise of professional designers, in which arranging element layouts is one of the most important procedures. However, this specific task has received limited attention, often overshadowed by broader layout generation tasks such as document or poster design. In this paper, we propose a Vision-Language Model (VLM)-based framework that generates content-aware text logo layouts by integrating multi-modal inputs with user-defined constraints, enabling more flexible and robust layout generation for real-world applications. We introduce two model techniques that reduce the computational cost for processing multiple glyph images simultaneously, without compromising performance. To support instruction tuning of our model, we construct two extensive text logo datasets that are five times larger than existing public datasets. In addition to geometric annotations (e.g., text masks and character recognition), our datasets include detailed layout descriptions in natural language, enabling the model to reason more effectively in handling complex designs and custom user inputs. Experimental results demonstrate the effectiveness of our proposed framework and datasets, outperforming existing methods on various benchmarks that assess geometric aesthetics and human preferences.
Junwen He, Yifan Wang 0004, Lijun Wang 0001, Huchuan Lu, Chenyang Li 0007, Jin-Peng Lan, Jun-Yan He, Bin Luo 0008, Yifeng Geng
ACM Multimedia4
2025 Regularizing Subspace Redundancy of Low-Rank Adaptation
abstract
Low-Rank Adaptation (LoRA) and its variants have delivered strong capability in Parameter-Efficient Transfer Learning (PETL) by minimizing trainable parameters and benefiting from reparameterization. However, their projection matrices remain unrestricted during training, causing high representation redundancy and diminishing the effectiveness of feature adaptation in the resulting subspaces. While existing methods mitigate this by manually adjusting the rank or implicitly applying channel-wise masks, they lack flexibility and generalize poorly across various datasets and architectures. Hence, we propose ReSoRA, a method that explicitly models redundancy between mapping subspaces and adaptively Regularizes Subspace redundancy of Low-Rank Adaptation. Specifically, it theoretically decomposes the low-rank submatrices into multiple equivalent subspaces and systematically applies de-redundancy constraints to the feature distributions across different projections. Extensive experiments validate that our proposed method consistently facilitates existing state-of-the-art PETL methods across various backbones and datasets in vision-language retrieval and standard visual classification benchmarks. Besides, as a training supervision, ReSoRA can be seamlessly integrated into existing approaches in a plug-and-play manner, with no additional inference costs. Code is publicly available at: https://github.com/Lucenova/ReSoRA.
Yue Zhu 0012, Haiwen Diao, Shang Gao 0012, Jiazuo Yu 0001, Jiawen Zhu 0003, Yunzhi Zhuge, Shuai Hao 0007, Xu Jia 0012, Lu Zhang 0053, Ying Zhang 0021, Huchuan Lu
ACM Multimedia11
2025 Rethinking Evaluation of Infrared Small Target Detection
abstract
As an essential vision task, infrared small target detection (IRSTD) has seen significant advancements through deep learning. However, critical limitations in current evaluation protocols impede further progress. First, existing methods rely on fragmented pixel- and target-level specific metrics, which fails to provide a comprehensive view of model capabilities. Second, an excessive emphasis on overall performance scores obscures crucial error analysis, which is vital for identifying failure modes and improving real-world system performance. Third, the field predominantly adopts dataset-specific training-testing paradigms, hindering the understanding of model robustness and generalization across diverse infrared scenarios. This paper addresses these issues by introducing a hybrid-level metric incorporating pixel- and target-level performance, proposing a systematic error analysis method, and emphasizing the importance of cross-dataset evaluation. These aim to offer a more thorough and rational hierarchical analysis framework, ultimately fostering the development of more effective and robust IRSTD models. An open-source toolkit has be released to facilitate standardized benchmarking.
Youwei Pang, Xiaoqi Zhao 0003, Lihe Zhang, Huchuan Lu, Georges El Fakhri, Xiaofeng Liu 0001, Shijian Lu
NeurIPS4
2025 End-to-End Vision Tokenizer Tuning
abstract
Existing vision tokenization isolates the optimization of vision tokenizers from downstream training, implicitly assuming the visual tokens can generalize well across various tasks, e.g., image generation and visual question answering. The vision tokenizer optimized for low-level reconstruction is agnostic to downstream tasks requiring varied representations and semantics. This decoupled paradigm introduces a critical misalignment: The loss of the vision tokenization can be the representation bottleneck for target tasks. For example, errors in tokenizing text in a given image lead to poor results when recognizing or generating them. To address this, we propose ETT, an end-to-end vision tokenizer tuning approach that enables joint optimization between vision tokenization and target autoregressive tasks. Unlike prior autoregressive models that use only discrete indices from a frozen vision tokenizer, ETT leverages the visual embeddings of the tokenizer codebook, and optimizes the vision tokenizers end-to-end with both reconstruction and caption objectives. ETT can be seamlessly integrated into existing training pipelines with minimal architecture modifications. Our ETT is simple to implement and integrate, without the need to adjust the original codebooks or architectures of the employed large language models. Extensive experiments demonstrate that our proposed end-to-end vision tokenizer tuning unlocks significant performance gains, i.e., 2-6% for multimodal understanding and visual generation tasks compared to frozen tokenizer baselines, while preserving the original reconstruction capability. We hope this very simple and strong method can empower multimodal foundation models besides image generation and understanding.
Wenxuan Wang 0002, Yufeng Cui, Haiwen Diao, Zhuoyan Luo, Huchuan Lu, Jing Liu 0001
NeurIPS6
2025 FineRS: Fine-grained Reasoning and Segmentation of Small Objects with Reinforcement Learning
abstract
Multi-modal Large Language Models (MLLMs) have shown remarkable capabilities across a wide range of vision-language tasks. However, due to the restricted input resolutions, MLLMs face significant challenges in precisely understanding and localizing visual details in high-resolution images---particularly when dealing with extra-small objects embedded in cluttered contexts. To address this issue, we propose FineRS, a two-stage MLLM-based reinforcement learning framework for jointly reasoning and segmenting extremely small objects within high-resolution scenes. FineRS adopts a coarse-to-fine pipeline comprising Global Semantic Exploration (GSE) and Localized Perceptual Refinement (LPR). Specifically, GSE performs instruction-guided reasoning to generate a textural response and a coarse target region, while LPR refines this region to produce an accurate bounding box and segmentation mask. To couple the two stages, we introduce a locate-informed retrospective reward, where LPR's outputs are used to optimize GSE for more robust coarse region exploration. Additionally, we present FineRS-4k, a new dataset for evaluating MLLMs on attribute-level reasoning and pixel-level segmentation on subtle, small-scale targets in complex high-resolution scenes. Experimental results on FineRS-4k and public datasets demonstrate that our method consistently outperforms state-of-the-art MLLM-based approaches on both instruction-guided segmentation and visual reasoning tasks.
Lu Zhang 0053, Jiazuo Yu 0001, Haomiao Xiong, Ping Hu 0001, Yunzhi Zhuge, Huchuan Lu, You He 0002
NeurIPS6
2025 From Forecasting to Planning: Policy World Model for Collaborative State-Action Prediction
abstract
Despite remarkable progress in driving world models, their potential for autonomous systems remains largely untapped: the world models are mostly learned for world simulation and decoupled from trajectory planning. While recent efforts aim to unify world modeling and planning in a single framework, the synergistic facilitation mechanism of world modeling for planning still requires further exploration. In this work, we introduce a new driving paradigm named Policy World Model (PWM), which not only integrates world modeling and trajectory planning within a unified architecture, but is also able to benefit planning using the learned world knowledge through the proposed action-free future state forecasting scheme. Through collaborative state-action prediction, PWM can mimic the human-like anticipatory perception, yielding more reliable planning performance. To facilitate the efficiency of video forecasting, we further introduce a parallel token generation mechanism, equipped with a context-guided tokenizer and an adaptive dynamic focal loss. Despite utilizing only front camera input, our method matches or exceeds state-of-the-art approaches that rely on multi-view and multi-modal inputs. Code will be released at https://github.com/6550Zhao/Policy-World-Model.
Zhida Zhao, Talas Fu, Yifan Wang 0004, Lijun Wang 0001, Huchuan Lu
NeurIPS5
2025 UniMRSeg: Unified Modality-Relax Segmentation via Hierarchical Self-Supervised Compensation
abstract
Multi-modal image segmentation faces real-world deployment challenges from incomplete/corrupted modalities degrading performance. While existing methods address training-inference modality gaps via specialized per-combination models, they introduce high deployment costs by requiring exhaustive model subsets and model-modality matching. In this work, we propose a unified modality-relax segmentation network (UniMRSeg) through hierarchical self-supervised compensation (HSSC). Our approach hierarchically bridges representation gaps between complete and incomplete modalities across input, feature and output levels. First, we adopt modality reconstruction with the hybrid shuffled-masking augmentation, encouraging the model to learn the intrinsic modality characteristics and generate meaningful representations for missing modalities through cross-modal fusion. Next, modality-invariant contrastive learning implicitly compensates the feature space distance among incomplete-complete modality pairs. Furthermore, the proposed lightweight reverse attention adapter explicitly compensates for the weak perceptual semantics in the frozen encoder. Last, UniMRSeg is fine-tuned under the hybrid consistency constraint to ensure stable prediction under all modality combinations without large performance fluctuations. Without bells and whistles, UniMRSeg significantly outperforms the state-of-the-art methods under diverse missing modality scenarios on MRI-based brain tumor segmentation, RGB-D semantic segmentation, RGB-D/T salient object segmentation. The code will be released at \url{https://github.com/Xiaoqi-Zhao-DLUT/UniMRSeg}.
Xiaoqi Zhao 0003, Youwei Pang, Chenyang Yu, Lihe Zhang, Huchuan Lu, Shijian Lu, Georges El Fakhri, Xiaofeng Liu 0001
NeurIPS5
2025 Bidirectional Spatial Semantics Correlation for Referring Image Segmentation
Lihe Zhang, Huchuan Lu
PRCV (5)4
2025 Towards Real-Time Open-Vocabulary Video Instance Segmentation
Bin Yan 0004, Martin Sundermeyer, David Joseph Tan, Huchuan Lu, Federico Tombari
WACV4
2025 Automated Evaluation of Large Vision-Language Models on Self-Driving Corner Cases
abstract
Large Vision-Language Models (LVLMs) have received widespread attentions for advancing the interpretable self-driving. Existing evaluations of LVLMs primarily focus on multi-faceted capabilities in natural circumstances, lacking automated and quantifiable assessment for self-driving, let alone the severe road corner cases. In this work, we propose CODA-LM, the very first benchmark for the automatic evaluation of LVLMs for self-driving corner cases. We adopt a hierarchical data structure and prompt powerful LVLMs to analyze complex driving scenes and generate high-quality pre-annotations for the human annotators, while for LVLM evaluation, we show that using the text-only large language models (LLMs) as judges reveals even better alignment with human preferences than the LVLM judges. Moreover, with our CODA-LM, we build CODA-VLM, a new driving LVLM surpassing all open-sourced counterparts on CODA-LM. Our CODA-VLM performs comparably with GPT-4V, even surpassing GPT-4V by +21.42% on the regional perception task. We hope CODA-LM can become the catalyst to promote interpretable self-driving empowered by LVLMs.
Kai Chen 0023, Yanxin Liu, Ruiyuan Gao 0001, Lanqing Hong, Xinhai Zhao, Zhenguo Li, Dit-Yan Yeung, Huchuan Lu, Xu Jia 0012
WACV12
2025 TrackDiffusion: Tracklet-Conditioned Video Generation via Diffusion Models
abstract
Despite remarkable achievements in video synthesis, achieving granular control over complex dynamics, such as nuanced movement among multiple interacting objects, still presents a significant hurdle for dynamic world modeling, compounded by the necessity to manage appearance and disappearance, drastic scale changes, and ensure consistency for instances across frames. These challenges hinder the development of video generation that can faithfully mimic real-world complexity, limiting utility for applications requiring high-level realism and controllability, including advanced scene simulation and training of perception systems. To address that, we propose TrackDiffusion, a novel video generation framework affording fine-grained trajectory-conditioned motion control via diffusion models, which facilitates the precise manipulation of the object trajectories and interactions, overcoming the prevalent limitation of scale and continuity disruptions. A pivotal component of TrackDiffusion is the instance enhancer, which explicitly ensures inter-frame consistency of multiple objects, a critical factor overlooked in the current literature. More-over, we demonstrate that generated video sequences by our TrackDiffusion can be used as training data for visual per-ception models. To the best of our knowledge, this is the first work to apply video diffusion models with tracklet conditions and demonstrate that generated frames can be beneficial for improving the performance of object trackers. 1
Kai Chen 0023, Zhili Liu, Ruiyuan Gao 0001, Lanqing Hong, Dit-Yan Yeung, Huchuan Lu, Xu Jia 0012
WACV7
2025 Self-calibrated region-level regression for crowd counting
Jiawen Zhu 0003, Wenda Zhao 0003, You He 0002, Huchuan Lu
Sci. China Inf. Sci.4
2025 Exploiting Lightweight Hierarchical ViT and Dynamic Framework for Efficient Visual Tracking
abstract
Abstract Transformer-based visual trackers have demonstrated significant advancements due to their powerful modeling capabilities. However, their practicality is limited on resource-constrained devices because of their slow processing speeds. To address this challenge, we present HiT, a novel family of efficient tracking models that achieve high performance while maintaining fast operation across various devices. The core innovation of HiT lies in its Bridge Module, which connects lightweight transformers to the tracking framework, enhancing feature representation quality. Additionally, we introduce a dual-image position encoding approach to effectively encode spatial information. HiT achieves an impressive speed of 61 frames per second (fps) on the NVIDIA Jetson AGX platform, alongside a competitive AUC of 64.6% on the LaSOT benchmark, outperforming all previous efficient trackers. Building on HiT, we propose DyHiT, an efficient dynamic tracker that flexibly adapts to scene complexity by selecting routes with varying computational requirements. DyHiT uses search area features extracted by the backbone network and inputs them into an efficient dynamic router to classify tracking scenarios. Based on the classification, DyHiT applies a divide-and-conquer strategy, selecting appropriate routes to achieve a superior trade-off between accuracy and speed. The fastest version of DyHiT achieves 111 fps on NVIDIA Jetson AGX while maintaining an AUC of 62.4% on LaSOT. Furthermore, we introduce a training-free acceleration method based on the dynamic routing architecture of DyHiT. This method significantly improves the execution speed of various high-performance trackers without sacrificing accuracy. For instance, our acceleration method enables the state-of-the-art tracker SeqTrack-B256 to achieve a $$2.68\times $$ 2.68 × speedup on an NVIDIA GeForce RTX 2080 Ti GPU while maintaining the same AUC of 69.9% on the LaSOT. Codes, models, and results are available at https://github.com/kangben258/HiT .
Ben Kang, Xin Chen 0032, Jie Zhao 0014, Chunjuan Bo, Dong Wang 0004, Huchuan Lu
Int. J. Comput. Vis.6
2025 Learning pose regression as reliable pixel-level matching for self-supervised depth estimation
Junwen He, Yifan Wang 0004, Lijun Wang 0001, Huchuan Lu
Neurocomputing4
2025 ComPtr: Toward Diverse Bi-Source Dense Prediction Tasks via a Simple Yet General Complementary Transformer
abstract
Deep learning (DL) has advanced the field of dense prediction, while gradually dissolving the inherent barriers between different tasks. However, most existing works focus on designing architectures and constructing visual cues only for the specific task, which ignores the potential uniformity introduced by the DL paradigm. In this paper, we attempt to construct a novel ComPlementary transformer, ComPtr, for diverse bi-source dense prediction tasks. Specifically, unlike existing methods that over-specialize in a single task or a subset of tasks, ComPtr starts from the more general concept of bi-source dense prediction. Based on the basic dependence on information complementarity, we propose consistency enhancement and difference awareness components with which ComPtr can evacuate and collect important visual semantic cues from different image sources for diverse tasks, respectively. ComPtr treats different inputs equally and builds an efficient dense interaction model in the form of sequence-to-sequence on top of the transformer. This task-generic design provides a smooth foundation for constructing the unified model that can simultaneously deal with various bi-source information. In extensive experiments across several representative vision tasks, i.e. remote sensing change detection, RGB-T crowd counting, RGB-D/T salient object detection, and RGB-D semantic segmentation, the proposed method consistently obtains favorable performance.
Youwei Pang, Xiaoqi Zhao 0003, Lihe Zhang, Huchuan Lu
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 MoE-Adapters++: Toward More Efficient Continual Learning of Vision-Language Models Via Dynamic Mixture-of-Experts Adapters
abstract
In this paper, we first propose MoE-Adapters, a parameter-efficient training framework to alleviate long-term forgetting issues in incremental learning with Vision-Language Models (VLM). Our MoE-Adapters leverages incrementally added routers to activate and integrate exclusive expert adapters from a pre-defined static expert set, enabling the pre-trained CLIP to efficiently adapt to new tasks. To preserve the zero-shot capability of VLM, a Distribution Discriminative Auto-Selector (DDAS) is introduced that automatically routes in-distribution and out-of-distribution inputs to the MoE-Adapters and the original CLIP, respectively. However, relying on a static expert set and a separate distribution selector can lead to parameter redundancy and increased training complexity. In response, we further extend an MoE-Adapters++ framework by introducing dynamic MoE-adapters, which allows experts to be adaptively involved during the continual learning process. Additionally, a Latent Embedding Auto-Selector (LEAS) is proposed that incorporates distribution selection within CLIP to create a more unified architecture. Extensive experiments across diverse settings demonstrate that the proposed method consistently surpasses previous state-of-the-art approaches while concurrently improving training efficiency.
Jiazuo Yu 0001, Zichen Huang 0004, Yunzhi Zhuge, Lu Zhang 0053, Ping Hu 0001, Dong Wang 0004, Huchuan Lu, You He 0002
IEEE Trans. Pattern Anal. Mach. Intell.7
2025 FreeFusion: Infrared and Visible Image Fusion via Cross Reconstruction Learning
abstract
Existing fusion methods empirically design elaborate fusion losses to retain the specific features from source images. Since image fusion has no ground truth, the hand-crafted losses may not make the fused images cover all the vital features, and then affect the performance of the high-level tasks. Here, there are two main challenges: domain discrepancy among source images and semantic mismatch at different-level tasks. This paper proposes an infrared and visible image fusion via cross reconstruction learning, which doesn't using any hand-crafted fusion losses, but prompts the network to adaptively fuse complementary information of source images. Firstly, we design a cross reconstruction learning model that decouples the fusion features to reconstruct another-modality source image. Thus, the fusion network is forced to learn the domain-adaptive representations of two modal features, which enables their domain alignment in a latent space. Secondly, we propose a dynamic interactive fusion strategy that builds a correlation matrix between fusion features and object semantic features to overcome the semantic mismatch. Further, we enhance the strong correlation features and suppress the weak correlation features to improve the interactive ability. Extensive experiments on three datasets demonstrate the superior fusion performance compared to the state-of-the-art methods, concurrently facilitating the segmentation accuracy.
Wenda Zhao 0003, Hengshuai Cui, You He 0002, Huchuan Lu
IEEE Trans. Pattern Anal. Mach. Intell.5
2025 Hybrid Gaussian Deformation for Efficient Remote Sensing Object Detection
abstract
Large-scale high-resolution remote sensing images (LSHR) are increasingly adopted for object detection, since they capture finer details. However, LSHR imposes a substantial computational cost. Existing methods explore lightweight backbones and advanced oriented bounding box regression mechanisms. Nevertheless, they still rely on high-resolution inputs to maintain detection accuracy. We observe that LSHR comprise extensive background areas that can be compressed to reduce unnecessary computation, while object regions contain details that can be reserved to improve detection accuracy. Thus, we propose a hybrid Gaussian deformation module that dynamically adjusts the sampling density at each location based on its relevance to the detection task, i.e., high-density sampling preserves more object regions and better retains detailed features, while low-density sampling diminishes the background proportion. Further, we introduce a bilateral deform-uniform detection framework to exploit the potential of the deformed sampled low-resolution images and original high-resolution images. Specifically, a deformed deep backbone takes the deformed sampled images as inputs to produce high-level semantic information, and a uniform shallow backbone takes the original high-resolution images as inputs to generate precise spatial location information. Moreover, we incorporate a deformation-aware feature registration module that calibrates the spatial information of deformed features, preventing regression degenerate solutions while maintaining feature activation. Subsequently, we introduce a feature relationship interaction fusion module to balance the contributions of features from both deformed and uniform backbones. Comprehensive experiments on three challenging datasets show that our method achieves superior performance compared with the state-of-the-art methods.
Wenda Zhao 0003, Xiao Zhang 0050, Huchuan Lu
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 Remote Sensing Image Generation via Object Text Decoupling
abstract
Remote sensing images usually reveal various objects with complex structures and different locations within vast ground area backgrounds. That leads to a major challenge for conventional generative models in handling remote sensing objects with correct shapes and clear textures. Integrating additional object-level controls can be a potential solution to improve generation quality, yet previous approaches inject the object-related conditions by specifying their locations, causing a limitation in object layout in generated results. To enable high object fidelity, high layout diversity and object customizable generation for remote sensing images, we propose a remote sensing image generation via object text decoupling, namely OTD-GAN. OTD-GAN takes advantage of the inherent text-to-image generation procedure and adaptively integrates the decoupled textual representations of visual objects into the global captions, thus achieving object-level controls without layout restrictions. Specifically, we design an object text decoupling module to predict a semantically consistent textual representation for each object. By decoupling the textual representation into a class invariant part and an object specific part, the converted representation is able to catch general semantic for similar objects as well as differentiated details for individual objects. After that, we use an object text semantic enhancement module to fuse the obtained object text representations with the global captions to enrich the object-related semantic within the textual modality. As a result, the generator will benefit from the object conditions and reinforce the generation quality while remaining flexibility to create diverse layouts. Extensive experiments on remote sensing image-caption datasets including NWPU-Captions and RSICD demonstrate that our method achieves leading performance compared to existing state-of-the-art approaches.
Wenda Zhao 0003, Zhepu Zhang, Fan Zhao 0005, You He 0002, Huchuan Lu
IEEE Trans. Pattern Anal. Mach. Intell.6
2025 Beyond mask: Rethinking guidance types in few-shot segmentation
Shijie Chang, Youwei Pang, Xiaoqi Zhao 0003, Huchuan Lu, Lihe Zhang
Pattern Recognit.4
2025 RVMamba: Selective Text-Vision Mamba for Referring Video Object Segmentation
abstract
Existing RVOS methods typically employ Transformers to model global cross-modal, temporal-spatial correspondences, but their quadratic complexity limits deployment on resource-constrained devices. To overcome this limitation, Mamba offers a sequence modeling framework with linear computational complexity. We introduceRVMamba, which utilizes weight modulation to selectively update hidden states across text-frame sequences, enabling effective linguistic context propagation, and a learning-based scanning strategy to efficiently capture spatio-temporal dependencies with linear memory consumption. Extensive experiments demonstrate thatRVMambaachieves state-of-the-art performance on public benchmarks, with significantly reduced memory growth, offering an efficient and scalable solution for long video processing.
Zhenyu Chen 0001, Jiawen Zhu 0003, Lu Zhang 0053, Ping Hu 0001, Yunzhi Zhuge, Huchuan Lu, You He 0002
IEEE Signal Process. Lett.6
2025 ProSegDiff: Prostate Segmentation Diffusion Network Based on Adaptive Adjustment of Injection Features
abstract
Recently, methods based on Diffusion Probability Models (DPM) have achieved notable success in the field of medical image segmentation. However, most of these methods do not perform well in segmenting ambiguous areas when dealing with prostate segmentation tasks due to the low distinguishability of prostate images and the high overlap of its boundary with adjacent organs. To address this issue, this paper introduces a diffusion-based framework named ProSegDiff, ProSegDiff employs an Adapter to dynamically adjust features from the conditional network to align with the denoising process of the denoising network. Furthermore, the denoising process is conducted in the latent space to minimize the consumption of computational resources, and a proposed selection strategy is employed to identify the better results from multiple inferences. Extensive comparative experiments on four benchmark datasets demonstrate the effectiveness of this method, which achieves superior performance across four evaluation metrics.
Jialong Zhong, Tingwei Liu, Yongri Piao, Weibing Sun, Huchuan Lu
IEEE Signal Process. Lett.5
2025 MambaVT: Spatio-Temporal Contextual Modeling for Robust RGB-T Tracking
Simiao Lai, Chang Liu 0071, Jiawen Zhu 0003, Ben Kang, Yang Liu 0066, Dong Wang 0004, Huchuan Lu
IEEE Trans. Circuits Syst. Video Technol.7
2025 MoBox: Enhancing Video Object Segmentation With Motion-Augmented Box Supervision
abstract
We propose MoBox, a low-cost solution for semi-supervised video object segmentation that requires only bounding boxes as manual annotations for training. Built upon a mature semi-supervised video object segmentation network, we redesign the training losses and employ a more stringent training strategy. Specifically, we introduce a well-designed constraint term that enhances traditional spatial projection by simultaneously leveraging the projections of both the ground-truth box and the predicted mask across two axes, rather than evaluating discrepancies along the x-axis and y-axis independently. To harness the intrinsic properties of videos, considering the underlying correspondence between motion represented by optical flow and the original image, we incorporate motion coherence information into the color consistency loss as supplementary information and propose a motion discrepancy loss to obtain accurate boundaries. Additionally, to mitigate the ambiguity of weak supervision, we further introduce the pseudo strict constraint during training, which significantly improves model performance. Our approach yields competitive scores on popular benchmarks, achieving a$\mathcal {J}\& \mathcal {F}$score of 78.6 on the DAVIS 2017 validation set and an Overall score of 78.0 on the YouTube-VOS 2018 validation set. These results highlight the efficacy of MoBox, demonstrating that the semi-supervised video object segmentation model can be effectively trained using only motion-augmented box supervision and intrinsic information of videos.
Xiaomin Li 0001, Dezhuang Li, Mengmeng Ge 0002, Xu Jia 0012, You He 0002, Huchuan Lu
IEEE Trans. Circuits Syst. Video Technol.7
2025 EMTrack: Efficient Multimodal Object Tracking
abstract
Multi-modal object tracking has received increasing attention, given the limitations the representation ability in certain challenging scenarios of single RGB modality. Recent prompt tuning techniques enable multimodal tracking to effectively inherit knowledge from foundation models trained with a large amount of RGB tracking data and achieve parameter-efficient training. However, few works focus on the efficient inference of multimodal tracking handling multiple RGB-X (RGB-Thermal, RGB-Depth, RGB-Event, etc.) tracking tasks simultaneously, especially on resource-limited devices such as CPU. In this work, we propose an efficient multimodal tracker named EMTrack. EMTrack follows a concise and unified multimodal tracking framework with simple knowledge distillation. RGB modality and auxiliary modality are added after patch-embedding layer for fusion, reducing the computational complexity of multimodal tracking compared with that of single modality. Before fusion operation, we introduce a modal-specific spatial modulation module to exploit and realize adaptive spatial adjustment of different modality features. Multiple modal-specific experts are adopted to capture specific information for different RGB-X tracking tasks, which assists in handling such tasks in a unified model with joint training. EMTrack achieves competitive performance on various RGB-X tracking benchmarks while reaching a good balance of performance and speed on different platforms. Especially on an Intel Core i9-10850K CPU device, EMTrack achieves 29.1 fps, a real-time speed, with only 2.0G MAC computation.
Chang Liu 0071, Ziqi Guan, Simiao Lai, Yang Liu 0066, Huchuan Lu, Dong Wang 0004
IEEE Trans. Circuits Syst. Video Technol.5
2025 CNN-Transformer Rectified Collaborative Learning for Medical Image Segmentation
abstract
Automatic and precise medical image segmentation (MIS) is of vital importance for clinical diagnosis and analysis. Current MIS methods mainly rely on the convolutional neural network (CNN) or self-attention mechanism (Transformer) for feature modeling. However, CNN-based methods suffer from the inaccurate localization owing to the limited global dependency while Transformer-based methods always present the coarse boundary for the lack of local emphasis. Although some CNN-Transformer hybrid methods are designed to synthesize the complementary local and global information for better performance, the combination of CNN and Transformer introduces numerous parameters and increases the computation cost. To this end, this paper proposes a CNN-Transformer rectified collaborative learning (CTRCL) framework to learn stronger CNN-based and Transformer-based models for MIS tasks via the bi-directional knowledge transfer between them. Specifically, we propose a rectified logit-wise collaborative learning (RLCL) strategy which introduces the ground truth to adaptively select and rectify the wrong regions in student soft labels for accurate knowledge transfer in the logit space. We also propose a class-aware feature-wise collaborative learning (CFCL) strategy to achieve effective knowledge transfer between CNN-based and Transformer-based models in the feature space by granting their intermediate features the similar capability of category perception. Extensive experiments on three popular MIS benchmarks demonstrate that our CTRCL outperforms most state-of-the-art collaborative learning methods under different evaluation metrics. The source code will be publicly available athttps://github.com/LanhooNg/CTRCL.
Lanhu Wu, Miao Zhang 0004, Yongri Piao, Zhenyan Yao, Weibing Sun, Feng Tian 0001, Huchuan Lu
IEEE Trans. Circuits Syst. Video Technol.7
2025 FocusCLIP: Focusing on Anomaly Regions by Visual-Text Discrepancies
abstract
Few-shot anomaly detection aims to detect defects with only a limited number of normal samples for training. Recent few-shot methods typically focus on object-level features rather than subtle defects within objects, as pretrained models are generally trained on classification or image-text matching datasets. However, object-level features are often insufficient to detect defects, which are characterized by fine-grained texture variations. To address this, we propose FocusCLIP, which consists of a vision-guided branch and a language-guided branch. FocusCLIP leverages the complementary relationship between visual and text modalities to jointly emphasize discrepancies in fine-grained textures of defect regions. Specifically, we design three modules to mine these discrepancies. In the vision-guided branch, we propose the Bidirectional Self-knowledge Distillation (BSD) structure, which identifies anomaly regions through inconsistent representations and accumulates these discrepancies. Within this structure, the Anomaly Capture Module (ACM) is designed to refine features and detect more comprehensive anomalies by leveraging semantic cues from multi-head self-attention. In the language-guided branch, Multi-level Adversarial Class Activation Mapping (MACAM) utilizes foreground-invariant responses to adversarial text prompts, reducing interference from object regions and further focusing on defect regions. Our approach outperforms the state-of-the-art methods in few-shot anomaly detection. Additionally, the language-guided branch within FocusCLIP also demonstrates competitive performance in zero-shot anomaly detection, further validating the effectiveness of our proposed method.
Yuan Zhao 0006, Lihe Zhang, Huchuan Lu
IEEE Trans. Circuits Syst. Video Technol.4
2025 Learning Language Prompt for Vision-Language Tracking
abstract
Vision-language object tracking integrates advanced linguistic information, enhancing its robustness and accuracy in complex scenarios. Nevertheless, current methods are constrained by a lack of sufficient vision-language data, making it challenging for the model to learn generalized knowledge. To alleviate this issue, we propose a new prompt-based framework for vision-language tracking, named ProVLT. This framework casts language information as a prompt for pretrained visionbased tracking models, thereby leveraging the knowledge from extensive tracking data. Experiments demonstrate that ProVLT achieves competitive performance while training only a fraction of parameters (approximately 29% of modal parameters). For instance, ProVLT achieves competitive performance, attaining AUC of 59.8% on TNL2K benchmark. Furthermore, we augment five mainstream vision-only tracking benchmarks with language annotations, and find that the inclusion of linguistic information consistently improves tracking performance. On these benchmarks, the linguistic information improves the performance by an average of 2.9% compared with the vision-based tracker. We will release the code, models, and benchmarks for the community.
ChengAo Zong, Jie Zhao 0014, Xin Chen 0032, Huchuan Lu, Dong Wang 0004
IEEE Trans. Circuits Syst. Video Technol.4
2025 Spatial-Frequency Enhanced Mamba for Multi-Modal Image Fusion
abstract
Multi-Modal Image Fusion (MMIF) aims to integrate complementary image information from different modalities to produce informative images. Previous deep learning-based MMIF methods generally adopt Convolutional Neural Networks (CNNs) or Transformers for feature extraction. However, these methods deliver unsatisfactory performances due to the limited receptive field of CNNs and the high computational cost of Transformers. Recently, Mamba has demonstrated a powerful potential for modeling long-range dependencies with linear complexity, providing a promising solution to MMIF. Unfortunately, Mamba lacks full spatial and frequency perceptions, which are very important for MMIF. Moreover, employing Image Reconstruction (IR) as an auxiliary task has been proven beneficial for MMIF. However, a primary challenge is how to leverage IR efficiently and effectively. To address the above issues, we propose a novel framework named Spatial-Frequency Enhanced Mamba Fusion (SFMFusion) for MMIF. More specifically, we first propose a three-branch structure to couple MMIF and IR, which can retain complete contents from source images. Then, we propose the Spatial-Frequency Enhanced Mamba Block (SFMB), which can enhance Mamba in both spatial and frequency domains for comprehensive feature extraction. Finally, we propose the Dynamic Fusion Mamba Block (DFMB), which can be deployed across different branches for dynamic feature fusion. Extensive experiments show that our method achieves better results than most state-of-the-art methods on six MMIF datasets. The source code is available at https://github.com/SunHui1216/SFMFusion.
Long Lv, Tongdan Tang, Feng Tian 0001, Weibing Sun, Huchuan Lu
IEEE Trans. Image Process.7
2025 CharacterFactory: Sampling Consistent Characters With GANs for Diffusion Models
abstract
Recent advances in text-to-image models have opened new frontiers in human-centric generation. However, these models cannot be directly employed to generate images with consistent newly coined identities. In this work, we propose CharacterFactory, a framework that allows sampling new characters with consistent identities in the latent space of GANs for diffusion models. More specifically, we consider the word embeddings of celeb names as ground truths for the identity-consistent generation task and train a GAN model to learn the mapping from a latent space to the celeb embedding space. In addition, we design a context-consistent loss to ensure that the generated identity embeddings can produce identity-consistent images in various contexts. Remarkably, the whole model only takes 10 minutes for training, and can sample infinite characters end-to-end during inference. Extensive experiments demonstrate excellent performance of the proposed CharacterFactory on character creation in terms of identity consistency and editability. Furthermore, the generated characters can be seamlessly combined with the off-the-shelf image/video/3D diffusion models. We believe that the proposed CharacterFactory is an important step for identity-consistent character generation. Code and Gradio demo are available at: https://qinghew.github.io/CharacterFactory/.
Baolu Li 0001, Xiaomin Li 0001, Bing Cao 0002, Liqian Ma, Huchuan Lu, Xu Jia 0012
IEEE Trans. Image Process.6
2025 Self-Adaptive Vision-Language Tracking With Context Prompting
abstract
Due to the substantial gap between vision and language modalities, along with the mismatch problem between fixed language descriptions and dynamic visual information, existing vision-language tracking methods exhibit performance on par with or slightly worse than vision-only tracking. Effectively exploiting the rich semantics of language to enhance tracking robustness remains an open challenge. To address these issues, we propose a self-adaptive vision-language tracking framework that leverages the pre-trained multi-modal CLIP model to obtain well-aligned visual-language representations. A novel context-aware prompting mechanism is introduced to dynamically adapt linguistic cues based on the evolving visual context during tracking. Specifically, our context prompter extracts dynamic visual features from the current search image and integrates them into the text encoding process, enabling self-updating language embeddings. Furthermore, our framework employs a unified one-stream Transformer architecture, supporting joint training for both vision-only and vision-language tracking scenarios. Our method not only bridges the modality gap but also enhances robustness by allowing language features to evolve with visual context. Extensive experiments on four vision-language tracking benchmarks demonstrate that our method effectively leverages the advantages of language to enhance visual tracking. Our large model can obtain 55.0% AUC on $\text {LaSOT}_{\text {EXT}}$ and 69.0% AUC on TNL2K. Additionally, our language-only tracking model achieves performance comparable to that of state-of-the-art vision-only tracking methods on TNL2K. Code is available at https://github.com/zj5559/SAVLT.
Jie Zhao 0014, Xin Chen 0032, Shengming Li, Chunjuan Bo, Dong Wang 0004, Huchuan Lu
IEEE Trans. Image Process.6
2025 Enhancing the Two-Stream Framework for Efficient Visual Tracking
abstract
Practical deployments, especially on resource-limited edge devices, necessitate high speed for visual object trackers. To meet this demand, we introduce a new efficient tracker with a Two-Stream architecture, named ToS. While the recent one-stream tracking framework, employing a unified backbone for simultaneous processing of both the template and search region, has demonstrated exceptional efficacy, we find the conventional two-stream tracking framework, which employs two separate backbones for the template and search region, offers inherent advantages. The two-stream tracking framework is more compatible with advanced lightweight backbones and can efficiently utilize benefits from large templates. We demonstrate that the two-stream setup can exceed the one-stream tracking model in both speed and accuracy through strategic designs. Our methodology rejuvenates the two-stream tracking paradigm with lightweight pre-trained backbones and the proposed three efficient strategies: 1) A feature-aggregation module that improves the representation capability of the backbone, 2) A channel-wise approach for feature fusion, presenting a more effective and lighter alternative to spatial concatenation techniques, and 3) An expanded template strategy to boost tracking accuracy with negligible additional computational cost. Extensive evaluations across multiple tracking benchmarks demonstrate that the proposed method sets a new state-of-the-art performance in efficient visual tracking.
ChengAo Zong, Xin Chen 0032, Jie Zhao 0014, Yang Liu 0066, Huchuan Lu, Dong Wang 0004
IEEE Trans. Image Process.5
2025 Unity Is Strength: Unifying Convolutional and Transformeral Features for Better Person Re-Identification
abstract
Person Re-identification (ReID) aims to retrieve the specific person across non-overlapping cameras, which greatly helps intelligent transportation systems. As we all know, Convolutional Neural Networks (CNNs) and Transformers have the unique strengths to extract local and global features, respectively. Considering this fact, we focus on the mutual fusion between them to learn more comprehensive representations for persons. In particular, we utilize the complementary integration of deep features from different model structures. We propose a novel fusion framework called FusionReID to unify the strengths of CNNs and Transformers for image-based person ReID. More specifically, we first deploy a Dual-branch Feature Extraction (DFE) to extract features through CNNs and Transformers from a single image. Moreover, we design a novel Dual-attention Mutual Fusion (DMF) to achieve sufficient feature fusions. The DMF comprises Local Refinement Units (LRU) and Heterogenous Transmission Modules (HTM). LRU utilizes depth-separable convolutions to align deep features in channel dimensions and spatial sizes. HTM consists of a Shared Encoding Unit (SEU) and two Mutual Fusion Units (MFU). Through the continuous stacking of HTM, deep features after LRU are repeatedly utilized to generate more discriminative features. Extensive experiments on three public ReID benchmarks demonstrate that our method can attain superior performances than most state-of-the-arts. The source code is available athttps://github.com/924973292/FusionReID.
Xuehu Liu, Zhengzheng Tu, Huchuan Lu
IEEE Trans. Intell. Transp. Syst.5
2025 AVS-Mamba: Exploring Temporal and Multi-Modal Mamba for Audio-Visual Segmentation
abstract
The essence of audio-visual segmentation (AVS) lies in locating and delineating sound-emitting objects within a video stream. While Transformer-based methods have shown promise, their handling of long-range dependencies struggles due to quadratic computational costs, presenting a bottleneck in complex scenarios. To overcome this limitation and facilitate complex multi-modal comprehension with linear complexity, we introduce AVS-Mamba, a selective state space model to address the AVS task. Our framework incorporates two key components for video understanding and cross-modal learning: Temporal Mamba Block for sequential video processing and Vision-to-Audio Fusion Block for advanced audio-vision integration. Building on this, we develop the Multi-scale Temporal Encoder, aimed at enhancing the learning of visual features across scales, facilitating the perception of intra- and inter-frame information. To perform multi-modal fusion, we propose the Modality Aggregation Decoder, leveraging the Vision-to-Audio Fusion Block to integrate visual features into audio features across both frame and temporal levels. Further, we adopt the Contextual Integration Pyramid to perform audio-to-vision spatial-temporal context collaboration. Through these innovative contributions, our approach achieves new state-of-the-art results on the AVSBench-object and AVSBench-semantic datasets.
Sitong Gong, Yunzhi Zhuge, Lu Zhang 0053, Yifan Wang 0004, Lijun Wang 0001, Huchuan Lu
IEEE Trans. Multim.7
2025 Complementary and Contrastive Learning for Audio-Visual Segmentation
abstract
Audio-Visual Segmentation (AVS) aims to generate pixel-wise segmentation maps that correlate with the auditory signals of objects. This field has seen significant progress with numerous CNN and Transformer-based methods enhancing the segmentation accuracy and robustness. Traditional CNN approaches manage audio-visual interactions through basic operations like padding and multiplications but are restricted by CNNs' limited local receptive field. More recently, Transformer-based methods treat auditory cues as queries, utilizing attention mechanisms to enhance audio-visual cooperation within frames. Nevertheless, they typically struggle to extract multimodal coefficients and temporal dynamics adequately. To overcome these limitations, we present the Complementary and Contrastive Transformer (CCFormer), a novel framework adept at processing both local and global information and capturing spatial-temporal context comprehensively. Our CCFormer initiates with the Early Integration Module (EIM) that employs a parallel bilateral architecture, merging multi-scale visual features with audio data to boost cross-modal complementarity. To extract the intra-frame spatial features and facilitate the perception of temporal coherence, we introduce the Multi-query Transformer Module (MTM), which dynamically endows audio queries with learning capabilities and models the frame and video-level relations simultaneously. Furthermore, we propose the Bi-modal Contrastive Learning (BCL) to promote the alignment across both modalities in the unified feature space. Through the effective combination of those designs, our method sets new state-of-the-art benchmarks across the S4, MS3 and AVSS datasets.
Sitong Gong, Yunzhi Zhuge, Lu Zhang 0053, Huchuan Lu
IEEE Trans. Multim.5
2025 StableIdentity: Inserting Anybody Into Anywhere at First Sight
abstract
Recent advances in large pretrained text-to-image generation models have shown unprecedented capabilities for high-quality human-centric generation, however, customizing face identity is still an intractable problem. Existing methods cannot ensure stable identity preservation and flexible editability, even with several images for each subject during training. In this work, we propose StableIdentity, which allows identity-consistent recontextualization with just one face image from a person seen for the first time. More specifically, we employ a face encoder with the identity prior to encode the input face, and then calibrate the face representation to align the distribution of a space with the editability prior, which is constructed from celeb names. By incorporating identity prior and editability prior, the learned identity can be injected anywhere with various contexts. In addition, we design a masked two-phase diffusion loss to boost the pixel-level perception of the input face and maintain the diversity of generation. Extensive experiments demonstrate our method outperforms previous customization methods. In addition, the learned identity can be flexibly combined with the off-theshelf modules such as ControlNet. Notably, to the best of our knowledge, we are the first to directly inject the identity learned from a single image into video/3D generation without finetuning. We believe that the proposed StableIdentity is an important step to unify image, video, and 3D customized generation models. The code is available: https://github.com/qinghew/StableIdentity.
Xu Jia 0012, Xiaomin Li 0001, Taiqing Li, Liqian Ma, Yunzhi Zhuge, Huchuan Lu
IEEE Trans. Multim.7
2025 3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding
abstract
Multi-modal Large Language Models (MLLMs) exhibit impressive capabilities in 2D tasks, yet encounter challenges in discerning the spatial positions, interrelations, and causal logic in scenes when transitioning from 2D to 3D representations. We find that the limitations mainly lie in: i) the high annotation cost restricting the scale-up of volumes of 3D scene data, and ii) the lack of a straightforward and effective way to perceive 3D information which results in prolonged training durations and complicates the streamlined framework. To this end, we develop a pipeline based on open-source 2D MLLMs and LLMs to generate high-quality 3D-text pairs and construct 3DS-160 K, to enhance the pre-training process. Leveraging this high-quality pre-training data, we introduce the 3UR-LLM model, an end-to-end 3D MLLM designed for precise interpretation of 3D scenes, showcasing exceptional capability in navigating the complexities of the physical world. 3UR-LLM directly receives 3D point cloud as input and project 3D features fused with text instructions into a manageable set of tokens. Considering the computation burden derived from these hybrid tokens, we design a 3D compressor module to cohesively compress the 3D spatial cues and textual narrative. 3UR-LLM achieves promising performance with respect to the previous SOTAs, for instance, 3UR-LLM exceeds its counterparts by 7.1% CIDEr on ScanQA, while utilizing fewer training resources. The code and model weights for 3UR-LLM and the 3DS-160 K benchmark are available at 3UR-LLM.
Haomiao Xiong, Yunzhi Zhuge, Jiawen Zhu 0003, Lu Zhang 0053, Huchuan Lu
IEEE Trans. Multim.5
2025 Semantics Alternating Enhancement and Bidirectional Aggregation for Referring Video Object Segmentation
abstract
Referring Video Object Segmentation (RVOS) aims at segmenting out the described object in a video clip according to given expression. The task requires methods to effectively fuse cross-modality features, communicate temporal information, and delineate referent appearance. However, existing solutions bias their focus to mainly mining one or two clues, causing their performance inferior. In this paper, we propose Semantics Alternating Enhancement (SAE) to achieve cross-modality fusion and temporal-spatial semantics mining in an alternate way that makes comprehensive exploit of three cues possible. During each update, SAE will generate a cross-modality and temporal-aware vector that guides vision feature to amplify its referent semantics while filtering out irrelevant contents. In return, the purified feature will provide the contextual soil to produce a more refined guider. Overall, cross-modality interaction and temporal communication are together interleaved into axial semantics enhancement steps. Moreover, we design a simplified SAE by dropping spatial semantics enhancement steps, and employ the variant in the early stages of vision encoder to further enhance usability. To integrate features of different scales, we propose Bidirectional Semantic Aggregation decoder (BSA) to obtain referent mask. The BSA arranges the comprehensively-enhanced features into two groups, and then employs difference awareness step to achieve intra-group feature aggregation bidirectionally and consistency constraint step to realize inter-group integration of semantics-dense and appearance-rich features. Extensive results on challenging benchmarks show that our method performs favorably against the state-of-the-art competitors.
Lihe Zhang, Huchuan Lu
IEEE Trans. Multim.3
2025 PMNet: Predator-Mimicking Network for Video Camouflaged Object Detection
abstract
The predator has the ability to quickly respond to the misjudged decision and hunt the camouflaged target by analyzing its movement. Those decision compensation and movement analysis for hunting are closely tied to temporal and spatial information. This can be mirrored in the video camouflaged object detection (VCOD) task where the captured temporal information may be misjudged as well as the spatial information tends to be inaccurate in complex scenes. Thus, two key factors should be considered in the VCOD task: How can a model cope with the misjudged temporal information; How can spatial features interact with the temporal information to understand dynamic scenes? To this end, we propose a predator-mimicking network (PMNet) equipped with a selective temporal alignment module (STAM) and a temporal-spatial feedback module (T-SFM). The STAM is designed to alleviate the influence of the misjudged motion trajectory by adopting our adaptive selection mechanism from a novel perspective. In T-SFM, the temporal information works as the self-knowledge to provide assistance and interact with spatial features, enabling the model to effectively detect the camouflaged object. Experimental results demonstrate that our method achieves state-of-the-art performance on VCOD benchmarks. Furthermore, our model can be generalized in the video salient object detection (VSOD) task and also outperforms existing state-of-the-art methods. The source code will be publicly available athttps://github.com/LiuTingWed/CriDiff.
Miao Zhang 0004, Beiqi Hu, Shunyu Yao 0004, Yongri Piao, Huchuan Lu
IEEE Trans. Multim.5
2025 MaskTrack: Auto-Labeling and Stable Tracking for Video Object Segmentation
abstract
Video object segmentation (VOS) has witnessed notable progress due to the establishment of video training datasets and the introduction of diverse, innovative network architectures. However, video mask annotation is a highly intricate and labor-intensive task, as meticulous frame-by-frame comparisons are needed to ascertain the positions and identities of targets in the subsequent frames. Current VOS benchmarks often annotate only a few instances in each video to save costs, which, however, hinders the model's understanding of the complete context of the video scenes. To simplify video annotation and achieve efficient dense labeling, we introduce a zero-shot auto-labeling strategy based on the segment anything model (SAM), enabling it to densely annotate video instances without access to any manual annotations. Moreover, although existing VOS methods demonstrate improving performance, segmenting long-term and complex video scenes remains challenging due to the difficulties in stably discriminating and tracking instance identities. To this end, we further introduce a new framework, MaskTrack, which excels in long-term VOS and also exhibits significant performance advantages in distinguishing instances in complex videos with densely packed similar objects. We conduct extensive experiments to demonstrate the effectiveness of the proposed method and show that without introducing image datasets for pretraining, it achieves excellent performance on both short-term (86.2% in YouTube-VOS val) and long-term (68.2% in LVOS val) VOS benchmarks. Our method also surprisingly demonstrates strong generalization ability and performs well in visual object tracking (VOT) (65.6% in VOTS2023) and referring VOS (RVOS) (65.2% in Ref YouTube VOS) challenges.
Zhenyu Chen 0001, Lu Zhang 0053, Ping Hu 0001, Huchuan Lu, You He 0002
IEEE Trans. Neural Networks Learn. Syst.4
2025 Refocus the Attention for Parameter-Efficient Thermal Infrared Object Tracking
abstract
Introducing deep trackers to thermal infrared (TIR) tracking is hampered by the scarcity of large training datasets. To alleviate the predicament, a common approach is full fine-tuning (FFT) based on pretrained RGB parameters. Nevertheless, due to its inefficient training pattern and representation collapse risk, some parameter-efficient fine-tuning (PEFT) alternatives have been promoted recently. However, the existing PEFT algorithms typically follow a bottom-up way, where their attention solely relies on the input and lacks the capability of task-guided top-down attention, which provides the task-relevant representation such as the human visual perception system. In this article, we introduce ReFocus, a new PEFT method that adapts the pretrained RGB foundation tracking model to the downstream TIR tracking task through the guidance of high-level task-specific signals in a top-down attention manner. By freezing the entire foundation model and only training query-guided feature selection and top-down blocks, ReFocus achieves state-of-the-art (SOTA) TIR tracking performance while keeping training efficiency. Extensive experiments on five TIR tracking benchmarks demonstrate that ReFocus significantly improves the performance of the foundation tracker. Besides, further ablation studies show the effectiveness and flexible adaptability of the proposed method to lighter foundation models and different tracking frameworks. Compared to FFT and other bottom-up PEFT paradigms, such as head probe, low-rank adaptation (LoRA), and adapter, our method achieves comparable or superior performance with fewer training parameters and reveals the advantage of learning stability.
Simiao Lai, Chang Liu 0071, Dong Wang 0004, Huchuan Lu
IEEE Trans. Neural Networks Learn. Syst.4
2025 Learning Discriminative Representation for Co-Salient Object Detection
abstract
Co-salient object detection (CoSOD) is the task of identifying and emphasizing the common salient objects in a collection of images. The current co-salient object detection frameworks often extract features and model interimage relations separately. Although these methods achieve promising performance in many scenes, separating the feature extraction and relation modeling falls short of obtaining discriminative features for co-salient objects, resulting in subperformance, especially in some complex and cluttered real-world scenes. In this article, we introduce a novel CoSOD framework to unify feature extraction and interimage relation modeling. We design an early token interaction module (ETIM) that bridges information flow between branches to simultaneously realize feature extraction and interimage information interaction. To further enhance our network's capability to distinguish co-salient objects from other irrelevant foreground objects, we introduce a pixel-to-group contrastive (PGC) learning method. This approach aids in eliminating the need for additional interaction modules while preserving features' discriminative power for co-salient objects. Our proposed CoSOD framework only includes a backbone embedded with ETIM, a decoder without interaction modules and a project head only used during the training phase. Extensive experiments on three challenging benchmarks, that is, CoCA, CoSOD3k, and Cosal2015, demonstrate that our proposed method can outperform current leading-edge models and achieve the new state-of-the-art. The source code is available at https://github.com/zhiwang98/LDRNet.
Yongri Piao, Tingwei Liu, Jihao Yin, Miao Zhang 0004, Huchuan Lu
IEEE Trans. Neural Networks Learn. Syst.6
2025 Real-Time Semantic Segmentation via a Densely Aggregated Bilateral Network
abstract
With the growing demands of applications on online devices, the speed-accuracy trade-off is critical in the semantic segmentation system. Recently, the bilateral segmentation network has shown promising capacity to achieve the balance between favorable accuracy and fast speed, and has become the mainstream backbone in real-time semantic segmentation. Segmentation of target objects relies on high-level semantics, whereas it requires detailed low-level features to model specific local patterns for accurate location. However, the lightweight backbone of bilateral architecture limits the extraction of semantic context and spatial details. And the late fusion of the bilateral streams incurs the insufficient aggregation of semantic context and spatial details. In this article, we propose a densely aggregated bilateral network (DAB-Net) for real-time semantic segmentation. In the context path, a patchwise context enhancement (PCE) module is proposed to efficiently capture the local semantic contextual information from spatialwise and channelwise, respectively. Meanwhile, a context-guided spatial path (CGSP) is designed to exploit more spatial information by encoding finer details from the raw image and the transition from the context path. Finally, with multiple interactions between bilateral branches, the intertwined outputs from bilateral streams are combined in a unified decoder for a final interaction to further enhance the feature representation, which generates the final segmentation prediction. Experimental results on three public benchmarks demonstrate that our proposed method achieves higher accuracy with a limited decay in speed, which performs favorably against state-of-the-art real-time approaches and runs at 31.1 frames/s (FPS) on the high resolution of . The source code is released at https://github.com/isyangshu/DABNet.
Shu Yang 0004, Lu Zhang 0053, Shuai Liu 0009, Huchuan Lu, Hao Chen 0011
IEEE Trans. Neural Networks Learn. Syst.4
2025 Event-Assisted Recurrent Network for Arbitrary-Temporal-Scale Blurry Image Unfolding
abstract
Recovering a sequence of latent sharp frames from a motion-blurred image is a challenging task. The bio-inspired event camera, which produces an event stream with high temporal resolution, has been exploited to promote the recovery performance. However, recovering sharp sequences with arbitrary temporal scales has been ignored for a long time. Existing works can only recover a fixed number of latent frames from a blurry image once they are trained. In this work, we propose an event-assisted blurry image unfolding framework that can work across arbitrary temporal scales. A bi-directional recurrent network is employed to encode events corresponding to each latent frame, which gathers information over all events in the exposure time. Features of both the blurry image and events are fused together and fed to a bi-directional latent sequence decoder (BiLSD) to produce a sequence of latent sharp frames. Extensive experiments show that the proposed method not only performs favorably against state-of-the-art methods in recovering a fixed number of frames from a blurry image but can be well generalized to arbitrary-temporal-scale blurry image unfolding.
Hao Ju 0004, Weihua He, Yaoyuan Wang, Shengming Li, Dong Wang 0004, Huchuan Lu, Xu Jia 0012
IEEE Trans. Neural Networks Learn. Syst.8
2025 Exploring Dynamic Transformer for Efficient Object Tracking
abstract
The speed-precision tradeoff is a critical problem in visual object tracking, as it typically requires low latency and is deployed on resource-constrained platforms. Existing solutions for efficient tracking primarily focus on lightweight backbones or modules, which, however, come at a sacrifice in precision. In this article, inspired by dynamic network routing, we propose DyTrack, a dynamic transformer framework for efficient tracking. Real-world tracking scenarios exhibit varying levels of complexity. We argue that a simple network is sufficient for easy video frames, while more computational resources should be assigned to difficult ones. DyTrack automatically learns to configure proper reasoning routes for different inputs, thereby improving the utilization of the available computational budget and achieving higher performance at the same running speed. We formulate instance-specific tracking as a sequential decision problem and incorporate terminating branches to intermediate layers of the model. Furthermore, we propose a feature recycling mechanism to maximize computational efficiency by reusing the outputs of predecessors. Additionally, a target-aware self-distillation strategy is designed to enhance the discriminating capabilities of early-stage predictions by mimicking the representation patterns of the deep model. Extensive experiments demonstrate that DyTrack achieves promising speed-precision tradeoffs with only a single model. For instance, DyTrack obtains 64.9% area under the curve (AUC) on LaSOT with a speed of 256 fps.
Jiawen Zhu 0003, Xin Chen 0032, Haiwen Diao, Shuai Li 0014, Jun-Yan He, Chenyang Li 0007, Bin Luo 0008, Dong Wang 0004, Huchuan Lu
IEEE Trans. Neural Networks Learn. Syst.9
2025 Learning Motion and Temporal Cues for Unsupervised Video Object Segmentation
abstract
In this article, we address the challenges in unsupervised video object segmentation (UVOS) by proposing an efficient algorithm, termed MTNet, which concurrently exploits motion and temporal cues. Unlike previous methods that focus solely on integrating appearance with motion or on modeling temporal relations, our method combines both aspects by integrating them within a unified framework. MTNet is devised by effectively merging appearance and motion features during the feature extraction process within encoders, promoting a more complementary representation. To capture the intricate long-range contextual dynamics and information embedded within videos, a temporal transformer module is introduced, facilitating efficacious interframe interactions throughout a video clip. Furthermore, we employ a cascade of decoders all feature levels across all feature levels to optimally exploit the derived features, aiming to generate increasingly precise segmentation masks. As a result, MTNet provides a strong and compact framework that explores both temporal and cross-modality knowledge to robustly localize and track the primary object accurately in various challenging scenarios efficiently. Extensive experiments across diverse benchmarks conclusively show that our method not only attains state-of-the-art performance in UVOS but also delivers competitive results in video salient object detection (VSOD). These findings highlight the method's robust versatility and its adeptness in adapting to a range of segmentation tasks. The source code is available at https://github.com/hy0523/MTNet.
Yunzhi Zhuge, Hongyu Gu, Lu Zhang 0053, Jinqing Qi, Huchuan Lu
IEEE Trans. Neural Networks Learn. Syst.5
2024 TOP-ReID: Multi-Spectral Object Re-identification with Token Permutation
abstract
Multi-spectral object Re-identification (ReID) aims to retrieve specific objects by leveraging complementary information from different image spectra. It delivers great advantages over traditional single-spectral ReID in complex visual environment. However, the significant distribution gap among different image spectra poses great challenges for effective multi-spectral feature representations. In addition, most of current Transformer-based ReID methods only utilize the global feature of class tokens to achieve the holistic retrieval, ignoring the local discriminative ones. To address the above issues, we step further to utilize all the tokens of Transformers and propose a cyclic token permutation framework for multi-spectral object ReID, dubbled TOP-ReID. More specifically, we first deploy a multi-stream deep network based on vision Transformers to preserve distinct information from different image spectra. Then, we propose a Token Permutation Module (TPM) for cyclic multi-spectral feature aggregation. It not only facilitates the spatial feature alignment across different image spectra, but also allows the class token of each spectrum to perceive the local details of other spectra. Meanwhile, we propose a Complementary Reconstruction Module (CRM), which introduces dense token-level reconstruction constraints to reduce the distribution gap across different image spectra. With the above modules, our proposed framework can generate more discriminative multi-spectral features for robust object ReID. Extensive experiments on three ReID benchmarks (i.e., RGBNT201, RGBNT100 and MSVR310) verify the effectiveness of our methods. The code is available at https://github.com/924973292/TOP-ReID.
Xuehu Liu, Hu Lu, Zhengzheng Tu, Huchuan Lu
AAAI6
2024 Hybrid-SORT: Weak Cues Matter for Online Multi-Object Tracking
abstract
Multi-Object Tracking (MOT) aims to detect and associate all desired objects across frames. Most methods accomplish the task by explicitly or implicitly leveraging strong cues (i.e., spatial and appearance information), which exhibit powerful instance-level discrimination. However, when object occlusion and clustering occur, spatial and appearance information will become ambiguous simultaneously due to the high overlap among objects. In this paper, we demonstrate this long-standing challenge in MOT can be efficiently and effectively resolved by incorporating weak cues to compensate for strong cues. Along with velocity direction, we introduce the confidence and height state as potential weak cues. With superior performance, our method still maintains Simple, Online and Real-Time (SORT) characteristics. Also, our method shows strong generalization for diverse trackers and scenarios in a plug-and-play and training-free manner. Significant and consistent improvements are observed when applying our method to 5 different representative trackers. Further, with both strong and weak cues, our method Hybrid-SORT achieves superior performance on diverse benchmarks, including MOT17, MOT20, and especially DanceTrack where interaction and severe occlusion frequently happen with complex motions. The code and models are available at https://github.com/ymzis69/HybridSORT.
Mingzhan Yang, Guangxin Han, Bin Yan 0004, Jinqing Qi, Huchuan Lu, Dong Wang 0004
AAAI6
2024 TF-CLIP: Learning Text-Free CLIP for Video-Based Person Re-identification
abstract
Large-scale language-image pre-trained models (e.g., CLIP) have shown superior performances on many cross-modal retrieval tasks. However, the problem of transferring the knowledge learned from such models to video-based person re-identification (ReID) has barely been explored. In addition, there is a lack of decent text descriptions in current ReID benchmarks. To address these issues, in this work, we propose a novel one-stage text-free CLIP-based learning framework named TF-CLIP for video-based person ReID. More specifically, we extract the identity-specific sequence feature as the CLIP-Memory to replace the text feature. Meanwhile, we design a Sequence-Specific Prompt (SSP) module to update the CLIP-Memory online. To capture temporal information, we further propose a Temporal Memory Diffusion (TMD) module, which consists of two key components: Temporal Memory Construction (TMC) and Memory Diffusion (MD). Technically, TMC allows the frame-level memories in a sequence to communicate with each other, and to extract temporal information based on the relations within the sequence. MD further diffuses the temporal memories to each token in the original features to obtain more robust sequence features. Extensive experiments demonstrate that our proposed method shows much better results than other state-of-the-art methods on MARS, LS-VID and iLIDS-VID.
Chenyang Yu, Xuehu Liu, Yingquan Wang, Huchuan Lu
AAAI5
2024 DME: Unveiling the Bias for Better Generalized Monocular Depth Estimation
abstract
This paper aims to design monocular depth estimation models with better generalization abilities. To this end, we have conducted quantitative analysis and discovered two important insights. First, the Simulation Correlation phenomenon, commonly seen in long-tailed classification problems, also exists in monocular depth estimation, indicating that the imbalanced depth distribution in training data may be the cause of limited generalization ability. Second, the imbalanced and long-tail distribution of depth values extends beyond the dataset scale, and also manifests within each individual image, further exacerbating the challenge of monocular depth estimation. Motivated by the above findings, we propose the Distance-aware Multi-Expert (DME) depth estimation model. Unlike prior methods that handle different depth range indiscriminately, DME adopts a divide-and-conquer philosophy where each expert is responsible for depth estimation of regions within a specific depth range. As such, the depth distribution seen by each expert is more uniform and can be more easily predicted. A pixel-level routing module is further designed and learned to stitch the prediction of all experts into the final depth map. Experiments show that DME achieves state-of-the-art performance on both NYU-Depth v2 and KITTI, and also delivers favorable zero-shot generalization capability on unseen datasets.
Songsong Yu, Yifan Wang 0004, Yunzhi Zhuge, Lijun Wang 0001, Huchuan Lu
AAAI5
2024 Large Occluded Human Image Completion via Image-Prior Cooperating
abstract
The completion of large occluded human body images poses a unique challenge for general image completion methods. The complex shape variations of human bodies make it difficult to establish a consistent understanding of their structures. Furthermore, as human vision is highly sensitive to human bodies, even slight artifacts can significantly compromise image fidelity. To address these challenges, we propose a large occluded human image completion (LOHC) model based on a novel image-prior cooperative completion strategy. Our model leverages human segmentation maps as a prior, and completes the image and prior simultaneously. Compared to the widely adopted prior-then-image completion strategy for object completion, this cooperative completion process fosters more effective interaction between the prior and image information. Our model consists of two stages. The first stage is a transformer-based auto-regressive network that predicts the overall structure of the missing area by generating a coarse completed image at a lower resolution. The second stage is a convolutional network that refines the coarse images. As the coarse result may not always be accurate, we propose a Dynamic Fusion Module (DFM) to selectively fuses the useful features from the coarse image with the original input at spatial and channel levels. Through extensive experiments, we demonstrate our method’s superior performance compared to state-of-the-art methods.
Hengrun Zhao, Yu Zeng 0001, Huchuan Lu, Lijun Wang 0001
AAAI3
2024 3D Prompt Learning for RGB-D Tracking
Bocen Li, Yunzhi Zhuge, Lijun Wang 0001, Yifan Wang 0004, Huchuan Lu
ACCV (2)6
2024 PsySafe: A Comprehensive Framework for Psychological-based Attack, Defense, and Evaluation of Multi-agent System Safety
abstract
Zaibin Zhang, Yongting Zhang, Lijun Li, Hongzhi Gao, Lijun Wang, Huchuan Lu, Feng Zhao, Yu Qiao, Jing Shao. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Zaibin Zhang, Yongting Zhang, Hongzhi Gao, Yu Qiao 0001, Lijun Wang 0001, Huchuan Lu, Feng Zhao 0004
ACL (1)8
2024 BGDiff: Boundary-Guided Injection Diffusion Framework for Prostate Segmentation
abstract
Recently, the Diffusion Probabilistic Model (DPM)-based methods have achieved substantial success in the field of medical image segmentation. However, most of these methods are not effective in addressing the issue of blurred edges in prostate segmentation tasks. To address this issue, This paper proposes a framework based on the diffusion model, named BGDiff, which is based on Boundary Guided Injection Module(BGIM) and Adaptive Boundary Loss for prostate segmentation. The BGIM can establish connections between the denoising processes of adjacent steps, thereby providing stable guidance for boundary areas as the denoising progresses step by step, while the Adaptive Boundary Loss adjusts the loss weights for more challenging boundaries based on the model’s active feedback. Extensive experiments on four benchmark datasets demonstrate the effectiveness of the proposed method and achieve state-of-the-art performance on four evaluation metrics. The source code will be publicly available at https://github.com/zjlGO/BGDiff
Jialong Zhong, Tingwei Liu, Miao Zhang 0004, Yongri Piao, Weibing Sun, Huchuan Lu
BIBM6
2024 UniPT: Universal Parallel Tuning for Transfer Learning with Efficient Parameter and Memory
abstract
Parameter-efficient transfer learning (PETL), i.e., finetuning a small portion of parameters, is an effective strategy for adapting pre-trained models to downstream domains. To further reduce the memory demand, recent PETL works focus on the more valuable memory-efficient characteristic. In this paper, we argue that the scalability, adaptability, and generalizability of state-of-the-art methods are hindered by structural dependency and pertinency on specific pretrained backbones. To this end, we propose a new memoryefficient PETL strategy, Universal Parallel Tuning (UniPT), to mitigate these weaknesses. Specifically, we facilitate the transfer process via a lightweight and learnable parallel network, which consists of: 1) A parallel interaction module that decouples the sequential connections and processes the intermediate activations detachedly from the pre-trained network. 2) A confidence aggregation module that learns optimal strategies adaptively for integrating cross-layer features. We evaluate UniPT with different backbones (e.g., T5 [69], VSE∞[12], CLIP4Clip [58], Clip-ViL [73], and MDETR [42]) on various vision-and-language and pure NLP tasks. Extensive ablations on 18 datasets have validated that UniPT can not only dramatically reduce memory consumption and outperform the best competitor, but also achieve competitive performance over other plain PETL methods with lower training memory overhead. Our code is publicly available at: https://github.com/Paranioar/UniPT.
Haiwen Diao, Ying Zhang 0021, Xu Jia 0012, Huchuan Lu, Long Chen 0016
CVPR5
2024 Multi-Modal Instruction Tuned LLMs with Fine-Grained Visual Perception
abstract
Multimodal Large Language Model (MLLMs) leverages Large Language Models as a cognitive framework for diverse visual-language tasks. Recent efforts have been made to equip MLLMs with visual perceiving and grounding capabilities. However, there still remains a gap in providing fine-grained pixel-level perceptions and extending interactions beyond text-specific inputs. In this work, we propose AnyRef, a general MLLM model that can generate pixel-wise object perceptions and natural language descriptions from multi-modality references, such as texts, boxes, images, or audio. This innovation empowers users with greater flexibility to engage with the model beyond textual and regional prompts, without modality-specific designs. Through our proposed refocusing mechanism, the generated grounding output is guided to better focus on the referenced object, implicitly incorporating additional pixel-level supervision. This simple modification utilizes attention scores generated during the inference of LLM, eliminating the need for extra computations while exhibiting performance enhancements in both grounding masks and referring expressions. With only publicly available training data, our model achieves state-of-the-art results across multiple benchmarks, including diverse modality referring segmentation and region-level referring expression generation. Code and models are available at https://github.com/jwh97nn/AnyRef
Junwen He, Yifan Wang 0004, Lijun Wang 0001, Huchuan Lu, Jun-Yan He, Jin-Peng Lan, Bin Luo 0008, Xuansong Xie
CVPR4
2024 Boosting Continual Learning of Vision-Language Models via Mixture-of-Experts Adapters
abstract
Continual learning can empower vision-language models to continuously acquire new knowledge, without the need for access to the entire historical dataset. However, mitigating the performance degradation in large-scale models is non-trivial due to (i) parameter shifts throughout life-long learning and (ii) significant computational burdens associated with full-model tuning. In this work, we present a parameter-efficient continual learning framework to alleviate long-term forgetting in incremental learning with vision-language models. Our approach involves the dynamic expansion of a pre-trained CLIP model, through the integration of Mixture-of-Experts (MoE) adapters in response to new tasks. To preserve the zero-shot recognition capability of vision-language models, we further introduce a Distribution Discriminative Auto-Selector (DDAS) that automatically routes in-distribution and out-of-distribution inputs to the MoE Adapter and the original CLIP, respectively. Through extensive experiments across various settings, our proposed method consistently outperforms previous state-of-the-art approaches while concurrently reducing parameter training burdens by 60%. Our code locates at https://github.com/JiazuoYu/MoE-Adapters4CL
Jiazuo Yu 0001, Yunzhi Zhuge, Lu Zhang 0053, Ping Hu 0001, Dong Wang 0004, Huchuan Lu, You He 0002
CVPR6
2024 Multi-View Aggregation Network for Dichotomous Image Segmentation
abstract
Dichotomous Image Segmentation (DIS) has recently emerged towards high-precision object segmentation from high-resolution natural images. When designing an effective DIS model, the main challenge is how to balance the semantic dispersion of high-resolution targets in the small receptive field and the loss of high-precision details in the large receptive field. Existing methods rely on tedious multiple encoder-decoder streams and stages to gradually complete the global localization and local refinement. Human visual system captures regions of interest by observing them from multiple views. Inspired by it, we model DIS as a multi-view object perception problem and provide a parsi-monious multi-view aggregation network (MVANet), which unifies the feature fusion of the distant view and close-up view into a single stream with one encoder-decoder structure. With the help of the proposed multi-view complementary localization and refinement modules, our approach established long-range, profound visual interactions across multiple views, allowing the features of the detailed close-up view to focus on highly slender structures. Experiments on the popular DIS-5K dataset show that our MVANet significantly outperforms state-of-the-art methods in both accuracy and speed. The source code and datasets will be publicly available at MVANet.
Qian Yu 0015, Xiaoqi Zhao 0003, Youwei Pang, Lihe Zhang, Huchuan Lu
CVPR5
2024 Magic Tokens: Select Diverse Tokens for Multi-modal Object Re-Identification
abstract
Single-modal object re-identification (ReID) faces great challenges in maintaining robustness within complex visual scenarios. In contrast, multi-modal object ReID utilizes complementary information from diverse modalities, showing great potentials for practical applications. How-ever, previous methods may be easily affected by irrele-vant backgrounds and usually ignore the modality gaps. To address above issues, we propose a novel learning frame-work named EDITOR to select diverse tokens from vision Transformers for multi-modal object ReID. We be-gin with a shared vision Transformer to extract tokenized features from different input modalities. Then, we intro-duce a Spatial-Frequency Token Selection (SFTS) module to adaptively select object-centric tokens with both spa-tial and frequency information. Afterwards, we employ a Hierarchical Masked Aggregation (HMA) module to fa-cilitate feature interactions within and across modalities. Finally, to further reduce the effect of backgrounds, we propose a Background Consistency Constraint (BCC) and an Object-Centric Feature Refinement (OCFR). They are formulated as two new loss functions, which improve the feature discrimination with background suppression. As a result, our framework can generate more discriminative features for multi-modal object ReID. Extensive ex-periments on three multi-modal ReID benchmarks verify the effectiveness of our methods. The code is available at https://github.com/924973292/EDITOR.
Zhengzheng Tu, Huchuan Lu
CVPR5
2024 Fantastic Animals and Where to Find Them: Segment Any Marine Animal with Dual SAM
abstract
As an important pillar of underwater intelligence, Marine Animal Segmentation (MAS) involves segmenting ani-mals within marine environments. Previous methods don't excel in extracting long-range contextual features and over-look the connectivity between discrete pixels. Recently, Segment Anything Model (SAM) offers a universal frame-workfor general segmentation tasks. Unfortunately, trained with natural images, SAM does not obtain the prior knowl-edge from marine images. In addition, the single-position prompt of SAM is very insufficient for prior guidance. To address these issues, we propose a novel feature learning framework, named Dual-SAM for high-performance MAS. To this end, we first introduce a dual structure with SAM's paradigm to enhance feature learning of marine images. Then, we propose a Multi-level Coupled Prompt (MCP) strategy to instruct comprehensive underwater prior infor-mation, and enhance the multi-level features of SAM's en-coder with adapters. Subsequently, we design a Dilated Fusion Attention Module (DFAM) to progressively inte-grate multi-level features from SAM's encoder. Finally, in-stead of directly predicting the masks of marine animals, we propose a Criss-Cross Connectivity Prediction (C3P) paradigm to capture the inter-connectivity between discrete pixels. With dual decoders, it generates pseudo-labels and achieves mutual supervision for complementary feature rep-resentations, resulting in considerable improvements over previous techniques. Extensive experiments verify that our proposed method achieves state-of-the-art performances on five widely-used MAS datasets. The code is available at https://github.con1IDrchip61IDual_SAM.
Tianyu Yan, Yang Liu 0346, Huchuan Lu
CVPR4
2024 Towards Automatic Power Battery Detection: New Challenge, Benchmark Dataset and Baseline
abstract
We conduct a comprehensive study on a new task named power battery detection (PBD), which aims to localize the dense cathode and anode plates endpoints from X-ray images to evaluate the quality of power batteries. Existing manufacturers usually rely on human eye observation to complete PBD, which makes it difficult to balance the accuracy and efficiency of detection. To address this issue and drive more attention into this meaningful task, we first elaborately collect a dataset, called X-ray PBD, which has 1,500 diverse X-ray images selected from thousands of power batteries of 5 manufacturers, with 7 different visual interference. Then, we propose a novel segmentation-based solution for PBD, termed multi-dimensional collaborative network (MDCNet). With the help of line and counting predictors, the representation of the point segmentation branch can be improved at both semantic and detail aspects. Besides, we design an effective distance-adaptive mask generation strategy, which can alleviate the visual challenge caused by the inconsistent distribution density of plates to provide MDCNet with stable supervision. Without any bells and whistles, our segmentation-based MDCNet consistently outperforms various other corner detection, crowd counting and general/tiny object detection-based so-lutions, making it a strong baseline that can help facilitate future research in PBD. Finally, we share some potential difficulties and works for future researches. The source code and datasets will be publicly available at X-ray PBD.
Xiaoqi Zhao 0003, Youwei Pang, Zhenyu Chen 0001, Qian Yu 0015, Lihe Zhang, Hanqi Liu, Jiaming Zuo, Huchuan Lu
CVPR8
2024 PIXART-Σ: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation
Junsong Chen, Chongjian Ge, Enze Xie, Yue Wu 0002, Lewei Yao, Xiaozhe Ren, Zhongdao Wang, Ping Luo 0002, Huchuan Lu, Zhenguo Li
ECCV (32)9
2024 SHERL: Synthesizing High Accuracy and Efficient Memory for Resource-Limited Transfer Learning
Haiwen Diao, Xu Jia 0012, Yunzhi Zhuge, Ying Zhang 0021, Huchuan Lu, Long Chen 0016
ECCV (44)6
2024 Spatial-Temporal Multi-level Association for Video Object Segmentation
Deshui Miao, Xin Li 0034, Zhenyu He 0001, Huchuan Lu, Ming-Hsuan Yang 0001
ECCV (67)4
2024 Open-Vocabulary Camouflaged Object Segmentation
Youwei Pang, Xiaoqi Zhao 0003, Jiaming Zuo, Lihe Zhang, Huchuan Lu
ECCV (47)5
2024 EvSign: Sign Language Recognition and Translation with Streaming Events
Zeren Wang, Wenyue Chen, Shengming Li, Dong Wang 0004, Huchuan Lu, Xu Jia 0012
ECCV (5)7
2024 Part Representation Learning with Teacher-Student Decoder for Occluded Person Re-Identification
abstract
Occluded person re-identification (ReID) is a very challenging task due to the occlusion disturbance and incomplete target information. Leveraging external cues such as human pose or parsing to locate and align part features has been proven to be very effective in occluded person ReID. Meanwhile, recent Transformer structures have a strong ability of long-range modeling. Considering the above facts, we propose a Teacher-Student Decoder (TSD) framework for occluded person ReID, which utilizes the Transformer decoder with the help of human parsing. More specifically, our proposed TSD consists of a Parsing-aware Teacher Decoder (PTD) and a Standard Student Decoder (SSD). PTD employs human parsing cues to restrict Transformer’s attention and imparts this information to SSD through feature distillation. Thereby, SSD can learn from PTD to aggregate information of body parts automatically. Moreover, a mask generator is designed to provide discriminative regions for better ReID. In addition, existing occluded person ReID benchmarks utilize occluded samples as queries, which will amplify the role of alleviating occlusion interference and underestimate the impact of the feature absence issue. Contrastively, we propose a new benchmark with non-occluded queries, serving as a complement to the existing benchmark. Extensive experiments demonstrate that our proposed method is superior and the new benchmark is essential. The source codes are available at https://github.com/hh23333/TSD.
Shang Gao 0012, Chenyang Yu, Huchuan Lu
ICASSP4
2024 PixArt-α: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis
Junsong Chen, Chongjian Ge, Lewei Yao, Enze Xie, Zhongdao Wang, James T. Kwok, Ping Luo 0002, Huchuan Lu, Zhenguo Li
ICLR9
2024 DepthRefiner: Adapting RGB Trackers to RGBD Scenes via Depth-Fused Refinement
abstract
The increasing availability of depth sensors has facilitated the acquisition of depth images, thereby driving advancements in RGBD tracking. However, compared to RGB benchmarks, deficient data hampers the sufficient learning of RGBD trackers. In this paper, instead of developing a new RGBD tracker from scratch, we aim to learn a depth-fused refinement module that enables existing RGB trackers to adapt to RGBD scenes. Specifically, we introduce a compact yet effective module, named DepthRefiner (DR), based on multi-head self-attention, a simple bimodal fusion technique, and the center-based head. This approach leverages the learned prior representations of RGB trackers from large-scale RGB data and can be flexibly integrated into various off-the-shelf trackers without modifying original pipelines. Comprehensive experiments on CDTB, DepthTrack, VOT-RGBD2022, and RGBD1K benchmarks with multiple base trackers validate that our approach significantly improves the base tracker’s performance while adding minimal computational overhead.
Simiao Lai, Dong Wang 0004, Huchuan Lu
ICME3
2024 Multi-Stage Fusion for Event-based Multimodal Tracker
abstract
Event cameras are bio-inspired sensors with high dynamic range and time resolution, which are favorable properties for visual object tracking. There are already some methods that fuse the event modality and RGB modality with cross-domain feature integrator to achieve improved tracking performance. Researchers have developed some architectures for event modality processing or fusion, successfully boosting the tracking performance. In this work, we design a RGB-E tracker with multi-stage fusion. In the early stage, frames are enhanced with aid of events to mitigate blur or under/over-exposure degradation. During the middle stage, we utilize a fusion module for feature-level integration. At the late stage, we carry out decision-level fusion by predicting tracking boxes based on frame features, event features, and fused features, and the one with highest score is taken as the final estimation. Our design thoroughly integrate information from various levels, allowing each modality to contribute to the tracking process as much as possible. Extensive experiments demonstrate that the proposed method performs favorably against state-of-the-art RGB-E trackers in both accuracy and efficiency.
Xinyu Zhang 0017, Hefei Huang, Xu Jia 0012, Wenyue Chen, Dong Wang 0004, Shengming Li, Huchuan Lu
ICME7
2024 Spider: A Unified Framework for Context-dependent Concept Segmentation
abstract
Different from the context-independent (CI) concepts such as human, car, and airplane, context-dependent (CD) concepts require higher visual understanding ability, such as camouflaged object and medical lesion. Despite the rapid advance of many CD understanding tasks in respective branches, the isolated evolution leads to their limited cross-domain generalisation and repetitive technique innovation. Since there is a strong coupling relationship between foreground and background context in CD tasks, existing methods require to train separate models in their focused domains. This restricts their real-world CD concept understanding towards artificial general intelligence (AGI). We propose a unified model with a single set of parameters, Spider, which only needs to be trained once. With the help of the proposed concept filter driven by the image-mask group prompt, Spider is able to understand and distinguish diverse strong context-dependent concepts to accurately capture the Prompter's intention. Without bells and whistles, Spider significantly outperforms the state-of-the-art specialized models in 8 different context-dependent segmentation tasks, including 4 natural scenes (salient, camouflaged, and transparent objects and shadow) and 4 medical lesions (COVID-19, polyp, breast, and skin lesion with color colonoscopy, CT, ultrasound, and dermoscopy modalities). Besides, Spider shows obvious advantages in continuous learning. It can easily complete the training of new tasks by fine-tuning parameters less than 1% and bring a tolerable performance degradation of less than 5% for all old tasks. The source code will be publicly available at https://github.com/Xiaoqi-Zhao-DLUT/Spider-UniCDSeg.
Xiaoqi Zhao 0003, Youwei Pang, Wei Ji 0011, Baicheng Sheng, Jiaming Zuo, Lihe Zhang, Huchuan Lu
ICML7
2024 DCPT: Darkness Clue-Prompted Tracking in Nighttime UAVs
abstract
Existing nighttime unmanned aerial vehicle (UAV) trackers follow an "Enhance-then-Track" architecture - first using a light enhancer to brighten the nighttime video, then employing a daytime tracker to locate the object. This separate enhancement and tracking fails to build an end-to-end trainable vision system. To address this, we propose a novel architecture called Darkness Clue-Prompted Tracking (DCPT) that achieves robust UAV tracking at night by efficiently learning to generate darkness clue prompts. Without a separate enhancer, DCPT directly encodes anti-dark capabilities into prompts using a darkness clue prompter (DCP). Specifically, DCP iteratively learns emphasizing and undermining projections for darkness clues. It then injects these learned visual prompts into a daytime tracker with fixed parameters across transformer layers. Moreover, a gated feature aggregation mechanism enables adaptive fusion between prompts and between prompts and the base model. Extensive experiments show state-of-the-art performance for DCPT on multiple dark scenario benchmarks. The unified end-to-end learning of enhancement and tracking in DCPT enables a more trainable system. The darkness clue prompting efficiently injects anti-dark knowledge without extra modules. Code is available at https://github.com/bearyi26/DCPT.
Jiawen Zhu 0003, Huayi Tang, Zhi-Qi Cheng, Jun-Yan He, Bin Luo 0008, Shihao Qiu, Shengming Li, Huchuan Lu
ICRA8
2024 MAS-SAM: Segment Any Marine Animal with Aggregated Features
Tianyu Yan, Zifu Wan, Xinhao Deng 0002, Yang Liu 0346, Huchuan Lu
IJCAI6
2024 Safety-First Tracker: A Trajectory Planning Framework for Omnidirectional Robot Tracking
abstract
This paper introduces a Safety-First Tracker (SF-Tracker) designed for omnidirectional autonomous tracking robots. The position and orientation of omnidirectional robots are decoupled for stepwise planning to ensure trajectory safety and maintain target visibility. SF-Tracker puts the trajectory safety in the first place. First, a collision-free and occlusion-free reference path is efficiently initialized by constructing a directed weighted graph. By building upon this path, safe trajectory optimization is implemented to ensure safe movement. Finally, an orientation planner is developed to achieve target visibility based on the safe trajectory. Extensive experimental evaluations in simulated environments and the real world demonstrate that the SF-Tracker outperforms state-of-the-art methods in terms trajectory safety and target visibility. Ablation experiments further demonstrate the significance of each step of the SF-Tracker. The source code and demonstration video can be found at https://github.com/Yue-0/SF-Tracker.
Yang Liu 0003, Xin Chen 0032, Dong Wang 0004, Huchuan Lu
IROS6
2024 CriDiff: Criss-Cross Injection Diffusion Framework via Generative Pre-train for Prostate Segmentation
Tingwei Liu, Miao Zhang 0004, Leiye Liu, Jialong Zhong, Shuyao Wang, Yongri Piao, Huchuan Lu
MICCAI (8)7
2024 Multi-Scale and Detail-Enhanced Segment Anything Model for Salient Object Detection
abstract
Salient Object Detection (SOD) aims to identify and segment the most prominent objects in images. Advanced SOD methods often utilize various Convolutional Neural Networks (CNN) or Transformers for deep feature extraction. However, these methods still deliver low performance and poor generalization in complex cases. Recently, Segment Anything Model (SAM) has been proposed as a visual fundamental model, which gives strong segmentation and generalization capabilities. Nonetheless, SAM requires accurate prompts of target objects, which are unavailable in SOD. Additionally, SAM lacks the utilization of multi-scale and multi-level information, as well as the incorporation of fine-grained details. To address these shortcomings, we propose a Multi-scale and Detail-enhanced SAM (MDSAM) for SOD. Specifically, we first introduce a Lightweight Multi-Scale Adapter (LMSA), which allows SAM to learn multi-scale information with very few trainable parameters. Then, we propose a Multi-Level Fusion Module (MLFM) to comprehensively utilize the multi-level information from the SAM's encoder. Finally, we propose a Detail Enhancement Module (DEM) to incorporate SAM with fine-grained details. Experimental results demonstrate the superior performance of our model on multiple SOD datasets and its strong generalization on other segmentation tasks. The source code is released at https://github.com/BellyBeauty/MDSAM
Shixuan Gao, Tianyu Yan, Huchuan Lu
ACM Multimedia4
2024 Customizing Text-to-Image Generation with Inverted Interaction
abstract
Subject-driven image generation, aimed at customizing user-specified subjects, has experienced rapid progress. However, most of them focus on transferring the customized appearance of subjects. In this work, we consider a novel concept customization task, that is, capturing the interaction between subjects in exemplar images and transferring the learned concept of interaction to achieve customized text-to-image generation. Intrinsically, the interaction between subjects is diverse and is difficult to describe in only a few words. In addition, typical exemplar images are about the interaction between humans, which further intensifies the challenge of interaction-driven image generation with various categories of subjects. To address this task, we adopt a divide-and-conquer strategy and propose a two-stage interaction inversion framework. The framework begins by learning a pseudo-word for a single pose of each subject in the interaction. This is then employed to promote the learning of the concept for the interaction. In addition, language prior and cross-attention loss are incorporated into the optimization process to encourage the modeling of interaction. Extensive experiments demonstrate that the proposed methods are able to effectively invert the interactive pose from exemplar images and apply it to the customized generation with user-specified interaction.
Mengmeng Ge 0002, Xu Jia 0012, Takashi Isobe, Xiaomin Li 0001, Dong Zhou 0003, Li Wang 0125, Huchuan Lu, Ashish Sirasao, Emad Barsoum
ACM Multimedia9
2024 Event-Guided Rolling Shutter Correction with Time-Aware Cross-Attentions
abstract
Many consumer cameras with rolling shutter (RS) CMOS would suffer undesired distortion and artifacts, particularly when objects experiences fast motion. The neuromorphic event camera, with high temporal resolution events, could bring much benefit to the RS correction process. In this work, we explore the characteristics of RS images and event data for the design of the rolling shutter correction (RSC) model. Specifically, the relationship between RS images and event data is modeled by incorporating time encoding to the computation of cross-attention in transformer encoder to achieve time-aware multi-modal information fusion. Features from RS images enhanced by event data are adopted as keys and values in transformer decoder, providing source for appearance, while features from event data enhanced by RS images are adopted as queries, providing spatial transition information. By embedding the time information of the desired global shutter (GS) image into the query, the transformer with deformable attention is capable of producing the target GS image.To enhance the model's generalization ability, we propose to further self-supervise the model by cycling between time coordinate systems corresponding to RS images and GS images. Extensive evaluations over both synthetic and real datasets demonstrate that the proposed method performs favorably against state-of-the-art approaches.
Hefei Huang, Xu Jia 0012, Xinyu Zhang 0017, Shengming Li, Huchuan Lu
ACM Multimedia5
2024 MoTrans: Customized Motion Transfer with Text-driven Video Diffusion Models
abstract
Existing pretrained text-to-video (T2V) models have demonstrated impressive abilities in generating realistic videos with basic motion or camera movement. However, these models exhibit significant limitations when generating intricate, human-centric motions. Current efforts primarily focus on fine-tuning models on a small set of videos containing a specific motion. They often fail to effectively decouple motion and the appearance in the limited reference videos, thereby weakening the modeling capability of motion patterns. To this end, we propose MoTrans, a customized motion transfer method enabling video generation of similar motion in new context. Specifically, we introduce a multimodal large language model (MLLM)-based recaptioner to expand the initial prompt to focus more on appearance and an appearance injection module to adapt appearance prior from video frames to the motion modeling process. These complementary multimodal representations from recaptioned prompt and video frames promote the modeling of appearance and facilitate the decoupling of appearance and motion. In addition, we devise a motion-specific embedding for further enhancing the modeling of the specific motion. Experimental results demonstrate that our method effectively learns specific motion pattern from singular or multiple reference videos, performing favorably against existing methods in customized video generation.
Xiaomin Li 0001, Xu Jia 0012, Haiwen Diao, Mengmeng Ge 0002, You He 0002, Huchuan Lu
ACM Multimedia8
2024 SelM: Selective Mechanism based Audio-Visual Segmentation
abstract
Audio-Visual Segmentation (AVS) aims to segment sound-producing objects in videos according to associated audio cues, where both modalities are affected by noise to different extents, such as the blending of background noises in audio or the presence of distracted objects in video. Most existing methods focus on learning interactions between modalities at high semantic levels but is incapable of filtering low-level noise or achieving fine-grained representational interactions during the early feature extraction phase. Consequently, they struggle with illusion issues, where nonexistent audio cues are erroneously linked to visual objects. In this paper, we present SelM, a novel architecture that leverages selective mechanisms to counteract these illusions. SelM employs State Space model for noise reduction and robust feature selection. By imposing additional bidirectional constraints on audio and visual embeddings, it is able to precisely identify crucial features corresponding to sound-emitting targets. To fill the existing gap in early fusion within AVS, SelM introduces a dual alignment mechanism specifically engineered to facilitate intricate spatio-temporal interactions between audio and visual streams, achieving more fine-grained representations. Moreover, we develop a cross-level decoder for layered reasoning, significantly enhancing segmentation precision by exploring the complex relationships between audio and visual information. SelM achieves state-of-the-art performance in AVS tasks, especially in the challenging Audio-Visual Semantic Segmentation subset. The code can be found at https://github.com/Cyyzpoi/SelM.
Songsong Yu, Yifan Wang 0004, Lijun Wang 0001, Huchuan Lu
ACM Multimedia5
2024 LOVD: Large-and-Open Vocabulary Object Detection
abstract
Existing open-vocabulary object detectors require an accurate and compact vocabulary pre-defined during inference. Their performance is largely degraded in real scenarios where the underlying vocabulary may be indeterminate and often exponentially large. To have a more comprehensive understanding of this phenomenon, we propose a new setting called Large-and-Open Vocabulary object Detection, which simulates real scenarios by testing detectors with large vocabularies containing thousands of unseen categories. The vast unseen categories inevitably lead to an increase in category distractors, severely impeding the recognition process and leading to unsatisfactory detection results. To address this challenge, We propose a Large and Open Vocabulary Detector (LOVD) with two core components, termed the Image-to-Region Filtering (IRF) module and Cross-View Verification (CV2) scheme. To relieve the category distractors of the given large vocabularies, IRF performs image-level recognition to build a compact vocabulary relevant to the image scene out of the large input vocabulary, followed by region-level classification upon the compact vocabulary. CV2 further enhances the IRF by conducting image-to-region filtering in both global and local views and produces the final detection categories through a two-branch voting mechanism. Compared to the prior works, our LOVD is more scalable and robust to large input vocabularies, and can be seamlessly integrated with predominant detection methods to improve their open-vocabulary performance. The code can be found at https://github.com/Altria-luo/LOVD.
Shiyu Tang, Zhaofan Luo, Yifan Wang 0004, Lijun Wang 0001, Huchuan Lu, Weibo Su
ACM Multimedia5
2024 MaskMentor: Unlocking the Potential of Masked Self-Teaching for Missing Modality RGB-D Semantic Segmentation
abstract
Existing RGB-D semantic segmentation methods struggle to handle modality missing input, where only RGB images or depth maps are available, leading to degenerated segmentation performance. We tackle this issue using MaskMentor, a new pre-training framework for modality missing segmentation, which advances its counterparts via two novel designs: Masked Modality and Image Modeling (M2IM), and Self-Teaching via Token-Pixel Joint reconstruction (STTP). M2IM simulates modality missing scenarios by combining both modality- and patch-level random masking. Meanwhile, STTP offers an effective self-teaching strategy, where the trained network assumes a dual role, simultaneously acting as both the teacher and the student. The student with modality missing input is supervised by the teacher with complete modality input through both token- and pixel-wise masked modeling, closing the gap between missing and complete input modalities. By integrating M2IM and STTP, MaskMentor significantly improves the generalization ability of the trained model across diverse input conditions and outperforms state-of-the-art methods on two popular benchmarks by a considerable margin. Extensive ablation studies further verify the effectiveness of the above contributions.
Zhida Zhao, Lijun Wang 0001, Yifan Wang 0004, Huchuan Lu
ACM Multimedia5
2024 Unveiling Encoder-Free Vision-Language Models
abstract
Existing vision-language models (VLMs) mostly rely on vision encoders to extract visual features followed by large language models (LLMs) for visual-language tasks. However, the vision encoders set a strong inductive bias in abstracting visual representation, e.g., resolution, aspect ratio, and semantic priors, which could impede the flexibility and efficiency of the VLMs. Training pure VLMs that accept the seamless vision and language inputs, i.e., without vision encoders, remains challenging and rarely explored. Empirical observations reveal that direct training without encoders results in slow convergence and large performance gaps. In this work, we bridge the gap between encoder-based and encoder-free models, and present a simple yet effective training recipe towards pure VLMs. Specifically, we unveil the key aspects of training encoder-free VLMs efficiently via thorough experiments: (1) Bridging vision-language representation inside one unified decoder; (2) Enhancing visual recognition capability via extra supervision. With these strategies, we launch EVE, an encoder-free vision-language model that can be trained and forwarded efficiently. Notably, solely utilizing 35M publicly accessible data, EVE can impressively rival the encoder-based VLMs of similar capacities across multiple vision-language benchmarks. It significantly outperforms the counterpart Fuyu-8B with mysterious training procedures and undisclosed training data. We believe that EVE provides a transparent and efficient route for developing pure decoder-only architecture across modalities.
Haiwen Diao, Yufeng Cui, Yueze Wang, Huchuan Lu
NeurIPS5
2024 LLMs Can Evolve Continually on Modality for X-Modal Reasoning
Jiazuo Yu 0001, Haomiao Xiong, Lu Zhang 0053, Haiwen Diao, Yunzhi Zhuge, Lanqing Hong, Dong Wang 0004, Huchuan Lu, You He 0002, Long Chen 0016
NeurIPS8
2024 Video Frame Interpolation for Large Motion with Generative Prior
Xu Jia 0012, Lu Zhang 0053, Xiaomin Li 0001, Huchuan Lu
PRCV (10)7
2024 Auto-USOD: Searching Topology for Underwater Salient Object Detection
Tingwei Liu, Runyu Wang, Miao Zhang 0004, Yongri Piao, Huchuan Lu
PRCV (2)5
2024 Edge-Guided Bidirectional-Attention Residual Network for Polyp Segmentation
Lanhu Wu, Miao Zhang 0004, Yongri Piao, Huchuan Lu
PRCV (14)5
2024 Focal Perception Transformer for Light Field Salient Object Detection
Miao Zhang 0004, Yongri Piao, Jihao Yin, Huchuan Lu
PRCV (8)5
2024 Leveraging the Power of Data Augmentation for Transformer-based Tracking
abstract
Due to long-distance correlation and powerful pretrained models, transformer-based methods have initiated a breakthrough in visual object tracking performance. Previous works focus on designing effective architectures suited for tracking, but ignore that data augmentation is equally crucial for training a well-performing model. In this paper, we first explore the impact of general data augmentations on transformer-based trackers via systematic experiments, and reveal the limited effectiveness of these common strategies. Motivated by experimental observations, we then propose two data augmentation methods customized for tracking. First, we optimize existing random cropping via a dynamic search radius mechanism and simulation for boundary samples. Second, we propose a token-level feature mixing augmentation strategy, which enables the model against challenges like background interference. Extensive experiments on two transformer-based trackers and six benchmarks demonstrate the effectiveness and data efficiency of our methods, especially under challenging settings, like one-shot tracking and small image resolutions. Code is available at https://github.com/zj5559/DATr.
Jie Zhao 0014, Johan Edstedt, Michael Felsberg, Dong Wang 0004, Huchuan Lu
WACV5
2024 Learning depth-aware decomposition for single image dehazing
Yumeng Kang, Lu Zhang 0053, Ping Hu 0001, Yu Liu 0005, Huchuan Lu, You He 0002
Comput. Vis. Image Underst.5
2024 Other tokens matter: Exploring global and local features of Vision Transformers for Object Re-Identification
Yingquan Wang, Dong Wang 0004, Huchuan Lu
Comput. Vis. Image Underst.4
2024 Neural image re-exposure
Xinyu Zhang 0017, Hefei Huang, Xu Jia 0012, Dong Wang 0004, Lihe Zhang, Bolun Zheng, Wei Zhou 0021, Huchuan Lu
Comput. Vis. Image Underst.8
2024 Class-conditional domain adaptation for semantic segmentation
abstract
Semantic segmentation is an important sub-task for many applications. However, pixel-level ground-truth labeling is costly, and there is a tendency to overfit to training data, thereby limiting the generalization ability. Unsupervised domain adaptation can potentially address these problems by allowing systems trained on labelled datasets from the source domain (including less expensive synthetic domain) to be adapted to a novel target domain. The conventional approach involves automatic extraction and alignment of the representations of source and target domains globally. One limitation of this approach is that it tends to neglect the differences between classes: representations of certain classes can be more easily extracted and aligned between the source and target domains than others, limiting the adaptation over all classes. Here, we address this problem by introducing a Class-Conditional Domain Adaptation (CCDA) method. This incorporates a class-conditional multi-scale discriminator and class-conditional losses for both segmentation and adaptation. Together, they measure the segmentation, shift the domain in a class-conditional manner, and equalize the loss over classes. Experimental results demonstrate that the performance of our CCDA method matches, and in some cases, surpasses that of state-of-the-art methods.
Yue Wang 0038, James H. Elder, Runmin Wu, Huchuan Lu
Comput. Vis. Media5
2024 Multi-modal visual tracking: Review and experimental comparison
abstract
Visual object tracking has been drawing increasing attention in recent years, as a fundamental task in computer vision. To extend the range of tracking applications, researchers have been introducing information from multiple modalities to handle specific scenes, with promising research prospects for emerging methods and benchmarks. To provide a thorough review of multi-modal tracking, different aspects of multi-modal tracking algorithms are summarized under a unified taxonomy, with specific focus on visible-depth (RGB-D) and visible-thermal (RGB-T) tracking. Subsequently, a detailed description of the related benchmarks and challenges is provided. Extensive experiments were conducted to analyze the effectiveness of trackers on five datasets: PTB, VOT19-RGBD, GTOT, RGBT234, and VOT19-RGBT. Finally, various future directions, including model design and dataset construction, are discussed from different perspectives for further research.
Dong Wang 0004, Huchuan Lu
Comput. Vis. Media3
2024 Adaptive Multi-Source Predictor for Zero-Shot Video Object Segmentation
Xiaoqi Zhao 0003, Shijie Chang, Youwei Pang, Lihe Zhang, Huchuan Lu
Int. J. Comput. Vis.6
2024 Towards Diverse Binary Segmentation via a Simple yet General Gated Network
Xiaoqi Zhao 0003, Youwei Pang, Lihe Zhang, Huchuan Lu, Lei Zhang 0006
Int. J. Comput. Vis.4
2024 CSRNet: Focusing on critical points for depth completion
Bocen Li, Yifan Wang 0004, Lijun Wang 0001, Huchuan Lu
Image Vis. Comput.4
2024 ITrans: generative image inpainting with transformers
abstract
Abstract Despite significant improvements, convolutional neural network (CNN) based methods are struggling with handling long-range global image dependencies due to their limited receptive fields, leading to an unsatisfactory inpainting performance under complicated scenarios. To address this issue, we propose the Inpainting Transformer (ITrans) network, which combines the power of both self-attention and convolution operations. The ITrans network augments convolutional encoder–decoder structure with two novel designs, i.e. , the global and local transformers. The global transformer aggregates high-level image context from the encoder in a global perspective, and propagates the encoded global representation to the decoder in a multi-scale manner. Meanwhile, the local transformer is intended to extract low-level image details inside the local neighborhood at a reduced computational overhead. By incorporating the above two transformers, ITrans is capable of both global relationship modeling and local details encoding, which is essential for hallucinating perceptually realistic images. Extensive experiments demonstrate that the proposed ITrans network outperforms favorably against state-of-the-art inpainting methods both quantitatively and qualitatively.
Wei Miao 0006, Lijun Wang 0001, Huchuan Lu, Kaining Huang, Xinchu Shi, Bocong Liu
Multim. Syst.3
2024 ZoomNeXt: A Unified Collaborative Pyramid Network for Camouflaged Object Detection
abstract
Recent camouflaged object detection (COD) attempts to segment objects visually blended into their surroundings, which is extremely complex and difficult in real-world scenarios. Apart from the high intrinsic similarity between camouflaged objects and their background, objects are usually diverse in scale, fuzzy in appearance, and even severely occluded. To this end, we propose an effective unified collaborative pyramid network that mimics human behavior when observing vague images and videos, i.e., zooming in and out. Specifically, our approach employs the zooming strategy to learn discriminative mixed-scale semantics by the multi-head scale integration and rich granularity perception units, which are designed to fully explore imperceptible clues between candidate objects and background surroundings. The former's intrinsic multi-head aggregation provides more diverse visual patterns. The latter's routing mechanism can effectively propagate inter-frame differences in spatiotemporal scenarios and be adaptively deactivated and output all-zero results for static representations. They provide a solid foundation for realizing a unified architecture for static and dynamic COD. Moreover, considering the uncertainty and ambiguity derived from indistinguishable textures, we construct a simple yet effective regularization, uncertainty awareness loss, to encourage predictions with higher confidence in candidate regions. Our highly task-friendly framework consistently outperforms existing state-of-the-art methods in image and video COD benchmarks.
Youwei Pang, Xiaoqi Zhao 0003, Tian-Zhu Xiang, Lihe Zhang, Huchuan Lu
IEEE Trans. Pattern Anal. Mach. Intell.5
2024 Efficient Adaptive Feature Fusion Network for Remote-Sensing Image Super-Resolution
abstract
Image super-resolution is a fundamental low-level vision task aimed at recovering high-resolution images with fine details. Deep learning has significantly enhanced the performance of super-resolution techniques for remote sensing imagery. However, increasing the depth of networks and the size of their parameters has resulted in substantial computational and storage burdens. To address this challenge, we propose an adaptive approach that learns both local and global information for each region. We introduce a lightweight hybrid model named the Efficient Adaptive Feature Fusion Network, which combines CNNs and Transformers to fully exploit the texture information in remote sensing images. This model leverages local details and long-range dependencies within images in an adaptive manner to achieve superior super-resolution. Specifically, a set of Transformers is employed to model the self-similarity between pixels and perform dense texture pattern predictions at each pixel, while a set of CNNs captures local details within the images. The computed global and local features serve as inputs to the proposed Adaptive Contextual Fusion Block, which learns to fuse local and global information across different regions to generate robust image super-resolution features. We conduct extensive experimental evaluations of the proposed method on the UCMerced and AID datasets, demonstrating its outstanding performance in terms of PSNR and SSIM metrics. Comprehensive experiments validate the effectiveness of our approach, showing that the proposed method achieves an excellent balance between performance and complexity.
Shuai Hao 0007, Shuai Liu 0009, Xu Jia 0012, Huchuan Lu, You He 0002
IEEE Signal Process. Lett.4
2024 LGTrack: Exploiting Local and Global Properties for Robust Visual Tracking
abstract
Re-detection is a necessary capability for long-term tracking. Target candidate proposals in the whole image can provide a chance of tracking reset when tracking fails due to tracking drift or target invisibility. In this paper, we propose a unified local-global tracker based on the same transformer architecture sharing weights, which can not only search in a continuous local region but also provide target candidates of the global image in every frame. The requirements of both long-term and short-term scenarios can be addressed using a unified model. A simple proposal selection scheme is adopted to properly select the candidate proposals of re-detection, to assist tracking and obtain better performance. The scheme performs reevaluation of all high-quality proposals based on a transformer-based embedding network, once the predicted state of the local tracking is not sufficient to be accurate. To capture appearance variations brought by online updates in minimum risks, a long-term-friendly dynamic template update scheme is also designed. Extensive experiments are conducted to demonstrate the effectiveness of our proposed tracker, including three short-term tracking benchmarks and six long-term benchmarks. Our tracker can achieve results comparable to that of the state-of-the-art. The proposed tracker can also work well in balancing the performance and speed, achieving an average speed of approximately 25 fps tested on LaSOT testing set.
Chang Liu 0071, Jie Zhao 0014, Chunjuan Bo, Shengming Li, Dong Wang 0004, Huchuan Lu
IEEE Trans. Circuits Syst. Video Technol.6
2024 Learning Local-Global Representation for Scribble-Based RGB-D Salient Object Detection via Transformer
abstract
Manual scribbles have been introduced to RGB-D Salient Object Detection (SOD) as a credible indicator for salient regions and backgrounds, helping to strike a balance between detection accuracy and labeling efficiency. Previous works address this task by constructing loss functions on semantics, edges, and structures to distinguish salient pixels from the background. However, using local representations extracted by CNNs or Transformers and the incomplete scribble annotations are ineffective in capturing the global contexts of salient objects, and thus cause inaccurate predictions in cluttered regions. In this paper, we propose a local-global representation learning framework by incorporating multi-perception information to boost scribble-based RGB-D SOD. Our system is composed of three sub-modules: Local Representation Aggregation (LRA), Global Representation Initialization (GRI) and Dual Transformer Decoder (DTD). The LRA module first conducts integration of multi-scale, multi-modal local representations extracted from RGB images and depth maps. The GRI module then learns inter- and intra-image representations to capture the global contexts of salient regions from different aspects. Finally, the DTD module alternately updates local-global representations through a dual Transformer architecture. Experimental results on six benchmarks demonstrate that the proposed method performs favorably against state-of-the-art scribble-based RGB-D SOD approaches and is competitive with the fully-supervised approaches.
Yue Wang 0038, Lu Zhang 0053, Yunzhi Zhuge, Huchuan Lu
IEEE Trans. Circuits Syst. Video Technol.7
2024 SRRT: Exploring Search Region Regulation for Visual Object Tracking
abstract
The dominant trackers generate a fixed-size rectangular region based on the previous prediction or initial bounding box as the model input, i.e., search region. While this manner obtains promising tracking efficiency, a fixed-size search region lacks flexibility and is likely to fail in some cases, e.g., fast motion and distractor interference. Trackers tend to lose the target object due to the limited search region or experience interference from distractors due to the excessive search region. Drawing inspiration from the pattern humans track an object, we propose a novel tracking paradigm, called Search Region Regulation Tracking (SRRT) that applies a small eyereach when the target is captured and zooms out the search field when the target is about to be lost. SRRT applies a proposed search region regulator to estimate an optimal search region dynamically for each frame, by which the tracker can flexibly respond to transient changes in the location of object occurrences. To adapt the object’s appearance variation during online tracking, we further propose a locking-state determined updating strategy for reference frame updating. The proposed SRRT is concise without bells and whistles, yet achieves evident improvements and competitive results with other state-of-the-art trackers on eight benchmarks. On the large-scale LaSOT benchmark, SRRT improves SiamRPN++ and TransT with absolute gains of 4.6% and 3.1% in terms of AUC. The code and models will be released.
Jiawen Zhu 0003, Xin Chen 0032, Xinying Wang 0005, Dong Wang 0004, Wenda Zhao 0003, Huchuan Lu
IEEE Trans. Circuits Syst. Video Technol.7
2024 Center-Wise Feature Consistency Learning for Long-Tailed Remote Sensing Object Recognition
abstract
Long-tailed distribution of remote sensing data generally limits the object recognition performance of deep neural networks. We notice that too many samples from head class will induce the neural network to learn features of tail class samples being biased towards the head. To solve this, we propose a novel center-wise feature consistency learning (CFCL) mechanism for long-tailed remote sensing object recognition. Firstly, we implement a head-tail center feature generation procedure that builds two teacher models to extract the knowledge from the head class and tail class samples respectively, so as to avoid the extracted tail class features being affected by the head classes. Secondly, a center-wise feature consistency learning strategy is introduced, which distills the central feature of each class to a student model, thereby making the classification boundaries more prominent. Especially, the central feature is estimated by referring to the features which are correctly classified by the teacher models, thus the inaccurate knowledge is abandoned. Extensive experiments on widely-adopted remote sensing recognition datasets including FGSC-23, DIOR, xView and HRSC2016 demonstrate that our method achieves superior performance compared to the state-of-the-art approaches.Code and data are available at: https://github.com/wdzhao123/CWFC.
Wenda Zhao 0003, Zhepu Zhang, Jiani Liu 0004, Yu Liu 0005, You He 0002, Huchuan Lu
IEEE Trans. Geosci. Remote. Sens.6
2024 Deep Boosting Learning: A Brand-New Cooperative Approach for Image-Text Matching
abstract
Image-text matching remains a challenging task due to heterogeneous semantic diversity across modalities and insufficient distance separability within triplets. Different from previous approaches focusing on enhancing multi-modal representations or exploiting cross-modal correspondence for more accurate retrieval, in this paper we aim to leverage the knowledge transfer between peer branches in a boosting manner to seek a more powerful matching model. Specifically, we propose a brand-new Deep Boosting Learning (DBL) algorithm, where an anchor branch is first trained to provide insights into the data properties, with a target branch gaining more advanced knowledge to develop optimal features and distance metrics. Concretely, an anchor branch initially learns the absolute or relative distance between positive and negative pairs, providing a foundational understanding of the particular network and data distribution. Building upon this knowledge, a target branch is concurrently tasked with more adaptive margin constraints to further enlarge the relative distance between matched and unmatched samples. Extensive experiments validate that our DBL can achieve impressive and consistent improvements based on various recent state-of-the-art models in the image-text matching field, and outperform related popular cooperative strategies, e.g., Conventional Distillation, Mutual Learning, and Contrastive Learning. Beyond the above, we confirm that DBL can be seamlessly integrated into their training scenarios and achieve superior performance under the same computational costs, demonstrating the flexibility and broad applicability of our proposed method.
Haiwen Diao, Ying Zhang 0021, Shang Gao 0012, Xiang Ruan, Huchuan Lu
IEEE Trans. Image Process.5
2024 GSSF: Generalized Structural Sparse Function for Deep Cross-Modal Metric Learning
abstract
Cross-modal metric learning is a prominent research topic that bridges the semantic heterogeneity between vision and language. Existing methods frequently utilize simple cosine or complex distance metrics to transform the pairwise features into a similarity score, which suffers from an inadequate or inefficient capability for distance measurements. Consequently, we propose a Generalized Structural Sparse Function to dynamically capture thorough and powerful relationships across modalities for pair-wise similarity learning while remaining concise but efficient. Specifically, the distance metric delicately encapsulates two formats of diagonal and block-diagonal terms, automatically distinguishing and highlighting the cross-channel relevancy and dependency inside a structured and organized topology. Hence, it thereby empowers itself to adapt to the optimal matching patterns between the paired features and reaches a sweet spot between model complexity and capability. Extensive experiments on cross-modal and two extra uni-modal retrieval tasks (image-text retrieval, person re-identification, fine-grained image retrieval) have validated its superiority and flexibility over various popular retrieval frameworks. More importantly, we further discover that it can be seamlessly incorporated into multiple application scenarios, and demonstrates promising prospects from Attention Mechanism to Knowledge Distillation in a plug-and-play manner.
Haiwen Diao, Ying Zhang 0021, Shang Gao 0012, Jiawen Zhu 0003, Long Chen 0016, Huchuan Lu
IEEE Trans. Image Process.6
2024 Event-Assisted Blurriness Representation Learning for Blurry Image Unfolding
abstract
The goal of blurry image deblurring and unfolding task is to recover a single sharp frame or a sequence from a blurry one. Recently, its performance is greatly improved with introduction of a bio-inspired visual sensor, event camera. Most existing event-assisted deblurring methods focus on the design of powerful network architectures and effective training strategy, while ignoring the role of blur modeling in removing various blur in dynamic scenes. In this work, we propose to implicitly model blur in an image by computing blurriness representation with an event-assisted blurriness encoder. The learning of blurriness representation is formulated as a ranking problem based on specially synthesized pairs. Blurriness-aware image unfolding is achieved by integrating blur relevant information contained in the representation into a base unfolding network. The integration is mainly realized by the proposed blurriness-guided modulation and multi-scale aggregation modules. Experiments on GOPRO and HQF datasets show favorable performance of the proposed method against state-of-the-art approaches. More results on real-world data validate its effectiveness in recovering a sequence of latent sharp frames from a blurry image.
Hao Ju 0004, Lei Yu 0006, Weihua He, Yaoyuan Wang, Qi Xu 0008, Shengming Li, Dong Wang 0004, Huchuan Lu, Xu Jia 0012
IEEE Trans. Image Process.10
2024 A Video Is Worth Three Views: Trigeminal Transformers for Video-Based Person Re-Identification
abstract
Video-based person Re-Identification (Re-ID) is a hot research topic in intelligent transportation systems, which aims to retrieve video sequences of the same person under non-overlapping surveillance cameras. Compared with static images, video sequences contain more visual information from multiple views, such as spatial and temporal views. However, previous Re-ID methods usually focus on single limited views, lacking diverse observations from different views. To capture richer perceptions and extract more comprehensive representations, we propose a novel learning framework namedTrigeminal Transformers (TMT)to tackle video-based person Re-ID. More specifically, we first design aView-wise Projector (VP)to jointly transform raw videos from spatial, temporal and spatial-temporal views. In addition, inspired by the great success of Vision Transformers (ViT), we introduce the Transformer structure for information enhancement and aggregation. In our work, threeSelf-view Transformers (ST)are proposed to exploit the relationships of local features for information enhancement in spatial, temporal and spatial-temporal. Moreover, aCross-view Transformer (CT)is proposed to aggregate the multi-view features for comprehensive representations. Experimental results indicate that our approach can obtain better performance than some other state-of-the-art approaches on four public Re-ID benchmarks.
Xuehu Liu, Chenyang Yu, Xuesheng Qian, Xiaoyun Yang, Huchuan Lu
IEEE Trans. Intell. Transp. Syst.6
2024 GroupMorph: Medical Image Registration via Grouping Network With Contextual Fusion
abstract
Pyramid-based deformation decomposition is a promising registration framework, which gradually decomposes the deformation field into multi-resolution subfields for precise registration. However, most pyramid-based methods directly produce one subfield per resolution level, which does not fully depict the spatial deformation. In this paper, we propose a novel registration model, called GroupMorph. Different from typical pyramid-based methods, we adopt the grouping-combination strategy to predict deformation field at each resolution. Specifically, we perform group-wise correlation calculation to measure the similarities of grouped features. After that, n groups of deformation subfields with different receptive fields are predicted in parallel. By composing these subfields, a deformation field with multi-receptive field ranges is formed, which can effectively identify both large and small deformations. Meanwhile, a contextual fusion module is designed to fuse the contextual features and provide the inter-group information for the field estimator of the next level. By leveraging the inter-group correspondence, the synergy among deformation subfields is enhanced. Extensive experiments on four public datasets demonstrate the effectiveness of GroupMorph. Code is available at https://github.com/TVayne/GroupMorph.
Zuopeng Tan, Lihe Zhang, Yanan Lv, Yili Ma, Huchuan Lu
IEEE Trans. Medical Imaging5
2024 Attacking Defocus Detection With Blur-Aware Transformation for Defocus Deblurring
abstract
Previous fully-supervised defocus deblurring has made significant progress. However, training such deep models requires abundant paired ground truth, which is expensive and error-prone. This paper makes an attempt to train a defocus deblurring model without using paired ground truth and any other unpaired data. Related reblur-to-deblur schemes generally use physics-based reblur or GAN-based reblur, suffering from the robustness of blur kernel and hallucination generated by GAN. Besides, the domain gap between the realistic blurred image and reblurred image hinders deblurring performance. Addressing these challenges, we propose a weakly-supervised defocus deblurring framework via defocus detection attack. On one hand, we build a focused area detection attack (FADA) to enforce the focused area to reblur, thereby reversing its detection result by a pretrained defocus blur detection network. Moreover, we introduce a blur-aware transfer modulated from the defocused region to help FADA render a robust reblurred region. On the other hand, we implement a defocused region detection attack to guide the realistic blurred region to deblur in the process of training deblurring network with simulated-paired areas. Extensive experiments on three widely-used datasets verify the effectiveness of our framework.
Wenda Zhao 0003, Fei Wei, You He 0002, Huchuan Lu
IEEE Trans. Multim.6
2024 Deformable Dynamic Sampling and Dynamic Predictable Mask Mining for Image Inpainting
abstract
Existing image inpainting methods often produce artifacts that are caused by using vanilla convolution layers as building blocks that treat all image regions equally and generate holes at random locations with equal probability. This design does not differentiate the missing regions and valid regions in inference and does not consider the predictability of missing regions in training. To address these issues, we propose a deformable dynamic sampling (DDS) mechanism which is built on deformable convolutions (DCs), and a constraint is proposed to avoid the deformably sampled elements falling into the corrupted regions. Furthermore, to select both valid sample locations and suitable kernels dynamically, we equip DCs with content-aware dynamic kernel selection (DKS). In addition, to further encourage the DDS mechanism to find meaningful sampling locations, we propose to train the inpainting model with mined predictable regions as holes. During training, we jointly train a mask generator with the inpainting network to generate hole masks dynamically for each training sample. Thus, the mask generator can find large yet predictable missing regions as a better alternative to random masks. Extensive experiments demonstrate the advantages of our method over state-of-the-art methods qualitatively and quantitatively.
Cai Cai, Yu Zeng 0001, Shu Yang 0004, Xu Jia 0012, Huchuan Lu, You He 0002
IEEE Trans. Neural Networks Learn. Syst.5
2024 Learning From Box Annotations for Referring Image Segmentation
abstract
Referring image segmentation (RIS) has obtained an impressive achievement by fully convolutional networks (FCNs). However, previous RIS methods require a large number of pixel-level annotations. In this article, we present a weakly supervised RIS method by using bounding box (BB) annotations. In the first stage, we introduce an adversarial boundary loss to extract the object contour from the BB, which is then used to select appropriate region proposals for pseudoground-truth (PGT) generation. In the second stage, we design a co-training (Co-T) strategy to purify the pseudolabels. Specifically, we train two networks and interactively guide them to pick clean labels for each other's networks, which can weaken the effect of noisy labels on model training. Experiment results on four benchmark datasets demonstrate that the proposed method can produce high-quality masks with a speed of 63 frames/s.
Lihe Zhang, Zhiwei Hu, Huchuan Lu
IEEE Trans. Neural Networks Learn. Syst.4
2024 Self-Supervised Tracking via Target-Aware Data Synthesis
abstract
While deep-learning-based tracking methods have achieved substantial progress, they entail large-scale and high-quality annotated data for sufficient training. To eliminate expensive and exhaustive annotation, we study self-supervised (SS) learning for visual tracking. In this work, we develop the crop-transform-paste operation, which is able to synthesize sufficient training data by simulating various appearance variations during tracking, including appearance variations of objects and background interference. Since the target state is known in all synthesized data, existing deep trackers can be trained in routine ways using the synthesized data without human annotation. The proposed target-aware data-synthesis method adapts existing tracking approaches within a SS learning framework without algorithmic changes. Thus, the proposed SS learning mechanism can be seamlessly integrated into existing tracking frameworks to perform training. Extensive experiments show that our method: 1) achieves favorable performance against supervised (Su) learning schemes under the cases with limited annotations; 2) helps deal with various tracking challenges such as object deformation, occlusion (OCC), or background clutter (BC) due to its manipulability; 3) performs favorably against the state-of-the-art unsupervised tracking methods; and 4) boosts the performance of various state-of-the-art Su learning frameworks, including SiamRPN++, DiMP, and TransT.
Xin Li 0034, Wenjie Pei, Yaowei Wang 0001, Zhenyu He 0001, Huchuan Lu, Ming-Hsuan Yang 0001
IEEE Trans. Neural Networks Learn. Syst.5
2024 Deeply Coupled Convolution-Transformer With Spatial-Temporal Complementary Learning for Video-Based Person Re-Identification
abstract
Advanced deep convolutional neural networks (CNNs) have shown great success in video-based person re-identification (Re-ID). However, they usually focus on the most obvious regions of persons with a limited global representation ability. Recently, it witnesses that Transformers explore the interpatch relationships with global observations for performance improvements. In this work, we take both the sides and propose a novel spatial-temporal complementary learning framework named deeply coupled convolution-transformer (DCCT) for high-performance video-based person Re-ID. First, we couple CNNs and Transformers to extract two kinds of visual features and experimentally verify their complementarity. Furthermore, in spatial, we propose a complementary content attention (CCA) to take advantages of the coupled structure and guide independent features for spatial complementary learning. In temporal, a hierarchical temporal aggregation (HTA) is proposed to progressively capture the interframe dependencies and encode temporal information. Besides, a gated attention (GA) is used to deliver aggregated temporal information into the CNN and Transformer branches for temporal complementary learning. Finally, we introduce a self-distillation training strategy to transfer the superior spatial-temporal knowledge to backbone networks for higher accuracy and more efficiency. In this way, two kinds of typical features from same videos are integrated mechanically for more informative representations. Extensive experiments on four public Re-ID benchmarks demonstrate that our framework could attain better performances than most state-of-the-art methods.
Xuehu Liu, Chenyang Yu, Huchuan Lu
IEEE Trans. Neural Networks Learn. Syst.4
2024 Referring Image Segmentation With Fine-Grained Semantic Funneling Infusion
abstract
Recently, referring image segmentation has attracted wide attention given its huge potential in human-robot interaction. Networks to identify the referred region must have a deep understanding of both the image and language semantics. To do so, existing works tend to design various mechanisms to achieve cross-modality fusion, for example, tile and concatenation and vanilla nonlocal manipulation. However, the plain fusion usually is either coarse or constrained by the exorbitant computation overhead, finally causing not enough understanding of the referent. In this work, we propose a fine-grained semantic funneling infusion (FSFI) mechanism to solve the problem. The FSFI introduces a constant spatial constraint on the querying entities from different encoding stages and dynamically infuses the gleaned language semantic into the vision branch. Moreover, it decomposes the features from different modalities into more delicate components, allowing the fusion to happen in multiple low-dimensional spaces. The fusion is more effective than the one only happening in one high-dimensional space, given its ability to sink more representative information along the channel dimension. Another problem haunting the task is that the instilling of high-abstract semantic will blur the details of the referent. Targetedly, we propose a multiscale attention-enhanced decoder (MAED) to alleviate the problem. We design a detail enhancement operator (DeEh) and apply it in a multiscale and progressive way. Features from the higher level are used to generate attention guidance to enlighten the lower-level features to more attend to the detail regions. Extensive results on the challenging benchmarks show that our network performs favorably against the state-of-the-arts (SOTAs).
Lihe Zhang, Huchuan Lu
IEEE Trans. Neural Networks Learn. Syst.3
2024 Defocus Blur Detection Attack via Mutual-Referenced Feature Transfer
abstract
Benefiting from deep learning, defocus blur detection (DBD) has made prominent progress. Existing DBD methods generally study multiscale and multilevel features to improve performance. In this article, from a different perspective, we explore to generate confrontational images to attack DBD network. Based on the observation that defocus area and focus region in an image can provide mutual feature reference to help improve the quality of the confrontational image, we propose a novel mutual-referenced attack framework. Firstly, we design a divide-and-conquer perturbation image generation model, where the focus region attack image and defocus area attack image are generated respectively. Then, we integrate mutual-referenced feature transfer (MRFT) models to improve attack performance. Comprehensive experiments are provided to verify the effectiveness of our method. Moreover, related applications of our study are presented, e.g., sample augmentation to improve DBD and paired sample generation to boost defocus deblurring.
Wenda Zhao 0003, Fei Wei, You He 0002, Huchuan Lu
IEEE Trans. Neural Networks Learn. Syst.6
2024 Interactive Feature Embedding for Infrared and Visible Image Fusion
abstract
General deep learning-based methods for infrared and visible image fusion rely on the unsupervised mechanism for vital information retention by utilizing elaborately designed loss functions. However, the unsupervised mechanism depends on a well-designed loss function, which cannot guarantee that all vital information of source images is sufficiently extracted. In this work, we propose a novel interactive feature embedding in a self-supervised learning framework for infrared and visible image fusion, attempting to overcome the issue of vital information degradation. With the help of a self-supervised learning framework, hierarchical representations of source images can be efficiently extracted. In particular, interactive feature embedding models are tactfully designed to build a bridge between self-supervised learning and infrared and visible image fusion learning, achieving vital information retention. Qualitative and quantitative evaluations exhibit that the proposed method performs favorably against state-of-the-art methods.
Fan Zhao 0005, Wenda Zhao 0003, Huchuan Lu
IEEE Trans. Neural Networks Learn. Syst.3
2024 Confusion Region Mining for Crowd Counting
abstract
Existing works mainly focus on crowd and ignore the confusion regions which contain extremely similar appearance to crowd in the background, while crowd counting needs to face these two sides at the same time. To address this issue, we propose a novel end-to-end trainable confusion region discriminating and erasing network called CDENet. Specifically, CDENet is composed of two modules of confusion region mining module (CRM) and guided erasing module (GEM). CRM consists of basic density estimation (BDE) network, confusion region aware bridge and confusion region discriminating network. The BDE network first generates a primary density map, and then the confusion region aware bridge excavates the confusion regions by comparing the primary prediction result with the ground-truth density map. Finally, the confusion region discriminating network learns the difference of feature representations in confusion regions and crowds. Furthermore, GEM gives the refined density map by erasing the confusion regions. We evaluate the proposed method on four crowd counting benchmarks, including ShanghaiTech Part_A, ShanghaiTech Part_B, UCF_CC_50, and UCF-QNRF, and our CDENet achieves superior performance compared with the state-of-the-arts.
Jiawen Zhu 0003, Wenda Zhao 0003, Libo Yao, You He 0002, Maodi Hu, Huchuan Lu
IEEE Trans. Neural Networks Learn. Syst.9
2023 Dual Memory Aggregation Network for Event-Based Object Detection with Learnable Representation
abstract
Event-based cameras are bio-inspired sensors that capture brightness change of every pixel in an asynchronous manner. Compared with frame-based sensors, event cameras have microsecond-level latency and high dynamic range, hence showing great potential for object detection under high-speed motion and poor illumination conditions. Due to sparsity and asynchronism nature with event streams, most of existing approaches resort to hand-crafted methods to convert event data into 2D grid representation. However, they are sub-optimal in aggregating information from event stream for object detection. In this work, we propose to learn an event representation optimized for event-based object detection. Specifically, event streams are divided into grids in the x-y-t coordinates for both positive and negative polarity, producing a set of pillars as 3D tensor representation. To fully exploit information with event streams to detect objects, a dual-memory aggregation network (DMANet) is proposed to leverage both long and short memory along event streams to aggregate effective information for object detection. Long memory is encoded in the hidden state of adaptive convLSTMs while short memory is modeled by computing spatial-temporal correlation between event pillars at neighboring time intervals. Extensive experiments on the recently released event-based automotive detection dataset demonstrate the effectiveness of the proposed method.
Xu Jia 0012, Xinyu Zhang 0017, Yaoyuan Wang, Dong Wang 0004, Huchuan Lu
AAAI8
2023 Universal Instance Perception as Object Discovery and Retrieval
abstract
All instance perception tasks aim at finding certain objects specified by some queries such as category names, language expressions, and target annotations, but this complete field has been split into multiple independent sub-tasks. In this work, we present a universal instance perception model of the next generation, termed UNINEXT. UNINEXT reformulates diverse instance perception tasks into a unified object discovery and retrieval paradigm and can flexibly perceive different types of objects by simply changing the input prompts. This unified formulation brings the following benefits: (1) enormous data from different tasks and label vocabularies can be exploited for jointly training general instance-level representations, which is especially beneficial for tasks lacking in training data. (2) the unified model is parameter-efficient and can save redundant computation when handling multiple tasks simultaneously. UNINEXT shows superior performance on 20 challenging benchmarks from 10 instance-level tasks including classical image-level tasks (object detection and instance segmentation), vision-and-language tasks (referring expression comprehension and segmentation), and six video-level object tracking tasks. Code is available at https://github.com/MasterBin-IIAU/UNINEXT.
Bin Yan 0004, Yi Jiang 0009, Jiannan Wu, Dong Wang 0004, Ping Luo 0002, Zehuan Yuan, Huchuan Lu
CVPR7
2023 SeqTrack: Sequence to Sequence Learning for Visual Object Tracking
abstract
In this paper, we present a new sequence-to-sequence learning framework for visual tracking, dubbed SeqTrack. It casts visual tracking as a sequence generation problem, which predicts object bounding boxes in an autoregressive fashion. This is different from prior Siamese trackers and transformer trackers, which rely on designing complicated head networks, such as classification and regression heads. SeqTrack only adopts a simple encoder-decoder transformer architecture. The encoder extracts visual features with a bidirectional transformer, while the decoder generates a sequence of bounding box values autoregressively with a causal transformer. The loss function is a plain cross-entropy. Such a sequence learning paradigm not only simplifies tracking framework, but also achieves competitive performance on benchmarks. For instance, SeqTrack gets 72.5% AUC on LaSOT, establishing a new state-of-the-art performance. Code and models are available at https://github.com/microsoft/VideoX.
Xin Chen 0032, Houwen Peng, Dong Wang 0004, Huchuan Lu, Han Hu 0001
CVPR4
2023 GM-NeRF: Learning Generalizable Model-Based Neural Radiance Fields from Multi-View Images
abstract
In this work, we focus on synthesizing high-fidelity novel view images for arbitrary human performers, given a set of sparse multi-view images. It is a challenging task due to the large variation among articulated body poses and heavy self-occlusions. To alleviate this, we introduce an effective generalizable framework Generalizable Model-based Neural Radiance Fields (GM-NeRF) to synthesize free-viewpoint images. Specifically, we propose a geometry-guided attention mechanism to register the appearance code from multi-view 2D images to a geometry proxy which can alleviate the misalignment between inaccurate geometry prior and pixel space. On top of that, we further conduct neural rendering and partial gradient back propagation for efficient perceptual supervision and improvement of the perceptual quality of synthesis. To evaluate our method, we conduct experiments on synthesized datasets THuman2.0 and Multi-garment, and real-world datasets Genebody and ZJUMocap. The results demonstrate that our approach outperforms state-of-the-art methods in terms of novel view synthesis and geometric reconstruction.
Jianchuan Chen, Wentao Yi, Liqian Ma, Xu Jia 0012, Huchuan Lu
CVPR5
2023 Compression-Aware Video Super-Resolution
abstract
Videos stored on mobile devices or delivered on the Internet are usually in compressed format and are of various unknown compression parameters, but most video super-resolution (VSR) methods often assume ideal inputs resulting in large performance gap between experimental settings and real-world applications. In spite of a few pioneering works being proposed recently to super-resolve the compressed videos, they are not specially designed to deal with videos of various levels of compression. In this paper, we propose a novel and practical compression-aware video super-resolution model, which could adapt its video enhancement process to the estimated compression level. A compression encoder is designed to model compression levels of input frames, and a base VSR model is then conditioned on the implicitly computed representation by inserting compression-aware modules. In addition, we propose to further strengthen the VSR model by taking full advantage of meta data that is embedded naturally in compressed video streams in the procedure of information fusion. Extensive experiments are conducted to demonstrate the effectiveness and efficiency of the proposed method on compressed VSR benchmarks. The codes will be available at https://github.com/aprBlue/CAVSR
Takashi Isobe, Xu Jia 0012, Xin Tao 0001, Huchuan Lu, Yu-Wing Tai
CVPR5
2023 ARKitTrack: A New Diverse Dataset for Tracking Using Mobile RGB-D Data
abstract
Compared with traditional RGB-only visual tracking, few datasets have been constructed for RGB-D tracking. In this paper, we propose ARKitTrack, a new RGB-D tracking dataset for both static and dynamic scenes captured by consumer-grade LiDAR scanners equipped on Apple's iPhone and iPad. ARKitTrack contains 300 RGB-D sequences, 455 targets, and 229.7K video frames in total. Along with the bounding box annotations and frame-level attributes, we also annotate this dataset with 123.9K pixel-level target masks. Besides, the camera intrinsic and camera pose of each frame are provided for future developments. To demonstrate the potential usefulness of this dataset, we further present a unified baseline for both box-level and pixel-level tracking, which integrates RGB features with bird's-eye-view representations to better explore cross-modality 3D geometry. In-depth empirical analysis has verified that the ARKitTrack dataset can significantly facilitate RGB-D tracking and that the proposed baseline method compares favorably against the state of the arts. The code and dataset is available at https://arkittrack.github.io.
Haojie Zhao, Junsong Chen, Lijun Wang 0001, Huchuan Lu
CVPR4
2023 Representation Learning for Visual Object Tracking by Masked Appearance Transfer
abstract
Visual representation plays an important role in visual object tracking. However, few works study the tracking-specified representation learning method. Most trackers directly use ImageNet pre-trained representations. In this paper, we propose masked appearance transfer, a simple but effective representation learning method for tracking, based on an encoder-decoder architecture. First, we encode the visual appearances of the template and search region jointly, and then we decode them separately. During decoding, the original search region image is reconstructed. However, for the template, we make the decoder reconstruct the target appearance within the search region. By this target appearance transfer, the tracking-specified representations are learned. We randomly mask out the inputs, thereby making the learned representations more discriminative. For sufficient evaluation, we design a simple and lightweight tracker that can evaluate the representation for both target localization and box regression. Extensive experiments show that the proposed method is effective, and the learned representations can enable the simple tracker to obtain state-of-the-art performance on six datasets. https://github.com/difhnp/MAT
Haojie Zhao, Dong Wang 0004, Huchuan Lu
CVPR3
2023 MetaFusion: Infrared and Visible Image Fusion via Meta-Feature Embedding from Object Detection
abstract
Fusing infrared and visible images can provide more texture details for subsequent object detection task. Conversely, detection task furnishes object semantic information to improve the infrared and visible image fusion. Thus, a joint fusion and detection learning to use their mutual promotion is attracting more attention. However, the feature gap between these two different-level tasks hinders the progress. Addressing this issue, this paper proposes an infrared and visible image fusion via meta-feature embedding from object detection. The core idea is that meta-feature embedding model is designed to generate object semantic features according to fusion network ability, and thus the semantic features are naturally compatible with fusion features. It is optimized by simulating a meta learning. Moreover, we further implement a mutual promotion learning between fusion and detection tasks to improve their performances. Comprehensive experiments on three public datasets demonstrate the effectiveness of our method. Code and model are available at: https://github.com/wdzhao123/MetaFusion.
Wenda Zhao 0003, Shigeng Xie, Fan Zhao 0005, You He 0002, Huchuan Lu
CVPR5
2023 Visual Prompt Multi-Modal Tracking
abstract
Visible-modal object tracking gives rise to a series of downstream multi-modal tracking tributaries. To inherit the powerful representations of the foundation model, a natural modus operandi for multi-modal tracking is full fine-tuning on the RGB-based parameters. Albeit effective, this manner is not optimal due to the scarcity of downstream data and poor transferability, etc. In this paper, inspired by the recent success of the prompt learning in language models, we develop Visual Prompt multi-modal Tracking (ViPT), which learns the modal-relevant prompts to adapt the frozen pre-trained foundation model to various downstream multi-modal tracking tasks. ViPT finds a better way to stimulate the knowledge of the RGB-based model that is pre-trained at scale, meanwhile only introducing a few trainable parameters (less than 1% of model parameters). ViPT outperforms the full fine-tuning paradigm on multiple downstream tracking tasks including RGB+Depth, RGB+Thermal, and RGB+Event tracking. Extensive experiments show the potential of visual prompt learning for multi-modal tracking, and ViPT can achieve state-of-the-art performance while satisfying parameter efficiency. Code and models are available at https://github.com/jiawen-zhu/ViPT.
Jiawen Zhu 0003, Simiao Lai, Xin Chen 0032, Dong Wang 0004, Huchuan Lu
CVPR5
2023 MetaBEV: Solving Sensor Failures for 3D Detection and Map Segmentation
abstract
Perception systems in modern autonomous driving vehicles typically take inputs from complementary multi-modal sensors, e.g., LiDAR and cameras. However, in real-world applications, sensor corruptions and failures lead to inferior performances, thus compromising autonomous safety. In this paper, we propose a robust framework, called MetaBEV, to address extreme real-world environments, involving overall six sensor corruptions and two extreme sensor-missing situations. In MetaBEV, signals from multiple sensors are first processed by modal-specific encoders. Subsequently, a set of dense BEV queries are initialized, termed meta-BEV. These queries are then processed iteratively by a BEV-Evolving decoder, which selectively aggregates deep features from either LiDAR, cameras, or both modalities. The updated BEV representations are further leveraged for multiple 3D prediction tasks. Additionally, we introduce a new M2oE structure to alleviate the performance drop on distinct tasks in multi-task joint learning. Finally, MetaBEV is evaluated on the nuScenes dataset with 3D object detection and BEV map segmentation tasks. Experiments show MetaBEV outperforms prior arts by a large margin on both full and corrupted modalities. For instance, when the LiDAR signal is missing, MetaBEV improves 35.5% detection NDS and 17.7% segmentation mIoU upon the vanilla BEVFusion [25] model; and when the camera signal is absent, MetaBEV still achieves 69.2% NDS and 53.7% mIoU, which is even higher than previous works that perform on full-modalities. Moreover, MetaBEV performs moderately against previous methods in both canonical perception and multi-task learning settings, refreshing state-of-the-art nuScenes BEV map segmentation with 70.4% mIoU.
Chongjian Ge, Junsong Chen, Enze Xie, Zhongdao Wang, Lanqing Hong, Huchuan Lu, Zhenguo Li, Ping Luo 0002
ICCV6
2023 Towards Deeply Unified Depth-aware Panoptic Segmentation with Bi-directional Guidance Learning
abstract
Depth-aware panoptic segmentation is an emerging topic in computer vision which combines semantic and geometric understanding for more robust scene interpretation. Recent works pursue unified frameworks to tackle this challenge but mostly still treat it as two individual learning tasks, which limits their potential for exploring cross-domain information. We propose a deeply unified framework for depth-aware panoptic segmentation, which performs joint segmentation and depth estimation both in a persegment manner with identical object queries. To narrow the gap between the two tasks, we further design a geometric query enhancement method, which is able to integrate scene geometry into object queries using latent representations. In addition, we propose a bi-directional guidance learning approach to facilitate cross-task feature learning by taking advantage of their mutual relations. Our method sets the new state of the art for depth-aware panoptic segmentation on both Cityscapes-DVPS and SemKITTI-DVPS datasets. Moreover, our guidance learning approach is shown to deliver performance improvement even under incomplete supervision labels. Code and models are available at https://github.com/jwh97nn/DeepDPS.
Junwen He, Yifan Wang 0004, Lijun Wang 0001, Huchuan Lu, Bin Luo 0008, Jun-Yan He, Jin-Peng Lan, Yifeng Geng, Xuansong Xie
ICCV4
2023 Exploring Lightweight Hierarchical Vision Transformers for Efficient Visual Tracking
abstract
Transformer-based visual trackers have demonstrated significant progress owing to their superior modeling capabilities. However, existing trackers are hampered by low speed, limiting their applicability on devices with limited computational power. To alleviate this problem, we propose HiT, a new family of efficient tracking models that can run at high speed on different devices while retaining high performance. The central idea of HiT is the Bridge Module, which bridges the gap between modern lightweight transformers and the tracking framework. The Bridge Module incorporates the high-level information of deep features into the shallow large-resolution features. In this way, it produces better features for the tracking head. We also propose a novel dual-image position encoding technique that simultaneously encodes the position information of both the search region and template images. The HiT model achieves promising speed with competitive performance. For instance, it runs at 61 frames per second (fps) on the Nvidia Jetson AGX edge device. Furthermore, HiT attains 64.6% AUC on the LaSOT benchmark, surpassing all previous efficient trackers. Code and models are available at https://github.com/kangben258/HiT.
Ben Kang, Xin Chen 0032, Dong Wang 0004, Houwen Peng, Huchuan Lu
ICCV5
2023 CiteTracker: Correlating Image and Text for Visual Tracking
abstract
Existing visual tracking methods typically take an image patch as the reference of the target to perform tracking. However, a single image patch cannot provide a complete and precise concept of the target object as images are limited in their ability to abstract and can be ambiguous, which makes it difficult to track targets with drastic variations. In this paper, we propose the CiteTracker to enhance target modeling and inference in visual tracking by connecting images and text. Specifically, we develop a text generation module to convert the target image patch into a descriptive text containing its class and attribute information, providing a comprehensive reference point for the target. In addition, a dynamic description module is designed to adapt to target variations for more effective target representation. We then associate the target description and the search image using an attention-based correlation module to generate the correlated features for target state reference. Extensive experiments on five diverse datasets are conducted to evaluate the proposed algorithm and the favorable performance against the state-of-the-art methods demonstrates the effectiveness of the proposed tracking method. The source code and trained models will be released at https://github.com/NorahGreen/CiteTracker.
Xin Li 0034, Yuqing Huang, Zhenyu He 0001, Yaowei Wang 0001, Huchuan Lu, Ming-Hsuan Yang 0001
ICCV5
2023 Adaptive Illumination Mapping for Shadow Detection in Raw Images
abstract
Shadow detection methods rely on multi-scale contrast, especially global contrast, information to locate shadows correctly. However, we observe that the camera image signal processor (ISP) tends to preserve more local contrast information by sacrificing global contrast information during the raw-to-sRGB conversion process. This often causes existing methods to fail in scenes with high global contrast but low local contrast in shadow regions. In this paper, we propose a novel method to detect shadows from raw images. Our key idea is that instead of performing a many-to-one mapping like the ISP process, we can learn a many-to-many mapping from the high dynamic range raw images to the sRGB images of different illumination, which is able to preserve multi-scale contrast for accurate shadow detection. To this end, we first construct a new shadow dataset with ~ 7000 raw images and shadow masks. We then propose a novel network, which includes a novel adaptive illumination mapping (AIM) module to project the input raw images into sRGB images of different intensity ranges and a shadow detection module to leverage the preserved multi-scale contrast information to detect shadows. To learn the shadow-aware adaptive illumination mapping process, we propose a novel feedback mechanism to guide the AIM during training. Experiments show that our method outperforms state- of-the-art shadow detectors. Code and dataset are available at https://github.com/jiayusun/SARA.
Ke Xu 0010, Youwei Pang, Lihe Zhang, Huchuan Lu, Gerhard P. Hancke 0002, Rynson W. H. Lau
ICCV5
2023 Segment Every Reference Object in Spatial and Temporal Spaces
abstract
The reference-based object segmentation tasks, namely referring image segmentation (RIS), referring video object segmentation (RVOS), and video object segmentation (VOS), aim to segment a specific object by utilizing either language or annotated masks as references. Despite significant progress in each respective field, current methods are task-specifically designed and developed in different directions, which hinders the activation of multi-task capabilities for these tasks. In this work, we end the current fragmented situation and propose UniRef to unify the three reference-based object segmentation tasks with a single architecture. At the heart of our approach is the multiway-fusion for handling different task with respect to their specified references. And a unified Transformer architecture is then adopted for performing instance-level segmentation. With the unified designs, UniRef can be jointly trained on a broad range of benchmarks and can flexibly perform multiple tasks at runtime by specifying the corresponding references. We evaluate the jointly trained network on various benchmarks. Extensive experimental results indicate that our proposed UniRef achieves state-of-the-art performance on RIS and RVOS, and performs competitively on VOS with a single network.
Jiannan Wu, Yi Jiang 0009, Bin Yan 0004, Huchuan Lu, Zehuan Yuan, Ping Luo 0002
ICCV4
2023 Exploring Transformers for Open-world Instance Segmentation
abstract
Open-world instance segmentation is a rising task, which aims to segment all objects in the image by learning from a limited number of base-category objects. This task is challenging, as the number of unseen categories could be hundreds of times larger than that of seen categories. Recently, the DETR-like models have been extensively studied in the closed world while stay unexplored in the open world. In this paper, we utilize the Transformer for open-world instance segmentation and present SWORD. Firstly, we introduce to attach the stop-gradient operation before classification head and further add IoU heads for discovering novel objects. We demonstrate that a simple stop-gradient operation not only prevents the novel objects from being suppressed as background, but also allows the network to enjoy the merit of heuristic label assignment. Secondly, we propose a novel contrastive learning framework to enlarge the representations between objects and background. Specifically, we maintain a universal object queue to obtain the object center, and dynamically select positive and negative samples from the object queries for contrastive learning. While the previous works only focus on pursuing average recall and neglect average precision, we show the prominence of SWORD by giving consideration to both criteria. Our models achieve state-of-the-art performance in various open-world cross-category and cross-dataset generalizations. Particularly, in VOC to non-VOC setup, our method sets new state-of-the-art results of 40.0% on ${\text{AR}}_{100}^{\text{b}}$ and 34.9% on ${\text{AR}}_{100}^{\text{m}}$. For COCO to UVO generalization, SWORD significantly outperforms the previous best open-world model by 5.9% on APmand 8.1% on ${\text{AR}}_{100}^{\text{m}}$.
Jiannan Wu, Yi Jiang 0009, Bin Yan 0004, Huchuan Lu, Zehuan Yuan, Ping Luo 0002
ICCV4
2023 Isomer: Isomerous Transformer for Zero-shot Video Object Segmentation
abstract
Recent leading zero-shot video object segmentation (ZVOS) works devote to integrating appearance and motion information by elaborately designing feature fusion modules and identically applying them in multiple feature stages. Our preliminary experiments show that with the strong long-range dependency modeling capacity of Transformer, simply concatenating the two modality features and feeding them to vanilla Transformers for feature fusion can distinctly benefit the performance but at a cost of heavy computation. Through further empirical analysis, we find that attention dependencies learned in Transformer in different stages exhibit completely different properties: global query-independent dependency in the low-level stages and semantic-specific dependency in the high-level stages. Motivated by the observations, we propose two Transformer variants: i) Context-Sharing Transformer (CST) that learns the global-shared contextual information within image frames with a lightweight computation. ii) Semantic Gathering-Scattering Transformer (SGST) that models the semantic correlation separately for the foreground and background and reduces the computation cost with a soft token merging mechanism. We apply CST and SGST for low-level and high-level feature fusions, respectively, formulating a level-isomerous Transformer framework for ZVOS task. Compared with the baseline that uses vanilla Transformers for multi-stage fusion, ours significantly increase the speed by 13× and achieves new state-of-the-art ZVOS performance. Code is available at https://github.com/DLUT-yyc/Isomer.
Yifan Wang 0004, Lijun Wang 0001, Xiaoqi Zhao 0003, Huchuan Lu, Yu Wang 0108, Weibo Su, Lei Zhang 0006
ICCV5
2023 Video-Based Person Re-Identification with Long Short-Term Representation Learning
Xuehu Liu, Huchuan Lu
ICIG (1)3
2023 Few-shot Semantic Segmentation by Exploiting Dynamic and Regional Contexts
abstract
Few-shot Semantic Segmentation (FSS) has received increasing interests recently. Modeling effective interaction be-tween support and query images is a crucial challenge in existing prototype based methods. In this paper, we propose a Dynamic and Regional Context Network (DRCNet) to achieve sufficient support-query interaction for accurate FSS. A Dynamic Context Module (DCM) is first proposed to capture the spatial details in query images by building dynamic convolutions in local views. To further alleviate the undesirable noises, a Regional Context Module (RCM) is proposed to mine and exclude the background and ambiguous objects in query images by modeling the prototypes for ambiguous regions. Experimental results on Pascal-5iand COCO-20idatasets demonstrate that our proposed DRCNet performs significantly superior against state-of-the-art methods.
Hongyu Gu, Yunzhi Zhuge, Lu Zhang 0053, Jinqing Qi, Huchuan Lu
ICME5
2023 Image Super-Resolution with Implicit Texture Pattern Modulation
abstract
Image super-resolution is one of the classical low-level vision tasks with the purpose of restoring a high-resolution image with fine details. Being aware of texture patterns with an image would benefit super-resolution performance a lot. However, it would be difficult to predict texture patterns for each region because of lack of annotations with different kinds of categories. In this work, we propose to implicitly model texture information with each region and take that as prior to promote super-resolution performance. In order to fully explore texture patterns, a hybrid model of convolutional neural networks and transformers is proposed. It is able to take advantage of both local and long range dependencies within an image to super-resolve an image. Specifically, a set of transformers are employed to model self-similarity among pixels within an image and to make dense texture patterns prediction at each pixel. The computed texture pattern representation then works as condition to modulate convolution-based residual blocks. In this way texture patterns could be integrated into the CNNs to obtain powerful features for image super-resolution. Extensive experiments on several benchmark datasets demonstrate its favorable performance against state-of-the-art methods and show its potential as a generic design for hybrid of transformers and CNNs.
Shuai Hao 0007, Xu Jia 0012, You He 0002, Huchuan Lu
ICME5
2023 Progressively Coupling Network for Brain MRI Registration in Few-Shot Situation
Zuopeng Tan, Feng Tian 0001, Lihe Zhang, Weibing Sun, Huchuan Lu
MICCAI (10)6
2023 Recurrent Multi-scale Transformer for High-Resolution Salient Object Detection
abstract
Salient Object Detection (SOD) aims to identify and segment the most conspicuous objects in an image or video. As an important pre-processing step, it has many potential applications in multimedia and vision tasks. With the advance of imaging devices, SOD with high-resolution images is of great demand, recently. However, traditional SOD methods are largely limited to low-resolution images, making them difficult to adapt to the development of High-Resolution SOD (HRSOD). Although some HRSOD methods emerge, there are no large enough datasets for training and evaluating. Besides, current HRSOD methods generally produce incomplete object regions and irregular object boundaries. To address above issues, in this work, we first propose a new HRS10K dataset, which contains 10,500 high-quality annotated images at 2K-8K resolution. As far as we know, it is the largest dataset for the HRSOD task, which will significantly help future works in training and evaluating models. Furthermore, to improve the HRSOD performance, we propose a novel Recurrent Multi-scale Transformer (RMFormer), which recurrently utilizes shared Transformers and multi-scale refinement architectures. Thus, high-resolution saliency maps can be generated with the guidance of lower-resolution predictions. Extensive experiments on both high-resolution and low-resolution benchmarks show the effectiveness and superiority of the proposed framework. The source code and dataset are released at: https://github.com/DrowsyMon/RMFormer.
Xinhao Deng 0002, Wei Liu 0044, Huchuan Lu
ACM Multimedia4
2023 A Simple Baseline for Open-World Tracking via Self-training
abstract
Open-World Tracking (OWT) presents a challenging yet emerging problem, aiming to track every object of any category. Different from traditional Multi-Object Tracking (MOT), OWT needs to additionally track targets beyond predefined categories in the training set. To address the problem, we propose a simple baseline, SimOWT. We simplify the recently proposed OWT algorithm by streamlining the association module and accelerating the inference speed. By leveraging the self-training paradigm, SimOWT can distinguish unknown-class targets from the background, fully unleashing the potential of TAO-OW dataset. Furthermore, we enhance SimOWT from the perspectives of Pseudo Boxes Merging and Re-Weighting, thereby discovering more targets belonging to unknown classes and reducing the sensitivity of the model to low-quality pseudo-labels. Benefiting from the proposed approaches, SimOWT demonstrates a significant improvement in tracking performance on unknown classes. Moreover, the comprehensive experiments on the TAO-OW benchmark demonstrate that our model outperforms the state-of-the-art OWT method, OWTB, with an absolute gain of 11.2% OWTA and 16.4% detection recall respectively on unknown classes. The code is released at https://github.com/22109095/SimOWT.
Bingyang Wang, Tanlin Li, Jiannan Wu, Yi Jiang 0009, Huchuan Lu, You He 0002
ACM Multimedia5
2023 Ped-Mix: Mix Pedestrians for Occluded Person Re-identification
Shang Gao 0012, Chenyang Yu, Huchuan Lu
PRCV (12)4
2023 Only Classification Head Is Sufficient for Medical Image Segmentation
Hongbin Wei, Zhiwei Hu, Zhilong Ji, Hongpeng Jia, Lihe Zhang, Huchuan Lu
PRCV (13)7
2023 Growth Simulation Network for Polyp Segmentation
Hongbin Wei, Xiaoqi Zhao 0003, Long Lv, Lihe Zhang, Weibing Sun, Huchuan Lu
PRCV (13)6
2023 Delving into Calibrated Depth for Accurate RGB-D Salient Object Detection
Wei Ji 0011, Miao Zhang 0004, Yongri Piao, Huchuan Lu, Li Cheng 0001
Int. J. Comput. Vis.5
2023 High-Performance Transformer Tracking
abstract
Correlation has a critical role in the tracking field, especially in recent popular Siamese-based trackers. The correlation operation is a simple fusion method that considers the similarity between the template and the search region. However, the correlation operation is a local linear matching process, losing semantic information and easily falling into a local optimum, which may be the bottleneck in designing high-accuracy tracking algorithms. In this work, to determine whether a better feature fusion method exists than correlation, a novel attention-based feature fusion network, inspired by the transformer, is presented. This network effectively combines the template and search region features using attention mechanism. Specifically, the proposed method includes an ego-context augment module based on self-attention and a cross-feature augment module based on cross-attention. First, we present a transformer tracking (named TransT) method based on the Siamese-like feature extraction backbone, the designed attention-based fusion mechanism, and the classification and regression heads. Based on the TransT baseline, we also design a segmentation branch to generate the accurate mask. Finally, we propose a stronger version of TransT by extending it with a multi-template scheme and an IoU prediction head, named TransT-M. Experiments show that our TransT and TransT-M methods achieve promising results on seven popular benchmarks. Code and models are available at https://github.com/chenxin-dlut/TransT-M.
Xin Chen 0032, Bin Yan 0004, Jiawen Zhu 0003, Huchuan Lu, Xiang Ruan, Dong Wang 0004
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 Referring Segmentation via Encoder-Fused Cross-Modal Attention Network
abstract
This paper focuses on referring segmentation, which aims to selectively segment the corresponding visual region in an image (or video) according to the referring expression. However, the existing methods usually consider the interaction between multi-modal features at the decoding end of the network. Specifically, they interact the visual features of each scale with language respectively, thus ignoring the correlation between multi-scale features. In this work, we present an encoder fusion network (EFN), which transfers the multi-modal feature learning process from the decoding end to the encoding end and realizes the gradual refinement of multi-modal features by the language. In EFN, we also adopt a co-attention mechanism to promote the mutual alignment of language and visual information in feature space. In the decoding stage, a boundary enhancement module (BEM) is proposed to enhance the network's attention to the details of the target. For video data, we introduce an asymmetric cross-frame attention module (ACFM) to effectively capture the temporal information from the video frames by computing the relationship between each pixel of the current frame and each pooled sub-region of the reference frames. Extensive experiments on referring image/video segmentation datasets show that our method outperforms the state-of-the-art performance.
Lihe Zhang, Zhiwei Hu, Huchuan Lu
IEEE Trans. Pattern Anal. Mach. Intell.5
2023 Effective Local and Global Search for Fast Long-Term Tracking
abstract
Compared with short-term tracking, long-term tracking remains a challenging task that usually requires the tracking algorithm to track targets within a local region and re-detect targets over the entire image. However, few works have been done and their performances have also been limited. In this paper, we present a novel robust and real-time long-term tracking framework based on the proposed local search module and re-detection module. The local search module consists of an effective bounding box regressor to generate a series of candidate proposals and a target verifier to infer the optimal candidate with its confidence score. For local search, we design a long short-term updated scheme to improve the target verifier. The verification capability of the tracker can be improved by using several templates updated at different times. Based on the verification scores, our tracker determines whether the tracked object is present or absent and then chooses the tracking strategies of local or global search, respectively, in the next frame. For global re-detection, we develop a novel re-detection module that can estimate the target position and target size for a given base tracker. We conduct a series of experiments to demonstrate that this module can be flexibly integrated into many other tracking algorithms for long-term tracking and that it can improve long-term tracking performance effectively. Numerous experiments and discussions are conducted on several popular tracking datasets, including VOT, OxUvA, TLP, and LaSOT. The experimental results demonstrate that the proposed tracker achieves satisfactory performance with a real-time speed. Code is available at https://github.com/difhnp/ELGLT.
Haojie Zhao, Bin Yan 0004, Dong Wang 0004, Xuesheng Qian, Xiaoyun Yang, Huchuan Lu
IEEE Trans. Pattern Anal. Mach. Intell.6
2023 Robust Online Tracking With Meta-Updater
abstract
In a sequence, the appearance of both the target and background often changes dramatically. Offline-trained models may not handle huge appearance variations well, causing tracking failures. Most discriminative trackers address this issue by introducing an online update scheme, making the model dynamically adapt the changes of the target and background. Although the online update scheme plays an important role in improving the tracker's accuracy, it inevitably pollutes the model with noisy observation samples. It is necessary to reduce the risk of the online update scheme for better tracking. In this work, we propose a novel offline-trained Meta-Updater to address an important but unsolved problem: Is the tracker ready for updating in the current frame? The proposed module can effectively integrate geometric, discriminative, and appearance cues in a sequential manner, and then mine the sequential information with a designed cascaded LSTM module. Moreover, we strengthen the effect of appearance information on the module, i.e., the additional local outlier factor is introduced to integrate into a newly designed network. We integrate our meta-updater into eight different types of online update trackers. Extensive experiments on four long-term and two short-term tracking benchmarks demonstrate that our meta-updater is effective and has strong generalization ability.
Jie Zhao 0014, Kenan Dai, Dong Wang 0004, Huchuan Lu
IEEE Trans. Pattern Anal. Mach. Intell.5
2023 A uniform transformer-based structure for feature fusion and enhancement for RGB-D saliency detection
Yue Wang 0038, Xu Jia 0012, Lu Zhang 0053, James H. Elder, Huchuan Lu
Pattern Recognit.6
2023 Lane Detection with Versatile AtrousFormer and Local Semantic Guidance
Lihe Zhang, Huchuan Lu
Pattern Recognit.3
2023 Transformer vision-language tracking via proxy token guided cross-modal fusion
Haojie Zhao, Xiao Wang 0014, Dong Wang 0004, Huchuan Lu, Xiang Ruan
Pattern Recognit. Lett.4
2023 PANet: Patch-Aware Network for Light Field Salient Object Detection
abstract
Most existing light field saliency detection methods have achieved great success by exploiting unique light field data-focus information in focal slices. However, they process light field data in a slicewise way, leading to suboptimal results because the relative contribution of different regions in focal slices is ignored. How we can comprehensively explore and integrate focused saliency regions that would positively contribute to accurate saliency detection. Answering this question inspires us to develop a new insight. In this article, we propose a patch-aware network to explore light field data in a regionwise way. First, we excavate focused salient regions with a proposed multisource learning module (MSLM), which generates a filtering strategy for integration followed by three guidances based on saliency, boundary, and position. Second, we design a sharpness recognition module (SRM) to refine and update this strategy and perform feature integration. With our proposed MSLM and SRM, we can obtain more accurate and complete saliency maps. Comprehensive experiments on three benchmark datasets prove that our proposed method achieves competitive performance over 2-D, 3-D, and 4-D salient object detection methods. The code and results of our method are available at https://github.com/OIPLab-DUT/IEEE-TCYB-PANet.
Yongri Piao, Yongyao Jiang, Miao Zhang 0004, Huchuan Lu
IEEE Trans. Cybern.5
2023 TransY-Net: Learning Fully Transformer Networks for Change Detection of Remote Sensing Images
abstract
In the remote sensing field, Change Detection (CD) aims to identify and localize the changed regions from dual-phase images over the same places. Recently, it has achieved great progress with the advances of deep learning. However, current methods generally deliver incomplete CD regions and irregular CD boundaries due to the limited representation ability of the extracted visual features. To relieve these issues, in this work we propose a novel Transformer-based learning framework named TransY-Net for remote sensing image CD, which improves the feature extraction from a global view and combines multi-level visual features in a pyramid manner. More specifically, the proposed framework first utilizes the advantages of Transformers in long-range dependency modeling. It can help to learn more discriminative global-level features and obtain complete CD regions. Then, we introduce a novel pyramid structure to aggregate multi-level visual features from Transformers for feature enhancement. The pyramid structure grafted with a Progressive Attention Module (PAM) can improve the feature representation ability with additional inter-dependencies through spatial and channel attentions. Finally, to better train the whole framework, we utilize the deeply-supervised learning with multiple boundary-aware loss functions. Extensive experiments demonstrate that our proposed method achieves a new state-of-the-art performance on four optical and two SAR image CD benchmarks. The source code is released at https://github.com/Drchip61/TransYNet.
Tianyu Yan, Zifu Wan, Gong Cheng 0003, Huchuan Lu
IEEE Trans. Geosci. Remote. Sens.5
2023 Weakly Correlated Distillation for Remote Sensing Object Recognition
abstract
Remote sensing object labels require high specialization, resulting in a limited number of labeled samples. Without large labeled samples to support training, general remote sensing object recognition models have limited accuracy. Addressing this issue, this paper proposes a weakly correlated distillation learning framework for remote sensing object recognition with small number of samples. Benefitting from large-scale natural image datasets, many recognition models achieve superior feature extraction capabilities. Thus, we use them as backbones to build teacher models, and then fine-tune the teacher models with a small-scale remote sensing dataset. However, due to the limited number of remote sensing samples, the teacher models may produce noisy features that reduce the performance of the student model. Therefore, we propose a weakly correlated distillation method that selects the weakly correlated features from teacher models to distill the student. Since the weakly correlated features contain different noise distributions which can be mutually suppressed, thereby improving the performance of the student. Extensive experiments on three widely-used datasets of DOTA, HRRSD and NWPU VHR-10 demonstrate the superior performance of our method compared with the state of the arts. Code is available at: https://github.com/wdzhao123/WCD.
Wenda Zhao 0003, Xiangzhu Lv, Yu Liu 0005, You He 0002, Huchuan Lu
IEEE Trans. Geosci. Remote. Sens.6
2023 Plug-and-Play Regulators for Image-Text Matching
abstract
Exploiting fine-grained correspondence and visual-semantic alignments has shown great potential in image-text matching. Generally, recent approaches first employ a cross-modal attention unit to capture latent region-word interactions, and then integrate all the alignments to obtain the final similarity. However, most of them adopt one-time forward association or aggregation strategies with complex architectures or additional information, while ignoring the regulation ability of network feedback. In this paper, we develop two simple but quite effective regulators which efficiently encode the message output to automatically contextualize and aggregate cross-modal representations. Specifically, we propose (i) a Recurrent Correspondence Regulator (RCR) which facilitates the cross-modal attention unit progressively with adaptive attention factors to capture more flexible correspondence, and (ii) a Recurrent Aggregation Regulator (RAR) which adjusts the aggregation weights repeatedly to increasingly emphasize important alignments and dilute unimportant ones. Besides, it is interesting that RCR and RAR are "plug-and-play": both of them can be incorporated into many frameworks based on cross-modal interaction to obtain significant benefits, and their cooperation achieves further improvements. Extensive experiments on MSCOCO and Flickr30K datasets validate that they can bring an impressive and consistent R@1 gain on multiple models, confirming the general effectiveness and generalization ability of the proposed methods.
Haiwen Diao, Ying Zhang 0021, Wei Liu 0005, Xiang Ruan, Huchuan Lu
IEEE Trans. Image Process.5
2023 CAVER: Cross-Modal View-Mixed Transformer for Bi-Modal Salient Object Detection
abstract
Most of the existing bi-modal (RGB-D and RGB-T) salient object detection methods utilize the convolution operation and construct complex interweave fusion structures to achieve cross-modal information integration. The inherent local connectivity of the convolution operation constrains the performance of the convolution-based methods to a ceiling. In this work, we rethink these tasks from the perspective of global information alignment and transformation. Specifically, the proposed cross-modal view-mixed transformer (CAVER) cascades several cross-modal integration units to construct a top-down transformer-based information propagation path. CAVER treats the multi-scale and multi-modal feature integration as a sequence-to-sequence context propagation and update process built on a novel view-mixed attention mechanism. Besides, considering the quadratic complexity w.r.t. the number of input tokens, we design a parameter-free patch-wise token re-embedding strategy to simplify operations. Extensive experimental results on RGB-D and RGB-T SOD datasets demonstrate that such a simple two-stream encoder-decoder framework can surpass recent state-of-the-art methods when it is equipped with the proposed components.
Youwei Pang, Xiaoqi Zhao 0003, Lihe Zhang, Huchuan Lu
IEEE Trans. Image Process.4
2023 Depth Injection Framework for RGBD Salient Object Detection
abstract
Depth data with a predominance of discriminative power in location is advantageous for accurate salient object detection (SOD). Existing RGBD SOD methods have focused on how to properly use depth information for complementary fusion with RGB data, having achieved great success. In this work, we attempt a far more ambitious use of the depth information by injecting the depth maps into the encoder in a single-stream model. Specifically, we propose a depth injection framework (DIF) equipped with an Injection Scheme (IS) and a Depth Injection Module (DIM). The proposed IS enhances the semantic representation of the RGB features in the encoder by directly injecting depth maps into the high-level encoder blocks, while helping our model maintain computational convenience. Our proposed DIM acts as a bridge between the depth maps and the hierarchical RGB features of the encoder and helps the information of two modalities complement and guide each other, contributing to a great fusion effect. Experimental results demonstrate that our proposed method can achieve state-of-the-art performance on six RGBD datasets. Moreover, our method can achieve excellent performance on RGBT SOD and our DIM can be easily applied to single-stream SOD models and the transformer architecture, proving a powerful generalization ability.
Shunyu Yao 0004, Miao Zhang 0004, Yongri Piao, Chaoyi Qiu, Huchuan Lu
IEEE Trans. Image Process.5
2023 Nowhere to Disguise: Spot Camouflaged Objects via Saliency Attribute Transfer
abstract
Both salient object detection (SOD) and camouflaged object detection (COD) are typical object segmentation tasks. They are intuitively contradictory, but are intrinsically related. In this paper, we explore the relationship between SOD and COD, and then borrow successful SOD models to detect camouflaged objects to save the design cost of COD models. The core insight is that both SOD and COD leverage two aspects of information: object semantic representations for distinguishing object and background, and context attributes that decide object category. Specifically, we start by decoupling context attributes and object semantic representations from both SOD and COD datasets through designing a novel decoupling framework with triple measure constraints. Then, we transfer saliency context attributes to the camouflaged images through introducing an attribute transfer network. The generated weakly camouflaged images can bridge the context attribute gap between SOD and COD, thereby improving the SOD models' performances on COD datasets. Comprehensive experiments on three widely-used COD datasets verify the ability of the proposed method. Code and model are available at: https://github.com/wdzhao123/SAT.
Wenda Zhao 0003, Shigeng Xie, Fan Zhao 0005, You He 0002, Huchuan Lu
IEEE Trans. Image Process.5
2023 Noise-Sensitive Adversarial Learning for Weakly Supervised Salient Object Detection
abstract
Weakly supervised salient object detection (WSOD) aims at training saliency detection models with weak supervision. Normally, the WSOD methods use pseudo labels converted from image-level classification labels to train the saliency network. However, the converted pseudo labels always contain noise information compared to ground truth. Previous methods are directly affected by pseudo label noise to generate error-prone predictions. To mitigate this problem, we design a noise-robust adversarial learning framework and propose a noise-sensitive training strategy for the framework. The framework consists of a saliency network and a noise-robust discriminator network. With the guidance of noise-robust discriminator network, our saliency network is robust to noise information in pseudo labels. The proposed noise-sensitive training strategy can make good use of both superior and inferior samples in the pseudo label dataset. With the noise-sensitive training strategy, our framework can further balance the learning of saliency information and the robustness of noise information. Comprehensive experiments on five public datasets demonstrate that our method outperforms the existing image-level classification label based WSOD methods.
Yongri Piao, Miao Zhang 0004, Yongyao Jiang, Huchuan Lu
IEEE Trans. Multim.5
2023 Full-Scene Defocus Blur Detection With DeFBD+ via Multi-Level Distillation Learning
abstract
Existing defocus blur detection (DBD) methods generally perform well on a single type of unfocused blur scene (e.g., foreground focus), thereby suffering from the performance degradation for the other types of unfocused blur scenes. In this paper, we present the first exploration on full-scene DBD, and propose a separate-and-combine framework to achieve excellent performance for diverse defocus blur scenes. We firstly structure full-scene DBD dataset (named as DeFBD+) through collecting more types of unfocused blur scenes (e.g., background focus, full focus and full out of focus) with pixel-level annotations. Then, to avoid performance degradation caused by mutual interference from local feature representation and global content perception, we implement a pixel-level DBD network and an image-level DBD classification network to learn these two abilities separately. After that, we propose an isomeric distillation mechanism to combine these two abilities. Extensive experiments show that the proposed approach achieves superior performance compared with state-of-the-art methods.
Wenda Zhao 0003, Fei Wei, You He 0002, Huchuan Lu
IEEE Trans. Multim.5
2023 Bidirectional Relationship Inferring Network for Referring Image Localization and Segmentation
abstract
Recently, referring image localization and segmentation has aroused widespread interest. However, the existing methods lack a clear description of the interdependence between language and vision. To this end, we present a bidirectional relationship inferring network (BRINet) to effectively address the challenging tasks. Specifically, we first employ a vision-guided linguistic attention module to perceive the keywords corresponding to each image region. Then, language-guided visual attention adopts the learned adaptive language to guide the update of the visual features. Together, they form a bidirectional cross-modal attention module (BCAM) to achieve the mutual guidance between language and vision. They can help the network align the cross-modal features better. Based on the vanilla language-guided visual attention, we further design an asymmetric language-guided visual attention, which significantly reduces the computational cost by modeling the relationship between each pixel and each pooled subregion. In addition, a segmentation-guided bottom-up augmentation module (SBAM) is utilized to selectively combine multilevel information flow for object localization. Experiments show that our method outperforms other state-of-the-art methods on three referring image localization datasets and four referring image segmentation datasets.
Zhiwei Hu, Lihe Zhang, Huchuan Lu
IEEE Trans. Neural Networks Learn. Syst.5
2022 You Only Infer Once: Cross-Modal Meta-Transfer for Referring Video Object Segmentation
abstract
We present YOFO (You Only inFer Once), a new paradigm for referring video object segmentation (RVOS) that operates in an one-stage manner. Our key insight is that the language descriptor should serve as target-specific guidance to identify the target object, while a direct feature fusion of image and language can increase feature complexity and thus may be sub-optimal for RVOS. To this end, we propose a meta-transfer module, which is trained in a learning-to-learn fashion and aims to transfer the target-specific information from the language domain to the image domain, while discarding the uncorrelated complex variations of language description. To bridge the gap between the image and language domains, we develop a multi-scale cross-modal feature mining block that aggregates all the essential features required by RVOS from both domains and generates regression labels for the meta-transfer module. The whole system can be trained in an end-to-end manner and shows competitive performance against state-of-the-art two-stage approaches.
Dezhuang Li, Ruoqi Li, Lijun Wang 0001, Yifan Wang 0004, Jinqing Qi, Lu Zhang 0053, Ting Liu 0018, Qingquan Xu, Huchuan Lu
AAAI9
2022 Self-Supervised Pretraining for RGB-D Salient Object Detection
abstract
Existing CNNs-Based RGB-D salient object detection (SOD) networks are all required to be pretrained on the ImageNet to learn the hierarchy features which helps provide a good initialization. However, the collection and annotation of large-scale datasets are time-consuming and expensive. In this paper, we utilize self-supervised representation learning (SSL) to design two pretext tasks: the cross-modal auto-encoder and the depth-contour estimation. Our pretext tasks require only a few and unlabeled RGB-D datasets to perform pretraining, which makes the network capture rich semantic contexts and reduce the gap between two modalities, thereby providing an effective initialization for the downstream task. In addition, for the inherent problem of cross-modal fusion in RGB-D SOD, we propose a consistency-difference aggregation (CDA) module that splits a single feature fusion into multi-path fusion to achieve an adequate perception of consistent and differential information. The CDA module is general and suitable for cross-modal and cross-level feature fusion. Extensive experiments on six benchmark datasets show that our self-supervised pretrained model performs favorably against most state-of-the-art methods pretrained on ImageNet. The source code will be publicly available at https://github.com/Xiaoqi-Zhao-DLUT/SSLSOD.
Xiaoqi Zhao 0003, Youwei Pang, Lihe Zhang, Huchuan Lu, Xiang Ruan
AAAI4
2022 Video Object Segmentation via Structural Feature Reconfiguration
Zhenyu Chen 0001, Ping Hu 0001, Lu Zhang 0053, Huchuan Lu, You He 0002, Maodi Hu
ACCV (7)4
2022 Blind Image Super-Resolution with Degradation-Aware Adaptation
Yue Wang 0038, Jiawen Ming, Xu Jia 0012, James H. Elder, Huchuan Lu
ACCV (3)5
2022 Semantics-Adding Flaw-Erasing Network for Semantic Human Matting
Zhanghan Ke, Ke Xu 0010, Fan Shao, Lihe Zhang, Huchuan Lu, Rynson W. H. Lau
BMVC6
2022 TimeReplayer: Unlocking the Potential of Event Cameras for Video Interpolation
abstract
Recording fast motion in a high FPS (frame-per-second) requires expensive high-speed cameras. As an alternative, interpolating low-FPS videos from commodity cameras has attracted significant attention. If only low-FPS videos are available, motion assumptions (linear or quadratic) are necessary to infer intermediate frames, which fail to model complex motions. Event camera, a new camera with pixels producing events of brightness change at the temporal resolution of μs (10–6second), is a game-changing device to enable video interpolation at the presence of arbitrarily complex motion. Since event camera is a novel sensor, its potential has not been fulfilled due to the lack of processing algorithms. The pioneering work Time Lens introduced event cameras to video interpolation by designing optical devices to collect a large amount of paired training data of high-speed frames and events, which is too costly to scale. To fully unlock the potential of event cameras, this paper proposes a novel TimeReplayer algorithm to interpolate videos captured by commodity cameras with events. It is trained in an unsupervised cycleconsistent style, canceling the necessity of high-speed training data and bringing the additional ability of video extrapolation. Its state-of-the-art results and demo videos in supplementary reveal the promising future of event-based vision.
Weihua He, Kaichao You, Zhendong Qiao, Xu Jia 0012, Wenhui Wang 0001, Huchuan Lu, Yaoyuan Wang, Jianxing Liao
CVPR7
2022 Look Back and Forth: Video Super-Resolution with Explicit Temporal Difference Modeling
abstract
Temporal modeling is crucial for video super-resolution. Most of the video super-resolution methods adopt the optical flow or deformable convolution for explicitly motion compensation. However, such temporal modeling techniques increase the model complexity and might fail in case of occlusion or complex motion, resulting in serious distortion and artifacts. In this paper, we propose to explore the role of explicit temporal difference modeling in both LR and HR space. Instead of directly feeding consecutive frames into a VSR model, we propose to compute the temporal difference between frames and divide those pixels into two subsets according to the level of difference. They are separately processed with two branches of different receptive fields in order to better extract complementary information. To further enhance the super-resolution result, not only spatial residual features are extracted, but the difference between consecutive frames in high-frequency domain is also computed. It allows the model to exploit intermediate SR results in both future and past to refine the current SR output. The difference at different time steps could be cached such that information from further distance in time could be propagated to the current frame for refinement. Experiments on several video super-resolution benchmark datasets demonstrate the effectiveness of the proposed method and its favorable performance against state-of-the-art methods.
Takashi Isobe, Xu Jia 0012, Xin Tao 0001, Ruihuang Li, Yongjie Shi, Huchuan Lu, Yu-Wing Tai
CVPR8
2022 Multi-Object Tracking Meets Moving UAV
abstract
Multi-object tracking in unmanned aerial vehicle (UAV) videos is an important vision task and can be applied in a wide range of applications. However, conventional multi-object trackers do not work well on UAV videos due to the challenging factors of irregular motion caused by moving camera and view change in 3D directions. In this paper, we propose a UAVMOT network specially for multi-object tracking in UAV views. The UAVMOT introduces an ID feature update module to enhance the object's feature association. To better handle the complex motions under UAV views, we develop an adaptive motion filter module. In addition, a gradient balanced focal loss is used to tackle the imbalance categories and small objects detection problem. Experimental results on the VisDrone2019 and UAVDT datasets demonstrate that the proposed UAVMOT achieves considerable improvement against the state-of-the-art tracking methods on UAV videos.
Shuai Liu 0009, Xin Li 0034, Huchuan Lu, You He 0002
CVPR3
2022 Zoom In and Out: A Mixed-scale Triplet Network for Camouflaged Object Detection
abstract
The recently proposed camouflaged object detection (COD) attempts to segment objects that are visually blended into their surroundings, which is extremely complex and difficult in real-world scenarios. Apart from high intrinsic similarity between the camouflaged objects and their background, the objects are usually diverse in scale, fuzzy in appearance, and even severely occluded. To deal with these problems, we propose a mixed-scale triplet network, Zoom- Net, which mimics the behavior of humans when observing vague images, i.e., zooming in and out. Specifically, our ZoomNet employs the zoom strategy to learn the discriminative mixed-scale semantics by the designed scale integration unit and hierarchical mixed-scale unit, which fully explores imperceptible clues between the candidate objects and background surroundings. Moreover, considering the uncertainty and ambiguity derived from indistinguishable textures, we construct a simple yet effective regularization constraint, uncertainty-aware loss, to promote the model to accurately produce predictions with higher confidence in candidate regions. Without bells and whistles, our proposed highly task-friendly model consistently surpasses the existing 23 state-of-the-art methods on four public datasets. Besides, the superior performance over the recent cutting-edge models on the SOD task also verifies the effectiveness and generality of our model. The code will be available at https://github.com/lartpang/ZoomNet.
Youwei Pang, Xiaoqi Zhao 0003, Tian-Zhu Xiang, Lihe Zhang, Huchuan Lu
CVPR5
2022 Multi-Source Uncertainty Mining for Deep Unsupervised Saliency Detection
abstract
Deep learning-based image salient object detection (SOD) heavily relies on large-scale training data with pixel-wise labeling. High-quality labels involve intensive labor and are expensive to acquire. In this paper, we propose a novel multi-source uncertainty mining method to facilitate unsupervised deep learning from multiple noisy labels generated by traditional handcrafted SOD methods. We design an Uncertainty Mining Network (UMNet) which consists of multiple Merge-and-Split (MS) modules to recursively analyze the commonality and difference among multiple noisy labels and infer pixel-wise uncertainty map for each label. Meanwhile, we model the noisy labels using Gibbs distribution and propose a weighted uncertainty loss to jointly train the UMNet with the SOD network. As a consequence, our UMNet can adaptively select reliable labels for SOD network learning. Extensive experiments on benchmark datasets demonstrate that our method not only outperforms existing unsupervised methods, but also is on par with fully-supervised state-of-the-art models.
Yifan Wang 0004, Lijun Wang 0001, Ting Liu 0018, Huchuan Lu
CVPR5
2022 Visible-Thermal UAV Tracking: A Large-Scale Benchmark and New Baseline
abstract
With the popularity of multi-modal sensors, visible-thermal (RGB-T) object tracking is to achieve robust performance and wider application scenarios with the guidance of objects' temperature information. However, the lack of paired training samples is the main bottleneck for unlocking the power of RGB-T tracking. Since it is laborious to collect high-quality RGB-T sequences, recent benchmarks only provide test sequences. In this paper, we construct a large-scale benchmark with high diversity for visible-thermal UAV tracking (VTUAV), including 500 sequences with 1.7 million high-resolution (1920* 1080 pixels) frame pairs. In addition, comprehensive applications (short-term tracking, long-term tracking and segmentation mask prediction) with diverse categories and scenes are considered for exhaustive evaluation. Moreover, we provide a coarse-to-fine attribute annotation, where frame-level attributes are provided to exploit the potential of challenge-specific trackers. In addition, we design a new RGB-T baseline, named Hierarchical Multi-modal Fusion Tracker (HMFT), which fuses RGB-T data in various levels. Numerous experiments on several datasets are conducted to reveal the effectiveness of HMFT and the complement of different fusion types. The project is available at here.
Jie Zhao 0014, Dong Wang 0004, Huchuan Lu, Xiang Ruan
CVPR4
2022 Adaptive Co-teaching for Unsupervised Monocular Depth Estimation
Weisong Ren, Lijun Wang 0001, Yongri Piao, Miao Zhang 0004, Huchuan Lu, Ting Liu 0018
ECCV (1)5
2022 Towards Grand Unification of Object Tracking
Bin Yan 0004, Yi Jiang 0009, Peize Sun, Dong Wang 0004, Zehuan Yuan, Ping Luo 0002, Huchuan Lu
ECCV (21)7
2022 United Defocus Blur Detection and Deblurring via Adversarial Promoting Learning
Wenda Zhao 0003, Fei Wei, You He 0002, Huchuan Lu
ECCV (30)4
2022 MVSalNet: Multi-view Augmentation for RGB-D Salient Object Detection
Jiayuan Zhou, Lijun Wang 0001, Huchuan Lu, Kaining Huang, Xinchu Shi, Bocong Liu
ECCV (29)3
2022 Depth-inspired Label Mining for Unsupervised RGB-D Salient Object Detection
abstract
Existing deep learning-based unsupervised Salient Object Detection (SOD) methods heavily rely on the pseudo labels predicted from handcrafted features. However, the pseudo ground truth obtained only from RGB space would easily bring undesirable noises, especially in some complex scenarios. This naturally leads to the incorporation of extra depth modality with RGB images for more robust object identification, namely RGB-D SOD. Compared with the well-studied unsupervised SOD in the RGB domain, deep unsupervised RGB-D SOD is a less explored direction in the literature. In this paper, we propose to tackle this task by introducing a novel systemic design for high-quality pseudo-label mining. Our framework consists of two key components, Depth-inspired Label Generation (DLG) and Multi-source Uncertainty-aware Label Optimization (MULO). In DLG, a lightweight deep network is designed for automatically producing pseudo labels from depth maps in a self-supervised manner. Then, MULO introduces an effective pseudo label optimization strategy by learning the uncertainty of the pseudo labels from the depth domain and heuristic features. Extensive experiments demonstrate that the proposed method significantly outperforms the state-of-the-art unsupervised methods on mainstream benchmarks.
Yue Wang 0038, Lu Zhang 0053, Jinqing Qi, Huchuan Lu
ACM Multimedia5
2022 PreyNet: Preying on Camouflaged Objects
abstract
Species often adopt various camouflage strategies to be seamlessly blended into the surroundings for self-protection. To figure out the concealment, predators have evolved excellent hunting skills. Exploring the intrinsic mechanisms of the predation behavior can offer more insightful glimpse into the task of camouflaged object detection (COD). In this work, we strive to seek answers for accurate COD and propose a PreyNet, which mimics the two processes of predation, namely, initial detection (sensory mechanism) and predator learning (cognitive mechanism). To exploit the sensory process, a bidirectional bridging interaction module (BBIM) is designed for selecting and aggregating initial features in an attentive manner. The predator learning process is formulated as a policy-and-calibration paradigm, with the goal of deciding on uncertain regions and encouraging targeted feature calibration. Besides, we obtain adaptive weight for multi-layer supervision during training via computing on the uncertainty estimation. Extensive experiments demonstrate that our model produces state-of-the-art results on several benchmarks. We further verify the scalability of the predator learning paradigm through applications on top-ranking salient object detection models. Our code is publicly available at \urlhttps://github.com/OIPLab-DUT/PreyNet.
Miao Zhang 0004, Yongri Piao, Dongxiang Shi, Shusen Lin, Huchuan Lu
ACM Multimedia6
2022 Semi-Supervised Video Salient Object Detection Based on Uncertainty-Guided Pseudo Labels
abstract
Semi-Supervised Video Salient Object Detection (SS-VSOD) is challenging because of the lack of temporal information in video sequences caused by sparse annotations. Most works address this problem by generating pseudo labels for unlabeled data. However, error-prone pseudo labels negatively affect the VOSD model. Therefore, a deeper insight into pseudo labels should be developed. In this work, we aim to explore 1) how to utilize the incorrect predictions in pseudo labels to guide the network to generate more robust pseudo labels and 2) how to further screen out the noise that still exists in the improved pseudo labels. To this end, we propose an Uncertainty-Guided Pseudo Label Generator (UGPLG), which makes full use of inter-frame information to ensure the temporal consistency of the pseudo labels and improves the robustness of the pseudo labels by strengthening the learning of difficult scenarios. Furthermore, we also introduce the adversarial learning to address the noise problems in pseudo labels, guaranteeing the positive guidance of pseudo labels during model training. Experimental results demonstrate that our methods outperform existing semi-supervised method and partial fully-supervised methods across five public benchmarks of DAVIS, FBMS, MCL, ViSal and SegTrack-V2.
Yongri Piao, Chenyang Lu 0007, Miao Zhang 0004, Huchuan Lu
NeurIPS4
2022 Road extraction from satellite images with iterative cross-task feature enhancement
Weiling Yin, Mingyang Qian, Lijun Wang 0001, Jinqing Qi, Huchuan Lu
Neurocomputing5
2022 Learning to Detect Salient Object With Multi-Source Weak Supervision
abstract
High-cost pixel-level annotations makes it appealing to train saliency detection models with weak supervision. However, a single weak supervision source hardly contain enough information to train a well-performing model. To this end, we introduce a unified two-stage framework to learn from category labels, captions, web images and unlabeled images. In the first stage, we design a classification network (CNet) and a caption generation network (PNet), which learn to predict object categories and generate captions, respectively, meanwhile highlights the potential foreground regions. We present an attention transfer loss to transmit supervisions between two tasks and an attention coherence loss to encourage the networks to detect generally salient regions instead of task-specific regions. In the second stage, we create two complementary training datasets using CNet and PNet, i.e., natural image dataset with noisy labels for adapting saliency prediction network (SNet) to natural image input, and synthesized image dataset by pasting objects on background images for providing SNet with accurate ground-truth. During the testing phases, we only need SNet to predict saliency maps. Experiments indicate the performance of our method compares favorably against unsupervised, weakly supervised methods and even some supervised methods.
Hongshuang Zhang, Yu Zeng 0001, Huchuan Lu, Lihe Zhang, Jinqing Qi
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 Encoder deep interleaved network with multi-scale aggregation for RGB-D salient object detection
Jinyu Meng, Lihe Zhang, Huchuan Lu
Pattern Recognit.4
2022 Salient object detection with image-level binary supervision
Pengjie Wang 0001, Ying Cao 0001, Xin Yang 0011, Huchuan Lu, Rynson W. H. Lau
Pattern Recognit.6
2022 Video Saliency Prediction via Joint Discrimination and Local Consistency
abstract
While saliency detection on static images has been widely studied, the research on video saliency detection is still in an early stage and requires more efforts due to the challenge to bring both local and global consistency of salient objects into full consideration. In this article, we propose a novel dynamic saliency network based on both local consistency and global discriminations, via which semantic features across video frames are simultaneously extracted and a recurrent feature optimization structure is designed to further enhance its performances. To ensure that the generated dynamic salient map is more concentrated, we design a lightweight discriminator with a local consistency loss LC to identify subtle differences between predicted maps and ground truths. As a result, the proposed network can be further stimulated to produce more realistic saliency maps with smoother boundaries and simpler layer transitions. The added LC loss forces the network to pay more attention to the local consistency between continuous saliency maps. Both qualitative and quantitative experiments are carried out on three large datasets, and the results demonstrate that our proposed network not only achieves improved performances but also shows good robustness.
Zheng Wang 0008, Ziqi Zhou 0002, Huchuan Lu, Qinghua Hu, Jianmin Jiang
IEEE Trans. Cybern.3
2022 Center-Boundary Dual Attention for Oriented Object Detection in Remote Sensing Images
abstract
Recently, anchor-free object detectors have shown promising performance in oriented object detection on remote sensing images. However, the objects in remote sensing images always have large variations in arbitrary orientations, sizes, and aspect ratios, which makes the existing anchor-free methods hard to obtain satisfactory results. In this article, we propose a novel anchor-free detector, center-boundary dual attention (CBDA) network (CBDA-Net), for fast and accurate oriented object detection on remote sensing images. In CBDA-Net, we construct a CBDA module, which utilizes a dual attention mechanism to extract attention features on the center and boundary regions of objects. The CBDA module can learn more essential features for rotating objects and reduce the interference from complex background. Besides, to resolve the influence of object aspect ratio on angle errors, we propose an aspect ratio weighted angle loss (arwLoss), where diffident penalties are assigned on the angle loss based on the aspect ratios of objects. This loss construction is effective in improving the detection accuracy of oriented objects, especially for slender objects. We conduct extensive experiments on two publish benchmarks, i.e., DOTA and HRSC2016. The experimental results demonstrate that our CBDA-Net achieves favorable performance against other anchor-free state of the arts with a real-time speed of 50 FPS.
Shuai Liu 0009, Lu Zhang 0053, Huchuan Lu, You He 0002
IEEE Trans. Geosci. Remote. Sens.3
2022 Teaching Teachers First and Then Student: Hierarchical Distillation to Improve Long-Tailed Object Recognition in Aerial Images
abstract
Remote sensing data distribution generally exposes the long-tail characteristic. This will limit the object recognition performance of existing deep models when they are trained with such unbalanced data. In this paper, we propose a novel hierarchical distillation framework to address the long-tailed object recognition in aerial images. Firstly, we notice that not only student model should learn feature representations from teachers, but also teacher models should learn feature representations from each other. Therefore, we build hierarchical teacher-wise distillation to improve the feature representations of the teacher models trained with middle and tail data, which is achieved by distilling the feature representations of the teacher model trained with head data. Secondly, we notice that the feature representations of the middle and tail classes can not be effectively distilled from the teacher to the student, since too little middle and tail data can be used to learn. Thus, we propose self-calibrated sampling learning that enforces the student to strengthen the learning of the middle and tail data, thereby improving the student’ feature learning ability. Extensive experiments on two widely-used DOTA and FGSC-23 datasets demonstrate superior performance of the proposed method compared with state-of-the-art methods. Model and code are publicly available at: https://github.com/wdzhao123/T2FTS.
Wenda Zhao 0003, Jiani Liu 0004, Yu Liu 0005, Fan Zhao 0005, You He 0002, Huchuan Lu
IEEE Trans. Geosci. Remote. Sens.6
2022 Diversity Consistency Learning for Remote-Sensing Object Recognition With Limited Labels
abstract
Annotating remote sensing object recognition needs high professionalism, and thus limited labeled samples are available. Suffering from this, general remote sensing object recognition methods are facing low recognition accuracy. Addressing this issue, this paper proposes a diversity consistency learning for remote sensing object recognition with limited labels. Specifically, diversity generation model is designed as a teacher model to generate diverse results, which is trained with labeled samples. Then, round consistency distillation model is introduced to distill the knowledge of diverse pseudo labels to a student network, which is trained with unlabeled samples. Especially, diverse pseudo labels are generated by the well-trained diversity generation model, which can improve recognition accuracy since diverse pseudo label errors can cancel each other out. Extensive experiments on two widely-used datasets of FS23 and HRSC2016 demonstrate the superior performance of our method compared with the state of the arts.
Wenda Zhao 0003, Tingting Tong, Fan Zhao 0005, You He 0002, Huchuan Lu
IEEE Trans. Geosci. Remote. Sens.6
2022 Feature Balance for Fine-Grained Object Classification in Aerial Images
abstract
Fine-grained object classification (FGOC) focuses on identifying subcategories of objects, which is crucial in military and civilian. Existing FGOC methods primarily focus on high-resolution aerial images, limiting their application on low-resolution (LR) FGOC that is a more realistic setting, especially on resource-constrained satellite devices. It is more challenging to deal with LR FGOC since objects’ details are blurred or missing. Addressing this issue, we make the first attempt to explore LR FGOC and propose a novel pipeline based on two technical insights: 1) feature balance strategy discriminatively integrates super-resolution weak and strong detailed presentations into coarse features of LR aerial images, achieving a feature balance to avoid that the weak detailed presentations are inhibited by the strong ones and 2) iterative interaction mechanism alternately refines feature details of the discriminative ship regions and optimizes the performance of FGOC. Moreover, we build a low-resolution fine-grained object (LFS) dataset to promote further study and evaluation. Extensive experiments on the proposed LFS dataset and the other three object datasets of DOTA, FS23, and HRSC2016 demonstrate that our method outperforms state-of-the-art algorithms. Dataset and code are publicly available athttps://github.com/wdzhao123/FBNet.
Wenda Zhao 0003, Tingting Tong, Libo Yao, Yu Liu 0005, Cong'an Xu, You He 0002, Huchuan Lu
IEEE Trans. Geosci. Remote. Sens.7
2022 Neighbor2Neighbor: A Self-Supervised Framework for Deep Image Denoising
abstract
In recent years, image denoising has benefited a lot from deep neural networks. However, these models need large amounts of noisy-clean image pairs for supervision. Although there have been attempts in training denoising networks with only noisy images, existing self-supervised algorithms suffer from inefficient network training, heavy computational burden, or dependence on noise modeling. In this paper, we proposed a self-supervised framework named Neighbor2Neighbor for deep image denoising. We develop a theoretical motivation and prove that by designing specific samplers for training image pairs generation from only noisy images, we can train a self-supervised denoising network similar to the network trained with clean images supervision. Besides, we propose a regularizer in the perspective of optimization to narrow the optimization gap between the self-supervised denoiser and the supervised denoiser. We present a very simple yet effective self-supervised training scheme based on the theoretical understandings: training image pairs are generated by random neighbor sub-samplers, and denoising networks are trained with a regularized loss. Moreover, we propose a training strategy named BayerEnsemble to adapt the Neighbor2Neighbor framework in raw image denoising. The proposed Neighbor2Neighbor framework can enjoy the progress of state-of-the-art supervised denoising networks in network architecture design. It also avoids heavy dependence on the assumption of the noise distribution. We evaluate the Neighbor2Neighbor framework through extensive experiments, including synthetic experiments with different noise distributions and real-world experiments under various scenarios. The code is available online: https://github.com/TaoHuang2018/Neighbor2Neighbor.
Tao Huang 0022, Songjiang Li, Xu Jia 0012, Huchuan Lu, Jianzhuang Liu
IEEE Trans. Image Process.4
2022 DMRA: Depth-Induced Multi-Scale Recurrent Attention Network for RGB-D Saliency Detection
abstract
In this work, we propose a novel depth-induced multi-scale recurrent attention network for RGB-D saliency detection, named as DMRA. It achieves dramatic performance especially in complex scenarios. There are four main contributions of our network that are experimentally demonstrated to have significant practical merits. First, we design an effective depth refinement block using residual connections to fully extract and fuse cross-modal complementary cues from RGB and depth streams. Second, depth cues with abundant spatial information are innovatively combined with multi-scale contextual features for accurately locating salient objects. Third, a novel recurrent attention module inspired by Internal Generative Mechanism of human brain is designed to generate more accurate saliency results via comprehensively learning the internal semantic relation of the fused feature and progressively optimizing local details with memory-oriented scene understanding. Finally, a cascaded hierarchical feature fusion strategy is designed to promote efficient information interaction of multi-level contextual features and further improve the contextual representability of model. In addition, we introduce a new real-life RGB-D saliency dataset containing a variety of complex scenarios that has been widely used as a benchmark dataset in recent RGB-D saliency detection research. Extensive empirical experiments demonstrate that our method can accurately identify salient objects and achieve appealing performance against 18 state-of-the-art RGB-D saliency models on nine benchmark datasets.
Wei Ji 0011, Ge Yan 0006, Yongri Piao, Shunyu Yao 0004, Miao Zhang 0004, Li Cheng 0001, Huchuan Lu
IEEE Trans. Image Process.8
2022 From Pixels to Semantics: Self-Supervised Video Object Segmentation With Multiperspective Feature Mining
abstract
Existing self-supervised methods pose one-shot video object segmentation (O-VOS) as pixel-level matching to enable segmentation mask propagation across frames. However, the two tasks are not fully equivalent since O-VOS is more reliant on semantic correspondence rather than accurate pixel matching. To remedy this issue, we explore a new self-supervised framework that integrates pixel-level correspondence learning with semantic-level adaptation. The pixel-level correspondence learning is performed through photometric reconstruction of adjacent RGB frames during offline training, while semantic-level adaption operates at test-time by enforcing a bi-directional agreement of the predicted segmentation masks. In addition, we further propose a new network architecture with multi-perspective feature mining mechanism which can not only enhance reliable features but also suppress noisy ones to facilitate more robust image matching. By training the network using the proposed self-supervised framework, we achieve state-of-the-art performance on widely adopted datasets, further closing up the gap between self-supervised learning methods and their fully supervised counterparts.
Ruoqi Li, Yifan Wang 0004, Lijun Wang 0001, Huchuan Lu, Xiaopeng Wei, Qiang Zhang 0008
IEEE Trans. Image Process.4
2022 Exploring Spatial Correlation for Light Field Saliency Detection: Expansion From a Single View
abstract
Previous 2D saliency detection methods extract salient cues from a single view and directly predict the expected results. Both traditional and deep-learning-based 2D methods do not consider geometric information of 3D scenes. Therefore the relationship between scene understanding and salient objects cannot be effectively established. This limits the performance of 2D saliency detection in challenging scenes. In this paper, we show for the first time that saliency detection problem can be reformulated as two sub-problems: light field synthesis from a single view and light-field-driven saliency detection. This paper first introduces a high-quality light field synthesis network to produce reliable 4D light field information. Then a novel light-field-driven saliency detection network is proposed, in which a Direction-specific Screening Unit (DSU) is tailored to exploit the spatial correlation among multiple viewpoints. The whole pipeline can be trained in an end-to-end fashion. Experimental results demonstrate that the proposed method outperforms the state-of-the-art 2D, 3D and 4D saliency detection methods. Our code is publicly available at https://github.com/OIPLab-DUT/ESCNet.
Miao Zhang 0004, Yongri Piao, Huchuan Lu
IEEE Trans. Image Process.4
2022 Joint Learning of Salient Object Detection, Depth Estimation and Contour Extraction
abstract
Benefiting from color independence, illumination invariance and location discrimination attributed by the depth map, it can provide important supplemental information for extracting salient objects in complex environments. However, high-quality depth sensors are expensive and can not be widely applied. While general depth sensors produce the noisy and sparse depth information, which brings the depth-based networks with irreversible interference. In this paper, we propose a novel multi-task and multi-modal filtered transformer (MMFT) network for RGB-D salient object detection (SOD). Specifically, we unify three complementary tasks: depth estimation, salient object detection and contour estimation. The multi-task mechanism promotes the model to learn the task-aware features from the auxiliary tasks. In this way, the depth information can be completed and purified. Moreover, we introduce a multi-modal filtered transformer (MFT) module, which equips with three modality-specific filters to generate the transformer-enhanced feature for each modality. The proposed model works in a depth-free style during the testing phase. Experiments show that it not only significantly surpasses the depth-based RGB-D SOD methods on multiple datasets, but also precisely predicts a high-quality depth map and salient contour at the same time. And, the resulted depth map can help existing RGB-D SOD methods obtain significant performance gain.
Xiaoqi Zhao 0003, Youwei Pang, Lihe Zhang, Huchuan Lu
IEEE Trans. Image Process.4
2021 Similarity Reasoning and Filtration for Image-Text Matching
abstract
Image-text matching plays a critical role in bridging the vision and language, and great progress has been made by exploiting the global alignment between image and sentence, or local alignments between regions and words. However, how to make the most of these alignments to infer more accurate matching scores is still underexplored. In this paper, we propose a novel Similarity Graph Reasoning and Attention Filtration (SGRAF) network for image-text matching. Specifically, the vector-based similarity representations are firstly learned to characterize the local and global alignments in a more comprehensive manner, and then the Similarity Graph Reasoning (SGR) module relying on one graph convolutional neural network is introduced to infer relation-aware similarities with both the local and global alignments. The Similarity Attention Filtration (SAF) module is further developed to integrate these alignments effectively by selectively attending on the significant and representative alignments and meanwhile casting aside the interferences of non-meaningful alignments. We demonstrate the superiority of the proposed method with achieving state-of-the-art performances on the Flickr30K and MSCOCO datasets, and the good interpretability of SGR and SAF with extensive qualitative experiments and analyses.
Haiwen Diao, Ying Zhang 0021, Lin Ma 0002, Huchuan Lu
AAAI4
2021 Transformer Tracking
abstract
Correlation acts as a critical role in the tracking field, especially in recent popular Siamese-based trackers. The correlation operation is a simple fusion manner to consider the similarity between the template and the search region. However, the correlation operation itself is a local linear matching process, leading to lose semantic information and fall into local optimum easily, which may be the bottleneck of designing high-accuracy tracking algorithms. Is there any better feature fusion method than correlation? To address this issue, inspired by Transformer, this work presents a novel attention-based feature fusion network, which effectively combines the template and search region features solely using attention. Specifically, the proposed method includes an ego-context augment module based on self-attention and a cross-feature augment module based on cross-attention. Finally, we present a Transformer tracking (named TransT) method based on the Siamese-like feature extraction backbone, the designed attention-based fusion mechanism, and the classification and regression head. Experiments show that our TransT achieves very promising results on six challenging datasets, especially on large-scale LaSOT, TrackingNet, and GOT-10k benchmarks. Our tracker runs at approximatively 50 fps on GPU. Code and models are available at https://github.com/chenxin-dlut/TransT.
Xin Chen 0032, Bin Yan 0004, Jiawen Zhu 0003, Dong Wang 0004, Xiaoyun Yang, Huchuan Lu
CVPR6
2021 Encoder Fusion Network With Co-Attention Embedding for Referring Image Segmentation
abstract
Recently, referring image segmentation has aroused widespread interest. Previous methods perform the multi-modal fusion between language and vision at the decoding side of the network. And, linguistic feature interacts with visual feature of each scale separately, which ignores the continuous guidance of language to multi-scale visual features. In this work, we propose an encoder fusion network (EFN), which transforms the visual encoder into a multi-modal feature learning network, and uses language to refine the multi-modal features progressively. Moreover, a co-attention mechanism is embedded in the EFN to realize the parallel update of multi-modal features, which can promote the consistent of the cross-modal information representation in the semantic space. Finally, we propose a boundary enhancement module (BEM) to make the network pay more attention to the fine structure. The experiment results on four benchmark datasets demonstrate that the proposed approach achieves the state-of-the-art performance under different evaluation metrics without any post-processing.
Zhiwei Hu, Lihe Zhang, Huchuan Lu
CVPR4
2021 Neighbor2Neighbor: Self-Supervised Denoising From Single Noisy Images
abstract
In the last few years, image denoising has benefited a lot from the fast development of neural networks. However, the requirement of large amounts of noisy-clean image pairs for supervision limits the wide use of these models. Although there have been a few attempts in training an image denoising model with only single noisy images, existing self-supervised denoising approaches suffer from inefficient network training, loss of useful information, or dependence on noise modeling. In this paper, we present a very simple yet effective method named Neighbor2Neighbor to train an effective image denoising model with only noisy images. Firstly, a random neighbor sub-sampler is proposed for the generation of training image pairs. In detail, input and target used to train a network are images sub-sampled from the same noisy image, satisfying the requirement that paired pixels of paired images are neighbors and have very similar appearance with each other. Secondly, a denoising network is trained on sub-sampled training pairs generated in the first stage, with a proposed regularizer as additional loss for better performance. The proposed Neighbor2Neighbor framework is able to enjoy the progress of state-of-the-art supervised denoising networks in network architecture design. Moreover, it avoids heavy dependence on the assumption of the noise distribution. We explain our approach from a theoretical perspective and further validate it through extensive experiments, including synthetic experiments with different noise distributions in sRGB space and real-world experiments on a denoising benchmark dataset in raw-RGB space.
Tao Huang 0022, Songjiang Li, Xu Jia 0012, Huchuan Lu, Jianzhuang Liu
CVPR4
2021 Multi-Target Domain Adaptation With Collaborative Consistency Learning
abstract
Recently unsupervised domain adaptation for the semantic segmentation task has become more and more popular due to high-cost of pixel-level annotation on real-world images. However, most domain adaptation methods are only restricted to single-source-single-target pair, and can not be directly extended to multiple target domains. In this work, we propose a collaborative learning framework to achieve unsupervised multi-target domain adaptation. An unsupervised domain adaptation expert model is first trained for each source-target pair and is further encouraged to collaborate with each other through a bridge built between different target domains. These expert models are further improved by adding the regularization of making the consistent pixel-wise prediction for each sample with the same structured context. To obtain a single model that works across multiple target domains, we propose to simultaneously learn a student model which is trained to not only imitate the output of each expert on the corresponding target domain, but also to pull different expert close to each other with regularization on their weights. Extensive experiments demonstrate that the proposed method can effectively exploit rich structured information contained in both labeled source domain and multiple unlabeled target domains. Not only does it perform well across multiple target domains but also performs favorably against state-of-the-art unsupervised domain adaptation methods specially trained on a single source-target pair. Code is available at https://github.com/junpan19/MTDA.
Takashi Isobe, Xu Jia 0012, Shuaijun Chen, Yongjie Shi, Jianzhuang Liu, Huchuan Lu, Shengjin Wang
CVPR7
2021 Calibrated RGB-D Salient Object Detection
abstract
Complex backgrounds and similar appearances between objects and their surroundings are generally recognized as challenging scenarios in Salient Object Detection (SOD). This naturally leads to the incorporation of depth information in addition to the conventional RGB image as input, known as RGB-D SOD or depth-aware SOD. Meanwhile, this emerging line of research has been considerably hindered by the noise and ambiguity that prevail in raw depth images. To address the aforementioned issues, we propose a Depth Calibration and Fusion (DCF) framework that contains two novel components: 1) a learning strategy to calibrate the latent bias in the original depth maps towards boosting the SOD performance; 2) a simple yet effective cross reference module to fuse features from both RGB and depth modalities. Extensive empirical experiments demonstrate that the proposed approach achieves superior performance against 27 state-of-the-art methods. Moreover, our depth calibration strategy alone can work as a preprocessing step; empirically it results in noticeable improvements when being applied to existing cutting-edge RGB-D SOD models. Source code is available at https://github.com/jiwei0921/DCF.
Wei Ji 0011, Miao Zhang 0004, Yongri Piao, Shunyu Yao 0004, Qi Bi, Kai Ma 0002, Yefeng Zheng 0001, Huchuan Lu, Li Cheng 0001
CVPR10
2021 Watching You: Global-Guided Reciprocal Learning for Video-Based Person Re-Identification
abstract
Video-based person re-identification (Re-ID) aims to automatically retrieve video sequences of the same person under non-overlapping cameras. To achieve this goal, it is the key to fully utilize abundant spatial and temporal cues in videos. Existing methods usually focus on the most conspicuous image regions, thus they may easily miss out fine-grained clues due to the person varieties in image sequences. To address above issues, in this paper, we propose a novel Global-guided Reciprocal Learning (GRL) framework for video-based person Re-ID. Specifically, we first propose a Global-guided Correlation Estimation (GCE) to generate feature correlation maps of local features and global features, which help to localize the high- and low-correlation regions for identifying the same person. After that, the discriminative features are disentangled into high-correlation features and low-correlation features under the guidance of the global representations. Moreover, a novel Temporal Reciprocal Learning (TRL) mechanism is designed to sequentially enhance the high-correlation semantic information and accumulate the low-correlation sub-critical clues. Extensive experiments are conducted on three public benchmarks. The experimental results indicate that our approach can achieve better performance than other state-of-the-art approaches. The code is released at https://github.com/flysnowtiger/GRL.
Xuehu Liu, Chenyang Yu, Huchuan Lu, Xiaoyun Yang
CVPR4
2021 LightTrack: Finding Lightweight Neural Networks for Object Tracking via One-Shot Architecture Search
abstract
Object tracking has achieved significant progress over the past few years. However, state-of-the-art trackers become increasingly heavy and expensive, which limits their deployments in resource-constrained applications. In this work, we present LightTrack, which uses neural architecture search (NAS) to design more lightweight and efficient object trackers. Comprehensive experiments show that our LightTrack is effective. It can find trackers that achieve superior performance compared to handcrafted SOTA trackers, such as SiamRPN++ [30] and Ocean [56], while using much fewer model Flops and parameters. Moreover, when deployed on resource-constrained mobile chipsets, the discovered trackers run much faster. For example, on Snapdragon 845 Adreno GPU, LightTrack runs 12× faster than Ocean, while using 13× fewer parameters and 38× fewer Flops. Such improvements might narrow the gap between academic models and industrial deployments in object tracking task. LightTrack is released at here.
Bin Yan 0004, Houwen Peng, Dong Wang 0004, Jianlong Fu, Huchuan Lu
CVPR6
2021 Alpha-Refine: Boosting Tracking Performance by Precise Bounding Box Estimation
abstract
Visual object tracking aims to precisely estimate the bounding box for the given target, which is a challenging problem due to factors such as deformation and occlusion. Many recent trackers adopt the multiple-stage strategy to improve bounding box estimation. These methods first coarsely locate the target and then refine the initial prediction in the following stages. However, existing approaches still suffer from limited precision, and the coupling of different stages severely restricts the method’s transferability. This work proposes a novel, flexible, and accurate refinement module called Alpha-Refine (AR), which can significantly improve the base trackers’ box estimation quality. By exploring a series of design options, we conclude that the key to successful refinement is extracting and maintaining detailed spatial information as much as possible. Following this principle, Alpha-Refine adopts a pixel-wise correlation, a corner prediction head, and an auxiliary mask head as the core components. Comprehensive experiments on TrackingNet, LaSOT, GOT-10K, and VOT2020 benchmarks with multiple base trackers show that our approach significantly improves the base tracker’s performance with little extra latency. The proposed Alpha-Refine method leads to a series of strengthened trackers, among which the ARSiamRPN (AR strengthened SiamRPNpp) and the ARDiMP50 (AR strengthened DiMP50) achieve good efficiency-precision trade-off, while the ARDiMPsuper (AR strengthened DiMPsuper) achieves very competitive performance at a realtime speed. Code and pretrained models are available at https://github.com/MasterBin-IIAU/AlphaRefine.
Bin Yan 0004, Xinyu Zhang 0017, Dong Wang 0004, Huchuan Lu, Xiaoyun Yang
CVPR4
2021 Self-Generated Defocus Blur Detection via Dual Adversarial Discriminators
abstract
Although existing fully-supervised defocus blur detection (DBD) models significantly improve performance, training such deep models requires abundant pixel-level manual annotation, which is highly time-consuming and error-prone. Addressing this issue, this paper makes an effort to train a deep DBD model without using any pixel-level annotation. The core insight is that a defocus blur region/focused clear area can be arbitrarily pasted to a given realistic full blurred image/full clear image without affecting the judgment of the full blurred image/full clear image. Specifically, we train a generator G in an adversarial manner against dual discriminators Dcand Db. G learns to produce a DBD mask that generates a composite clear image and a composite blurred image through copying the focused area and unfocused region from corresponding source image to another full clear image and full blurred image. Then, Dcand Dbcan not distinguish them from realistic full clear image and full blurred image simultaneously, achieving a self-generated DBD by an implicit manner to define what a defocus blur area is. Besides, we propose a bilateral triplet-excavating constraint to avoid the degenerate problem caused by the case one discriminator defeats the other one. Comprehensive experiments on two widely-used DBD datasets demonstrate the superiority of the proposed approach. Source codes are available at: https://github.com/shangcai1/SG.
Wenda Zhao 0003, Cai Shang, Huchuan Lu
CVPR3
2021 Learning Spatio-Temporal Transformer for Visual Tracking
abstract
In this paper, we present a new tracking architecture with an encoder-decoder transformer as the key component. The encoder models the global spatio-temporal feature dependencies between target objects and search regions, while the decoder learns a query embedding to predict the spatial positions of the target objects. Our method casts object tracking as a direct bounding box prediction problem, without using any proposals or predefined anchors. With the encoder-decoder transformer, the prediction of objects just uses a simple fully-convolutional network, which estimates the corners of objects directly. The whole method is end-to-end, does not need any postprocessing steps such as cosine window and bounding box smoothing, thus largely simplifying existing tracking pipelines. The proposed tracker achieves state-of-the-art performance on multiple challenging short-term and long-term benchmarks, while running at real-time speed, being 6× faster than Siam R-CNN [54]. Code and models are open-sourced at https://github.com/researchmm/Stark.
Bin Yan 0004, Houwen Peng, Jianlong Fu, Dong Wang 0004, Huchuan Lu
ICCV5
2021 Video Annotation for Visual Tracking via Selection and Refinement
abstract
Deep learning based visual trackers entail offline pre-training on large volumes of video datasets with accurate bounding box annotations that are labor-expensive to achieve. We present a new framework to facilitate bounding box annotations for video sequences, which investigates a selection-and-refinement strategy to automatically improve the preliminary annotations generated by tracking algorithms. A temporal assessment network (T-Assess Net) is proposed which is able to capture the temporal coherence of target locations and select reliable tracking results by measuring their quality. Meanwhile, a visual-geometry refinement network (VG-Refine Net) is also designed to further enhance the selected tracking results by considering both target appearance and temporal geometry constraints, allowing inaccurate tracking results to be corrected. The combination of the above two networks provides a principled approach to ensure the quality of automatic video annotation. Experiments on large scale tracking benchmarks demonstrate that our method can deliver highly accurate bounding box annotations and significantly reduce human labor by 94.0%, yielding an effective means to further boost tracking performance with augmented training data.
Kenan Dai, Jie Zhao 0014, Lijun Wang 0001, Dong Wang 0004, Huchuan Lu, Xuesheng Qian, Xiaoyun Yang
ICCV6
2021 MFNet: Multi-filter Directive Network for Weakly Supervised Salient Object Detection
abstract
Weakly supervised salient object detection (WSOD) targets to train a CNNs-based saliency network using only low-cost annotations. Existing WSOD methods take various techniques to pursue single "high-quality" pseudo label from low-cost annotations and then develop their saliency networks. Though these methods have achieved good performance, the generated single label is inevitably affected by adopted refinement algorithms and shows prejudiced characteristics which further influence the saliency networks. In this work, we introduce a new multiple-pseudo-label framework to integrate more comprehensive and accurate saliency cues from multiple labels, avoiding the aforementioned problem. Specifically, we propose a multi-filter directive network (MFNet) including a saliency network as well as multiple directive filters. The directive filter (DF) is designed to extract and filter more accurate saliency cues from the noisy pseudo labels. The multiple accurate cues from multiple DFs are then simultaneously propagated to the saliency network with a multi-guidance loss. Extensive experiments on five datasets over four metrics demonstrate that our method outperforms all the existing con-generic methods. Moreover, it is also worth noting that our framework is flexible enough to apply to existing methods and improve their performance. The code and results of our method are available at https://github.com/OIPLab-DUT/MFNet.
Yongri Piao, Miao Zhang 0004, Huchuan Lu
ICCV4
2021 Can Scale-Consistent Monocular Depth Be Learned in a Self-Supervised Scale-Invariant Manner?
abstract
Geometric constraints are shown to enforce scale consistency and remedy the scale ambiguity issue in self-supervised monocular depth estimation. Meanwhile, scale-invariant losses focus on learning relative depth, leading to accurate relative depth prediction. To combine the best of both worlds, we learn scale-consistent self-supervised depth in a scale-invariant manner. Towards this goal, we present a scale-aware geometric (SAG) loss, which enforces scale consistency through point cloud alignment. Compared to prior arts, SAG loss takes relative scale into consideration during relative motion estimation, enabling more precise alignment and explicit supervision for scale inference. In addition, a novel two-stream architecture for depth estimation is designed, which disentangles scale from depth estimation and allows depth to be learned in a scale-invariant manner. The integration of SAG loss and two-stream network enables more consistent scale inference and more accurate relative depth estimation. Our method achieves state-of-the-art performance under both scale-invariant and scale-dependent evaluation settings.
Lijun Wang 0001, Yifan Wang 0004, Linzhao Wang, Yunlong Zhan, Huchuan Lu
ICCV6
2021 Learning Motion-Appearance Co-Attention for Zero-Shot Video Object Segmentation
abstract
How to make the appearance and motion information interact effectively to accommodate complex scenarios is a fundamental issue in flow-based zero-shot video object segmentation. In this paper, we propose an Attentive Multi-Modality Collaboration Network (AMC-Net) to utilize appearance and motion information uniformly. Specifically, AMC-Net fuses robust information from multi-modality features and promotes their collaboration in two stages. First, we propose a Multi-Modality Co-Attention Gate (MCG) on the bilateral encoder branches, in which a gate function is used to formulate co-attention scores for balancing the contributions of multi-modality features and suppressing the redundant and misleading information. Then, we propose a Motion Correction Module (MCM) with a visualmotion attention mechanism, which is constructed to emphasize the features of foreground objects by incorporating the spatio-temporal correspondence between appearance and motion cues. Extensive experiments on three public challenging benchmark datasets verify that our proposed network performs favorably against existing state-of-the-art methods via training with fewer data. The code is released at https://github.com/isyangshu/AMC-Net.
Shu Yang 0004, Lu Zhang 0053, Jinqing Qi, Huchuan Lu
ICCV4
2021 CR-Fill: Generative Image Inpainting with Auxiliary Contextual Reconstruction
abstract
Recent deep generative inpainting methods use attention layers to allow the generator to explicitly borrow feature patches from the known region to complete a missing region. Due to the lack of supervision signals for the correspondence between missing regions and known regions, it may fail to find proper reference features, which often leads to artifacts in the results. Also, it computes pair-wise similarity across the entire feature map during inference bringing a significant computational overhead. To address this issue, we propose to teach such patch-borrowing behavior to an attention-free generator by joint training of an auxiliary contextual reconstruction task, which encourages the generated output to be plausible even when reconstructed by surrounding regions. The auxiliary branch can be seen as a learnable loss function, i.e. named as contextual reconstruction (CR) loss, where query-reference feature similarity and reference-based reconstructor are jointly optimized with the inpainting generator. The auxiliary branch ( i.e. CR loss) is required only during training, and only the inpainting generator is required during the inference. Experimental results demonstrate that the proposed inpainting model compares favourably against the state-of-the-art in terms of quantitative and visual performance. Code is available at https://github.com/zengxianyu/crfill.
Yu Zeng 0001, Zhe Lin 0001, Huchuan Lu, Vishal M. Patel
ICCV3
2021 Dynamic Context-Sensitive Filtering Network for Video Salient Object Detection
abstract
The ability to capture inter-frame dynamics has been critical to the development of video salient object detection (VSOD). While many works have achieved great success in this field, a deeper insight into its dynamic nature should be developed. In this work, we aim to answer the following questions: How can a model adjust itself to dynamic variations as well as perceive fine differences in the real-world environment; How are the temporal dynamics well introduced into spatial information over time? To this end, we propose a dynamic context-sensitive filtering network (DCFNet) equipped with a dynamic context-sensitive filtering module (DCFM) and an effective bidirectional dynamic fusion strategy. The proposed DCFM sheds new light on dynamic filter generation by extracting location-related affinities between consecutive frames. Our bidirectional dynamic fusion strategy encourages the interaction of spatial and temporal information in a dynamic manner. Experimental results demonstrate that our proposed method can achieve state-of-the-art performance on most VSOD datasets while ensuring a real-time speed of 28 fps. The source code is publicly available at https://github.com/OIPLab-DUT/DCFNet.
Miao Zhang 0004, Jie Liu 0044, Yongri Piao, Shunyu Yao 0004, Wei Ji 0011, Huchuan Lu, Zhongxuan Luo
ICCV8
2021 Automatic Polyp Segmentation via Multi-scale Subtraction Network
Xiaoqi Zhao 0003, Lihe Zhang, Huchuan Lu
MICCAI (1)3
2021 Weakly-Supervised Temporal Action Localization via Cross-Stream Collaborative Learning
abstract
Weakly supervised temporal action localization (WTAL) is a challenging task as only video-level category labels are available during training stage. Without precise temporal annotations, most approaches rely on complementary RGB and optical flow features to predict the start and end frame of each action category in a video. However, existing approaches simply resort to either concatenation or weighted sum to learn how to take advantages of these two modalities for accurate action localization, which ignore the substantial variance between such two modalities. In this paper, we present Cross-Stream Collaborative Learning (CSCL) to address these issues. The proposed CSCL introduce a cross-stream weighting module to identify which modality is more robust during training and take advantage of the robust modality to guide the weaker one. Furthermore, we suppress the snippets which has high action-ness scores in both modalities to further exploiting the complementary property between two modalities. In addition, we bring the concept of co-training for WTAL and take both modalities into account for pseudo label generation to help training a stronger model. Extensive experiments conducted on THUMOS14 and ActivityNet dataset demonstrate that CSCL achieves a favorable performance against state-of-the-arts methods.
Xu Jia 0012, Huchuan Lu, Xiang Ruan
ACM Multimedia3
2021 Polar Ray: A Single-stage Angle-free Detector for Oriented Object Detection in Aerial Images
abstract
Oriented bounding boxes are widely used for object detection in aerial images. Existing oriented object detection methods typically follow the general object detection paradigm by adding an extra rotation angle on the horizontal bounding boxes. However, the angular periodicity incurs the difficulty in angle regression and rotation sensitivity on bounding boxes. In this paper, we propose a new anchor-free oriented object detector, Polar Ray Network (PRNet), where object keypoints are represented by polar coordinates without angle regression. Our PRNet learns a set of polar rays from the object center to boundary with predefined equal-distributed angles. We introduce a dynamic PointConv module to optimize the regression of polar ray by incorporating object corner features. Furthermore, a classification feature guidance module is presented to improve the classification accuracy by incorporating more spatial contents from polar rays. Experimental results on two public datasets, i.e., DOTA and HRSC2016, demonstrate that the proposed PRNet significantly outperforms existing anchor-free detectors, and shows highly competitiveness with the state-of-the-art two-stage anchor-based methods.
Shuai Liu 0009, Lu Zhang 0053, Shuai Hao 0007, Huchuan Lu, You He 0002
ACM Multimedia4
2021 Auto-MSFNet: Search Multi-scale Fusion Network for Salient Object Detection
abstract
Multi-scale features fusion plays a critical role in salient object detection. Most of existing methods have achieved remarkable performance by exploiting various multi-scale features fusion strategies. However, an elegant fusion framework requires expert knowledge and experience, heavily relying on laborious trial and error. In this paper, we propose a multi-scale features fusion framework based on Neural Architecture Search (NAS), named Auto-MSFNet. First, we design a novel search cell, named FusionCell to automatically decide multi-scale features aggregation. Rather than searching one repeatable cell stacked, we allow different FusionCells to flexibly integrate multi-level features. Simultaneously, considering features generated from CNNs are naturally spatial and channel-wise, we propose a new search space for efficiently focusing on the most relevant information. The search space mitigates incomplete object structures or over-predicted foreground regions caused by progressive fusion. Second, we propose a progressive polishing loss to further obtain exquisite boundaries by penalizing misalignment of salient object boundaries. Extensive experiments on five benchmark datasets demonstrate the effectiveness of the proposed method and achieve state-of-the-art performance on four evaluation metrics. The code and results of our method are available at https://github.com/OIPLab-DUT/Auto-MSFNet.
Miao Zhang 0004, Tingwei Liu, Yongri Piao, Shunyu Yao 0004, Huchuan Lu
ACM Multimedia5
2021 HAT: Hierarchical Aggregation Transformers for Person Re-identification
abstract
Recently, with the advance of deep Convolutional Neural Networks (CNNs), person Re-Identification (Re-ID) has witnessed great success in various applications.However, with limited receptive fields of CNNs, it is still challenging to extract discriminative representations in a global view for persons under non-overlapped cameras.Meanwhile, Transformers demonstrate strong abilities of modeling long-range dependencies for spatial and sequential data.In this work, we take advantages of both CNNs and Transformers, and propose a novel learning framework named Hierarchical Aggregation Transformer (HAT) for image-based person Re-ID with high performance.To achieve this goal, we first propose a Deeply Supervised Aggregation (DSA) to recurrently aggregate hierarchical features from CNN backbones.With multi-granularity supervision, the DSA can enhance multi-scale features for person retrieval, which is very different from previous methods.Then, we introduce a Transformer-based Feature Calibration (TFC) to integrate low-level detail information as the global prior for high-level semantic information.The proposed TFC is inserted to each level of hierarchical features, resulting in great performance improvements.To our best knowledge, this work is the first to take advantages of both CNNs and Transformers for image-based person Re-ID.Comprehensive experiments on four large-scale Re-ID benchmarks demonstrate that our method shows better results than several state-of-the-art methods.The code is released at https://github.com/AI-Zhpp/HAT.
Guowen Zhang, Jinqing Qi, Huchuan Lu
ACM Multimedia4
2021 Multi-Source Fusion and Automatic Predictor Selection for Zero-Shot Video Object Segmentation
abstract
Location and appearance are the key cues for video object segmentation. Many sources such as RGB, depth, optical flow and static saliency can provide useful information about the objects. However, existing approaches only utilize the RGB or RGB and optical flow. In this paper, we propose a novel multi-source fusion network for zero-shot video object segmentation. With the help of interoceptive spatial attention module (ISAM), spatial importance of each source is highlighted. Furthermore, we design a feature purification module (FPM) to filter the inter-source incompatible features. By the ISAM and FPM, the multi-source features are effectively fused. In addition, we put forward an automatic predictor selection network (APS) to select the better prediction of either the static saliency predictor or the moving object predictor in order to prevent over-reliance on the failed results caused by low-quality optical flow maps. Extensive experiments on three challenging public benchmarks (i.e. DAVIS$_16 $, Youtube-Objects and FBMS) show that the proposed model achieves compelling performance against the state-of-the-arts. The source code will be publicly available at https://github.com/Xiaoqi-Zhao-DLUT/Multi-Source-APS-ZVOS
Xiaoqi Zhao 0003, Youwei Pang, Lihe Zhang, Huchuan Lu
ACM Multimedia5
2021 Joint Semantic Mining for Weakly Supervised RGB-D Salient Object Detection
abstract
Training saliency detection models with weak supervisions, e.g., image-level tags or captions, is appealing as it removes the costly demand of per-pixel annotations. Despite the rapid progress of RGB-D saliency detection in fully-supervised setting, it however remains an unexplored territory when only weak supervision signals are available. This paper is set to tackle the problem of weakly-supervised RGB-D salient object detection. The key insight in this effort is the idea of maintaining per-pixel pseudo-labels with iterative refinements by reconciling the multimodal input signals in our joint semantic mining (JSM). Considering the large variations in the raw depth map and the lack of explicit pixel-level supervisions, we propose spatial semantic modeling (SSM) to capture saliency-specific depth cues from the raw depth and produce depth-refined pseudo-labels. Moreover, tags and captions are incorporated via a fill-in-the-blank training in our textual semantic modeling (TSM) to estimate the confidences of competing pseudo-labels. At test time, our model involves only a light-weight sub-network of the training pipeline, i.e., it requires only an RGB image as input, thus allowing efficient inference. Extensive evaluations demonstrate the effectiveness of our approach under the weakly-supervised setting. Importantly, our method could also be adapted to work in both fully-supervised and unsupervised paradigms. In each of these scenarios, superior performance has been attained by our approach with comparing to the state-of-the-art dedicated methods. As a by-product, a CapS dataset is constructed by augmenting existing benchmark training set with additional image tags and captions.
Wei Ji 0011, Qi Bi, Miao Zhang 0004, Yongri Piao, Huchuan Lu, Li Cheng 0001
NeurIPS7
2021 IPE Transformer for Depth Completion with Input-Aware Positional Embeddings
Bocen Li, Guozhen Li, Haiting Wang, Lijun Wang 0001, Zhenfei Gong, Huchuan Lu
PRCV (4)7
2021 Online visual tracking via cross-similarity-based siamese network
abstract
Summary Among deep‐learning‐based trackers, the siamese‐based method inspires many researchers due to its effectiveness and simplicity. However, the traditional siamese tracker has not achieved satisfactory performance due to the limited representation ability and the lack of appropriate model update strategy. To cover the shortage of siamese models, we proposed a cross‐similarity‐based siamese network with three contributions. First, we introduce a novel cross similarity module into the SiameseFC framework, which could improve the matching ability of fully convolutional networks during the tracking process. Second, we propose a novel attention weighting layer to emphasize various contributions of matching scores in different positions. This adaptive attention weighting scheme makes our tracker well adapt to the appearance change caused by pose variation, partial occlusion, and so on. Third, we develop a simple yet effective model update strategy, which exploits an independent classification model to invoke the model fine‐tuning process. Experimental results on the standard tracking benchmark show that our tracker performs much better than the baseline SiameseFC method and also achieves promising results in comparisons to other state‐of‐the‐art algorithms.
Huchuan Lu
Concurr. Comput. Pract. Exp.2
2021 Learning Adaptive Attribute-Driven Representation for Real-Time RGB-T Tracking
Dong Wang 0004, Huchuan Lu, Xiaoyun Yang
Int. J. Comput. Vis.3
2021 Learning Regression and Verification Networks for Robust Long-term Tracking
Lijun Wang 0001, Dong Wang 0004, Jinqing Qi, Huchuan Lu
Int. J. Comput. Vis.5
2021 Self-attention guided representation learning for image-text matching
Xuefei Qi, Ying Zhang 0021, Jinqing Qi, Huchuan Lu
Neurocomputing4
2021 Deeply supervised group recursive saliency prediction
Lingwei Kong, Lu Zhang 0053, Huchuan Lu
Neurocomputing4
2021 Temporal consistent portrait video segmentation
Yifan Wang 0004, Lijun Wang 0001, Fenghua Yang, Huchuan Lu
Pattern Recognit.5
2021 Residual multi-task learning for facial landmark localization and expression recognition
Wenlong Guan, Peixia Li, Naoki Ikeda, Kosuke Hirasawa, Huchuan Lu
Pattern Recognit.6
2021 Spatial context-aware network for salient object detection
Yuqiu Kong, Mengyang Feng, Xin Li 0003, Huchuan Lu, Xiuping Liu
Pattern Recognit.4
2021 Deep mutual learning for visual object tracking
Haojie Zhao, Gang Yang 0002, Dong Wang 0004, Huchuan Lu
Pattern Recognit.4
2021 Looking for the Detail and Context Devils: High-Resolution Salient Object Detection
abstract
In recent years, Salient Object Detection (SOD) has shown great success with the achievements of large-scale benchmarks and deep learning techniques. However, existing SOD methods mainly focus on natural images with low-resolutions, e.g., 400×400 or less. This drawback hinders them for advanced practical applications, which need high-resolution, detail-aware results. Besides, lacking of the boundary detail and semantic context of salient objects is also a key concern for accurate SOD. To address these issues, in this work we focus on the High-Resolution Salient Object Detection (HRSOD) task. Technically, we propose the first end-to-end learnable framework, named Dual ReFinement Network (DRFNet), for fully automatic HRSOD. More specifically, the proposed DRFNet consists of a shared feature extractor and two effective refinement heads. By decoupling the detail and context information, one refinement head adopts a global-aware feature pyramid. Without increasing too much computational burden, it can boost the spatial detail information, which narrows the gap between high-level semantics and low-level details. In parallel, the other refinement head adopts hybrid dilated convolutional blocks and group-wise upsamplings, which are very efficient in extracting contextual information. Based on the dual refinements, our approach can enlarge receptive fields and obtain more discriminative features from high-resolution images. Experimental results on high-resolution benchmarks (the public DUT-HRSOD and the proposed DAVIS-SOD) demonstrate that our method is not only efficient but also performs more accurate than other state-of-the-arts. Besides, our method generalizes well on typical low-resolution benchmarks.
Wei Liu 0044, Yi Zeng 0006, Yinjie Lei, Huchuan Lu
IEEE Trans. Image Process.5
2021 Jointly Modeling Motion and Appearance Cues for Robust RGB-T Tracking
abstract
In this study, we propose a novel RGB-T tracking framework by jointly modeling both appearance and motion cues. First, to obtain a robust appearance model, we develop a novel late fusion method to infer the fusion weight maps of both RGB and thermal (T) modalities. The fusion weights are determined by using offline-trained global and local multimodal fusion networks, and then adopted to linearly combine the response maps of RGB and T modalities. Second, when the appearance cue is unreliable, we comprehensively take motion cues, i.e., target and camera motions, into account to make the tracker robust. We further propose a tracker switcher to switch the appearance and motion trackers flexibly. Numerous results on three recent RGB-T tracking datasets show that the proposed tracker performs significantly better than other state-of-the-art algorithms.
Jie Zhao 0014, Chunjuan Bo, Dong Wang 0004, Huchuan Lu, Xiaoyun Yang
IEEE Trans. Image Process.5
2021 Defocus Blur Detection via Boosting Diversity of Deep Ensemble Networks
abstract
Existing defocus blur detection (DBD) methods usually explore multi-scale and multi-level features to improve performance. However, defocus blur regions normally have incomplete semantic information, which will reduce DBD's performance if it can't be used properly. In this paper, we address the above problem by exploring deep ensemble networks, where we boost diversity of defocus blur detectors to force the network to generate diverse results that some rely more on high-level semantic information while some ones rely more on low-level information. Then, diverse result ensemble makes detection errors cancel out each other. Specifically, we propose two deep ensemble networks (e.g., adaptive ensemble network (AENet) and encoder-feature ensemble network (EFENet)), which focus on boosting diversity while costing less computation. AENet constructs different light-weight sequential adapters for one backbone network to generate diverse results without introducing too many parameters and computation. AENet is optimized only by the self- negative correlation loss. On the other hand, we propose EFENet by exploring the diversity of multiple encoded features and ensemble strategies of features (e.g., group-channel uniformly weighted average ensemble and self-gate weighted ensemble). Diversity is represented by encoded features with less parameters, and a simple mean squared error loss can achieve the superior performance. Experimental results demonstrate the superiority over the state-of-the-arts in terms of accuracy and speed. Codes and models are available at: https://github.com/wdzhao123/DENets.
Wenda Zhao 0003, Xueqing Hou, You He 0002, Huchuan Lu
IEEE Trans. Image Process.4
2021 Semantic Scene Labeling via Deep Nested Level Set
abstract
Semantic scene labeling plays a very important role in intelligent transportation tasks, such as autonomous driving and advanced driver assistance. Recently, thanks to the advances of deep learning, significant improvements have been achieved for this pixel-wise labeling task. Although effective, current methods lack of explicitly modeling the boundary of objects, resulting in inaccurate labeling results. Meanwhile, traditional level set based methods perform better to capture the evolution of boundaries. However, they are sensitive to the model initialization. To address these issues, in this work we propose a novel deep learning framework, named deep nested level set (DNLS) for boundary-aware semantic scene labeling. Different from previous works, our proposed framework explicitly takes deep learned features and object boundary information into account. More specifically, our proposed framework first predicts semantic probability maps and boundary locations of objects using a bifurcated fully convolutional network (BFCN). Then, these probability maps are seamlessly integrated into a nested level set function for accurate scene labeling. As a result, our approach can automatically initialize the nested level set function, and the whole framework can be trained in an end-to-end manner, providing a new solution for accurate semantic scene parsing. Extensive experiments on public CamVid and Cityscapes datasets demonstrate that our proposed framework produces high-quality predictions with clear object boundaries and spatial consistency.
Wei Liu 0044, Yinjie Lei, Huchuan Lu
IEEE Trans. Intell. Transp. Syst.4
2020 Exploit and Replace: An Asymmetrical Two-Stream Architecture for Versatile Light Field Saliency Detection
abstract
Light field saliency detection is becoming of increasing interest in recent years due to the significant improvements in challenging scenes by using abundant light field cues. However, high dimension of light field data poses computation-intensive and memory-intensive challenges, and light field data access is far less ubiquitous as RGB data. These may severely impede practical applications of light field saliency detection. In this paper, we introduce an asymmetrical two-stream architecture inspired by knowledge distillation to confront these challenges. First, we design a teacher network to learn to exploit focal slices for higher requirements on desktop computers and meanwhile transfer comprehensive focusness knowledge to the student network. Our teacher network is achieved relying on two tailor-made modules, namely multi-focusness recruiting module (MFRM) and multi-focusness screening module (MFSM), respectively. Second, we propose two distillation schemes to train a student network towards memory and computation efficiency while ensuring the performance. The proposed distillation schemes ensure better absorption of focusness knowledge and enable the student to replace the focal slices with a single RGB image in an user-friendly way. We conduct the experiments on three benchmark datasets and demonstrate that our teacher network achieves state-of-the-arts performance and student network (ResNet18) achieves Top-1 accuracies on HFUT-LFSD dataset and Top-4 on DUT-LFSD, which tremendously minimizes the model size by 56% and boosts the Frame Per Second (FPS) by 159%, compared with the best performing method.
Yongri Piao, Zhengkun Rong, Miao Zhang 0004, Huchuan Lu
AAAI4
2020 Multi-Type Self-Attention Guided Degraded Saliency Detection
abstract
Existing saliency detection techniques are sensitive to image quality and perform poorly on degraded images. In this paper, we systematically analyze the current status of the research on detecting salient objects from degraded images and then propose a new multi-type self-attention network, namely MSANet, for degraded saliency detection. The main contributions include: 1) Applying attention transfer learning to promote semantic detail perception and internal feature mining of the target network on degraded images; 2) Developing a multi-type self-attention mechanism to achieve the weight recalculation of multi-scale features. By computing global and local attention scores, we obtain the weighted features of different scales, effectively suppress the interference of noise and redundant information, and achieve a more complete boundary extraction. The proposed MSANet converts low-quality inputs to high-quality saliency maps directly in an end-to-end fashion. Experiments on seven widely-used datasets show that our approach produces good performance on both clear and degraded images.
Ziqi Zhou 0002, Zheng Wang 0008, Huchuan Lu, Song Wang 0002, Meijun Sun
AAAI3
2020 Synergistic Saliency and Depth Prediction for RGB-D Saliency Detection
Yue Wang 0038, James H. Elder, Runmin Wu, Huchuan Lu, Lu Zhang 0053
ACCV (2)5
2020 High-Performance Long-Term Tracking With Meta-Updater
abstract
Long-term visual tracking has drawn increasing attention because it is much closer to practical applications than short-term tracking. Most top-ranked long-term trackers adopt the offline-trained Siamese architectures, thus,they cannot benefit from great progress of short-term trackers with online update. However, it is quite risky to straightforwardly introduce online-update-based trackers to solve the long-term problem, due to long-term uncertain and noisy observations. In this work, we propose a novel offline-trained Meta-Updater to address an important but unsolved problem: Is the tracker ready for updating in the current frame? The proposed meta-updater can effectively integrate geometric, discriminative, and appearance cues in a sequential manner, and then mine the sequential information with a designed cascaded LSTM module. Our meta-updater learns a binary output to guide the tracker’s update and can be easily embedded into different trackers. This work also introduces a long-term tracking framework consisting of an online local tracker, an online verifier, a SiamRPN-based re-detector, and our meta-updater. Numerous experimental results on the VOT2018LT,VOT2019LT, OxUvALT, TLP, and LaSOT benchmarks show that our tracker performs remarkably better than other competing algorithms. Our project is available on the website: https://github.com/Daikenan/LTMU.
Kenan Dai, Dong Wang 0004, Huchuan Lu, Xiaoyun Yang
CVPR5
2020 Pose-Guided Visible Part Matching for Occluded Person ReID
abstract
Occluded person re-identification is a challenging task as the appearance varies substantially with various obstacles, especially in the crowd scenario. To address this issue, we propose a Pose-guided Visible Part Matching (PVPM) method that jointly learns the discriminative features with pose-guided attention and self-mines the part visibility in an end-to-end framework. Specifically, the proposed PVPM includes two key components: 1) pose-guided attention (PGA) method for part feature pooling that exploits more discriminative local features; 2) pose-guided visibility predictor (PVP) that estimates whether a part suffers the occlusion or not. As there are no ground truth training annotations for the occluded part, we turn to utilize the characteristic of part correspondence in positive pairs and self-mining the correspondence scores via graph matching. The generated correspondence scores are then utilized as pseudo-labels for visibility predictor (PVP). Experimental results on three reported occluded benchmarks show that the proposed method achieves competitive performance to state-of-the-art methods. The source codes are available at https://github.com/hh23333/PVPM.
Shang Gao 0012, Jingya Wang 0001, Huchuan Lu, Zimo Liu
CVPR3
2020 Bi-Directional Relationship Inferring Network for Referring Image Segmentation
abstract
Most existing methods do not explicitly formulate the mutual guidance between vision and language. In this work, we propose a bi-directional relationship inferring network (BRINet) to model the dependencies of cross-modal information. In detail, the vision-guided linguistic attention is used to learn the adaptive linguistic context corresponding to each visual region. Combining with the language-guided visual attention, a bi-directional cross-modal attention module (BCAM) is built to learn the relationship between multi-modal features. Thus, the ultimate semantic context of the target object and referring expression can be represented accurately and consistently. Moreover, a gated bi-directional fusion module (GBFM) is designed to integrate the multi-level features where a gate function is used to guide the bi-directional flow of multi-level information. Extensive experiments on four benchmark datasets demonstrate that the proposed method outperforms other state-of-the-art methods under different evaluation metrics.
Zhiwei Hu, Lihe Zhang, Huchuan Lu
CVPR5
2020 Multi-Scale Interactive Network for Salient Object Detection
abstract
Deep-learning based salient object detection methods achieve great progress. However, the variable scale and unknown category of salient objects are great challenges all the time. These are closely related to the utilization of multi-level and multi-scale features. In this paper, we propose the aggregate interaction modules to integrate the features from adjacent levels, in which less noise is introduced because of only using small up-/down-sampling rates. To obtain more efficient multi-scale features from the integrated features, the self-interaction modules are embedded in each decoder unit. Besides, the class imbalance issue caused by the scale variation weakens the effect of the binary cross entropy loss and results in the spatial inconsistency of the predictions. Therefore, we exploit the consistency-enhanced loss to highlight the fore-/back-ground difference and preserve the intra-class consistency. Experimental results on five benchmark datasets demonstrate that the proposed method without any post-processing performs favorably against 23 state-of-the-art approaches. The source code will be publicly available at https://github.com/lartpang/MINet.
Youwei Pang, Xiaoqi Zhao 0003, Lihe Zhang, Huchuan Lu
CVPR4
2020 A2dele: Adaptive and Attentive Depth Distiller for Efficient RGB-D Salient Object Detection
abstract
Existing state-of-the-art RGB-D salient object detection methods explore RGB-D data relying on a two-stream architecture, in which an independent subnetwork is required to process depth data. This inevitably incurs extra computational costs and memory consumption, and using depth data during testing may hinder the practical applications of RGB-D saliency detection. To tackle these two dilemmas, we propose a depth distiller (A2dele) to explore the way of using network prediction and attention as two bridges to transfer the depth knowledge from the depth stream to the RGB stream. First, by adaptively minimizing the differences between predictions generated from the depth stream and RGB stream, we realize the desired control of pixel-wise depth knowledge transferred to the RGB stream. Second, to transfer the localization knowledge to RGB features, we encourage consistencies between the dilated prediction of the depth stream and the attention map from the RGB stream. As a result, we achieve a lightweight architecture without use of depth data at test time by embedding our A2dele. Our extensive experimental evaluation on five benchmarks demonstrate that our RGB stream achieves state-of-the-art performance, which tremendously minimizes the model size by 76% and runs 12 times faster, compared with the best performing method. Furthermore, our A2dele can be applied to existing RGB-D networks to significantly improve their efficiency while maintaining performance (boosts FPS by nearly twice for DMRA and 3 times for CPFP).
Yongri Piao, Zhengkun Rong, Miao Zhang 0004, Weisong Ren, Huchuan Lu
CVPR5
2020 SDC-Depth: Semantic Divide-and-Conquer Network for Monocular Depth Estimation
abstract
Monocular depth estimation is an ill-posed problem, and as such critically relies on scene priors and semantics. Due to its complexity, we propose a deep neural network model based on a semantic divide-and-conquer approach. Our model decomposes a scene into semantic segments, such as object instances and background stuff classes, and then predicts a scale and shift invariant depth map for each semantic segment in a canonical space. Semantic segments of the same category share the same depth decoder, so the global depth prediction task is decomposed into a series of category-specific ones, which are simpler to learn and easier to generalize to new scene types. Finally, our model stitches each local depth segment by predicting its scale and shift based on the global context of the image. The model is trained end-to-end using a multi-task loss for panoptic segmentation and depth prediction, and is therefore able to leverage large-scale panoptic segmentation datasets to boost its semantic understanding. We validate the effectiveness of our approach and show state-of-the-art performance on three benchmark datasets.
Lijun Wang 0001, Jianming Zhang 0001, Oliver Wang, Zhe Lin 0001, Huchuan Lu
CVPR5
2020 Cooling-Shrinking Attack: Blinding the Tracker With Imperceptible Noises
abstract
Adversarial attack of CNN aims at deceiving models to misbehave by adding imperceptible perturbations to images. This feature facilitates to understand neural networks deeply and to improve the robustness of deep learning models. Although several works have focused on attacking image classifiers and object detectors, an effective and efficient method for attacking single object trackers of any target in a model-free way remains lacking. In this paper, a cooling-shrinking attack method is proposed to deceive state-of-the-art SiameseRPN-based trackers. An effective and efficient perturbation generator is trained with a carefully designed adversarial loss, which can simultaneously cool hot regions where the target exists on the heatmaps and force the predicted bounding box to shrink, making the tracked target invisible to trackers. Numerous experiments on OTB100, VOT2018, and LaSOT datasets show that our method can effectively fool the state-of-the-art SiameseRPN++ tracker by adding small perturbations to the template or the search regions. Besides, our method has good transferability and is able to deceive other top-performance trackers such as DaSiamRPN, DaSiamRPN-UpdateNet, and DiMP. The source codes are available at https://github.com/MasterBin-IIAU/CSA.
Bin Yan 0004, Dong Wang 0004, Huchuan Lu, Xiaoyun Yang
CVPR3
2020 Select, Supplement and Focus for RGB-D Saliency Detection
abstract
Depth data containing a preponderance of discriminative power in location have been proven beneficial for accurate saliency prediction. However, RGB-D saliency detection methods are also negatively influenced by randomly distributed erroneous or missing regions on the depth map or along the object boundaries. This offers the possibility of achieving more effective inference by well designed models. In this paper, we propose a new framework for accurate RGB-D saliency detection taking account of global location and local detail complementarities from two modalities. This is achieved by designing a complimentary interaction module (CIM) to discriminatively select useful representation from the RGB and depth data, and effectively integrate cross-modal features. Benefiting from the proposed CIM, the fused features can accurately locate salient objects with fine edge details. Moreover, we propose a compensation-aware loss to improve the network's confidence in detecting hard samples. Comprehensive experiments on six public datasets demonstrate that our method outperforms 18 state-of-the-art methods.
Miao Zhang 0004, Weisong Ren, Yongri Piao, Zhengkun Rong, Huchuan Lu
CVPR5
2020 Accurate RGB-D Salient Object Detection via Collaborative Learning
Wei Ji 0011, Miao Zhang 0004, Yongri Piao, Huchuan Lu
ECCV (18)5
2020 Hierarchical Dynamic Filtering Network for RGB-D Salient Object Detection
Youwei Pang, Lihe Zhang, Xiaoqi Zhao 0003, Huchuan Lu
ECCV (25)4
2020 CLIFFNet for Monocular Depth Estimation with Hierarchical Embedding Loss
Lijun Wang 0001, Jianming Zhang 0001, Yifan Wang 0004, Huchuan Lu, Xiang Ruan
ECCV (5)4
2020 High-Resolution Image Inpainting with Iterative Confidence Feedback and Guided Upsampling
Yu Zeng 0001, Zhe Lin 0001, Jimei Yang, Jianming Zhang 0001, Eli Shechtman, Huchuan Lu
ECCV (19)6
2020 Asymmetric Two-Stream Architecture for Accurate RGB-D Saliency Detection
Miao Zhang 0004, Sun Xiao Fei, Jie Liu 0044, Yongri Piao, Huchuan Lu
ECCV (28)6
2020 Unsupervised Video Object Segmentation with Joint Hotspot Tracking
Lu Zhang 0053, Jianming Zhang 0001, Zhe Lin 0001, Radomír Mech, Huchuan Lu, You He 0002
ECCV (14)5
2020 Suppress and Balance: A Simple Gated Network for Salient Object Detection
Xiaoqi Zhao 0003, Youwei Pang, Lihe Zhang, Huchuan Lu, Lei Zhang 0006
ECCV (2)4
2020 A Single Stream Network for Robust and Real-Time RGB-D Salient Object Detection
Xiaoqi Zhao 0003, Lihe Zhang, Youwei Pang, Huchuan Lu, Lei Zhang 0006
ECCV (22)4
2020 Feature Reintegration over Differential Treatment: A Top-down and Adaptive Fusion Network for RGB-D Salient Object Detection
abstract
Most methods for RGB-D salient object detection (SOD) utilize the same fusion strategy to explore the cross-modal complementary information at each level. However, this may ignore different feature contributions from two modalities on different levels towards prediction. In this paper, we propose a novel top-down multi-level fusion structure where different fusion strategies are utilized to effectively explore the low-level and high-level features. This is achieved by designing the interweave fusion module (IFM) to effectively integrate the global information and designing the gated select fusion module (GSFM) to discriminatively select useful local information by filtering out the unnecessary one from RGB and depth data. Moreover, we propose an adaptive fusion module (AFM) to reintegrate the fused cross-modal features of each level to predict a more accurate result. Comprehensive experiments on 7 challenging benchmark datasets demonstrate that our method achieves the competitive performance over 14 state-of-the-art RGB-D alternative methods.
Miao Zhang 0004, Yu Zhang 0165, Yongri Piao, Beiqi Hu, Huchuan Lu
ACM Multimedia5
2020 Online Filtering Training Samples for Robust Visual Tracking
abstract
In recent years, discriminative trackers show its great tracking performance, that is mainly due to the online updating using samples collected during tracking. The model could adapt appearance changes of objects and the background well after updating. But these trackers have a serious disadvantage that wrong samples may cause severe model degradation. Most of the training samples in the tracking phase are obtained according to the tracking result of the current frame. Wrong training samples will be collected when the tracking result is inaccurate, seriously affecting the discrimination ability of the model. Besides, partial occlusion also leads to the same problem. In this paper, we propose an optimization module named MetricNet for online filtering training samples. It applies a matching network containing the classification and distance branches, and uses multiple metric methods for different type samples. MetricNet optimizes the training sample set by recognizing wrong and redundant samples, thereby improving the tracking performance. The proposed MetricNet can be regarded as an independent optimization module and integrated into all discriminative trackers updated online. Extensive experiments on three tracking datasets show its effectiveness and generalization ability. After applying MetricNet to MDNet, the tracking result is increased by 5.3% in terms of the success plot on the LaSOT dataset. Our project is available at https://github.com/zj5559/MetricNet.
Jie Zhao 0014, Kenan Dai, Dong Wang 0004, Huchuan Lu, Xiaoyun Yang
ACM Multimedia4
2020 Dynamically-Passed Contextual Information Network for Saliency Detection
Abdelhafid Dakhia, Tiantian Wang 0002, Huchuan Lu
PRCV (3)3
2020 CACNet: Salient object detection via context aggregation and contrast embedding
Hongguang Bo, Lihe Zhang, Huchuan Lu
Neurocomputing5
2020 Multi-attention guided feature fusion network for salient object detection
Anni Li, Jinqing Qi, Huchuan Lu
Neurocomputing3
2020 Segmentation based rotated bounding boxes prediction and image synthesizing for object detection of high resolution aerial images
Lijun Wang 0001, Huchuan Lu, You He 0002
Neurocomputing3
2020 Salient object detection via double random walks with dual restarts
Lihe Zhang, Huchuan Lu, Guohua Wei
Image Vis. Comput.4
2020 Defocus Blur Detection via Multi-Stream Bottom-Top-Bottom Network
abstract
Defocus blur detection (DBD) is aimed to estimate the probability of each pixel being in-focus or out-of-focus. This process has been paid considerable attention due to its remarkable potential applications. Accurate differentiation of homogeneous regions and detection of low-contrast focal regions, as well as suppression of background clutter, are challenges associated with DBD. To address these issues, we propose a multi-stream bottom-top-bottom fully convolutional network (BTBNet), which is the first attempt to develop an end-to-end deep network to solve the DBD problems. First, we develop a fully convolutional BTBNet to gradually integrate nearby feature levels of bottom to top and top to bottom. Then, considering that the degree of defocus blur is sensitive to scales, we propose multi-stream BTBNets that handle input images with different scales to improve the performance of DBD. Finally, a cascaded DBD map residual learning architecture is designed to gradually restore finer structures from the small scale to the large scale. To promote further study and evaluation of the DBD models, we construct a new database of 1100 challenging images and their pixel-wise defocus blur annotations. Experimental results on the existing and our new datasets demonstrate that the proposed method achieves significantly better performance than other state-of-the-art algorithms.
Wenda Zhao 0003, Fan Zhao 0006, Dong Wang 0004, Huchuan Lu
IEEE Trans. Pattern Anal. Mach. Intell.4
2020 Dynamic imposter based online instance matching for person search
Ju Dai, Huchuan Lu, Hongyu Wang 0001
Pattern Recognit.3
2020 Visual tracking by dynamic matching-classification network switching
Peixia Li, Dong Wang 0004, Huchuan Lu
Pattern Recognit.4
2020 Blind single image super-resolution with a mixture of deep networks
Yifan Wang 0004, Lijun Wang 0001, Hongyu Wang 0001, Peihua Li, Huchuan Lu
Pattern Recognit.5
2020 Global and local sensitivity guided key salient object re-augmentation for video saliency detection
Zheng Wang 0008, Ziqi Zhou 0002, Huchuan Lu, Jianmin Jiang
Pattern Recognit.3
2020 Non-rigid object tracking via deep multi-scale spatial-temporal discriminative saliency maps
Wei Liu 0044, Dong Wang 0004, Yinjie Lei, Hongyu Wang 0001, Huchuan Lu
Pattern Recognit.6
2020 Introduction to the Special Section on Deep Learning in Video Enhancement and Evaluation: The New Frontier
abstract
Although video enhancement and evaluation have been studied for many years, they are still challenging due to the evolutions of video acquiring and processing techniques. While the development of deep learning is undoubtedly exciting and has demonstrated its superior performance in a variety of applications, it is important to further investigate advanced technologies and solutions to bring seemingly endless possibilities for video enhancement and evaluation. This Special Section intends to collect some recent solutions toward video enhancement and evaluation.
Zhenzhong Chen 0001, Huchuan Lu, Junwei Han 0001
IEEE Trans. Circuits Syst. Video Technol.2
2020 Reverse Attention-Based Residual Network for Salient Object Detection
abstract
Benefiting from the quick development of deep convolutional neural networks, especially fully convolutional neural networks (FCNs), remarkable progresses have been achieved on salient object detection recently. Nevertheless, these FCNs based methods are still challenging to generate high resolution saliency maps, and also not applicable for subsequent applications due to their heavy model weights. In this paper, we propose a compact and efficient deep network with high accuracy for salient object detection. Firstly, we propose two strategies for initial prediction, one is a new designed multi-scale context module, the other is incorporating hand-crafted saliency priors. Secondly, we employ residual learning to refine it progressively by only learning the residual in each side-output, which can be achieved with few convolutional parameters, therefore leads to high compactness and high efficiency. Finally, we further design a novel reverse attention block to guide side-output residual learning in a top-down manner. Specifically, the current predicted salient regions are erased from each side-output feature, thus the missing object parts and details can be efficiently learned from these unerased regions, which results in high resolution and accuracy. Extensive experimental results on seven benchmark datasets demonstrate that the proposed network performs favorably against the state-of-the-art methods, and with advantages in terms of simplicity, efficiency and model size.
Shuhan Chen, Xiuli Tan, Huchuan Lu, Xuelong Hu, Yun Fu 0001
IEEE Trans. Image Process.4
2020 Residual Learning for Salient Object Detection
abstract
Recent deep learning based salient object detection methods improve the performance by introducing multi-scale strategies into fully convolutional neural networks (FCNs). The final result is obtained by integrating all the predictions at each scale. However, the existing multi-scale based methods suffer from several problems: 1) it is difficult to directly learn discriminative features and filters to regress high-resolution saliency masks for each scale; 2) rescaling the multi-scale features could pull in many redundant and inaccurate values, and this weakens the representational ability of the network. In this paper, we propose a residual learning strategy and introduce to gradually refine the coarse prediction scale-by-scale. Concretely, instead of directly predicting the finest-resolution result at each scale, we learn to predict residuals to remedy the errors between coarse saliency map and scale-matching ground truth masks. We employ a Dilated Convolutional Pyramid Pooling (DCPP) module to generate the coarse prediction and guide the the residual learning process through several novel Attentional Residual Modules (ARMs). We name our network as Residual Refinement Network (R2Net). We demonstrate the effectiveness of the proposed method against other state-of-the-art algorithms on five released benchmark datasets. Our R2Net is a fully convolutional network which does not need any post-processing and achieves a real-time speed of 33 FPS when it is run on one GPU.
Mengyang Feng, Huchuan Lu, Yizhou Yu
IEEE Trans. Image Process.2
2020 Saliency Detection via Depth-Induced Cellular Automata on Light Field
abstract
Incorrect saliency detection such as false alarms and missed alarms may lead to potentially severe consequences in various application areas. Effective separation of salient objects in complex scenes is a major challenge in saliency detection. In this paper, we propose a new method for saliency detection on light field to improve the saliency detection in challenging scenes. We construct an object-guided depth map, which acts as an inducer to efficiently incorporate the relations among light field cues, by using abundant light field cues. Furthermore, we enforce spatial consistency by constructing an optimization model, named Depth-induced Cellular Automata (DCA), in which the saliency value of each superpixel is updated by exploiting the intrinsic relevance of its similar regions. Additionally, the proposed DCA model enables inaccurate saliency maps to achieve a high level of accuracy. We analyze our approach on one publicly available dataset. Experiments show the proposed method is robust to a wide range of challenging scenes and outperforms the state-of-the-art 2D/3D/4D (light-field) saliency detection approaches.
Yongri Piao, Miao Zhang 0004, Jingyi Yu 0001, Huchuan Lu
IEEE Trans. Image Process.5
2020 LFNet: Light Field Fusion Network for Salient Object Detection
abstract
In this work, we propose a novel light field fusion network-LFNet, a CNNs-based light field saliency model using 4D light field data containing abundant spatial and contextual information. The proposed method can reliably locate and identify salient objects even in a complex scene. Our LFNet contains a light field refinement module (LFRM) and a light field integration module (LFIM) which can fully refine and integrate focusness, depths and objectness cues from light field image. The LFRM learns the light field residual between light field and RGB images for refining features with useful light field cues, and then the LFIM weights each refined light field feature and learns spatial correlation between them to predict saliency maps. Our method can take full advantage of light field information and achieve excellent performance especially in complex scenes, e.g., similar foreground and background, multiple or transparent objects and low-contrast environment. Experiments show our method outperforms the state-of-the-art 2D, 3D and 4D methods across three light field datasets.
Miao Zhang 0004, Wei Ji 0011, Yongri Piao, Yu Zhang 0165, Huchuan Lu
IEEE Trans. Image Process.7
2020 Deep Multiphase Level Set for Scene Parsing
abstract
Recently, Fully Convolutional Network (FCN) seems to be the go-to architecture for image segmentation, including semantic scene parsing. However, it is difficult for a generic FCN to predict semantic labels around the object boundaries, thus FCN-based methods usually produce parsing results with inaccurate boundaries. Meanwhile, many works have demonstrate that level set based active contours are superior to the boundary estimation in sub-pixel accuracy. However, they are quite sensitive to initial settings. To address these limitations, in this paper we propose a novel Deep Multiphase Level Set (DMLS) method for semantic scene parsing, which efficiently incorporates multiphase level sets into deep neural networks. The proposed method consists of three modules, i.e., recurrent FCNs, adaptive multiphase level set, and deeply supervised learning. More specifically, recurrent FCNs learn multi-level representations of input images with different contexts. Adaptive multiphase level set drives the discriminative contour for each semantic class, which makes use of the advantages of both global and local information. In each time-step of the recurrent FCNs, deeply supervised learning is incorporated for model training. Extensive experiments on three public benchmarks have shown that our proposed method achieves new state-of-the-art performances. The source codes will be released at https://github.com/Pchank/DMLS-for-SSP.
Wei Liu 0044, Yinjie Lei, Hongyu Wang 0001, Huchuan Lu
IEEE Trans. Image Process.5
2020 RAPNet: Residual Atrous Pyramid Network for Importance-Aware Street Scene Parsing
abstract
Street Scene Parsing (SSP) is a fundamental and important step for autonomous driving and traffic scene understanding. Recently, Fully Convolutional Network (FCN) based methods have delivered expressive performances with the help of large-scale dense-labeling datasets. However, in urban traffic environments, not all the labels contribute equally for making the control decision. Certain labels such as pedestrian, car, bicyclist, road lane or sidewalk would be more important in comparison with labels for vegetation, sky or building. Based on this fact, in this paper we propose a novel deep learning framework, named Residual Atrous Pyramid Network (RAPNet), for importance-aware SSP. More specifically, to incorporate the importance of various object classes, we propose an Importance-Aware Feature Selection (IAFS) mechanism which automatically selects the important features for label predictions. The IAFS can operate in each convolutional block, and the semantic features with different importance are captured in different channels so that they are automatically assigned with corresponding weights. To enhance the labeling coherence, we also propose a Residual Atrous Spatial Pyramid (RASP) module to sequentially aggregate global-to-local context information in a residual refinement manner. Extensive experiments on two public benchmarks have shown that our approach achieves new state-of-the-art performances, and can consistently obtain more accurate results on the semantic classes with high importance levels.
Wei Liu 0044, Yinjie Lei, Hongyu Wang 0001, Huchuan Lu
IEEE Trans. Image Process.5
2020 Visual Saliency Detection via Kernelized Subspace Ranking With Active Learning
abstract
Saliency detection task has witnessed a booming interest for years, due to the growth of the computer vision community. In this paper, we introduce a new saliency model that performs active learning with kernelized subspace ranker (KSR) referred to as KSR-AL. This pool-based active learning algorithm ranks the informativeness of unlabeled data by considering both uncertainty sampling and information density, thereby minimizing the cost of labeling. The informative images are selected to train the KSR iteratively and incrementally. The learning model of this algorithm is designed on object-level proposals and region-based convolutional neural network (R-CNN) features, by jointly learning a Rank-SVM classifier and a subspace projection. When the active learning process meets its stopping criteria, the saliency map of each image is generated by a weight fusion of its top-ranked proposals, whose ranking scores are graded by the learned ranker. We show that the KSR-AL achieves a reduction in annotation, as well as improvement in performance, compared with the supervised learning scheme. Besides, the proposed algorithm also outperforms the state-of-the-art methods. These improvements are demonstrated by extensive experiments on six publicly available benchmark datasets.
Lihe Zhang, Tiantian Wang 0002, Yifan Min, Huchuan Lu
IEEE Trans. Image Process.5
2020 A Multistage Refinement Network for Salient Object Detection
abstract
Deep convolutional neural networks (CNNs) have been successfully applied to a wide variety of problems in computer vision, including salient object detection. To accurately detect and segment salient objects, it is necessary to extract and combine high-level semantic features with low-level fine details simultaneously. This is challenging for CNNs because repeated subsampling operations such as pooling and convolution lead to a significant decrease in the feature resolution, which results in the loss of spatial details and finer structures. Therefore, we propose augmenting feedforward neural networks by using the multistage refinement mechanism. In the first stage, a master net is built to generate a coarse prediction map in which most detailed structures are missing. In the following stages, the refinement net with layerwise recurrent connections to the master net is equipped to progressively combine local context information across stages to refine the preceding saliency maps in a stagewise manner. Furthermore, the pyramid pooling module and channel attention module are applied to aggregate different-region-based global contexts. Extensive evaluations over six benchmark datasets show that the proposed method performs favorably against the state-of-the-art approaches.
Lihe Zhang, Jie Wu 0028, Tiantian Wang 0002, Ali Borji, Guohua Wei, Huchuan Lu
IEEE Trans. Image Process.6
2020 Towards Weakly-Supervised Focus Region Detection via Recurrent Constraint Network
abstract
Recent state-of-the-art methods on focus region detection (FRD) rely on deep convolutional networks trained with costly pixel-level annotations. In this study, we propose a FRD method that achieves competitive accuracies but only uses easily obtained bounding box annotations. Box-level tags provide important cues of focus regions but lose the boundary delineation of the transition area. A recurrent constraint network (RCN) is introduced for this challenge. In our static training, RCN is jointly trained with a fully convolutional network (FCN) through box-level supervision. The RCN can generate a detailed focus map to locate the boundary of the transition area effectively. In our dynamic training, we iterate between fine-tuning FCN and RCN with the generated pixel-level tags and generate finer new pixel-level tags. To boost the performance further, a guided conditional random field is developed to improve the quality of the generated pixel-level tags. To promote further study of the weakly supervised FRD methods, we construct a new dataset called FocusBox, which consists of 5000 challenging images with bounding box-level labels. Experimental results on existing datasets demonstrate that our method not only yields comparable results than fully supervised counterparts but also achieves a faster speed.
Wenda Zhao 0003, Xueqing Hou, You He 0002, Huchuan Lu
IEEE Trans. Image Process.5
2019 Deep Embedding Features for Salient Object Detection
abstract
Benefiting from the rapid development of Convolutional Neural Networks (CNNs), some salient object detection methods have achieved remarkable results by utilizing multi-level convolutional features. However, the saliency training datasets is of limited scale due to the high cost of pixel-level labeling, which leads to a limited generalization of the trained model on new scenarios during testing. Besides, some FCN-based methods directly integrate multi-level features, ignoring the fact that the noise in some features are harmful to saliency detection. In this paper, we propose a novel approach that transforms prior information into an embedding space to select attentive features and filter out outliers for salient object detection. Our network firstly generates a coarse prediction map through an encorder-decorder structure. Then a Feature Embedding Network (FEN) is trained to embed each pixel of the coarse map into a metric space, which incorporates much attentive features that highlight salient regions and suppress the response of non-salient regions. Further, the embedded features are refined through a deep-to-shallow Recursive Feature Integration Network (RFIN) to improve the details of prediction maps. Moreover, to alleviate the blurred boundaries, we propose a Guided Filter Refinement Network (GFRN) to jointly optimize the predicted results and the learnable guidance maps. Extensive experiments on five benchmark datasets demonstrate that our method outperforms state-of-the-art results. Our proposed method is end-to-end and achieves a realtime speed of 38 FPS.
Yunzhi Zhuge, Yu Zeng 0001, Huchuan Lu
AAAI3
2019 Visual Tracking via Adaptive Spatially-Regularized Correlation Filters
abstract
In this work, we propose a novel adaptive spatially-regularized correlation filters (ASRCF) model to simultaneously optimize the filter coefficients and the spatial regularization weight. First, this adaptive spatial regularization scheme could learn an effective spatial weight for a specific object and its appearance variations, and therefore result in more reliable filter coefficients during the tracking process. Second, our ASRCF model can be effectively optimized based on the alternating direction method of multipliers, where each subproblem has the closed-from solution. Third, our tracker applies two kinds of CF models to estimate the location and scale respectively. The location CF model exploits ensembles of shallow and deep features to determine the optimal position accurately. The scale CF model works on multi-scale shallow features to estimate the optimal scale efficiently. Extensive experiments on five recent benchmarks show that our tracker performs favorably against many state-of-the-art algorithms, with real-time performance of 28fps.
Kenan Dai, Dong Wang 0004, Huchuan Lu
CVPR3
2019 Attentive Feedback Network for Boundary-Aware Salient Object Detection
abstract
Recent deep learning based salient object detection methods achieve gratifying performance built upon Fully Convolutional Neural Networks (FCNs). However, most of them have suffered from the boundary challenge. The state-of-the-art methods employ feature aggregation tech- nique and can precisely find out wherein the salient object, but they often fail to segment out the entire object with fine boundaries, especially those raised narrow stripes. So there is still a large room for improvement over the FCN based models. In this paper, we design the Attentive Feedback Modules (AFMs) to better explore the structure of objects. A Boundary-Enhanced Loss (BEL) is further employed for learning exquisite boundaries. Our proposed deep model produces satisfying results on the object boundaries and achieves state-of-the-art performance on five widely tested salient object detection benchmarks. The network is in a fully convolutional fashion running at a speed of 26 FPS and does not need any post-processing.
Mengyang Feng, Huchuan Lu, Errui Ding
CVPR2
2019 ROI Pooled Correlation Filters for Visual Tracking
abstract
The ROI (region-of-interest) based pooling method performs pooling operations on the cropped ROI regions for various samples and has shown great success in the object detection methods. It compresses the model size while preserving the localization accuracy, thus it is useful in the visual tracking field. Though being effective, the ROI-based pooling operation is not yet considered in the correlation filter formula. In this paper, we propose a novel ROI pooled correlation filter (RPCF) algorithm for robust visual tracking. Through mathematical derivations, we show that the ROI-based pooling can be equivalently achieved by enforcing additional constraints on the learned filter weights, which makes the ROI-based pooling feasible on the virtual circular samples. Besides, we develop an efficient joint training formula for the proposed correlation filter algorithm, and derive the Fourier solvers for efficient model training. Finally, we evaluate our RPCF tracker on OTB-2013, OTB-2015 and VOT-2017 benchmark datasets. Experimental results show that our tracker performs favourably against other state-of-the-art trackers.
Yuxuan Sun 0003, Dong Wang 0004, You He 0002, Huchuan Lu
CVPR5
2019 A Mutual Learning Method for Salient Object Detection With Intertwined Multi-Supervision
abstract
Though deep learning techniques have made great progress in salient object detection recently, the predicted saliency maps still suffer from incomplete predictions due to the internal complexity of objects and inaccurate boundaries caused by strides in convolution and pooling operations. To alleviate these issues, we propose to train saliency detection networks by exploiting the supervision from not only salient object detection, but also foreground contour detection and edge detection. First, we leverage salient object detection and foreground contour detection tasks in an intertwined manner to generate saliency maps with uniform highlight. Second, the foreground contour and edge detection tasks guide each other simultaneously, thereby leading to preciser foreground contour prediction and reducing the local noises for edge prediction. In addition, we develop a novel mutual learning module (MLM) which serves as the building block of our method. Each MLM consists of multiple network branches trained in a mutual learning manner, which improves the performance by a large margin. Extensive experiments on seven challenging datasets demonstrate that the proposed method has delivered state-of-the-art results in both salient object detection and edge detection.
Runmin Wu, Mengyang Feng, Wenlong Guan, Dong Wang 0004, Huchuan Lu, Errui Ding
CVPR5
2019 Multi-Source Weak Supervision for Saliency Detection
abstract
The high cost of pixel-level annotations makes it appealing to train saliency detection models with weak supervision. However, a single weak supervision source usually does not contain enough information to train a well-performing model. To this end, we propose a unified framework to train saliency detection models with diverse weak supervision sources. In this paper, we use category labels, captions, and unlabelled data for training, yet other supervision sources can also be plugged into this flexible framework. We design a classification network (CNet) and a caption generation network (PNet), which learn to predict object categories and generate captions, respectively, meanwhile highlight the most important regions for corresponding tasks. An attention transfer loss is designed to transmit supervision signal between networks, such that the network designed to be trained with one supervision source can benefit from another. An attention coherence loss is defined on unlabelled data to encourage the networks to detect generally salient regions instead of task-specific regions. We use CNet and PNet to generate pixel-level pseudo labels to train a saliency prediction network (SNet). During the testing phases, we only need SNet to predict saliency maps. Experiments demonstrate the performance of our method compares favourably against unsupervised and weakly supervised methods and even some supervised methods.
Yu Zeng 0001, Yunzhi Zhuge, Huchuan Lu, Lihe Zhang, Mingyang Qian, Yizhou Yu
CVPR3
2019 CapSal: Leveraging Captioning to Boost Semantics for Salient Object Detection
abstract
Detecting salient objects in cluttered scenes is a big challenge. To address this problem, we argue that the model needs to learn discriminative semantic features for salient objects. To this end, we propose to leverage captioning as an auxiliary semantic task to boost salient object detection in complex scenarios. Specifically, we develop a CapSal model which consists of two sub-networks, the Image Captioning Network (ICN) and the Local-Global Perception Network (LGPN). ICN encodes the embedding of a generated caption to capture the semantic information of major objects in the scene, while LGPN incorporates the captioning embedding with local-global visual contexts for predicting the saliency map. ICN and LGPN are jointly trained to model high-level semantics as well as visual saliency. Extensive experiments demonstrate the effectiveness of image captioning in boosting the performance of salient object detection. In particular, our model performs significantly better than the state-of-the-art methods on several challenging datasets of complex scenarios.
Lu Zhang 0053, Jianming Zhang 0001, Zhe Lin 0001, Huchuan Lu, You He 0002
CVPR4
2019 Enhancing Diversity of Defocus Blur Detectors via Cross-Ensemble Network
abstract
Defocus blur detection (DBD) is a fundamental yet challenging topic, since the homogeneous region is obscure and the transition from the focused area to the unfocused region is gradual. Recent DBD methods make progress through exploring deeper or wider networks with the expense of high memory and computation. In this paper, we propose a novel learning strategy by breaking DBD problem into multiple smaller defocus blur detectors and thus estimate errors can cancel out each other. Our focus is the diversity enhancement via cross-ensemble network. Specifically, we design an end-to-end network composed of two logical parts: feature extractor network (FENet) and defocus blur detector cross-ensemble network (DBD-CENet). FENet is constructed to extract low-level features. Then the features are fed into DBD-CENet containing two parallel-branches for learning two groups of defocus blur detectors. For each individual, we design cross-negative and self-negative correlations and an error function to enhance ensemble diversity and balance individual accuracy. Finally, the multiple defocus blur detectors are combined with a uniformly weighted average to obtain the final DBD map. Experimental results indicate the superiority of our method in terms of accuracy and speed when compared with several state-of-the-art methods.
Wenda Zhao 0003, Qiuhua Lin, Huchuan Lu
CVPR4
2019 Language Person Search with Mutually Connected Classification Loss
abstract
In this work, we develop an effective person search algorithm with natural language descriptions. The contributions of this work mainly include two aspects. First, we design a baseline language person search framework including three basic components: a deep CNN model to extract visual features, a bi-directional LSTM to encode language descriptions and the triplet loss to conduct cross-modal feature embedding. Second, we propose a novel mutually connected classification loss to fully exploit the identity-level information, which not only introduces the identification information into both image and language descriptions but also encourages the cross-modal classification probabilities of the same identity to be more similar. The experimental results on the CUHK-PEDES dataset demonstrate that our method achieves significantly better performance than other state-of-the-art algorithms.
Chunjuan Bo, Dong Wang 0004, Yunwei Qi, Huchuan Lu
ICASSP6
2019 Online Single Person Tracking for Unmanned Aerial Vehicles: Benchmark and New Baseline
abstract
Online tracking a specific person from a low-altitude unmanned aerial vehicle (UAV) is a very interesting and challenging problem to be solved. However, there exists no large-scale aerial video dataset regarding this online single person tracking (OSPT) task. To promote the study of the OSPT problem in UAV, we first construct a new benchmark dataset including 100 fully annotated aerial videos with nearly 130K frames and 11 challenging factors. Second, we evaluate several state-of-the-art online trackers with real-time performance using our dataset, considering the potential applications in the UAV platform. In addition, with respect to the OSPT problem, we attempt to design a new baseline method with the combination of tracking, detection and re-identification and conduct detailed analysis of different components. This method achieves much better performance than the existing online trackers, which will serve as a new baseline for our benchmark.
Zhihui Wang 0001, Dong Wang 0004, Yunwei Qi, Huchuan Lu
ICASSP6
2019 Deep Learning for Light Field Saliency Detection
abstract
Recent research in 4D saliency detection is limited by the deficiency of a large-scale 4D light field dataset. To address this, we introduce a new dataset to assist the subsequent research in 4D light field saliency detection. To the best of our knowledge, this is to date the largest light field dataset in which the dataset provides 1465 all-focus images with human-labeled ground truth masks and the corresponding focal stacks for every light field image. To verify the effectiveness of the light field data, we first introduce a fusion framework which includes two CNN streams where the focal stacks and all-focus images serve as the input. The focal stack stream utilizes a recurrent attention mechanism to adaptively learn to integrate every slice in the focal stack, which benefits from the extracted features of the good slices. Then it is incorporated with the output map generated by the all-focus stream to make the saliency prediction. In addition, we introduce adversarial examples by adding noise intentionally into images to help train the deep network, which can improve the robustness of the proposed network. The noise is designed by users, which is imperceptible but can fool the CNNs to make the wrong prediction. Extensive experiments show the effectiveness and superiority of the proposed model on the popular evaluation metrics. The proposed method performs favorably compared with the existing 2D, 3D and 4D saliency detection methods on the proposed dataset and existing LFSD light field dataset. The code and results can be found at https://github.com/OIPLab-DUT/ ICCV2019_Deeplightfield_Saliency. Moreover, to facilitate research in this field, all images we collected are shared in a ready-to-use manner.
Tiantian Wang 0002, Yongri Piao, Huchuan Lu, Lihe Zhang
ICCV3
2019 GradNet: Gradient-Guided Network for Visual Object Tracking
abstract
The fully-convolutional siamese network based on template matching has shown great potentials in visual tracking. During testing, the template is fixed with the initial target feature and the performance totally relies on the general matching ability of the siamese network. However, this manner cannot capture the temporal variations of targets or background clutter. In this work, we propose a novel gradient-guided network to exploit the discriminative information in gradients and update the template in the siamese network through feed-forward and backward operations. To be specific, the algorithm can utilize the information from the gradient to update the template in the current frame. In addition, a template generalization training method is proposed to better use gradient information and avoid overfitting. To our knowledge, this work is the first attempt to exploit the information in the gradient for template update in siamese-based trackers. Extensive experiments on recent benchmarks demonstrate that our method achieves better performance than other state-of-the-art trackers.
Peixia Li, Wanli Ouyang, Dong Wang 0004, Xiaoyun Yang, Huchuan Lu
ICCV6
2019 Deep Reinforcement Active Learning for Human-in-the-Loop Person Re-Identification
abstract
Most existing person re-identification(Re-ID) approaches achieve superior results based on the assumption that a large amount of pre-labelled data is usually available and can be put into training phrase all at once. However, this assumption is not applicable to most real-world deployment of the Re-ID task. In this work, we propose an alternative reinforcement learning based human-in-the-loop model which releases the restriction of pre-labelling and keeps model upgrading with progressively collected data. The goal is to minimize human annotation efforts while maximizing Re-ID performance. It works in an iteratively updating framework by refining the RL policy and CNN parameters alternately. In particular, we formulate a Deep Reinforcement Active Learning (DRAL) method to guide an agent (a model in a reinforcement learning process) in selecting training samples on-the-fly by a human user/annotator. The reinforcement learning reward is the uncertainty value of each human selected sample. A binary feedback (positive or negative) labelled by the human annotator is used to select the samples of which are used to fine-tune a pre-trained CNN Re-ID model. Extensive experiments demonstrate the superiority of our DRAL method for deep reinforcement learning based human-in-the-loop person Re-ID when compared to existing unsupervised and transfer learning models as well as active learning models.
Zimo Liu, Jingya Wang 0001, Shaogang Gong, Dacheng Tao, Huchuan Lu
ICCV5
2019 Depth-Induced Multi-Scale Recurrent Attention Network for Saliency Detection
abstract
In this work, we propose a novel depth-induced multi-scale recurrent attention network for saliency detection. It achieves dramatic performance especially in complex scenarios. There are three main contributions of our network that are experimentally demonstrated to have significant practical merits. First, we design an effective depth refinement block using residual connections to fully extract and fuse multi-level paired complementary cues from RGB and depth streams. Second, depth cues with abundant spatial information are innovatively combined with multi-scale context features for accurately locating salient objects. Third, we boost our model's performance by a novel recurrent attention module inspired by Internal Generative Mechanism of human brain. This module can generate more accurate saliency results via comprehensively learning the internal semantic relation of the fused feature and progressively optimizing local details with memory-oriented scene understanding. In addition, we create a large scale RGB-D dataset containing more complex scenarios, which can contribute to comprehensively evaluating saliency models. Extensive experiments on six public datasets and ours demonstrate that our method can accurately identify salient objects and achieve consistently superior performance over 16 state-of-the-art RGB and RGB-D approaches.
Yongri Piao, Wei Ji 0011, Miao Zhang 0004, Huchuan Lu
ICCV5
2019 'Skimming-Perusal' Tracking: A Framework for Real-Time and Robust Long-Term Tracking
abstract
Compared with traditional short-term tracking, long-term tracking poses more challenges and is much closer to realistic applications. However, few works have been done and their performance have also been limited. In this work, we present a novel robust and real-time long-term tracking framework based on the proposed skimming and perusal modules. The perusal module consists of an effective bounding box regressor to generate a series of candidate proposals and a robust target verifier to infer the optimal candidate with its confidence score. Based on this score, our tracker determines whether the tracked object being present or absent, and then chooses the tracking strategies of local search or global search respectively in the next frame. To speed up the image-wide global search, a novel skimming module is designed to efficiently choose the most possible regions from a large number of sliding windows. Numerous experimental results on the VOT-2018 long-term and OxUvA long-term benchmarks demonstrate that the proposed method achieves the best performance and runs in real-time. The source codes are available at https://github.com/iiau-tracker/SPLT.
Bin Yan 0004, Haojie Zhao, Dong Wang 0004, Huchuan Lu, Xiaoyun Yang
ICCV4
2019 Joint Learning of Saliency Detection and Weakly Supervised Semantic Segmentation
abstract
Existing weakly supervised semantic segmentation (WSSS) methods usually utilize the results of pre-trained saliency detection (SD) models without explicitly modelling the connections between the two tasks, which is not the most efficient configuration. Here we propose a unified multi-task learning framework to jointly solve WSSS and SD using a single network, i.e. saliency and segmentation network (SSNet). SSNet consists of a segmentation network (SN) and a saliency aggregation module (SAM). For an input image, SN generates the segmentation result and, SAM predicts the saliency of each category and aggregating the segmentation masks of all categories into a saliency map. The proposed network is trained end-to-end with image-level category labels and class-agnostic pixel-level saliency labels. Experiments on PASCAL VOC 2012 segmentation dataset and four saliency benchmark datasets show the performance of our method compares favorably against state-of-the-art weakly supervised segmentation methods and fully supervised saliency detection methods.
Yu Zeng 0001, Yunzhi Zhuge, Huchuan Lu, Lihe Zhang
ICCV3
2019 Towards High-Resolution Salient Object Detection
abstract
Deep neural network based methods have made a significant breakthrough in salient object detection. However, they are typically limited to input images with low resolutions (400×400 pixels or less). Little effort has been made to train neural networks to directly handle salient object segmentation in high-resolution images. This paper pushes forward high-resolution saliency detection, and contributes a new dataset, named High-Resolution Salient Object Detection (HRSOD) dataset. To our best knowledge, HRSOD is the first high-resolution saliency detection dataset to date. As another contribution, we also propose a novel approach, which incorporates both global semantic information and local high-resolution details, to address this challenging task. More specifically, our approach consists of a Global Semantic Network (GSN), a Local Refinement Network (LRN) and a Global-Local Fusion Network (GLFN). The GSN extracts the global semantic information based on downsampled entire image. Guided by the results of GSN, the LRN focuses on some local regions and progressively produces high-resolution predictions. The GLFN is further proposed to enforce spatial consistency and boost performance. Experiments illustrate that our method outperforms existing state-of-the-art methods on high-resolution saliency datasets by a large margin, and achieves comparable or even better performance than them on some widely used saliency benchmarks.
Yi Zeng 0006, Zhe Lin 0001, Jianming Zhang 0001, Huchuan Lu
ICCV5
2019 Cascaded Context Pyramid for Full-Resolution 3D Semantic Scene Completion
abstract
Semantic Scene Completion (SSC) aims to simultaneously predict the volumetric occupancy and semantic category of a 3D scene. It helps intelligent devices to understand and interact with the surrounding scenes. Due to the high-memory requirement, current methods only produce low-resolution completion predictions, and generally lose the object details. Furthermore, they also ignore the multi-scale spatial contexts, which play a vital role for the 3D inference. To address these issues, in this work we propose a novel deep learning framework, named Cascaded Context Pyramid Network (CCPNet), to jointly infer the occupancy and semantic labels of a volumetric 3D scene from a single depth image. The proposed CCPNet improves the labeling coherence with a cascaded context pyramid. Meanwhile, based on the low-level features, it progressively restores the fine-structures of objects with Guided Residual Refinement (GRR) modules. Our proposed framework has three outstanding advantages: (1) it explicitly models the 3D spatial context for performance improvement; (2) full-resolution 3D volumes are produced with structure-preserving details; (3) light-weight models with low-memory requirements are captured with a good extensibility. Extensive experiments demonstrate that in spite of taking a single-view depth map, our proposed framework can generate high-quality SSC results, and outperforms state-of-the-art approaches on both the synthetic SUNCG and real NYU datasets.
Wei Liu 0044, Yinjie Lei, Huchuan Lu, Xiaoyun Yang
ICCV4
2019 Fast Video Object Segmentation via Dynamic Targeting Network
abstract
We propose a new model for fast and accurate video object segmentation. It consists of two convolutional neural networks, a Dynamic Targeting Network (DTN) and a Mask Refinement Network (MRN). DTN locates the object by dynamically focusing on regions of interest surrounding the target object. The target region is predicted by DTN via two sub-streams, Box Propagation (BP) and Box Re-identification (BR). The BP stream is faster but less effective at objects with large deformation or occlusion. The BR stream performs better in difficult scenarios at a higher computation cost. We propose a Decision Module (DM) to adaptively determine which sub-stream to use for each frame. Finally, MRN is exploited to predict segmentation within the target region. Experimental results on two public datasets demonstrate that the proposed model significantly outperforms existing methods without online training in both accuracy and efficiency, and is comparable to online training-based methods in accuracy with an order of magnitude faster speed.
Lu Zhang 0053, Zhe Lin 0001, Jianming Zhang 0001, Huchuan Lu, You He 0002
ICCV4
2019 Lightweight Deep Neural Network for Real-Time Visual Tracking with Mutual Learning
abstract
In this work, we develop a real-time tracking algorithm with a lightweight deep neural network. The contributions of this work mainly include two aspects. First, we reformulate the discriminative correlation filter (DCF) based tracker as a fully convolutional neural network and design an effective end-to-end tracking framework. Second, we build our tracker with a pruned convolutional neural network, which is trained by a mutual learning approach to further improve the location accuracy. The proposed tracking algorithm can track objects at 60 FPS. Extensive experiments on OTB2013, OTB2015 and VOT2017 Real-time demonstrate that the proposed tracker performs favorably against state-of-the-art methods.
Haojie Zhao, Gang Yang 0002, Dong Wang 0004, Huchuan Lu
ICIP4
2019 Deep Light-field-driven Saliency Detection from a Single View
abstract
Previous 2D saliency detection methods extract salient cues from a single view and directly predict the expected results. Both traditional and deep-learning-based 2D methods do not consider geometric information of 3D scenes. Therefore the relationship between scene understanding and salient objects cannot be effectively established. This limits the performance of 2D saliency detection in challenging scenes. In this paper, we show for the first time that saliency detection problem can be reformulated as two sub-problems: light field synthesis from a single view and light-field-driven saliency detection. We propose a high-quality light field synthesis network to produce reliable 4D light field information. Then we propose a novel light-field-driven saliency detection network with two purposes, that is, i) richer saliency features can be produced for effective saliency detection; ii) geometric information can be considered for integration of multi-view saliency maps in a view-wise attention fashion. The whole pipeline can be trained in an end-to-end fashion. For training our network, we introduce the largest light field dataset for saliency detection, containing 1580 light fields that cover a wide variety of challenging scenes. With this new formulation, our method is able to achieve state-of-the-art performance.
Yongri Piao, Zhengkun Rong, Miao Zhang 0004, Huchuan Lu
IJCAI5
2019 Memory-oriented Decoder for Light Field Salient Object Detection
abstract
Light field data have been demonstrated in favor of many tasks in computer vision, but existing works about light field saliency detection still rely on hand-crafted features. In this paper, we present a deep-learning-based method where a novel memory-oriented decoder is tailored for light field saliency detection. Our goal is to deeply explore and comprehensively exploit internal correlation of focal slices for accurate prediction by designing feature fusion and integration mechanisms. The success of our method is demonstrated by achieving the state of the art on three datasets. We present this problem in a way that is accessible to members of the community and provide a large-scale light field dataset that facilitates comparisons across algorithms. The code and dataset will be made publicly available.
Miao Zhang 0004, Ji Wei, Yongri Piao, Huchuan Lu
NeurIPS5
2019 Multi-scale Pyramid Pooling Network for salient object detection
Abdelhafid Dakhia, Tiantian Wang 0002, Huchuan Lu
Neurocomputing3
2019 A hybrid-backward refinement model for salient object detection
Abdelhafid Dakhia, Tiantian Wang 0002, Huchuan Lu
Neurocomputing3
2019 Salient Object Detection with Recurrent Fully Convolutional Networks
abstract
Deep networks have been proved to encode high-level features with semantic meaning and delivered superior performance in salient object detection. In this paper, we take one step further by developing a new saliency detection method based on recurrent fully convolutional networks (RFCNs). Compared with existing deep network based methods, the proposed network is able to incorpor- ate saliency prior knowledge for more accurate inference. In addition, the recurrent architecture enables our method to automatically learn to refine the saliency map by iteratively correcting its previous errors, yielding more reliable final predictions. To train such a netw- ork with numerous parameters, we propose a pre-training strategy using semantic segmentation data, which simultaneously leverages the strong supervision of segmentation tasks for effective training and enables the network to capture generic representations to chara- cterize category-agnostic objects for saliency detection. Extensive experimental evaluations demonstrate that the proposed method compares favorably against state-of-the-art saliency detection approaches. Additional validations are also performed to study the impact of the recurrent architecture and pre-training strategy on both saliency detection and semantic segmentation, which provides important knowledge for network design and training in the future research.
Linzhao Wang, Lijun Wang 0001, Huchuan Lu, Xiang Ruan
IEEE Trans. Pattern Anal. Mach. Intell.3
2019 Multi attention module for visual tracking
Peixia Li, Dong Wang 0004, Gang Yang 0002, Huchuan Lu
Pattern Recognit.6
2019 Language-aware weak supervision for salient object detection
Mingyang Qian, Jinqing Qi, Lihe Zhang, Mengyang Feng, Huchuan Lu
Pattern Recognit.5
2019 Hyperfusion-Net: Hyper-densely reflective feature fusion for salient object detection
Wei Liu 0044, Yinjie Lei, Huchuan Lu
Pattern Recognit.4
2019 Deep gated attention networks for large-scale street-level scene segmentation
Wei Liu 0044, Hongyu Wang 0001, Yinjie Lei, Huchuan Lu
Pattern Recognit.5
2019 Edge-Aware Convolution Neural Network Based Salient Object Detection
abstract
Salient object detection has received great amount of attention in recent years. In this letter, we propose a novel salient object detection algorithm, which combines the global contextual information along with the low-level edge features. First, we train an edge detection stream based on the state-of-the-art holistically-nested edge detection (HED) model and extract hierarchical boundary information from each VGG block. Then, the edge contours are served as the complementary edge-aware information and integrated with the saliency detection stream to depict continuous boundary for salient objects. Finally, we combine pyramid pooling modules with auxiliary side output supervision to form the multi-scale pyramid-based supervision module, providing multi-scale global contextual information for the saliency detection network. Compared with the previous methods, the proposed network contains more explicit edge-aware features and exploit the multi-scale global information more effectively. Experiments demonstrate the effectiveness of the proposed method, which achieves the state-of-the-art performance on five popular benchmarks.
Wenlong Guan, Tiantian Wang 0002, Jinqing Qi, Lihe Zhang, Huchuan Lu
IEEE Signal Process. Lett.5
2019 Subspace Clustering Under Complex Noise
abstract
In this paper, we study the subspace clustering problem under complex noise. A wide class of reconstruction-based methods model the subspace clustering problem by combining a quadratic data-fidelity term and a regularization term. In a statistical framework, the data-fidelity term assumes to be contaminated by a unimodal Gaussian noise, which is a popular setting in most current subspace clustering models. However, the realistic noise is much more complex than our assumptions. Besides, the coarse representation of the data-fidelity term may depress the clustering accuracy, which is often used to evaluate the models. To address this issue, we propose the mixture of Gaussian regression (MoG Regression) for subspace clustering. The MoG Regression seeks a valid way to model the unknown noise distribution, which approaches the real one as far as possible, so that the desired affinity matrix is better at characterizing the structure of data in the real world, and furthermore, improving the performance. Theoretically, the proposed model enjoys the grouping effect, which encourages the coefficients of highly correlated points are nearly equal. Drawing upon the ideal of the minimum message length, a model selection strategy is proposed to estimate the numbers of the Gaussian components that shows a way how to seek the number of Gaussian components besides determining it by empirical value. In addition, the asymptotic property of our model is investigated. The proposed model is evaluated on the challenging datasets. The experimental results show that the proposed MoG Regression model significantly outperforms several state-of-the-art subspace clustering methods.
Baohua Li, Huchuan Lu, Ying Zhang 0021, Zhouchen Lin, Wei Wu 0010
IEEE Trans. Circuits Syst. Video Technol.2
2019 Multi-Focus Image Fusion With a Natural Enhancement via a Joint Multi-Level Deeply Supervised Convolutional Neural Network
abstract
Common non-focused areas are often present in multi-focus images due to the limitation of the number of focused images. This factor severely degrades the fusion quality of multi-focus images. To address this problem, we propose a novel end-to-end multi-focus image fusion with a natural enhancement method based on deep convolutional neural network (CNN). Several end-to-end CNN architectures that are specifically adapted to this task are first designed and researched. On the basis of the observation that low-level feature extraction can capture low-frequency content, whereas high-level feature extraction effectively captures high-frequency details, we further combine multi-level outputs such that the most visually distinctive features can be extracted, fused, and enhanced. In addition, the multi-level outputs are simultaneously supervised during training to boost the performance of image fusion and enhancement. Extensive experiments show that the proposed method can deliver superior fusion and enhancement performance than the state-of-the-art methods in the presence of multi-focus images with common non-focused areas, anisotropic blur, and misregistration.
Wenda Zhao 0003, Dong Wang 0004, Huchuan Lu
IEEE Trans. Circuits Syst. Video Technol.3
2019 Video Person Re-Identification by Temporal Residual Learning
abstract
In this paper, we propose a novel feature learning framework for video person re-identification (re-ID). The proposed framework largely aims to exploit the adequate temporal information of video sequences and tackle the poor spatial alignment of moving pedestrians. More specifically, for exploiting the temporal information, we design a temporal residual learning (TRL) module to simultaneously extract the generic and specific features of consecutive frames. The TRL module is equipped with two bi-directional LSTM (BiLSTM), which are respectively responsible to describe a moving person in different aspects, providing complementary information for better feature representations. To deal with the poor spatial alignment in video re- ID datasets, we propose a spatial-temporal transformer network (ST2N) module. Transformation parameters in the ST2N module are learned by leveraging the high-level semantic information of the current frame as well as the temporal context knowledge from other frames. The proposed ST2N module with less learnable parameters allows effective person alignments under significant appearance changes. Extensive experimental results on the largescale MARS, PRID2011, ILIDS-VID and SDU-VID datasets demonstrate that the proposed method achieves consistently superior performance and outperforms most of the very recent state-of-the-art methods.
Ju Dai, Dong Wang 0004, Huchuan Lu, Hongyu Wang 0001
IEEE Trans. Image Process.4
2019 Tensor Completion From One-Bit Observations
abstract
The tensor completion issues have obtained a great deal of attention in the past few years. However, the data fidelity part minimizes a squared loss function, which may be inappropriate for the case of noisy one-bit observations. In this paper, we alleviate the mentioned difficulty by drawing on the experience of matrix scenarios. Based on the convex relation to $\ell _{1}$ norm of the tensor multi-rank, we propose a novel optimization model trying to recover the underlying tensor in case of one-bit observations. The feasibility of this model is proved by theoretical derivations. Furthermore, an alternating direction method of multipliers based algorithm is designed to find the solution. The numerical experiments demonstrate the effectiveness of our method.
Baohua Li, Xiaoli Li 0011, Huchuan Lu
IEEE Trans. Image Process.4
2019 Salient Object Detection With Lossless Feature Reflection and Weighted Structural Loss
abstract
Salient object detection (SOD), which aims to identify and locate the most salient pixels or regions in images, has been attracting more and more interest due to its various realworld applications. However, this vision task is quite challenging, especially under complex image scenes. Inspired by the intrinsic reflection of natural images, in this paper we propose a novel feature learning framework for large-scale salient object detection. Specifically, we design a symmetrical fully convolutional network (SFCN) to effectively learn complementary saliency features under the guidance of lossless feature reflection. The location information, together with contextual and semantic information, of salient objects are jointly utilized to supervise the proposed network for more accurate saliency predictions. In addition, to overcome the blurry boundary problem, we propose a new weighted structural loss function to ensure clear object boundaries and spatially consistent saliency. The coarse prediction results are effectively refined by these structural information for performance improvements. Extensive experiments on seven saliency detection datasets demonstrate that our approach achieves consistently superior performance and outperforms the very recent state-of-the-art methods with a large margin.
Wei Liu 0044, Huchuan Lu, Chunhua Shen
IEEE Trans. Image Process.3
2019 Pose-Invariant Embedding for Deep Person Re-Identification
abstract
Pedestrian misalignment, which mainly arises from detector errors and pose variations, is a critical problem for a robust person re-identification (re-ID) system. With poor alignment, the feature learning and matching process might be largely compromised. To address this problem, this paper introduces the pose invariant embedding (PIE) as a pedestrian descriptor. First, in order to align pedestrians to a standard pose, the PoseBox structure is introduced, which is generated through pose estimation followed by affine transformations. Second, to reduce the impact of pose estimation errors and information loss during PoseBox construction, we design a PoseBox fusion (PBF) CNN architecture that takes the original image, the PoseBox, and the pose estimation confidence as input. The proposed PIE descriptor is thus defined as the fully connected layer of the PBF network for the retrieval task. Experiments are conducted on the Market-1501, CUHK03-NP, and DukeMTMC-reID datasets. We show that PoseBox alone yields decent re-ID accuracy, and that when integrated in the PBF network, the learned PIE descriptor produces competitive performance compared with the state-of-the-art approaches.
Liang Zheng 0001, Yujia Huang, Huchuan Lu, Yi Yang 0001
IEEE Trans. Image Process.3
2019 Person Reidentification by Joint Local Distance Metric and Feature Transformation
abstract
Person reidentification is of great importance in visual surveillance and multiperson tracking across multiple camera views. Two fundamental problems are critical for person reidentification: 1) how to account for appearance variation or feature transformation caused by viewpoint changes and 2) how to learn a discriminative distance metric for reidentification. In this paper, we propose an algorithm in which both feature transformation and metric learning are exploited and jointly optimized. We learn local models from subsets of training samples with regularization imposed by the global model which is trained among the entire data set. The learned local models enhance the discriminative strength and generalization ability. Experimental results on the Viewpoint Invariant PEdestrian Eecognition, Queen Mary University of London ground reidentification, CUHK01, and CUHK03 benchmark data sets show that the proposed sample-specific view-invariant approach performs favorably against the state-of-the-art person reidentification methods.
Zimo Liu, Huchuan Lu, Xiang Ruan, Ming-Hsuan Yang 0001
IEEE Trans. Neural Networks Learn. Syst.2
2018 Learning Dual Convolutional Neural Networks for Low-Level Vision
abstract
In this paper, we propose a general dual convolutional neural network (DualCNN) for low-level vision problems, e.g., super-resolution, edge-preserving filtering, deraining and dehazing. These problems usually involve the estimation of two components of the target signals: structures and details. Motivated by this, our proposed DualCNN consists of two parallel branches, which respectively recovers the structures and details in an end-to-end manner. The recovered structures and details can generate the target signals according to the formation model for each particular application. The DualCNN is a flexible framework for low-level vision tasks and can be easily incorporated into existing CNNs. Experimental results show that the DualCNN can be effectively applied to numerous low-level vision tasks with favorable performance against the state-of-the-art methods.
Jinshan Pan, Sifei Liu, Deqing Sun, Jiawei Zhang 0002, Yang Liu 0119, Jimmy S. J. Ren, Zechao Li, Jinhui Tang 0001, Huchuan Lu, Yu-Wing Tai, Ming-Hsuan Yang 0001
CVPR9
2018 Correlation Tracking via Joint Discrimination and Reliability Learning
abstract
For visual tracking, an ideal filter learned by the correlation filter (CF) method should take both discrimination and reliability information. However, existing attempts usually focus on the former one while pay less attention to reliability learning. This may make the learned filter be dominated by the unexpected salient regions on the feature map, thereby resulting in model degradation. To address this issue, we propose a novel CF-based optimization problem to jointly model the discrimination and reliability information. First, we treat the filter as the element-wise product of a base filter and a reliability term. The base filter is aimed to learn the discrimination information between the target and backgrounds, and the reliability term encourages the final filter to focus on more reliable regions. Second, we introduce a local response consistency regular term to emphasize equal contributions of different regions and avoid the tracker being dominated by unreliable regions. The proposed optimization problem can be solved using the alternating direction method and speeded up in the Fourier domain. We conduct extensive experiments on the OTB-2013, OTB-2015 and VOT-2016 datasets to evaluate the proposed tracker. Experimental results show that our tracker performs favorably against other state-of-the-art trackers.
Dong Wang 0004, Huchuan Lu, Ming-Hsuan Yang 0001
CVPR3
2018 Learning Spatial-Aware Regressions for Visual Tracking
abstract
In this paper, we analyze the spatial information of deep features, and propose two complementary regressions for robust visual tracking. First, we propose a kernelized ridge regression model wherein the kernel value is defined as the weighted sum of similarity scores of all pairs of patches between two samples. We show that this model can be formulated as a neural network and thus can be efficiently solved. Second, we propose a fully convolutional neural network with spatially regularized kernels, through which the filter kernel corresponding to each output channel is forced to focus on a specific region of the target. Distance transform pooling is further exploited to determine the effectiveness of each output channel of the convolution layer. The outputs from the kernelized ridge regression model and the fully convolutional neural network are combined to obtain the ultimate response. Experimental results on two benchmark datasets validate the effectiveness of the proposed method.
Dong Wang 0004, Huchuan Lu, Ming-Hsuan Yang 0001
CVPR3
2018 Detect Globally, Refine Locally: A Novel Approach to Saliency Detection
abstract
Effective integration of contextual information is crucial for salient object detection. To achieve this, most existing methods based on 'skip' architecture mainly focus on how to integrate hierarchical features of Convolutional Neural Networks (CNNs). They simply apply concatenation or element-wise operation to incorporate high-level semantic cues and low-level detailed information. However, this can degrade the quality of predictions because cluttered and noisy information can also be passed through. To address this problem, we proposes a global Recurrent Localization Network (RLN) which exploits contextual information by the weighted response map in order to localize salient objects more accurately. Particularly, a recurrent module is employed to progressively refine the inner structure of the CNN over multiple time steps. Moreover, to effectively recover object boundaries, we propose a local Boundary Refinement Network (BRN) to adaptively learn the local contextual information for each spatial position. The learned propagation coefficients can be used to optimally capture relations between each pixel and its neighbors. Experiments on five challenging datasets show that our approach performs favorably against all existing methods in terms of the popular evaluation metrics.
Tiantian Wang 0002, Lihe Zhang, Huchuan Lu, Gang Yang 0002, Xiang Ruan, Ali Borji
CVPR4
2018 Learning to Promote Saliency Detectors
abstract
The categories and appearance of salient objects vary from image to image, therefore, saliency detection is an image-specific task. Due to lack of large-scale saliency training data, using deep neural networks (DNNs) with pretraining is difficult to precisely capture the image-specific saliency cues. To solve this issue, we formulate a zero-shot learning problem to promote existing saliency detectors. Concretely, a DNN is trained as an embedding function to map pixels and the attributes of the salient/background regions of an image into the same metric space, in which an image-specific classifier is learned to classify the pixels. Since the image-specific task is performed by the classifier, the DNN embedding effectively plays the role of a general feature extractor. Compared with transferring the learning to a new recognition task using limited data, this formulation makes the DNN learn more effectively from small data. Extensive experiments on five data sets show that our method significantly improves accuracy of existing methods and compares favorably against state-of-the-art approaches.
Yu Zeng 0001, Huchuan Lu, Lihe Zhang, Mengyang Feng, Ali Borji
CVPR2
2018 A Bi-Directional Message Passing Model for Salient Object Detection
abstract
Recent progress on salient object detection is beneficial from Fully Convolutional Neural Network (FCN). The saliency cues contained in multi-level convolutional features are complementary for detecting salient objects. How to integrate multi-level features becomes an open problem in saliency detection. In this paper, we propose a novel bi-directional message passing model to integrate multi-level features for salient object detection. At first, we adopt a Multi-scale Context-aware Feature Extraction Module (MCFEM) for multi-level feature maps to capture rich context information. Then a bi-directional structure is designed to pass messages between multi-level features, and a gate function is exploited to control the message passing rate. We use the features after message passing, which simultaneously encode semantic information and spatial details, to predict saliency maps. Finally, the predicted results are efficiently combined to generate the final saliency map. Quantitative and qualitative experiments on five benchmark datasets demonstrate that our proposed model performs favorably against the state-of-the-art methods under different evaluation metrics.
Lu Zhang 0053, Ju Dai, Huchuan Lu, You He 0002, Gang Wang 0012
CVPR3
2018 Progressive Attention Guided Recurrent Network for Salient Object Detection
abstract
Effective convolutional features play an important role in saliency estimation but how to learn powerful features for saliency is still a challenging task. FCN-based methods directly apply multi-level convolutional features without distinction, which leads to sub-optimal results due to the distraction from redundant details. In this paper, we propose a novel attention guided network which selectively integrates multi-level contextual information in a progressive manner. Attentive features generated by our network can alleviate distraction of background thus achieve better performance. On the other hand, it is observed that most of existing algorithms conduct salient object detection by exploiting side-output features of the backbone feature extraction network. However, shallower layers of backbone network lack the ability to obtain global semantic information, which limits the effective feature learning. To address the problem, we introduce multi-path recurrent feedback to enhance our proposed progressive attention driven framework. Through multi-path recurrent connections, global semantic information from the top convolutional layer is transferred to shallower layers, which intrinsically refines the entire network. Experimental results on six benchmark datasets demonstrate that our algorithm performs favorably against the state-of-the-art approaches.
Tiantian Wang 0002, Jinqing Qi, Huchuan Lu, Gang Wang 0012
CVPR4
2018 Deep Mutual Learning
abstract
Model distillation is an effective and widely used technique to transfer knowledge from a teacher to a student network. The typical application is to transfer from a powerful large network or ensemble to a small network, in order to meet the low-memory or fast execution requirements. In this paper, we present a deep mutual learning (DML) strategy. Different from the one-way transfer between a static pre-defined teacher and a student in model distillation, with DML, an ensemble of students learn collaboratively and teach each other throughout the training process. Our experiments show that a variety of network architectures benefit from mutual learning and achieve compelling results on both category and instance recognition tasks. Surprisingly, it is revealed that no prior powerful teacher network is necessary - mutual learning of a collection of simple student networks works, and moreover outperforms distillation from a more powerful yet static teacher.
Ying Zhang 0021, Tao Xiang 0002, Timothy M. Hospedales, Huchuan Lu
CVPR4
2018 Defocus Blur Detection via Multi-Stream Bottom-Top-Bottom Fully Convolutional Network
abstract
Defocus blur detection (DBD) is the separation of in-focus and out-of-focus regions in an image. This process has been paid considerable attention because of its remarkable potential applications. Accurate differentiation of homogeneous regions and detection of low-contrast focal regions, as well as suppression of background clutter, are challenges associated with DBD. To address these issues, we propose a multi-stream bottom-top-bottom fully convolutional network (BTBNet), which is the first attempt to develop an end-to-end deep network for DBD. First, we develop a fully convolutional BTBNet to integrate low-level cues and high-level semantic information. Then, considering that the degree of defocus blur is sensitive to scales, we propose multi-stream BTBNets that handle input images with different scales to improve the performance of DBD. Finally, we design a fusion and recurrent reconstruction network to recurrently refine the preceding blur detection maps. To promote further study and evaluation of the DBD models, we construct a new database of 500 challenging images and their pixel-wise defocus blur annotations. Experimental results on the existing and our new datasets demonstrate that the proposed method achieves significantly better performance than other state-of-the-art algorithms.
Wenda Zhao 0003, Fan Zhao 0006, Dong Wang 0004, Huchuan Lu
CVPR4
2018 Real-Time 'Actor-Critic' Tracking
Dong Wang 0004, Peixia Li, Huchuan Lu
ECCV (7)5
2018 Deep Cross-Modal Projection Learning for Image-Text Matching
Ying Zhang 0021, Huchuan Lu
ECCV (1)2
2018 Structured Siamese Network for Real-Time Visual Tracking
Lijun Wang 0001, Jinqing Qi, Dong Wang 0004, Mengyang Feng, Huchuan Lu
ECCV (9)6
2018 Salient Object Detection by Lossless Feature Reflection
abstract
Salient object detection, which aims to identify and locate the most salient pixels or regions in images, has been attracting more and more interest due to its various real-world applications. However, this vision task is quite challenging, especially under complex image scenes. Inspired by the intrinsic reflection of natural images, in this paper we propose a novel feature learning framework for large-scale salient object detection. Specifically, we design a symmetrical fully convolutional network (SFCN) to learn complementary saliency features under the guidance of lossless feature reflection. The location information, together with contextual and semantic information, of salient objects are jointly utilized to supervise the proposed network for more accurate saliency predictions. In addition, to overcome the blurry boundary problem, we propose a new structural loss function to learn clear object boundaries and spatially consistent saliency. The coarse prediction results are effectively refined by these structural information for performance improvements. Extensive experiments on seven saliency detection datasets demonstrate that our approach achieves consistently superior performance and outperforms the very recent state-of-the-art methods.
Wei Liu 0044, Huchuan Lu, Chunhua Shen
IJCAI3
2018 Hierarchical Cellular Automata for Visual Saliency
Yao Qin 0001, Mengyang Feng, Huchuan Lu, Garrison W. Cottrell
Int. J. Comput. Vis.3
2018 Human body segmentation in static images by models with shape as guidance
Shifeng Li, Huchuan Lu
Neurocomputing3
2018 Deep multi-level networks with multi-task learning for saliency detection
Lihe Zhang, Hongguang Bo, Tiantian Wang 0002, Huchuan Lu
Neurocomputing5
2018 Spectral-spatial K-Nearest Neighbor approach for hyperspectral image classification
Chunjuan Bo, Huchuan Lu, Dong Wang 0004
Multim. Tools Appl.2
2018 Cross-view semantic projection learning for person re-identification
Ju Dai, Ying Zhang 0021, Huchuan Lu, Hongyu Wang 0001
Pattern Recognit.3
2018 Deep visual tracking: Review and experimental comparison
Peixia Li, Dong Wang 0004, Lijun Wang 0001, Huchuan Lu
Pattern Recognit.4
2018 Predicting human gaze with multi-level information
Jinqing Qi, Huchuan Lu
Signal Process.5
2018 Boundary-Guided Feature Aggregation Network for Salient Object Detection
abstract
Fully convolutional networks (FCN) has significantly improved the performance of many pixel-labeling tasks, such as semantic segmentation and depth estimation. However, it still remains nontrivial to thoroughly utilize the multilevel convolutional feature maps and boundary information for salient object detection. In this letter, we propose a novel FCN framework to integrate multilevel convolutional features recurrently with the guidance of object boundary information. First, a deep convolutional network is used to extract multilevel feature maps and separately aggregate them into multiple resolutions, which can be used to generate coarse saliency maps. Meanwhile, another boundary information extraction branch is proposed to generate boundary features. Finally, an attention-based feature fusion module is designed to fuse boundary information into salient regions to achieve accurate boundary inference and semantic enhancement. The final saliency maps are the combination of the predicted boundary maps and integrated saliency maps, which are more closer to the ground truths. Experiments and analysis on four large-scale benchmarks verify that our framework achieves new state-of-the-art results.
Yunzhi Zhuge, Gang Yang 0002, Huchuan Lu
IEEE Signal Process. Lett.4
2018 Subspace Clustering With K-Support Norm
abstract
Subspace clustering aims to cluster a collection of data points lying in a union of subspaces. Based on the assumption that each point can be approximately represented as a linear combination of other points, extensive efforts have been made to compute an affinity matrix in a self-expressive framework for describing the similarity between points. However, the existing clustering methods consider the average feature solutions, which would not be powerful enough to capture the intrinsic relationship between points. In this paper, we present the k-support norm subspace clustering (KSC) method by utilizing k-support norm regularization. The k-support norm trades off the sparsity of ℓ1norm and the uniform shrinkage of ℓ2norm to yield better predictive performance on the data connection. The theoretical analysis of KSC makes up a large proportion of paper. In the noise-free case, we provide the kEBD condition, which ensures the coefficient matrix to be block diagonal. If the data are corrupted, we prove the incompletion-grouping effect for KSC. Moreover, we provide the statistical recovery guarantee for both noise-free and noise cases. The theory analyses show the validity and feasibility of KSC, and the experimental results on multiple challenging databases demonstrate the effectiveness of the proposed algorithm.
Baohua Li, Huchuan Lu, Fu Li 0003, Wei Wu 0010
IEEE Trans. Circuits Syst. Video Technol.2
2018 Tracking With Static and Dynamic Structured Correlation Filters
abstract
Tracking methods based on correlation filters have recently attracted attention for achieving fast tracking. However, their performance is somewhat limited in long-term tracking tasks, especially in an occlusion situation. To address this issue, we propose a novel structured correlation filter, which depends on coupled interactions between a static model and a dynamic model. Specifically, the static model exploits the star graph to capture spatial information and provides an initial estimation. The dynamic model based on Bayesian inference uses the rough location as a reference to estimate the final target state. Then, the dynamic model provides a feedback to the static regarding their updates. Finally, the dynamic model provides a scale adaptivity mechanism, which makes the proposed tracker effectively deal with not only partial occlusion but also scale variation. Both qualitative and quantitative evaluations on challenging image sequences demonstrate that the proposed method performs favorably against the state-of-the-art tracking algorithms.
Dong Wang 0004, Huchuan Lu
IEEE Trans. Circuits Syst. Video Technol.3
2018 Constrained Superpixel Tracking
abstract
In this paper, we propose a constrained graph labeling algorithm for visual tracking where nodes denote superpixels and edges encode the underlying spatial, temporal, and appearance fitness constraints. First, the spatial smoothness constraint, based on a transductive learning method, is enforced to leverage the latent manifold structure in feature space by investigating unlabeled superpixels in the current frame. Second, the appearance fitness constraint, which measures the probability of a superpixel being contained in the target region, is developed to incrementally induce a long-term appearance model. Third, the temporal smoothness constraint is proposed to construct a short-term appearance model of the target, which handles the drastic appearance change between consecutive frames. All these three constraints are incorporated in the proposed graph labeling algorithm such that induction and transduction, short- and long-term appearance models are combined, respectively. The foreground regions inferred by the proposed graph labeling method are used to guide the tracking process. Tracking results, in turn, facilitate more accurate online update by filtering out potential contaminated training samples. Both quantitative and qualitative evaluations on challenging tracking data sets show that the proposed constrained tracking algorithm performs favorably against the state-of-the-art methods.
Lijun Wang 0001, Huchuan Lu, Ming-Hsuan Yang 0001
IEEE Trans. Cybern.2
2018 Exemplar-Aided Salient Object Detection via Joint Latent Space Embedding
abstract
Traditional unsupervised salient object detection methods majorly rely on pre-defined assumptions about saliency. However, these assumptions may not be sufficient for handling test images of varied content and context. Meanwhile, supervised models learn saliency knowledge from thousands of annotated images, which are usually expensive to obtain. In this paper, we propose an exemplar-aided salient object detection method, which can complement heuristic saliency assumptions by leveraging only a few exemplar images. This is a challenging task since the appearances between the query images and the exemplars can be quite different. We handle it by learning the matching relationship of the intra-class instances in a latent embedding space in an online fashion. Given a test image and an annotated reference image (retrieved from several exemplar images), our method transfers the foreground and background information of the reference image to the test image via a joint latent embedding of image superpixels. Extensive experiments show that our method can easily improve the performance of existing unsupervised methods even when a very small reference image dataset (e.g. one image) is used. In addition, our method is able to attain competitive performance against fully supervised methods.
Yuqiu Kong, Jianming Zhang 0001, Huchuan Lu, Xiuping Liu
IEEE Trans. Image Process.3
2018 An Unsupervised Game-Theoretic Approach to Saliency Detection
abstract
We propose a novel unsupervised game-theoretic salient object detection algorithm that does not require labeled training data. First, saliency detection problem is formulated as a non-cooperative game, hereinafter referred to as Saliency Game, in which image regions are players who choose to be "background" or "foreground" as their pure strategies. A payoff function is constructed by exploiting multiple cues and combining complementary features. Saliency maps are generated according to each region's strategy in the Nash equilibrium of the proposed Saliency Game. Second, we explore the complementary relationship between color and deep features and propose an Iterative Random Walk algorithm to combine saliency maps produced by the Saliency Game using different features. Iterative random walk allows sharing information across feature spaces, and detecting objects that are otherwise very hard to detect. Extensive experiments over 6 challenging datasets demonstrate the superiority of our proposed unsupervised algorithm compared to several state of the art supervised algorithms.
Yu Zeng 0001, Mengyang Feng, Huchuan Lu, Gang Yang 0002, Ali Borji
IEEE Trans. Image Process.3
2018 Saliency Detection via Absorbing Markov Chain With Learnt Transition Probability
abstract
In this paper, we propose a bottom-up saliency model based on absorbing Markov chain (AMC). First, a sparsely connected graph is constructed to capture the local context information of each node. All image boundary nodes and other nodes are, respectively, treated as the absorbing nodes and transient nodes in the absorbing Markov chain. Then, the expected number of times from each transient node to all other transient nodes can be used to represent the saliency value of this node. The absorbed time depends on the weights on the path and their spatial coordinates, which are completely encoded in the transition probability matrix. Considering the importance of this matrix, we adopt different hierarchies of deep features extracted from fully convolutional networks and learn a transition probability matrix, which is called learnt transition probability matrix. Although the performance is significantly promoted, salient objects are not uniformly highlighted very well. To solve this problem, an angular embedding technique is investigated to refine the saliency results. Based on pairwise local orderings, which are produced by the saliency maps of AMC and boundary maps, we rearrange the global orderings (saliency value) of all nodes. Extensive experiments demonstrate that the proposed algorithm outperforms the state-of-the-art methods on six publicly available benchmark data sets.
Lihe Zhang, Jianwu Ai, Huchuan Lu, Xiukui Li
IEEE Trans. Image Process.4
2018 DeepLens: shallow depth of field from a single image
abstract
We aim to generate high resolution shallow depth-of-field (DoF) images from a single all-in-focus image with controllable focal distance and aperture size. To achieve this, we propose a novel neural network model comprised of a depth prediction module, a lens blur module, and a guided upsampling module. All modules are differentiable and are learned from data. To train our depth prediction module, we collect a dataset of 2462 RGB-D images captured by mobile phones with a dual-lens camera, and use existing segmentation datasets to improve border prediction. We further leverage a synthetic dataset with known depth to supervise the lens blur and guided upsampling modules. The effectiveness of our system and training strategies are verified in the experiments. Our method can generate high-quality shallow DoF images at high resolution, and produces significantly fewer artifacts than the baselines and existing solutions for single image shallow DoF synthesis. Compared with the iPhone portrait mode, which is a state-of-the-art shallow DoF solution based on a dual-lens depth camera, our method generates comparable results, while allowing for greater flexibility to choose focal points and aperture size, and is not limited to one capture setup.
Lijun Wang 0001, Xiaohui Shen, Jianming Zhang 0001, Oliver Wang, Zhe Lin 0001, Chih-Yao Hsieh, Sarah Kong, Huchuan Lu
ACM Trans. Graph.8
2017 Learning to Detect Salient Objects with Image-Level Supervision
abstract
Deep Neural Networks (DNNs) have substantially improved the state-of-the-art in salient object detection. However, training DNNs requires costly pixel-level annotations. In this paper, we leverage the observation that image-level tags provide important cues of foreground salient objects, and develop a weakly supervised learning method for saliency detection using image-level tags only. The Foreground Inference Network (FIN) is introduced for this challenging task. In the first stage of our training method, FIN is jointly trained with a fully convolutional network (FCN) for image-level tag prediction. A global smooth pooling layer is proposed, enabling FCN to assign object category tags to corresponding object regions, while FIN is capable of capturing all potential foreground regions with the predicted saliency maps. In the second stage, FIN is fine-tuned with its predicted saliency maps as ground truth. For refinement of ground truth, an iterative Conditional Random Field is developed to enforce spatial label consistency and further boost performance. Our method alleviates annotation efforts and allows the usage of existing large scale training sets with image-level tags. Our model runs at 60 FPS, outperforms unsupervised ones with a large margin, and achieves comparable or even superior performance than fully supervised counterparts.
Lijun Wang 0001, Huchuan Lu, Yifan Wang 0004, Mengyang Feng, Dong Wang 0004, Xiang Ruan
CVPR2
2017 Stepwise Metric Promotion for Unsupervised Video Person Re-identification
abstract
The intensive annotation cost and the rich but unlabeled data contained in videos motivate us to propose an unsupervised video-based person re-identification (re-ID) method. We start from two assumptions: 1) different video tracklets typically contain different persons, given that the tracklets are taken at distinct places or with long intervals; 2) within each tracklet, the frames are mostly of the same person. Based on these assumptions, this paper propose a stepwise metric promotion approach to estimate the identities of training tracklets, which iterates between cross-camera tracklet association and feature learning. Specifically, We use each training tracklet as a query, and perform retrieval in the cross-camera training set. Our method is built on reciprocal nearest neighbor search and can eliminate the hard negative label matches, i.e., the cross-camera nearest neighbors of the false matches in the initial rank list. The tracklet that passes the reciprocal nearest neighbor check is considered to have the same ID with the query. Experimental results on the PRID 2011, ILIDS-VID, and MARS datasets show that the proposed method achieves very competitive re-ID accuracy compared with its supervised counterparts.
Zimo Liu, Dong Wang 0004, Huchuan Lu
ICCV3
2017 A Stagewise Refinement Model for Detecting Salient Objects in Images
Tiantian Wang 0002, Ali Borji, Lihe Zhang, Huchuan Lu
ICCV5
2017 Amulet: Aggregating Multi-level Convolutional Features for Salient Object Detection
abstract
Fully convolutional neural networks (FCNs) have shown outstanding performance in many dense labeling problems. One key pillar of these successes is mining relevant information from features in convolutional layers. However, how to better aggregate multi-level convolutional feature maps for salient object detection is underexplored. In this work, we present Amulet, a generic aggregating multi-level convolutional feature framework for salient object detection. Our framework first integrates multi-level feature maps into multiple resolutions, which simultaneously incorporate coarse semantics and fine details. Then it adaptively learns to combine these feature maps at each resolution and predict saliency maps with the combined features. Finally, the predicted results are efficiently fused to generate the final saliency map. In addition, to achieve accurate boundary inference and semantic enhancement, edge-aware feature maps in low-level layers and the predicted results of low resolution features are recursively embedded into the learning framework. By aggregating multi-level convolutional features in this efficient and flexible manner, the proposed saliency model provides accurate salient object labeling. Comprehensive experiments demonstrate that our method performs favorably against state-of-the-art approaches in terms of near all compared evaluation metrics.
Dong Wang 0004, Huchuan Lu, Hongyu Wang 0001, Xiang Ruan
ICCV3
2017 Learning Uncertain Convolutional Features for Accurate Saliency Detection
abstract
Deep convolutional neural networks (CNNs) have delivered superior performance in many computer vision tasks. In this paper, we propose a novel deep fully convolutional network model for accurate salient object detection. The key contribution of this work is to learn deep uncertain convolutional features (UCF), which encourage the robustness and accuracy of saliency detection. We achieve this via introducing a reformulated dropout (R-dropout) after specific convolutional layers to construct an uncertain ensemble of internal feature units. In addition, we propose an effective hybrid upsampling method to reduce the checkerboard artifacts of deconvolution operators in our decoder network. The proposed methods can also be applied to other deep convolutional networks. Compared with existing saliency detection methods, the proposed UCF model is able to incorporate uncertainties for more accurate object boundary inference. Extensive experiments demonstrate that our proposed saliency model performs favorably against state-of-the-art approaches. The uncertain feature learning mechanism as well as the upsampling method can significantly improve performance on other pixel-wise vision tasks.
Dong Wang 0004, Huchuan Lu, Hongyu Wang 0001
ICCV3
2017 CNN for saliency detection with low-level feature integration
Hongyang Li 0001, Huchuan Lu, Zhizhen Chi
Neurocomputing3
2017 Saliency detection via joint modeling global shape and local consistency
Jinqing Qi, Shijing Dong, Huchuan Lu
Neurocomputing4
2017 Visual tracking with structured patch-based model
Fu Li 0003, Xu Jia 0012, Cheng Xiang 0001, Huchuan Lu
Image Vis. Comput.4
2017 Ranking Saliency
abstract
Most existing bottom-up algorithms measure the foreground saliency of a pixel or region based on its contrast within a local context or the entire image, whereas a few methods focus on segmenting out background regions and thereby salient objects. Instead of only considering the contrast between salient objects and their surrounding regions, we consider both foreground and background cues in this work. We rank the similarity of image elements with foreground or background cues via graph-based manifold ranking. The saliency of image elements is defined based on their relevances to the given seeds or queries. We represent an image as a multi-scale graph with fine superpixels and coarse regions as nodes. These nodes are ranked based on the similarity to background and foreground queries using affinity matrices. Saliency detection is carried out in a cascade scheme to extract background regions and foreground salient objects efficiently. Experimental results demonstrate the proposed method performs well against the state-of-the-art methods in terms of accuracy and speed. We also propose a new benchmark dataset containing 5,168 images for large-scale performance evaluation of saliency detection methods.
Lihe Zhang, Huchuan Lu, Xiang Ruan, Ming-Hsuan Yang 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2017 Interactive Video Segmentation via Local Appearance Model
abstract
In numerous video segmentation algorithms, shape and color priors from previous frames are propagated to successive frames for processing. One prime issue of the existing algorithms is how priors are modeled and propagated effectively. This paper proposes a novel algorithm for accurate and robust foreground prediction via a local appearance model based on shape and color. In the shape estimation process, instead of performing global matching, a local search mechanism is developed to capture complex motions. In addition, global information is used to facilitate local matching for the final shape estimation. Color cues from multiple frames are used to estimate the foreground (background) pixel distribution in the current image. Furthermore, the contribution of each frame is weighted based on its reliability and discriminative strength. Based on the proposed local appearance model, a graph cut algorithm is used to generate segmentation results. The experimental results show that accurate video segmentation can be obtained by the proposed algorithm with fewer user interactions.
Huchuan Lu
IEEE Trans. Circuits Syst. Video Technol.2