Bineng Zhong 0001

dblp:25/1637-1 · DBLP profile ↗
← Back
136ranked-venue papers
13as first author
83since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 96 · 7 first-author · 61 since 2021Artificial intelligence and machine learning · 67 · 9 first-author · 39 since 2021Computer networks · 6 · 2 since 2021Databases, data management, data science and information retrieval · 4 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Systems, architecture and hardware · 2 · 2 since 2021
YearPublicationVenuePosition
2026 MUTrack: A Memory-Aware Unified Representation Framework for Visual Tracking
abstract
Building a unified target representation that simultaneously achieves short-term adaptability and long-term stability is crucial for robust visual tracking. However, existing trackers typically face an inherent trade-off. Methods primarily relying on short-term appearance and motion cues achieve rapid adaptation, but they often struggle with long-term identity consistency. Conversely, trackers that emphasize extensive temporal context provide strong robustness, yet this approach can compromise their short-term adaptability. To bridge this gap, we propose a novel tracker, MUTrack, which comprehensively integrates both long-term and short-term memories into a unified target representation for more robust tracking. Specifically, we design a unified memory bank that stores and manages long-term memory for maintaining long-term identity consistency, and short-term memory for adapting to instantaneous appearance changes. To fully leverage the complementary nature of both long-term and short-term temporal information, we introduce a perception interaction module that dynamically fuses these memory types through deep and bidirectional interactions, enabling mutual refinement where one guides the other. This ultimately generates a highly adaptive target representation, which effectively balances adaptability to instantaneous changes with robustness against long-term identity drift. Extensive experiments on GOT10k, TrackingNet, LaSOT, LaSOT_ext, NfS, and OTB100 consistently demonstrate that MUTrack achieves SOTA performance.
Weijing Wu, Qihua Liang, Bineng Zhong 0001, Yufei Tan, Ning Li 0044, Yuanliang Xue
AAAI3
2026 Motion-Aware Object Tracking via Motion and Geometry-Aware Cues
abstract
Understanding motion is essential for visual object tracking, especially in complex and dynamic scenarios. Yet, many existing methods rely on simplistic strategies such as template updates or temporal feature propagation, often overlooking the deeper modeling of motion information. To mitigate this limitation, we introduce a motion-aware spatio-temporal framework that enhances motion perception by explicitly matching motion patterns and modeling inter-frame motion relationships. Central to our design is a motion pattern dictionary, which encodes a diverse set of representative motion cues as learnable features. During tracking, features from the search region interact with the dictionary to retrieve the most relevant motion patterns, allowing the model to adapt to the current motion state. A dedicated decoder further incorporates temporal correlations to refine motion awareness. To complement motion modeling, we embed geometric cues into the search region features, which strengthens spatial perception, reduces ambiguity under occlusion, and improves foreground-background separation. Extensive evaluations on seven challenging benchmarks demonstrate the effectiveness of our design. In particular, MoDTrack_384 surpasses recent SOTA trackers on LaSOT by 1.2% in AUC, highlighting the benefits of motion pattern modeling and geometry-guided enhancement in mitigating tracking drift.
Bineng Zhong 0001, Qihua Liang, Xiantao Hu, Yufei Tan, Haiying Xia, Shuxiang Song 0001
AAAI2
2026 Aware Distillation for Robust Vision-Language Tracking Under Linguistic Sparsity
abstract
Vision-language object tracking overcomes the limitations of relying solely on visual features by leveraging language descriptions of objects to provide cross-modal semantic information, thereby enhancing model robustness in complex scenarios. However, most existing high-performance vision-language trackers are trained jointly on pure visual data and vision-language multimodal data. Due to the relative sparsity of language annotations in the data, the trackers tend to prioritize the localization role of visual features, diminishing the model's attention to language information. To mitigate this issue, we propose a novel vision-language tracker: Aware Distillation for Robust Vision-Language Tracking under Linguistic Sparsity (ADTrack). We introduce a knowledge distillation framework employing a knowledge-rich teacher model and a lightweight student model to establish modality correlations between vision and language, enabling efficient modeling between visual information and language descriptions. Specifically, our lightweight student module simultaneously distills language encoding capabilities from large language models through teacher-guided learning on input language, while performing target-aware perception on template images using language descriptions to generate more effective template features for subsequent visual extraction. Furthermore, to ensure perceptual robustness in linguistically sparse scenarios, we simulate language-deficient conditions during training and employ contrastive learning to enhance model adaptability. Extensive experiments demonstrate that ADTrack reduces parameters by over 50% while achieving state-of-the-art (SOTA) performance and speed on vision-language tracking benchmarks, including LaSOT, LaSOText, TNL2K, OTB-Lang and MGIT.
Guangtong Zhang, Bineng Zhong 0001, Shirui Yang, Tian Bai 0002
AAAI2
2026 PCT: Pose Convolutional Transformer for skeleton-based action recognition
Bineng Zhong 0001, Chuanqi Li
Comput. Vis. Image Underst.2
2026 IGEFusion: Low-light infrared and visible image fusion via infrared-guided enhancement
Shengjia An, Zhi Li 0017, Shaorong Zhang, Bineng Zhong 0001
Image Vis. Comput.5
2026 Updatable one-stream Vision-Language tracking via multilayer perceptual memory network
Peipei Song, Huanlong Zhang, Bin Jiang 0007, Bineng Zhong 0001
Pattern Recognit.6
2026 Curriculum adaptation for one-stream RGB-T tracking
Xiantao Hu, Fansheng Zeng, Bineng Zhong 0001, Zhangyong Tang, Wenxuan Fang 0001, Jun Li 0027, Ying Tai, Jian Yang 0003
Pattern Recognit.3
2026 Cascaded Image Fusion Bridged RGB-T Tracking
Shengjia An, Zhi Li 0017, Shaorong Zhang, Zhenjie Li, Bineng Zhong 0001
IEEE Signal Process. Lett.5
2026 IPEFusion: Infrared Prior-Enhanced Low-Light Infrared and Visible Image Fusion Framework
Shengjia An, Zhi Li 0017, Shaorong Zhang, Bineng Zhong 0001
IEEE Signal Process. Lett.5
2026 Unified Multi-Modal Tracking via Proxy Visual Prompts
abstract
Recent unified multi-modal tracking frameworks often encounter high computational overhead due to complex fusion operations. In this paper, we propose PoATrack, a proxy-based bidirectional fusion framework designed to facilitate efficient multi-modal tracking via interactive proxy prompts. Specifically, the framework treats features from all modalities equally and introduces an adaptive feature enhancement module to improve spatial representations within the search region. To further reduce the fusion cost, a lightweight proxy prompt fusion module is developed following a two-stage strategy: (1) representing modality-specific features through proxy visual prompts, and (2) dynamically learning hierarchical cross-modal relationships via proxy-guided fusion for enriched contextual modeling. Extensive experiments on six public benchmarks (RGB-T, RGB-E, and RGB-D) demonstrate that PoATrack achieves competitive accuracy while operating 1.8× faster than SDSTrack [17] under identical hardware settings.
Bineng Zhong 0001, Yaozong Zheng, Qihua Liang, Zhiruo Zhu, Shuxiang Song 0001
IEEE Trans Autom. Sci. Eng.2
2026 Text-Guided Vision Token Reduction With Low-Rank Adaptation for Efficient Visual Grounding
abstract
Transformer-based pretrained models have advanced Visual Grounding (VG) significantly, but their scaling up has caused soaring training and inference costs. While efforts adapting Parameter-Efficient Fine-Tuning to VG have cut training expenses to some extent, inference costs are over-looked, stemming from Transformer’s computation costs growing quadratically with input token length. To address the above challenge, we propose a text-guided token prune method based on a one-stream architecture for VG, named OneSVG. Specifically, OneSVG transfers low-rank adaptation LoRA to VG, and calculates the relevance between each vision token and text semantics in multiple stages, and gradually prunes the vision tokens with low relevance. In addition, traditional VG methods use a single [REG] token to predict bounding boxes, and the [REG] token relies on complete vision tokens during training, pruning large number of vision tokens will disrupt the spatial orientation information of images. Therefore, the prediction head in OneSVG first restores the square structure of images by padding the missing area, and then accepts complete vision tokens. Experimental results on three widely-used benchmarks demonstrate that our OneSVG achieves state-of-the-art real-time speed while maintaining the best accuracy with only 34.3% tokens. Codes and models are available at https://github.com/ltShi/OneSVG.
Liangtao Shi, Ting Liu 0018, Jinxia Xie, Ning Li 0044, Bineng Zhong 0001, Richang Hong
IEEE Trans. Circuits Syst. Video Technol.5
2026 FMTrack: Frequency-Aware Interaction and Multi-Expert Fusion for RGB-T Tracking
abstract
Recently, RGB-T tracking has received increasing attention due to its robustness. However, existing RGB-T trackers mainly use cross-attention for modal feature interaction, limiting the utilization of complementary information. In addition, these trackers employ fixed dominant-auxiliary paradigms for feature fusion, ignoring modal quality fluctuations. To address these issues, we propose FMTrack, an effective framework for fully capturing complementary information. FMTrack consists of two key components, a frequency-aware interaction network (FIN) and a multi-expert fusion module (MEFM). To emphasize the valuable information in each modality, FIN utilizes frequency masks to perform high-pass and low-pass filtering on RGB and TIR data. FIN explicitly establishes cross-modal interactions via frequency domain learning, which facilitates the sharing of complementary information. Besides, MEFM extracts diverse features via the differentiated expert network and then adjusts feature combinations according to modal reliability, achieving deep understanding and flexible fusion of multimodal data. With FIN and MEFM, FMTrack makes full use of the advantageous information of each modality to highlight target representations, thus improving performance in complex scenes. Extensive experiments on four popular RGBT tracking datasets (LasHeR, VTUAV, RGBT234, and RGBT210) show that our FMTrack achieves leading performance. The code is available at https://github.com/xyl-507/FMTrack.
Yuanliang Xue, Guodong Jin, Bineng Zhong 0001, Lining Tan, Chaocan Xue, Yaozong Zheng
IEEE Trans. Circuits Syst. Video Technol.3
2026 Mamba-Driven Diffusion Model for Salient Object Detection in Optical Remote Sensing Images
abstract
Existing Optical Remote Sensing Image Salient Object Detection (ORSI-SOD) methods mainly rely on a semantic segmentation paradigm, which relies on pixel-wise probabilities, leading to overconfident mispredictions. In contrast, the random sampling process of the diffusion model allows multiple possible predictions to be drawn from the mask distribution, effectively alleviating this problem. However, existing diffusion models mainly use Transformers as conditional feature extraction networks. Although they are good at global modeling, they have limited ability to handle long-range dependencies due to computational complexity. To overcome these challenges, we introduce MambaDif, an innovative diffusion model architecture based on Mamba. Specifically, we regard ORSI-SOD as a conditional mask generation task leveraging the diffusion model and achieving target distribution matching by adding noise to the mask and iteratively denoising it to match the target distribution. Then, we adopt Mamba to extract global features, efficiently process long sequences, and capture global contextual information with linear complexity. In addition, we introduce the global-local feature collaborative completion module (GLM), which combines the ability of convolutional layers to extract local features with the advantage of Mamba in capturing long-range dependencies, thereby achieving excellent denoising performance. Extensive experiments show that MambaDif outperforms SOTA methods in eight evaluation metrics on two standard datasets (EORSSD and ORSSD). We also report the generalization performance of the model on the challenging ORSI-4199 to evaluate its robustness.
Bineng Zhong 0001, Qihua Liang, Yufei Tan, Haiying Xia, Shuxiang Song 0001
IEEE Trans. Circuits Syst. Video Technol.2
2026 Robust RGB-T Tracking via Learnable Visual Fourier Prompt Fine-Tuning and Modality Fusion Prompt Generation
abstract
Recently, visual prompt tuning is introduced to RGB-Thermal (RGB-T) tracking as a parameter-efficient finetuning (PEFT) method. However, these PEFT-based RGB-T tracking methods typically rely solely on spatial domain information as prompts for feature extraction. As a result, they often fail to achieve optimal performance by overlooking the crucial role of frequency-domain information in prompt learning. To address this issue, we propose an efficient Visual Fourier Prompt Tracking (named VFPTrack) method to learn modality-related prompts via Fast Fourier Transform (FFT). Our method consists of symmetric feature extraction encoder with shared parameters, visual-fourier prompts, and Modality Fusion Prompt Generator that generates bidirectional interaction prompts through multi-modal feature fusion. Specifically, we first use a frozen feature extraction encoder to extract RGB and thermal infrared (TIR) modality features. Then, we combine the visual prompts in the spatial domain with the frequency domain prompts obtained from the FFT, which allows for the full extraction and understanding of modality features from different domain information. Finally, unlike previous fusion methods, the modality fusion prompt generation module we use combines features from different modalities to generate a fused modality prompt. This modality prompt is interacted with each individual modality to fully enable feature interaction across different modalities. Extensive experiments conducted on three popular RGB-T tracking benchmarks show that our method demonstrates outstanding performance.
Bineng Zhong 0001, Qihua Liang, Zhiruo Zhu, Yaozong Zheng, Ning Li 0044
IEEE Trans. Multim.2
2026 RWKV-Inspired Multi-Modal Relation Modeling for Vision-Language Tracking
abstract
Vision-language object tracking can provide more state representations for targets by introducing the language modality, achieving more robust tracking and localization. Therefore, designing multi-modal interactions to achieve feature alignment between vision and language has been one of the research hotspots. However, existing multi-modal interaction methods face two key issues: on the one hand, they lack effective exploration of modeling the relationship between the contextual information of language sequences and visual features; on the other hand, the introduction of modalities leads to increased computational time costs in multi-modal interactions, which severely affects the real-time performance of vision-language tracking algorithms. To address these challenges, we propose a vision-language tracking framework called RWKV-Inspired Multi-modal Relation Modeling for Vision-Language Tracking (RrmTrack). We introduce a novel modality interaction method specific to vision-language object tracking based on RWKV, providing customized interaction for different modalities in vision-language tracking and effectively reducing the computational time cost of cross-modal interaction. Specifically, this method uses a time mixing module to model the relationship between language information and image features, and a channel mixing module to facilitate information interaction between images. By combining parallelized training with a linear attention mechanism and efficient RNN inference, it enables accurate and fast target localization in vision-language tracking. Additionally, we propose a novel feature extraction structure that integrates Siamese and One-stream architectures. An information restoration module is designed to reduce the information interference introduced by the search image to the template image during interaction. RrmTrack achieves state-of-the-art results and speed on multiple vision-language object tracking benchmarks, including TNL2k, LaSOT, OTB-Lang, LaSOText, and MGIT.
Guangtong Zhang, Bineng Zhong 0001, Yuhao Mu, Tian Bai 0002
IEEE Trans. Multim.3
2025 Exploiting Multimodal Spatial-temporal Patterns for Video Object Tracking
abstract
Multimodal tracking has garnered widespread attention as a result of its ability to effectively address the inherent limitations of traditional RGB tracking. However, existing multimodal trackers mainly focus on the fusion and enhancement of spatial features or merely leverage the sparse temporal relationships between video frames. These approaches do not fully exploit the temporal correlations in multimodal videos, making it difficult to capture the dynamic changes and motion information of targets in complex scenarios. To alleviate this problem, we propose a unified multimodal spatial-temporal tracking approach named STTrack. In contrast to previous paradigms that solely relied on updating reference information, we introduced a temporal state generator (TSG) that continuously generates a sequence of tokens containing multimodal temporal information. These temporal information tokens are used to guide the localization of the target in the next time state, establish long-range contextual relationships between video frames, and capture the temporal trajectory of the target. Furthermore, at the spatial level, we introduced the mamba fusion and background suppression interactive (BSI) modules. These modules establish a dual-stage mechanism for coordinating information interaction and fusion between modalities. Extensive comparisons on five benchmark datasets illustrate that STTrack achieves state-of-the-art performance across various multimodal tracking scenarios.
Xiantao Hu, Ying Tai, Xu Zhao 0001, Chen Zhao 0002, Zhenyu Zhang 0005, Jun Li 0027, Bineng Zhong 0001, Jian Yang 0003
AAAI7
2025 MambaLCT: Boosting Tracking via Long-term Context State Space Model
abstract
Effectively constructing context information with long-term dependencies from video sequences is crucial for object tracking. However, the context length constructed by existing work is limited, only considering object information from adjacent frames or video clips, leading to insufficient utilization of contextual information. To address this issue, we propose MambaLCT, which constructs and utilizes target variation cues from the first frame to the current frame for robust tracking. First, a novel unidirectional Context Mamba module is designed to scan frame features along the temporal dimension, gathering target change cues throughout the entire sequence. Specifically, target-related information in frame features is compressed into a hidden state space through a selective scanning mechanism. The target information across the entire video is continuously aggregated into target variation cues. Next, we inject the target change cues into the attention mechanism, providing temporal information for modeling the relationship between the template and search frames. The advantage of MambaLCT is its ability to continuously extend the length of the context, capturing complete target change cues, which enhances the stability and robustness of the tracker. Extensive experiments show that long-term context information enhances the model's ability to perceive targets in complex scenarios. MambaLCT achieves new SOTA performance on six benchmarks while maintaining real-time runing speeds.
Xiaohai Li, Bineng Zhong 0001, Qihua Liang, Guorong Li, Zhiyi Mo, Shuxiang Song 0001
AAAI2
2025 Robust Tracking via Mamba-based Context-aware Token Learning
abstract
How to make a good trade-off between performance and computational cost is crucial for a tracker. However, current famous methods typically focus on complicated and time-consuming learning that combining temporal and appearance information by input more and more images (or features). Consequently, these methods not only increase the model's computational source and learning burden but also introduce much useless and potentially interfering information. To alleviate the above issues, we propose a simple yet robust tracker that separates temporal information learning from appearance modeling and extracts temporal relations from a set of representative tokens rather than several images (or features). Specifically, we introduce one track token for each frame to collect the target's appearance information in the backbone. Then, we design a mamba-based Temporal Module for track tokens to be aware of context by interacting with other track tokens within a sliding window. This module consists of a mamba layer with autoregressive characteristic and a cross-attention layer with strong global perception ability, ensuring sufficient interaction for track tokens to perceive the appearance changes and movement trends of the target. Finally, track tokens serve as a guidance to adjust the appearance feature for the final prediction in the head. Experiments show our method is effective and achieves competitive performance on multiple benchmarks at a real-time speed.
Jinxia Xie, Bineng Zhong 0001, Qihua Liang, Ning Li 0044, Zhiyi Mo, Shuxiang Song 0001
AAAI2
2025 Less Is More: Token Context-Aware Learning for Object Tracking
abstract
Recently, several studies have shown that utilizing contextual information to perceive target states is crucial for object tracking. They typically capture context by incorporating multiple video frames. However, these naive frame-context methods fail to consider the importance of each patch within a reference frame, making them susceptible to noise and redundant tokens, which deteriorates tracking performance. To address this challenge, we propose a new token context-aware tracking pipeline named LMTrack, designed to automatically learn high-quality reference tokens for efficient visual tracking. Embracing the principle of Less is More, the core idea of LMTrack is to analyze the importance distribution of all reference tokens, where important tokens are collected, continually attended to, and updated. Specifically, a novel Token Context Memory module is designed to dynamically collect high-quality spatio-temporal information of a target in an autoregressive manner, eliminating redundant background tokens from the reference frames. Furthermore, an effective Unidirectional Token Attention mechanism is designed to establish dependencies between reference tokens and search frame, enabling robust cross-frame association and target localization. Extensive experiments demonstrate the superiority of our tracker, achieving state-of-the-art results on tracking benchmarks such as GOT-10K, TrackingNet, and LaSOT.
Chenlong Xu, Bineng Zhong 0001, Qihua Liang, Yaozong Zheng, Guorong Li, Shuxiang Song 0001
AAAI2
2025 Decoupled Spatio-Temporal Consistency Learning for Self-Supervised Tracking
abstract
The success of visual tracking has been largely driven by datasets with manual box annotations. However, these box annotations require tremendous human effort, limiting the scale and diversity of existing tracking datasets. In this work, we present a novel Self-Supervised Tracking framework, named SSTrack, designed to eliminate the need of box annotations. Specifically, a decoupled spatio-temporal consistency training framework is proposed to learn rich target information across timestamps through global spatial localization and local temporal association. This allows for the simulation of appearance and motion variations of instances in real-world scenarios. Furthermore, an instance contrastive loss is designed to learn instance-level correspondences from a multi-view perspective, offering robust instance supervision without additional labels. This new design paradigm enables SSTrack to effectively learn generic tracking representations in a self-supervised manner, while reducing reliance on extensive box annotations. Extensive experiments on nine benchmark datasets demonstrate that SSTrack surpasses SOTA self-supervised tracking methods, achieving an improvement of more than 25.3%, 20.4%, and 14.8% in AUC (AO) score on the GOT10K, LaSOT, TrackingNet datasets, respectively.
Yaozong Zheng, Bineng Zhong 0001, Qihua Liang, Ning Li 0044, Shuxiang Song 0001
AAAI2
2025 Dynamic Updates for Language Adaptation in Visual-Language Tracking
abstract
The consistency between the semantic information provided by the multi-modal reference and the tracked object is crucial for visual-language (VL) tracking. However, existing VL tracking frameworks rely on static multi-modal references to locate dynamic objects, which can lead to semantic discrepancies and reduce the robustness of the tracker. To address this issue, we propose a novel vision-language tracking framework, named DUTrack, which captures the latest state of the target by dynamically updating multimodal references to maintain consistency. Specifically, we introduce a Dynamic Language Update Module, which leverages a large language model to generate dynamic language descriptions for the object based on visual features and object category information. Then, we design a Dynamic Template Capture Module, which captures the regions in the image that highly match the dynamic language descriptions. Furthermore, to ensure the efficiency of description generation, we design an update strategy that assesses changes in target displacement, scale, and other factors to decide on updates. Finally, the dynamic template and language descriptions that record the latest state of the target are used to update the multi-modal references, providing more accurate reference information for subsequent inference and enhancing the robustness of the tracker. DUTrack achieves new state-of-the-art performance on five mainstream vision-language and two vision-only tracking benchmarks, including LaSOT, LaSOText, TNL2K, OTB99-Lang, MGIT, GOT-10K, and UAV123. Code and models are available at https://github.com/GXNU-ZhongLab/DUTrack.
Xiaohai Li, Bineng Zhong 0001, Qihua Liang, Zhiyi Mo, Jian Nong, Shuxiang Song 0001
CVPR2
2025 Similarity-Guided Layer-Adaptive Vision Transformer for UAV Tracking
abstract
Vision transformers (ViTs) have emerged as a popular backbone for visual tracking. However, complete ViT architectures are too cumbersome to deploy for unmanned aerial vehicle (UAV) tracking which extremely emphasizes efficiency. In this study, we discover that many layers within lightweight ViT-based trackers tend to learn relatively redundant and repetitive target representations. Based on this observation, we propose a similarity-guided layer adaptation approach to optimize the structure of ViTs. Our approach dynamically disables a large number of representation-similar layers and selectively retains only a single optimal layer among them, aiming to achieve a better accuracy-speed trade-off. By incorporating this approach into existing ViTs, we tailor previously complete ViT architectures into an efficient similarity-guided layer-adaptive framework, namely SGLATrack, for real-time UAV tracking. Extensive experiments on six tracking benchmarks verify the effectiveness of the proposed approach, and show that our SGLATrack achieves a state-of-the-art real-time speed while maintaining competitive tracking precision. Codes and models are available at https://github.com/GXNU-ZhongLab/SGLATrack.
Chaocan Xue, Bineng Zhong 0001, Qihua Liang, Yaozong Zheng, Ning Li 0044, Yuanliang Xue, Shuxiang Song 0001
CVPR2
2025 Explicit Context Reasoning with Supervision for Visual Tracking
abstract
Contextual reasoning with constraints is crucial for enhancing temporal consistency in cross-frame modeling for visual tracking. However, mainstream tracking algorithms typically associate context by merely stacking historical information without explicitly supervising the association process, making it difficult to effectively model the target's evolving dynamics. To alleviate this problem, we propose RSTrack, which explicitly models and supervises context reasoning via three core mechanisms. 1) Context Reasoning Mechanism : Constructs a target state reasoning pipeline, converting unconstrained contextual associations into a temporal reasoning process that predicts the current representation based on historical target states, thereby enhancing temporal consistency. 2) Forward Supervision Strategy : Utilizes true target features as anchors to constrain the reasoning pipeline, guiding the predicted output toward the true target distribution and suppressing drift in the context reasoning process. 3) Efficient State Modeling : Employs a compression-reconstruction mechanism to extract the core features of the target, removing redundant information across frames and preventing ineffective contextual associations. These three mechanisms collaborate to effectively alleviate the issue of contextual association divergence in traditional temporal modeling. Experimental results show that RSTrack achieves state-of-the-art performance on multiple benchmark datasets while maintaining real-time running speeds. Our code is available at https://github.com/GXNU-ZhongLab/RSTrack.
Fansheng Zeng, Bineng Zhong 0001, Haiying Xia, Yufei Tan, Xiantao Hu, Liangtao Shi, Shuxiang Song 0001
ACM Multimedia2
2025 SAM-Assisted Temporal-Location Enhanced Transformer Segmentation for Object Tracking with Online Motion Inference
Huanlong Zhang, Xiangbo Yang, Xin Wang 0137, Weiqiang Fu, Bineng Zhong 0001
Neurocomputing5
2025 Lightweight completion with high-order semantic attributes for heterogeneous sparse attribute graph learning
Yuanjun Yang, Weihua Ou, Yunshun Wu, Jianping Gou, Bineng Zhong 0001
Knowl. Based Syst.8
2025 Target-background interaction modeling transformer for object tracking
Huanlong Zhang, Weiqiang Fu, Bineng Zhong 0001, Xin Wang 0137, Yanfeng Wang 0002
Knowl. Based Syst.4
2025 Towards Universal Modal Tracking With Online Dense Temporal Token Learning
abstract
We propose a universal video-level modality-awareness tracking model with online dense temporal token learning (called UM-ODTrack). It is designed to support various tracking tasks, including RGB, RGB+Thermal, RGB+Depth, and RGB+Event, utilizing the same model architecture and parameters. Specifically, our model is designed with three core goals: Video-level Sampling. We expand the model's inputs to a video sequence level, aiming to see a richer video context from an near-global perspective. Video-level Association. Furthermore, we introduce two simple yet effective online dense temporal token association mechanisms to propagate the appearance and motion trajectory information of target via a video stream manner. Modality Scalable. We propose two novel gated perceivers that adaptively learn cross-modal representations via a gated attention mechanism, and subsequently compress them into the same set of model parameters via a one-shot training manner for multi-task inference. This new solution brings the following benefits: (i) The purified token sequences can serve as temporal prompts for the inference in the next video frames, whereby previous information is leveraged to guide future inference. (ii) Unlike multi-modal trackers that require independent training, our one-shot training scheme not only alleviates the training burden, but also improves model representation. Extensive experiments on visible and multi-modal benchmarks show that our UM-ODTrack achieves a new SOTA performance.
Yaozong Zheng, Bineng Zhong 0001, Qihua Liang, Shengping Zhang, Guorong Li, Xianxian Li, Rongrong Ji
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 OSFusion: A One-Stream Infrared and Visible Image Fusion Framework
abstract
The current popular two-stream two-stage image fusion framework extracts features of infrared and visible images separately and then performs feature fusion. The extracted features lack interaction between the source images and have limited cross-modal complementary capability. To address these issues, we propose a novel one-stream infrared and visible image fusion (OSFusion) framework that connects a source image pair to achieve bidirectional information flow. In this way, the fused features with cross-modal complementary information can be dynamically extracted by mutual guidance. To further improve the inference efficiency and obtain high-quality fused images, a feature extraction and fusion module (FEFM) is proposed based on Transformer structure. The combination of feature extraction and feature fusion is realized by using it. Since there is no need for an extra feature interaction module and the implementation is highly parallel, the speed of image fusion is extremely fast. Benefiting from the one-stream structure and FEFM, OSFusion achieves promising infrared and visible image fusion performance on MSRS, M3FD, and RoadScene datasets. Besides, our method achieves a good balance in the trade-off between performance and complexity, and also shows a faster convergence trend.
Shengjia An, Zhi Li 0017, Shaorong Zhang, Bineng Zhong 0001
IEEE Signal Process. Lett.5
2025 SIEVL-Track: Exploring Semantic Information Enhancement for Visual-Language Object Tracking
abstract
With the assistance of language descriptions, Visual-Language (VL) object tracking can obtain more accurate semantic information compared to traditional Visual-Only object tracking. However, the ability of current VL trackers to obtain target semantic information has not been fully developed due to limitations such as wasted modeling capabilities and insufficient utilization of historical temporal information. On the one hand, the modeling output from Transformer shallow encoders often does not directly participate in the prediction of tracking results, resulting in a certain degree of model capability waste. On the other hand, the semantic information of historical tracking results has also not been fully utilized in the tracking process, resulting in a certain degree of lack of semantic assistance capability. Therefore, we propose a novel hierarchical multi-stage VL tracker called SIEVL-Track to enhance target semantic information. Specifically, we first design a multi-stage visual language tracking framework for modeling multi-scale semantic information in Visual-Language tracking pipeline. Secondly, we propose a selective deep and shallow semantic information fusion module (S-DSFM) that explicitly integrates shallow output features into deep output features, so to reduce the waste of modeling capabilities and obtain more high-frequency semantic information related to the target. Finally, we design a temporal cue modeling module based on linguistic classification and multi-frame historical information(MHLS-TCM), with the aim of more comprehensive utilization of historical temporal semantic information. Benefit from the above designs, our VL tracker can obtain stronger target semantic information. Competitive performance from extensive experimental results on five popular vision-language tracking benchmarks, including LaSOT, OTB99-Lang, WebUAV-3M, LaSOText and TNL2K, have demonstrated the superiority and effectiveness of our SIEVL-Track.
Ning Li 0044, Bineng Zhong 0001, Qihua Liang, Zhiyi Mo, Jian Nong, Shuxiang Song 0001
IEEE Trans. Circuits Syst. Video Technol.2
2025 Mamba Adapter: Efficient Multi-Modal Fusion for Vision-Language Tracking
abstract
Utilizing the high-level semantic information of language to compensate for the limitations of vision information is a highly regarded approach in single-object tracking. However, most existing vision-language (VL) trackers employ full-parameter fine-tuning, which can easily lead to catastrophic forgetting. Therefore, they fail to fully exploit the prior knowledge of pre-trained models from upstream tasks, resulting in unsatisfactory tracking performance. To alleviate the above problem, we propose a simple yet effective Vision-Language Tracking pipeline based on Mamba Adapter, named MAVLT, which adopts the idea of parameter-efficient fine-tuning (PEFT) to realize the interaction between vision-language modalities. This novel approach offers the following advantages: (1)The knowledge of the upstream pre-trained model is efficiently inherited by freezing its parameters. This ensures that the VL tracking framework only learns the modules for vision and language interaction, with a focus on the fusion between modalities. (2)The modal interaction between language and vision encoders is flexibly bridged in each encoder layer via proposed mamba adapter, enabling efficient interaction of visual and language information at multiple levels. Extensive experiments on five popular vision-language tracking benchmarks validate the effectiveness of the proposed MAVLT. Particularly, the MAVLT achieves 73.4% AUC score on the LaSOT benchmarks with only 0.18%(0.32M) of the total parameters updates. Code and models are available at https://github.com/GXNU-ZhongLab/MAVLT.
Liangtao Shi, Bineng Zhong 0001, Qihua Liang, Xiantao Hu, Zhiyi Mo, Shuxiang Song 0001
IEEE Trans. Circuits Syst. Video Technol.2
2025 AVLTrack: Dynamic Sparse Learning for Aerial Vision-Language Tracking
abstract
The introduction of natural language for vision-language (VL) tracking has been proven to improve performance. However, natural language remains under-explored in existing aerial trackers. Moreover, existing VL trackers ignore the misalignment of language with dynamic target states, which is prominent in complex UAV scenarios. In this work, we present AVLTrack, a flexible framework for aerial vision-language tracking. It consists of three key components, a dynamic sparse learning (DSL) module, an efficient Transformer backbone, and a multi-level language perception (MLP) strategy. First, DSL sparsely connects language and images via dynamic sparse attention, providing accurate multi-modal prompts. To adapt to target state variations, the sparsity in DSL is dynamically adjusted based on semantic information, flexibly highlighting target-specific tokens. Next, the Transformer backbone follows highly parallelized one-stream architectures, allowing efficient multi-modal feature extraction and interaction. Finally, MLP enables the iterative interaction of language and visual information, aiming to utilize language priori to guide the generation of discriminative visual features. Moreover, we construct the DTB70-NLP dataset to facilitate UAV vision-language tracking. Extensive experiments on WebUAV-3M and DTB70-NLP demonstrate the leading performance of AVLTrack compared to existing outstanding trackers while maintaining a high running speed of 80.5 FPS. The dataset and codes are available athttps://github.com/xyl-507/AVLTrack.
Yuanliang Xue, Bineng Zhong 0001, Guodong Jin, Lining Tan, Ning Li 0044, Yaozong Zheng
IEEE Trans. Circuits Syst. Video Technol.2
2025 Unifying Motion and Appearance Cues for Visual Tracking via Shared Queries
abstract
The rich motion and appearance cues between consecutive frames are crucial for robust visual tracking. However, most existing tracking methods are still limited in designing different components to separately employ corresponding cues and even ignore one of them. This makes them difficult to maintain effective interaction between different cues, thus hindering the models from fostering a comprehensive understanding of the target objects. To address these issues, we propose a unified spatio-temporal cues learning framework (named USCLTrack) that comprehensively mines the variation patterns of targets between consecutive frames in complex video streams. Specifically, USCLTrack firstly aggregates motion and appearance cues into shared queries to provide the bridge of interaction between both cues. Then, it directly generates object locations on the condition of these shared queries in an autoregressive manner, unifying different cues to guide future inferences. To effectively learn multiple spatio-temporal cues aggregated in the shared queries, we develop a spatio-temporal attention mechanism. This mechanism integrates motion cues with appearance cues according to the time steps for ensuring temporal consistency. Moreover, it concurrently captures motion trends and appearance changes to facilitate the understanding of the target objects. Extensive experiments on eight popular tracking benchmarks validate the effectiveness of the proposed USCLTrack.
Chaocan Xue, Bineng Zhong 0001, Qihua Liang, Haiying Xia, Shuxiang Song 0001
IEEE Trans. Circuits Syst. Video Technol.2
2025 Adaptive Expert Decision for RGB-T Tracking
abstract
The features provided by RGB and Thermal Infrared (TIR) images have their own characteristics. Therefore, how to adaptively fuse multi-modal features according to different tracking scenarios is crucial for RGB-T tracking. However, current mainstream RGB-T tracking algorithms often use fixed fusion operations for modal interaction in different scenarios. Consequently, their tracking permanence is deteriorated due to they are unable to dynamically adjust the fused multi-modal features based on the current scenes. To address this issue, we propose a novel RGB-T tracking algorithm called AETrack, which can dynamically extract effective modal features in different scenarios for adaptive fusion. Firstly, we design an adaptive expert decision mechanism that employs multiple experts to process the input features. Each expert focuses on and learns different relevant features. Based on this mechanism, we then propose a feature-guided method that leverages the correlations between modalities to provide cross-modal information. This guidance enables the adaptive expert mechanism to adaptively select the most suitable expert to output effective features based on different scenarios, ensuring that our proposed AETrack prioritizes effective features and thus alleviates interference from irrelevant information. Finally, we design a Progressive Cross-modal Fusion operation to achieve multi-level adaptive fusion of effective features across different modalities. Benefiting from this adaptive fusion process, we can effectively achieve multi-modal interaction in different scenarios to guide robust tracking. Extensive experiments on three popular benchmarks (i.e., LasHeR, RGBT210, RGBT234) show that our proposed AETrack can significantly improve tracking performance.
Zhiruo Zhu, Bineng Zhong 0001, Qihua Liang, Yaozong Zheng, Ning Li 0044
IEEE Trans. Circuits Syst. Video Technol.2
2025 Vision-language discriminative fusion network for object tracking
Huanlong Zhang, Liusen Xu, Bineng Zhong 0001
J. Supercomput.6
2025 Robust Multi-Stage Tracking via Multi-Scale and Multi-Level Representation Learning
abstract
How to learn multi-scale and multi-level representations is crucial for robust tracking. However, most current one-stream structure based trackers with visual transformers (dubbed ViTs) cannot effectively capture multi-scale representations due to the structure of their adopted ViTs is non-hierarchical. Meanwhile, they often only use the output features from the final layer for predicting results (i.e., ignoring the utilization of low-level features from the shallow layers) which may result in a certain degree of lacking multi-level representation learning ability. To address these issues, we propose a robust multi-stage tracker that effectively combines the advantages of both hierarchical and one-stream structured ViT as a tracking backbone to improve the multi-scale and multi-level representation learning abilities. Specifically, first of all, we design a hierarchical tracker with a three-stage backbone. In the first two stages of our tracker, we utilize a dual-branch structure to obtain multi-scale features of the template and search region separately. Especially, We design the local scale awareness modules based on simple MLP layers to capture multi-scale features. These modules remove complex operations such as convolutions or shifted window attentions, thus avoiding the performance degradation caused by traditional hierarchical ViTs. In the third stage (i.e. the main stage), we construct a global encoder based on the one-stream ViT to achieve efficient feature extraction and feature interaction for our tracker. Then, we design a multi-level feature integration module in the main stage to explicitly utilize the representation information learned from the shallow layers and fuse them with the features of the final layer to obtain multi-level representation information. Lastly, benefit from the these designs, our tracker can effectively capture more multi-scale and multi-level representations for robust tracking. Comprehensive experiments on GOT-10k, LaSOT, LaSOT$_{ext}$, TNL2K, UAV123, TrackingNet and VOT2020 benchmarks validate the effectiveness and robustness of our method.
Ning Li 0044, Bineng Zhong 0001, Qihua Liang, Zhiyi Mo, Shuxiang Song 0001
IEEE Trans. Multim.2
2025 Uncertainty-Guided Diffusion Model for Camouflaged Object Detection
abstract
Recently, diffusion models have significantly improved the performance of Camouflaged Object Detection (COD) by adding noise to a mask and iteratively denoising it to match the target distributions. Due to the direct extraction of features from noisy masks and the lack of conditional constraints on a prediction area, the diffusion model may deviate from a correct prediction range and produces mispredictions in regions with high uncertainty. To address this issue, we propose an uncertainty-guided diffusion model (UGDNet) for COD, which explicitly quantifies uncertainty and integrates it as an anchor condition into the diffusion models to provide an initialization of the diffusion regions. The core idea is first to utilize a probability representation and transformer to explicitly model uncertainty, aiming to identify areas where a model may generate overconfident mispredictions. Then, we use the uncertainty as an anchor condition to provide a reference prediction range for the diffusion model, guiding each step of the diffusion process. Furthermore, we use uncertainty to guide feature aggregation, prompting the model to pay extra attention to the semantic features of regions with high uncertainty to refine the segmentation results further. The experimental results indicate that our proposed UGDNet achieves higher accuracy than existing state-of-the-art models on five COD benchmarks, including COD10K, NC4K, CAMO, CHAMELEON, and CDS2K.
Bineng Zhong 0001, Qihua Liang, Zhiyi Mo, Shengping Zhang, Shuxiang Song 0001
IEEE Trans. Multim.2
2024 Explicit Visual Prompts for Visual Object Tracking
abstract
How to effectively exploit spatio-temporal information is crucial to capture target appearance changes in visual tracking. However, most deep learning-based trackers mainly focus on designing a complicated appearance model or template updating strategy, while lacking the exploitation of context between consecutive frames and thus entailing the when-and-how-to-update dilemma. To address these issues, we propose a novel explicit visual prompts framework for visual tracking, dubbed EVPTrack. Specifically, we utilize spatio-temporal tokens to propagate information between consecutive frames without focusing on updating templates. As a result, we cannot only alleviate the challenge of when-to-update, but also avoid the hyper-parameters associated with updating strategies. Then, we utilize the spatio-temporal tokens to generate explicit visual prompts that facilitate inference in the current frame. The prompts are fed into a transformer encoder together with the image tokens without additional processing. Consequently, the efficiency of our model is improved by avoiding how-to-update. In addition, we consider multi-scale information as explicit visual prompts, providing multiscale template features to enhance the EVPTrack's ability to handle target scale changes. Extensive experimental results on six benchmarks (i.e., LaSOT, LaSOText, GOT-10k, UAV123, TrackingNet, and TNL2K.) validate that our EVPTrack can achieve competitive performance at a real-time speed by effectively exploiting both spatio-temporal and multi-scale information. Code and models are available at https://github.com/GXNU-ZhongLab/EVPTrack.
Liangtao Shi, Bineng Zhong 0001, Qihua Liang, Ning Li 0044, Shengping Zhang, Xianxian Li
AAAI2
2024 ODTrack: Online Dense Temporal Token Learning for Visual Tracking
abstract
Online contextual reasoning and association across consecutive video frames are critical to perceive instances in visual tracking. However, most current top-performing trackers persistently lean on sparse temporal relationships between reference and search frames via an offline mode. Consequently, they can only interact independently within each image-pair and establish limited temporal correlations. To alleviate the above problem, we propose a simple, flexible and effective video-level tracking pipeline, named ODTrack, which densely associates the contextual relationships of video frames in an online token propagation manner. ODTrack receives video frames of arbitrary length to capture the spatio-temporal trajectory relationships of an instance, and compresses the discrimination features (localization information) of a target into a token sequence to achieve frame-to-frame association. This new solution brings the following benefits: 1) the purified token sequences can serve as prompts for the inference in the next video frame, whereby past information is leveraged to guide future inference; 2) the complex online update strategies are effectively avoided by the iterative propagation of token sequences, and thus we can achieve more efficient model representation and computation. ODTrack achieves a new SOTA performance on seven benchmarks, while running at real-time speed. Code and models are available at https://github.com/GXNU-ZhongLab/ODTrack.
Yaozong Zheng, Bineng Zhong 0001, Qihua Liang, Zhiyi Mo, Shengping Zhang, Xianxian Li
AAAI2
2024 Fourier Priors-Guided Diffusion for Zero-Shot Joint Low-Light Enhancement and Deblurring
abstract
Existing joint low-light enhancement and deblurring methods learn pixel-wise mappings from paired synthetic data, which results in limited generalization in real-world scenes. While some studies explore the rich generative prior of pre-trained diffusion models, they typically rely on the assumed degradation process and cannot handle unknown real-world degradations well. To address these problems, we propose a novel zero-shot framework, FourierDiff, which embeds Fourier priors into a pre-trained diffusion model to harmoniously handle the joint degradation of luminance and structures. FourierDiff is appealing in its relaxed requirements on paired training data and degradation assumptions. The key zero-shot insight is motivated by image characteristics in the Fourier domain: most luminance information concentrates on amplitudes while structure and content information are closely related to phases. Based on this observation, we decompose the sampled results of the reverse diffusion process in the Fourier domain and take advantage of the amplitude of the generative prior to align the enhanced brightness with the distribution of natural images. To yield a sharp and content-consistent enhanced result, we further design a spatial-frequency alternating optimization strategy to progressively refine the phase of the input. Extensive experiments demonstrate the superior effectiveness of the proposed method, especially in real-world scenes. The code is available at https://github.com/aipixel/FourierDiff.
Xiaoqian Lv, Shengping Zhang, Chenyang Wang 0002, Yichen Zheng, Bineng Zhong 0001, Chongyi Li, Liqiang Nie
CVPR5
2024 DiffPerformer: Iterative Learning of Consistent Latent Guidance for Diffusion-Based Human Video Generation
abstract
Existing diffusion models for pose-guided human video generation mostly suffer from temporal inconsistency in the generated appearance and poses due to the inherent randomization nature of the generation process. In this paper, we propose a novel framework, DiffPerformer, to synthesize high-fidelity and temporally consistent human video. Without complex architecture modification or costly training, DiffPerformer finetunes a pre-trained diffusion model on a single video of the target character and introduces an implicit video representation as a proxy to learn temporally consistent guidance for the diffusion model. The guidance is encoded into VAE latent space and an iterative optimization loop is constructed between the implicit video representation and the diffusion model, allowing to harness the smooth property of the implicit video representation and the generative capabilities of the diffusion model in a mutually beneficial way. Moreover, we propose 3D-aware human flow as a temporal constraint during the optimization to explicitly model the correspondence between driving poses and human appearance. This alleviates the mis-alignment between driving poses and target performer and therefore maintains the appearance coherence under various motions. Extensive experiments demonstrate that our method outperforms the state-of-the-art methods. The code is available at https://github.com/aipixel/DiffPerformer.
Chenyang Wang 0002, Zerong Zheng, Tao Yu 0007, Xiaoqian Lv, Bineng Zhong 0001, Shengping Zhang, Liqiang Nie
CVPR5
2024 Autoregressive Queries for Adaptive Tracking with Spatio-Temporal Transformers
abstract
The rich spatio-temporal information is crucial to capture the complicated target appearance variations in visual tracking. However, most top-performing tracking algorithms rely on many hand-crafted components for spatio-temporal information aggregation. Consequently, the spatio-temporal information is far away from being fully explored. To alleviate this issue, we propose an adaptive tracker with spatio-temporal transformers (named AQA-Track), which adopts simple autoregressive queries to effectively learn spatio-temporal information without many hand-designed components. Firstly, we introduce a set of learnable and autoregressive queries to capture the instantaneous target appearance changes in a sliding window fashion. Then, we design a novel attention mechanism for the interaction of existing queries to generate a new query in current frame. Finally, based on the initial target template and learnt autoregressive queries, a spatio-temporal information fusion module (STM) is designed for spatiotemporal formation aggregation to locate a target object. Benefiting from the STM, we can effectively combine the static appearance and instantaneous changes to guide robust tracking. Extensive experiments show that our method significantly improves the tracker's performance on six popular tracking benchmarks: LaSOT, LaSOText, TrackingNet, GOT-10k, TNL2K, and UAV123. Code and models will be https://github.com/orgs/GXNU-ZhongLab.
Jinxia Xie, Bineng Zhong 0001, Zhiyi Mo, Shengping Zhang, Liangtao Shi, Shuxiang Song 0001, Rongrong Ji
CVPR2
2024 Visual Adapt for RGBD Tracking
abstract
Recent RGBD trackers have employed cueing techniques by overlaying Depth modality images as cues onto RGB modality images, which are then fed into the RGB-based model for tracking. However, the direct overlaying interaction method between modalities not only introduces more noise into the feature space but also exhibits the inadaptability of the RGB-based model to mixed-modality inputs. To address these issues, we introduce Visual Adapt for RGBD Tracking (VADT). Specifically, we maintain the input of the RGB-based model as the RGB modality. Additionally, we have devised a fusion module to enable modality interaction between depth and RGB features. Subsequently, a Depth Adapt module has been formulated to facilitate image interaction with the fused features. This module involves cross-attending to the obtained depth-assisted features and the RGB search frame features produced by the RGB-based model’s output. Experimental results indicate that our proposed tracker achieves state-of-the-art results on various RGBD benchmark tests.
Guangtong Zhang, Qihua Liang, Zhiyi Mo, Ning Li 0044, Bineng Zhong 0001
ICASSP5
2024 Diffusion Mask-Driven Visual-language Tracking
Guangtong Zhang, Bineng Zhong 0001, Qihua Liang, Zhiyi Mo, Shuxiang Song 0001
IJCAI2
2024 ICSGD-Momentum: SGD Momentum Based on Inter-gradient Collision
abstract
Deep neural networks (DNNs) are widely used in fields like computer vision and natural language processing. A key component of DNN training is the optimizer. SGD-Momentum is popular in many DNN methodologies, such as ResNet and DenseNet, due to its simplicity and effectiveness. However, its slow convergence rate limits its use. To overcome this, we introduce inter-gradient collision into SGD-Momentum, inspired by the elastic collision model in physics. This new method, called ICSGD-Momentum, aims to improve convergence. We provide theoretical proof of convergence and establish a regret bound for ICSGD-Momentum. Experiments on benchmarks including function optimization, CIFAR-100, ImageNet, Penn Treebank, COCO, and YCB-Video show that ICSGD-Momentum accelerates training and enhances the generalization performance of DNNs compared to optimizers like SGD-Momentum, Adam, Radam, Adabound, and AdaBelief.
Weidong Zou, Weipeng Cao, Yuanqing Xia, Bineng Zhong 0001, Dachuan Li
INDIN4
2024 Dual-stream Multi-modal Interactive Vision-language Tracking
Zhiyi Mo, Guangtong Zhang, Jian Nong, Bineng Zhong 0001, Zhi Li 0017
MMAsia4
2024 Top-Down Cross-Modal Guidance for Robust RGB-T Tracking
abstract
Most RGB-T trackers heavily rely on bottom-up attention and thus overlook top-down cross-modal guidance for learning target features. Consequently, the discriminative power of the learnt target features is weak. To address this issue, we propose a novel RGB-T tracker (called TGTrack) that designs a Top-down Cross-modal Guidance mechanism to learn target features in two stages. In the first stage, our TGTrack effectively generates top-down cross-modal guidance signals with multi-modal encoders-decoders and prior vectors. In the second stage, these signals are transmitted and integrated to improve the discriminative power of our target features by the attention layers of the cross-modal encoders. Moreover, we introduce an Attention-Driven Spatio-Temporal Updater for updating discriminative target features. Through cross-frame attention guidance, it can effectively eliminates irrelevant features within the search region. As a result, our TGTrack can effectively avoid the complex multi-modal fusion modules and thus achieve robust RGB-T tracking. Extensive experiments on three popular RGB-T tracking benchmarks (i.e., LasHeR, RGBT234, and RGBT210) demonstrate that our TGTrack achieves new state-of-the-art performances.
Bineng Zhong 0001, Qihua Liang, Yaozong Zheng, Zhiyi Mo, Shuxiang Song 0001
IEEE Trans. Circuits Syst. Video Technol.2
2024 Toward Modalities Correlation for RGB-T Tracking
abstract
Recently, RGB-T tracking methods have made significant progress, demonstrating remarkable capabilities in addressing the complexities of tracking tasks within demanding environments. However, these methods overlook instability of modal validity in real-world scenarios. This limits the model’s ability to understand the correlation between modalities, thereby hindering the model’s ability to fully leverage the synergistic effects of RGB and TIR. To address this challenge, we propose a novel RGB-T tracking model named MCTrack, from the perspective of leveraging correlation among modalities. First, during the feature extraction stage, we design a novel module based on channel matching modeling to construct bidirectional channel context information flow for two modalities. By leveraging information flow, specific modalities correlation information can be transmitted to two modes, augmenting the correlation between the two modes adaptively. Subsequently, after the feature extraction network, the features of each modality are decoded and transformed to generate more correlated feature representations. During this stage, we extract distinctive and collective features by leveraging the correlation among modalities. Then fusing these features and generated search region features specifically for localization. This aids the model in comprehending the correlation between RGB and TIR under complex scenarios, thereby enhancing its ability to capture and utilize key features. Based on extensive experiments conducted on four popular RGB-T tracking benchmarks, our model demonstrates superior performance, particularly showcasing impressive results on the LasHeR dataset with an achieved Precision of 71.6%.
Xiantao Hu, Bineng Zhong 0001, Qihua Liang, Shengping Zhang, Ning Li 0044, Xianxian Li
IEEE Trans. Circuits Syst. Video Technol.2
2024 Transformer Tracking via Frequency Fusion
abstract
Transformer has achieved impressive progress in visual tracking due to their capability of global modeling, which enables them to learn low-frequency features(i.e., high-level semantic information). However, it seems to overlook the high-frequency features(i.e., low-level texture and edge information) which are crucial to identify different intra-class object instances in the tracking task. To address this issue, we propose a transformer based tracker via frequency fusion perspective that investigated whether high-frequency and low-frequency features can be effectively combined to achieve robust tracking. Specifically, we design a simple yet effective two-stage fusion strategy and use an appropriate frequency fusion strategy in tracking process of each stage so as to make full use of frequency domain information. In the feature extraction stage, we use wavelet decomposition of high-frequency subbands to solve the performance loss caused by the transformer’s catastrophic forgetting of high-frequency information. In the prediction head stage, we use a variety of wavelet decomposition subbands to model the multi-frequency information. The two-stage fusion strategy makes our model extract more balanced and beneficial multi-frequency information, enabling it to effectively capture target texture information and local edge information while also being sensitive to global information. Extensive experiments on six challenging benchmarks (i.e., LaSOT$_{ext}$, UAV123, TNL2K, LaSOT, TrackingNet, and GOT-10k) demonstrates the superior performance of our tracker.
Xiantao Hu, Bineng Zhong 0001, Qihua Liang, Shengping Zhang, Ning Li 0044, Xianxian Li, Rongrong Ji
IEEE Trans. Circuits Syst. Video Technol.2
2024 Hybrid Transformers With Attention-Guided Spatial Embeddings for Makeup Transfer and Removal
abstract
Existing makeup transfer methods typically transfer simple makeup colors in a well-conditioned face image and fail to handle makeup style details (e.g., complicated colors and shapes) and facial occlusion. To address these problems, this paper proposes Hybrid Transformers with Attention-guided Spatial Embeddings (named HT-ASE) for makeup transfer and removal. Specifically, a makeup context extractor adopts makeup context global-local interactions to aggregate the high-level context and low-level detail features of the makeup styles, which obtains the context-aware makeup features that encode the complicated colors and shapes of the makeup styles. A face identity extractor adopts a face identity local interaction to aggregate the identity-relevant features of shallow layers into identity semantic features, which refines the identity features. A spatially similarity-aware fusion network introduces a spatially-adaptive layer-instance normalization with attention-guided spatial embeddings to perform semantic alignment and fusion between the makeup and identity features, yielding precise and robust transfer results even with large spatial misalignment and facial occlusion. Extensive experimental results demonstrate that the proposed method outperforms the state-of-the-art methods, especially in the preservation of makeup style details and handling facial occlusion.
Mingxiu Li, Wei Yu 0002, Qinglin Liu, Zonglin Li 0004, Ru Li 0002, Bineng Zhong 0001, Shengping Zhang
IEEE Trans. Circuits Syst. Video Technol.6
2024 Identity-Aware Variational Autoencoder for Face Swapping
abstract
Face swapping aims to transfer the identity of a source face to a target face image while preserving the target attributes (e.g., facial expression, head pose, illumination, and background). Most existing methods use a face recognition model to extract global features from the source face and directly fuse them with the target to generate a swapping result. However, identity-irrelevant attributes (e.g., hairstyle and facial appearances) contribute a lot to the recognition task, and thus swapping this task-specific feature inevitably interfuses source attributes with target ones. In this paper, we propose an identity-aware variational autoencoder (ID-VAE) based face swapping framework, dubbed VAFSwap, which learns disentangled identity and attribute representations for high-fidelity face swapping. In particular, we overcome the unpaired training barrier of VAE and impose a proxy identity on the latent space by exploiting the weak supervision from an auxiliary image set whose identity is averaged from multiple collected face images. To explicitly guide the identity fusion, we further devise an identity-associated matrix that corresponds different face regions with their identity representations to perform identity-related feature interactions. Finally, we incorporate spatial dimensions into the latent space and exploit the generative priors of a pre-trained face generator, allowing the effective elimination of noticeable swapping artifacts. Extensive experiments on the FaceForensics++ and CelebA-HQ datasets demonstrate that our method outperforms the state-of-the-art significantly.
Zonglin Li 0004, Shengfeng He, Quanling Meng, Shengping Zhang, Bineng Zhong 0001, Rongrong Ji
IEEE Trans. Circuits Syst. Video Technol.6
2024 Robust Tracking via Combing Top-Down and Bottom-Up Attention
abstract
Transformer attention plays an important role in current top-performing trackers. However, it is bottom-up, driven by stimulus and lacks intrinsic prior guidance. This bottom-up attention mechanism leads to an emphasis on all objects in the input images, rather than the task related objects. As a result, the performance of the bottom-up attention based trackers is deteriorated in complicated scenes. To address this issue, we propose a robust tracker that combines bottom-up attention with top-down attention to comply with the existing ViT framework, named TBTrack. TBTrack can not only utilize the existing bottom-up attention mechanisms to model the long-range relationship of input tokens, but also utilize a newly added top-down attention mechanism to pay more attention to task related object and further eliminate interference from similar objects and backgrounds. Specifically, we firstly design a top-down prior generation module using an adaptive learning parameter combined with the template inputs to obtain top-down task guided signals. Then, we inject the prior signals into a bottom-up attention module to obtain a top-down and bottom-up attention combination block (TB-Block). Finally, we stack these TB-Blocks to construct our tracker (TBTrack) with top-down prior guidance capability, which focuses more on the task related object. Through extensive experiments, our TBTrack achieves impressive performance on multiple tracking benchmarks, including GOT-10k, LaSOT, LaSOText, TNL2K, TrackingNet, UAV123 and so on. The code and trained models will be publicly available.
Ning Li 0044, Bineng Zhong 0001, Yaozong Zheng, Qihua Liang, Zhiyi Mo, Shuxiang Song 0001
IEEE Trans. Circuits Syst. Video Technol.2
2024 End-to-End Human Instance Matting
abstract
Human instance matting aims to estimate an alpha matte for each human instance in an image, which is extremely challenging and has rarely been studied so far. Despite some efforts to use instance segmentation to generate a trimap for each instance and apply trimap-based matting methods, the resulting alpha mattes are often inaccurate due to inaccurate segmentation. In addition, this approach is computationally inefficient due to multiple executions of the matting method. To address these problems, this paper proposes a novel End-to-End Human Instance Matting (E2E-HIM) framework for simultaneous multiple instance matting in a more efficient manner. Specifically, a general perception network first extracts image features and decodes instance contexts into latent codes. Then, a united guidance network exploits spatial attention and semantics embedding to generate united semantics guidance, which encodes the locations and semantic correspondences of all instances. Finally, an instance matting network decodes the image features and united semantics guidance to predict all instance-level alpha mattes. In addition, we construct a large-scale human instance matting dataset (HIM-100K) comprising over 100,000 human images with instance alpha matte labels. Experiments on HIM-100K demonstrate the proposed E2E-HIM outperforms the existing methods on human instance matting with 50% lower errors and 5× faster speed (6 instances in a 640 × 640 image). Experiments on the PPM-100, RWP-636, and P3M datasets demonstrate that E2E-HIM also achieves competitive performance on traditional human matting.
Qinglin Liu, Shengping Zhang, Quanling Meng, Bineng Zhong 0001, Peiqiang Liu, Hongxun Yao
IEEE Trans. Circuits Syst. Video Technol.4
2024 Motion-Driven Tracking via End-to-End Coarse-to-Fine Verifying
abstract
Target appearance and motion variations are the primary challenges in visual tracking. To tackle these challenges, top-performing trackers commonly rely on constructing complex appearance or motion models. However, the efficacy of these models in enhancing track performance can be limited by the lack of effective and seamless integration. The utilization of simplistic handcrafted fusion methods may even exacerbate the issue, resulting in a decline in tracking performance. To address this issue, we propose an end-to-end coarse-to-fine verifying approach in our motion-driven tracker. At the coarse level, we developed a motion prediction module (MPM) that efficiently extracts and utilizes motion information by leveraging the differences between adjacent frames. The MPM constructs not only a position prior for the decoder but also hybrid features that combine both motion and appearance. At the fine level, we employ a deformable transformer-based appearance model to accurately verify a local region centered on the predicted locations from the MPM. To further enhance the generalization capability of our tracker, we propose the use of an instance domain discriminator (IDD) during the training phase. This discriminator is based on domain adaptation theory and aims to sharpen the distinction between the target and other instances, thereby improving the robustness of tracking. Experimental results on five popular benchmarks, including GOT10k, LaSOT, TrackingNet, OTB, and VOT, validate the effectiveness of our proposed tracker.
Bineng Zhong 0001, Yan Chen 0017
IEEE Trans. Circuits Syst. Video Technol.2
2024 Positive-Sample-Free Object Tracking via a Soft Constraint
abstract
Most of the existing bounding box-based trackers rely on a classification subnetwork and a regression subnetwork to predict the location and scale of the bounding box. They learn the classification subnetwork by processing each sample individually and applying the suggested classification confidence to produce the final prediction. They typically involve heuristic positive sample configurations, which inevitably introduce mislabelled training samples and therefore deteriorate their tracking performance. Moreover, the parallel prediction of the bounding box position and scale may lead to misalignment of classification and regression. To address these issues,we propose a simple yet effective soft constraint-based tracking framework without positive samples (named SoftCT). SoftCT adaptively senses the target’s pixel position through a soft constraint mechanism, which eliminates potential performance gaps caused by artificially marking the target’s pixel position. In addition, SoftCT computes the state of the bounding box by aggregating such positional information, thereby allowing the tracker to avoid misalignment in classification and regression due to uninformed communication. Specifically, SoftCT directly senses the position of the target pixel and fuses this information into the bounding box prediction, rather than requiring explicit annotation or regression of the target pixel. Extensive experiments on six tracking benchmarks including GOT-10k, TrackingNet, LaSOT, UAV123, LaSOText and TNL2K demonstrate that our tracker achieves state-of-the-art performance, confirming its effectiveness and efficiency.
Jiaxin Ye, Bineng Zhong 0001, Qihua Liang, Shengping Zhang, Xianxian Li, Rongrong Ji
IEEE Trans. Circuits Syst. Video Technol.2
2024 One-Stream Stepwise Decreasing for Vision-Language Tracking
abstract
Based on the fixed language descriptions in the initial frames, a vision-language tracker typically adopts a two-stream model structure to align vision and language features at the feature fusion stages. However, this paradigm may degrade the tracking performance due to inaccurate language descriptions and lacks further modal interaction. To address these issues, we propose a one-stream vision-language model called One-stream Stepwise Decreasing for Vision-Language Tracking (OSDT). Specifically, we first encode the language description using a language encoder. The obtained language features are then combined with visual images and entered jointly into a visual encoder, in which the encoder’s self-attention mechanism is utilized to facilitate more interactions between language and visual features. Moreover, to mitigate the problems caused by inaccurate language descriptions, we design a stepwise decreasing multi-modal interaction framework, in which a Feature Filter Module (FFM) is introduced to select language features that are more relevant to visual information to provide semantic guidance for visual feature extraction. Furthermore, without additional feature fusion modules, our one-stream model framework can efficiently utilize the proposed feature filtering module for feature selection. Consequently, our tracker can achieve fast tracking speed in the vision-language tracking domain compared to existing state-of-the-art methods. We extensively evaluate our tracker on three benchmarks, i.e. TNL2K, LaSOT, and OTB99, demonstrating competing performance compared to state-of-the-art vision-language tracking methods.
Guangtong Zhang, Bineng Zhong 0001, Qihua Liang, Zhiyi Mo, Ning Li 0044, Shuxiang Song 0001
IEEE Trans. Circuits Syst. Video Technol.2
2024 Toward Unified Token Learning for Vision-Language Tracking
abstract
In this paper, we present a simple, flexible and effective vision-language (VL) tracking pipeline, termed MMTrack, which casts VL tracking as a token generation task. Traditional paradigms address VL tracking task indirectly with sophisticated prior designs, making them over-specialize on the features of specific architectures or mechanisms. In contrast, our proposed framework serializes language description and bounding box into a sequence of discrete tokens. In this new design paradigm, all token queries are required to perceive the desired target and directly predict spatial coordinates of the target in an auto-regressive manner. The design without other prior modules avoids multiple sub-tasks learning and hand-designed loss functions, significantly reducing the complexity of VL tracking modeling and allowing our tracker to use a simple cross-entropy loss as unified optimization objective for VL tracking task. Extensive experiments on TNL2K, LaSOT, LaSOT$_{\mathrm{ext}}$and OTB99-Lang benchmarks show that our approach achieves promising results, compared to other state-of-the-arts.
Yaozong Zheng, Bineng Zhong 0001, Qihua Liang, Guorong Li, Rongrong Ji, Xianxian Li
IEEE Trans. Circuits Syst. Video Technol.2
2024 Robust Tracking via Bidirectional Transduction With Mask Information
abstract
In the tracking literature, foreground and background information have been extensively investigated to discriminate a target from its surrounding background. However, both foreground and background possess their own spatial-temporal correlation relationship that provide significant information to separate the target from its surrounding background, which has been usually ignored by existing work. To address this issue, we propose a bidirectional transductive network based tracker, which incorporates long-range spatial-temporal and bidirectional constraints. Specifically, our tracker consists of two modules, namely the mask generation module (MGM) and the transduction attention module (TAM). MGM aggregates long-range interdependencies of a target along the history frames for generating accurate target masks. TAM retrieves back to the history frames to find patches similar to the current frame, which are then forwarded along with the target masks generated by MGM. In this manner, each position in the current frame can determine its own identity, whether belonging to either the background or the foreground, hence accurately distinguishing the target from its distractors. We conduct systematically experiments and achieve state-of-the-art performance on several benchmarks, obtaining 69.2% AO on GOT-10k and 82.1% on TrackingNet.
TianYu Ning, Bineng Zhong 0001, Qihua Liang, Zhenjun Tang, Xianxian Li
IEEE Trans. Multim.2
2024 One-Stream Vision-Language Memory Network for Object Tracking
abstract
Most existing tracking methods try to represent the target by exploiting visual information as much as possible based on the various deep networks. However, the appearance model hardly describes the attribute feature of the target well, which makes the trackers fail to adapt to the complex visual surrounding. In this article, inspired by brain-like intelligence, we propose an One-stream Vision-Language Memory network (OVLM) for object tracking. Firstly, we use the combination of vision and language to build the target model and use the semantic information in the language to compensate for the instability of visual information, making the target model more stable in the face of complex appearance changes. Secondly, to build a more compact target model, we propose a memory token selection mechanism that utilizes linguistic information to eliminate tokens that do not contain target information. Furthermore, to provide better visual information for target modeling, we propose a language-based evaluation method to select high-quality target samples to be stored in the memory. Finally, OVLM achieves a 64.7% success rate on the large-scale tracking benchmark dataset TNL2K, outperforming the previous best result (VLT) by 11.6%. By exposing the possibility of the vision-language memory network, we aim to draw greater attention to it and open up new avenues for vision-language tracking.
Huanlong Zhang, Jianwei Zhang 0014, Tianzhu Zhang 0001, Bineng Zhong 0001
IEEE Trans. Multim.5
2024 Self Supervised Progressive Network for High Performance Video Object Segmentation
abstract
Recently, self-supervised video object segmentation (VOS) has attracted much interest. However, most proxy tasks are proposed to train only a single backbone, which relies on a point-to-point correspondence strategy to propagate masks through a video sequence. Due to its simple pipeline, the performance of the single backbone paradigm is still unsatisfactory. Instead of following the previous literature, we propose our self-supervised progressive network (SSPNet) which consists of a memory retrieval module (MRM) and collaborative refinement module (CRM). The MRM can perform point-to-point correspondence and produce a propagated coarse mask for a query frame through self-supervised pixel-level and frame-level similarity learning. The CRM, which is trained via cycle consistency region tracking, aggregates the reference & query information and learns the collaborative relationship among them implicitly to refine the coarse mask. Furthermore, to learn semantic knowledge from unlabeled data, we also design two novel mask-generation strategies to provide the training data with meaningful semantic information for the CRM. Extensive experiments conducted on DAVIS-17, YouTube- VOS and SegTrack v2 demonstrate that our method surpasses the state-of-the-art self-supervised methods and narrows the gap with the fully supervised methods.
Guorong Li, Dexiang Hong, Kai Xu 0013, Bineng Zhong 0001, Li Su 0003, Zhenjun Han, Qingming Huang
IEEE Trans. Neural Networks Learn. Syst.4
2024 Visualizing and Understanding Patch Interactions in Vision Transformer
abstract
Vision transformer (ViT) has become a leading tool in various computer vision tasks, owing to its unique self-attention mechanism that learns visual representations explicitly through cross-patch information interactions. Despite having good success, the literature seldom explores the explainability of ViT, and there is no clear picture of how the attention mechanism with respect to the correlation across comprehensive patches will impact the performance and what is the further potential. In this work, we propose a novel explainable visualization approach to analyze and interpret the crucial attention interactions among patches for ViT. Specifically, we first introduce a quantification indicator to measure the impact of patch interaction and verify such quantification on attention window design and indiscriminative patches removal. Then, we exploit the effective responsive field of each patch in ViT and devise a window-free transformer (WinfT) architecture accordingly. Extensive experiments on ImageNet demonstrate that the exquisitely designed quantitative method is shown able to facilitate ViT model learning, leading the top-1 accuracy by 4.28% at most. More remarkably, the results on downstream fine-grained recognition tasks further validate the generalization of our proposal.
Jie Ma 0006, Yalong Bai, Bineng Zhong 0001, Wei Zhang 0031, Ting Yao 0003, Tao Mei 0001
IEEE Trans. Neural Networks Learn. Syst.3
2023 Interactive Object Placement with Reinforcement Learning
abstract
Object placement aims to insert a foreground object into a background image with a suitable location and size to create a natural composition. To predict a diverse distribution of placements, existing methods usually establish a one-to-one mapping from random vectors to the placements. However, these random vectors are not interpretable, which prevents users from interacting with the object placement process. To address this problem, we propose an Interactive Object Placement method with Reinforcement Learning, dubbed IOPRE, to make sequential decisions for producing a reasonable placement given an initial location and size of the foreground. We first design a novel action space to flexibly and stably adjust the location and size of the foreground while preserving its aspect ratio. Then, we propose a multi-factor state representation learning method, which integrates composition image features and sinusoidal positional embeddings of the foreground to make decisions for selecting actions. Finally, we design a hybrid reward function that combines placement assessment and the number of steps to ensure that the agent learns to place objects in the most visually pleasing and semantically appropriate location. Experimental results on the OPA dataset demonstrate that the proposed method achieves state-of-the-art performance in terms of plausibility and diversity.
Shengping Zhang, Quanling Meng, Qinglin Liu, Liqiang Nie, Bineng Zhong 0001, Xiaopeng Fan 0001, Rongrong Ji
ICML5
2023 Unambiguous Object Tracking by Exploiting Target Cues
abstract
Siamese tracking exploits the template and the search region features to adaptively locate arbitrary objects in the tracking. A noteworthy issue is that both foreground and background mix in the template, and thus a tracker needs to learn what the target is and which pixels belong to it. However, existing trackers cannot effectively exploit the template information, resulting in a deficiency of target information and causing confusion for the tracker regarding which pixels belong to the target. To alleviate this issue, we propose UTrack, a simple and effective algorithm for unambiguous object tracking. UTrack utilizes long-term contextual information to propagate the appearance state of the target so as to explicitly model the apparent information of the target. Additionally, UTrack can resist the appearance change of the target by leveraging the target cues. Moreover, the proposed method uses the refined template to obtain more detailed information about the target and better understand which pixels belong to the target. Extensive experiments and comparisons with competitive trackers on challenging large-scale benchmarks show that our tracker can achieve state-of-the-art performances with real-time running. In particular, UTrack achieves 77.7% AO on GOT-10k.
Jie Gao 0021, Bineng Zhong 0001, Yan Chen 0017
ACM Multimedia2
2023 Robust Tracking via Unifying Pretrain-Finetuning and Visual Prompt Tuning
abstract
The finetuning paradigm has been a widely used methodology for the supervised training of top-performing trackers. However, the finetuning paradigm faces one key issue: it is unclear how best to perform the finetuning method to adapt a pretrained model to tracking tasks while alleviating the catastrophic forgetting problem. To address this problem, we propose a novel partial finetuning paradigm for visual tracking via unifying pretrain-finetuning and visual prompt tuning (named UPVPT), which can not only efficiently learn knowledge from the tracking task but also reuse the prior knowledge learned by the pre-trained model for effectively handling various challenges in tracking task. Firstly, to maintain the pre-trained prior knowledge, we design a Prompt-style method to freeze some parameters of the pretrained network. Then, to learn knowledge from the tracking task, we update the parameters of the prompt and MLP layers. As a result, we cannot only retain useful prior knowledge of the pre-trained model by freezing the backbone network but also effectively learn target domain knowledge by updating the Prompt and MLP layer. Furthermore, the proposed UPVPT can easily be embedded into existing Transformer trackers (e.g., OSTracker and SwinTracker) by adding only a small number of model parameters (less than 1% of a Backbone network). Extensive experiments on five tracking benchmarks (i.e., UAV123, GOT-10k, LaSOT, TNL2K, and TrackingNet) demonstrate that the proposed UPVPT can improve the robustness and effectiveness of the model, especially in complex scenarios.
Guangtong Zhang, Qihua Liang, Ning Li 0044, Zhiyi Mo, Bineng Zhong 0001
MMAsia5
2023 SpectralTracker: Jointly High and Low-Frequency Modeling for Tracking
Yimin Rong, Qihua Liang, Ning Li 0044, Zhiyi Mo, Bineng Zhong 0001
PRCV (12)5
2023 SiamBAN: Target-Aware Tracking With Siamese Box Adaptive Network
abstract
Variation of scales or aspect ratios has been one of the main challenges for tracking. To overcome this challenge, most existing methods adopt either multi-scale search or anchor-based schemes, which use a predefined search space in a handcrafted way and therefore limit their performance in complicated scenes. To address this problem, recent anchor-free based trackers have been proposed without using prior scale or anchor information. However, an inconsistency problem between classification and regression degrades the tracking performance. To address the above issues, we propose a simple yet effective tracker (named Siamese Box Adaptive Network, SiamBAN) to learn a target-aware scale handling schema in a data-driven manner. Our basic idea is to predict the target boxes in a per-pixel fashion through a fully convolutional network, which is anchor-free. Specifically, SiamBAN divides the tracking problem into classification and regression tasks, which directly predict objectiveness and regress bounding boxes, respectively. A no-prior box design is proposed to avoid tuning hyper-parameters related to candidate boxes, which makes SiamBAN more flexible. SiamBAN further uses a target-aware branch to address the inconsistency problem. Experiments on benchmarks including VOT2018, VOT2019, OTB100, UAV123, LaSOT and TrackingNet show that SiamBAN achieves promising performance and runs at 35 FPS.
Zedu Chen, Bineng Zhong 0001, Guorong Li, Shengping Zhang, Rongrong Ji, Zhenjun Tang, Xianxian Li
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Robust Tracking via Learning Model Update With Unsupervised Anomaly Detection Philosophy
abstract
Template tracking is a typical paradigm to adaptively locate arbitrary objects in the tracking literature. Although existing works present diverse template updating approaches, one of the essential problems of template updating has not been solved effectively, i.e., when and how to update a template. In this work, we treat the updating time as an abnormal moment that indicates the previous template cannot depict the target accurately any more. Thus, we introduce an effective State-Edge Awareness (SEA) module that detect such abnormal moments via unsupervised anomaly detection. To be specific, by retaining multi search frames of a video, SEA firstly analysis the correlation features that generated by the template and search images. Then, it estimates the measurement for abnormal degree that is regarded as the sign for template updating. As a result, our method can not only capture the updating time automatically, but also update the templates effectively. Furthermore, the effectiveness of the proposed method has been verified on a representative CNN-based and Transformer-based tracker, respectively. The experimental results on five popular benchmarks show that our tracker can achieve the state-of-the-art performance.
Jie Gao 0021, Bineng Zhong 0001, Yan Chen 0017
IEEE Trans. Circuits Syst. Video Technol.2
2023 Robust Tracking via Uncertainty-Aware Semantic Consistency
abstract
Robust tracking has a variety of practical applications. Despite many years of progress, it is still a difficult problem due to enormous uncertainties in real-world scenes. To address this issue, we propose a robust anchor-free based tracking model with uncertainty estimation. Within the model, a new data-driven uncertainty estimation strategy is proposed to generate uncertainty-aware features with promising discriminative and descriptive power. Then, a simple yet effective pyramid-wise cross correlation operation is constructed to extract multi-scale semantic features that provide rich correlation information for uncertainty-aware estimation and thus enhances the tracking robustness. Finally, a semantic consistency checking branch is designed to further estimate uncertainty of output results from the classification and regression branches by adaptively generating semantically consistent labels. Experiments on six benchmarks (i.e., OTB100, VOT2018, VOT2020, TrackingNet, GOT-10K and LaSOT) show the competing performance of our tracker with 130 FPS.
Jie Ma 0006, Xiangyuan Lan, Bineng Zhong 0001, Guorong Li, Zhenjun Tang, Xianxian Li, Rongrong Ji
IEEE Trans. Circuits Syst. Video Technol.3
2023 Leveraging Local and Global Cues for Visual Tracking via Parallel Interaction Network
abstract
Despite that both local and context information are crucial for robust tracking, existing CNN-based and transformer-based methods mainly focus on one of these aspects. Consequently, the former fails to exploit rich global context information due to the limited receptive field, while the latter suffers from the deficiencies in constructing the local relationship among neighboring regions. To address this issue, we propose the SiamPIN tracker, based on our Parallel Interaction Network. It consists of two effective modules, namely Global Aggregation Block (GAB) and Local Process Block (LPB). GAB perceives the global context to capture the long-range spatial dependency through a transformer-based architecture. Meanwhile, LPB performs local information extraction using a CNN model to retain the detailed appearance information of the target. These two modules are connected consecutively to compose a Trans-Conv unit block, which transmits the global context information to the local feature extraction procedure, hence enables the interaction of global-local information flow. Several such blocks are cascaded so that our model can learn to aggregate local and context information interactively. The proposed tracker achieves state-of-the-art performance on six benchmark datasets, while maintaining a real time running speed.
Yaozong Zheng, Bineng Zhong 0001, Qihua Liang, Zhenjun Tang, Rongrong Ji, Xianxian Li
IEEE Trans. Circuits Syst. Video Technol.2
2023 Robust Long-Term Tracking via Localizing Occluders
abstract
Occlusion is known as one of the most challenging factors in long-term tracking because of its unpredictable shape. Existing works devoted into the design of loss functions, training strategies or model architectures, which are considered to have not directly touched the key point. Alternatively, we came up with a direct and natural idea that is discarding things that covers the target. We propose a novel occluder-aware representation learning framework to develop this idea. First, we design a local occluders detection module (LODM) to localize the occluders, which works on the principle that discriminates the non-noumenal part from a target based on the general knowledge of this category. An extra dataset and a clustering strategy is proposed to support this general knowledge. Second, we devise a feature reconstruction module to guide the occluder-aware representation learning. With the help of above methods, our localizing occluders tracker, called LOTracker, can learn an occluder-free representation and promote the performance that tracks with occlusion scenarios. Extensive experimental results show that our LOTracker achieves a state-of-the-art performance in multiple benchmarks such as LaSOT, VOTLT2018, VOTLT2019, and OxUvALT.
Binfei Chu, Bineng Zhong 0001, Zhenjun Tang, Xianxian Li, Jing Wang 0049
ACM Trans. Multim. Comput. Commun. Appl.3
2023 Unifying Dual-Attention and Siamese Transformer Network for Full-Reference Image Quality Assessment
abstract
Image Quality Assessment (IQA) is a critical task of computer vision. Most Full-Reference (FR) IQA methods have limitation in the accurate prediction of perceptual qualities of the traditional distorted images and the Generative Adversarial Networks (GANs) based distorted images. To address this issue, we propose a novel method by Unifying Dual-Attention and Siamese Transformer Network (UniDASTN) for FR-IQA. An important contribution is the spatial attention module composed of a Siamese Transformer Network and a feature fusion block. It can focus on significant regions and effectively maps the perceptual differences between the reference and distorted images to a latent distance for distortion evaluation. Another contribution is the dual-attention strategy that exploits channel attention and spatial attention to aggregate features for enhancing distortion sensitivity. In addition, a novel loss function is designed by jointly exploiting Mean Square Error (MSE), bidirectional Kullback–Leibler divergence, and rank order of quality scores. The designed loss function can offer stable training and thus enables the proposed UniDASTN to effectively learn visual perceptual image quality. Extensive experiments on standard IQA databases are conducted to validate the effectiveness of the proposed UniDASTN. The IQA results demonstrate that the proposed UniDASTN outperforms some state-of-the-art FR-IQA methods on the LIVE, CSIQ, TID2013, and PIPAL databases.
Zhenjun Tang, Zhixin Li 0001, Bineng Zhong 0001, Xianquan Zhang, Xinpeng Zhang 0001
ACM Trans. Multim. Comput. Commun. Appl.4
2022 BacklitNet: A dataset and network for backlit image enhancement
Xiaoqian Lv, Shengping Zhang, Qinglin Liu, Haozhe Xie, Bineng Zhong 0001, Huiyu Zhou 0001
Comput. Vis. Image Underst.5
2022 Teacher-student knowledge distillation for real-time correlation tracking
Qihuang Chen, Bineng Zhong 0001, Qihua Liang, Qingyong Deng, Xianxian Li
Neurocomputing2
2022 Perceptual Hashing With Complementary Color Wavelet Transform and Compressed Sensing for Reduced-Reference Image Quality Assessment
abstract
Image quality assessment (IQA) is an important task of image processing and has diverse applications, such as image super-resolution reconstruction, image transmission and monitoring systems. This paper proposes a perceptual hashing algorithm with complementary color wavelet transform (CCWT) and compressed sensing (CS) for reduced-reference (RR) IQA. The CCWT is exploited to decompose input color image into different sub-bands. Since the calculation of CCWT uses all color channels without discarding any information, the distortions introduced by digital operations on color channels are preserved in the CCWT sub-bands. The block-based CS is used to extract features from the CCWT sub-bands. As the Euclidean distance between the block-based CS features is slightly influenced by content-preserving operations, perceptual features constructed by Euclidean distances are robust, discriminative and compact. Hash sequence is finally determined by quantifying the perceptual features. Effectiveness of the proposed hashing is verified by various experiments on four open image databases. Experimental results demonstrate that the proposed hashing is superior to some state-of-the-art algorithms in terms of classification and RR IQA application.
Mengzhu Yu, Zhenjun Tang, Xianquan Zhang, Bineng Zhong 0001, Xinpeng Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2022 Large-Scale Subspace Clustering by Independent Distributed and Parallel Coding
abstract
Subspace clustering is a popular method to discover underlying low-dimensional structures of high-dimensional multimedia data (e.g., images, videos, and texts). In this article, we consider a large-scale subspace clustering (LS2C) problem, that is, partitioning million data points with a millon dimensions. To address this, we explore an independent distributed and parallel framework by dividing big data/variable matrices and regularization by both columns and rows. Specifically, LS2C is independently decomposed into many subproblems by distributing those matrices into different machines by columns since the regularization of the code matrix is equal to a sum of that of its submatrices (e.g., square-of-Frobenius/$\ell _{1}$-norm). Consensus optimization is designed to solve these subproblems in a parallel way for saving communication costs. Moreover, we provide theoretical guarantees that LS2C can recover consensus subspace representations of high-dimensional data points under broad conditions. Compared with the state-of-the-art LS2C methods, our approach achieves better clustering results in public datasets, including a million images and videos.
Jun Li 0027, Zhiqiang Tao, Yue Wu 0008, Bineng Zhong 0001, Yun Fu 0001
IEEE Trans. Cybern.4
2022 Representative Task Self-Selection for Flexible Clustered Lifelong Learning
abstract
Consider the lifelong machine learning paradigm whose objective is to learn a sequence of tasks depending on previous experiences, e.g., knowledge library or deep network weights. However, the knowledge libraries or deep networks for most recent lifelong learning models are of prescribed size and can degenerate the performance for both learned tasks and coming ones when facing with a new task environment (cluster). To address this challenge, we propose a novel incremental clustered lifelong learning framework with two knowledge libraries: feature learning library and model knowledge library, called Flexible Clustered Lifelong Learning (FCL3). Specifically, the feature learning library modeled by an autoencoder architecture maintains a set of representation common across all the observed tasks, and the model knowledge library can be self-selected by identifying and adding new representative models (clusters). When a new task arrives, our FCL3 model firstly transfers knowledge from these libraries to encode the new task, i.e., effectively and selectively soft-assigning this new task to multiple representative models over feature learning library. Then: 1) the new task with a higher outlier probability will be judged as a new representative, and used to redefine both feature learning library and representative models over time; or 2) the new task with lower outlier probability will only refine the feature learning library. For model optimization, we cast this lifelong learning problem as an alternating direction minimization problem as a new task comes. Finally, we evaluate the proposed framework by analyzing several multitask data sets, and the experimental results demonstrate that our FCL3 model can achieve better performance than most lifelong learning frameworks, even batch clustered multitask learning models.
Gan Sun, Yang Cong, Qianqian Wang 0001, Bineng Zhong 0001, Yun Fu 0001
IEEE Trans. Neural Networks Learn. Syst.4
2022 Attention-Based Neural Architecture Search for Person Re-Identification
abstract
Recent years have witnessed significant progress of person reidentification (reID) driven by expert-designed deep neural network architectures. Despite the remarkable success, such architectures often suffer from high model complexity and time-consuming pretraining process, as well as the mismatches between the image classification-driven backbones and the reID task. To address these issues, we introduce neural architecture search (NAS) into automatically designing person reID backbones, i.e., reID-NAS, which is achieved via automatically searching attention-based network architectures from scratch. Different from traditional NAS approaches that originated for image classification, we design a reID-based search space as well as a search objective to fit NAS for the reID tasks. In terms of the search space, reID-NAS includes a lightweight attention module to precisely locate arbitrary pedestrian bounding boxes, which is automatically added as attention to the reID architectures. In terms of the search objective, reID-NAS introduces a new retrieval objective to search and train reID architectures from scratch. Finally, we propose a hybrid optimization strategy to improve the search stability in reID-NAS. In our experiments, we validate the effectiveness of different parts in reID-NAS, and show that the architecture searched by reID-NAS achieves a new state of the art, with one order of magnitude fewer parameters on three-person reID datasets. As a concomitant benefit, the reliance on the pretraining process is vastly reduced by reID-NAS, which facilitates one to directly search and train a lightweight reID model from scratch.
Qinqin Zhou 0001, Bineng Zhong 0001, Xin Liu 0011, Rongrong Ji
IEEE Trans. Neural Networks Learn. Syst.2
2021 Learning To Filter: Siamese Relation Network for Robust Tracking
abstract
Despite the great success of Siamese-based trackers, their performance under complicated scenarios is still not satisfying, especially when there are distractors. To this end, we propose a novel Siamese relation network, which introduces two efficient modules, i.e. Relation Detector (RD) and Refinement Module (RM). RD performs in a meta-learning way to obtain a learning ability to filter the distractors from the background while RM aims to effectively integrate the proposed RD into the Siamese framework to generate accurate tracking result. Moreover, to further improve the discriminability and robustness of the tracker, we introduce a contrastive training strategy that attempts not only to learn matching the same target but also to learn how to distinguish the different objects. Therefore, our tracker can achieve accurate tracking results when facing background clutters, fast motion, and occlusion. Experimental results on five popular benchmarks, including VOT2018, VOT2019, OTB100, LaSOT, and UAV123, show that the proposed method is effective and can achieve state-of-the-art results. The code will be available at https://github.com/hqucv/siamrn
Siyuan Cheng 0003, Bineng Zhong 0001, Guorong Li, Xin Liu 0011, Zhenjun Tang, Xianxian Li, Jing Wang 0049
CVPR2
2021 Discover Cross-Modality Nuances for Visible-Infrared Person Re-Identification
abstract
Visible-infrared person re-identification (Re-ID) aims to match the pedestrian images of the same identity from different modalities. Existing works mainly focus on alleviating the modality discrepancy by aligning the distributions of features from different modalities. However, nuanced but discriminative information, such as glasses, shoes, and the length of clothes, has not been fully explored, especially in the infrared modality. Without discovering nuances, it is challenging to match pedestrians across modalities using modality alignment solely, which inevitably reduces feature distinctiveness. In this paper, we propose a joint Modality and Pattern Alignment Network (MPANet) to discover cross-modality nuances in different patterns for visible-infrared person Re-ID, which introduces a modality alleviation module and a pattern alignment module to jointly extract discriminative features. Specifically, we first propose a modality alleviation module to dislodge the modality information from the extracted feature maps. Then, We devise a pattern alignment module, which generates multiple pattern maps for the diverse patterns of a person, to discover nuances. Finally, we introduce a mutual mean learning fashion to alleviate the modality discrepancy and propose a center cluster loss to guide both identity learning and nuances discovering. Extensive experiments on the public SYSU-MM01 and RegDB datasets demonstrate the superiority of MPANet over state-of-the-arts.
Qiong Wu 0012, Pingyang Dai, Jie Chen 0001, Chia-Wen Lin, Yongjian Wu 0001, Feiyue Huang, Bineng Zhong 0001, Rongrong Ji
CVPR7
2021 Distractor-Aware Fast Tracking via Dynamic Convolutions and MOT Philosophy
abstract
A practical long-term tracker typically contains three key properties, i.e. an efficient model design, an effective global re-detection strategy and a robust distractor awareness mechanism. However, most state-of-the-art long-term trackers (e.g., Pseudo and re-detecting based ones) do not take all three key properties into account and therefore may either be time-consuming or drift to distractors. To address the issues, we propose a two-task tracking framework (named DMTrack), which utilizes two core components (i.e., one-shot detection and re-identification (re-id) association) to achieve distractor-aware fast tracking via Dynamic convolutions (d-convs) and Multiple object tracking (MOT) philosophy. To achieve precise and fast global detection, we construct a lightweight one-shot detector using a novel dynamic convolutions generation method, which provides a unified and more flexible way for fusing target information into the search field. To distinguish the target from distractors, we resort to the philosophy of MOT to reason distractors explicitly by maintaining all potential similarities’ tracklets. Benefited from the strength of high recall detection and explicit object association, our tracker achieves state-of-the-art performance on the LaSOT, Ox-UvA, TLP, VOT2018LT and VOT2019LT benchmarks and runs in real-time (3x faster than comparisons)1.
Zikai Zhang 0003, Bineng Zhong 0001, Shengping Zhang, Zhenjun Tang, Xin Liu 0011, Zhaoxiang Zhang 0001
CVPR2
2021 EC-DARTS: Inducing Equalized and Consistent Optimization into DARTS
abstract
Based on the relaxed search space, differential architecture search (DARTS) is efficient in searching for a high-performance architecture. However, the unbalanced competition among operations that have different trainable parameters causes the model collapse. Besides, the inconsistent structures in the search and retraining stages causes cross-stage evaluation to be unstable. In this paper, we call these issues as an operation gap and a structure gap in DARTS. To shrink these gaps, we propose to induce equalized and consistent optimization in differentiable architecture search (EC-DARTS). EC-DARTS decouples different operations based on their categories to optimize the operation weights so that the operation gap between them is shrinked. Besides, we introduce an induced structural transition to bridge the structure gap between the model structures in the search and retraining stages. Extensive experiments on CIFAR10 and ImageNet demonstrate the effectiveness of our method. Specifically, on CIFAR10, we achieve a test error of 2.39%, while only 0.3 GPU days on NVIDIA TITAN V. On ImageNet, our method achieves a top-1 error of 23.6% under the mobile setting.
Qinqin Zhou 0001, Xiawu Zheng, Liujuan Cao, Bineng Zhong 0001, Teng Xi, Errui Ding, Mingliang Xu 0001, Rongrong Ji
ICCV4
2021 Long-Range Feature Propagating for Natural Image Matting
abstract
Natural image matting estimates the alpha values of unknown regions in the trimap. Recently, deep learning based methods propagate the alpha values from the known regions to unknown regions according to the similarity between them. However, we find that more than 50% pixels in the unknown regions cannot be correlated to pixels in known regions due to the limitation of small effective reception fields of common convolutional neural networks, which leads to inaccurate estimation when the pixels in the unknown regions cannot be inferred only with pixels in the reception fields. To solve this problem, we propose Long-Range Feature Propagating Network (LFPNet), which learns the long-range context features outside the reception fields for alpha matte estimation. Specifically, we first design the propagating module which extracts the context features from the downsampled image. Then, we present Center-Surround Pyramid Pooling (CSPP) that explicitly propagates the context features from the surrounding context image patch to the inner center image patch. Finally, we use the matting module which takes the image, trimap and context features to estimate the alpha matte. Experimental results demonstrate that the proposed method performs favorably against the state-of-the-art methods on the AlphaMatting and Adobe Image Matting datasets.
Qinglin Liu, Haozhe Xie, Shengping Zhang, Bineng Zhong 0001, Rongrong Ji
ACM Multimedia4
2021 Real-time video dehazing via incremental transmission learning and spatial-temporally coherent regularization
Shu-Juan Peng, Xin Liu 0011, Wentao Fan 0001, Bineng Zhong 0001, Jixiang Du
Neurocomputing5
2021 Residual Dense Network for Image Restoration
abstract
Recently, deep convolutional neural network (CNN) has achieved great success for image restoration (IR) and provided hierarchical features at the same time. However, most deep CNN based IR models do not make full use of the hierarchical features from the original low-quality images; thereby, resulting in relatively-low performance. In this work, we propose a novel and efficient residual dense network (RDN) to address this problem in IR, by making a better tradeoff between efficiency and effectiveness in exploiting the hierarchical features from all the convolutional layers. Specifically, we propose residual dense block (RDB) to extract abundant local features via densely connected convolutional layers. RDB further allows direct connections from the state of preceding RDB to all the layers of current RDB, leading to a contiguous memory mechanism. To adaptively learn more effective features from preceding and current local features and stabilize the training of wider network, we proposed local feature fusion in RDB. After fully obtaining dense local features, we use global feature fusion to jointly and adaptively learn global hierarchical features in a holistic way. We demonstrate the effectiveness of RDN with several representative IR applications, single image super-resolution, Gaussian image denoising, image compression artifact reduction, and image deblurring. Experiments on benchmark and real-world datasets show that our RDN achieves favorable performance against state-of-the-art methods for each IR task quantitatively and visually.
Yulun Zhang 0001, Yapeng Tian, Yu Kong 0001, Bineng Zhong 0001, Yun Fu 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2020 Siamese Box Adaptive Network for Visual Tracking
abstract
Most of the existing trackers usually rely on either a multi-scale searching scheme or pre-defined anchor boxes to accurately estimate the scale and aspect ratio of a target. Unfortunately, they typically call for tedious and heuristic configurations. To address this issue, we propose a simple yet effective visual tracking framework (named Siamese Box Adaptive Network, SiamBAN) by exploiting the expressive power of the fully convolutional network (FCN). SiamBAN views the visual tracking problem as a parallel classification and regression problem, and thus directly classifies objects and regresses their bounding boxes in a unified FCN. The no-prior box design avoids hyper-parameters associated with the candidate boxes, making SiamBAN more flexible and general. Extensive experiments on visual tracking benchmarks including VOT2018, VOT2019, OTB100, NFS, UAV123, and LaSOT demonstrate that SiamBAN achieves state-of-the-art performance and runs at 40 FPS, confirming its effectiveness and efficiency. The code will be available at https://github.com/hqucv/siamban.
Zedu Chen, Bineng Zhong 0001, Guorong Li, Shengping Zhang, Rongrong Ji
CVPR2
2020 What Can Be Transferred: Unsupervised Domain Adaptation for Endoscopic Lesions Segmentation
abstract
Unsupervised domain adaptation has attracted growing research attention on semantic segmentation. However, 1) most existing models cannot be directly applied into lesions transfer of medical images, due to the diverse appearances of same lesion among different datasets; 2) equal attention has been paid into all semantic representations instead of neglecting irrelevant knowledge, which leads to negative transfer of untransferable knowledge. To address these challenges, we develop a new unsupervised semantic transfer model including two complementary modules (i.e., T_D and T_F ) for endoscopic lesions segmentation, which can alternatively determine where and how to explore transferable domain-invariant knowledge between labeled source lesions dataset (e.g., gastroscope) and unlabeled target diseases dataset (e.g., enteroscopy). Specifically, T_D focuses on where to translate transferable visual information of medical lesions via residual transferability-aware bottleneck, while neglecting untransferable visual characterizations. Furthermore, T_F highlights how to augment transferable semantic features of various lesions and automatically ignore untransferable representations, which explores domain-invariant knowledge and in return improves the performance of T_D. To the end, theoretical analysis and extensive experiments on medical endoscopic dataset and several non-medical public datasets well demonstrate the superiority of our proposed model.
Jiahua Dong 0001, Yang Cong, Gan Sun, Bineng Zhong 0001, Xiaowei Xu 0001
CVPR4
2020 Projection & Probability-Driven Black-Box Attack
abstract
Generating adversarial examples in a black-box setting retains a significant challenge with vast practical application prospects. In particular, existing black-box attacks suffer from the need for excessive queries, as it is non-trivial to find an appropriate direction to optimize in the high-dimensional space. In this paper, we propose Projection & Probability-driven Black-box Attack (PPBA) to tackle this problem by reducing the solution space and providing better optimization. For reducing the solution space, we first model the adversarial perturbation optimization problem as a process of recovering frequency-sparse perturbations with compressed sensing, under the setting that random noise in the low-frequency space is more likely to be adversarial. We then propose a simple method to construct a low-frequency constrained sensing matrix, which works as a plug-and-play projection matrix to reduce the dimensionality. Such a sensing matrix is shown to be flexible enough to be integrated into existing methods like NES and BanditsTD. For better optimization, we perform a random walk with a probability-driven strategy, which utilizes all queries over the whole progress to make full use of the sensing matrix for a less query budget. Extensive experiments show that our method requires at most 24% fewer queries with a higher attack success rate compared with state-of-the-art approaches. Finally, the attack method is evaluated on the real-world online service, i.e., Google Cloud Vision API, which further demonstrates our practical potentials.
Jie Li 0052, Rongrong Ji, Hong Liu 0009, Jianzhuang Liu, Bineng Zhong 0001, Cheng Deng 0002, Qi Tian 0001
CVPR5
2020 Hearing like Seeing: Improving Voice-Face Interactions and Associations via Adversarial Deep Semantic Matching Network
abstract
Many cognitive researches have shown that human may 'see voices' or 'hear faces', and such ability can be potentially associated by machine vision and intelligence. However, this research is still under early stage. In this paper, we present a novel adversarial deep semantic matching network for efficient voice-face interactions and associations, which can well learn the correspondence between voices and faces for various cross-modal matching and retrieval tasks. Within the proposed framework, we exploit a simple and efficient adversarial learning architecture to learn the cross-modal embeddings between faces and voices, which consists of two subnetworks, respectively, for generator and discriminator. The former subnetwork is designed to adaptively discriminate the high-level semantical features between voices and faces, in which the triplet loss and multi-modal center loss are in tandem utilized to explicitly regularize the correspondences among them. The latter subnetwork is further leveraged to maximally bridge the semantic gap between the representations of voice and face data, featuring on maintaining the semantic consistency. Through the joint exploitation of the above, the proposed framework can well push representations of voice-face data from the same person closer while pulling those representations of different person away. Extensive experiments empirically show that the proposed approach involves fewer parameters and calculations, adapts various cross-modal matching tasks for voice-face data and brings substantial improvements over the state-of-the-art methods.
Xin Liu 0011, Yiu-Ming Cheung, Xing Xu 0001, Bineng Zhong 0001
ACM Multimedia6
2020 A Cooperative Tracker by Fusing Correlation Filter and Siamese Network
Bin Zhou 0004, Xin Liu 0011, Bineng Zhong 0001
PRCV (2)3
2020 Semi-supervised discrete hashing for efficient cross-modal retrieval
Xingzhi Wang, Xin Liu 0011, Shu-Juan Peng, Bineng Zhong 0001, Yewang Chen, Jixiang Du
Multim. Tools Appl.4
2020 Conditional GAN based individual and global motion fusion for multiple object tracking in UAV videos
Hongyang Yu 0001, Guorong Li, Li Su 0003, Bineng Zhong 0001, Hongxun Yao, Qingming Huang
Pattern Recognit. Lett.4
2020 Fine-Grained Spatial Alignment Model for Person Re-Identification With Focal Triplet Loss
abstract
Recent advances of person re-identification have well advocated the usage of human body cues to boost performance. However, most existing methods still retain on exploiting a relatively coarse-grained local information. Such information may include redundant backgrounds that are sensitive to the apparently similar persons when facing challenging scenarios like complex poses, inaccurate detection, occlusion and misalignment. In this paper we propose a novel Fine-Grained Spatial Alignment Model (FGSAM) to mine fine-grained local information to handle the aforementioned challenge effectively. In particular, we first design a pose resolve net with channel parse blocks (CPB) to extract pose information in pixel-level. This network allows the proposed model to be robust to complex pose variations while suppressing the redundant backgrounds caused by inaccurate detection and occlusion. Given the extracted pose information, a locally reinforced alignment mode is further proposed to address the misalignment problem between different local parts by considering different local parts along with attribute information in a fine-grained way. Finally, a focal triplet loss is designed to effectively train the entire model, which imposes a constraint on the intra-class and an adaptively weight adjustment mechanism to handle the hard sample problem. Extensive evaluations and analysis on Market1501, DukeMTMC-reid and PETA datasets demonstrate the effectiveness of FGSAM in coping with the problems of misalignment, occlusion and complex poses.
Qinqin Zhou 0001, Bineng Zhong 0001, Xiangyuan Lan, Gan Sun, Yulun Zhang 0001, Baochang Zhang 0001, Rongrong Ji
IEEE Trans. Image Process.2
2020 Corse-to-Fine Road Extraction Based on Local Dirichlet Mixture Models and Multiscale-High-Order Deep Learning
abstract
Road extraction from remote sensing images is an attractive but difficult task. Gray-value distribution and structure feature information are both crucial for road extraction task. However, existing methods mainly focus on structure feature information which contains morphological shape features and machine learning features, suffering from lots of false positives which are generated at positions having similar structure features but different gray-value distribution with roads. To effectively fuse the two complementary gray-value distribution and structure feature information, we propose a coarse-to-fine road extraction algorithm from remote sensing images. First, at the coarse level, we introduce a local Dirichlet mixture models (LDMM) which utilizing gray-value distribution information to pre-segment images into potential roads and backgrounds. Thus, most backgrounds having different gray-value distribution with roads can be removed firstly. Compared with original Dirichlet mixture models, the LDMM is much faster and more accurate. Next, at the fine level, we introduce a multiscale-high-order deep learning strategy based on ResNet model which can learn robust structure context features for final road extraction step. Based on the results of LDMM, the multiscale-high-order strategy can further remove false positives which have different structure features with roads. Compared with a single scanning size ResNet, our multiscale-high-order strategy can learn higher-order context information, leading to better performances. We test our algorithm on Shaoshan dataset. Experiments illustrate our better performance compared with other six state-of-the-art methods.
Ziyi Chen 0001, Wentao Fan 0001, Bineng Zhong 0001, Jonathan Li 0001, Jixiang Du, Cheng Wang 0003
IEEE Trans. Intell. Transp. Syst.3
2019 Structured and Sparse Annotations for Image Emotion Distribution Learning
abstract
Label distribution learning methods effectively address the label ambiguity problem and have achieved great success in image emotion analysis. However, these methods ignore structured and sparse information naturally contained in the annotations of emotions. For example, emotions can be grouped and ordered due to their polarities and degrees. Meanwhile, emotions have the character of intensity and are reflected in different levels of sparse annotations. Motivated by these observations, we present a convolutional neural network based framework called Structured and Sparse annotations for image emotion Distribution Learning (SSDL) to tackle two challenges. In order to utilize structured annotations, the Earth Mover’s Distance is employed to calculate the minimal cost required to transform one distribution to another for ordered emotions and emotion groups. Combined with Kullback-Leibler divergence, we design the loss to penalize the mispredictions according to the dissimilarities of same emotions and different emotions simultaneously. Moreover, in order to handle sparse annotations, sparse regularization based on emotional intensity is adopted. Through combined loss and sparse regularization, SSDL could effectively leverage structured and sparse annotations for predicting emotion distribution. Experiment results demonstrate that our proposed SSDL significantly outperforms the state-of-the-art methods.
Haitao Xiong, Hongfu Liu 0001, Bineng Zhong 0001, Yun Fu 0001
AAAI3
2019 Residual Non-local Attention Networks for Image Restoration
Yulun Zhang 0001, Kai Li 0012, Bineng Zhong 0001, Yun Fu 0001
ICLR (Poster)4
2019 LRDNN: Local-refining based Deep Neural Network for Person Re-Identification with Attribute Discerning
abstract
Recently, pose or attribute information has been widely used to solve person re-identification (re-ID) problem. However, the inaccurate output from pose or attribute modules will impair the final person re-ID performance. Since re-ID, pose estimation and attribute recognition are all based on the person appearance information, we propose a Local-refining based Deep Neural Network (LRDNN) to aggregate pose estimation and attribute recognition to improve the re-ID performance. To this end, we add a pose branch to extract the local spatial information and optimize the whole network on both person identity and attribute objectives. To diminish the negative affect from unstable pose estimation, a novel structure called channel parse block (CPB) is introduced to learn weights on different feature channels in pose branch. Then two branches are combined with compact bilinear pooling. Experimental results on Market1501 and DukeMTMC-reid datasets illustrate the effectiveness of the proposed method.
Qinqin Zhou 0001, Bineng Zhong 0001, Xiangyuan Lan, Gan Sun, Yulun Zhang 0001, Mengran Gou
IJCAI2
2019 Hierarchical Tracking by Reinforcement Learning-Based Searching and Coarse-to-Fine Verifying
abstract
A class-agnostic tracker typically consists of three key components, i.e., its motion model, its target appearance model, and its updating strategy. However, most recent topperforming trackers mainly focus on constructing complicated appearance models and updating strategies, while using comparatively simple and heuristic motion models that may result in an inefficient search and degrade the tracking performance. To address this issue, we propose a hierarchical tracker that learns to move and track based on the combination of data-driven search at the coarse level, and coarse-to-fine verification at the fine level. At the coarse level, a data-driven motion model learned from deep recurrent reinforcement learning provides our tracker with coarse localization of an object. By formulating motion search as an action-decision problem in reinforcement learning, our tracker utilizes a recurrent convolutional neural network based deep Q-network to effectively learn data-driven searching policies. The learned motion model cannot only significantly reduce the search space, but also provide more reliable interested regions for further verifying. At the fine level, a kernelized correlation filter (KCF) based appearance model is adopted to densely yet efficiently verify a local region centered on the predicted location from the motion model. Through using of circulant matrices and fast Fourier transformation, a large number of candidate samples in the local region can be efficiently and effectively evaluated by the KCF based appearance model. Finally, a simple yet robust estimator is designed to analyze possible tracking failure. The experiments on OTB50 and OTB100 illustrate that our tracker achieves better performance than the state-of-the-art trackers.
Bineng Zhong 0001, Jun Li 0027, Yulun Zhang 0001, Yun Fu 0001
IEEE Trans. Image Process.1
2019 Deep Alignment Network Based Multi-Person Tracking With Occlusion and Motion Reasoning
abstract
Tracking-by-detection is one of the typical paradigms for multi-person tracking, due to the availability of automatic pedestrian detectors. However, existing multi-person trackers are greatly challenged by misalignment in the pedestrian detectors (i.e., excessive background and part missing) and occlusion. To effectively handle these problems, we propose a deep alignment network-based multi-person tracking method with occlusion and motion reasoning. Specifically, the inaccurate detections are first corrected via a deep alignment network, in which an alignment estimation module is used to automatically learn the spatial transformation of these detections. As a result, the deep features from our alignment network will have better representation power and, thus, lead to more consistent tracks. Then, a coarse-to-fine schema is designed for construing a discriminative association cost matrix with spatial, motion, and appearance information. Meanwhile, a principled approach is developed to allow our method to handle occlusion with motion reasoning and the reidentification ability of the pedestrian alignment network. Finally, the association problem is solved via a simple yet real-time Hungarian algorithm. Comprehensive experiments on MOT16, ISSIA soccer, PETS09, and TUD datasets validate the effectiveness and robustness of our proposed tracker.
Qinqin Zhou 0001, Bineng Zhong 0001, Yulun Zhang 0001, Jun Li 0027, Yun Fu 0001
IEEE Trans. Multim.2
2018 Residual Dense Network for Image Super-Resolution
abstract
A very deep convolutional neural network (CNN) has recently achieved great success for image super-resolution (SR) and offered hierarchical features as well. However, most deep CNN based SR models do not make full use of the hierarchical features from the original low-resolution (LR) images, thereby achieving relatively-low performance. In this paper, we propose a novel residual dense network (RDN) to address this problem in image SR. We fully exploit the hierarchical features from all the convolutional layers. Specifically, we propose residual dense block (RDB) to extract abundant local features via dense connected convolutional layers. RDB further allows direct connections from the state of preceding RDB to all the layers of current RDB, leading to a contiguous memory (CM) mechanism. Local feature fusion in RDB is then used to adaptively learn more effective features from preceding and current local features and stabilizes the training of wider network. After fully obtaining dense local features, we use global feature fusion to jointly and adaptively learn global hierarchical features in a holistic way. Experiments on benchmark datasets with different degradation models show that our RDN achieves favorable performance against state-of-the-art methods.
Yulun Zhang 0001, Yapeng Tian, Yu Kong 0001, Bineng Zhong 0001, Yun Fu 0001
CVPR4
2018 Image Super-Resolution Using Very Deep Residual Channel Attention Networks
Yulun Zhang 0001, Kai Li 0012, Lichen Wang, Bineng Zhong 0001, Yun Fu 0001
ECCV (7)5
2018 Semi-Convex Hull Tree: Fast Nearest Neighbor Queries for Large Scale Data on GPUs
abstract
A fast exact nearest neighbor search algorithm over large scale data is proposed based on semi-convex hull tree, where each node represents a semi-convex hull, which is made of a set of hyper planes. When performing the task of nearest neighbor queries, unnecessary distance computations can be greatly reduced by quadratic programming. GPUs are also used to accelerate the query process. Experiments conducted on both Intel(R) HD Graphics 4400 and Nvidia Geforce GTX1050 TI, as well as theoretical analysis show that the proposed algorithm yields significant improvements and outperforms current k-d tree based nearest neighbor query algorithms and others.
Yewang Chen, Lida Zhou, Nizar Bouguila, Bineng Zhong 0001, Zhen Lei 0001, Jixiang Du, Hailin Li
ICDM4
2018 Robust feature learning for online discriminative tracking without large-scale pre-training
Jun Zhang 0011, Bineng Zhong 0001, Cheng Wang 0020, Jixiang Du
Frontiers Comput. Sci.2
2018 Kernel correlation filters for visual tracking with adaptive fusion of heterogeneous cues
Bineng Zhong 0001, Gu Ouyang, Xin Liu 0011, Ziyi Chen 0001, Cheng Wang 0020
Neurocomputing2
2018 Coarse-to-fine visual tracking with PSR and scale driven expert-switching
Yan Chen 0017, Bineng Zhong 0001, Gu Ouyang, Jixiang Du
Neurocomputing3
2017 Automatic facial flaw detection and retouching via discriminative structure tensor
abstract
Facial retouching has been increasingly applied in current social media and entertainment industries. In this study, the authors propose an efficient approach to automatically detect and retouch the facial flaws by using discriminative structure tensor. First, a non‐linear structure tensor associated with saliency model is exploited to discriminatively and automatically detect the significant facial flaws. Then, a Gaussian skin model is constructed in YCbCr space and the OSTU operation is simultaneously utilised to precisely mark the facial skin regions, in which the mouth, eyebrows and nostril parts are excluded. Subsequently, diverse structure tensor is employed to discriminatively adjust the inpainting priority and propose a structure tensor‐based inpainting algorithm to retouch the detected flaws. Without manual intervention, the extensive experiments have shown its effectiveness in marking the freckles, blemishes and moles in face images, and the retouching performance is visually pleasing in comparison with state‐of‐the‐art counterparts.
Xin Liu 0011, Lu Xie, Bineng Zhong 0001, Jixiang Du, Qinmu Peng
IET Image Process.3
2017 Sparse Representation-Based Semi-Supervised Regression for People Counting
abstract
Label imbalance and the insufficiency of labeled training samples are major obstacles in most methods for counting people in images or videos. In this work, a sparse representation-based semi-supervised regression method is proposed to count people in images with limited data. The basic idea is to predict the unlabeled training data, select reliable samples to expand the labeled training set, and retrain the regression model. In the algorithm, the initial regression model, which is learned from the labeled training data, is used to predict the number of people in the unlabeled training dataset. Then, the unlabeled training samples are regarded as an over-complete dictionary. Each feature of the labeled training data can be expressed as a sparse linear approximation of the unlabeled data. In turn, the labels of the labeled training data can be estimated based on a sparse reconstruction in feature space. The label confidence in labeling an unlabeled sample is estimated by calculating the reconstruction error. The training set is updated by selecting unlabeled samples with minimal reconstruction errors, and the regression model is retrained on the new training set. A co-training style method is applied during the training process. The experimental results demonstrate that the proposed method has a low mean square error and mean absolute error compared with those of state-of-the-art people-counting benchmarks.
Hongbo Zhang 0002, Bineng Zhong 0001, Jixiang Du, Jialin Peng, Duansheng Chen, Xiao Ke
ACM Trans. Multim. Comput. Commun. Appl.2
2016 Probability-based method for boosting human action recognition using scene context
abstract
In this study, the authors investigate the possibility of boosting action recognition performance by exploiting the associated scene context. Towards this end, the authors model a scene as a mid‐level ‘middle layer’ in order to bridge action descriptors and action categories. This is achieved via a scene topic model, in which hybrid visual descriptors, including spatial–temporal action features and scene descriptors, are first extracted from a video sequence. Then, the authors learn a joint probability distribution between scene and action using a naive Bayes nearest neighbour algorithm, which is adopted to jointly infer the action categories online by combining off‐the‐shelf action recognition algorithms. The authors demonstrate the advantages of their approach by comparing it with state‐of‐the‐art approaches using several action recognition benchmarks.
Hongbo Zhang 0002, Duansheng Chen, Bineng Zhong 0001, Jialin Peng, Jixiang Du, Songzhi Su
IET Comput. Vis.4
2016 Higher order partial least squares for object tracking: A 4D-tracking method
Bineng Zhong 0001, Xiangnan Yang, Yingju Shen, Cheng Wang 0020, Tian Wang 0001, Zhen Cui 0001, Hongbo Zhang 0002, Xiaopeng Hong, Duansheng Chen
Neurocomputing1
2016 Network in network based weakly supervised learning for visual tracking
Yan Chen 0017, Xiangnan Yang, Bineng Zhong 0001, Huizhen Zhang, Changlong Lin
J. Vis. Commun. Image Represent.3
2015 Understanding image structure via hierarchical shape parsing
abstract
Exploring image structure is a long-standing yet important research subject in the computer vision community. In this paper, we focus on understanding image structure inspired by the “simple-to-complex” biological evidence. A hierarchical shape parsing strategy is proposed to partition and organize image components into a hierarchical structure in the scale space. To improve the robustness and flexibility of image representation, we further bundle the image appearances into hierarchical parsing trees. Image descriptions are subsequently constructed by performing a structural pooling, facilitating efficient matching between the parsing trees. We leverage the proposed hierarchical shape parsing to study two exemplar applications including edge scale refinement and unsupervised “objectness” detection. We show competitive parsing performance comparing to the state-of-the-arts in above scenarios with far less proposals, which thus demonstrates the advantage of the proposed parsing scheme.
Xianming Liu 0005, Rongrong Ji, Changhu Wang, Wei Liu 0005, Bineng Zhong 0001, Thomas S. Huang
CVPR5
2015 Detecting Targets Based on a Realistic Detection and Decision Model in Wireless Sensor Networks
Tian Wang 0001, Zhen Peng 0003, Junbin Liang, Yiqiao Cai, Hui Tian 0002, Bineng Zhong 0001
WASA7
2015 Maximizing real-time streaming services based on a multi-servers networking framework
Tian Wang 0001, Yiqiao Cai, Weijia Jia 0001, Sheng Wen, Guojun Wang 0001, Hui Tian 0002, Bineng Zhong 0001
Comput. Networks8
2015 Online learning 3D context for robust visual tracking
Bineng Zhong 0001, Yingju Shen, Yan Chen 0017, Weibo Xie, Zhen Cui 0001, Hongbo Zhang 0002, Duansheng Chen, Tian Wang 0001, Xin Liu 0011, Shu-Juan Peng, Jin Gou, Jixiang Du, Jing Wang 0049, Wenming Zheng
Neurocomputing1
2015 Robust infrared target tracking based on particle filter with embedded saliency detection
Fanglin Wang, Yi Zhen, Bineng Zhong 0001, Rongrong Ji
Inf. Sci.3
2015 3D object tracking via image sets and depth-based occlusion detection
Yan Chen 0017, Yingju Shen, Xin Liu 0011, Bineng Zhong 0001
Signal Process.4
2014 Deep Network Cascade for Image Super-resolution
Zhen Cui 0001, Hong Chang 0001, Shiguang Shan, Bineng Zhong 0001, Xilin Chen 0001
ECCV (5)4
2014 Robust tracking via patch-based appearance model and local background estimation
Bineng Zhong 0001, Yan Chen 0017, Yingju Shen, Yewang Chen, Zhen Cui 0001, Rongrong Ji, Xiao-Tong Yuan, Duansheng Chen
Neurocomputing1
2014 Structured partial least squares for simultaneous object tracking and segmentation
Bineng Zhong 0001, Xiao-Tong Yuan, Rongrong Ji, Yan Yan 0001, Zhen Cui 0001, Xiaopeng Hong, Yan Chen 0017, Tian Wang 0001, Duansheng Chen
Neurocomputing1
2014 Visual tracking via weakly supervised learning from multiple imperfect oracles
Bineng Zhong 0001, Hongxun Yao, Sheng Chen 0007, Rongrong Ji, Tat-Jun Chin, Hanzi Wang
Pattern Recognit.1
2014 Automatic motion capture data denoising via filtered subspace clustering and low rank matrix approximation
Xin Liu 0011, Yiu-Ming Cheung, Shu-Juan Peng, Zhen Cui 0001, Bineng Zhong 0001, Jixiang Du
Signal Process.5
2013 Online structured hough forests for visual tracking
abstract
Segmentation-based tracking methods are popular in alleviating the model drift problem during online-learning of visual trackers. However, one of the limitations of those methods is that tracking results guide the process of segmentation. The model drift problem in tracking may have significant influence on segmentation. In this paper, we propose an online structured Hough Forests to address this limitation. The results of object tracking do not have significant influence on the process of segmentation. Our algorithm shows more robust results on several challenging sequences.
Bineng Zhong 0001, Hanzi Wang
ICASSP2
2013 Seeing actions through scene context
abstract
Recognizing human actions is not alone, as hinted by the scene herein. In this paper, we investigate the possibility to boost the action recognition performance by exploiting their scene context associated. To this end, we model the scene as a mid-level “hidden layer” to bridge action descriptors and action categories. This is achieved via a scene topic model, in which hybrid visual descriptors including spatiotemporal action features and scene descriptors are first extracted from the video sequence. Then, we learn a joint probability distribution between scene and action by a Naive Bayesian N-earest Neighbor algorithm, which is adopted to jointly infer the action categories online by combining off-the-shelf action recognition algorithms. We demonstrate our merits by comparing to state-of-the-arts in several action recognition benchmarks.
Hongbo Zhang 0002, Songzhi Su, Shaozi Li, Duansheng Chen, Bineng Zhong 0001, Rongrong Ji
VCIP5
2013 An effective unconstrained correlation filter and its kernelization for face recognition
Yan Yan 0001, Hanzi Wang, Cuihua Li, Chenhui Yang, Bineng Zhong 0001
Neurocomputing5
2013 Background subtraction driven seeds selection for moving objects segmentation and matting
Bineng Zhong 0001, Yan Chen 0017, Yewang Chen, Rongrong Ji, Duansheng Chen, Hanzi Wang
Neurocomputing1
2012 Matting-driven online learning of Hough forests for object tracking
Bineng Zhong 0001, Tat-Jun Chin, Hanzi Wang
ICPR2
2011 Complex background modeling based on Texture Pattern Flow with adaptive threshold propagation
Baochang Zhang 0001, Bineng Zhong 0001, Yao Cao
J. Vis. Commun. Image Represent.2
2011 Kernel Similarity Modeling of Texture Pattern Flow for Motion Detection in Complex Background
abstract
This paper proposes a novel kernel similarity modeling of texture pattern flow (KSM-TPF) for background modeling and motion detection in complex and dynamic environments. The texture pattern flow encodes the binary pattern changes in both spatial and temporal neighborhoods. The integral histogram of texture pattern flow is employed to extract the discriminative features from the input videos. Different from existing uniform threshold based motion detection approaches which are only effective for simple background, the kernel similarity modeling is proposed to produce an adaptive threshold for complex background. The adaptive threshold is computed from the mean and variance of an extended Gaussian mixture model. The proposed KSM-TPF approach incorporates machine learning method with feature extraction method in a homogenous way. Experimental results on the publicly available video sequences demonstrate that the proposed approach provides an effective and efficient way for background modeling and motion detection.
Baochang Zhang 0001, Yongsheng Gao 0001, Sanqiang Zhao, Bineng Zhong 0001
IEEE Trans. Circuits Syst. Video Technol.4
2011 Mining flickr landmarks by modeling reconstruction sparsity
abstract
In recent years, there have been ever-growing geographical tagged photos on the community Web sites such as Flickr. Discovering touristic landmarks from these photos can help us to make better sense of our visual world. In this article, we report our work on mining landmarks from geotagged Flickr photos for city scene summarization and touristic recommendations. We begin by exploring the geographical and visual statistics of the Web users' photographing manner, based on which we conduct landmark mining in two steps: First, we propose to partition each city into geographical regions based on spectral clustering over the geotags of Flickr photos. Second, in each landmark region, we present a representative photo mining scheme based on sparse representation. Our main idea is to regard the landmark mining problem as a process to find photos whose visual signatures can be reconstructed using other photos of this landmark region with a minimal coding length. This sparse reconstruction scheme offers a general perspective to mine the representative photos. Indeed, by simplifying the data correlation constraints in our scheme, several previous works in representative photo discovery and landmark mining can be derived. Finally, we introduce a Hyperlink-Induced Topic Search model to refine our landmark ranking, which incorporates the community knowledge to simulate the landmark ranking problem as a dynamic page ranking problem. We have deployed our proposed landmark mining framework on a city scene summarization and navigation system, which works on one million geotagged Flickr photos coming from twenty worldwide metropolises. We have also quantitatively compared our scheme with several state-of-the-art works.
Rongrong Ji, Yue Gao 0002, Bineng Zhong 0001, Hongxun Yao, Qi Tian 0001
ACM Trans. Multim. Comput. Commun. Appl.3
2010 Towards semantic embedding in visual vocabulary
abstract
Visual vocabulary serves as a fundamental component in many computer vision tasks, such as object recognition, visual search, and scene modeling. While state-of-the-art approaches build visual vocabulary based solely on visual statistics of local image patches, the correlative image labels are left unexploited in generating visual words. In this work, we present a semantic embedding framework to integrate semantic information from Flickr labels for supervised vocabulary construction. Our main contribution is a Hidden Markov Random Field modeling to supervise feature space quantization, with specialized considerations to label correlations: Local visual features are modeled as an Observed Field, which follows visual metrics to partition feature space. Semantic labels are modeled as a Hidden Field, which imposes generative supervision to the Observed Field with WordNet-based correlation constraints as Gibbs distribution. By simplifying the Markov property in the Hidden Field, both unsupervised and supervised (label independent) vocabularies can be derived from our framework. We validate our performances in two challenging computer vision tasks with comparisons to state-of-the-arts: (1) Large-scale image search on a Flickr 60,000 database; (2) Object recognition on the PASCAL VOC database.
Rongrong Ji, Hongxun Yao, Xiaoshuai Sun, Bineng Zhong 0001, Wen Gao 0001
CVPR4
2010 Visual tracking via weakly supervised learning from multiple imperfect oracles
abstract
Long-term persistent tracking in ever-changing environments is a challenging task, which often requires addressing difficult object appearance update problems. To solve them, most top-performing methods rely on online learning-based algorithms. Unfortunately, one inherent problem of online learning-based trackers is drift, a gradual adaptation of the tracker to non-targets. To alleviate this problem, we consider visual tracking in a novel weakly supervised learning scenario where (possibly noisy) labels but no ground truth are provided by multiple imperfect oracles (i.e., trackers), some of which may be mediocre. A probabilistic approach is proposed to simultaneously infer the most likely object position and the accuracy of each tracker. Moreover, an online evaluation strategy of trackers and a heuristic training data selection scheme are adopted to make the inference more effective and fast. Consequently, the proposed method can avoid the pitfalls of purely single tracking approaches and get reliable labeled samples to incrementally update each tracker (if it is an appearance-adaptive tracker) to capture the appearance changes. Extensive comparing experiments on challenging video sequences demonstrate the robustness and effectiveness of the proposed method.
Bineng Zhong 0001, Hongxun Yao, Sheng Chen 0007, Rongrong Ji, Xiao-Tong Yuan, Shaohui Liu, Wen Gao 0001
CVPR1
2010 Robust background modeling via standard variance feature
abstract
In this paper, a novel standard variance feature is proposed for background modeling in dynamic scenes involving waving trees and ripples in water. The standard variance feature is the standard variance of a set of pixels' feature values, which captures mainly co-occurrence statistics of neighboring pixels in an image patch. The background modeling method based on standard variance feature includes two main components. First, we divide image into patches and represent each image patch as a standard variance feature. Then, assuming that standard variance feature fits a mixture of Gaussians distribution, we use mixture of Gaussians models to model it. Experimental results on several challenging video sequences demonstrate the effectiveness of our method.
Bineng Zhong 0001, Hongxun Yao, Shaohui Liu
ICASSP1
2010 Sigma Set Based Implicit Online Learning for Object Tracking
abstract
This letter presents a novel object tracking approach within the Bayesian inference framework through implicit online learning. In our approach, the target is represented by multiple patches, each of which is encoded by a powerful and efficient region descriptor called Sigma set. To model each target patch, we propose to utilize the online one-class support vector machine algorithm, named Implicit online Learning with Kernels Model (ILKM). ILKM is simple, efficient, and capable of learning a robust online target predictor in the presence of appearance changes. Responses of ILKMs related to multiple target patches are fused by an arbitrator with an inference of possible partial occlusions, to make the decision and trigger the model update. Experimental results demonstrate that the proposed tracking approach is effective and efficient in ever-changing and cluttered scenes.
Xiaopeng Hong, Hong Chang 0001, Shiguang Shan, Bineng Zhong 0001, Xilin Chen 0001, Wen Gao 0001
IEEE Signal Process. Lett.4
2009 Local Spatial Co-occurrence for Background Subtraction via Adaptive Binned Kernel Estimation
Bineng Zhong 0001, Shaohui Liu, Hongxun Yao
ACCV (3)1
2009 Multl-resolution background subtraction for dynamic scenes
abstract
Dynamic scenes (e.g. waving trees, ripples in water, illumination changes, camera jitters etc.) challenge many traditional background subtraction methods. In this paper, we present a novel background subtraction approach for dynamic scenes, in which the background is modeled in a multi-resolution framework. First, for each level of the pyramid, we run an independent mixture of Gaussians Models (GMM) that outputs a background subtraction map. Second, these background subtraction maps are combined via AND operator to finally get a more robust and accurate background subtraction map. This is a natural fusion because the original resolution and low resolution images have complementary strengths, which original resolution image contains rich information and low resolution image is insensitive to the noises and the small movement of dynamic scene. Experimental result shows that this real-time algorithm is able to detect moving objects accurately even in dynamic scenes.
Bineng Zhong 0001, Shaohui Liu, Hongxun Yao, Baochang Zhang 0001
ICIP1
2009 Neighboring Image Patches Embedding for background modeling
abstract
We present a novel feature extraction framework, Neighboring Image Patches Embedding (NIPE), for robust and efficient background modeling. We divide image into patches and represent each image patch as a NIPE vector. Then, the background model of each image patch is constructed as a group of weighted adaptive NIPE vectors. The NIPE feature vector, whose components are similarities between current image patch and its neighbors, describes mainly the mutual relationship between neighboring patches. Since neighboring image patches tend to be similarly affected by environmental effects (e.g., dynamic background), the NIPE vectors are more robust in these conditions comparing with the conventional method. Experimental results demonstrate the efficiency and effectiveness of our proposed NIPE method.
Bineng Zhong 0001, Hongxun Yao, Shaohui Liu
ICIP1
2008 Complex background modeling and motion detection based on Texture Pattern Flow
abstract
This paper proposes a novel texture pattern flow (TPF) for complex background modeling and motion detection. The pattern flow is proposed to encode the binary pattern changes among the neighborhoods in the space-time domain. To model the distribution of the TPF, the TPF integral histograms are used to extract the discriminative features to represent the input video. Experimental results on the public videos testify the effectiveness of the proposed method in comparison to LBP and GMM based background modeling methods.
Baochang Zhang 0001, Yongsheng Gao 0001, Bineng Zhong 0001
ICPR3
2008 Hierarchical background subtraction using local pixel clustering
abstract
We propose a robust hierarchical background subtraction technique which takes the spatial relations of neighboring pixels in a local region into account to detect objects in difficult conditions. Our algorithm combines a per-pixel with a per-region background model in a hierarchical manner, which accentuates the advantages of each. This is a natural combination because the two models have complementary strengths. The per-pixel background model is achieved by mixture of Gaussians Models (GMM) with RGB feature. Although precisely describing background change in high resolution, it suffers from the sensitivity to quick variations in dynamic environment. To tolerate these quick variations, we further develop a novel GMM based per-region background model, which is updated by the cluster centers obtained from a k-means clustering of the pixels’ RGB feature in the region. Numerical and qualitative experimental results on challenging videos demonstrate the robustness of the proposed method.
Bineng Zhong 0001, Hongxun Yao, Shiguang Shan, Xilin Chen 0001, Wen Gao 0001
ICPR1