EDBT 2026 Demo / reviewers in the wild / expert
Zhenyu He 0001
dblp:57/6240-1
· DBLP profile ↗
109ranked-venue papers
8as first author
67since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 65 · 6 first-author · 36 since 2021Graphics, computer vision, multimedia, augmented reality and games · 51 · 1 first-author · 39 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 3 since 2021Computer networks · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorSecurity and privacy · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SGAFuse: Semantic-guided adaptive fusion for RGB-thermal images via dynamic gating
Deshui Miao, Zhenyu He 0001 |
Neural Networks | 5 |
| 2026 | Adversarial flow-based generative models for visible-to-Infrared person re-Identification
Honghu Pan, Yongyong Chen, Xin Li 0034, Zhenyu He 0001 |
Pattern Recognit. | 4 |
| 2026 | Harnessing Vision-Language Pretrained Models With Temporal-Aware Adaptation for Referring Video Object SegmentationabstractReferring Video Object Segmentation (RVOS) is a task that involves segmenting target objects in a video based on the given referring expressions. It is critical for video editing and analysis. The crux of RVOS is to model dense text-video relations to associate abstract linguistic concepts with dynamic visual contents at pixel-level. Most RVOS methods typically use vision and language models pretrained independently as backbones, mapping images and texts to uncoupled feature spaces. As a result, they must learn Vision-Language (VL) relation modeling from scratch. Vision-Language Pretrained (VLP) models have achieved remarkable success. Inspired by this, we propose to explore relation modeling for RVOS based on their aligned VL feature space. Nevertheless, transferring VLP models to RVOS is deceptively challenging, due to the gap between static image/region-level pretraining and dynamic pixel-level prediction. To bridge this gap, we introduce a framework named VLP-RVOS, which harnesses VLP models for RVOS through temporal-aware adaptation. We first propose temporal-aware prompt-tuning to adapt pretrained representations for pixel-level prediction and empower the vision encoder to model temporal contexts. We further customize a cube-frame attention mechanism for robust spatial-temporal reasoning. Besides, we propose to perform multi-stage VL relation modeling while and after feature extraction for comprehensive understanding. Extensive experiments demonstrate that VLP-RVOS performs favorably against state-of-the-art algorithms and generalizes well. Our codes are available at https://github.com/xwt909090/VLP-RVOS. Zikun Zhou, Wentao Xiong, Li Zhou 0017, Xin Li 0034, Zhenyu He 0001, Yaowei Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | Interactive Classification and Regression for Visual Tracking with Dual Update StrategyabstractThe current two-stage tracking method locates the target using the position with the highest confidence score, and updates the template using a carefully designed template update strategy. However, we identify two key issues with these trackers: (1) the update strategy lacks continuous, cost-free template adaptation, leading to suboptimal tracking under appearance changes and (2) the location with the highest confidence score does not always yield accurate bounding boxes, potentially resulting in incomplete target coverage. In this article, we propose a novel tracker that incorporates two key innovations. First, the tracker employs a dual update strategy that performs online template updates at both the image and feature levels. This strategy enables continuous adaptation to target appearance changes without introducing additional computational overhead. Second, we enhance the existing loss function by introducing a Classification–Regression Interaction (CRI) loss, which guides the training process to produce confidence scores that more accurately reflect the quality of the predicted bounding boxes. Extensive experiments are conducted to evaluate the performance of our tracker and the effectiveness of the proposed methods. The experimental results show that our method has achieved a comprehensive improvement over the baseline on five datasets, and achieves competitive performance compared to state-of-the-art trackers. Di Yuan 0002, Gu Geng, Qiao Liu 0001, Xiaojun Chang, Zhenyu He 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2025 | MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language TrackingabstractThe vision-language tracking task aims to perform object tracking based on various modality references. Existing Transformer-based vision-language tracking methods have made remarkable progress by leveraging the global modeling ability of self-attention. However, current approaches still face challenges in effectively exploiting the temporal information and dynamically updating reference features during tracking. Recently, the State Space Model (SSM), known as Mamba, has shown astonishing ability in efficient long-sequence modeling. Particularly, its state space evolving process demonstrates promising capabilities in memorizing multimodal temporal information with linear complexity. Witnessing its success, we propose a Mamba-based vision-language tracking model to exploit its state space evolving ability in temporal space for robust multimodal tracking, dubbed MambaVLT. In particular, our approach mainly integrates a time-evolving hybrid state space block and a selective locality enhancement block, to capture contextual information for multimodal modeling and adaptive reference feature update. Besides, we introduce a modality-selection module that dynamically adjusts the weighting between visual and language references, mitigating potential ambiguities from either reference type. Extensive experimental results show that our method performs favorably against state-of-the-art trackers across diverse benchmarks. Xinqi Liu, Li Zhou 0017, Zikun Zhou, Jianqiu Chen, Zhenyu He 0001 |
CVPR | 5 |
| 2025 | Learning Spatial-Semantic Features for Robust Video Object SegmentationabstractTracking and segmenting multiple similar objects with distinct or complex parts in long-term videos is particularly challenging due to the ambiguity in identifying target components and the confusion caused by occlusion, background clutter, and changes in appearance or environment over time. In this paper, we propose a robust video object segmentation framework that learns spatial-semantic features and discriminative object queries to address the above issues. Specifically, we construct a spatial-semantic block comprising a semantic embedding component and a spatial dependency modeling part for associating global semantic features and local spatial features, providing a comprehensive target representation. In addition, we develop a masked cross-attention module to generate object queries that focus on the most discriminative parts of target objects during query propagation, alleviating noise accumulation to ensure effective long-term query propagation. The experimental results show that the proposed method sets new state-of-the-art performance on multiple data sets, including the DAVIS2017 test (\textbf{87.8\%}), YoutubeVOS 2019 (\textbf{88.1\%}), MOSE val (\textbf{74.0\%}), and LVOS test (\textbf{73.0\%}), which demonstrate the effectiveness and generalization capacity of the proposed method. We will make all the source code and trained models publicly available. Xin Li 0034, Deshui Miao, Zhenyu He 0001, Yaowei Wang 0001, Huchuan Lu, Ming-Hsuan Yang 0001 |
ICLR | 3 |
| 2025 | CMRFusion: Efficient Feature Decomposition for RGB-T Fusion via Cross Modality Mask ReconstructionabstractThe objective of infrared-visible image fusion is to effectively combine the complementary information from both modalities and preserve the significant features within their respective imaging domains. In this paper, we propose a novel fusion method for infrared-visible images named CMRFusion, which employs mask reconstruction to disentangle cross-modality features. To be specific, we firstly transform the infrared and visual images to the modality-common (the semantic information of the scene) and modality-unique (the intensity attributes of the imaging sensors) feature subspaces separately by reconstructing the masked patches from those two modalities under meticulously designed constraints. A simple yet effective generator is then trained to integrate the above decomposed features into the final fused image. Extensive quantitative experiments on the public TNO and Roadscene datasets demonstrate that CMRFusion mostly outperforms the existing state-of-the-art methods. The visualization results also prove the effectiveness of CMRFusion in terms of decomposing the shared and distinctive features present within the two modalities. Qiang Wang 0051, Zhenyu He 0001 |
ICME | 5 |
| 2025 | ZeroBP: Learning Position-Aware Correspondence for Zero-Shot 6D Pose Estimation in Bin-PickingabstractBin-picking is a practical and challenging robotic manipulation task, where accurate 6D pose estimation plays a pivotal role. The workpieces in bin-picking are typically texture-less and randomly stacked in a bin, which poses a significant challenge to 6D pose estimation. Existing solutions are typically learning-based methods, which require object-specific training. Their efficiency of practical deployment for novel workpieces is highly limited by data collection and model retraining. Zero-shot 6D pose estimation is a potential approach to address the issue of deployment efficiency. Nevertheless, existing zero-shot 6D pose estimation methods are designed to leverage feature matching to establish point-to-point correspondences for pose estimation, which is less effective for workpieces with textureless appearances and ambiguous local regions. In this paper, we propose ZeroBP, a zero-shot pose estimation frame-work designed specifically for the bin-picking task. ZeroBP learns Position-Aware Correspondence (PAC) between the scene instance and its CAD model, leveraging both local features and global positions to resolve the mismatch issue caused by ambiguous regions with similar shapes and appearances. Extensive experiments on the ROBI dataset demonstrate that ZeroBP outperforms state-of-the-art zero-shot pose estimation methods, achieving an improvement of 9.1 % in average recall of correct poses. Jianqiu Chen, Zikun Zhou, Xin Li 0034, Tianpeng Bao, Zhenyu He 0001 |
ICRA | 6 |
| 2025 | CMFS: CLIP-Guided Modality Interaction for Mitigating Noise in Multi-Modal Image Fusion and SegmentationabstractInfrared-visible image fusion and semantic segmentation are pivotal tasks for robust scene understanding under challenging conditions such as low light. However, existing methods often struggle with high noise, modality inconsistencies, and inefficient cross-modal interactions, limiting fusion quality and segmentation accuracy. To this end, we propose CMFS, a unified framework that leverages CLIP-guided modality interaction to mitigate noise in multi-modal image fusion and segmentation. Our approach features a region-aware Modal Interaction Alignment module that combines a VMamba-based encoder with an additional shuffle layer to obtain more robust features and a CLIP-guided, regionally constrained multi-modal feature interaction block to emphasize foreground targets while suppressing low-light noise. Additionally, a Frequency-Spatial Collaboration module uses selective scanning and integrates wavelet-, spatial-, and Fourier-domain features to achieve adaptive denoising and balanced feature allocation. Furthermore, we employ a low-rank mixture-of-experts with dynamic routing to improve region-specific fusion and enhance pixel-level accuracy. Extensive experiments on several benchmarks show that, compared with state-of-the-art methods, the proposed approach demonstrates effectiveness in both image fusion quality and semantic segmentation accuracy, especially in complex environments. The source code will be released at IJCAI2025-CMFS. Guilin Su, Yuqing Huang, Zhenyu He 0001 |
IJCAI | 4 |
| 2025 | A Multi-Stream Visual-Spectral-Spatial Adaptive Hyperspectral Object TrackingabstractHyperspectral videos contain rich spectral and spatial information, which has enormous potential for object tracking compared to traditional RGB videos. However, the limited hyperspectral data, spectral differences across different bands, and high computational costs result in existing trackers being unable to effectively build connections between spectral and spatial information, leading to suboptimal tracking performance. To address these issues, this paper proposes a multi-stream Visual-Spectral-Spatial adaptive hyperspectral object tracking model(VSS). First, we design a multi-stream hyperspectral Transformer module to receive spectral data from different bands and extract Visual-Spectral-Spatial features across those bands. Then, to address the band differences between different modalities, we introduce the Bidirectional Visual-Spectral-Spatial Adapter module, which adaptively fuses visual, spectral, and spatial information from different bands. Finally, experiments on the HOTC dataset demonstrate the excellent performance of the VSS model. Qiao Liu 0001, Zhenyu He 0001, Di Yuan 0002 |
ICMR | 3 |
| 2025 | SCFusion: Enhance Infrared and Visible Modality Fusion by Preserving Salient Object ConsistencyabstractInfrared and visible images are captured using different sensors, resulting in various differences between the two modalities. However, current image fusion methods mainly focus on retaining global information, while neglecting to preserve the salient objects of the two source images. Additionally, existing evaluation metrics fail to measure whether the salient objects are preserved in the fused image from the two original modalities. To this end, we propose a novel image fusion method for infrared and visible images called SCFusion, which maintains salient objects consistency between the original two modalities and the fused image. Specifically, we designed a new module called the saliency decision (SD) to separate the unique and common saliency maps from the infrared and visible images for target enhancement in the final fused image. We then introduce a new metric called saliency information weight (SIW) to evaluate the preservation of salient objects by calculating the overlap between the saliency map of the fused image and those of the original modalities. To validate the practical application of our fusion algorithm, we establish a physical visible-infrared fusion system integrating SCFusion to provide real-time service, including a dual-sensor camera and an AI edge platform. Quantitative and qualitative experiments demonstrate the superiority of SCFusion over state-of-the-art methods in terms of salient objects preservation from the original two modalities. Qiang Wang 0001, Zhenyu He 0001, Xiaowen Chu 0001 |
IEEE Internet Things J. | 6 |
| 2025 | Learning a robust RGB-Thermal detector for extreme modality imbalance
Qiang Wang 0001, Zhenyu He 0001 |
Pattern Recognit. Lett. | 5 |
| 2025 | ZeroPose: CAD-Prompted Zero-Shot Object 6D Pose Estimation in Cluttered ScenesabstractMany robotics and industry applications have a high demand for the capability to estimate the 6D pose of novel objects from the cluttered scene. However, existing classic pose estimation methods are object-specific, which can only handle the specific objects seen during training. When applied to a novel object, these methods necessitate a cumbersome onboarding process, which involves extensive dataset preparation and model retraining. The extensive duration and resource consumption of onboarding limit their practicality in real-world applications In this paper, we introduce ZeroPose, a novel zero-shot framework that performs pose estimation following a Discovery-Orientation-Registration (DOR) inference pipeline. This framework generalizes to novel objects without requiring model retraining. Given the CAD model of a novel object, ZeroPose enables in seconds onboarding time to extract visual and geometric embeddings from the CAD model as a prompt. With the prompting of the above embeddings, DOR can discover all related instances and estimate their 6D poses without additional human interaction or presupposing scene conditions. Compared with existing zero-shot methods solved by the render-and-compare paradigm, the DOR pipeline formulates the object pose estimation into a feature-matching problem, which avoids time-consuming online rendering and improves efficiency. Experimental results on the seven datasets show that ZeroPose as a zero-shot method achieves comparable performance with object-specific training methods and outperforms the state-of-the-art zero-shot method with 50x inference speed improvement. Jianqiu Chen, Zikun Zhou, Mingshan Sun, Rui Zhao 0001, Tianpeng Bao, Zhenyu He 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2025 | Fine-Grained Feature and Template Reconstruction for TIR Object TrackingabstractThermal infrared (TIR) object tracking is a significant subject within the field of computer vision. Currently, TIR object tracking faces challenges such as insufficient representation of object texture information and underutilization of temporal information, which severely affects the tracking accuracy of TIR tracking methods. To address these issues, we propose a TIR object tracking method (called: FFTR) based on fine-grained feature and template reconstruction. Specifically, aiming at the fine-grained information of the TIR object, we employ a frequency channel attention mechanism that transforms TIR images into the frequency domain using discrete cosine transform features. By capturing the fine-grained feature of TIR images from the frequency domain, we enhance the model’s ability to comprehend these images. To better leverage temporal information, we utilize a template region reconstruction method. This method reconstructs the template from the previous frame based on the search area of the current frame, which is then incorporated into the attention computation for the subsequent frame, thereby improving the tracking capability of TIR objects. Extensive quantitative and qualitative experiments show that our method achieves competitive tracking performance on the TIR benchmarks. Donghai Liao, Xiu Shu, Zhihui Li 0001, Qiao Liu 0001, Di Yuan 0002, Xiaojun Chang, Zhenyu He 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2025 | Adaptive Trajectory Correction for Underwater Object TrackingabstractMost extant underwater object tracking (UOT) utilize generic tracking algorithms, which lack applicability to underwater tracking scenarios. Moreover, these algorithms primarily emphasize minimizing interference from various challenging tasks to prevent target drift, but pay less attention to the strategies for mitigating target drift once it occurs. To alleviate the above problems, we propose a simple, effective, and UOT-focused adaptive trajectory correction framework, named ATCTrack. From the perspective of tracking failure, this methodology aims to promptly identify and rectify unreasonable target drift through accurate trajectory coordinate correction and trajectory template updates. Additionally, to mitigate the adverse effects of potential erroneous corrections, we implement an adaptive strategy that corrects only significant target drift, allowing for self-correction within a certain margin. Finally, we introduce an adaptive underwater image enhancement technique to improve the underwater image quality and maintain the trajectory’s stability and clarity. Our tracker achieves state-of-the-art performance on the currently prevalent UOT tracking benchmarks compared to other trackers. Di Yuan 0002, Xiu Shu, Qiao Liu 0001, Xiaojun Chang, Zhenyu He 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | Geo6D: Geometric-Constraints-Guided Direct Object 6D Pose Estimation NetworkabstractDirect pose estimation networks aim to directly regress the 6D poses of target objects in the scene image using a neural network. These direct methods offer efficiency and an optimal optimization target, presenting significant potential for practical applications. However, due to the complex and implicit mappings between input features and target pose parameters, direct methods are challenging to train and prone to overfitting on mappings seen during training, resulting in limited effectiveness and generalization capability on unseen mappings. Existing methods focus primarily on improvements of the network architecture and training strategies, with less attention given to mappings. In this work, we propose a geometric constraints learning approach, which enables networks to explicitly capture and utilize the geometric mappings between inputs and optimization targets for pose estimation. Specifically, we introduce a residual pose transformation formula that preserves pose transformation constraints within both the 2D image plane and the 3D space while decoupling the absolute pose distribution, thereby addressing the pose distribution gap issue. We further design a Geo6D mechanism based on the formula, which enables the network to explicitly utilize geometric constraints for pose estimation by reconstructing the inputs and outputs. We select two different methods as our baseline and extensive experiments show that Geo6D enhances the performance and reduces the dependence on extensive training data, remaining effective even with only 10% of the typical data volume. Jianqiu Chen, Mingshan Sun, Tianpeng Bao, Zhenyu He 0001, Donghai Li, Guoqiang Jin, Rui Zhao 0001, Xiaoke Jiang |
IEEE Trans. Multim. | 5 |
| 2025 | Reliability-Guided Hierarchical Memory Network for Scribble-Supervised Video Object SegmentationabstractThis article aims to solve the video object segmentation (VOS) task in a scribble-supervised manner, in which VOS models are not only initialized with sparse target scribbles for inference but also trained by sparse scribble annotations. Thus, the annotation burdens for both initialization and training can be substantially lightened. The difficulties of scribble-supervised VOS lie in two aspects: 1) it demands a strong reasoning ability to carefully segment the target given only a sparse initial target scribble and 2) it necessitates learning dense prediction from sparse scribble annotations during training, requiring powerful learning capability. In this work, we propose a reliability-guided hierarchical memory network (RHMNet) for this task, which segments the target in a stepwise expanding strategy w.r.t. the memory reliability level. To be specific, RHMNet maintains a reliability-guided memory bank. It first uses the high-reliability memory to locate the region with high reliability belonging to the target, i.e., highly similar to the initial target scribble. Then, it expands the located high-reliability region to the entire target conditioned on the region itself and all existing memories. In addition, we propose a scribble-supervised learning mechanism to facilitate the model learning for dense prediction. It exploits the pixel-level relations within a single frame and the instance-level variations across multiple frames to take full advantage of the scribble annotations in sequence training samples. The favorable performance on four popular benchmarks demonstrates that our method is promising. Our project is available at: https://github.com/mkg1204/RHMNet-for-SSVOS. Zikun Zhou, Kaige Mao, Wenjie Pei, Hongpeng Wang 0002, Yaowei Wang 0001, Zhenyu He 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2024 | Block Image Compressive Sensing with Local and Global Information InteractionabstractBlock image compressive sensing methods, which divide a single image into small blocks for efficient sampling and reconstruction, have achieved significant success. However, these methods process each block locally and thus disregard the global communication among different blocks in the reconstruction step. Existing methods have attempted to address this issue with local filters or by directly reconstructing the entire image, but they have only achieved insufficient communication among adjacent pixels or bypassed the problem. To directly confront the communication problem among blocks and effectively resolve it, we propose a novel approach called Block Reconstruction with Blocks' Communication Network (BRBCN). BRBCN focuses on both local and global information, while further taking their interactions into account. Specifically, BRBCN comprises dual CNN and Transformer architectures, in which CNN is used to reconstruct each block for powerful local processing and Transformer is used to calculate the global communication among all the blocks. Moreover, we propose a global-to-local module (G2L) and a local-to-global module (L2G) to effectively integrate the representations of CNN and Transformer, with which our BRBCN network realizes the bidirectional interaction between local and global information. Extensive experiments show our BRBCN method outperforms existing state-of-the-art methods by a large margin. The code is available at https://github.com/kongxiuxiu/BRBCN Xiaoyu Kong, Yongyong Chen, Feng Zheng 0001, Zhenyu He 0001 |
AAAI | 4 |
| 2024 | RTracker: Recoverable Tracking via PN Tree Structured MemoryabstractExisting tracking methods mainly focus on learning better target representation or developing more robust prediction models to improve tracking performance. While tracking performance has significantly improved, the target loss issue occurs frequently due to tracking failures, complete occlusion, or out-of-view situations. However, con-siderably less attention is paid to the self-recovery issue of tracking methods, which is crucial for practical applications. To this end, we propose a recoverable tracking framework, RTracker, that uses a tree-structured memory to dynamically associate a tracker and a detector to enable self-recovery ability. Specifically, we propose a Positive-Negative Tree-structured memory to chronologically store and maintain positive and negative target samples. Upon the PN tree memory, we develop corresponding walking rules for determining the state of the target and define a set of control flows to unite the tracker and the detector in different tracking scenarios. Our core idea is to use the support samples of positive and negative target categories to establish a relative distance-based criterion for a reliable assessment of target loss. The favorable performance in comparison against the state-of-the-art methods on nu-merous challenging benchmarks demonstrates the effectiveness of the proposed algorithm. All the source code and trained models will be released at https://github.com/NorahGreen/RTracker. Yuqing Huang, Xin Li 0034, Zikun Zhou, Yaowei Wang 0001, Zhenyu He 0001, Ming-Hsuan Yang 0001 |
CVPR | 5 |
| 2024 | GRAPE: Generalizable and Robust Multi-view Facial Capture
Jing Li 0071, Zhenyu He 0001 |
ECCV (47) | 3 |
| 2024 | Spatial-Temporal Multi-level Association for Video Object Segmentation
Deshui Miao, Xin Li 0034, Zhenyu He 0001, Huchuan Lu, Ming-Hsuan Yang 0001 |
ECCV (67) | 3 |
| 2024 | Simplifying Cross-modal Interaction via Modality-Shared Features for RGBT TrackingabstractThermal infrared(TIR) data exhibits higher tolerance to extreme environments, making it a valuable complement to RGB data in tracking tasks. RGBT tracking aims to leverage information from RGB and TIR images for stable and robust tracking. However, existing RGBT tracking methods face challenges due to significant modality differences and selective emphasis on interactive information, leading to inefficiencies in the cross-modal interaction. To address these issues, we propose a novel Integrating Interaction into Modality-shared Features with ViT(IIMF) framework, which is a simplified cross-modal interaction network including modality-shared, RGB modality-specific, and TIR modality-specific branches. The Modality-shared branch aggregates modality-shared information and implements inter-modal interaction. Specifically, our approach first extracts modality-shared features from RGB and TIR features with a cross-attention mechanism. Furthermore, we design a Cross-Attention-based Modality-shared Information Aggregation(CAMIA) module to further aggregate modality-shared information with modality-shared tokens. We evaluate our model on three widely-used benchmark datasets and extensive experiments demonstrate that our method achieves state-of-the-art performance. All the source code are released at https://github.com/Liqiu-Chen/IIMF. Liqiu Chen, Yuqing Huang, Zikun Zhou, Zhenyu He 0001 |
ACM Multimedia | 5 |
| 2024 | Data Generation Scheme for Thermal Modality with Edge-Guided Adversarial Conditional Diffusion ModelabstractIn challenging low-light and adverse weather conditions, thermal vision algorithms, especially object detection, have exhibited remarkable potential, contrasting with the frequent struggles encountered by visible vision algorithms. Nevertheless, the efficacy of thermal vision algorithms driven by deep learning models remains constrained by the paucity of available training data samples. To this end, this paper introduces a novel approach termed the edge-guided conditional diffusion model (ECDM). This framework aims to produce meticulously aligned pseudo thermal images at the pixel level, leveraging edge information extracted from visible images. By utilizing edges as contextual cues from the visible domain, the diffusion model achieves meticulous control over the delineation of objects within the generated images. To alleviate the impacts of those visible-specific edge information that should not appear in the thermal domain, a two-stage modality adversarial training (TMAT) strategy is proposed to filter them out from the generated images by differentiating the visible and thermal modality. Extensive experiments on LLVIP demonstrate ECDM's superiority over existing state-of-the-art approaches in terms of image generation quality. The pseudo thermal images generated by ECDM also help to boost the performance of various thermal object detectors by up to 7.1 mAP. Code is available at https://github.com/lengmo1996/ECDM. Honghu Pan, Qiang Wang 0001, Zhenyu He 0001 |
ACM Multimedia | 6 |
| 2024 | Learning diverse fine-grained features for thermal infrared tracking
Qiao Liu 0001, Gaojun Li, Honghu Pan, Zhenyu He 0001 |
Expert Syst. Appl. | 5 |
| 2024 | Self-supervised discriminative model prediction for visual tracking
Di Yuan 0002, Gu Geng, Xiu Shu, Qiao Liu 0001, Xiaojun Chang, Zhenyu He 0001, Guangming Shi |
Neural Comput. Appl. | 6 |
| 2024 | Unified Conditional Image Generation for Visible-Infrared Person Re-IdentificationabstractThis paper proposes a unified multi-modal image generation method to address two critical challenges in visible-infrared (VI) person re-identification (ReID): the insufficiency of training samples and the large cross-modality discrepancy. To be specific, we propose to generate cross-modal and middle-modal images to explicitly reduce the modality discrepancy, and generate intra-modal images to serve as training samples for datasets augmentation. To this end, we adapt the conditional diffusion model for multi-modal image generation. The condition includes a binary modality indicator and modal-irrelative pedestrian contour to control the target modality and pedestrian identity, respectively. For the intra-modality and cross-modality image generation, we modify the structure of UNet to take as input the conditions, and estimate the conditional probability density by optimizing its variational lower bound. Furthermore, we devise modal discriminators and adversarial training strategies to achieve modality alignment. The middle-modality image generation method shares the same network architecture with intra- and cross-modality generation, but has specific training objectives. We define the middle modality as the distribution equidistant from the visible modality and infrared modality. We employ the adversarial training to measure the distance from the visible or infrared modality to the middle modality, and thus minimize the difference between these two adversarial losses, serving as an equidistant constraint. Experimental results on SYSU-MM01 and RegDB demonstrate the effectiveness and generalization of the intra-modality, cross-modality, and middle-modality image generation. Honghu Pan, Wenjie Pei, Xin Li 0034, Zhenyu He 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2024 | When Channel Correlation Meets Sparse Prior: Keeping Interpretability in Image Compressive SensingabstractImage compressive sensing (CS), recovering an unknown image by resorting to a small number of its measurements, has become an increasingly popular topic in multimedia technology and applications. For a better reconstruction, diverse priors, from the original sparse prior to the new deep prior, have been exploited. Despite the powerful learning capability and satisfactory reconstruction performance, the deep prior is known as a black box and loses clear interpretability. In this article, we first revisit image CS with different priors and observe that the method with hand-crafted sparse prior could still outperform state-of-art methods with deep prior or no prior when under the same settings, while the interpretability is well preserved. Then, towards a better performance of the sparse-prior-based method, we propose a Channel Adaptive Thresholding Network, namely CAT-Net. CAT-Net draws the support from channel correlation calculation to extend the single thresholding in the iterative soft thresholding algorithm (ISTA) into channel-wise thresholding. The channel adaptive thresholding conducts soft thresholding operation in each channel of the image features and can be adjusted adaptively to the inputs, which can reconstruct more precisely than a single static thresholding. The careful CAT operation can preserve patterns both in detail and holistically well. Experimental results demonstrate the proposed method outperforms the state-of-the-art image CS methods with both traditional and deep priors. Xiaoyu Kong, Yongyong Chen, Zhenyu He 0001 |
IEEE Trans. Multim. | 3 |
| 2024 | Robust Tensor Recovery for Incomplete Multi-View ClusteringabstractIncomplete multi-view clustering is gaining increased attention owing to its great success in mining underlying information from the missing views. However, the existing approaches still encounter two issues: 1) They generally do not give sufficient consideration to the robustness of incomplete multi-view data with noise; 2) They only exploit the low-rank structures in the intra-view graphs, while the low-rank priors embedded in inter-view graphs are ignored. To this end, we propose a Robust Tensor Recovery for Incomplete Multi-view Clustering (RIMC) method, which transforms the view-missing problem into the tensor graph recovery problem by manipulating the comprehensive low-rank priors. Specifically, RIMC first employs a marginalized denoising operation to construct robust graphs and further builds a tensor graph by stacking these robust graphs. Then, we develop a novel tensor completion to recover the tensor graph by performing comprehensive low-rank priors: low-rank structures in the inter-view graphs (i.e., horizontal and lateral slices); low-rank structures in the intra-view graphs (i.e., frontal slices). Meanwhile, we integrate the tensor completion and spectral clustering to learn a unified indicator matrix. Extensive experiments show the promising performance of our method. Qiangqiang Shen, Yongsheng Liang 0001, Yongyong Chen, Zhenyu He 0001 |
IEEE Trans. Multim. | 5 |
| 2024 | Cross-Modality Proposal-Guided Feature Mining for Unregistered RGB-Thermal Pedestrian DetectionabstractRGB-Thermal (RGB-T) pedestrian detection aims to locate pedestrians in RGB-T image pairs to exploit the complementation between the two modalities for improving detection robustness in extreme conditions. Most existing algorithms assume that the RGB-T image pairs are well registered, while in the real world, they are not ideally aligned due to parallax or different field-of-view of the cameras. The pedestrians in misaligned image pairs may be located at different positions in two images, which results in two challenges: 1) how to achieve inter-modality complementation using spatially misaligned RGB-T pedestrian patches and 2) how to recognize unpaired pedestrians at the boundary. To address these issues, we propose a new paradigm for unregistered RGB-T pedestrian detection, which predicts two separate pedestrian locations in RGB and thermal images. Specifically, we propose a cross-modality proposal-guided feature mining (CPFM) mechanism to extract two precise fusion features for representing a pedestrian in the two modalities, even if the given RGB-T image pair is unaligned. It enables us to effectively exploit the complementation between the two modalities. With the CPFM mechanism, we build a two-stream dense detector that predicts two pedestrian locations in the two modalities based on the corresponding fusion features mined by the CPFM mechanism. In addition, we design a data augmentation method, named Homography, to simulate the discrepancy in scales and views between images. We also investigate two non-maximum suppression (NMS) methods for post-processing purposes. Favorable experimental results demonstrate the effectiveness and robustness of our method in addressing unregistered pedestrians with different shifts. Zikun Zhou, Yuqing Huang, Gaojun Li, Zhenyu He 0001 |
IEEE Trans. Multim. | 5 |
| 2024 | Exploring the Temporal Consistency of Arbitrary Style Transfer: A Channelwise PerspectiveabstractArbitrary image stylization by neural networks has become a popular topic, and video stylization is attracting more attention as an extension of image stylization. However, when image stylization methods are applied to videos, unsatisfactory results that suffer from severe flickering effects appear. In this article, we conducted a detailed and comprehensive analysis of the cause of such flickering effects. Systematic comparisons among typical neural style transfer approaches show that the feature migration modules for state-of-the-art (SOTA) learning systems are ill-conditioned and could lead to a channelwise misalignment between the input content representations and the generated frames. Unlike traditional methods that relieve the misalignment via additional optical flow constraints or regularization modules, we focus on keeping the temporal consistency by aligning each output frame with the input frame. To this end, we propose a simple yet efficient multichannel correlation network (MCCNet), to ensure that output frames are directly aligned with inputs in the hidden feature space while maintaining the desired style patterns. An inner channel similarity loss is adopted to eliminate side effects caused by the absence of nonlinear operations such as softmax for strict alignment. Furthermore, to improve the performance of MCCNet under complex light conditions, we introduce an illumination loss during training. Qualitative and quantitative evaluations demonstrate that MCCNet performs well in arbitrary video and image style transfer tasks. Code is available at https://github.com/kongxiuxiu/MCCNetV2. Xiaoyu Kong, Yingying Deng, Fan Tang, Weiming Dong, Chongyang Ma, Yongyong Chen, Zhenyu He 0001, Changsheng Xu |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2024 | Self-Supervised Tracking via Target-Aware Data SynthesisabstractWhile deep-learning-based tracking methods have achieved substantial progress, they entail large-scale and high-quality annotated data for sufficient training. To eliminate expensive and exhaustive annotation, we study self-supervised (SS) learning for visual tracking. In this work, we develop the crop-transform-paste operation, which is able to synthesize sufficient training data by simulating various appearance variations during tracking, including appearance variations of objects and background interference. Since the target state is known in all synthesized data, existing deep trackers can be trained in routine ways using the synthesized data without human annotation. The proposed target-aware data-synthesis method adapts existing tracking approaches within a SS learning framework without algorithmic changes. Thus, the proposed SS learning mechanism can be seamlessly integrated into existing tracking frameworks to perform training. Extensive experiments show that our method: 1) achieves favorable performance against supervised (Su) learning schemes under the cases with limited annotations; 2) helps deal with various tracking challenges such as object deformation, occlusion (OCC), or background clutter (BC) due to its manipulability; 3) performs favorably against the state-of-the-art unsupervised tracking methods; and 4) boosts the performance of various state-of-the-art Su learning frameworks, including SiamRPN++, DiMP, and TransT. Xin Li 0034, Wenjie Pei, Yaowei Wang 0001, Zhenyu He 0001, Huchuan Lu, Ming-Hsuan Yang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | LSOTB-TIR: A Large-Scale High-Diversity Thermal Infrared Single Object Tracking BenchmarkabstractUnlike visual object tracking, thermal infrared (TIR) object tracking methods can track the target of interest in poor visibility such as rain, snow, and fog, or even in total darkness. This feature brings a wide range of application prospects for TIR object-tracking methods. However, this field lacks a unified and large-scale training and evaluation benchmark, which has severely hindered its development. To this end, we present a large-scale and high-diversity unified TIR single object tracking benchmark, called LSOTB-TIR, which consists of a tracking evaluation dataset and a general training dataset with a total of 1416 TIR sequences and more than 643 K frames. We annotate the bounding box of objects in every frame of all sequences and generate over 770 K bounding boxes in total. To the best of our knowledge, LSOTB-TIR is the largest and most diverse TIR object tracking benchmark to date. We spilt the evaluation dataset into a short-term tracking subset and a long-term tracking subset to evaluate trackers using different paradigms. What's more, to evaluate a tracker on different attributes, we also define four scenario attributes and 12 challenge attributes in the short-term tracking evaluation subset. By releasing LSOTB-TIR, we encourage the community to develop deep learning-based TIR trackers and evaluate them fairly and comprehensively. We evaluate and analyze 40 trackers on LSOTB-TIR to provide a series of baselines and give some insights and future research directions in TIR object tracking. Furthermore, we retrain several representative deep trackers on LSOTB-TIR, and their results demonstrate that the proposed training dataset significantly improves the performance of deep TIR trackers. Codes and dataset are available at https://github.com/QiaoLiuHit/LSOTB-TIR. Qiao Liu 0001, Xin Li 0034, Di Yuan 0002, Xiaojun Chang, Zhenyu He 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2024 | Active Learning for Deep Visual TrackingabstractConvolutional neural networks (CNNs) have been successfully applied to the single target tracking task in recent years. Generally, training a deep CNN model requires numerous labeled training samples, and the number and quality of these samples directly affect the representational capability of the trained model. However, this approach is restrictive in practice, because manually labeling such a large number of training samples is time-consuming and prohibitively expensive. In this article, we propose an active learning method for deep visual tracking, which selects and annotates the unlabeled samples to train the deep CNN model. Under the guidance of active learning, the tracker based on the trained deep CNN model can achieve competitive tracking performance while reducing the labeling cost. More specifically, to ensure the diversity of selected samples, we propose an active learning method based on multiframe collaboration to select those training samples that should be and need to be annotated. Meanwhile, considering the representativeness of these selected samples, we adopt a nearest-neighbor discrimination method based on the average nearest-neighbor distance to screen isolated samples and low-quality samples. Therefore, the training samples' subset selected based on our method requires only a given budget to maintain the diversity and representativeness of the entire sample set. Furthermore, we adopt a Tversky loss to improve the bounding box estimation of our tracker, which can ensure that the tracker achieves more accurate target states. Extensive experimental results confirm that our active-learning-based tracker (ALT) achieves competitive tracking accuracy and speed compared with state-of-the-art trackers on the seven most challenging evaluation benchmarks. Project website: https://sites.google.com/view/altrack/. Di Yuan 0002, Xiaojun Chang, Qiao Liu 0001, Yi Yang 0001, Minglei Shu, Zhenyu He 0001, Guangming Shi |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2024 | Audio2Gestures: Generating Diverse Gestures From AudioabstractPeople may perform diverse gestures affected by various mental and physical factors when speaking the same sentences. This inherent one-to-many relationship makes co-speech gesture generation from audio particularly challenging. Conventional CNNs/RNNs assume one-to-one mapping, and thus tend to predict the average of all possible target motions, easily resulting in plain/boring motions during inference. So we propose to explicitly model the one-to-many audio-to-motion mapping by splitting the cross-modal latent code into shared code and motion-specific code. The shared code is expected to be responsible for the motion component that is more correlated to the audio while the motion-specific code is expected to capture diverse motion information that is more independent of the audio. However, splitting the latent code into two parts poses extra training difficulties. Several crucial training losses/strategies, including relaxed motion loss, bicycle constraint, and diversity loss, are designed to better train the VAE. Experiments on both 3D and 2D motion datasets verify that our method generates more realistic and diverse motions than previous state-of-the-art methods, quantitatively and qualitatively. Besides, our formulation is compatible with discrete cosine transformation (DCT) modeling and other popular backbones (i.e., RNN, Transformer). As for motion losses and quantitative motion evaluation, we find structured losses/metrics (e.g. STFT) that consider temporal and/or spatial context complement the most commonly used point-wise losses (e.g. PCK), resulting in better motion dynamics and more nuanced motion details. Finally, we demonstrate that our method can be readily used to generate motion sequences with user-specified motion clips on the timeline. Jing Li 0071, Wenjie Pei, Xuefei Zhe, Ying Zhang 0021, Linchao Bao, Zhenyu He 0001 |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2023 | Skating-Mixer: Long-Term Sport Audio-Visual Modeling with MLPsabstractFigure skating scoring is challenging because it requires judging players’ technical moves as well as coordination with the background music. Most learning-based methods struggle for two reasons: 1) each move in figure skating changes quickly, hence simply applying traditional frame sampling will lose a lot of valuable information, especially in 3 to 5 minutes lasting videos; 2) prior methods rarely considered the critical audio-visual relationship in their models. Due to these reasons, we introduce a novel architecture, named Skating-Mixer. It extends the MLP framework into a multimodal fashion and effectively learns long-term representations through our designed memory recurrent unit (MRU). Aside from the model, we collected a high-quality audio-visual FS1000 dataset, which contains over 1000 videos on 8 types of programs with 7 different rating metrics, overtaking other datasets in both quantity and diversity. Experiments show the proposed method achieves SOTAs over all major metrics on the public Fis-V and our FS1000 dataset. In addition, we include an analysis applying our method to the recent competitions in Beijing 2022 Winter Olympic Games, proving our method has strong applicability. Jingfei Xia, Mingchen Zhuge, Tiantian Geng, Shun Fan, Yuantai Wei, Zhenyu He 0001, Feng Zheng 0001 |
AAAI | 6 |
| 2023 | Joint Visual Grounding and Tracking with Natural Language SpecificationabstractTracking by natural language specification aims to locate the referred target in a sequence based on the natural language description. Existing algorithms solve this issue in two steps, visual grounding and tracking, and accordingly deploy the separated grounding model and tracking model to implement these two steps, respectively. Such a separated framework overlooks the link between visual grounding and tracking, which is that the natural language descriptions provide global semantic cues for localizing the target for both two steps. Besides, the separated framework can hardly be trained end-to-end. To handle these issues, we propose a joint visual grounding and tracking framework, which reformulates grounding and tracking as a unified task: localizing the referred target based on the given visual-language references. Specifically, we propose a multi-source relation modeling module to effectively build the relation between the visual-language references and the test image. In addition, we design a temporal modeling module to provide a temporal clue with the guidance of the global semantic information for our model, which effectively improves the adaptability to the appearance variations of the target. Extensive experimental results on TNL2K, LaSOT, OTB99, and RefCOCOg demonstrate that our method performs favorably against state-of-the-art algorithms for both tracking and grounding. Code is available at https://github.com/lizhou-cs/JointNLT. Li Zhou 0017, Zikun Zhou, Kaige Mao, Zhenyu He 0001 |
CVPR | 4 |
| 2023 | Transferable Decoding with Visual Entities for Zero-Shot Image CaptioningabstractImage-to-text generation aims to describe images using natural language. Recently, zero-shot image captioning based on pre-trained vision-language models (VLMs) and large language models (LLMs) has made significant progress. However, we have observed and empirically demonstrated that these methods are susceptible to modality bias induced by LLMs and tend to generate descriptions containing objects (entities) that do not actually exist in the image but frequently appear during training (i.e., object hallucination). In this paper, we propose ViECap, a transferable decoding model that leverages entity-aware decoding to generate descriptions in both seen and unseen scenarios. ViECap incorporates entity-aware hard prompts to guide LLMs’ attention toward the visual entities present in the image, enabling coherent caption generation across diverse scenes. With entity-aware hard prompts, ViECap is capable of maintaining performance when transferring from in-domain to out-of-domain scenarios. Extensive experiments demonstrate that ViECap sets a new state-of-the-art cross-domain (transferable) captioning and performs competitively in-domain captioning compared to previous VLMs-based zero-shot methods. Our code is available at: https://github.com/FeiElysia/ViECap Junjie Fei, Teng Wang 0007, Zhenyu He 0001, Chengjie Wang 0001, Feng Zheng 0001 |
ICCV | 4 |
| 2023 | CiteTracker: Correlating Image and Text for Visual TrackingabstractExisting visual tracking methods typically take an image patch as the reference of the target to perform tracking. However, a single image patch cannot provide a complete and precise concept of the target object as images are limited in their ability to abstract and can be ambiguous, which makes it difficult to track targets with drastic variations. In this paper, we propose the CiteTracker to enhance target modeling and inference in visual tracking by connecting images and text. Specifically, we develop a text generation module to convert the target image patch into a descriptive text containing its class and attribute information, providing a comprehensive reference point for the target. In addition, a dynamic description module is designed to adapt to target variations for more effective target representation. We then associate the target description and the search image using an attention-based correlation module to generate the correlated features for target state reference. Extensive experiments on five diverse datasets are conducted to evaluate the proposed algorithm and the favorable performance against the state-of-the-art methods demonstrates the effectiveness of the proposed tracking method. The source code and trained models will be released at https://github.com/NorahGreen/CiteTracker. Xin Li 0034, Yuqing Huang, Zhenyu He 0001, Yaowei Wang 0001, Huchuan Lu, Ming-Hsuan Yang 0001 |
ICCV | 3 |
| 2023 | Siamese residual network for efficient visual tracking
Nana Fan, Qiao Liu 0001, Xin Li 0034, Zikun Zhou, Zhenyu He 0001 |
Inf. Sci. | 5 |
| 2023 | Robust thermal infrared tracking via an adaptively multi-feature fusion model
Di Yuan 0002, Xiu Shu, Qiao Liu 0001, Zhenyu He 0001 |
Neural Comput. Appl. | 5 |
| 2023 | Multi-granularity graph pooling for video-based person re-identification
Honghu Pan, Yongyong Chen, Zhenyu He 0001 |
Neural Networks | 3 |
| 2023 | Pose-Aided Video-Based Person Re-Identification via Recurrent Graph Convolutional NetworkabstractExisting methods for video-based person re- identification (ReID) mainly learn the appearance feature of a given pedestrian via a feature extractor and a feature aggregator. However, the appearance models would fail to learn a large inter-class variance when different pedestrians have similar appearances. Considering that different pedestrians have different walking postures and body proportions, we propose to learn the discriminative pose feature beyond the appearance feature for video retrieval. Specifically, we implement a two-branch architecture to separately learn the appearance feature and pose feature, and then concatenate them together for inference. To learn the pose feature, we first detect the pedestrian pose in each frame through an off-the-shelf pose detector, and construct a temporal graph using the pose sequence. We then exploit a recurrent graph convolutional network (RGCN) to learn the node embeddings of the temporal pose graph, which devises a global information propagation mechanism to simultaneously achieve the neighborhood aggregation of intra-frame nodes and message passing among inter-frame graphs. Finally, we propose a dual-attention method (DAM) consisting of node-attention and time-attention to obtain the temporal graph representation from the node embeddings, where the self-attention mechanism is employed to learn the importance of each node and each frame. We verify the proposed method on three video-based ReID datasets, i.e., Mars, DukeMTMC and iLIDS-VID, whose experimental results demonstrate that the learned pose feature can effectively improve the performance of existing appearance models. Honghu Pan, Qiao Liu 0001, Yongyong Chen, Yunqi He, Yuan Zheng 0002, Feng Zheng 0001, Zhenyu He 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2023 | Toward Complete-View and High-Level Pose-Based Gait RecognitionabstractModel-based gait recognition methods usually adopt the pedestrian walking postures to identify human beings. However, existing methods did not explicitly resolve the large intra-class variance of human pose due to changes in camera view. In this paper, we propose a lower-upper generative adversarial network (LUGAN) to generate multi-view pose sequences for each single-view sample to reduce the cross-view variance. Based on the prior of camera imaging, we prove that the spatial coordinates between cross-view poses satisfy a linear transformation of a full-rank matrix. Hence, LUGAN employs the adversarial training to learn full-rank transformation matrices from the source pose and target views to obtain the target pose sequences. The generator of LUGAN is composed of graph convolutional (GCN) layers, fully connected (FC) layers and two-branch convolutional (CNN) layers: GCN layers and FC layers encode the source pose sequence and target view, then CNN layers take as input the encoded features to learn a lower triangular matrix and an upper one, finally the transformation matrix is formulated by multiplying the lower and upper triangular matrices. For the purpose of adversarial training, we develop a conditional discriminator that distinguishes whether the pose sequence is true or generated. Furthermore, to facilitate the high-level correlation learning, we propose a plug-and-play module, named multi-scale hypergraph convolution (HGC), to replace the spatial graph convolutional layer in baseline, which can simultaneously model the joint-level, part-level and body-level correlations. Extensive experiments on three large gait recognition datasets (i.e., CASIA-B, OUMVLP-Pose and NLPR) demonstrate that our method outperforms the baseline model by a large margin. Honghu Pan, Yongyong Chen, Tingyang Xu, Yunqi He, Zhenyu He 0001 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2023 | Learning Dual-Level Deep Representation for Thermal Infrared TrackingabstractThe feature models used by existing Thermal InfraRed (TIR) tracking methods are usually learned from RGB images due to the lack of a large-scale TIR image training dataset. However, these feature models are less effective in representing TIR objects and they are difficult to effectively distinguish distractors because they do not contain fine-grained discriminative information. To this end, we propose a dual-level feature model containing the TIR-specific discriminative feature and fine-grained correlation feature for robust TIR object tracking. Specifically, to distinguish inter-class TIR objects, we first design an auxiliary multi-classification network to learn the TIR-specific discriminative feature. Then, to recognize intra-class TIR objects, we propose a fine-grained aware module to learn the fine-grained correlation feature. These two kinds of features complement each other and represent TIR objects in the levels of inter-class and intra-class respectively. These two feature models are constructed using a multi-task matching framework and are jointly optimized on the TIR object tracking task. In addition, we develop a large-scale TIR image dataset to train the network for learning TIR-specific feature patterns. To the best of our knowledge, this is the largest TIR tracking training dataset with the richest object class and scenario. To verify the effectiveness of the proposed dual-level feature model, we propose an offline TIR tracker (MMNet) and an online TIR tracker (ECO-MM) based on the feature model and evaluate them on three TIR tracking benchmarks. Extensive experimental results on these benchmarks demonstrate that the proposed algorithms perform favorably against the state-of-the-art methods. Qiao Liu 0001, Di Yuan 0002, Nana Fan, Peng Gao 0005, Xin Li 0034, Zhenyu He 0001 |
IEEE Trans. Multim. | 6 |
| 2022 | GuidedMix-Net: Semi-supervised Semantic Segmentation by Using Labeled Images as ReferenceabstractSemi-supervised learning is a challenging problem which aims to construct a model by learning from limited labeled examples. Numerous methods for this task focus on utilizing the predictions of unlabeled instances consistency alone to regularize networks. However, treating labeled and unlabeled data separately often leads to the discarding of mass prior knowledge learned from the labeled examples. In this paper, we propose a novel method for semi-supervised semantic segmentation named GuidedMix-Net, by leveraging labeled information to guide the learning of unlabeled instances. Specifically, GuidedMix-Net employs three operations: 1) interpolation of similar labeled-unlabeled image pairs; 2) transfer of mutual information; 3) generalization of pseudo masks. It enables segmentation models can learning the higher-quality pseudo masks of unlabeled data by transfer the knowledge from labeled samples to unlabeled data. Along with supervised learning for labeled data, the prediction of unlabeled data is jointly learned with the generated pseudo masks from the mixed data. Extensive experiments on PASCAL VOC 2012, and Cityscapes demonstrate the effectiveness of our GuidedMix-Net, which achieves competitive segmentation accuracy and significantly improves the mIoU over 7$\%$ compared to previous approaches. Peng Tu, Yawen Huang, Feng Zheng 0001, Zhenyu He 0001, Liujuan Cao, Ling Shao 0001 |
AAAI | 4 |
| 2022 | Global Tracking via Ensemble of Local TrackersabstractThe crux of long-term tracking lies in the difficulty of tracking the target with discontinuous moving caused by out-of-view or occlusion. Existing long-term tracking methods follow two typical strategies. The first strategy employs a local tracker to perform smooth tracking and uses another re-detector to detect the target when the target is lost. While it can exploit the temporal context like historical appearances and locations of the target, a potential limitation of such strategy is that the local tracker tends to misidentify a nearby distractor as the target instead of activating the re-detector when the real target is out of view. The other long-term tracking strategy tracks the target in the entire image globally instead of local tracking based on the previous tracking results. Unfortunately, such global tracking strategy cannot leverage the temporal context effectively. In this work, we combine the advantages of both strategies: tracking the target in a global view while exploiting the temporal context. Specifically, we perform global tracking via ensemble of local trackers spreading the full image. The smooth moving of the target can be handled steadily by one local tracker. When the local tracker accidentally loses the target due to suddenly discontinuous moving, another local tracker close to the target is then activated and can readily take over the tracking to locate the target. While the activated local tracker performs tracking locally by leveraging the temporal context, the ensemble of local trackers renders our model the global view for tracking. Extensive experiments on six datasets demonstrate that our method performs favorably against state-of-the-art algorithms. Zikun Zhou, Jianqiu Chen, Wenjie Pei, Kaige Mao, Hongpeng Wang 0002, Zhenyu He 0001 |
CVPR | 6 |
| 2022 | Structural target-aware model for thermal infrared tracking
Di Yuan 0002, Xiu Shu, Qiao Liu 0001, Zhenyu He 0001 |
Neurocomputing | 4 |
| 2022 | Accurate bounding-box regression with distance-IoU loss for visual tracking
Di Yuan 0002, Xiu Shu, Nana Fan, Xiaojun Chang, Qiao Liu 0001, Zhenyu He 0001 |
J. Vis. Commun. Image Represent. | 6 |
| 2022 | AAGCN: Adjacency-aware Graph Convolutional Network for person re-identification
Honghu Pan, Zhenyu He 0001, Chunkai Zhang |
Knowl. Based Syst. | 3 |
| 2022 | STCDesc: Learning deep local descriptor using similar triangle constraint
Qiao Liu 0001, Fanyang Meng, Zhenyu He 0001 |
Knowl. Based Syst. | 4 |
| 2022 | Noise-Suppressing Deep TrackingabstractIn visual tracking, it is challenging to distinguish the target from similar objects called noises in the background. As deep trackers use convolutional neural networks for image classification as feature extractors, the extracted features are insensitive to different instances in the same class, which is prone to make prediction models confuse the target and the similar noises in the background. To this end, we propose a noise-suppressing algorithm to learn the discriminative representation for distinguishing the target from the noises in the background. First, we learn polynomial kernels for a search patch under the semantic guidance to increase the difference between representations of the target and the noises in the background. Second, we formulate the online foreground-background functions for the target and the noises in the background to learn an adaptive kernel, which suppresses the features positive for the noises and promotes the features positive for the target. We evaluate the proposed method on seven public datasets including OTB-2013, OTB-2015, VOT-2018, LaSOT, TrackingNet, GOT10k, and NFS. The comprehensive experimental results show that the proposed algorithm performs favorably against state-of-the-art methods, while running at real-time speed. Nana Fan, Xin Li 0034, Zikun Zhou, Qiao Liu 0001, Zhenyu He 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | TCDesc: Learning Topology Consistent Descriptors for Image MatchingabstractThe triplet loss is widely used in learning the local descriptors for image matching. However, existing triplet loss-based methods, like HardNet and DSM, employ the point-to-point distance metric, which neglects the neighborhood information of descriptors. Considering the fact that local neighborhood structures of matching descriptors would be similar under the ideal condition, this paper aims to learn the neighborhood topology-consistent descriptors (TCDesc). To this end, we first propose the linear combination weight as the topology weight to depict the neighborhood topology for each descriptor, where the difference between the center descriptor and the linear combination of its neighbors is minimized. For the global comparison, we then define a global topology vector by using the local topology weights. Next, beyond the Euclidean distance, we define a topology distance with the topology vectors to indicate the topological difference between the matching descriptors. Furthermore, we propose an adaptive weighting strategy to jointly minimize the topology distance and Euclidean distance in triplet loss. Experimental results on four widely-used datasets, i.e., UBC PhotoTourism, HPatches, W1BS and Oxford, demonstrate that our method can effectively improve the performance of both HardNet and DSM. Honghu Pan, Yongyong Chen, Zhenyu He 0001, Fanyang Meng, Nana Fan |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Target-Aware State Estimation for Visual TrackingabstractTrackers based on the IoU prediction network (IoU-Net) have shown superior performance, which refines a coarse bounding box to an accurate one by maximizing the IoU between the target and the coarse box. However, the traditional IoU-Net is less effective in exploiting the limited but crucial supervision information contained in the initial frame, including the discriminative information between the target and backgrounds and the structure information of the initial target. Missing such information makes the IoU-Net less robust to background distractors and diverse variations of the target appearance. To address this issue, we propose a target-aware state estimation network for visual tracking. A gradient-guided feature adjustment module is built on an online discriminative model to generate target-aware features for constructing the state estimation network; it conveys the online learned discriminative information into the offline trained state estimation network. In addition, we propose a structure-aware integration module and embed it into the state estimation network, enabling the tracker to explicitly model the structure information of the initial target. Extensive experimental results on the VOT2018, OTB2015, UAV123, NFS30, TC128, TrackingNet, LaSOT, and VOT2018-LT datasets demonstrate that the proposed approach performs favorably against state-of-the-art trackers. Zikun Zhou, Xin Li 0034, Nana Fan, Hongpeng Wang 0002, Zhenyu He 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | Object Tracking via Spatial-Temporal Memory NetworkabstractTemporal and spatial contexts, characterizing target appearance variations and target-background differences, respectively, are crucial for improving the online adaptive ability and instance-level discriminative ability of object tracking. However, most existing trackers focus on either the temporal context or the spatial context during tracking and have not exploited these contexts simultaneously and effectively. In this paper, we propose a Spatial-TEmporal Memory (STEM) network to exploit these contexts jointly for object tracking. Specifically, we develop a key-value structured memory model equipped with a key-value index-based memory reading mechanism to model the spatial and temporal contexts simultaneously. To update the memory with new target states and ensure the diversity of the memory, we introduce a similarity-aware memory update scheme. In addition, we construct an entropy-guided ensemble strategy to fuse the prediction models based on these two contexts, such that these two contexts can be exploited to estimate the target state jointly. Extensive experimental results on eight challenging datasets, including OTB2015, TC128, UAV123, VOT2018, LaSOT, TrackingNet, GOT-10k, and OxUvA, demonstrate that the proposed method performs favorably against state-of-the-art trackers. Zikun Zhou, Xin Li 0034, Tianzhu Zhang 0001, Hongpeng Wang 0002, Zhenyu He 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | Fast Extended Inductive Robust Principal Component Analysis With Optimal MeanabstractInspired by the mean calculation of RPCA_OM and inductiveness of IRPCA, we first propose an inductive robust principal component analysis method with removing the optimal mean automatically, which is shorted as IRPCA_OM. Furthermore, IRPCA_OM is extended to Schatten-$p$norm and a more general framework (i.e., EIRPCA_OM) is presented. The objective function of EIRPCA_OM includes two terms, the first term is a robust reconstruction error term constrained by an$\ell _{2,1}$-norm and the second term is a regularization term constrained by a Schatten-$p$norm. The proposed EIRPCA_OM method is robust, inductive and accurate. However, on the high-dimensional data, it would spend a large computation cost in training stage. To this end, a fast version of EIRPCA_OM called as FEIRPCA_OM is proposed, and its basic idea is to eliminate the zero eigenvalues of data matrix. More importantly, an effective theoretical proof is presented to ensure that FEIRPCA_OM has faster processing speed than EIRPCA_OM when processing high-dimensional data, but without any performance loss. Based on it, we also can exchange the less performance loss for the higher computation efficiency by removing the small eigenvalues of data matrix. Experimental results on the public datasets demonstrate that FEIRPCA_OM works efficiently on the high-dimensional data. Shuangyan Yi, Feiping Nie 0001, Yongsheng Liang 0001, Wei Liu 0065, Zhenyu He 0001, Qingmin Liao |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2022 | SiamCorners: Siamese Corner Networks for Visual TrackingabstractThe current Siamese network based on region proposal network (RPN) has attracted great attention in visual tracking due to its excellent accuracy and high efficiency. However, the design of the RPN involves the selection of the number, scale, and aspect ratios of anchor boxes, which will affect the applicability and convenience of the model. Furthermore, these anchor boxes require complicated calculations, such as calculating their intersection-over-union (IoU) with ground truth bounding boxes. Due to the problems related to anchor boxes, we propose a simple yet effective anchor-free tracker (named Siamese corner networks, SiamCorners), which is end-to-end trained offline on large-scale image pairs. Specifically, we introduce a modified corner pooling layer to convert the bounding box estimate of the target into a pair of corner predictions (the bottom-right and the top-left corners). By tracking a target as a pair of corners, we avoid the need to design the anchor boxes. This will make the entire tracking algorithm more flexible and simple than anchor-based trackers. In our network design, we further introduce a layer-wise feature aggregation strategy that enables the corner pooling module to predict multiple corners for a tracking target in deep networks. We then introduce a new penalty term that is used to select an optimal tracking box in these candidate corners. Finally, SiamCorners achieves experimental results that are comparable to the state-of-art tracker while maintaining a high running speed. In particular, SiamCorners achieves a 53.7% AUC on NFS30 and a 61.4% AUC on UAV123, while still running at 42 frames per second (FPS). Kai Yang 0018, Zhenyu He 0001, Wenjie Pei, Zikun Zhou, Xin Li 0034, Di Yuan 0002, Haijun Zhang 0002 |
IEEE Trans. Multim. | 2 |
| 2022 | Learning Adaptive Spatial-Temporal Context-Aware Correlation Filters for UAV TrackingabstractTracking in the unmanned aerial vehicle (UAV) scenarios is one of the main components of target-tracking tasks. Different from the target-tracking task in the general scenarios, the target-tracking task in the UAV scenarios is very challenging because of factors such as small scale and aerial view. Although the discriminative correlation filter (DCF)-based tracker has achieved good results in tracking tasks in general scenarios, the boundary effect caused by the dense sampling method will reduce the tracking accuracy, especially in UAV-tracking scenarios. In this work, we propose learning an adaptive spatial-temporal context-aware (ASTCA) model in the DCF-based tracking framework to improve the tracking accuracy and reduce the influence of boundary effect, thereby enabling our tracker to more appropriately handle UAV-tracking tasks. Specifically, our ASTCA model can learn a spatial-temporal context weight, which can precisely distinguish the target and background in the UAV-tracking scenarios. Besides, considering the small target scale and the aerial view in UAV-tracking scenarios, our ASTCA model incorporates spatial context information within the DCF-based tracker, which could effectively alleviate background interference. Extensive experiments demonstrate that our ASTCA method performs favorably against state-of-the-art tracking methods on some standard UAV datasets. Di Yuan 0002, Xiaojun Chang, Zhihui Li 0001, Zhenyu He 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2021 | Audio2Gestures: Generating Diverse Gestures from Speech Audio with Conditional Variational AutoencodersabstractGenerating conversational gestures from speech audio is challenging due to the inherent one-to-many mapping be-tween audio and body motions. Conventional CNNs/RNNs assume one-to-one mapping, and thus tend to predict the average of all possible target motions, resulting in plain/boring motions during inference. In order to over-come this problem, we propose a novel conditional variational autoencoder (VAE) that explicitly models one-to-many audio-to-motion mapping by splitting the cross-modal latent code into shared code and motion-specific code. The shared code mainly models the strong correlation between audio and motion (such as the synchronized audio and motion beats), while the motion-specific code captures diverse motion information independent of the audio. However, splitting the latent code into two parts poses training difficulties for the VAE model. A mapping network facilitating random sampling along with other techniques including relaxed motion loss, bicycle constraint, and diversity loss are designed to better train the VAE. Experiments on both 3D and 2D motion datasets verify that our method generates more realistic and diverse motions than state-of-the-art methods, quantitatively and qualitatively. Finally, we demonstrate that our method can be readily used to generate motion sequences with user-specified motion clips on the timeline. Code and more results are at https://jingli513.github.io/audio2gestures. Jing Li 0071, Wenjie Pei, Xuefei Zhe, Ying Zhang 0021, Zhenyu He 0001, Linchao Bao |
ICCV | 6 |
| 2021 | Saliency-Associated Object TrackingabstractMost existing trackers based on deep learning perform tracking in a holistic strategy, which aims to learn deep representations of the whole target for localizing the target. It is arduous for such methods to track targets with various appearance variations. To address this limitation, another type of methods adopts a part-based tracking strategy which divides the target into equal patches and tracks all these patches in parallel. The target state is inferred by summarizing the tracking results of these patches. A potential limitation of such trackers is that not all patches are equally informative for tracking. Some patches that are not discriminative may have adverse effects. In this paper, we propose to track the salient local parts of the target that are discriminative for tracking. In particular, we propose a fine-grained saliency mining module to capture the local saliencies. Further, we design a saliency-association modeling module to associate the captured saliencies together to learn effective correlation representations between the exemplar and the search image for state estimation. Extensive experiments on five diverse datasets demonstrate that the proposed method performs favorably against state-of-the-art trackers. Zikun Zhou, Wenjie Pei, Xin Li 0034, Hongpeng Wang 0002, Feng Zheng 0001, Zhenyu He 0001 |
ICCV | 6 |
| 2021 | Deep 3D human pose estimation: A reviewabstractThree-dimensional (3D) human pose estimation involves estimating the articulated 3D joint locations of a human body from an image or video. Due to its widespread applications in a great variety of areas, such as human motion analysis, human–computer interaction, robots, 3D human pose estimation has recently attracted increasing attention in the computer vision community, however, it is a challenging task due to depth ambiguities and the lack of in-the-wild datasets. A large number of approaches, with many based on deep learning, have been developed over the past decade, largely advancing the performance on existing benchmarks. To guide future development, a comprehensive literature review is highly desired in this area. However, existing surveys on 3D human pose estimation mainly focus on traditional methods and a comprehensive review on deep learning based methods remains lacking in the literature. In this paper, we provide a thorough review of existing deep learning based works for 3D pose estimation, summarize the advantages and disadvantages of these methods and provide an in-depth understanding of this area. Furthermore, we also explore the commonly-used benchmark datasets on which we conduct a comprehensive study for comparison and analysis. Our study sheds light on the state of research development in 3D human pose estimation and provides insights that can facilitate the future design of models and algorithms. Jinbao Wang 0001, Shujie Tan, Xiantong Zhen, Feng Zheng 0001, Zhenyu He 0001, Ling Shao 0001 |
Comput. Vis. Image Underst. | 6 |
| 2021 | Interactive convolutional learning for visual tracking
Nana Fan, Qiao Liu 0001, Xin Li 0034, Zikun Zhou, Zhenyu He 0001 |
Knowl. Based Syst. | 5 |
| 2021 | Learning dual-margin model for visual tracking
Nana Fan, Xin Li 0034, Zikun Zhou, Qiao Liu 0001, Zhenyu He 0001 |
Neural Networks | 5 |
| 2021 | Adaptive ensemble perception tracking
Zikun Zhou, Nana Fan, Kai Yang 0018, Hongpeng Wang 0002, Zhenyu He 0001 |
Neural Networks | 5 |
| 2021 | Semi-Supervised Multi-View Deep Discriminant Representation LearningabstractLearning an expressive representation from multi-view data is a key step in various real-world applications. In this paper, we propose a semi-supervised multi-view deep discriminant representation learning (SMDDRL) approach. Unlike existing joint or alignment multi-view representation learning methods that cannot simultaneously utilize the consensus and complementary properties of multi-view data to learn inter-view shared and intra-view specific representations, SMDDRL comprehensively exploits the consensus and complementary properties as well as learns both shared and specific representations by employing the shared and specific representation learning network. Unlike existing shared and specific multi-view representation learning methods that ignore the redundancy problem in representation learning, SMDDRL incorporates the orthogonality and adversarial similarity constraints to reduce the redundancy of learned representations. Moreover, to exploit the information contained in unlabeled data, we design a semi-supervised learning framework by combining deep metric learning and density clustering. Experimental results on three typical multi-view learning tasks, i.e., webpage classification, image classification, and document classification demonstrate the effectiveness of the proposed approach. Xiaodong Jia 0005, Xiaoyuan Jing, Xiaoke Zhu, Songcan Chen, Bo Du 0001, Ziyun Cai, Zhenyu He 0001, Dong Yue 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2021 | Self-Supervised Deep Correlation TrackingabstractThe training of a feature extraction network typically requires abundant manually annotated training samples, making this a time-consuming and costly process. Accordingly, we propose an effective self-supervised learning-based tracker in a deep correlation framework (named: self-SDCT). Motivated by the forward-backward tracking consistency of a robust tracker, we propose a multi-cycle consistency loss as self-supervised information for learning feature extraction network from adjacent video frames. At the training stage, we generate pseudo-labels of consecutive video frames by forward-backward prediction under a Siamese correlation tracking framework and utilize the proposed multi-cycle consistency loss to learn a feature extraction network. Furthermore, we propose a similarity dropout strategy to enable some low-quality training sample pairs to be dropped and also adopt a cycle trajectory consistency loss in each sample pair to improve the training loss function. At the tracking stage, we employ the pre-trained feature extraction network to extract features and utilize a Siamese correlation tracking framework to locate the target using forward tracking alone. Extensive experimental results indicate that the proposed self-supervised deep correlation tracker (self-SDCT) achieves competitive tracking performance contrasted to state-of-the-art supervised and unsupervised tracking methods on standard evaluation benchmarks. Di Yuan 0002, Xiaojun Chang, Po-Yao Huang 0001, Qiao Liu 0001, Zhenyu He 0001 |
IEEE Trans. Image Process. | 5 |
| 2021 | Learning Deep Multi-Level Similarity for Thermal Infrared Object TrackingabstractExisting deep Thermal InfraRed (TIR) trackers only use semantic features to represent the TIR object, which lack the sufficient discriminative capacity for handling distractors. This becomes worse when the feature extraction network is only trained on RGB images. To address this issue, we propose a multi-level similarity model under a Siamese framework for robust TIR object tracking. Specifically, we compute different pattern similarities using the proposed multi-level similarity network. One of them focuses on the global semantic similarity and the other computes the local structural similarity of the TIR object. These two similarities complement each other and hence enhance the discriminative capacity of the network for handling distractors. In addition, we design a simple while effective relative entropy based ensemble subnetwork to integrate the semantic and structural similarities. This subnetwork can adaptive learn the weights of the semantic and structural similarities at the training stage. To further enhance the discriminative capacity of the tracker, we propose a large-scale TIR video sequence dataset for training the proposed model. To the best of our knowledge, this is the first and the largest TIR object tracking training dataset to date. The proposed TIR dataset not only benefits the training for TIR object tracking but also can be applied to numerous TIR visual tasks. Extensive experimental results on three benchmarks demonstrate that the proposed algorithm performs favorably against the state-of-the-art methods. Qiao Liu 0001, Xin Li 0034, Zhenyu He 0001, Nana Fan, Di Yuan 0002, Hongpeng Wang 0002 |
IEEE Trans. Multim. | 3 |
| 2021 | Similarity-Maintaining Privacy Preservation and Location-Aware Low-Rank Matrix Factorization for QoS Prediction Based Web Service RecommendationabstractWeb service recommendation plays an important role in building service-oriented systems. QoS-based Web service recommendation has recently gained much attention for providing a promising way to help users find high-quality services. To accurately predict the QoS values of candidate Web services, Web service recommendation systems usually need to collect historical QoS data from users, which will potentially pose a threat to the user's privacy. However, how to simultaneously protect user's privacy and make an accurate prediction has not been well studied. By taking these two aspects into consideration, we propose a novel QoS prediction approach for Web service recommendation in this paper. Specifically, we first design a similarity-maintaining privacy preservation (SPP) strategy, which aims to protect the user's privacy and maintain the utility of user data in the meanwhile. Then, we propose a location-aware low-rank matrix factorization (LLMF) algorithm, which employs the L1L1-norm low-rank matrix factorization to improve the model's robustness, and combines the matrix factorization model with two kinds of location information (continent, longitude and latitude) in the prediction process. Experimental results on two publicly available real-world Web service QoS datasets demonstrate the effectiveness of our privacy-preserving QoS prediction approach. Xiaoke Zhu, Xiaoyuan Jing, Di Wu 0014, Zhenyu He 0001, Jicheng Cao, Dong Yue 0001, Lina Wang 0001 |
IEEE Trans. Serv. Comput. | 4 |
| 2020 | Multi-Task Driven Feature Models for Thermal Infrared TrackingabstractExisting deep Thermal InfraRed (TIR) trackers usually use the feature models of RGB trackers for representation. However, these feature models learned on RGB images are neither effective in representing TIR objects nor taking fine-grained TIR information into consideration. To this end, we develop a multi-task framework to learn the TIR-specific discriminative features and fine-grained correlation features for TIR tracking. Specifically, we first use an auxiliary classification network to guide the generation of TIR-specific discriminative features for distinguishing the TIR objects belonging to different classes. Second, we design a fine-grained aware module to capture more subtle information for distinguishing the TIR objects belonging to the same class. These two kinds of features complement each other and recognize TIR objects in the levels of inter-class and intra-class respectively. These two feature models are learned using a multi-task matching framework and are jointly optimized on the TIR tracking task. In addition, we develop a large-scale TIR training dataset to train the network for adapting the model to the TIR domain. Extensive experimental results on three benchmarks show that the proposed algorithm achieves a relative gain of 10% over the baseline and performs favorably against the state-of-the-art methods. Codes and the proposed TIR dataset are available at https://github.com/QiaoLiuHit/MMNet. Qiao Liu 0001, Xin Li 0034, Zhenyu He 0001, Nana Fan, Di Yuan 0002, Wei Liu 0065, Yongsheng Liang 0001 |
AAAI | 3 |
| 2020 | LSOTB-TIR: A Large-Scale High-Diversity Thermal Infrared Object Tracking BenchmarkabstractIn this paper, we present a Large-Scale and high-diversity general Thermal InfraRed (TIR) Object Tracking Benchmark, called LSOTB-TIR, which consists of an evaluation dataset and a training dataset with a total of 1,400 TIR sequences and more than 600K frames. We annotate the bounding box of objects in every frame of all sequences and generate over 730K bounding boxes in total. To the best of our knowledge, LSOTB-TIR is the largest and most diverse TIR object tracking benchmark to date. To evaluate a tracker on different attributes, we define 4 scenario attributes and 12 challenge attributes in the evaluation dataset. By releasing LSOTB-TIR, we encourage the community to develop deep learning based TIR trackers and evaluate them fairly and comprehensively. We evaluate and analyze more than 30 trackers on LSOTB-TIR to provide a series of baselines, and the results show that deep trackers achieve promising performance. Furthermore, we re-train several representative deep trackers on LSOTB-TIR, and their results demonstrate that the proposed training dataset significantly improves the performance of deep TIR trackers. Codes and dataset are available at https://github.com/QiaoLiuHit/LSOTB-TIR. Qiao Liu 0001, Xin Li 0034, Zhenyu He 0001, Chenglong Li 0002, Zikun Zhou, Di Yuan 0002, Jing Li 0071, Kai Yang 0018, Nana Fan, Feng Zheng 0001 |
ACM Multimedia | 3 |
| 2020 | SRHEN: Stepwise-Refining Homography Estimation Network via Parsing Geometric Correspondences in Deep Latent SpaceabstractThe crux of homography estimation is that the homography is characterized by the geometric correspondences between two related images rather than appearance features, which differs from typical image recognition tasks. Existing methods either decompose the task of homography estimation into several individual sub-problems and optimize them sequentially, or attempt to tackle it in an end-to-end manner by delegating the whole task to deep convolutional networks (CNNs). However, it is quite arduous for CNNs to learn the mapping function from appearance features of related images to the homography directly. In this paper, we propose to parse the geometric correspondences between related images explicitly to bridge the gap between deep appearance features and the homography. Furthermore, we propose a coarse-to-fine estimation framework to capture different scale of homography transformations and thus predict the homography in a stepwise-refining manner. Additionally, we propose a pyramidal supervision scheme to leverage an important prior concerning the homography estimation. Extensive experiments on two large-scale datasets demonstrate that our model advances the state-of-the-art performance significantly. Yi Li 0042, Wenjie Pei, Zhenyu He 0001 |
ACM Multimedia | 3 |
| 2020 | Group sparse additive machine with average top-k loss
Peipei Yuan, Xinge You, Hong Chen 0004, Qinmu Peng, Zhou Xu 0003, Xiaoyuan Jing, Zhenyu He 0001 |
Neurocomputing | 8 |
| 2020 | TRBACF: Learning temporal regularized correlation filters for high performance online visual object tracking
Di Yuan 0002, Xiu Shu, Zhenyu He 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2020 | SiamAtt: Siamese attention network for visual tracking
Kai Yang 0018, Zhenyu He 0001, Zikun Zhou, Nana Fan |
Knowl. Based Syst. | 2 |
| 2020 | Learning target-focusing convolutional regression model for visual object tracking
Di Yuan 0002, Nana Fan, Zhenyu He 0001 |
Knowl. Based Syst. | 3 |
| 2020 | Robust visual tracking with correlation filters and metric learning
Di Yuan 0002, Zhenyu He 0001 |
Knowl. Based Syst. | 3 |
| 2020 | Visual object tracking with adaptive structural convolutional network
Di Yuan 0002, Xin Li 0034, Zhenyu He 0001, Qiao Liu 0001, Shuwei Lu |
Knowl. Based Syst. | 3 |
| 2020 | Dual-regression model for visual tracking
Xin Li 0034, Qiao Liu 0001, Nana Fan, Zikun Zhou, Zhenyu He 0001, Xiaoyuan Jing |
Neural Networks | 5 |
| 2020 | PTB-TIR: A Thermal Infrared Pedestrian Tracking BenchmarkabstractThermal infrared (TIR) pedestrian tracking is one of the important components among numerous applications of computer vision, which has a major advantage: it can track pedestrians in total darkness. The ability to evaluate the TIR pedestrian tracker fairly, on a benchmark dataset, is significant for the development of this field. However, there is not a benchmark dataset. In this paper, we develop a TIR pedestrian tracking dataset for the TIR pedestrian tracker evaluation. The dataset includes 60 thermal sequences with manual annotations. Each sequence has nine attribute labels for the attribute based evaluation. In addition to the dataset, we carry out the large-scale evaluation experiments on our benchmark dataset using nine publicly available trackers. The experimental results help us understand the strengths and weaknesses of these trackers. In addition, in order to gain more insight into the TIR pedestrian tracker, we divide its functions into three components: feature extractor, motion model, and observation model. Then, we conduct three comparison experiments on our benchmark dataset to validate how each component affects the tracker's performance. The findings of these experiments provide some guidelines for future research. Qiao Liu 0001, Zhenyu He 0001, Xin Li 0034, Yuan Zheng 0002 |
IEEE Trans. Multim. | 2 |
| 2020 | Adaptive Weighted Sparse Principal Component Analysis for Robust Unsupervised Feature SelectionabstractCurrent unsupervised feature selection methods cannot well select the effective features from the corrupted data. To this end, we propose a robust unsupervised feature selection method under the robust principal component analysis (PCA) reconstruction criterion, which is named the adaptive weighted sparse PCA (AW-SPCA). In the proposed method, both the regularization term and the reconstruction error term are constrained by the ℓ2,1-norm: the ℓ2,1-norm regularization term plays a role in the feature selection, while the ℓ2,1-norm reconstruction error term plays a role in the robust reconstruction. The proposed method is in a convex formulation, and the selected features by it can be used for robust reconstruction and clustering. Experimental results demonstrate that the proposed method can obtain better reconstruction and clustering performance, especially for the corrupted data. Shuangyan Yi, Zhenyu He 0001, Xiaoyuan Jing, Yi Li 0042, Yiu-Ming Cheung, Feiping Nie 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2019 | Target-Aware Deep TrackingabstractExisting deep trackers mainly use convolutional neural networks pre-trained for the generic object recognition task for representations. Despite demonstrated successes for numerous vision tasks, the contributions of using pre-trained deep features for visual tracking are not as significant as that for object recognition. The key issue is that in visual tracking the targets of interest can be arbitrary object class with arbitrary forms. As such, pre-trained deep features are less effective in modeling these targets of arbitrary forms for distinguishing them from the background. In this paper, we propose a novel scheme to learn target-aware features, which can better recognize the targets undergoing significant appearance variations than pre-trained deep features. To this end, we develop a regression loss and a ranking loss to guide the generation of target-active and scale-sensitive features. We identify the importance of each convolutional filter according to the back-propagated gradients and select the target-aware features based on activations for representing the targets. The target-aware features are integrated with a Siamese matching network for visual tracking. Extensive experimental results show that the proposed algorithm performs favorably against the state-of-the-art methods in terms of accuracy and speed. Xin Li 0034, Chao Ma 0004, Baoyuan Wu, Zhenyu He 0001, Ming-Hsuan Yang 0001 |
CVPR | 4 |
| 2019 | Multiple pedestrian tracking by combining particle filter and network flow model
Zhenyu He 0001, Jianjun Hu |
Neurocomputing | 3 |
| 2019 | Low-Rank Projection Learning via Graph Embedding
Yingyi Liang, Xiaohuan Lu, Zhenyu He 0001, Hongpeng Wang 0002 |
Neurocomputing | 4 |
| 2019 | Distracter-aware tracking via correlation filter
Xiaohuan Lu, Jing Li 0071, Zhenyu He 0001, Wei Wang 0011, Hongzhi Wang 0001 |
Neurocomputing | 3 |
| 2019 | Region-filtering correlation tracking
Nana Fan, Jing Li 0071, Zhenyu He 0001, Chunkai Zhang, Xin Li 0034 |
Knowl. Based Syst. | 3 |
| 2019 | Hierarchical spatial-aware Siamese network for thermal infrared object tracking
Xin Li 0034, Qiao Liu 0001, Nana Fan, Zhenyu He 0001, Hongzhi Wang 0001 |
Knowl. Based Syst. | 4 |
| 2019 | Dual Pursuit for Subspace LearningabstractIn general, low-rank representation (LRR) aims to find the lowest rank representation with respect to a dictionary. In fact, the dictionary is a key aspect of low-rank representation. However, a lot of low-rank representation methods usually use the data itself as a dictionary (i.e., a fixed dictionary), which may degrade their performances due to the lack of clustering ability of a fixed dictionary. To this end, we propose learning a locality-preserving dictionary instead of the fixed dictionary for low-rank representation, where the locality-preserving dictionary is constructed by using a graph regularization technique to capture the intrinsic geometric structure of the dictionary and, hence, the locality-preserving dictionary has an underlying clustering ability. In this way, the obtained low-rank representation via the locality-preserving dictionary has a better grouping-effect representation. Inversely, a better grouping-effect representation can help to learn a good dictionary. The locality-preserving dictionary and the grouping-effect representation interact with each other, where dual pursuit is called. The proposed method, namely, Dual Pursuit for Subspace Learning, provides us with a robust method for clustering and classification simultaneously, and compares favorably with the other state-of-the-art methods. Shuangyan Yi, Yingyi Liang, Zhenyu He 0001, Yi Li 0042, Yiu-Ming Cheung |
IEEE Trans. Multim. | 3 |
| 2018 | Enhancing social network privacy with accumulated non-zero prior knowledge
Xiaofeng Zhang 0002, Zhenyu He 0001 |
Inf. Sci. | 5 |
| 2018 | Unified Sparse Subspace Learning via Self-Contained RegressionabstractIn order to improve the interpretation of principal components, many sparse principal component analysis (PCA) methods have been proposed by in the form of self-contained regression-type. In this paper, we generalize the steps needed to move from PCA-like methods to its self-contained regression-type, and propose a joint sparse pixel weighted PCA method. More specifically, we generalize a self-contained regression-type framework of graph embedding. Unlike the regression-type of graph embedding relying on the regular low-dimensional data, the self-contained regression-type framework does not rely on the regular low-dimensional data of graph embedding. The learned low-dimensional data in the form of self-contained regression theoretically approximates to the regular low-dimensional data. Under this self-contained regression-type, sparse regularization term can be arbitrarily added, and hence, the learned sparse regression coefficients can interpret the low-dimensional data. By using the joint sparse ℓ2,1-norm regularizer, a sparse self-contained regression-type of pixel weighted PCA can be produced. Experiments on six data sets demonstrate that the proposed method is both feasible and effective. Shuangyan Yi, Zhenyu He 0001, Yiu-Ming Cheung |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2017 | Deep convolutional neural networks for thermal infrared object tracking
Qiao Liu 0001, Xiaohuan Lu, Zhenyu He 0001, Chunkai Zhang |
Knowl. Based Syst. | 3 |
| 2017 | Joint sparse principal component analysis
Shuangyan Yi, Zhihui Lai 0001, Zhenyu He 0001, Yiu-Ming Cheung, Yang Liu 0007 |
Pattern Recognit. | 3 |
| 2017 | Robust Object Tracking via Key Patch Sparse RepresentationabstractMany conventional computer vision object tracking methods are sensitive to partial occlusion and background clutter. This is because the partial occlusion or little background information may exist in the bounding box, which tends to cause the drift. To this end, in this paper, we propose a robust tracker based on key patch sparse representation (KPSR) to reduce the disturbance of partial occlusion or unavoidable background information. Specifically, KPSR first uses patch sparse representations to get the patch score of each patch. Second, KPSR proposes a selection criterion of key patch to judge the patches within the bounding box and select the key patch according to its location and occlusion case. Third, KPSR designs the corresponding contribution factor for the sampled patches to emphasize the contribution of the selected key patches. Comparing the KPSR with eight other contemporary tracking methods on 13 benchmark video data sets, the experimental results show that the KPSR tracker outperforms classical or state-of-the-art tracking methods in the presence of partial occlusion, background clutter, and illumination change. Zhenyu He 0001, Shuangyan Yi, Yiu-Ming Cheung, Xinge You, Yuan Yan Tang |
IEEE Trans. Cybern. | 1 |
| 2016 | Simultaneous Dual-Views Reconstruction with Adaptive Dictionary and Low-Rank RepresentationabstractLow-Rank Representation (LRR) is an effective self-expressiveness method, which uses the observed data itself as the dictionary to reconstruct the original data. LRR focuses on representing the global low-dimensional information, but ignores the real fact that data often resides on low-dimensional manifolds embedded in a high-dimensional data. Therefore, LRR can not capture the non-linear geometric structures within data. As well known, locality preserving projections (LPP) is able to preserve the intrinsic geometry structure embedded in high-dimensional data. To this end, we treat the projected data by LPP as an adaptive dictionary, and such a dictionary can capture the intrinsic geometry structures of data. In this way, our method is in favor of the global low-rank representation. Specifically, the proposed method provides a way to reconstruct the original data from two views, and hence we call this proposed method as Simultaneous Dual-Views Reconstruction with Adaptive Dictionary and Low-Rank Representation. The proposed method can be used for unsupervised feature extraction and subspace clustering. Experiments on benchmark databases show the excellent performance of this proposed method in comparison with other state-of-the-art methods. Shuangyan Yi, Zhenyu He 0001, Yi Li 0042, Yiu-Ming Cheung |
ICPR | 2 |
| 2016 | A multi-view model for visual tracking via correlation filters
Xin Li 0034, Qiao Liu 0001, Zhenyu He 0001, Hongpeng Wang 0002, Chunkai Zhang |
Knowl. Based Syst. | 3 |
| 2016 | Visual tracking via exemplar regression model
Qiao Liu 0001, Zhenyu He 0001, Xiaofeng Zhang 0002 |
Knowl. Based Syst. | 3 |
| 2016 | Connected Component Model for Multi-Object TrackingabstractIn multi-object tracking, it is critical to explore the data associations by exploiting the temporal information from a sequence of frames rather than the information from the adjacent two frames. Since straightforwardly obtaining data associations from multi-frames is an NP-hard multi-dimensional assignment (MDA) problem, most existing methods solve this MDA problem by either developing complicated approximate algorithms, or simplifying MDA as a 2D assignment problem based upon the information extracted only from adjacent frames. In this paper, we show that the relation between associations of two observations is the equivalence relation in the data association problem, based on the spatial-temporal constraint that the trajectories of different objects must be disjoint. Therefore, the MDA problem can be equivalently divided into independent subproblems by equivalence partitioning. In contrast to existing works for solving the MDA problem, we develop a connected component model (CCM) by exploiting the constraints of the data association and the equivalence relation on the constraints. Based upon CCM, we can efficiently obtain the global solution of the MDA problem for multi-object tracking by optimizing a sequence of independent data association subproblems. Experiments on challenging public data sets demonstrate that our algorithm outperforms the state-of-the-art approaches. Zhenyu He 0001, Xin Li 0034, Xinge You, Dacheng Tao, Yuan Yan Tang |
IEEE Trans. Image Process. | 1 |
| 2015 | One global optimization method in network flow model for multiple object tracking
Zhenyu He 0001, Hongpeng Wang 0002, Xinge You, C. L. Philip Chen |
Knowl. Based Syst. | 1 |
| 2015 | Single object tracking via robust combination of particle filter and sparse representation
Shuangyan Yi, Zhenyu He 0001, Xinge You, Yiu-Ming Cheung |
Signal Process. | 2 |
| 2015 | A robust local sparse tracker with global consistency constraint
Xinhua You, Xin Li 0034, Zhenyu He 0001, Xiaofeng Zhang 0002 |
Signal Process. | 3 |
| 2014 | A Sentence Vector Based Over-Sampling Method for Imbalanced Emotion Classification
Tao Chen 0026, Ruifeng Xu 0001, Qin Lu 0001, Bin Liu 0014, Jun Xu 0007, Zhenyu He 0001 |
CICLing (2) | 7 |
| 2014 | Reader Emotion Prediction Using Concept and Concept Sequence Features in News Headlines
Yuanlin Yao, Ruifeng Xu 0001, Qin Lu 0001, Bin Liu 0014, Jun Xu 0007, Chengtian Zou, Li Yuan 0002, Shuwei Wang, Zhenyu He 0001 |
CICLing (2) | 10 |
| 2014 | A novel joint tracker based on occlusion detection
Xin Li 0034, Zhenyu He 0001, Xinge You, C. L. Philip Chen |
Knowl. Based Syst. | 2 |
| 2008 | Writer identification using global wavelet-based features
Zhenyu He 0001, Xinge You, Yuan Yan Tang |
Neurocomputing | 1 |
| 2008 | Writer identification of Chinese handwriting documents using hidden Markov tree model
Zhenyu He 0001, Xinge You, Yuan Yan Tang |
Pattern Recognit. | 1 |
| 2006 | Handwriting-based personal identificationabstractHandwriting-based personal identification, which is also called handwriting-based writer identification, is an active research topic in pattern recognition. Despite continuous effort, offline handwriting-based writer identification still remains as a challenging problem because writing features can only be extracted from the handwriting image. As a result, plenty of dynamic writing information, which is very valuable for writer identification, is unavailable for offline writer identification. In this paper, we present a novel wavelet-based Generalized Gaussian Density (GGD) method for offline writer identification. Compared with the 2-D Gabor model, which is currently widely acknowledged as a good method for offline handwriting identification, GGD method not only achieves a better identification accuracy but also greatly reduces the elapsed time on calculation in our experiments. Zhenyu He 0001, Xinge You, Yuan Yan Tang, Bin Fang 0001, Jianwei Du |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2005 | A Novel Method for Off-line Handwriting-based Writer IdentificationabstractHandwriting-based writer identification is a hot research topic in the pattern recognition field. Nowadays, online handwriting-based writer identification is steadily growing toward its maturity. On the contrary, offline handwriting-based writer identification still remains as a challenging problem because writing features only can be extracted from the handwriting image in this situation. As a result, plenty of dynamic writing information, which is very valuable for writer identification, is lost. At present, 2D Gabor filter method is widely acknowledged as a good method for offline handwriting identification, however it still suffers from some inherent disadvantages, such as the high computational cost. In this paper, we present a novel wavelet-based GGD method to replace the traditional 2D Gabor filters. Shown in our experiments, this novel method not only achieves better experiment results but also greatly reduces the elapsed time on calculation. Zhenyu He 0001, Yuan Yan Tang, Bin Fang 0001, Jianwei Du, Xinge You |
ICDAR | 1 |
| 2005 | Similarity Measurement for Off-Line Signature Verification
Xinge You, Bin Fang 0001, Zhenyu He 0001, Yuan Yan Tang |
ICIC (1) | 3 |
| 2005 | Locating Vessel Centerlines in Retinal Images Using Wavelet Transform: A Multilevel Approach
Xinge You, Bin Fang 0001, Yuan Yan Tang, Zhenyu He 0001, Jian Huang 0009 |
ICIC (1) | 4 |
| 2005 | A contourlet-based method for writer identificationabstractHandwriting-based writer identification is a hot research topic in the field of pattern recognition. Typically, there are four modes of writer identification: on-line text-dependent, on-line text-independent, off-line text-dependent, off-line text-independent; and off-line text-independent is the most challenging problem among them because many valuable writing features are not available in this case, such as shape features, dynastic writing information and etc. In this paper, we focus on the text-independent writer identification based on off-line Chinese handwriting and present a new contourlet-based GGD (Generalized Gaussian Density) method. This novel method achieves a good experiment result in our experiments. Zhenyu He 0001, Yuan Yan Tang, Xinge You |
SMC | 1 |
| 2005 | A wavelet-based approach to ridge thinning in fingerprint imagesabstractAs a global feature of fingerprints, the thinning of ridges, extraction of minutiae and computation of orientation field are very important for automatic fingerprint recognition. Many algorithms have been proposed for their computation and estimation, but their results are unsatisfactory, especially for poor quality fingerprint images. In this paper, a robust wavelet-based method to create thinned ridge map of fingerprint for automatic recognition is proposed. Properties of modulus minima based on the spline wavelet function are substantially investigated. Desirable characteristics show that this method is suitable to describe the skeleton of the ridge of the fingerprint image. A multi-scale thinning algorithm based on the modulus minima of wavelet transform is presented. The proposed algorithm is able to improve the skeleton representation of the ridge of the fingerprint without side-effects and limitations of the existing methods. The thinned ridge map can facilitate the extraction of the minutiae for matching in fingerprint recognition. Experiments have been conducted to validate the effectiveness and efficiency of the proposed method. Xinge You, Bin Fang 0001, Yuan Yan Tang, Zhenyu He 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 4 |