VLDB 2026 Research / reviewers in the wild / expert
Zikun Zhou
dblp:271/8084
· DBLP profile ↗
34ranked-venue papers
9as first author
31since 2021 · last 2026
0000-0002-2687-7762ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 22 · 7 first-author · 21 since 2021Artificial intelligence and machine learning · 19 · 5 first-author · 17 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Deeply-conditioned image compression via self-generated priors
Zhineng Zhao, Zhihai He, Zikun Zhou, Siwei Ma 0001, Yaowei Wang 0001 |
Neurocomputing | 3 |
| 2026 | Harnessing Vision-Language Pretrained Models With Temporal-Aware Adaptation for Referring Video Object SegmentationabstractReferring Video Object Segmentation (RVOS) is a task that involves segmenting target objects in a video based on the given referring expressions. It is critical for video editing and analysis. The crux of RVOS is to model dense text-video relations to associate abstract linguistic concepts with dynamic visual contents at pixel-level. Most RVOS methods typically use vision and language models pretrained independently as backbones, mapping images and texts to uncoupled feature spaces. As a result, they must learn Vision-Language (VL) relation modeling from scratch. Vision-Language Pretrained (VLP) models have achieved remarkable success. Inspired by this, we propose to explore relation modeling for RVOS based on their aligned VL feature space. Nevertheless, transferring VLP models to RVOS is deceptively challenging, due to the gap between static image/region-level pretraining and dynamic pixel-level prediction. To bridge this gap, we introduce a framework named VLP-RVOS, which harnesses VLP models for RVOS through temporal-aware adaptation. We first propose temporal-aware prompt-tuning to adapt pretrained representations for pixel-level prediction and empower the vision encoder to model temporal contexts. We further customize a cube-frame attention mechanism for robust spatial-temporal reasoning. Besides, we propose to perform multi-stage VL relation modeling while and after feature extraction for comprehensive understanding. Extensive experiments demonstrate that VLP-RVOS performs favorably against state-of-the-art algorithms and generalizes well. Our codes are available at https://github.com/xwt909090/VLP-RVOS. Zikun Zhou, Wentao Xiong, Li Zhou 0017, Xin Li 0034, Zhenyu He 0001, Yaowei Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2026 | Enhancing Federated Domain Generalization by Data Influences on Global Model UpdateabstractWith the popularity of federated learning, federated domain generalization (FedDG) has attracted more and more attention. Existing works of federated learning indicate that the generalization performance of the global model can be improved when the global model is obtained by aggregating local models according to suitable weights. However, existing methods to calculate weights do not fully utilize the data influences on the global model update, which gives us an opportunity to improve the generalization performance of the global model further. In this paper, we propose the method DI (data influences), which utilizes data influences on the global model update to calculate dynamical weights of local model in each round of training. Specifically, the first component data influence calculator (DIC) of DI calculates local weights of local model from the influences of data on the global model update and we introduce the influence function to complete the calculation process. The second component data influence adjuster (DIA) of DI calculates global weights (which are used in the aggregation process of the global model) from local weights. Extensive experiments indicate that our method improves the generalization performance of models significantly. In particular, our method improves model accuracy on benchmark datasets PACS, OfficeHome, and Office-31 by 1.79%, 1.61%, and 2.39% on average, respectively. Source code is publicly available at github-https://github.com/zikunZHOUHH/Fed-DI. Wen Huang 0002, Zikun Zhou, Weixin Zhao, Xingyi Wang, Jian Peng 0002 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2026 | Prototype Perturbation for Relaxing Alignment Constraints in Backward-Compatible LearningabstractThe traditional paradigm to update retrieval models requires re-computing the embeddings of the gallery data, a time-consuming and computationally intensive process known as backfilling. To circumvent backfilling, Backward-Compatible Learning (BCL) has been widely explored, which aims to train a new model compatible with the old one. Many previous works focus on effectively aligning the embeddings of the new model with those of the old one to enhance backward compatibility. Nevertheless, such strong alignment constraints would compromise the discriminative ability of the new model, particularly when different classes are closely clustered and hard to distinguish in the old feature space. To address this issue, we propose to relax the constraints by introducing perturbations to the old feature prototypes. This allows us to align the new feature space with a pseudo-old feature space defined by these perturbed prototypes, thereby preserving the discriminative ability of the new model in backward-compatible learning. We have developed two approaches for calculating the perturbations: Neighbor-Driven Prototype Perturbation (NDPP) and Optimization-Driven Prototype Perturbation (ODPP). Particularly, they take into account the feature distributions of not only the old but also the new models to obtain proper perturbations along with new model updating. Extensive experiments on the landmark and commodity datasets demonstrate that our approaches perform favorably against state-of-the-art BCL algorithms. Zikun Zhou, Yushuai Sun, Wenjie Pei, Xin Li 0034, Yaowei Wang 0001 |
IEEE Trans. Multim. | 1 |
| 2025 | Instruction-guided Multi-Granularity Segmentation and Captioning with Large Multimodal ModelabstractLarge Multimodal Models (LMMs) have significantly progressed by extending large language models. Building on this progress, the latest developments in LMMs demonstrate the ability to generate dense pixel-wise segmentation by integrating segmentation models. Despite the innovations, existing works’ textual responses and segmentation masks remain at the instance level, showing limited ability to perform fine-grained understanding and segmentation even provided with detailed textual cues. To overcome this limitation, we introduce a Multi-Granularity Large Multimodal Model (MGLMM), which is capable of seamlessly adjusting the granularity of Segmentation and Captioning (SegCap) following user instructions, from panoptic SegCap to fine-grained SegCap. We name such a new task Multi-Granularity Segmentation and Captioning (MGSC). Observing the lack of a benchmark for model training and evaluation over the MGSC task, we establish a benchmark with aligned masks and captions in multi-granularity using our customized automated annotation pipeline. This benchmark comprises 10K images and more than 30K image-question pairs. We will release our dataset along with the implementation of our automated dataset annotation pipeline for further research. Besides, we propose a novel unified SegCap data format to unify heterogeneous segmentation datasets; it effectively facilitates learning to associate object concepts with visual features during multi-task training. Extensive experiments demonstrate that our MGLMM excels at tackling more than eight downstream tasks and achieves state-of-the-art performance in MGSC, GCG, image captioning, referring segmentation, multiple/empty segmentation, and reasoning segmentation. The great properties and versatility of MGLMM underscore its potential impact on advancing multimodal research. Li Zhou 0017, Zenghui Sun, Zikun Zhou, Jinsong Lan |
AAAI | 4 |
| 2025 | Improving Federated Domain Generalization Through Dynamical Weights Calculated from Data Influences on Global Model UpdateabstractWith the popularity of federated learning, federated domain generalization (FedDG) has attracted more and more attentions. Existing works of federated learning indicate that the generalization performance of the global model can be improved when the global model is obtained by aggregating local models according to a suitable weights. However, the existing methods to calculate weights do not fully utilize the data influences on the global model update, which gives us an opportunity to improve the generalization performance of the global model further. In this paper, we propose the method DI (data influences), which utilizes the data influences on the global model update to calculate dynamical weights of local model in each round of training. Specifically, the first component data influences calculator (DIC) of DI calculates the local weights of local model from the influences of each data on the global model update and we introduce the influences function to complete the calculation process. The second component data influences adjuster (DIA) of DI calculates the global weights (which are used in the aggregation process of the global model) from local weights. Extensive experiments indicate that our method improves the generalization performance of models significantly. In particular, our method improves model accuracy on benchmark datasets PACS, OfficeHome, and Office-31 by 1.79%, 1.61%, and 2.39% on average, respectively. Source code is publicly available at github. Zikun Zhou, Wen Huang 0002, Xingyi Wang, Jian Peng 0002, Feihu Huang 0002 |
AAAI | 1 |
| 2025 | MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language TrackingabstractThe vision-language tracking task aims to perform object tracking based on various modality references. Existing Transformer-based vision-language tracking methods have made remarkable progress by leveraging the global modeling ability of self-attention. However, current approaches still face challenges in effectively exploiting the temporal information and dynamically updating reference features during tracking. Recently, the State Space Model (SSM), known as Mamba, has shown astonishing ability in efficient long-sequence modeling. Particularly, its state space evolving process demonstrates promising capabilities in memorizing multimodal temporal information with linear complexity. Witnessing its success, we propose a Mamba-based vision-language tracking model to exploit its state space evolving ability in temporal space for robust multimodal tracking, dubbed MambaVLT. In particular, our approach mainly integrates a time-evolving hybrid state space block and a selective locality enhancement block, to capture contextual information for multimodal modeling and adaptive reference feature update. Besides, we introduce a modality-selection module that dynamically adjusts the weighting between visual and language references, mitigating potential ambiguities from either reference type. Extensive experimental results show that our method performs favorably against state-of-the-art trackers across diverse benchmarks. Xinqi Liu, Li Zhou 0017, Zikun Zhou, Jianqiu Chen, Zhenyu He 0001 |
CVPR | 3 |
| 2025 | Learning Compatible Multi-Prize Subnetworks for Asymmetric RetrievalabstractAsymmetric retrieval is a typical scenario in real-world retrieval systems, where compatible models of varying capacities are deployed on platforms with different resource configurations. Existing methods generally train pre-defined networks or subnetworks with capacities specifically designed for pre-determined platforms, using compatible learning. Nevertheless, these methods suffer from limited flexibility for multi-platform deployment. For example, when introducing a new platform into the retrieval systems, developers have to train an additional model at an appropriate capacity that is compatible with existing models via backward-compatible learning. In this paper, we propose a Prunable Network with self-compatibility, which allows developers to generate compatible subnetworks at any desired capacity through post-training pruning. Thus it allows the creation of a sparse subnetwork matching the resources of the new platform without additional training. Specifically, we optimize both the architecture and weight of subnetworks at different capacities within a dense network in compatible learning. We also design a conflict-aware gradient integration scheme to handle the gradient conflicts between the dense network and subnetworks during compatible learning. Extensive experiments on diverse benchmarks and visual backbones demonstrate the effectiveness of our method. The code will be made publicly available. Yushuai Sun, Zikun Zhou, Dongmei Jiang, Yaowei Wang 0001, Jun Yu 0002, Guangming Lu 0002, Wenjie Pei |
CVPR | 2 |
| 2025 | ZeroBP: Learning Position-Aware Correspondence for Zero-Shot 6D Pose Estimation in Bin-PickingabstractBin-picking is a practical and challenging robotic manipulation task, where accurate 6D pose estimation plays a pivotal role. The workpieces in bin-picking are typically texture-less and randomly stacked in a bin, which poses a significant challenge to 6D pose estimation. Existing solutions are typically learning-based methods, which require object-specific training. Their efficiency of practical deployment for novel workpieces is highly limited by data collection and model retraining. Zero-shot 6D pose estimation is a potential approach to address the issue of deployment efficiency. Nevertheless, existing zero-shot 6D pose estimation methods are designed to leverage feature matching to establish point-to-point correspondences for pose estimation, which is less effective for workpieces with textureless appearances and ambiguous local regions. In this paper, we propose ZeroBP, a zero-shot pose estimation frame-work designed specifically for the bin-picking task. ZeroBP learns Position-Aware Correspondence (PAC) between the scene instance and its CAD model, leveraging both local features and global positions to resolve the mismatch issue caused by ambiguous regions with similar shapes and appearances. Extensive experiments on the ROBI dataset demonstrate that ZeroBP outperforms state-of-the-art zero-shot pose estimation methods, achieving an improvement of 9.1 % in average recall of correct poses. Jianqiu Chen, Zikun Zhou, Xin Li 0034, Tianpeng Bao, Zhenyu He 0001 |
ICRA | 2 |
| 2025 | TorVIA: A novel encrypted video identification method based on Tor transmission characteristics
Juncheng Lu, Zikun Zhou, Hua Wu 0004, Guang Cheng 0001 |
Comput. Networks | 2 |
| 2025 | ZeroPose: CAD-Prompted Zero-Shot Object 6D Pose Estimation in Cluttered ScenesabstractMany robotics and industry applications have a high demand for the capability to estimate the 6D pose of novel objects from the cluttered scene. However, existing classic pose estimation methods are object-specific, which can only handle the specific objects seen during training. When applied to a novel object, these methods necessitate a cumbersome onboarding process, which involves extensive dataset preparation and model retraining. The extensive duration and resource consumption of onboarding limit their practicality in real-world applications In this paper, we introduce ZeroPose, a novel zero-shot framework that performs pose estimation following a Discovery-Orientation-Registration (DOR) inference pipeline. This framework generalizes to novel objects without requiring model retraining. Given the CAD model of a novel object, ZeroPose enables in seconds onboarding time to extract visual and geometric embeddings from the CAD model as a prompt. With the prompting of the above embeddings, DOR can discover all related instances and estimate their 6D poses without additional human interaction or presupposing scene conditions. Compared with existing zero-shot methods solved by the render-and-compare paradigm, the DOR pipeline formulates the object pose estimation into a feature-matching problem, which avoids time-consuming online rendering and improves efficiency. Experimental results on the seven datasets show that ZeroPose as a zero-shot method achieves comparable performance with object-specific training methods and outperforms the state-of-the-art zero-shot method with 50x inference speed improvement. Jianqiu Chen, Zikun Zhou, Mingshan Sun, Rui Zhao 0001, Tianpeng Bao, Zhenyu He 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Reliability-Guided Hierarchical Memory Network for Scribble-Supervised Video Object SegmentationabstractThis article aims to solve the video object segmentation (VOS) task in a scribble-supervised manner, in which VOS models are not only initialized with sparse target scribbles for inference but also trained by sparse scribble annotations. Thus, the annotation burdens for both initialization and training can be substantially lightened. The difficulties of scribble-supervised VOS lie in two aspects: 1) it demands a strong reasoning ability to carefully segment the target given only a sparse initial target scribble and 2) it necessitates learning dense prediction from sparse scribble annotations during training, requiring powerful learning capability. In this work, we propose a reliability-guided hierarchical memory network (RHMNet) for this task, which segments the target in a stepwise expanding strategy w.r.t. the memory reliability level. To be specific, RHMNet maintains a reliability-guided memory bank. It first uses the high-reliability memory to locate the region with high reliability belonging to the target, i.e., highly similar to the initial target scribble. Then, it expands the located high-reliability region to the entire target conditioned on the region itself and all existing memories. In addition, we propose a scribble-supervised learning mechanism to facilitate the model learning for dense prediction. It exploits the pixel-level relations within a single frame and the instance-level variations across multiple frames to take full advantage of the scribble annotations in sequence training samples. The favorable performance on four popular benchmarks demonstrates that our method is promising. Our project is available at: https://github.com/mkg1204/RHMNet-for-SSVOS. Zikun Zhou, Kaige Mao, Wenjie Pei, Hongpeng Wang 0002, Yaowei Wang 0001, Zhenyu He 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Robust 3D Tracking with Quality-Aware Shape Completionabstract3D single object tracking remains a challenging problem due to the sparsity and incompleteness of the point clouds. Existing algorithms attempt to address the challenges in two strategies. The first strategy is to learn dense geometric features based on the captured sparse point cloud. Nevertheless, it is quite a formidable task since the learned dense geometric features are with high uncertainty for depicting the shape of the target object. The other strategy is to aggregate the sparse geometric features of multiple templates to enrich the shape information, which is a routine solution in 2D tracking. However, aggregating the coarse shape representations can hardly yield a precise shape representation. Different from 2D pixels, 3D points of different frames can be directly fused by coordinate transform, i.e., shape completion. Considering that, we propose to construct a synthetic target representation composed of dense and complete point clouds depicting the target shape precisely by shape completion for robust 3D tracking. Specifically, we design a voxelized 3D tracking framework with shape completion, in which we propose a quality-aware shape completion mechanism to alleviate the adverse effect of noisy historical predictions. It enables us to effectively construct and leverage the synthetic target representation. Besides, we also develop a voxelized relation modeling module and box refinement module to improve tracking performance. Favorable performance against state-of-the-art algorithms on three benchmarks demonstrates the effectiveness and generalization ability of our method. Zikun Zhou, Guangming Lu 0002, Jiandong Tian, Wenjie Pei |
AAAI | 2 |
| 2024 | Bilateral Event Mining and Complementary for Event Stream Super-ResolutionabstractEvent Stream Super-Resolution (ESR) aims to address the challenge of insufficient spatial resolution in event streams, which holds great significance for the application of event cameras in complex scenarios. Previous works for ESR often process positive and negative events in a mixed paradigm. This paradigm limits their ability to effectively model the unique characteristics of each event and mutually refine each other by considering their correlations. In this paper, we propose a bilateral event mining and complementary network (BMCNet) to fully leverage the potential of each event and capture the shared information to complement each other simultaneously. Specifically, we resort to a two-stream network to accomplish comprehensive mining of each type of events individually. To facilitate the exchange of information between two streams, we propose a bilateral information exchange (BIE) module. This module is layer-wisely embedded between two streams, enabling the effective propagation of hierarchical global information while alleviating the impact of invalid information brought by inherent characteristics of events. The experimental results demonstrate that our approach outperforms the previous state-of-the-art methods in ESR, achieving performance improvements of over 11% on both real and synthetic datasets. Moreover, our method significantly enhances the performance of event-based downstream tasks such as object recognition and video reconstruction. Our code is available at https://github.com/Lqm26/BMCNet-ESR. Zhilin Huang, Quanmin Liang, Yijie Yu 0001, Chujun Qin, Xiawu Zheng, Kai Huang 0001, Zikun Zhou, Wenming Yang |
CVPR | 7 |
| 2024 | RTracker: Recoverable Tracking via PN Tree Structured MemoryabstractExisting tracking methods mainly focus on learning better target representation or developing more robust prediction models to improve tracking performance. While tracking performance has significantly improved, the target loss issue occurs frequently due to tracking failures, complete occlusion, or out-of-view situations. However, con-siderably less attention is paid to the self-recovery issue of tracking methods, which is crucial for practical applications. To this end, we propose a recoverable tracking framework, RTracker, that uses a tree-structured memory to dynamically associate a tracker and a detector to enable self-recovery ability. Specifically, we propose a Positive-Negative Tree-structured memory to chronologically store and maintain positive and negative target samples. Upon the PN tree memory, we develop corresponding walking rules for determining the state of the target and define a set of control flows to unite the tracker and the detector in different tracking scenarios. Our core idea is to use the support samples of positive and negative target categories to establish a relative distance-based criterion for a reliable assessment of target loss. The favorable performance in comparison against the state-of-the-art methods on nu-merous challenging benchmarks demonstrates the effectiveness of the proposed algorithm. All the source code and trained models will be released at https://github.com/NorahGreen/RTracker. Yuqing Huang, Xin Li 0034, Zikun Zhou, Yaowei Wang 0001, Zhenyu He 0001, Ming-Hsuan Yang 0001 |
CVPR | 3 |
| 2024 | Interaction-based Retrieval-augmented Diffusion Models for Protein-specific 3D Molecule GenerationabstractGenerating ligand molecules that bind to specific protein targets via generative models holds substantial promise for advancing structure-based drug design. Existing methods generate molecules from scratch without reference or template ligands, which poses challenges in model optimization and may yield suboptimal outcomes. To address this problem, we propose an innovative interaction-based retrieval-augmented diffusion model named IRDiff to facilitate target-aware molecule generation. IRDiff leverages a curated set of ligand references, i.e., those with desired properties such as high binding affinity, to steer the diffusion model towards synthesizing ligands that satisfy design criteria. Specifically, we utilize a protein-molecule interaction network (PMINet), which is pretrained with binding affinity signals to: (i) retrieve target-aware ligand molecules with high binding affinity to serve as references, and (ii) incorporate essential protein-ligand binding structures for steering molecular diffusion generation with two effective augmentation mechanisms, i.e., retrieval augmentation and self augmentation. Empirical studies on CrossDocked2020 dataset show IRDiff can generate molecules with more realistic 3D structures and achieve state-of-the-art binding affinities towards the protein targets, while maintaining proper molecular properties. The codes and models are available at https://github.com/YangLing0818/IRDiff Zhilin Huang, Ling Yang 0006, Xiangxin Zhou, Chujun Qin, Yijie Yu 0001, Xiawu Zheng, Zikun Zhou, Wentao Zhang 0001, Yu Wang 0008, Wenming Yang |
ICML | 7 |
| 2024 | Simplifying Cross-modal Interaction via Modality-Shared Features for RGBT TrackingabstractThermal infrared(TIR) data exhibits higher tolerance to extreme environments, making it a valuable complement to RGB data in tracking tasks. RGBT tracking aims to leverage information from RGB and TIR images for stable and robust tracking. However, existing RGBT tracking methods face challenges due to significant modality differences and selective emphasis on interactive information, leading to inefficiencies in the cross-modal interaction. To address these issues, we propose a novel Integrating Interaction into Modality-shared Features with ViT(IIMF) framework, which is a simplified cross-modal interaction network including modality-shared, RGB modality-specific, and TIR modality-specific branches. The Modality-shared branch aggregates modality-shared information and implements inter-modal interaction. Specifically, our approach first extracts modality-shared features from RGB and TIR features with a cross-attention mechanism. Furthermore, we design a Cross-Attention-based Modality-shared Information Aggregation(CAMIA) module to further aggregate modality-shared information with modality-shared tokens. We evaluate our model on three widely-used benchmark datasets and extensive experiments demonstrate that our method achieves state-of-the-art performance. All the source code are released at https://github.com/Liqiu-Chen/IIMF. Liqiu Chen, Yuqing Huang, Zikun Zhou, Zhenyu He 0001 |
ACM Multimedia | 4 |
| 2024 | Motion-aware Latent Diffusion Models for Video Frame InterpolationabstractWith the advancement of AIGC, video frame interpolation (VFI) has become a crucial component in existing video generation frameworks, attracting widespread research interest. For the VFI task, the motion estimation between neighboring frames plays a crucial role in avoiding motion ambiguity. However, existing VFI methods always struggle to accurately predict the motion information between consecutive frames, and this imprecise estimation leads to blurred and visually incoherent interpolated frames. In this paper, we propose a novel diffusion framework, Motion-Aware latent Diffusion models (MADiff), which is specifically designed for the VFI task. By incorporating motion priors between the conditional neighboring frames with the target interpolated frame predicted throughout the diffusion sampling procedure, MADiff progressively refines the intermediate outcomes, culminating in generating both visually smooth and realistic results. Extensive experiments conducted on benchmark datasets demonstrate that our method achieves state-of-the-art performance significantly outperforming existing approaches, especially under challenging scenarios involving dynamic textures with complex motion. Zhilin Huang, Yijie Yu 0001, Ling Yang 0006, Chujun Qin, Xiawu Zheng, Zikun Zhou, Yaowei Wang 0001, Wenming Yang |
ACM Multimedia | 7 |
| 2024 | Robust Tracking via Fully Exploring Background Prior KnowledgeabstractTypical Siamese-based trackers focus on the target region and pay less attention to the background area. However, the background area can provide the tracker with prior knowledge about the target surroundings. Nonetheless, since the tracker can naturally utilize the target template for localization, importing additional background knowledge requires proper design so that the background area prior knowledge can be fully explored. Furthermore, the introduction of the entire background regions is redundant. Instead, the part background distractors in the regions are more meaningful for the discrimination of the tracker. In this work, we propose a background prior knowledge fully explored tracker for robust tracking. Firstly, we present a Transformer-based explicitly and fully background-utilizing scheme by boosting the tracker to independently exploit the background for localization. Specifically, a target-distractor independent decoder explicitly utilizes the background knowledge by making the target and the distractors independently perform fusion with the search feature. Secondly, we design a simple yet efficient discriminative distractors mining module to refine the background prior knowledge by replacing the whole background region with the mined background distractors. Extensive experiments demonstrate that the proposed method performs favorably against state-of-the-art trackers on nine benchmarks. Zheng'ao Wang, Zikun Zhou, Fanglin Chen 0001, Jun Xu 0008, Wenjie Pei, Guangming Lu 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Cross-Modality Proposal-Guided Feature Mining for Unregistered RGB-Thermal Pedestrian DetectionabstractRGB-Thermal (RGB-T) pedestrian detection aims to locate pedestrians in RGB-T image pairs to exploit the complementation between the two modalities for improving detection robustness in extreme conditions. Most existing algorithms assume that the RGB-T image pairs are well registered, while in the real world, they are not ideally aligned due to parallax or different field-of-view of the cameras. The pedestrians in misaligned image pairs may be located at different positions in two images, which results in two challenges: 1) how to achieve inter-modality complementation using spatially misaligned RGB-T pedestrian patches and 2) how to recognize unpaired pedestrians at the boundary. To address these issues, we propose a new paradigm for unregistered RGB-T pedestrian detection, which predicts two separate pedestrian locations in RGB and thermal images. Specifically, we propose a cross-modality proposal-guided feature mining (CPFM) mechanism to extract two precise fusion features for representing a pedestrian in the two modalities, even if the given RGB-T image pair is unaligned. It enables us to effectively exploit the complementation between the two modalities. With the CPFM mechanism, we build a two-stream dense detector that predicts two pedestrian locations in the two modalities based on the corresponding fusion features mined by the CPFM mechanism. In addition, we design a data augmentation method, named Homography, to simulate the discrepancy in scales and views between images. We also investigate two non-maximum suppression (NMS) methods for post-processing purposes. Favorable experimental results demonstrate the effectiveness and robustness of our method in addressing unregistered pedestrians with different shifts. Zikun Zhou, Yuqing Huang, Gaojun Li, Zhenyu He 0001 |
IEEE Trans. Multim. | 2 |
| 2023 | Joint Visual Grounding and Tracking with Natural Language SpecificationabstractTracking by natural language specification aims to locate the referred target in a sequence based on the natural language description. Existing algorithms solve this issue in two steps, visual grounding and tracking, and accordingly deploy the separated grounding model and tracking model to implement these two steps, respectively. Such a separated framework overlooks the link between visual grounding and tracking, which is that the natural language descriptions provide global semantic cues for localizing the target for both two steps. Besides, the separated framework can hardly be trained end-to-end. To handle these issues, we propose a joint visual grounding and tracking framework, which reformulates grounding and tracking as a unified task: localizing the referred target based on the given visual-language references. Specifically, we propose a multi-source relation modeling module to effectively build the relation between the visual-language references and the test image. In addition, we design a temporal modeling module to provide a temporal clue with the guidance of the global semantic information for our model, which effectively improves the adaptability to the appearance variations of the target. Extensive experimental results on TNL2K, LaSOT, OTB99, and RefCOCOg demonstrate that our method performs favorably against state-of-the-art algorithms for both tracking and grounding. Code is available at https://github.com/lizhou-cs/JointNLT. Li Zhou 0017, Zikun Zhou, Kaige Mao, Zhenyu He 0001 |
CVPR | 2 |
| 2023 | Siamese residual network for efficient visual tracking
Nana Fan, Qiao Liu 0001, Xin Li 0034, Zikun Zhou, Zhenyu He 0001 |
Inf. Sci. | 4 |
| 2022 | Global Tracking via Ensemble of Local TrackersabstractThe crux of long-term tracking lies in the difficulty of tracking the target with discontinuous moving caused by out-of-view or occlusion. Existing long-term tracking methods follow two typical strategies. The first strategy employs a local tracker to perform smooth tracking and uses another re-detector to detect the target when the target is lost. While it can exploit the temporal context like historical appearances and locations of the target, a potential limitation of such strategy is that the local tracker tends to misidentify a nearby distractor as the target instead of activating the re-detector when the real target is out of view. The other long-term tracking strategy tracks the target in the entire image globally instead of local tracking based on the previous tracking results. Unfortunately, such global tracking strategy cannot leverage the temporal context effectively. In this work, we combine the advantages of both strategies: tracking the target in a global view while exploiting the temporal context. Specifically, we perform global tracking via ensemble of local trackers spreading the full image. The smooth moving of the target can be handled steadily by one local tracker. When the local tracker accidentally loses the target due to suddenly discontinuous moving, another local tracker close to the target is then activated and can readily take over the tracking to locate the target. While the activated local tracker performs tracking locally by leveraging the temporal context, the ensemble of local trackers renders our model the global view for tracking. Extensive experiments on six datasets demonstrate that our method performs favorably against state-of-the-art algorithms. Zikun Zhou, Jianqiu Chen, Wenjie Pei, Kaige Mao, Hongpeng Wang 0002, Zhenyu He 0001 |
CVPR | 1 |
| 2022 | Noise-Suppressing Deep TrackingabstractIn visual tracking, it is challenging to distinguish the target from similar objects called noises in the background. As deep trackers use convolutional neural networks for image classification as feature extractors, the extracted features are insensitive to different instances in the same class, which is prone to make prediction models confuse the target and the similar noises in the background. To this end, we propose a noise-suppressing algorithm to learn the discriminative representation for distinguishing the target from the noises in the background. First, we learn polynomial kernels for a search patch under the semantic guidance to increase the difference between representations of the target and the noises in the background. Second, we formulate the online foreground-background functions for the target and the noises in the background to learn an adaptive kernel, which suppresses the features positive for the noises and promotes the features positive for the target. We evaluate the proposed method on seven public datasets including OTB-2013, OTB-2015, VOT-2018, LaSOT, TrackingNet, GOT10k, and NFS. The comprehensive experimental results show that the proposed algorithm performs favorably against state-of-the-art methods, while running at real-time speed. Nana Fan, Xin Li 0034, Zikun Zhou, Qiao Liu 0001, Zhenyu He 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Target-Aware State Estimation for Visual TrackingabstractTrackers based on the IoU prediction network (IoU-Net) have shown superior performance, which refines a coarse bounding box to an accurate one by maximizing the IoU between the target and the coarse box. However, the traditional IoU-Net is less effective in exploiting the limited but crucial supervision information contained in the initial frame, including the discriminative information between the target and backgrounds and the structure information of the initial target. Missing such information makes the IoU-Net less robust to background distractors and diverse variations of the target appearance. To address this issue, we propose a target-aware state estimation network for visual tracking. A gradient-guided feature adjustment module is built on an online discriminative model to generate target-aware features for constructing the state estimation network; it conveys the online learned discriminative information into the offline trained state estimation network. In addition, we propose a structure-aware integration module and embed it into the state estimation network, enabling the tracker to explicitly model the structure information of the initial target. Extensive experimental results on the VOT2018, OTB2015, UAV123, NFS30, TC128, TrackingNet, LaSOT, and VOT2018-LT datasets demonstrate that the proposed approach performs favorably against state-of-the-art trackers. Zikun Zhou, Xin Li 0034, Nana Fan, Hongpeng Wang 0002, Zhenyu He 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Object Tracking via Spatial-Temporal Memory NetworkabstractTemporal and spatial contexts, characterizing target appearance variations and target-background differences, respectively, are crucial for improving the online adaptive ability and instance-level discriminative ability of object tracking. However, most existing trackers focus on either the temporal context or the spatial context during tracking and have not exploited these contexts simultaneously and effectively. In this paper, we propose a Spatial-TEmporal Memory (STEM) network to exploit these contexts jointly for object tracking. Specifically, we develop a key-value structured memory model equipped with a key-value index-based memory reading mechanism to model the spatial and temporal contexts simultaneously. To update the memory with new target states and ensure the diversity of the memory, we introduce a similarity-aware memory update scheme. In addition, we construct an entropy-guided ensemble strategy to fuse the prediction models based on these two contexts, such that these two contexts can be exploited to estimate the target state jointly. Extensive experimental results on eight challenging datasets, including OTB2015, TC128, UAV123, VOT2018, LaSOT, TrackingNet, GOT-10k, and OxUvA, demonstrate that the proposed method performs favorably against state-of-the-art trackers. Zikun Zhou, Xin Li 0034, Tianzhu Zhang 0001, Hongpeng Wang 0002, Zhenyu He 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | SiamCorners: Siamese Corner Networks for Visual TrackingabstractThe current Siamese network based on region proposal network (RPN) has attracted great attention in visual tracking due to its excellent accuracy and high efficiency. However, the design of the RPN involves the selection of the number, scale, and aspect ratios of anchor boxes, which will affect the applicability and convenience of the model. Furthermore, these anchor boxes require complicated calculations, such as calculating their intersection-over-union (IoU) with ground truth bounding boxes. Due to the problems related to anchor boxes, we propose a simple yet effective anchor-free tracker (named Siamese corner networks, SiamCorners), which is end-to-end trained offline on large-scale image pairs. Specifically, we introduce a modified corner pooling layer to convert the bounding box estimate of the target into a pair of corner predictions (the bottom-right and the top-left corners). By tracking a target as a pair of corners, we avoid the need to design the anchor boxes. This will make the entire tracking algorithm more flexible and simple than anchor-based trackers. In our network design, we further introduce a layer-wise feature aggregation strategy that enables the corner pooling module to predict multiple corners for a tracking target in deep networks. We then introduce a new penalty term that is used to select an optimal tracking box in these candidate corners. Finally, SiamCorners achieves experimental results that are comparable to the state-of-art tracker while maintaining a high running speed. In particular, SiamCorners achieves a 53.7% AUC on NFS30 and a 61.4% AUC on UAV123, while still running at 42 frames per second (FPS). Kai Yang 0018, Zhenyu He 0001, Wenjie Pei, Zikun Zhou, Xin Li 0034, Di Yuan 0002, Haijun Zhang 0002 |
IEEE Trans. Multim. | 4 |
| 2021 | Saliency-Associated Object TrackingabstractMost existing trackers based on deep learning perform tracking in a holistic strategy, which aims to learn deep representations of the whole target for localizing the target. It is arduous for such methods to track targets with various appearance variations. To address this limitation, another type of methods adopts a part-based tracking strategy which divides the target into equal patches and tracks all these patches in parallel. The target state is inferred by summarizing the tracking results of these patches. A potential limitation of such trackers is that not all patches are equally informative for tracking. Some patches that are not discriminative may have adverse effects. In this paper, we propose to track the salient local parts of the target that are discriminative for tracking. In particular, we propose a fine-grained saliency mining module to capture the local saliencies. Further, we design a saliency-association modeling module to associate the captured saliencies together to learn effective correlation representations between the exemplar and the search image for state estimation. Extensive experiments on five diverse datasets demonstrate that the proposed method performs favorably against state-of-the-art trackers. Zikun Zhou, Wenjie Pei, Xin Li 0034, Hongpeng Wang 0002, Feng Zheng 0001, Zhenyu He 0001 |
ICCV | 1 |
| 2021 | Interactive convolutional learning for visual tracking
Nana Fan, Qiao Liu 0001, Xin Li 0034, Zikun Zhou, Zhenyu He 0001 |
Knowl. Based Syst. | 4 |
| 2021 | Learning dual-margin model for visual tracking
Nana Fan, Xin Li 0034, Zikun Zhou, Qiao Liu 0001, Zhenyu He 0001 |
Neural Networks | 3 |
| 2021 | Adaptive ensemble perception tracking
Zikun Zhou, Nana Fan, Kai Yang 0018, Hongpeng Wang 0002, Zhenyu He 0001 |
Neural Networks | 1 |
| 2020 | LSOTB-TIR: A Large-Scale High-Diversity Thermal Infrared Object Tracking BenchmarkabstractIn this paper, we present a Large-Scale and high-diversity general Thermal InfraRed (TIR) Object Tracking Benchmark, called LSOTB-TIR, which consists of an evaluation dataset and a training dataset with a total of 1,400 TIR sequences and more than 600K frames. We annotate the bounding box of objects in every frame of all sequences and generate over 730K bounding boxes in total. To the best of our knowledge, LSOTB-TIR is the largest and most diverse TIR object tracking benchmark to date. To evaluate a tracker on different attributes, we define 4 scenario attributes and 12 challenge attributes in the evaluation dataset. By releasing LSOTB-TIR, we encourage the community to develop deep learning based TIR trackers and evaluate them fairly and comprehensively. We evaluate and analyze more than 30 trackers on LSOTB-TIR to provide a series of baselines, and the results show that deep trackers achieve promising performance. Furthermore, we re-train several representative deep trackers on LSOTB-TIR, and their results demonstrate that the proposed training dataset significantly improves the performance of deep TIR trackers. Codes and dataset are available at https://github.com/QiaoLiuHit/LSOTB-TIR. Qiao Liu 0001, Xin Li 0034, Zhenyu He 0001, Chenglong Li 0002, Zikun Zhou, Di Yuan 0002, Jing Li 0071, Kai Yang 0018, Nana Fan, Feng Zheng 0001 |
ACM Multimedia | 6 |
| 2020 | SiamAtt: Siamese attention network for visual tracking
Kai Yang 0018, Zhenyu He 0001, Zikun Zhou, Nana Fan |
Knowl. Based Syst. | 3 |
| 2020 | Dual-regression model for visual tracking
Xin Li 0034, Qiao Liu 0001, Nana Fan, Zikun Zhou, Zhenyu He 0001, Xiaoyuan Jing |
Neural Networks | 4 |