Jiahao Nie 0001

dblp:319/4607 · DBLP profile ↗
← Back
25ranked-venue papers
8as first author
25since 2021 · last 2026
0000-0002-1474-1817ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 4 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 5 first-author · 10 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 CompTrack: Information Bottleneck-Guided Low-Rank Dynamic Token Compression for Point Cloud Tracking
abstract
3D single object tracking (SOT) in LiDAR point clouds is a critical task in computer vision and autonomous driving. Despite great success having been achieved, the inherent sparsity of point clouds introduces a dual-redundancy challenge that limits existing trackers: (1) vast spatial redundancy from background noise impairs accuracy, and (2) informational redundancy within the foreground hinders efficiency. To tackle these issues, we propose CompTrack, a novel end-to-end framework that systematically eliminates both forms of redundancy in point clouds. First, CompTrack incorporates a Spatial Foreground Predictor (SFP) module to filter out irrelevant background noise based on information entropy, addressing spatial redundancy. Subsequently, its core is an Information Bottleneck-guided Dynamic Token Compression (IB-DTC) module that eliminates the informational redundancy within the foreground. Theoretically grounded in low-rank approximation, this module leverages an online SVD analysis to adaptively compress the redundant foreground into a compact and highly informative set of proxy tokens. Extensive experiments on KITTI, nuScenes and Waymo datasets demonstrate that CompTrack achieves top-performing tracking performance with superior efficiency, running at a real-time 90 FPS on a single RTX 3090 GPU.
Sifan Zhou, Yichao Cao, Jiahao Nie 0001, Yuqian Fu, Xiaobo Lu, Shuo Wang 0030
AAAI3
2026 LGSA: Label Geometry Structuring and Aligning for Hierarchical Text Classification
abstract
Existing hierarchical text classification (HTC) methods typically use prompt tuning or contrastive learning to inject the label hierarchy into a model as prior knowledge to implicitly learn label embeddings for classification.However, such implicit learning fails to accurately reflect label geometry (i.e., feature spatial distribution of label embeddings), as it does not model hierarchy-aware geometric relations among labels.To address this issue, we propose a novel two-stage label geometry structuring and aligning framework, termed LGSA, which transforms the label hierarchy from an implicit prior into an explicit embedding.First, we propose a hierarchical geometric structuring (HGS) module that leverages a general orthogonal frame (GOF) to reconstruct an explicit label geometry conforming to the label hierarchy.The label geometry is then treated as a label prototype to guide model training.To facilitate the guidance, we thereby propose a hierarchical geometric aligning (HGA) module as a regularization term to align label geometry learned by the model with the explicit label prototype.Experiments on three realworld HTC datasets confirm that LGSA consistently outperforms existing state-of-the-art methods.
Shuai Zhang 0002, Weibo Xu, Jiahao Nie 0001, Kecheng Huang
ACL (1)3
2026 Dual-aware collaboration and localized semantic calibration for label-scarce vertical federated learning
Wenyu Zhang 0001, Jiahao Nie 0001, Zixuan Dai, Qingjun Mao
Inf. Sci.3
2026 G2CL: Gradient-guided graph contrastive learning for eliminating the message contrastive conflict
Shuai Zhang 0002, Shan Yang 0002, Wenyu Zhang 0001, Jiahao Nie 0001, Shan Ji
Neural Networks4
2025 Mamba-Adaptor: State Space Model Adaptor for Visual Recognition
abstract
Recent State Space Models (SSM), especially Mamba, have demonstrated impressive performance in visual modeling and possess superior model efficiency. However, the application of Mamba to visual tasks suffers inferior performance due to three main constraints existing in the sequential model: 1) Casual computing is incapable of accessing global context; 2) Long-range forgetting when computing the current hidden states; 3) Weak spatial structural modeling due to the transformed sequential input. To address these issues, we investigate a simple yet powerful vision task adapter for Mamba models, which consists of two functional modules: Adaptor-T and Adapator-S. When solving the hidden states for SSM, we apply a lightweight prediction module Adaptor-T to select a set of learnable locations as memory augmentations to ease long-range forgetting issues. Moreover, we leverage Adapator-S, composed of multi-scale dilated convolutional kernels, to enhance the spatial modeling and introduce the image inductive bias into the feature output. Both modules can enlarge the context modeling in casual computing, as the output is enhanced by the inaccessible features. We explore three usages of Mamba-Adaptor: A general visual backbone for various vision tasks; A booster module to raise the performance of pretrained backbones; A highly efficient fine-tuning module that adapts the base model for transfer learning tasks. Extensive experiments verify the effectiveness of Mamba-Adapter in three settings. Notably, our Mamba-Adapter achieves state-of-the-art performance on the ImageNet and COCO benchmarks.
Jiahao Nie 0001, Yujin Tang, Hongshen Zhao
CVPR2
2025 FocusTrack: One-Stage Focus-and-Suppress Framework for 3D Point Cloud Object Tracking
Sifan Zhou, Jiahao Nie 0001, Yichao Cao, Xiaobo Lu
ACM Multimedia2
2025 A multiple aging factor interactive learning framework for lithium-ion battery state-of-health estimation
Zhengyi Bao, Tingting Luo, Mingyu Gao 0002, Zhiwei He 0001, Yuxiang Yang 0001, Jiahao Nie 0001
Eng. Appl. Artif. Intell.6
2025 CDRM: Controllable diffusion restoration model for realistic image deblurring
Guangmang Cui, Jufeng Zhao, Jiahao Nie 0001
Expert Syst. Appl.4
2025 P2P: Part-to-Part Motion Cues Guide a Strong Tracking Framework for LiDAR Point Clouds
Jiahao Nie 0001, Sifan Zhou, Xueyi Zhou, Dong-Kyu Chae, Zhiwei He 0001
Int. J. Comput. Vis.1
2025 Specific Task-Guided Collaborative Domain Generalization Network for Intelligent Fault Diagnosis Under Unseen Conditions
abstract
Domain generalization-based methods perform cross-domain fault diagnosis by learning fault-discriminative and domain-invariant diagnostic knowledge between available working conditions (source domains) and unseen conditions (target domains). However, existing approaches ignore domain specificity within discriminative knowledge, preventing optimal fault discrimination across diverse domains. Moreover, since the target domain is inaccessible during model training, obtaining sufficient domain-invariant knowledge from the limited source domains poses a significant challenge. For the weaknesses, a specific task-guided collaborative domain generalization network (STCDGN) is proposed to enhance bearing fault diagnosis under unseen working conditions. Specifically, we construct a channel attention-guided multi-scale feature extractor and task-specific classifiers to establish adaptive decision boundaries for domain specificity. The boundaries interact with the extracted features to enhance fault-discriminative representations. To further mine domain invariance within these representations, we propose a two-stage training strategy through decision boundary divergence maximization and multi-scale hierarchical feature discrepancy minimization, effectively alleviating intra-class domain shift for improved diagnostic generalization. Finally, we propose an entropy-guided decision selection strategy for reliable inference diagnostic results. The average accuracies on the two public datasets reach 98.72% and 89.13%, respectively, indicating significant diagnostic generalization to the unseen target samples. The visualized experimental results further demonstrate the method’s effectiveness in learning fault-discriminative and domain-invariant diagnostic knowledge.
Xiaorong Zheng, Jiahao Nie 0001, Zhiwei He 0001, Mingyu Gao 0002
IEEE Internet Things J.2
2025 Context Matching-Guided Motion Modeling for 3D Point Cloud Object Tracking
abstract
LiDAR-based single object tracking plays a key role in intelligent vehicles. Current methods typically follow appearance matching or motion-centric frameworks. However, point clouds are usually sparse and incomplete, providing insufficient appearance information for matching. While the motion-centric framework predicts inter-frame motion of targets instead of performing appearance matching for tracking, it neglects contextual information matching of consecutive frames that is conducive to target motion modeling. In this paper, we propose an elegant and effective framework by leveraging Context Matching to guide motion modeling for accurate Tracking (CMTrack). The novel framework possesses two attractive properties: 1) It incorporates a context matching encoder-decoder network to match contextual information of consecutive frames, fully exploring informative cues relevant to target motion. 2) Benefiting from informative motion cues being modeling, CMTrack allows for accurate prediction of inter-frame motion of targets in a one-stage manner. Extensive experiments are conducted on several widely-adopted datasets, i.e., KITTI, NuScenes and Waymo Open Dataset. Without bells and whistles, our CMTrack demonstrates competitive tracking accuracy (e.g., 87.3% and 69.3% precision on KITTI and NuScenes, respectively) compared to state-of-the-art methods, while running at a high speed of 48 Fps on a single Titan Xp GPU.
Jiahao Nie 0001, Zhengyi Bao, Zhiwei He 0001, Mingyu Gao 0002
IEEE Trans. Circuits Syst. Video Technol.1
2025 Exploring Informative and Highly-Transferable Features for Cross-Machine Fault Diagnosis by ConvFormer-Based Biconditional Domain Adaptation Method
abstract
Domain adaptation-based methods have been proved success for cross-machine fault diagnosis. However, such methods suffer from limited diagnosis performance because the utilized networks typically rely on convolution layers with local receptive fields, failing to extract informative fault features, and the information on machine domain and fault category is not fully utilized, which prevents the transferability of fault features across machines. Towards these issues, a novel ConvFormer-based biconditional domain adaptation method (CFBDAM) is proposed to explore informative and highly-transferable fault features for accurate diagnosis. The proposed ConvFormer network first extracts global-local fault features in a parallel manner via a linear transformer and a separable shuffled CNN, respectively. The resulting features are then fed into a cross-attention feature fusion module to form informative diagnostic knowledge. Our ConvFormer is deployment-friendly owing to lightweight designs, such as linear and separation operations. To enhance cross-machine transferability of the informative fault features extracted by ConvFormer, a biconditional domain adaptation strategy is designed. It imposes biconditional constraints by using the information of both machine domain and fault category, thereby leading to highly-transferable fault features with domain insensitivity and category discriminability. Comprehensive experiments are conducted on six transfer diagnosis tasks across three machines. The experimental results show that CFBDAM achieves potential cross-machine diagnostic performance.
Xiaorong Zheng, Jiahao Nie 0001, Zhiwei He 0001, Mingyu Gao 0002
IEEE Trans. Ind. Informatics2
2024 Towards Category Unification of 3D Single Object Tracking on Point Clouds
abstract
Category-specific models are provenly valuable methods in 3D single object tracking (SOT) regardless of Siamese or motion-centric paradigms. However, such over-specialized model designs incur redundant parameters, thus limiting the broader applicability of 3D SOT task. This paper first introduces unified models that can simultaneously track objects across all categories using a single network with shared model parameters. Specifically, we propose to explicitly encode distinct attributes associated to different object categories, enabling the model to adapt to cross-category data. We find that the attribute variances of point cloud objects primarily occur from the varying size and shape (e.g., large and square vehicles v.s. small and slender humans). Based on this observation, we design a novel point set representation learning network inheriting transformer architecture, termed AdaFormer, which adaptively encodes the dynamically varying shape and size information from cross-category data in a unified manner. We further incorporate the size and shape prior derived from the known template targets into the model’s inputs and learning objective, facilitating the learning of unified representation. Equipped with such designs, we construct two category-unified models SiamCUT and MoCUT. Extensive experiments demonstrate that SiamCUT and MoCUT exhibit strong generalization and training stability. Furthermore, our category-unified models outperform the category-specific counterparts by a significant margin (e.g., on KITTI dataset, $\sim$12\% and $\sim$3\% performance gains on the Siamese and motion paradigms).
Jiahao Nie 0001, Zhiwei He 0001, Xueyi Zhou, Dong-Kyu Chae
ICLR1
2024 VoxelTrack: Exploring Multi-level Voxel Representation for 3D Point Cloud Object Tracking
abstract
Current LiDAR point cloud-based 3D single object tracking (SOT) methods typically rely on point-based representation network. Despite demonstrated success, such networks suffer from some fundamental problems: 1) It contains pooling operation to cope with inherently disordered point clouds, hindering the capture of 3D spatial information that is useful for tracking, a regression task. 2) The adopted set abstraction operation hardly handles density-inconsistent point clouds, also preventing 3D spatial information from being modeled. To solve these problems, we introduce a novel tracking framework, termed VoxelTrack. By voxelizing inherently disordered point clouds into 3D voxels and extracting their features via sparse convolution blocks, VoxelTrack effectively models precise and robust 3D spatial information, thereby guiding accurate position prediction for tracked objects. Moreover, VoxelTrack incorporates a dual-stream encoder with cross-iterative feature fusion module to further explore fine-grained 3D spatial information for tracking. Benefiting from accurate 3D spatial information being modeled, our VoxelTrack simplifies tracking pipeline with a single regression loss. Extensive experiments are conducted on three widely-adopted datasets including KITTI, NuScenes and Waymo Open Dataset. The experimental results confirm that VoxelTrack achieves state-of-the-art performance (88.3%, 71.4% and 63.6% mean precision on the three datasets, respectively), and outperforms the existing trackers with a real-time speed of 36 Fps on a single TITAN RTX GPU. The source code and model will be released.
Yuxuan Lu 0007, Jiahao Nie 0001, Zhiwei He 0001, Hongjie Gu
ACM Multimedia2
2024 SAR-SLAM: Self-Attentive Rendering-based SLAM with Neural Point Cloud Encoding
abstract
Neural implicit representations have recently revolutionized simultaneous localization and mapping (SLAM), giving rise to a groundbreaking paradigm known as NeRF-based SLAM. However, existing methods often fall short in accurately estimating poses and reconstructing scenes. This limitation largely stems from their reliance on volume rendering techniques, which oversimplify the modeling process. In this paper, we introduce a novel neural implicit SLAM system named SAR-SLAM to address these shortcomings. Our approach reconstructs Neural Radiance Fields (NeRFs) using a self-attentive architecture and represents scenes through neural point cloud encoding. Unlike previous NeRF-based SLAM methods, which depend on traditional volume rendering equations for scene representation and view synthesis, our method employs a self-attentive rendering framework with the Transformer architecture during mapping and tracking stages. To enable incremental mapping, we anchor scene features within a neural point cloud, striking a balance between estimation accuracy and computational cost. Experimental results on three challenging datasets show the superior performance and robustness of our SAR-SLAM compared to recent NeRF-based SLAM systems. The code will be released.
Zhiwei He 0001, Yuxiang Yang 0001, Jiahao Nie 0001, Jing Zhang 0037
ACM Multimedia4
2024 MPFormer: Multipatch Transformer for Multivariate Time-Series Anomaly Detection With Contrastive Learning
abstract
Recent unsupervised framework-based anomaly detection methods in Internet of Things (IoT) have been proved successful by learning temporal pattern representations from normal time-series data. However, such methods have long suffered from weakened normal-abnormal boundary incurred by discrimination-insufficient temporal pattern representations, due to: 1) normal time series often suffer from some unprocessed abnormal noise, interfering with model’s discriminative representations between normal and abnormal points and 2) previous methods rarely focus on contextual semantic features of time series that are also crucial for constructing discriminative representations. In this article, we propose multipatch contrastive learning framework for multivariate time-series anomaly detection (MPFormer), a transformer-based multipatch contrastive learning framework. This novel framework leverages contrastive learning to guide discriminative feature learning. It incorporates a data augmentation strategy to generate positive samples and maximizes the similarity between positive sample pairs, thereby capturing inherent representations of normal temporal patterns and enhancing robust discrimination ability to abnormal noise. To further facilitate discriminative representation modeling, we divide input time series into multiple patches and design a transformer-based dual-attention module to explore contextual semantic features in both interpatch and intrapatch views. Benefiting from discriminative representations being modeled, MPFormer effectively strengthens the normal-abnormal boundary, thus improving detection accuracy. Experimental results on six widely adopted data sets demonstrate that our proposed MPFormer outperforms existing baseline methods.
Shenhui Ma, Jiahao Nie 0001, Siwei Guan, Zhiwei He 0001, Mingyu Gao 0002
IEEE Internet Things J.2
2024 MSF-SLAM: Multi-Sensor-Fusion-Based Simultaneous Localization and Mapping for Complex Dynamic Environments
abstract
We proposed a multi-sensor fusion-based localization and scene reconstruction method for a complex dynamic scene. The multi-level fusion between multiple sensors was implemented by fusing data collected from different sensors in different system modules. In the front-end of the system, the camera and the LiDAR assisted each other. The LiDAR point clouds provided 3D information for the feature points in the image. The moving objects elimination method based on the image can remove the points on the moving objects in the LiDAR point clouds for localization accuracy improvement and static 3D scene reconstruction. To further improve the localization accuracy, a combination of visual loop closure detection and LiDAR loop closure detection was utilized to ensure the global consistency of scene reconstruction. At the system’s back-end, the observation model of different sensors was integrated to construct a multiple constraint factor graph with nonlinear optimization to obtain the optimal system states. Experimental results demonstrated that the proposed multi-sensor fusion-based localization and scene reconstruction algorithm could operate robustly in multiple complex dynamic scenes.
Zhiwei He 0001, Yuxiang Yang 0001, Jiahao Nie 0001, Zhekang Dong, Shuo Wang 0030, Mingyu Gao 0002
IEEE Trans. Intell. Transp. Syst.4
2023 GLT-T: Global-Local Transformer Voting for 3D Single Object Tracking in Point Clouds
abstract
Current 3D single object tracking methods are typically based on VoteNet, a 3D region proposal network. Despite the success, using a single seed point feature as the cue for offset learning in VoteNet prevents high-quality 3D proposals from being generated. Moreover, seed points with different importance are treated equally in the voting process, aggravating this defect. To address these issues, we propose a novel global-local transformer voting scheme to provide more informative cues and guide the model pay more attention on potential seed points, promoting the generation of high-quality 3D proposals. Technically, a global-local transformer (GLT) module is employed to integrate object- and patch-aware prior into seed point features to effectively form strong feature representation for geometric positions of the seed points, thus providing more robust and accurate cues for offset learning. Subsequently, a simple yet effective training strategy is designed to train the GLT module. We develop an importance prediction branch to learn the potential importance of the seed points and treat the output weights vector as a training constraint term. By incorporating the above components together, we exhibit a superior tracking method GLT-T. Extensive experiments on challenging KITTI and NuScenes benchmarks demonstrate that GLT-T achieves state-of-the-art performance in the 3D single object tracking task. Besides, further ablation studies show the advantages of the proposed global-local transformer voting scheme over the original VoteNet. Code and models will be available at https://github.com/haooozi/GLT-T.
Jiahao Nie 0001, Zhiwei He 0001, Yuxiang Yang 0001, Mingyu Gao 0002, Jing Zhang 0037
AAAI1
2023 OSP2B: One-Stage Point-to-Box Network for 3D Siamese Tracking
abstract
Two-stage point-to-box network acts as a critical role in the recent popular 3D Siamese tracking paradigm, which first generates proposals and then predicts corresponding proposal-wise scores. However, such a network suffers from tedious hyper-parameter tuning and task misalignment, limiting the tracking performance. Towards these concerns, we propose a simple yet effective one-stage point-to-box network for point cloud-based 3D single object tracking. It synchronizes 3D proposal generation and center-ness score prediction by a parallel predictor without tedious hyper-parameters. To guide a task-aligned score ranking of proposals, a center-aware focal loss is proposed to supervise the training of the center-ness branch, which enhances the network's discriminative ability to distinguish proposals of different quality. Besides, we design a binary target classifier to identify target-relevant points. By integrating the derived classification scores with the center-ness scores, the resulting network can effectively suppress interference proposals and further mitigate task misalignment. Finally, we present a novel one-stage Siamese tracker OSP2B equipped with the designed network. Extensive experiments on challenging benchmarks including KITTI and Waymo SOT Dataset show that our OSP2B achieves leading performance with a considerable real-time speed.
Jiahao Nie 0001, Zhiwei He 0001, Yuxiang Yang 0001, Zhengyi Bao, Mingyu Gao 0002, Jing Zhang 0037
IJCAI1
2023 FAML-RT: Feature alignment-based multi-level similarity metric learning network for a two-stage robust tracker
Jiahao Nie 0001, Zhekang Dong, Zhiwei He 0001, Mingyu Gao 0002
Inf. Sci.1
2023 Learning task-specific discriminative representations for multiple object tracking
Jiahao Nie 0001, Zhiwei He 0001, Mingyu Gao 0002
Neural Comput. Appl.2
2023 Leveraging temporal-aware fine-grained features for robust multiple object tracking
Jiahao Nie 0001, Zhiwei He 0001, Mingyu Gao 0002
J. Supercomput.2
2023 Learning Localization-Aware Target Confidence for Siamese Visual Tracking
abstract
Siamese tracking paradigm has achieved great success, providing effective appearance discrimination and size estimation by classification and regression. While such a paradigm typically optimizes the classification and regression independently, leading to task misalignment (accurate prediction boxes have no high target confidence scores). In this paper, to alleviate this misalignment, we propose a novel tracking paradigm, called SiamLA. Within this paradigm, a series of simple, yet effective localization-aware components are introduced to generate localization-aware target confidence scores. Specifically, with the proposedlocalization-aware dynamic label(LADL) loss andlocalization-aware label smoothing(LALS) strategy, collaborative optimization between the classification and regression is achieved, enabling classification scores to be aware of location state, not just appearance similarity. Besides, we propose a separatelocalization-aware quality prediction(LAQP) branch to produce location quality scores to further modify the classification scores. To guide a more reliable modification, a novellocalization-aware feature aggregation(LAFA) module is designed and embedded into this branch. Consequently, the resulting target confidence scores are more discriminative for the location state, allowing accurate prediction boxes tend to be predicted as high scores. Extensive experiments are conducted on six challenging benchmarks, including GOT-10 k, TrackingNet, LaSOT, TNL2K, OTB100 and VOT2018. Our SiamLA achieves competitive performance in terms of both accuracy and efficiency. Furthermore, a stability analysis reveals that our tracking paradigm is relatively stable, implying that the paradigm is potential for real-world applications.
Jiahao Nie 0001, Zhiwei He 0001, Yuxiang Yang 0001, Mingyu Gao 0002, Zhekang Dong
IEEE Trans. Multim.1
2022 Hierarchical Feature Fusion based Reconstruction Network for Unsupervised Anomaly Detection
abstract
With the wide application of deep learning in the hydropower industry, many anomaly detection methods based on deep neural networks have been proposed to improve detection accuracy in electromechanical systems. However, they typically utilize recurrent neural networks to spontaneously learn the properties of multidimensional time-series data, which rarely using the feature dimension and time dimension. To solve this issue, in this paper, we propose a hierarchical feature fusion reconstruction network (HFFRN) to detect anomaly using feature dimension and time dimension information. Specifically, we first construct a feature extraction layer to strengthen the feature interaction among different layers, in which the shallow feature has rich detail information, and the deep feature has rich semantic information. Then, the hierarchical feature fusion layer is designed to fuse the information of the shallow and deep layers, allowing the model to perceive more information of the feature dimension and time dimension. We conduct a series of experiments on the SWaT dataset. The experimental results show that HFFRN outperforms the baseline method. Notably, HFFRN improves the F1 performance by 13% compared to the baseline. In addition, to prove the generalization of the HFFRN model, we also test the model on Mammography time series dataset.
Binjie Zhao, Jiahao Nie 0001, Siwei Guan, Zhiwei He 0001, Mingyu Gao 0002
ETFA2
2022 Spreading Fine-Grained Prior Knowledge for Accurate Tracking
abstract
With the widespread use of deep learning in single object tracking task, mainstream tracking algorithms treat tracking as a combined classification and regression problem. Classification aims at locating an arbitrary target, and regression aims at estimating the corresponding bounding box. In this paper, we focus on regression and propose a novel box estimation network, which consists of a transformer encoder target pyramid guide (TPG) and transformer decoder target pyramid spread (TPS). Specifically, the transformer encoder TPG is designed to generate fine-grained prior knowledge with explicit representation for template targets. In contrast to the raw transformer encoder, we capture the visual dependence through local-global self-attention and deem the multi-scale target regions as the “local” region. Using this fine-grained prior knowledge, we design the transformer decoder TPS to spread it to the subsequent search regions with high affinity to accurately estimate the bounding boxes. Considering that self-attention fails to model information interaction across channels between the template target and search regions, we develop a channel-wise cross-attention block within the TPS as compensation. Extensive experiments on the OTB100, UAV123, NFS, VOT2020, VOT2021, LaSOT, LaSOT_ext, TrackingNet and GOT-10k benchmarks show that the proposed box estimation network outperforms most existing box estimation methods. Furthermore, our trackers based on this estimation network exhibit a competitive performance against state-of-the-art trackers.
Jiahao Nie 0001, Zhiwei He 0001, Mingyu Gao 0002, Zhekang Dong
IEEE Trans. Circuits Syst. Video Technol.1