Zhiwei He 0001

dblp:52/6077-1 · DBLP profile ↗
← Back
32ranked-venue papers
0as first author
27since 2021 · last 2026
0000-0001-7264-2019ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 9 since 2021Artificial intelligence and machine learning · 8 · 8 since 2021Systems, architecture and hardware · 7 · 3 since 2021Computer networks · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CMIF-Calib: A Novel LiDAR-Camera Extrinsic Calibration System With Cross-Modal Information Fusion
abstract
The complexity of unmanned system operating environments necessitates multi-sensor fusion for robust, long-term perception. Precise extrinsic calibration across heterogeneous sensors is fundamental to effective fusion. While data-driven calibration methods with artificial intelligence offer high accuracy and efficiency, their generalization to unseen scenarios and new sensor configurations remains limited. To address this, we propose CMIF-Calib, a novel LiDAR-Camera extrinsic calibration framework. Our approach employs an encoder-decoder network with cascaded multi-attention modules to learn reliable cross-modal correspondences between LiDAR point clouds and camera images, enabling indirect estimation of extrinsic parameters. By decoupling sensor intrinsics from the calibration network, CMIF-Calib supports simultaneous calibration of diverse sensor types (varying camera intrinsics/LiDAR channels) and multiple sensor combinations (including multi-camera and multi-LiDAR setups). Extensive experiments across diverse datasets and platforms demonstrate that CMIF-Calib achieves higher calibration accuracy and superior generalization compared to existing methods. The implementation of CMIF-Calib will be made publicly available upon publication at https://github.com/LvXudong-HIT/CMIF-Calib.
Shuo Wang 0030, Zhekang Dong, Mingyu Gao 0002, Zhiwei He 0001
IEEE Internet Things J.5
2025 UAWTrack: Universal 3D Single Object Tracking in Adverse Weather
abstract
3D single object tracking (3D SOT) in LiDAR point clouds is essential for autonomous driving. Most existing 3D SOT methods focus on clear weather, where point clouds are more defined. However, adverse weather conditions lead to sparser and noisier point clouds, significantly degrading tracking performance and posing safety risks. In this study, we introduce UAWTrack, a universal 3D SOT model designed to perform effectively across diverse real-world weather conditions. UAWTrack comprises three key modules: 1) Voxel Feature Extraction, which mitigates the perturbations in point clouds caused by adverse weather; 2) Motion-centric Spatial-temporal Aggregation and Motion-guided Feature Fusion, capturing motion clues and sampling dense BEV motion features to address the issue of sparsity; and 3) Weather-Specific Tracker, which efficiently handles tracking in various weather conditions. To fill the gap of lacking benchmarks for 3D SOT in adverse weather, we simulate physically valid adverse weather conditions on the KITTI and NuScenes datasets, creating two benchmarks: KITTI-A and NuScenes-A. Extensive experiments demonstrate that UAWTrack achieves state-of-the-art performance under all weather conditions.
Yuxiang Yang 0001, Hongjie Gu, Yingqi Deng, Zhekang Dong, Zhiwei He 0001, Jing Zhang 0037
AAAI5
2025 A multiple aging factor interactive learning framework for lithium-ion battery state-of-health estimation
Zhengyi Bao, Tingting Luo, Mingyu Gao 0002, Zhiwei He 0001, Yuxiang Yang 0001, Jiahao Nie 0001
Eng. Appl. Artif. Intell.4
2025 P2P: Part-to-Part Motion Cues Guide a Strong Tracking Framework for LiDAR Point Clouds
Jiahao Nie 0001, Sifan Zhou, Xueyi Zhou, Dong-Kyu Chae, Zhiwei He 0001
Int. J. Comput. Vis.6
2025 MTSBV Algorithm Design for Real-Time Battery Data Segmentation
abstract
With the increasing number of new energy vehicles, the safety of power batteries has gained significant attention. A prerequisite for accurate estimation of battery state is effective segmentation of battery data. In response to this issue, this paper develops a frequency-domain feature-based real-time segmentation methodology for battery data analysis, specifically proposing a Multi-Threshold Spectral Band Variance (MTSBV) approach. To verify the effectiveness and robustness of the approach in real-world driving applications, we used two cycles of data, the Federal Urban Driving Program (FUDS) and US06, from publicly available battery datasets at the University of Maryland. In addition to analyzing the performance of the algorithm under ideal noise-free conditions, we also evaluated its performance when the signal is corrupted by additive white Gaussian noise (AWGN), salt pepper noise (SPN), and colored noise (CN). The performance of the algorithm was analyzed from two perspectives: the algorithm’s performance under different Signal-to-Noise Ratios (SNRs) and its comparison with other commonly used segmentation algorithms. Experimental results show that MTSBV algorithm has a good effect on solving the problem of battery data segmentation under low SNR conditions. Under FUDS and UD06 conditions, when SNR=15dB, Root Mean Square Error (RMSE) is reduced by 32.1% and 21.8% compared with other optimal conditions, respectively.
Ping Li 0032, Yuxiang Yang 0001, Zhiwei He 0001, Mingyu Gao 0002
IEEE Internet Things J.4
2025 Specific Task-Guided Collaborative Domain Generalization Network for Intelligent Fault Diagnosis Under Unseen Conditions
abstract
Domain generalization-based methods perform cross-domain fault diagnosis by learning fault-discriminative and domain-invariant diagnostic knowledge between available working conditions (source domains) and unseen conditions (target domains). However, existing approaches ignore domain specificity within discriminative knowledge, preventing optimal fault discrimination across diverse domains. Moreover, since the target domain is inaccessible during model training, obtaining sufficient domain-invariant knowledge from the limited source domains poses a significant challenge. For the weaknesses, a specific task-guided collaborative domain generalization network (STCDGN) is proposed to enhance bearing fault diagnosis under unseen working conditions. Specifically, we construct a channel attention-guided multi-scale feature extractor and task-specific classifiers to establish adaptive decision boundaries for domain specificity. The boundaries interact with the extracted features to enhance fault-discriminative representations. To further mine domain invariance within these representations, we propose a two-stage training strategy through decision boundary divergence maximization and multi-scale hierarchical feature discrepancy minimization, effectively alleviating intra-class domain shift for improved diagnostic generalization. Finally, we propose an entropy-guided decision selection strategy for reliable inference diagnostic results. The average accuracies on the two public datasets reach 98.72% and 89.13%, respectively, indicating significant diagnostic generalization to the unseen target samples. The visualized experimental results further demonstrate the method’s effectiveness in learning fault-discriminative and domain-invariant diagnostic knowledge.
Xiaorong Zheng, Jiahao Nie 0001, Zhiwei He 0001, Mingyu Gao 0002
IEEE Internet Things J.3
2025 Unified Volumetric Avatar: Enabling flexible editing and rendering of neural human representations
abstract
Neural Radiance Field (NeRF) has emerged as a leading method for reconstructing 3D human avatars with exceptional rendering capabilities, particularly for novel view and pose synthesis. However, current approaches for editing these avatars are limited, typically allowing only global geometry adjustments or texture modifications via neural texture maps. This paper introduces Unified Volumetric Avatar, a novel framework enabling independent and simultaneous global and local editing of both geometry and texture of 3D human avatars and user-friendly manipulation. The proposed approach seamlessly integrates implicit neural fields with an explicit polygonal mesh , leveraging distinct geometry and appearance latent codes attached to the body mesh for precise local edits. These trackable latent codes permeate through the 3D space via barycentric interpolation, mitigating spatial ambiguity with the aid of a local signed height indicator. Furthermore, our method enhances surface illumination representation across different poses by incorporating a pose-dependent shading factor instead of relying on view-dependent radiance color. Experimental results on multiple human avatars demonstrate its efficacy in achieving competitive results for novel view synthesis and novel pose rendering, showcasing its potential for versatile human representation. The source code will be made publicly available.
Jinlong Fan 0001, Xuepu Zeng, Zhengyi Bao, Zhiwei He 0001, Mingyu Gao 0002
Image Vis. Comput.5
2025 Context Matching-Guided Motion Modeling for 3D Point Cloud Object Tracking
abstract
LiDAR-based single object tracking plays a key role in intelligent vehicles. Current methods typically follow appearance matching or motion-centric frameworks. However, point clouds are usually sparse and incomplete, providing insufficient appearance information for matching. While the motion-centric framework predicts inter-frame motion of targets instead of performing appearance matching for tracking, it neglects contextual information matching of consecutive frames that is conducive to target motion modeling. In this paper, we propose an elegant and effective framework by leveraging Context Matching to guide motion modeling for accurate Tracking (CMTrack). The novel framework possesses two attractive properties: 1) It incorporates a context matching encoder-decoder network to match contextual information of consecutive frames, fully exploring informative cues relevant to target motion. 2) Benefiting from informative motion cues being modeling, CMTrack allows for accurate prediction of inter-frame motion of targets in a one-stage manner. Extensive experiments are conducted on several widely-adopted datasets, i.e., KITTI, NuScenes and Waymo Open Dataset. Without bells and whistles, our CMTrack demonstrates competitive tracking accuracy (e.g., 87.3% and 69.3% precision on KITTI and NuScenes, respectively) compared to state-of-the-art methods, while running at a high speed of 48 Fps on a single Titan Xp GPU.
Jiahao Nie 0001, Zhengyi Bao, Zhiwei He 0001, Mingyu Gao 0002
IEEE Trans. Circuits Syst. Video Technol.4
2025 Exploring Informative and Highly-Transferable Features for Cross-Machine Fault Diagnosis by ConvFormer-Based Biconditional Domain Adaptation Method
abstract
Domain adaptation-based methods have been proved success for cross-machine fault diagnosis. However, such methods suffer from limited diagnosis performance because the utilized networks typically rely on convolution layers with local receptive fields, failing to extract informative fault features, and the information on machine domain and fault category is not fully utilized, which prevents the transferability of fault features across machines. Towards these issues, a novel ConvFormer-based biconditional domain adaptation method (CFBDAM) is proposed to explore informative and highly-transferable fault features for accurate diagnosis. The proposed ConvFormer network first extracts global-local fault features in a parallel manner via a linear transformer and a separable shuffled CNN, respectively. The resulting features are then fed into a cross-attention feature fusion module to form informative diagnostic knowledge. Our ConvFormer is deployment-friendly owing to lightweight designs, such as linear and separation operations. To enhance cross-machine transferability of the informative fault features extracted by ConvFormer, a biconditional domain adaptation strategy is designed. It imposes biconditional constraints by using the information of both machine domain and fault category, thereby leading to highly-transferable fault features with domain insensitivity and category discriminability. Comprehensive experiments are conducted on six transfer diagnosis tasks across three machines. The experimental results show that CFBDAM achieves potential cross-machine diagnostic performance.
Xiaorong Zheng, Jiahao Nie 0001, Zhiwei He 0001, Mingyu Gao 0002
IEEE Trans. Ind. Informatics3
2024 Towards Category Unification of 3D Single Object Tracking on Point Clouds
abstract
Category-specific models are provenly valuable methods in 3D single object tracking (SOT) regardless of Siamese or motion-centric paradigms. However, such over-specialized model designs incur redundant parameters, thus limiting the broader applicability of 3D SOT task. This paper first introduces unified models that can simultaneously track objects across all categories using a single network with shared model parameters. Specifically, we propose to explicitly encode distinct attributes associated to different object categories, enabling the model to adapt to cross-category data. We find that the attribute variances of point cloud objects primarily occur from the varying size and shape (e.g., large and square vehicles v.s. small and slender humans). Based on this observation, we design a novel point set representation learning network inheriting transformer architecture, termed AdaFormer, which adaptively encodes the dynamically varying shape and size information from cross-category data in a unified manner. We further incorporate the size and shape prior derived from the known template targets into the model’s inputs and learning objective, facilitating the learning of unified representation. Equipped with such designs, we construct two category-unified models SiamCUT and MoCUT. Extensive experiments demonstrate that SiamCUT and MoCUT exhibit strong generalization and training stability. Furthermore, our category-unified models outperform the category-specific counterparts by a significant margin (e.g., on KITTI dataset, $\sim$12\% and $\sim$3\% performance gains on the Siamese and motion paradigms).
Jiahao Nie 0001, Zhiwei He 0001, Xueyi Zhou, Dong-Kyu Chae
ICLR2
2024 VoxelTrack: Exploring Multi-level Voxel Representation for 3D Point Cloud Object Tracking
abstract
Current LiDAR point cloud-based 3D single object tracking (SOT) methods typically rely on point-based representation network. Despite demonstrated success, such networks suffer from some fundamental problems: 1) It contains pooling operation to cope with inherently disordered point clouds, hindering the capture of 3D spatial information that is useful for tracking, a regression task. 2) The adopted set abstraction operation hardly handles density-inconsistent point clouds, also preventing 3D spatial information from being modeled. To solve these problems, we introduce a novel tracking framework, termed VoxelTrack. By voxelizing inherently disordered point clouds into 3D voxels and extracting their features via sparse convolution blocks, VoxelTrack effectively models precise and robust 3D spatial information, thereby guiding accurate position prediction for tracked objects. Moreover, VoxelTrack incorporates a dual-stream encoder with cross-iterative feature fusion module to further explore fine-grained 3D spatial information for tracking. Benefiting from accurate 3D spatial information being modeled, our VoxelTrack simplifies tracking pipeline with a single regression loss. Extensive experiments are conducted on three widely-adopted datasets including KITTI, NuScenes and Waymo Open Dataset. The experimental results confirm that VoxelTrack achieves state-of-the-art performance (88.3%, 71.4% and 63.6% mean precision on the three datasets, respectively), and outperforms the existing trackers with a real-time speed of 36 Fps on a single TITAN RTX GPU. The source code and model will be released.
Yuxuan Lu 0007, Jiahao Nie 0001, Zhiwei He 0001, Hongjie Gu
ACM Multimedia3
2024 SAR-SLAM: Self-Attentive Rendering-based SLAM with Neural Point Cloud Encoding
abstract
Neural implicit representations have recently revolutionized simultaneous localization and mapping (SLAM), giving rise to a groundbreaking paradigm known as NeRF-based SLAM. However, existing methods often fall short in accurately estimating poses and reconstructing scenes. This limitation largely stems from their reliance on volume rendering techniques, which oversimplify the modeling process. In this paper, we introduce a novel neural implicit SLAM system named SAR-SLAM to address these shortcomings. Our approach reconstructs Neural Radiance Fields (NeRFs) using a self-attentive architecture and represents scenes through neural point cloud encoding. Unlike previous NeRF-based SLAM methods, which depend on traditional volume rendering equations for scene representation and view synthesis, our method employs a self-attentive rendering framework with the Transformer architecture during mapping and tracking stages. To enable incremental mapping, we anchor scene features within a neural point cloud, striking a balance between estimation accuracy and computational cost. Experimental results on three challenging datasets show the superior performance and robustness of our SAR-SLAM compared to recent NeRF-based SLAM systems. The code will be released.
Zhiwei He 0001, Yuxiang Yang 0001, Jiahao Nie 0001, Jing Zhang 0037
ACM Multimedia2
2024 Multivariate time series anomaly detection with variational autoencoder and spatial-temporal graph network
Siwei Guan, Zhiwei He 0001, Shenhui Ma, Mingyu Gao 0002
Comput. Secur.2
2024 MPFormer: Multipatch Transformer for Multivariate Time-Series Anomaly Detection With Contrastive Learning
abstract
Recent unsupervised framework-based anomaly detection methods in Internet of Things (IoT) have been proved successful by learning temporal pattern representations from normal time-series data. However, such methods have long suffered from weakened normal-abnormal boundary incurred by discrimination-insufficient temporal pattern representations, due to: 1) normal time series often suffer from some unprocessed abnormal noise, interfering with model’s discriminative representations between normal and abnormal points and 2) previous methods rarely focus on contextual semantic features of time series that are also crucial for constructing discriminative representations. In this article, we propose multipatch contrastive learning framework for multivariate time-series anomaly detection (MPFormer), a transformer-based multipatch contrastive learning framework. This novel framework leverages contrastive learning to guide discriminative feature learning. It incorporates a data augmentation strategy to generate positive samples and maximizes the similarity between positive sample pairs, thereby capturing inherent representations of normal temporal patterns and enhancing robust discrimination ability to abnormal noise. To further facilitate discriminative representation modeling, we divide input time series into multiple patches and design a transformer-based dual-attention module to explore contextual semantic features in both interpatch and intrapatch views. Benefiting from discriminative representations being modeled, MPFormer effectively strengthens the normal-abnormal boundary, thus improving detection accuracy. Experimental results on six widely adopted data sets demonstrate that our proposed MPFormer outperforms existing baseline methods.
Shenhui Ma, Jiahao Nie 0001, Siwei Guan, Zhiwei He 0001, Mingyu Gao 0002
IEEE Internet Things J.4
2024 MSF-SLAM: Multi-Sensor-Fusion-Based Simultaneous Localization and Mapping for Complex Dynamic Environments
abstract
We proposed a multi-sensor fusion-based localization and scene reconstruction method for a complex dynamic scene. The multi-level fusion between multiple sensors was implemented by fusing data collected from different sensors in different system modules. In the front-end of the system, the camera and the LiDAR assisted each other. The LiDAR point clouds provided 3D information for the feature points in the image. The moving objects elimination method based on the image can remove the points on the moving objects in the LiDAR point clouds for localization accuracy improvement and static 3D scene reconstruction. To further improve the localization accuracy, a combination of visual loop closure detection and LiDAR loop closure detection was utilized to ensure the global consistency of scene reconstruction. At the system’s back-end, the observation model of different sensors was integrated to construct a multiple constraint factor graph with nonlinear optimization to obtain the optimal system states. Experimental results demonstrated that the proposed multi-sensor fusion-based localization and scene reconstruction algorithm could operate robustly in multiple complex dynamic scenes.
Zhiwei He 0001, Yuxiang Yang 0001, Jiahao Nie 0001, Zhekang Dong, Shuo Wang 0030, Mingyu Gao 0002
IEEE Trans. Intell. Transp. Syst.2
2024 SpikeTOD: A Biologically Interpretable Spike-Driven Object Detection in Challenging Traffic Scenarios
abstract
Artificial neural networks (ANN) have shown remarkable performance in intelligent transportation systems (ITS), especially for the traffic object detection. However, as the ITS is applied to a wider range of traffic scenarios, the increasing demand for the trade-off between detection performance and power resources has become inevitable. A biologically interpretable spike-driven traffic object detector for challenging scenarios is proposed in this paper, named SpikeTOD, achieving the trade-off between the accuracy and power consumption. Firstly, the spike neural network (SNN) is employed to realize energy-efficient object detection in traffic scenarios. And a local modulation-based integrate-and-fire (IF) neuron is designed, which provides an efficient way to convert the traffic detection model from ANN to SNN. Secondly, a biology-inspired detail-guided context-aware network (DCNet) is proposed to improve the detection performance. The integration of detail coherence and global priors is leveraged to selectively emphasize object features and improve the detection capabilities within challenging conditions. As far as we know, this is the first application of SNN in traffic object detection tasks. SpikeTOD achieved a mAP@50 of 46.11% on the BDD100K dataset with a power consumption of 4.73E-03J, demonstrating a more efficient trade-off in detection accuracy and power consumption. Notably, SpikeTOD maintained an average missed detection rate of 44.56%, further contributing to its overall efficacy in traffic object detection. Further, we conducted on road test by deploying SpikeTOD on Jetson Xavier NX and Loihi to demonstrate that model achieves a better balance between accuracy and power consumption.
Junfan Wang, Xiaoyue Ji, Zhekang Dong, Mingyu Gao 0002, Zhiwei He 0001
IEEE Trans. Intell. Transp. Syst.6
2023 GLT-T: Global-Local Transformer Voting for 3D Single Object Tracking in Point Clouds
abstract
Current 3D single object tracking methods are typically based on VoteNet, a 3D region proposal network. Despite the success, using a single seed point feature as the cue for offset learning in VoteNet prevents high-quality 3D proposals from being generated. Moreover, seed points with different importance are treated equally in the voting process, aggravating this defect. To address these issues, we propose a novel global-local transformer voting scheme to provide more informative cues and guide the model pay more attention on potential seed points, promoting the generation of high-quality 3D proposals. Technically, a global-local transformer (GLT) module is employed to integrate object- and patch-aware prior into seed point features to effectively form strong feature representation for geometric positions of the seed points, thus providing more robust and accurate cues for offset learning. Subsequently, a simple yet effective training strategy is designed to train the GLT module. We develop an importance prediction branch to learn the potential importance of the seed points and treat the output weights vector as a training constraint term. By incorporating the above components together, we exhibit a superior tracking method GLT-T. Extensive experiments on challenging KITTI and NuScenes benchmarks demonstrate that GLT-T achieves state-of-the-art performance in the 3D single object tracking task. Besides, further ablation studies show the advantages of the proposed global-local transformer voting scheme over the original VoteNet. Code and models will be available at https://github.com/haooozi/GLT-T.
Jiahao Nie 0001, Zhiwei He 0001, Yuxiang Yang 0001, Mingyu Gao 0002, Jing Zhang 0037
AAAI2
2023 OSP2B: One-Stage Point-to-Box Network for 3D Siamese Tracking
abstract
Two-stage point-to-box network acts as a critical role in the recent popular 3D Siamese tracking paradigm, which first generates proposals and then predicts corresponding proposal-wise scores. However, such a network suffers from tedious hyper-parameter tuning and task misalignment, limiting the tracking performance. Towards these concerns, we propose a simple yet effective one-stage point-to-box network for point cloud-based 3D single object tracking. It synchronizes 3D proposal generation and center-ness score prediction by a parallel predictor without tedious hyper-parameters. To guide a task-aligned score ranking of proposals, a center-aware focal loss is proposed to supervise the training of the center-ness branch, which enhances the network's discriminative ability to distinguish proposals of different quality. Besides, we design a binary target classifier to identify target-relevant points. By integrating the derived classification scores with the center-ness scores, the resulting network can effectively suppress interference proposals and further mitigate task misalignment. Finally, we present a novel one-stage Siamese tracker OSP2B equipped with the designed network. Extensive experiments on challenging benchmarks including KITTI and Waymo SOT Dataset show that our OSP2B achieves leading performance with a considerable real-time speed.
Jiahao Nie 0001, Zhiwei He 0001, Yuxiang Yang 0001, Zhengyi Bao, Mingyu Gao 0002, Jing Zhang 0037
IJCAI2
2023 FAML-RT: Feature alignment-based multi-level similarity metric learning network for a two-stage robust tracker
Jiahao Nie 0001, Zhekang Dong, Zhiwei He 0001, Mingyu Gao 0002
Inf. Sci.3
2023 GCEVT: Learning Global Context Embedding for Vehicle Tracking in Unmanned Aerial Vehicle Videos
abstract
Vehicle tracking in the unmanned aerial vehicle (UAV) videos is a fundamental but vital computer vision task. It mainly consists of two key components, that is, detection and reidentification (ReID). Recently, one-shot trackers, which integrate detection and ReID in a unified network, have received significant attention for their fast-tracking speed. However, existing one-shot trackers typically utilize local information to distinguish the detected targets. Due to the lack of global relations, which are key cues for tracking, these methods struggle to identify targets in UAV videos accurately. To alleviate the above issue, we deliberately design an ReID head that combines nonlocal blocks and the transformer layer to capture the global semantic relation in this letter. First, we propose a novel pyramid fusion network (PFN) to obtain the pixel-wise relations of features at multiple levels and aggregate them into features with richer semantic information. Then, we present a channel-wise transformer enhancer (CTE) to model the dependencies among the channels of the feature map and predict fine-grained identity embeddings. Extensive experiments on VisDrone2021 and UAVDT benchmarks demonstrate that our tracker, namely global context embedding for vehicle tracking (GCEVT), achieves state-of-the-art tracking performance.
Zhiwei He 0001, Mingyu Gao 0002
IEEE Geosci. Remote. Sens. Lett.2
2023 Learning task-specific discriminative representations for multiple object tracking
Jiahao Nie 0001, Zhiwei He 0001, Mingyu Gao 0002
Neural Comput. Appl.4
2023 Leveraging temporal-aware fine-grained features for robust multiple object tracking
Jiahao Nie 0001, Zhiwei He 0001, Mingyu Gao 0002
J. Supercomput.4
2023 Learning Localization-Aware Target Confidence for Siamese Visual Tracking
abstract
Siamese tracking paradigm has achieved great success, providing effective appearance discrimination and size estimation by classification and regression. While such a paradigm typically optimizes the classification and regression independently, leading to task misalignment (accurate prediction boxes have no high target confidence scores). In this paper, to alleviate this misalignment, we propose a novel tracking paradigm, called SiamLA. Within this paradigm, a series of simple, yet effective localization-aware components are introduced to generate localization-aware target confidence scores. Specifically, with the proposedlocalization-aware dynamic label(LADL) loss andlocalization-aware label smoothing(LALS) strategy, collaborative optimization between the classification and regression is achieved, enabling classification scores to be aware of location state, not just appearance similarity. Besides, we propose a separatelocalization-aware quality prediction(LAQP) branch to produce location quality scores to further modify the classification scores. To guide a more reliable modification, a novellocalization-aware feature aggregation(LAFA) module is designed and embedded into this branch. Consequently, the resulting target confidence scores are more discriminative for the location state, allowing accurate prediction boxes tend to be predicted as high scores. Extensive experiments are conducted on six challenging benchmarks, including GOT-10 k, TrackingNet, LaSOT, TNL2K, OTB100 and VOT2018. Our SiamLA achieves competitive performance in terms of both accuracy and efficiency. Furthermore, a stability analysis reveals that our tracking paradigm is relatively stable, implying that the paradigm is potential for real-world applications.
Jiahao Nie 0001, Zhiwei He 0001, Yuxiang Yang 0001, Mingyu Gao 0002, Zhekang Dong
IEEE Trans. Multim.2
2022 Hierarchical Feature Fusion based Reconstruction Network for Unsupervised Anomaly Detection
abstract
With the wide application of deep learning in the hydropower industry, many anomaly detection methods based on deep neural networks have been proposed to improve detection accuracy in electromechanical systems. However, they typically utilize recurrent neural networks to spontaneously learn the properties of multidimensional time-series data, which rarely using the feature dimension and time dimension. To solve this issue, in this paper, we propose a hierarchical feature fusion reconstruction network (HFFRN) to detect anomaly using feature dimension and time dimension information. Specifically, we first construct a feature extraction layer to strengthen the feature interaction among different layers, in which the shallow feature has rich detail information, and the deep feature has rich semantic information. Then, the hierarchical feature fusion layer is designed to fuse the information of the shallow and deep layers, allowing the model to perceive more information of the feature dimension and time dimension. We conduct a series of experiments on the SWaT dataset. The experimental results show that HFFRN outperforms the baseline method. Notably, HFFRN improves the F1 performance by 13% compared to the baseline. In addition, to prove the generalization of the HFFRN model, we also test the model on Mammography time series dataset.
Binjie Zhao, Jiahao Nie 0001, Siwei Guan, Zhiwei He 0001, Mingyu Gao 0002
ETFA5
2022 An improved normalized PLL-based high-order SMO for Sensorless Control of PMSM
abstract
When the permanent magnet synchronous motor (PMSM) is reversed, using the traditional sliding-mode observer (SMO), there is a 180° phase difference between the actual electrical angle and the estimated value. This paper proposes an improved normalized phase-locked loop-based high-order sliding-mode observer (HSMO-INPLL) method to solve this problem. First, HSMO is used to observe the current and back-EMF to reduce the chattering of SMO and the phase delay caused by the low-pass filter (LPF). Secondly, the 180° phase difference in the reverse rotation of traditional SMO is analyzed. Finally, the angle and speed information are extracted by INPLL to solve this problem. The simulation model and ARM experimental platform are built. Simulation and experimental results show that HSMO-INPLL can significantly improve the velocity observation range and signal-to-noise ratio (SNR) of SMO, and it is highly efficient and feasible.
Jiaxin Qian, Mingyu Gao 0002, Zhiwei He 0001, Huipin Lin
IECON4
2022 Spreading Fine-Grained Prior Knowledge for Accurate Tracking
abstract
With the widespread use of deep learning in single object tracking task, mainstream tracking algorithms treat tracking as a combined classification and regression problem. Classification aims at locating an arbitrary target, and regression aims at estimating the corresponding bounding box. In this paper, we focus on regression and propose a novel box estimation network, which consists of a transformer encoder target pyramid guide (TPG) and transformer decoder target pyramid spread (TPS). Specifically, the transformer encoder TPG is designed to generate fine-grained prior knowledge with explicit representation for template targets. In contrast to the raw transformer encoder, we capture the visual dependence through local-global self-attention and deem the multi-scale target regions as the “local” region. Using this fine-grained prior knowledge, we design the transformer decoder TPS to spread it to the subsequent search regions with high affinity to accurately estimate the bounding boxes. Considering that self-attention fails to model information interaction across channels between the template target and search regions, we develop a channel-wise cross-attention block within the TPS as compensation. Extensive experiments on the OTB100, UAV123, NFS, VOT2020, VOT2021, LaSOT, LaSOT_ext, TrackingNet and GOT-10k benchmarks show that the proposed box estimation network outperforms most existing box estimation methods. Furthermore, our trackers based on this estimation network exhibit a competitive performance against state-of-the-art trackers.
Jiahao Nie 0001, Zhiwei He 0001, Mingyu Gao 0002, Zhekang Dong
IEEE Trans. Circuits Syst. Video Technol.3
2022 A novel kernelized correlation filter by fusing multiple feature response maps, enhanced target re-detection, and improved model updating for visual tracking
Chenjie Du, Zhongping Ji, Zhekang Dong, Mingyu Gao 0002, Zhiwei He 0001
Vis. Comput.6
2019 SOC estimation of Lithium-ion battery based on an Extended H-infinity filter
abstract
The accuracy of the state of charge (SOC) is of great significance to electric vehicles and is one of the key technical issues in the development of electric vehicles. Therefore, an accurate SOC estimation strategy is important for the research of electric vehicle battery management. In this paper, the extended H-infinity filter (EHF) algorithm obtained by combined the extended Kalman filter (EKF) with the H- infinity filter (HIF) is chosen as the SOC estimation method. The algorithm can compensate for the error caused by the accuracy of the sensor and the lack of accuracy of the initial value to some extent. The experimental results show that this is an ideal SOC estimation method for electric vehicle batteries.
Tiantian Cai, Zhiwei He 0001, Mingyu Gao 0002, Jingbiao Liu
INDIN3
2019 State of Charge Estimation for Lithium-Ion Batteries Based on NARX Neural Network and UKF
abstract
Lithium-ion batteries have been widely used as the energy storage systems in electric vehicles. State of Charge (SOC) is one of the most important characteristics of battery system. It is essential for efficient use of the battery and ensures the safety of the electric vehicle. In this paper, a nonlinear autoregressive with exogenous inputs neural network (NARXNN) architecture is designed to estimate the SOC. The proposed method requires no model or knowledge of battery's internal parameters, but rather uses the battery's voltage, charge/discharge currents in various ambient temperature to accurately estimate battery's SOC. An unscented Kalman filter is used to reduce the errors in the neural network-based SOC estimation. The adaptability, efficiency, and robustness of the model are evaluated using the FUDS, US06 and DST driving cycles at varying temperatures conditions. The results prove that the proposed NARXNN-UKF model achieves higher accuracy with less computational time under different temperature conditions and electric vehicle driving cycles.
Xiaohan Qin, Mingyu Gao 0002, Zhiwei He 0001
INDIN3
2018 Battery charging and discharging feature extraction method based on the best u-shapelets
abstract
With the advancement of science and technology,batteries have become an indispensable item in our daily life. At the same time, the study of the charge-discharging curve of the battery plays an important role. The problem of battery charging and discharging curve can be regarded as a time series data mining problem. We utilize the unsupervised shape u-shapelets for time series data mining, which is a newly emerging tiny local feature that has been widely used in many fields, e.g., battery grouping. Experimental results show the practicability and effectiveness of the battery charge/discharge feature extraction method using the best u-shapelets, the ability of the local characteristics of u-shapelets to provide more insights for the data, and the sensitivity to irrelevant data in the charging and discharging curve of the battery is reduced. Extracting local feature u-shapelets from battery charging and discharging curves is helpful for battery grouping.
Jingping Chen, Mingyu Gao 0002, Zhiwei He 0001, Zhongfei Yu
INDIN4
2016 A machine vision based sealing rings automatic grabbing and putting system
abstract
In order to allow the robot to automatically grab the front sealing rings, in this paper, a recognition and grabbing system of sealing rings based on machine vision is presented. Machine vision technologies are first utilized to help locate the sealing rings and also determine their orientations, an industrial robot is then utilized to accomplish the grabbing of them. The whole system consists of a light source, a camera, an image processing machine and a 4-degrees of freedom industrial robot. Specifically, the system uses the method of the non calibration to accurately locate the positions of sealing rings. Then, sealing rings with a front view are automatically identified by the Hough transform algorithm. Finally, the industrial robot grabs the one of the sealing rings and put it at the proper position of a battery lid. Experimental results show that the system can recognize, grab and put the sealing rings successfully. The proposed system can improve the efficiency for the assembly line of the battery manufacturing and enhance the flexibility and adaptability of the robot.
Guojin Ma, Yanting Lou, Mingyu Gao 0002, Yuxiang Yang 0001, Zhiwei He 0001, Hongjuan Zhu
INDIN7
2016 A robust vision inspection system for detecting surface defects of film capacitors
Yuxiang Yang 0001, Zhengjun Zha, Mingyu Gao 0002, Zhiwei He 0001
Signal Process.4