VLDB 2026 Research / reviewers in the wild / expert
Zhipeng Wang 0002
dblp:56/5818-2
· DBLP profile ↗
15ranked-venue papers
2as first author
13since 2021 · last 2026
0000-0003-3039-7582ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 12 · 2 first-author · 12 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RFIDet: Visual Prior-Guided Rail Fastener Integrity Detection for UAV-Based Aerial Railroad InspectionabstractRail fasteners are crucial railroad infrastructure and their health status is directly connected to the safety of traveling trains. UAV-based rail fastener visual inspection has shown strong advantages over traditional inspection techniques. However, existing general-purpose object detection architectures inevitably suffer from the limitation that they can only detect objects that are clearly visible. Consequently, they tend to neglect objects affected by occlusion or shadows and lead to missed detection problem. The industrial application of such models can bring huge safety risks to long-term operations of safety-sensitive railroad systems. Concerning the issues, this paper proposes a visual prior-guided rail fastener integrity detection architecture (RFIDet) to realize coarse-to-fine detection of all rail fasteners, whether normally visible or visually obscured. RFIDet employs a two-stage pipeline: the visual prior guidance (VPG) stage generates standard rail fastener layout representation (SRFLR) for coarse priors, while the precise location search (PLS) stage enables NMS-free refinement using adaptive anchors designed from actual physical distance priors. SRFLR takes full advantage of inherent spatial priors of all rail fasteners to perceive a unified, interconnected, and coarse location distribution. Then all rail fastener candidates activated by those coarse locations are further trained to search and regress refined offsets to the final bounding boxes. Structural loss functions for both stages are customized to facilitate the detection of individual fasteners while constraining the overall spatial distribution of all fasteners. Experiments have verified the effectiveness and better robustness of the proposed RFIDet with the mAP50value increased by at least 5.9% compared to a series of general-purpose SOTA YOLO detectors. RFIDet outperforms the comparing algorithms especially when coming across unexpected occlusions or shadows. Limin Jia 0002, Honggui Han, Yong Qin 0002, Haonan Zhang 0002, Zhipeng Wang 0002 |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2025 | Dual-stage manifold preserving mixed supervised learning for bogie fault diagnosis under variable conditions
Ning Wang 0034, Limin Jia 0002, Yong Qin 0002, Dechen Yao, Zhipeng Wang 0002 |
Eng. Appl. Artif. Intell. | 6 |
| 2025 | Both Objects and Relationships Matter: Cross Coupled Transformer for Railway Scene PerceptionabstractRailway scene status perception is the foundation for ensuring the efficient and safe operation of trains. Achieving comprehensive railway scene perception requires not only detecting objects but also understanding the relationships between them. However, current railway scene perception technologies struggle with the latter, while general object and relationship perception methods perform poorly when facing railway scene with significant structural features and small object characteristics. Thus, this paper proposes Rail-former, the first status perception framework specifically designed for railway scene. Rail-former operates based on two key components: the Railway Feature Enhancement Module (RFEM) and the Cross-Coupled Transformer (CCTF). RFEM first leverages a cluster-based normalization evaluation method to condense railway scene structural features in spatial domain. It then applies high-dimensional information compensation to enhance small-object feature representation. The CCTF implements a cross-coupled structure between object decoder and triplet decoder based on a novel interactive attention mechanism, rather than conventional single-decoder or parallel dual-decoder structures, enabling guided enhancement of both object features and relationship features in railway scene. Additionally, we introduce a multimodal matching loss function that incorporates masks, categories, and bounding boxes to mitigate overfitting to a single modality during training. To the best of our knowledge, Rail-former is the first work in the railway domain that leverages panoptic parsing results to reveal relationships. Extensive experiments conducted on both our self-constructed railway dataset and a public dataset demonstrate the superior performance of our approach, achieving accuracy rates of 49.9% and 37.4%, surpassing the best competing methods by 7.8% and 1.1%, respectively. Dingyuan Bai, Baoqing Guo, Xingfang Zhou, Hongwei Wang 0008, Zhipeng Wang 0002 |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2025 | SRLF: Sparse Representation Learning Framework for Railroad Surrounding Potential Risk Perception Using UAV ImageryabstractRegular inspection of potential risks in railroad surroundings is essential for operational safety. Uncrewed aerial vehicles (UAVs) offer an effective solution with aerial mobility and long-distance coverage. However, existing methods struggle with rare but extremely high risks characterized by limited samples and complex feature distributions. To address this, we propose SRLF (Sparse Representation Learning Framework), which decomposes sparse risks (SR) perception into three components: capture, excavation, and learning. First, Buffer Decouple Learning (BDL) decouples objectness from classification to capture and enhance foreground perception. Second, Feature Space Dynamic Sampling (FSDS) leverages adaptive quantity sampling from multivariate Gaussian distributions to excavate discriminative SR representations. Third, Triple Similarity Loss (TSL) constructs a triple comparison mechanism to contrastively shape uncertainty surfaces between SRs and common safety hazards (CSHs). Finally, extensive experiments conducted on the UAV-based railroad surroundings dataset demonstrate that SRLF can achieve a high detection rate of CSHs (95.6% mAP) while maintaining low miss-detection rate for SRs (81.9% Recall and 0.5% FPR95). Fanteng Meng, Yong Qin 0002, Yunpeng Wu, Mingyang Chen 0001, Ninghai Qiu, Zhipeng Wang 0002, Chongchong Yu, Huaizhi Yang |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2024 | TriRNet: Real-Time Rail Recognition Network for UAV-Based Railway InspectionabstractUAVs have a broad application prospect in the field of railway inspection due to their excellent mobility and flexibility. However, it still faces challenges, such as high human labor costs and low intelligence levels. Therefore, it is of great significance to develop a real-time intelligent rail recognition algorithm that can be deployed on the onboard computing device to guide the UAV’s camera to follow the target rail area and complete the inspection automatically. However, a significant issue is that rails from the perspective of UAVs may appear with changing pixel widths and various inclination angles. Concerning the issue, a general and adaptive rail representation method based on projection length discrimination (RRM-PLD) is proposed. It can always select the optimal representation direction, horizontal or vertical, to represent any kind of rails. With the RRM-PLD, a novel architecture (Real-Time Rail Recognition Network, TriRNet) is proposed. In TriRNet, a designed inter-rail attention (IRA) mechanism is presented to fuse local features of single rails and global features of other rails to accurately discriminate the geometric distribution of all rails in the image in a regressive way and thus improve the final recognition accuracy. Further, one-to-one mapping from anchor points to final feature maps is established. It greatly simplifies the model design process and improves the model’s interpretability. Besides, detailed model training strategies are also presented. Extensive experiments have verified the effectiveness and superiority of the proposed formulation in terms of both network reasoning latency and recognition accuracy. Zhipeng Wang 0002, Limin Jia 0002, Yong Qin 0002, Donghai Song, Bidong Miao, Yixuan Geng |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2023 | An Online Health Monitoring Framework for Traction Motors in High-Speed Trains Using Temperature SignalsabstractThe health monitoring of traction motors is crucial for the prognostics and health management of high-speed trains. The temperature signal is an outstanding health indicator. Due to the representing of the traction motor's health conditions and low cost, accurate prediction for the motor temperature is conducive to early detection of abnormalities. However, the traditional prediction models are trained offline with high dependency on training data and cannot adapt to varying distributions of real data timely. Therefore, over time, the accuracies of these models always decrease noticeably. Concerning this issue, we propose an online health monitoring framework for traction motors using temperature signals. First, in the offline phase, multisensor signals are utilized to develop a generalized prediction model to absorb extensive information from temperature and relevant signals. Second, during the online phase, the training parameters are dynamically estimated to fulfill individualized learning by adopting a combination of the sample complexity and real-time prediction errors so as to fulfill individualized training according to the monitored data samples. Furthermore, a low-regret strategy is also presented in the online phase to determine the optimization target of the model to make the online update adaptive enough to the online prediction task. Consequently, the model can obtain new knowledge and greater understanding about the real data by online-learning continuously. Finally, the proposed framework is verified by actual data collected from Chinese high-speed trains. Compared with the conventional multilayer perceptron, gated recurrent unit, and long short-term memory, new patterns of stream data can be captured and adapted by using our framework, and the average root mean square errors of prediction results are reduced by 5%, 12%, and 11%, the average mean absolute percentage errors are reduced by 10%, 12%, and 11%, respectively. It is proven that our framework has high prediction accuracy and well-performed adaptability on real datasets. Honghui Dong, Zhipeng Wang 0002, Jie Man, Limin Jia 0002, Yong Qin 0002 |
IEEE Trans. Ind. Informatics | 3 |
| 2023 | 3DGraphSeg: A Unified Graph Representation- Based Point Cloud Segmentation Framework for Full-Range High-Speed Railway EnvironmentsabstractPoint cloud semantic segmentation (PCSS) is crucial for digital twins of high-speed railways. By now, the concerned subjects are confined within the interior infrastructures of railways. However, the surrounding environments are also important for the safe operation. Concerning this issue, a full-range high-speed railway scanning scheme based on unmanned-aerial-vehicle-borne LiDAR is utilized. However, the massive data volume and data distribution imbalance pose great challenges for PCSS. To address these issues, a novel PCSS framework called 3DGraphSeg is proposed in this article. To cope with the massive data volume, a structural representation algorithm named local embedding super-point graph is proposed to represent the vast point cloud into a concise graph while retain the data's inherent topology structure by local spatial embedding. Then, the gated integration graph convolutional network (GIGCN) is proposed to contextual segment the graph. In the GIGCN, to prevent the gradients from vanishing or exploding, the hidden states of gated recurrent units in every layer are integrated using a new layer named gated hidden states integration (GHSI). GHSI strengthens the back propagation by giving the loss function direct access to each layer and absorbs the features of different layers comprehensively, which enables the network to produce a smoother decision boundary and prevents the overfitting problem. Besides, to enhance its robustness to data imbalance, we propose a loss function: adaptive weighted cross entropy. Finally, five experiments are designed for verification. The proposed framework has excelled in different datasets and outperforms state-of-the-art approaches on the SemanticRail dataset. Yixuan Geng, Zhipeng Wang 0002, Limin Jia 0002, Yong Qin 0002, Yuanyuan Chai, Keyan Liu |
IEEE Trans. Ind. Informatics | 2 |
| 2023 | Self-Attentive Local Aggregation Learning With Prototype Guided Regularization for Point Cloud Semantic Segmentation of High-Speed RailwaysabstractPoint cloud semantic segmentation for railway infrastructures is an essential step towards establishing railway digital twins. Deep learning-based methods have shown great potential in this field compared to traditional methods that rely on hand-crafted features. However, deep learning-based methods for railway point clouds still face typical challenges that need to be addressed. In this regard, we propose a novel learning framework named SALAProNet, which consists of a set of effective and concise modular solutions. The first challenge addressed is the massive data scale of railway point clouds, which makes it difficult to directly process large-scale point clouds due to memory limitations. To solve this problem, we adapt efficient random sampling in the network and propose the Self-Attentive Aggregation (SAA) module based on an attention mechanism to greatly expand the receptive field, which covers the unsampled points and successfully retains information in a high-dimensional feature space. The second challenge is fine-grained segmentation, where we propose the Local Geometry Embedding (LGE) module to embed local geometry. With the help of context information provided by SAA, the network can perform fine-grained segmentation for railway infrastructures. The third challenge is the insufficient generalization ability of the network, where we propose a Prototype Guided Regularization (PGR) method to guide the network to segment the point cloud among railways with different construction standards. This method enhances the network’s interpretability and improves its generalization ability. We have validated our proposed framework through experiments on different datasets, and it outperforms state-of-the-art approaches. Zhipeng Wang 0002, Yixuan Geng, Limin Jia 0002, Yong Qin 0002, Yuanyuan Chai, Keyan Liu |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2023 | Manifold-Contrastive Broad Learning System for Wheelset Bearing Fault DiagnosisabstractNewly deployed trains have massive normal data and scarce faulty data for training, which limits the diagnosis accuracy with class imbalance problem of small samples. Considering that there are a lot unutilized information hidden in the abundant unlabeled monitoring data, this paper proposes a novel method named manifold-contrastive broad learning system, which utilizes the online updating approach for dealing with the class imbalance problem of small samples. This method constructs a novel one-class broad-learning classifier based on an inherency-guided comparison mechanism, which can classify and annotate unlabeled data online. This classifier employs contrastive manifold matrices to maintain the inherent structures, which is not affected to the overfitting caused by imbalanced samples. Secondly, inspired by the active learning, this classifier proposes the minimum-error strategy to annotate the samples by classifying the modes, which solves the problem of insufficient training data. Thirdly, this method applies an incremental learning strategy that continuously absorbs the newly annotated data to update the model online, which improves the model accuracy under the data imbalanced condition. Finally, the feasibility and effectiveness of the proposed method are verified by wheelset bearing data collected from a test rig of a Chinese rolling stock company. Ning Wang 0034, Limin Jia 0002, Huiyue Zhang, Yong Qin 0002, Xuejun Zhao, Zhipeng Wang 0002 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2023 | Segmentalized mRMR Features and Cost-Sensitive ELM With Fixed Inputs for Fault Diagnosis of High-Speed Railway TurnoutsabstractTurnouts are crucial to the safety of high-speed railways. Due to the intensive use and complex environment, breakdowns caused by different faults occur frequently in practice. Considering that the operation of turnouts is a multi-stage process during which each stage has its specific health characteristics, this paper proposes segmentalized maximal-relevancy and minimal-redundancy (mRMR) for feature extraction from each stage separately. Based on mathematical analysis of the turnout mechanism, the electric power curve is segmented into four stages, from which time-domain analysis and mRMR are combined to extract valid features corresponding to different movements respectively. Then, a novel classifier named cost-sensitive Extreme Learning Machine with fixed inputs (cf-ELM) is proposed for fault classification. We modify the inputs of ELM and define a new formula to limit the input weights and biases for the sake of stability of the network structure. Besides, a cost-sensitive optimization method is also presented in this classifier to embed the failure degree and data proportion into cost calculation rules to deal with data imbalance. To verify our proposed method, real data collected from a turnout of Beijing-Shanghai high-speed railway is used. It is proven by comparisons that the accuracy of our method has achieved 100% with fast running speed and also outperforms traditional methods in terms of stability and generalization remarkably. Zhipeng Wang 0002, Ning Wang 0034, Huiyue Zhang, Limin Jia 0002, Yong Qin 0002, Yakun Zuo, Yusheng Zhang, Honghui Dong |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2022 | UAV-LiDAR-Based Measuring Framework for Height and Stagger of High-Speed Railway Contact WireabstractThe height and stagger of the contact wire directly affect the energy supply of high-speed trains. To ensure the operation safety, there is an urgent demand for high-speed railways to measure the static parameters of contact wires all over the line with high precision and efficiency. However, this issue is barely discussed. Concerning the issue, this paper proposes a UAV-LiDAR-based measuring framework for the static height and stagger of high-speed railway contact wire. By mounting LiDAR on the UAV, the framework can efficiently collect data from the lines in service without occupying the train operating-diagrams. It is extremely significant for the high-speed and high-density railways. Then, we present self-adaptive extraction algorithms to extract critical infrastructures (rails, contact wires, masts and other suspensions) based on their specific geometric characteristics as well as the continuity and consistency of the spatial distributions along the line. Finally, the height and stagger are calculated by formulas automatically. To verify the framework in practice, we tested it on Beijing-Shanghai high-speed railway, which is the busiest high-speed railway in China. It is shown that the measurement error is within 9mm and the framework has potential to reform the inspection of high-speed railways. Yixuan Geng, Fengjun Pan, Limin Jia 0002, Zhipeng Wang 0002, Yong Qin 0002, Shiqi Li 0003 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2022 | Fully Decoupled Residual ConvNet for Real-Time Railway Scene Parsing of UAV Aerial ImagesabstractUAV-based automatic railway inspection is expected to have the potential to reform the inspection of railways. In this area, real-time railway scene parsing is quite essential. However, the limited computation resources of the UAV onboard computer pose a huge challenge for the algorithm to juggle a precise prediction with strong timeliness. Concerning this issue, this paper proposes a novel algorithm named deep fully decoupled residual convolutional network, which consists of fully decoupled residual blocks (Non-bottleneck-FDs) to deal with the dilemma between the high demand of real-time and limited resources. The residual block is constructed based on a new convolution which divides the standard convolution into three sequential convolutions to decouple the conventional operational correlations fully. Furthermore, a customized auxiliary line loss (LL) function is proposed to constrain the segmentation of railway and non-railway simultaneously without increasing the computation complexity. The proposed LL can force the predicted railway areas to concentrate in long strip areas precisely and inhibit their appearances in other impossible local areas. Subsequently, an integrated loss backpropagation strategy of the LL and cross-entropy function is presented. A comprehensive set of experiments are conducted for verification. Experiments demonstrate the superior performance of our approach with a more than$2\times $reduction in parameters and computation cost. Moreover, our approach also has a faster inference speed than the most existing lightweight architectures while providing comparable or higher accuracy. It is proven that our approach can reconcile the precise prediction with strong timeliness for railway scene parsing within the limitation of onboard computers. Besides, the results also imply its highest performance in terms of local details and edges of railway areas. Zhipeng Wang 0002, Limin Jia 0002, Yong Qin 0002, Yanbin Wei, Huaizhi Yang, Yixuan Geng |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2022 | Fully Decomposed Singular Value and Fixed Dictionary Extreme Learning Machine for Bogie Fault DiagnosisabstractAs an essential part in the rail train, the bogie plays an important role in the safety of the train operation. However, the fluctuant wheel-rail connection, as well as the structure and complex operating environment of the bogie always lead to low signal-to-noise ratio condition and complicated wheel-rail dynamic coupling relationship. The existing fault diagnosis methods can hardly perform well in this scenario. Concerning this issue, a novel feature extraction method named fully decomposed singular value (FdSV) is proposed in this paper. FdSV can decompose singular value characteristics of signals completely and increase the divergence of features to extract weak fault features effectively. Then, inspired by the theory of compressed perception and Hierarchy-ELM, a fixed dictionary extreme learning machine (FD-ELM) is also proposed for fault identification. This method calculates the weight matrix by formulas without randomization and removes the bias matrix. Therefore, it can easily discover the internal laws of data and improve the running speed and accuracy rapidly. Finally, the proposed algorithms have been verified by actual bogie data collected from bogies under low SNR and variable working conditions. Compared with SVD, the FdSV features are 1%-6% higher in testing accuracies. The accuracies of FD-ELM are 2-20% higher than the conventional ELM, H-ELM and SVM. Yakun Zuo, Ning Wang 0034, Limin Jia 0002, Huiyue Zhang, Zhipeng Wang 0002, Yong Qin 0002 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2020 | Densely pyramidal residual network for UAV-based railway images dehazing
Yunpeng Wu, Yong Qin 0002, Zhipeng Wang 0002 |
Neurocomputing | 3 |
| 2020 | An Improved Faster R-CNN for UAV-Based Catenary Support Device InspectionabstractThe catenary support device inspection is of crucial importance for ensuring safety and reliability of railway systems. At present, visual detection tasks of catenary support devices defect are performed by trained personnel based on the images taken periodically by industrial cameras installed on inspection vehicle in a limited period of time at midnight. However, the inspection mean is inappropriate for low efficiency and high cost. This paper presents a novel network based on unmanned aerial vehicle (UAV) images for catenary support device inspection and focuses on small object detection and the imbalanced dataset. With regards to the first aspect, based on a pyramid network structure, the improved Faster R-CNN consists of a top-down-top feature pyramid fusion structure, which heavily fuses high-level semantic information and low-level detail information. The feature map fusions of three different pooling scales are employed for improving detection accuracy of predicted bounding boxes. With regards to the second, we copy and paste the small proportion objects of dataset for avoiding category imbalance. Finally, quantitative and qualitative evaluations illustrate that the improved Faster-RCNN achieves better performance over the classic methods, yet remains convenient and efficient. Zhipeng Wang 0002, Yunpeng Wu, Yong Qin 0002, Xianbin Cao 0003 |
Int. J. Softw. Eng. Knowl. Eng. | 2 |