Wei Zhou 0088

dblp:69/5011-88 · DBLP profile ↗
← Back
14ranked-venue papers
9as first author
14since 2021 · last 2026
0000-0003-3225-0576ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 10 · 7 first-author · 10 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 IoT-Enabled Cooperative Autonomous Driving: A Hierarchical Spatial-Temporal Transformer Framework for Trajectory Prediction
abstract
Accurate trajectory prediction for surrounding agents is of paramount importance for autonomous vehicles (AVs). Given that a single AV’s perception systems may face challenges in detecting and predicting agents that are occluded or located at relatively long distances, the Internet of Things (IoT) has enabled cooperative autonomous driving through networked sensing and communication. In this context, vehicle-infrastructure cooperative (VIC) solutions have emerged as a promising approach for accurate trajectory prediction. Therefore, this study proposes a hierarchical spatial-temporal transformer framework for VIC trajectory prediction. The framework employs an Encoder-Decoder architecture that takes trajectory sequences from different views and vectorized maps as inputs. The encoder leverages multi-head attention mechanisms and VIC fusion module to extract spatial-temporal interaction features and aggregate cross-view information, while the decoder generates multimodal trajectory predictions in a sequential manner. Specifically, the spatial-temporal transformer in this study adopts a two-level hierarchical structure: agent-level and block-level. The agent-level transformer executes temporal encoding and interaction feature extraction from each agent’s historical trajectory sequences while aggregating features from both ego-vehicle and infrastructure views. The block-level transformer alternately extracts spatial-temporal interaction features across blocks, capturing long-range temporal dependencies and maintaining longterm spatial-temporal coherence. Experimental analysis on real-world vehicle-infrastructure cooperative datasets demonstrates that the proposed method effectively utilizes VIC data to enhance trajectory prediction accuracy. Moreover, our proposed method achieves favorable performance even in challenging scenarios involving occlusion, long distances, and sparse trajectories, which could serve as an efficient approach for cooperative autonomous driving.
Lei Zhao 0019, Wei Zhou 0088, Sixuan Xu, Chen Wang 0085
IEEE Internet Things J.2
2026 A Data-Driven Path Choice Estimation Method Based on Passengers' Time-Dependent Travel Characteristics
abstract
The complex patterns of passenger travel in metropolitan rail systems pose substantial challenges for daily operations. To allocate revenues and optimize train schedules, operators require accurate data on passenger flows and path choices. We propose a data-driven method for path choice estimation that captures time-dependent segmented travel information in public transport systems. The method combines mobile signals, automatic fare collection (AFC), and train operation data to infer passenger travel behaviors within the metro system. Aspatiotemporal metro hypernetwork models passenger travel and train operations along feasible paths. The proposed travel-chain reconstruction method estimates segmented travel time distributions for each station across multiple time periods. A nonlinear programming optimization model calibrates path choice probabilities by minimizing discrepancies between theoretical and observed origin-destination (OD) journey time distributions. Case studies using both synthetic and real world data validate the model’s performance for the Nanjing metro system, China. Synthetic data validation shows the proposed model achieves approximately 85%-90% prediction accuracy for OD paths with ascending transfers, and remains robust to changes in path travel times. Real data results show that the TIDE-PCE model accurately estimates true path fractions and observed exit demand, with superior performance during peak hours.
Zhuangbin Shi, Wei Zhou 0088
IEEE Trans. Intell. Transp. Syst.5
2025 Data-Efficient Object Detection on Construction Sites Using Reweighting Mechanism and Cross-Batch Contrastive Learning
abstract
Deep learning-based object detectors face unique challenges in construction sites, including severe occlusions from complex spatial arrangements, similar appearances between different construction elements, and computational constraints of edge devices. While few-shot detection methods reduce annotation requirements, existing approaches using computationally intensive architectures [e.g., Faster region-based convolutional neural network (R-CNN) or detection transformer (DETR)] struggle to address these construction-specific challenges while maintaining real-time performance. To tackle these issues, we propose a lightweight few-shot detection framework based on CenterNet2, specifically designed for construction environments. Our framework introduces, first a dual reweighting mechanism to handle severe occlusions and complex spatial relationships, second a cross-batch contrastive learning strategy to enhance feature discrimination between similar construction objects, and third an efficient architecture suitable for edge deployment. Experimental results demonstrate our framework achieves 56.6% mean average precision while maintaining real-time performance, significantly outperforming existing methods in construction-specific scenarios.
Wei Zhou 0088, Lei Zhao 0019, Hongpu Huang, Chen Wang 0085
IEEE Trans. Ind. Informatics1
2025 Soft Actor-Critic Deep Reinforcement Learning for Train Timetable Collaborative Optimization of Large-Scale Urban Rail Transit Network Under Dynamic Demand
abstract
To address the collaborative issue in large-scale urban rail transit (URT) network operations, this paper proposes an adaptive real-time control framework based on the Soft Actor-Critic (SAC) deep reinforcement learning (DRL) method, featuring flexible train scheduling capabilities. First, by analyzing dynamic passenger travel behavior (e.g., entering/exiting stations, transferring) and train operation events (e.g., dispatching, interstation running, station dwelling), the control problem is modeled as a Markov Decision Process (MDP) and an efficient URT simulation environment is constructed. Then, considering constraints such as train capacity and dispatch intervals, a train scheduling model is developed to minimize both passenger costs and operational costs. Subsequently, the real-time state of the URT system is represented by the overall number of passengers present at every platform, and train dispatch intervals on all lines are used as decision variables. A solving algorithm based on the SAC framework is developed. Finally, experimental results on a large-scale URT network comprising 10 lines demonstrate the effectiveness of the proposed framework, showing superior performance compared to other reinforcement learning algorithms and traditional heuristic optimization algorithms. The proposed approach achieves a 1.63% reduction in average passenger waiting time, equivalent to 2.09 seconds, while utilizing 49 fewer trains, representing a 2.97% decrease, compared to the second-best TD3 algorithm.
Longhui Wen, Liyang Hu, Wei Zhou 0088, Gang Ren 0005
IEEE Trans. Intell. Transp. Syst.3
2024 Fine-Grained Vehicle Make and Model Recognition Framework Based on Magnetic Fingerprint
abstract
Fine-grained vehicle classification of make and model recognition is of practical significance and can be widely applied in the tasks such as traffic perception and control. To date, vision-based vehicle make and model recognition methods have improved through the years; however, they often fail in low-visibility scenes. This paper proposes a fine-grained vehicle model recognition framework based on magnetic fingerprints instead of videos from surveillance cameras. The magnetic fingerprint refers to the pattern of magnetic field data obtained when vehicles pass magnetic sensors, which have less sensitivity and more robustness towards low-visibility scenes than surveillance cameras. Specifically, the proposed framework first utilizes adversarial autoencoders to extract the magnetic fingerprint from the raw magnetic field data, keeping detailed information for fine-grained vehicle model recognition as complete as possible. Then, an AdaBoostSVM algorithm is used to identify vehicle models from the magnetic fingerprint and further alleviate the hard-negative issue to reduce false classifications. Massive experimental results show that the proposed fine-grained vehicle magnetic fingerprint/AdaBoostSVM method achieves state-of-the-art results on fine-grained vehicle make and model recognition. The average precision reached 94.3% and the average recall reached 94.2% in the dataset containing 36 vehicle models in total. Notably, the proposed method outperforms surveillance camera methods, especially in low-visibility scenes, making it a valuable addition to future intelligent transportation systems.
Hancheng Zhang, Wei Zhou 0088, Zuoxv Wang, Zhendong Qian
IEEE Trans. Intell. Transp. Syst.2
2024 Teaching Segment-Anything-Model Domain-Specific Knowledge for Road Crack Segmentation From On-Board Cameras
abstract
Road crack segmentation from on-board cameras is a highly desirable yet challenging task for road condition inspection and maintenance. However, existing methods trained on small-scale datasets present limited performance from such challenging perspectives due to the lack of sufficient prior knowledge and effective generalizability. To address this limitation, this paper incorporates a vision foundation model named Segment-Anything-Model (SAM) and fully leverages its rich prior knowledge and strong generalizability to achieve crack segmentation. Also, we construct a customized crack segmentation dataset shot from on-board cameras. Considering the direct use of SAM might not correctly segment cracks, some lightweight and learnable crack adaptation layers are developed and integrated into SAM’s image encoder, which take the patch embeddings of the input image and its high-frequency components within road regions as joint inputs. During training, the parameters of the crack adaptation layers are fine-tuned to acquire domain-specific knowledge, while the parameters of the image encoder remain frozen, preserving SAM’s rich prior knowledge. Additionally, a sparse prompt generation method is proposed based on the high-frequency components within road regions, which guides the SAM model to better focus on high-frequency regions that may contain cracks. Experimental results demonstrate that the proposed framework achieves state-of-the-art performance, with an improved average precision of 8.14% compared to Mask2Former. Furthermore, the parameter-efficient transfer learning framework significantly reduces the number of parameters requiring fine-tuning, thereby improving efficiency and reducing training costs. The dataset proposed in this paper is available athttps://github.com/TRMetaGroup/CrackSeg.
Wei Zhou 0088, Hongpu Huang, Hancheng Zhang, Chen Wang 0085
IEEE Trans. Intell. Transp. Syst.1
2024 Pedestrian Crossing Intention Prediction From Surveillance Videos for Over-the-Horizon Safety Warning
abstract
Pedestrian crossing intention prediction could effectively prevent traffic injuries and improve pedestrian safety. This paper focuses on pedestrian crossing intention prediction from surveillance cameras, which could provide over-the-horizon safety warnings and has the potential to better ensure pedestrian safety, compared with that from on-board cameras. However, most prediction-based methods are designed with a fundamental assumption that the visual data is collected from an on-board camera rather than a bird-eye-view one, thus the prevalent methods in this research domain do not match surveillance scenarios. To deal with this issue, an automated learning framework is proposed, in which a pedestrian-centric environment graph is primarily constructed to reflect visual variations and spatiotemporal relationships between pedestrians and their surroundings. After that, a Graph Convolutional Network (GCN) based environment encoder and a pedestrian-state encoder are designed to extract prominent environment features and pedestrian behavior features, respectively. Finally, an intention prediction decoder is developed to extrapolate the probability of crossing intention. Experimental results demonstrate that each component in the framework contributes to performance improvement and their combination obtains state-of-the-art performance, suggesting the effectiveness and superiority of our framework.
Wei Zhou 0088, Lei Zhao 0019, Sixuan Xu, Chen Wang 0085
IEEE Trans. Intell. Transp. Syst.1
2024 All-Day Vehicle Detection From Surveillance Videos Based on Illumination-Adjustable Generative Adversarial Network
abstract
Vehicle detection from surveillance videos is of great significance for various Intelligent Transportation System (ITS) applications. However, existing deep learning methods oftentimes fail under nighttime conditions on account of the lack of sufficient labeled nighttime data. To fill this gap, this paper proposes a novel framework for all-day vehicle detection, by introducing an illumination-adjustable GAN (IA-GAN). The IA-GAN transforms labeled daytime images into multiple nighttime images with diverse illumination, using an adjustable illumination vector as input. Notably, we utilize gray histogram distributions to automatically generate illumination labels, by which IA-GAN gains the knowledge of simulating lights. Following that, we construct a large dataset containing both labeled daytime images and all generated synthetic nighttime images with bounding box labels. Finally, a detector named Day-Night Balanced EfficientDet (DNBED) is developed for all-day vehicle detection. The experiments show that the proposed framework yields promising performance for all-day vehicle detection and competitive results for nighttime vehicle detection compared to existing GANs, indicating the effectiveness of proposed framework. The privacy-sanitized image data and its corresponding labels will be made publicly available at https://github.com/vvgoder/SEU_PML_Dataset.
Wei Zhou 0088, Chen Wang 0085, Yiran Ge, Longhui Wen, Yunfei Zhan
IEEE Trans. Intell. Transp. Syst.1
2024 Monitoring-Based Traffic Participant Detection in Urban Mixed Traffic: A Novel Dataset and A Tailored Detector
abstract
Monitoring-based traffic participant detection (TPD) is a highly desirable but challenging task. So far, deep learning-based methods have attained significant improvements on the TPD task, but oftentimes fail in urban mixed traffic due to the lack of relevant datasets and suitable detectors. In this study, we propose a large and detailed dataset named SEU_PML specialized for monitoring-based TPD in urban mixed traffic. This dataset contains a total of 270,684 objects annotated with 2D bounding box and covers 13 sub-categories, having (i) high-resolution images (from$1920\times 1080$to$4096\times 2160$pixels), (ii) high-quality annotation (annotation accuracy reaches 98%), and (iii) rich traffic scenarios covering diverse traffic scenes as well as different weather and illumination conditions. The mixed traffic along with high-quality annotation bring about a variety of small objects. To further address the issue on small object detection, we propose a novel detector named YOLO SOD, which embeds a super-resolution feature extraction module and uses knowledge distillation to learn the knowledge how the detector with high-resolution inputs perceives small objects. Moreover, a novel loss function named S-IoU is designed to enable YOLO SOD to focus more on small objects. Experimental results show that (1) the YOLO SOD detector has an increased mAP of 1.58% and operates approximately four times faster when compared to a state-of-art detector; (2) the detectors trained on the SEU_PML dataset have a strong transferability and could be well applied to traffic participant detection in urban mixed traffic. Our dataset is now available athttps://github.com/vvgoder/SEU_PML_Dataset.
Wei Zhou 0088, Chen Wang 0085, Jingxin Xia, Zhendong Qian
IEEE Trans. Intell. Transp. Syst.1
2023 Automatic waste detection with few annotated samples: Improving waste management efficiency
Wei Zhou 0088, Lei Zhao 0019, Hongpu Huang, Yuzhi Chen, Sixuan Xu, Chen Wang 0085
Eng. Appl. Artif. Intell.1
2023 Multi-weighted graph 3D convolution network for traffic prediction
Chen Wang 0085, Sixuan Xu, Wei Zhou 0088, Yuzhi Chen
Neural Comput. Appl.4
2023 An Appearance-Motion Network for Vision-Based Crash Detection: Improving the Accuracy in Congested Traffic
abstract
Crash detection is of great significance for traffic emergency management. Video-based approaches can effectively save the manpower monitoring cost and have achieved promising results in recent studies. However, they sometimes fail to correctly identify crashes in congested traffic. To fill the gap, this paper proposes a novel appearance-motion network for improving video-based crash detection performance in congested traffic. The appearance-motion network utilizes two paralleled convolutional networks (i.e., an appearance network and a motion network) to extract both appearance features and motion features of crashes. To learn discriminative appearance features for differentiating crashes in congested traffic scene (CCT) with non-crashes in congested traffic scene (NCCT), an auxiliary network combined with a triplet loss are introduced to train the appearance network. To better capture crash motion features in congested traffic, an optical flow learner is built in the motion network and trained to extract more fine-grained motion information. Moreover, a temporal attention module is applied to enable the motion network to focus on valuable frames. Experimental results show that the proposed network achieves a state-of-the-art result on crash detection and the introduction of the three components (i.e., the auxiliary network, the optical flow learner and the temporal attention module) effectively reduces false alarm rate by 28.07% and miss rate by 27.08% on crash detection in congested traffic. Our dataset will be available athttps://github.com/vvgoder/Dataset_for_crashdetection.
Wei Zhou 0088, Longhui Wen, Yunfei Zhan, Chen Wang 0085
IEEE Trans. Intell. Transp. Syst.1
2022 A vision-based abnormal trajectory detection framework for online traffic incident alert on freeways
Wei Zhou 0088, Yunhong Yu, Yunfei Zhan, Chen Wang 0085
Neural Comput. Appl.1
2022 An Automated Learning Framework With Limited and Cross-Domain Data for Traffic Equipment Detection From Surveillance Videos
abstract
Traffic equipment detection from surveillance videos is of practical significance for temporary traffic element update in high-precision maps. However, there is little relative research developed due to limited labeled data. Based on a detector dubbed Faster R-CNN, we propose an automated learning framework that utilizes easy-to-obtain Internet images containing traffic equipment to acquire the capability of detecting traffic equipment from surveillance videos. In this framework, an appearance weighting module using a comprehensive feature aggregation method is designed to allow Faster R-CNN to converge and generalize quickly by taking limited data (i.e., less than 30 images per class) as input. To further address the cross-domain issue brought by the domain gap between the Internet images and the surveillance video frames, a domain adaptation learning scheme is developed, which aims to align the two domains and guide the framework to learn more robust domain-invariant features. Experimental results show that both the appearance weighting module and the domain adaptation learning scheme could bring a great performance improvement. Moreover, the combination of the two results in a state-of-the-art performance (mAP of 44.6%) even if only 30 training images per class are provided. To sum up, the proposed framework is suitable for traffic equipment detection from surveillance videos and provides an inspiration for other detection tasks with limited and cross-domain data, allowing humans to reduce their efforts and time required for arduous data collection and annotation.
Wei Zhou 0088, Chen Wang 0085, Yunfei Zhan, Yulu Dai
IEEE Trans. Intell. Transp. Syst.1