VLDB 2026 Research / reviewers in the wild / expert
Zhiyong Cui
dblp:157/2959
· DBLP profile ↗
31ranked-venue papers
1as first author
29since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 14 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 1 first-author · 9 since 2021Computer networks · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FLAMD: Assembling Multilevel Feature Learning and Multiscale Motion Decoding for Trajectory Prediction in Autonomous VehiclesabstractAccurate motion prediction is a cornerstone of safe autonomous driving and a key enabler for building efficient intelligent transportation systems in Internet of Things environments. Mainstream methods typically follow a divided paradigm: first modeling scene dependencies and then generating trajectories in a one-shot manner. Despite their promising performance, persistent challenges still exist, especially regarding insufficient future representation learning due to constrained information and suboptimal motion decoding from overlooking varying temporal correlations. In this paper, we propose FLAMD, a novel trajectory prediction framework that integrates multi-level feature learning with multi-scale motion decoding. Different from previous methods with decoupled design, FLAMD introduces a joint optimization scheme for multimodal motion representation learning and scene context modeling. Equipped with the carefully designed motion-aware feature interaction module and cross-consistency enhancement mechanism, FLAMD promotes long-range interactions and feature co-evolution between historical observations and anticipated scenarios, guided by diverse future-oriented queries. This mutual interaction enables deeper reasoning about scene dynamics, leading to context-aware and future-guided representations. In the decoding phase, FLAMD proposes to extend traditional single-stage decoding into a multi-scale framework. Multiple horizon-specific decoders operate at different temporal resolutions, each custom-made for the complexity of its assigned prediction window. This structure helps to capture temporal correlations and allocates appropriate feature granularity across varying future steps, thereby improving the capacity of trajectory generation. Extensive experiments on two real-world benchmarks demonstrate the superiority of our approach in generating accurate and reliable predictions in diverse traffic scenarios. Zhengxing Lan, Lingshan Liu, Haiyang Yu 0002, Zhiyong Cui, Yilong Ren |
IEEE Internet Things J. | 4 |
| 2026 | LLM-enabled universal traffic signal control across different intersections and traffic flows
Haiyang Yu 0002, Han Jiang 0003, Minda Li, Zhiyong Cui, Yilong Ren |
Knowl. Based Syst. | 5 |
| 2026 | Fusion Filtering for RIS-Assisted Vehicle Localization With Unknown Inputs: A Privacy-Preserving Output-Mask StrategyabstractThis article investigates the privacy-preserving fusion filtering problem for vehicle localization subject to unknown inputs. High-accuracy and privacy-preserving localization is essential for the safe and reliable operation of intelligent transportation systems. In practice, vehicle localization is often challenged by measurement bias caused by nonline-of-sight (NLOS) propagation, unknown inputs arising from uncertainties or acceleration/deceleration maneuvers, and risks of signal and location privacy leakage. To address these issues, a privacy-preserving fusion filtering framework is proposed by integrating the reconfigurable intelligent surface (RIS) technique, an unknown-input estimation method, and an output-mask mechanism. A unified measurement model is first developed to represent both line-of-sight (LOS) and NLOS scenarios, and RISs are employed to construct virtual LOS paths to mitigate NLOS effects. An output-mask-based privacy-preserving strategy is then designed to prevent eavesdroppers from inferring vehicle locations or signal characteristics while maintaining the required filtering performance. Based on this model, an unknown-input estimator and a privacy-preserving filter are constructed, and the impacts of NLOS propagation, unknown inputs, and masking on filtering performance are analyzed. The associated gain matrix parameters are obtained by solving the corresponding optimization problems. Furthermore, an RIS-assisted privacy-preserving fusion filtering algorithm is developed to exploit multisource measurements and enhance localization robustness and accuracy. Simulation results demonstrate the effectiveness of the proposed method. Kaiqun Zhu, Zidong Wang 0001, Xinhu Zheng, Zhiyong Cui, Zhenning Li 0001, Keqiang Li 0002 |
IEEE Trans. Ind. Informatics | 4 |
| 2026 | OccSora: 4D Occupancy Generation Models as World Simulators for Autonomous DrivingabstractUnderstanding the evolution of 3D scenes is crucial for autonomous driving. While conventional methods describe scene development through individual instance motions, world models provide a generative framework for modeling overall scene dynamics. However, most existing approaches rely on autoregressive next-token prediction, which suffers from error accumulation and limited global spatiotemporal reasoning, leading to degraded long-term consistency. To address these issues, we propose a diffusion-based 4D occupancy generation model, OccSora, to simulate 3D world evolution for autonomous driving. A 4D scene tokenizer is introduced to obtain compact spatiotemporal representations and enable high-quality reconstruction of long occupancy sequences. We then train a diffusion transformer on these representations to generate 4D occupancy conditioned on trajectory prompts. Experiments on the nuScenes dataset with Occ3D annotations show that OccSora can generate 16s videos with authentic 3D layout and strong temporal consistency. With trajectory-aware 4D generation, OccSora has the potential to serve as a world simulator for autonomous driving decision-making. Project page: https://wzzheng.net/OccSora. Wenzhao Zheng, Yilong Ren, Han Jiang 0003, Zhiyong Cui, Haiyang Yu 0002, Jiwen Lu |
IEEE Trans. Image Process. | 5 |
| 2026 | Privacy-Constrained Edge Computing for Driver Expression Recognition: What and How to Offload?
Bo Fan 0003, Chongwei Zhou, Tongfei Li, Zhiyong Cui |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2026 | OAlight: Overflow-Aware Adaptive Traffic Signal Control via State-Wise Reinforcement LearningabstractAdaptive traffic signal control dynamically adapts to real-time traffic flow variations, effectively mitigating traffic congestion through efficient decision-making. Current ATSC methods primarily focus on vehicle mobility metrics, such as average waiting time and average speed. However, urban road networks require not only high traffic efficiency but also proactive prevention of critical scenarios like overflow, which, if not intervened, may degrade network-wide traffic conditions and compromise safety. Therefore, in order to a balance between overflow prevention and efficiency, we model traffic signal control as a Constrained Markov Decision Process (CMDP), and propose OAlight, an overflow-aware adaptive traffic signal control framework based on reinforcement learning (RL). Unlike existing RL approaches that merely incorporate overflow-related features into the reward function, OAlight integrates overflow awareness throughout the entire RL lifecycle — environment interaction, state representation, and policy learning. Specifically, considering the explicit overflow factor of excessive queue lengths and the implicit overflow factor of excessive waiting time for individual vehicles, we construct a dual safeguard mechanism for overflow prevention through environmental interactions by developing a comprehensive reward-cost function framework. In this framework, the reward function simultaneously captures both explicit and implicit factors, while two specialized cost functions assess behaviour that violates threshold performance metrics of explicit and implicit factors, respectively. We then design a state encoder and a state-action joint encoder to extract overflow-aware representations from the state space. Finally, we employ a Proximal Policy Optimization with Lagrangian method to guide policy learning, ensuring compliance with overflow constraints while optimizing control actions. The experimental results indicate that our approach achieves superior performance on our proposed overflow cost metric without compromising traffic efficiency metrics on both synthetic and real-world datasets, compared to the current state-of-the-art traffic signal control methods. Furthermore, the environment interaction module exhibits broad compatibility with diverse RL methods, significantly reducing overflow risks. Haiyang Yu 0002, Han Jiang 0003, Zhiyong Cui, Yilong Ren |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2026 | Emergency Events Traffic Flow Forecasting Using Text-Prompt-Guided Multimodal Large Language ModelsabstractEmergency events like traffic accidents and natural disasters frequently cause severe disruptions to urban traffic patterns, posing challenges to conventional forecasting methods based on historical data. Recent progress has explored integrating auxiliary textual information from social media, news reports, and incident details to enhance forecasting models. However, these approaches often fail to effectively align semantic context from textual data with the spatio-temporal patterns in urban traffic flow. To bridge this disparity, we propose a novel framework called TPGM-LLM, which leverages text-prompt-guided multimodal large language models to dynamically assimilate real-time data, including incident reports, traffic sensors, and weather conditions. Specifically, the proposed framework integrates a pre-trained LLM-based text encoder to interpret descriptions of emergency events as semantic prompts. Furthermore, it incorporates a dynamic spatio-temporal hypergraph module that employs FastDTW to capture non-local dependencies among different road segments. Additionally, a multimodal feature extraction LLM is utilized to merge textual guidance with traffic dynamics in a hierarchical manner, enhancing the accuracy of long-term traffic forecasting. Experimental results on the Beijing Text-Traffic dataset (BjTT) demonstrate that the proposed model outperforms existing methods, highlighting its effectiveness in complex and dynamic traffic situations. Our code is available athttps://github.com/luyaxuan/TPGM-LLM Yaxuan Lu, Guangyu Huo, Xiaohui Cui, Boyue Wang, Yong Zhang 0029, Zhiyong Cui |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2025 | NEST: A Neuromodulated Small-world Hypergraph Trajectory Prediction Model for Autonomous DrivingabstractAccurate trajectory prediction is essential for the safety and efficiency of autonomous driving. Traditional models often struggle with real-time processing, capturing non-linearity and uncertainty in traffic environments, efficiency in dense traffic, and modeling temporal dynamics of interactions. We introduce NEST (Neuromodulated Small-world Hypergraph Trajectory Prediction), a novel framework that integrates Small-world Networks and hypergraphs for superior interaction modeling and prediction accuracy. This integration enables the capture of both local and extended vehicle interactions, while the Neuromodulator component adapts dynamically to changing traffic conditions. We validate the NEST model on several real-world datasets, including nuScenes, MoCAD, and HighD. The results consistently demonstrate that NEST outperforms existing methods in various traffic scenarios, showcasing its exceptional generalization capability, efficiency, and temporal foresight. Our comprehensive evaluation illustrates that NEST significantly improves the reliability and operational efficiency of autonomous driving systems, making it a robust solution for trajectory prediction in complex traffic environments. Chengyue Wang 0001, Haicheng Liao, Bonan Wang, Yanchen Guan, Bin Rao 0003, Ziyuan Pu, Zhiyong Cui, Cheng-Zhong Xu 0001, Zhenning Li 0001 |
AAAI | 7 |
| 2025 | Authentic 4D Driving Simulation with a Video Generation Model
Wenzhao Zheng, Dalong Du, Yilong Ren, Han Jiang 0003, Zhiyong Cui, Haiyang Yu 0002, Jie Zhou 0001, Shanghang Zhang |
ICCV | 7 |
| 2025 | SAH-Drive: A Scenario-Aware Hybrid Planner for Closed-Loop Vehicle Trajectory GenerationabstractReliable planning is crucial for achieving autonomous driving. Rule-based planners are efficient but lack generalization, while learning-based planners excel in generalization yet have limitations in real-time performance and interpretability. In long-tail scenarios, these challenges make planning particularly difficult. To leverage the strengths of both rule-based and learning-based planners, we proposed the Scenario-Aware Hybrid Planner (SAH-Drive) for closed-loop vehicle trajectory planning. Inspired by human driving behavior, SAH-Drive combines a lightweight rule-based planner and a comprehensive learning-based planner, utilizing a dual-timescale decision neuron to determine the final trajectory. To enhance the computational efficiency and robustness of the hybrid planner, we also employed a diffusion proposal number regulator and a trajectory fusion module. The experimental results show that the proposed method significantly improves the generalization capability of the planning system, achieving state-of-the-art performance in interPlan, while maintaining computational efficiency without incurring substantial additional runtime. Yuqi Fan 0003, Zhiyong Cui, Zhenning Li 0001, Yilong Ren, Haiyang Yu 0002 |
ICML | 2 |
| 2025 | Reusing Attention for One-stage Lane Topology UnderstandingabstractUnderstanding lane topology relationships accurately is critical for safe autonomous driving. However, existing two-stage methods suffer from inefficiencies due to error propagations and increased computational overheads. To address these challenges, we propose a one-stage architecture that simultaneously predicts traffic elements, lane centerlines and topology relationship, improving both the accuracy and inference speed of lane topology understanding for autonomous driving. Our key innovation lies in reusing intermediate attention resources within distinct transformer decoders. This approach effectively leverages the inherent relational knowledge within the element detection module to enable the modeling of topology relationships among traffic elements and lanes without requiring additional computationally expensive graph networks. Furthermore, we are the first to demonstrate that knowledge can be distilled from models that utilize standard definition (SD) maps to those operates without using SD maps, enabling superior performance even in the absence of SD maps. Extensive experiments on the OpenLane-V2 dataset show that our approach outperforms baseline methods in both accuracy and efficiency, achieving superior results in lane detection, traffic element identification, and topology reasoning. Our code is available at https://github.com/Yang-Li-2000/one-stage.git. Yang Li 0178, Zongzheng Zhang, Xuchong Qiu, Xinrun Li, Leichen Wang, Ruikai Li, Zhenxin Zhu, Huan-ang Gao, Xiaojian Lin, Zhiyong Cui, Hang Zhao 0021, Hao Zhao 0002 |
IROS | 11 |
| 2025 | Vulnerability-aware and Curiosity-driven Adversarial Reinforcement Learning Policy for Safety-Critical Scenario GenerationabstractAutonomous vehicles (AVs) face significant threats to their safe operation in complex traffic environments. Adver-sarial policy for scenario generation has been established as a robust paradigm for enhancing AV resilience against adver-sarial perturbations through proactive exposure to synthetically engineered safety-critical scenarios. Training an attacker within an adversarial policy, allowing the target AV to expose vulnerabilities through interaction with this attacker. However, adversarial policies in existing methodologies often get stuck in a loop of over-exploiting established vulnerabilities, resulting in poor exploration for AVs. To overcome the limitations, we introduce a pioneering framework termed the vulnerability-aware and curiosity-driven adversarial reinforcement learning policy. Specifically, during the traffic vehicle attacker training phase, a surrogate network is employed to fit the value function of the AV victim, providing dense information about the victim's inherent vulnerabilities. Subsequently, random network distillation is used to characterize the novelty of the scenario, constructing an intrinsic reward to guide the attacker in exploring unexplored territories. Experimental results demonstrated that the adversarial policy embedded within the attacker exhibited robustness in convergence and significantly enhanced the activation of policy exposure in learning-based AVs, outperforming both other adversarial modalities and alternative reinforcement learning approaches, with a notable reduction in crash rates. The code is available at https://github.com/caixxuan/VCAT. Xuan Cai, Zhiyong Cui, Xuesong Bai, Ruimin Ke, Haiyang Yu 0002, Yilong Ren, Zechang Ye |
IV | 2 |
| 2025 | CLlight: Enhancing representation of multi-agent reinforcement learning with contrastive learning for cooperative traffic signal control
Yilong Ren, Han Jiang 0003, Zhiyong Cui, Haiyang Yu 0002 |
Expert Syst. Appl. | 5 |
| 2025 | Brimory: Bringing Humanoid Memory Into Trajectory Prediction Model for Autonomous DrivingabstractPredicting the future trajectories of moving agents serves as a pivotal endeavor in advancing safe autonomous driving forward within the context of the evolving Internet of Things technology. Despite significant progress driven by deep learning methods, a gap in predictive capability remains when compared to experienced human drivers, particularly in scenarios that demand heightened perception and broader situational understanding. In this study, we propose Brimory, a novel trajectory prediction approach inspired by human driving, designed to equip autonomous vehicles with humanoid memory systems. The core concept of Brimory is to emulate human driving processes by enabling the model to deeply comprehend traffic scenes, establish working memory, and iteratively accumulate to form driving experience. Brimory first introduces the working memory generator, which processes multiple interactions between the scene elements in the driving environment. It is noted that unlike traditional models confined to the time domain, Brimory also extracts underlying dependencies in the frequency domain. The long-term memory builder module is then developed to adaptively retrieve information from working memory and gradually accumulate driving experience with the help of the designed iterative updating mechanism. To mitigate the impact of irrelevant components on long-term memory effectiveness, we devise a co-modulation strategy that retains meaningful representations in a learnable manner. Finally, the combined humanoid memories are leveraged by the multimodal trajectory predictor to forecast future motions. Extensive experiments on benchmark datasets demonstrate the superiority, feasibility, and efficiency of Brimory. The results underscore the potential of humanoid memory frameworks in trajectory prediction, offering a promising path toward the realization of safer autonomous vehicles. Zhengxing Lan, Lingshan Liu, Yilong Ren, Zhiyong Cui, Haiyang Yu 0002 |
IEEE Internet Things J. | 4 |
| 2025 | MLB-Traj: Map-Free Trajectory Prediction With Local Behavior Query for Autonomous DrivingabstractPredicting future motions of target agents is crucial to ensuring the safety of autonomous vehicles in Internet of Things environments. Although significant progress has been made in this field, most mainstream approaches rely heavily on high-definition (HD) maps, which may not always be available or accurate owing to the high costs of map construction and the potential localization errors. Without the explicit guidance of HD maps, trajectory prediction would become more challenging. To address this challenge, we present MLB-Traj, an innovative framework for map-free motion prediction based on local behavior queries. MLB-Traj leverages the observation that agents often follow local behavior patterns in specific traffic scenarios, where these local behaviors reveal the potential trajectories of the targets and contain scenario-consistent information. It starts with a hierarchical dynamic modal query paradigm that first captures the scene’s general modal characteristics and then models target-specific properties. A dual Transformer query mechanism aggregates multiscale relationships to facilitate this process. To tackle potential inconsistency in map-free forecasting, we introduce a trajectory consistency module. It ensures the continuity of inferred trajectories by utilizing patch-wise interaction representations to capture local temporal dependencies, while also learning more robust representations by simulating the model’s response to spatial inconsistency in its predictions. Extensive experiments conducted on real-world datasets validate the effectiveness of MLB-Traj. The results indicate that our framework outperforms existing methods, highlighting its superiority in generating accurate predictions in map-free settings. Yilong Ren, Lingshan Liu, Zhengxing Lan, Zhiyong Cui, Haiyang Yu 0002 |
IEEE Internet Things J. | 4 |
| 2025 | Toward City-Scale Vehicular Crowd Sensing: A Decentralized Framework for Online Participant RecruitmentabstractAs an emerging urban computing paradigm, vehicle crowd sensing (VCS) leverages ubiquitous vehicles as basic sensing units to achieve more efficient data collection. However, with the expansion of the sensing range, the tens of thousands of vehicles and the openness of urban road networks pose a huge challenge for real-time participant recruitment in online VCS systems. To achieve efficient city-scale VCS, this paper proposes Dec-Recruiter, a decentralized framework for online participant recruitment. Specifically, Dec-Recruiter adopts a novel decision-making mode based on virtual grid agents, where vehicles traveling in the same direction within the same grid are considered homogeneous, simplifying the recruitment of specific vehicles to the selection of the number of vehicles in each direction. Meanwhile, through policy sharing among grid agents with the same geographic features, the complexity of city-scale VCS participant recruitment is further reduced. The core of Dec-Recruiter is a multi-agent contextual double-deep Q-network algorithm, which enables grid agents with different geographic features to collaborate on network-wide sensing tasks through their asynchronous decision-making. In this process, the Gaussian function is employed to adjust the reward distribution to address cold-start and data integrity issues in VCS. In addition, to ensure the convergence and training efficiency of the model on large-scale road networks, a pre-training-based transfer learning paradigm is also introduced. We conduct extensive experiments on both synthetic and real-world datasets. The results demonstrate that Dec-Recruiter can effectively recruit appropriate participants in the large-scale VCS and outperforms all baselines. Han Jiang 0003, Yilong Ren, Yanan Zhao 0002, Zhiyong Cui, Haiyang Yu 0002 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2025 | RM2Occ: Re-Projection Multi-Task Multi-Sensor Fusion for Autonomous Driving 3D Object Detection and Occupancy PerceptionabstractOccupancy prediction plays a crucial role in supporting autonomous driving planning and decision-making. Existing methods typically rely on modular stacking and fusion techniques of object detection, semantic segmentation, and depth estimation to achieve 3D occupancy. However, they fail to deeply explore the transformation relationships between 2D and 3D spaces and to efficiently fuse the different characteristics of multi-source sensors. We propose R$M^{2}$Occ, the first 3D occupancy perception network that integrates multi-sensor fusion based on different sensor principles and achieves multi-task learning. To leverage the rich 2D semantic information captured by cameras and elevate it to the 3D domain, we begin by querying and populating predefined empty voxels with multi-view image features. Subsequently, we progressively fuse 3D LiDAR point clouds with these populated voxels through an unbalanced fusion strategy that effectively supplements missing information and suppresses noise. Leveraging IMU data and calibration parameters, we then re-project the enriched voxels back onto the 2D image plane according to camera coordinates, performing a secondary query using the semantic segmentation results to recover semantic details potentially lost due to radar fusion limitations and incomplete voxel querying. Finally, supported by a multi-task detection head, R$M^{2}$Occ simultaneously accomplishes 3D object detection, semantic segmentation, Bird’s Eye View (BEV) detection, and full-scene grid occupancy prediction, enabling comprehensive multi-task output. Extensive experiments and ablation studies on the nuScenes dataset demonstrate that R$M^{2}$Occ significantly outperforms existing state-of-the-art methods, establishing a new paradigm for accurate and efficient multi-sensor fusion and multi-task perception in autonomous driving scenarios. Yilong Ren, Minda Li, Han Jiang 0003, Zhiyong Cui, Mengmeng Yang 0001, Haiyang Yu 0002, Diange Yang |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2024 | Less is More: Efficient Brain-Inspired Learning for Autonomous Driving Trajectory PredictionabstractAccurately and safely predicting the trajectories of surrounding vehicles is essential for fully realizing autonomous driving (AD). This paper presents the Human-Like Trajectory Prediction model (HLTP++), which emulates human cognitive processes to improve trajectory prediction in AD. HLTP++ incorporates a novel teacher-student knowledge distillation framework. The “teacher” model, equipped with an adaptive visual sector, mimics the dynamic allocation of attention human drivers exhibit based on factors like spatial orientation, proximity, and driving speed. On the other hand, the “student” model focuses on real-time interaction and human decision-making, drawing parallels to the human memory storage mechanism. Furthermore, we improve the model’s efficiency by introducing a new Fourier Adaptive Spike Neural Network (FA-SNN), allowing for faster and more precise predictions with fewer parameters. Evaluated using the NGSIM, HighD, and MoCAD benchmarks, HLTP++ demonstrates superior performance compared to existing models, which reduces the predicted trajectory error with over 11% on the NGSIM dataset and 25% on the HighD datasets. Moreover, HLTP++ demonstrates strong adaptability in challenging environments with incomplete input data. This marks a significant stride in the journey towards fully AD systems. Haicheng Liao, Yongkang Li 0003, Zhenning Li 0001, Chengyue Wang 0001, Guofa Li, Chunlin Tian, Zilin Bian, Kaiqun Zhu, Zhiyong Cui, Jia Hu 0003 |
ECAI | 9 |
| 2024 | A Cognitive-Driven Trajectory Prediction Model for Autonomous Driving in Mixed Autonomy Environments
Haicheng Liao, Zhenning Li 0001, Chengyue Wang 0001, Bonan Wang, Hanlin Kong, Yanchen Guan, Guofa Li, Zhiyong Cui |
IJCAI | 8 |
| 2024 | AccidentGPT: A V2X Environmental Perception Multi-modal Large Model for Accident Analysis and PreventionabstractTraffic accidents are a significant factor leading to injuries and property losses, prompting extensive research in the field of traffic safety. However, previous studies, whether focused on static environment assessment, dynamic driving analysis, pre-accident prediction, or post-accident rule checks, have often been conducted independently. Our introduces V2X Environmental Perception Multi-modal Large Model AccidentGPT for accident analysis and prevention. AccidentGPT establishes a multi-modal information interaction framework based on multisensory perception. It adopts a holistic approach to address traffic safety issues, providing environmental perception for autonomous vehicles to avoid collisions and maintain control. In human-driven vehicles, it offers proactive safety warnings, blind spot alerts, and driving suggestions through human-machine dialogue. Additionally, it aids traffic police and management agencies in considering factors such as pedestrians, vehicles, roads, and the environment for intelligent real-time analysis of traffic safety. The system also conducts a thorough analysis of accident causes and post-accident liabilities, making it the first large-scale model to integrate comprehensive scene understanding into traffic safety research. Project page: https://accidentgpt.github.io Yilong Ren, Han Jiang 0003, Pinlong Cai, Daocheng Fu, Zhiyong Cui, Haiyang Yu 0002, Xuesong Wang 0006, Hanchu Zhou, Helai Huang, Yinhai Wang |
IV | 7 |
| 2024 | Performance Limit: Fuzzy Logic Based Anti-Collision Algorithm for Industrial Internet of Things ApplicationsabstractPassive RFID has the advantages of rapid identification of multi-target objects and low implementation cost. It is the most critical technology in the Industrial Internet of Things information-gathering layer and is extensively applied in various industries, such as smart production, asset management, and monitoring. The signal collision caused by the communication between the reader/writer and tags sharing the same wireless channel has caused a series of problems, such as the reduction of the identification efficiency of the reader/writer and the increment of the missed reading rate, thus restricting the further development of RFID. At present, many hybrid anti-collision algorithms integrate the advantages of Aloha and TS algorithms to optimize RFID system performance, but these solutions also suffer from performance bottlenecks. In order to break through such performance bottleneck, based on the ISE-BS algorithm, we combined the sub-frame observation mechanism and the Q value adjustment strategy and proposed two hybrid anti-collision algorithms. The experimental results show that the two algorithms proposed in this paper have obvious advantages in system throughput, time efficiency and other metrics, surpassing existing UHF RFID anti-collision algorithms. Dongbo Zhong, Zhiyong Cui, Yufei Xie |
Int. J. Uncertain. Fuzziness Knowl. Based Syst. | 2 |
| 2024 | MuGIL: A Multi-Graph Interaction Learning Network for Multi-Task Traffic Prediction
Haiyang Yu 0002, Han Jiang 0003, Zhenliang Ma, Zhiyong Cui, Yilong Ren |
Knowl. Based Syst. | 5 |
| 2023 | Bandit-based data poisoning attack against federated learning for autonomous driving models
Shuo Wang 0017, Qianmu Li, Zhiyong Cui, Jun Hou 0002, Chanying Huang |
Expert Syst. Appl. | 3 |
| 2023 | Distributed Multiagent Deep Reinforcement Learning for Multiline Dynamic Bus Timetable OptimizationabstractAs a primary countermeasure to mitigate traffic congestion and air pollution, promoting public transit has become a global census. Designing a robust and reliable bus timetable is a pivotal step to increase ridership and reduce operating cost for transit authorities. However, most previous studies on bus timetabling rely on historical passenger count and travel time data to generate static schedules, which often yield biased results in these uncertain scenarios, such as demand surge or adverse weather. In addition, acquiring real-time passenger origin/destination from a limited number of running buses is not feasible. This article considers the multiline dynamic bus timetable optimization problem as a Markov decision process model to address the aforementioned issues, and proposes a multiagent deep reinforcement learning framework to ensure effective learning from the imperfect-information game, where the passenger demand and traffic condition are not always known in advance. Moreover, a distributed reinforcement learning algorithm is applied to overcome the limitation of high computational cost and low efficiency. A case study of multiple bus lines in Beijing, China, confirms the effectiveness and efficiency of the proposed model. The results demonstrate that our method outperforms heuristic and state-of-the-art reinforcement learning algorithms by reducing 20.30% of operating and passenger costs compared with actual timetables. Haoyang Yan, Zhiyong Cui, Xinqiang Chen, Xiaolei Ma |
IEEE Trans. Ind. Informatics | 2 |
| 2022 | Toward a real-time Smart Parking Data Management and Prediction (SPDMP) system by attributes representation learningabstractManaging and estimating the availability information of parking lots is of great importance to travelers and managers. However, the task is very challenging since the occupancy rate is affected by various factors, including spatial-temporal features, parking lot attributes features, and environmental changes. Previous studies mostly focus on the short-term prediction by capturing the historical sequential dependencies among inputs and outputs, which leads to low estimation accuracy for long-term prediction and limited scalability for real deployment in parking lots. To address the challenges, a comprehensive framework for real-time Smart Parking Data Management and Prediction (SPDMP) system is proposed. Three types of data sources, including historical sequential data, real-time sequential data, and attributes category data, are sufficiently integrated into a customized Parking Availability Prediction (PAP) neural network by representation learning and heterogeneous feature embedding. Specifically, instead of using parking lot property information and environmental data directly, the authors design an Attributes Tensor Embedding Component (ATEC) to integrate the intra-class affection and interclass representations and correlations by two steps of customized embedding process. To balance the impact of the indiscriminate features for various prediction targets, this study proposes a multi-factor attention mechanism to learn the weights and help the PAP achieve a stable performance for time-various tasks. From extensive experiments on two real-world large-scale data sets collected in China and USA (including 19 urban parking and 2 truck parking lots), PAP achieves superior prediction results on both urban parking (6.69% average MAPE) and truck parking prediction (7.63% average MAPE) for five prediction time slots (5 min, 10 min, 30 min, 1 h, and 2 h). Furthermore, even with limited training data, PAP still shows better transferability and adaptability for various types of lots. The proposed SPDMP is selected and used as a pilot test bed by the Washington State Department of Transportation (WSDOT) for truck parking information promotion. Hao (Frank) Yang, Ruimin Ke, Zhiyong Cui, Yinhai Wang, Karthik Murthy |
Int. J. Intell. Syst. | 3 |
| 2022 | Multimodal Traffic Speed Monitoring: A Real-Time System Based on Passive Wi-Fi and Bluetooth Sensing TechnologyabstractTraffic speed is one of the critical indicators reflecting traffic status of roadway networks. The abnormality and sudden changes of traffic speed indicate the occurrence of traffic congestions, accidents, and events. Traffic control and management systems usually take the spatiotemporal variations of traffic speed as the critical evidence to dynamically adjust the traffic signal timing plan, broadcast traffic accidents, and form a management strategy. Meanwhile, transport is multimodal in most cities, including vehicles, pedestrians, and bicyclists. Traffic states of different traffic modes are usually used simultaneously as the significant input of advanced traffic control systems, e.g., multiobjective traffic signal control system, connected vehicles, and autonomous driving. In previous studies, Wi-Fi and Bluetooth passive sensing technology was demonstrated as an effective method for obtaining traffic speed data. However, there are some challenges that greatly affect the accuracy the estimated traffic speed, e.g., traffic mode uncertainty and the errors caused by sensors’ detection range. Thus, this study develops a real-time method for estimating the multimodal traffic speed of road networks covered by Wi-Fi and Bluetooth passive sensors. To address the two identified challenges, an algorithm is developed to correct the biased estimated traffic speed based on the received signal strength indicator of Wi-Fi and Bluetooth signals, and a novel semisupervised Possibilistic Fuzzy$C$-Means clustering algorithm is proposed for identifying traffic modes of Wi-Fi and Bluetooth device owners. The performance of the proposed algorithms is evaluated by comparing with the selected baseline algorithms. The experimental results indicate the superiority of the proposed algorithm. The proposed method of this study can provide accurate and real-time multimodal traffic speed information for supporting traffic control and management, and, thus, improving the operational performance of the whole road network. Ziyuan Pu, Zhiyong Cui, Jinjun Tang, Shuo Wang 0017, Yinhai Wang |
IEEE Internet Things J. | 2 |
| 2021 | Monitoring Public Transit Ridership Flow by Passively Sensing Wi-Fi and Bluetooth Mobile DevicesabstractReal-time public transit ridership flow and origin-destination (O-D) information is essential for improving transit service quality and optimizing transit networks in smart cities. The effectiveness and accuracy of the traditional survey-based methods and smart card data-driven methods for O-D information inference have multiple disadvantages in terms of biased results, high latency, insufficient sample size, and the high cost of time and energy. By considering the ubiquity of smart mobile devices in the world, monitoring public transit ridership flow can be accomplished by passively sensing Wi-Fi and Bluetooth (BT) mobile devices of passengers. This study proposed a system for monitoring real-time public transit passenger ridership flow and O-D information based on customized Wi-Fi and BT sensing device. By combining the consideration of the assumed overlapping feature spaces of passenger and nonpassenger media access control address data, a three-step data-driven algorithm framework for estimating transit ridership flow and O-D information is proposed. The observed ridership flow is used as the ground truth for evaluating the performance of the proposed algorithm. According to the evaluation results, the proposed algorithm outperformed all selected baseline models and the existing filtering methods. The findings of this study can help to provide real time and precise transit ridership flow and O-D information for supporting transit vehicle management and the quality of service enhancement. Ziyuan Pu, Meixin Zhu, Zhiyong Cui, Xiaoyu Guo 0006, Yinhai Wang |
IEEE Internet Things J. | 4 |
| 2021 | Forecasting Transportation Network Speed Using Deep Capsule Networks With Nested LSTM ModelsabstractAccurate and reliable traffic forecasting for complicated transportation networks is of vital importance to modern transportation management. The complicated spatial dependencies of roadway links and the dynamic temporal patterns of traffic states make it particularly challenging. To address these challenges, we propose a new capsule network (CapsNet) to extract the spatial features of traffic networks and utilize a nested LSTM (NLSTM) structure to capture the hierarchical temporal dependencies in traffic sequence data. A framework for network-level traffic forecasting is also proposed by sequentially connecting CapsNet and NLSTM. On the basis of literature review, our study is the first to adopt CapsNet and NLSTM in the field of traffic forecasting. An experiment on a Beijing transportation network with 278 links shows that the proposed framework with the capability of capturing complicated spatiotemporal traffic patterns outperforms multiple state-of-the-art traffic forecasting baseline models. The superiority and feasibility of CapsNet and NLSTM are also demonstrated, respectively, by visualizing and quantitatively evaluating the experimental results. Xiaolei Ma, Houyue Zhong, Yi Li 0046, Junyan Ma, Zhiyong Cui, Yinhai Wang |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2021 | Deep Learning Architecture for Short-Term Passenger Flow Forecasting in Urban Rail TransitabstractShort-term passenger flow forecasting is an essential component in urban rail transit operation. Emerging deep learning models provide good insight into improving prediction precision. Therefore, we propose a deep learning architecture combining the residual network (ResNet), graph convolutional network (GCN), and long short-term memory (LSTM) (called “ResLSTM”) to forecast short-term passenger flow in urban rail transit on a network scale. First, improved methodologies of the ResNet, GCN, and attention LSTM models are presented. Then, the model architecture is proposed, wherein ResNet is used to capture deep abstract spatial correlations between subway stations, GCN is applied to extract network topology information, and attention LSTM is used to extract temporal correlations. The model architecture includes four branches for inflow, outflow, graph-network topology, as well as weather conditions and air quality. To the best of our knowledge, this is the first time that air-quality indicators have been taken into account, and their influences on prediction precision quantified. Finally, ResLSTM is applied to the Beijing subway using three time granularities (10, 15, and 30 min) to conduct short-term passenger flow forecasting. A comparison of the prediction performance of ResLSTM with those of many state-of-the-art models illustrates the advantages and robustness of ResLSTM. Moreover, a comparison of the prediction precisions obtained for time granularities of 10, 15, and 30 min indicates that prediction precision increases with increasing time granularity. This study can provide subway operators with insight into short-term passenger flow forecasting by leveraging deep learning models. Jinlei Zhang, Feng Chen 0029, Zhiyong Cui, Yinan Guo 0002, Yadi Zhu |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2020 | Traffic Graph Convolutional Recurrent Neural Network: A Deep Learning Framework for Network-Scale Traffic Learning and ForecastingabstractTraffic forecasting is a particularly challenging application of spatiotemporal forecasting, due to the time-varying traffic patterns and the complicated spatial dependencies on road networks. To address this challenge, we learn the traffic network as a graph and propose a novel deep learning framework, Traffic Graph Convolutional Long Short-Term Memory Neural Network (TGC-LSTM), to learn the interactions between roadways in the traffic network and forecast the network-wide traffic state. We define the traffic graph convolution based on the physical network topology. The relationship between the proposed traffic graph convolution and the spectral graph convolution is also discussed. An L1-norm on graph convolution weights and an L2-norm on graph convolution features are added to the model's loss function to enhance the interpretability of the proposed model. Experimental results show that the proposed model outperforms baseline methods on two real-world traffic state datasets. The visualization of the graph convolution weights indicates that the proposed framework can recognize the most influential road segments in real-world traffic networks. Zhiyong Cui, Kristian Henrickson, Ruimin Ke, Yinhai Wang |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2017 | Real-Time Bidirectional Traffic Flow Parameter Estimation From Aerial VideosabstractUnmanned aerial vehicles (UAVs) are gaining popularity in traffic monitoring due to their low cost, high flexibility, and wide view range. Traffic flow parameters such as speed, density, and volume extracted from UAV-based traffic videos are critical for traffic state estimation and traffic control and have recently received much attention from researchers. However, different from stationary surveillance videos, the camera platforms move with UAVs, and the background motion in aerial videos makes it very challenging to process for data extraction. To address this problem, a novel framework for real-time traffic flow parameter estimation from aerial videos is proposed. The proposed system identifies the directions of traffic streams and extracts traffic flow parameters of each traffic stream separately. Our method incorporates four steps that make use of the Kanade-Lucas-Tomasi (KLT) tracker, k-means clustering, connected graphs, and traffic flow theory. The KLT tracker and k-means clustering are used for interest-point-based motion analysis; then, four constraints are proposed to further determine the connectivity of interest points belonging to one traffic stream cluster. Finally, the average speed of a traffic stream as well as density and volume can be estimated using outputs from previous steps and reference markings. Our method was tested on five videos taken in very different scenarios. The experimental results show that in our case studies, the proposed method achieves about 96% and 87% accuracy in estimating average traffic stream speed and vehicle count, respectively. The method also achieves a fast processing speed that enables real-time traffic information estimation. Ruimin Ke, Zhibin Li 0003, John Ash, Zhiyong Cui, Yinhai Wang |
IEEE Trans. Intell. Transp. Syst. | 5 |