Yongkang Liu 0005

dblp:52/8970-5 · DBLP profile ↗
← Back
22ranked-venue papers
3as first author
19since 2021 · last 2026
0000-0001-6677-5122ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 2 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Computer networks · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2026 HONEST-CAV: Hierarchical Optimization of Network Signals and Trajectories for Connected and Automated Vehicles with Multi-Agent Reinforcement Learning
Changxin Wan, Peng Hao 0001, Kanok Boriboonsomsin, Matthew J. Barth, Yongkang Liu 0005, Seyhan Ucar, Guoyuan Wu 0001
IV6
2025 Generalizing Vehicle Energy Consumption Models: Introducing the Large Energy Model Concept
Shinhoon Kim, Vishnu Raghupathy, Yongkang Liu 0005, Keisuke Niimi, Takahiro Mochihara, Jiawei Yong
IEEE Big Data3
2025 MamBEV: Enabling State Space Models to Learn Birds-Eye-View Representations
abstract
3D visual perception tasks, such as 3D detection from multi-camera images, are essential components of autonomous driving and assistance systems. However, designing computationally efficient methods remains a significant challenge. In this paper, we propose a Mamba-based framework called MamBEV, which learns unified Bird's Eye View (BEV) representations using linear spatio-temporal SSM-based attention. This approach supports multiple 3D perception tasks with significantly improved computational and memory efficiency. Furthermore, we introduce SSM based cross-attention, analogous to standard cross attention, where BEV query representations can interact with relevant image features. Extensive experiments demonstrate MamBEV's promising performance across diverse visual perception metrics, highlighting its advantages in input scaling efficiency compared to existing benchmark models.
Hongyu Ke, Jack Morris, Kentaro Oguchi 0001, Xiaofei Cao, Yongkang Liu 0005, Haoxin Wang 0003, Yi Ding 0010
ICLR5
2025 TinyBEV: Compact Temporal Fusion for Multi-View 3D Perception
abstract
Multi-view camera-based 3D object detection through unified Bird's Eye View (BEV) representation has become popular for autonomous driving due to its low cost, but efficiently inferring precise spatial and temporal information from cameras alone remains a significant challenge. Transformer-based approaches have shown substantial performance improvements but have the drawback of quadratic memory complexity — making these architectures ill-suited for edge deployment. Recently, State Space Models (SSMs) offer a more favorable balance of computational efficiency and performance in 2D vision, suggesting that they could help here as well. We present TinyBEV, an efficient BEV framework for multi-view 3D perception. For spatial modeling, we replace cross attention with SSMs that fusing BEV and camera images with linear complexity. For temporal modeling, we adopt a lightweight, linear-complexity history-fusion scheme that uses explicit time conditioning and channel-level aggregation instead of cross-frame attention. Both fusion strategies follow small constant scaling with respect to history length and enabling edge-friendly deployment. Experiments on NuScenes datasets demonstrate that TinyBEV is comparable with other state-of-the-art methods across diverse visual perception metrics with advantages in computational efficiency.
Hongyu Ke, Jack Morris, Yongkang Liu 0005, Satoshi Kitai, Kentaro Oguchi 0001, Yi Ding 0041, Haoxin Wang 0003
SEC3
2025 Evaluating the Impacts of Connected Eco-Driving Technologies on Electrified Traffic System at the City Scale: A Simulation Approach
abstract
Eco-approach and Departure (EAD) strategies optimize energy efficiency in Connected and Automated Vehicles (CAVs) by leveraging Signal Phase and Timing (SPaT) data to adjust trajectories. This paper extends the EAD strategy to incorporate actuated signals in a city-scale electrified network and introduces the Eco-Pedaling (EP) strategy to further reduce electricity consumption. First, we analyze off-peak and peak scenarios, applying eco-driving technologies to Connected and Automated Electric Vehicles (CAEVs) in the city network, assuming all vehicles are electric. Results show that CAEVs with eco-driving technologies achieve a 32.24% reduction in electricity consumption per mile during off-peak hours and 31.50% during peak hours, with a 15% CAEV penetration rate, corresponding to Toyota's market share in North America. Second, a 24-hour evaluation with the same CAEV penetration rate highlights a 6.02% reduction in total electricity consumption across the network, at the cost of a 3.58% decrease in network efficiency. Finally, analysis of varying CAEV penetration rates reveals that the system-level electricity consumption decreases as the proportion of CAEVs increases.
Guoyuan Wu 0001, Peng Hao 0001, Kanok Boriboonsomsin, Matthew J. Barth, Yongkang Liu 0005, Yashar Zeiynali Farid
IV6
2025 Poster: Connected Vehicle Surveillance
abstract
Modern cars monitor their surroundings and record video to deter intruders, but current surveillance systems operate independently without communication. Connected vehicles can share detected features of suspicious individuals, improving tracking and alerting approaching drivers before they park in vulnerable spots. This paper explores this use case, where connected vehicles exchange features over the network upon detecting suspicious individuals. Real-world data analysis shows that connected vehicle surveillance improves detection accuracy by 38%.
Can Cui 0009, Seyhan Ucar, Yongkang Liu 0005, Ahmadreza Moradipari, Akin Sisbot, Kentaro Oguchi 0001
MobiHoc3
2024 End-to-End Trip Energy Predictions Developed from Real-World Driving Data using Recurrent Neural Networks
abstract
Battery electric vehicle (BEV) routing in a navigation system has become an essential tool for optimizing long road trips by minimizing the number of charging sessions en route to destinations. BEV routing is designed to provide accurate and realistic estimates of the electric driving range, considering the current battery state of charge (SOC) as well as current and future weather and traffic conditions. This is an improvement over the range estimates shown on a vehicle’s dashboard, typically calculated based on the current SOC and historical driving efficiencies stored in an electronic control unit (ECU). BEV routing enhances existing navigation by incorporating estimates of the required battery energy for the shortest or fastest routes to a destination. It uses predefined vehicle energy models that account for traffic, road, and weather attributes, such as speed limits, traffic speed, elevation gain, and ambient temperature, along the selected routes. Additionally, the vehicle energy models can be integrated into route optimizations to provide a comprehensive BEV routing solution, suggesting routes based on the driver’s preferences, such as energy-efficient or time-efficient routes. This paper addresses the shortcomings of conventional BEV routing, where the vehicle energy models do not accurately represent real-world energy consumption due to varying road, weather, and air conditioning (AC) conditions. The authors first utilize a state-of-the-art distributed data processing platform to extract, transform, and load petabytes worth of raw time-series data into position-based data suitable for model development. The authors then propose a novel method for training end-to-end multi-input vehicle energy models directly from this position-based data, along with efficient calculation schemes to enable accurate energy predictions for long-distance trips. The proposed method was validated with real-world data and demonstrated superior performance compared to two benchmark BEV routing solutions currently available.
Shinhoon Kim, Vishnu Raghupathy, Yongkang Liu 0005, Keisuke Niimi, Takahiro Mochihara
IEEE Big Data3
2024 Pillar Attention Encoder for Adaptive Cooperative Perception
abstract
Interest in cooperative perception is growing quickly due to its remarkable performance in improving perception capabilities for connected and automated vehicles. This improvement is crucial, especially for automated driving scenarios in which perception performance is one of the main bottlenecks to the development of safety and efficiency. However, current cooperative perception methods typically assume that all collaborating vehicles have enough communication bandwidth to share all features with an identical spatial size, which is impractical for real-world scenarios. In this paper, we propose Adaptive Cooperative Perception, a new cooperative perception framework that is not limited by the aforementioned assumptions, aiming to enable cooperative perception under more realistic and challenging conditions. To support this, a novel feature encoder is proposed and named Pillar Attention Encoder. A pillar attention mechanism is designed to extract the feature data while considering its significance for the perception task. An adaptive feature filter is proposed to adjust the size of the feature data for sharing by considering the importance value of the feature. Experiments are conducted for cooperative object detection from multiple vehicle-based and infrastructure-based LiDAR sensors under various communication conditions. Results demonstrate that our method can successfully handle dynamic communication conditions and improve the mean Average Precision by 10.18% when compared with the state-of-the-art feature encoder.
Zhengwei Bai, Guoyuan Wu 0001, Matthew J. Barth, Hang Qiu 0001, Yongkang Liu 0005, Akin Sisbot, Kentaro Oguchi 0001
IEEE Internet Things J.5
2024 A Survey and Framework of Cooperative Perception: From Heterogeneous Singleton to Hierarchical Cooperation
abstract
Perceiving the environment is one of the most fundamental keys to enabling Cooperative Driving Automation, which is regarded as the revolutionary solution to addressing the safety, mobility, and sustainability issues of contemporary transportation systems. Although an unprecedented evolution is now happening in the area of computer vision for object perception, state-of-the-art perception methods are still struggling with sophisticated real-world traffic environments due to the inevitable physical occlusion and limited receptive field of single-vehicle systems. Based on multiple spatially separated perception nodes, Cooperative Perception (CP) is born to unlock the bottleneck of perception for driving automation. In this paper, we comprehensively review and analyze the research progress on CP, and we propose a unified CP framework. The architectures and taxonomy of CP systems based on different types of sensors are reviewed to show a high-level description of the workflow and different structures for CP systems. The node structure, sensing modality, and fusion schemes are reviewed and analyzed with detailed explanations for CP. A Hierarchical Cooperative Perception (HCP) framework is proposed, followed by a review of existing open-source tools that support CP development. The discussion highlights the current opportunities, open challenges, and anticipated future trends.
Zhengwei Bai, Guoyuan Wu 0001, Matthew J. Barth, Yongkang Liu 0005, Akin Sisbot, Kentaro Oguchi 0001, Zhitong Huang
IEEE Trans. Intell. Transp. Syst.4
2023 Poster: Towards Realistic Federated Learning Evaluations for Connected and Automated Vehicles
abstract
Federated learning (FL) is widely recognized as a valuable approach for Connected and Automated Vehicles (CAVs) because it facilitates collaborative model development across a multitude of vehicles in a decentralized manner. However, numerous studies on FL algorithms only assessed their performance through experiments conducted in simulated client-server configurations (e.g., where both server and clients run on the same machine) or simplified scenarios that do not account for client downtime. In this paper, we aim to conduct more realistic evaluations for CAV applications leveraging FL. We present a preliminary experimental study as well as offer insights into potential future directions.
Yongkang Liu 0005, Chianing Johnny Wang, Kentaro Oguchi 0001
SEC1
2023 Hierarchical Federated Learning with Mean Field Game Device Selection for Connected Vehicle Applications
abstract
In this paper, a client-edge-cloud hierarchical federated learning (FL) model has been developed for connected vehicle applications. Generalized models are aggregated on the cloud server, while customized models trained on local data with similar data distribution are aggregated on the edge server, which mitigates the impact of data heterogeneity. To reduce the communication overhead of FL, clients will periodically update to the edge server and edge servers will periodically update to the cloud server. Moreover, we propose a mean field game-based probabilistic device selection scheme. Jointly considering their contributions and the population diversity, a fraction of devices will be selected to join the FL iteration. Taking driving range estimation as an example of connected vehicle applications in the experiment, we have shown that the proposed FL frameworks can increase the prediction accuracy by 28.9% with 5 times fewer clients’ participation, compared with the vanilla FL.
Hao Gao 0008, Yongkang Liu 0005, Akin Sisbot, Yashar Zeiynali Farid, Kentaro Oguchi 0001, Zhu Han 0001
IV2
2023 Field Experiments: Rear Vehicle Behavior Awareness to Avoid Rear-End Collisions
abstract
Modern cars can monitor rear vehicles and detect unsafe driving before the risk of rear-end collision becomes maximum. In this paper, we focus on this use case. We propose a Rear Vehicle Behavior Awareness (RVBA) system to prevent rear-end collisions. RVBA detects unsafe driving of rear vehicles and alerts the driver with guidance to reduce the risk of rearend collisions. We tested the RVBA in field experiments using a test vehicle. Experimental results show that RVBA can detect unsafe driving of rear vehicles in 4 seconds on average, with 70% accuracy, before the risk of rear-end collision becomes maximum. Index Terms–erratic movement patterns, rear-end collisions, rear vehicle behavior awareness, unsafe driving detection
Seyhan Ucar, Sachin Sharma 0003, Yongkang Liu 0005, Akin Sisbot, Kentaro Oguchi 0001, Richard T. Meyer
SECON3
2023 Cyber Mobility Mirror: A Deep Learning-Based Real-World Object Perception Platform Using Roadside LiDAR
abstract
Object perception plays a fundamental role in Cooperative Driving Automation (CDA) which is regarded as a revolutionary promoter for next-generation transportation systems. However, the vehicle-based perception may suffer from the limited sensing range and occlusion as well as low penetration rates in connectivity. In this paper, we propose Cyber Mobility Mirror (CMM), a next-generation real-world object perception system for 3D object detection, tracking, localization, and reconstruction, to explore the potential of roadside sensors for enabling CDA in the real world. The CMM system consists of six main components: i) the data pre-processor to retrieve and preprocess the raw data; ii) the roadside 3D object detector to generate 3D detection results; iii) the multi-object tracker to identify detected objects; iv) the global locator to generate geo-localization information; v) the mobile-edge-cloud-based communicator to transmit perception information to equipped vehicles, and vi) the onboard advisor to reconstruct and display the real-time traffic conditions. An automatic perception evaluation approach is proposed to support the assessment of data-driven models without human-labeling requirements and a CMM field-operational system is deployed at a real-world intersection to assess the performance of the CMM. Results from field tests demonstrate that our CMM prototype system can achieve 96.99% precision and 83.62% recall for detection and 73.55% ID-recall for tracking. High-fidelity real-time traffic conditions (at the object level) can be geo-localized with a root-mean-square error (RMSE) of$0.69m$and$0.33m$for lateral and longitudinal direction, respectively, and displayed on the GUI of the equipped vehicle with a frequency of$3-4 Hz$.
Zhengwei Bai, Saswat Priyadarshi Nayak, Xuanpeng Zhao, Guoyuan Wu 0001, Matthew J. Barth, Xuewei Qi, Yongkang Liu 0005, Akin Sisbot, Kentaro Oguchi 0001
IEEE Trans. Intell. Transp. Syst.7
2022 Demo: Nearby Aggressive Driving Detection
abstract
Aggressive driving is the leading cause of many fatal crashes. Ego vehicles should detect such dangerous driving behavior on other cars and guide drivers to mitigate collision risk. In this paper, we focus on that use case. We demonstrate a nearby aggressive driving detection system. In nearby aggressive driving detection, the ego vehicle observes the follower vehicle and detects aggressive driving behavior on the follower vehicle. It notifies its driver whenever the follower vehicle exhibits aggressive driving.
Tomohiro Matsuda, Seyhan Ucar, Yongkang Liu 0005, Akin Sisbot, Kentaro Oguchi 0001
SEC3
2022 Infrastructure-Based Object Detection and Tracking for Cooperative Driving Automation: A Survey
abstract
Object detection and tracking play a fundamental role in enabling Cooperative Driving Automation (CDA), which is regarded as the revolutionary solution to addressing safety, mobility, and sustainability issues of contemporary transportation systems. Although current computer vision technologies can provide satisfactory object detection results in occlusion-free scenarios, the perception performance of onboard sensors is inevitably limited by the range and occlusion. Owing to the flexible location and pose for sensor installation, infrastructure-based detection, and tracking systems can enhance the perception capability of connected vehicles; as such, they have quickly become a popular research topic. In this survey paper, we review the research progress for infrastructure-based object detection and tracking systems. Architectures of roadside perception systems based on different types of sensors are reviewed to show a high-level description of the workflows for infrastructure-based perception systems. Roadside sensors and different perception methodologies are reviewed and analyzed with detailed literature to provide a low-level explanation for specific methods followed by Datasets and Simulators to draw an overall landscape of infrastructure-based object detection and tracking methods. We highlight current opportunities, open problems, and anticipated future trends.
Zhengwei Bai, Guoyuan Wu 0001, Xuewei Qi, Yongkang Liu 0005, Kentaro Oguchi 0001, Matthew J. Barth
IV4
2022 Multi-Agent Trajectory Prediction with Graph Attention Isomorphism Neural Network
abstract
Multi-agent trajectory prediction is a challenging task because of the uncertainty of agents’ behaviors, interactions between agents, complex road geometry in urban environments, and imperfect/noisy agent histories. Although accurate prediction results are critical for safe and reliable intelligent driving applications (e.g., decision making, motion planning), some other applications may prefer light-weight and computation-efficient trajectory prediction models to handle dynamically changed environments. In this work, we propose a multi-agent, multi-modal Graph Attention Isomorphism Network (GAIN) based trajectory prediction framework to effectively understand and aggregate long-term interactions across agents. We also take the model complexity and computation efficiency into consideration. Experiments on both pedestrian and vehicle datasets demonstrated the effectiveness of our proposed method.
Yongkang Liu 0005, Xuewei Qi, Akin Sisbot, Kentaro Oguchi 0001
IV1
2022 An enhanced driver's risk perception modeling based on gate recurrent unit network
abstract
Risk perception is one of the most important driving skills for drivers to detect potential traffic accident and make correct risk avoidance behaviors. Accurate model and evaluation of risk perception can effectively identify driver’s perception deficiencies and serve as an important human factor for the design of advance driver assistance systems. Most traditional perception ability assessing methods are based on macroscopic statistical results and lack effective mathematical models, thus making it difficult to evaluate risk perception quantitatively. To this end, this paper first obtains semantic understanding of traffic scenes through a semantic segmentation method based on deep learning. Then, the semantic understanding information is fused with the driver’s risk perception results to form the time series data reflecting the risk perception ability. Finally, we use gate recurrent unit (GRU) network-based learning framework to learn the time series data features and thus obtain the behavioral model which responds to the driver’s risk perception ability. By verifying the classification performance on multiple drivers, the risk perception model based on traffic scene understanding can accurately predict the actual cognitive behavior of individual driver. Besides, the proposed GRU-based method outperforms the traditional machine learning algorithm in modeling the driver’s risk perception behavior.
Peng Ping, Weiping Ding 0001, Yongkang Liu 0005, Kazuya Takeda
IV3
2022 Spatiotemporal Transformer Attention Network for 3D Voxel Level Joint Segmentation and Motion Prediction in Point Cloud
abstract
Environment perception including detection, classification, tracking, and motion prediction are key enablers for automated driving systems and intelligent transportation applications. Fueled by the advances in sensing technologies and machine learning techniques, LiDAR-based sensing systems have become a promising solution. The current challenges of this solution are how to effectively combine different perception tasks into a single backbone and how to efficiently learn the spatiotemporal features directly from point cloud sequences. In this research, we propose a novel spatiotemporal attention network based on a transformer self-attention mechanism for joint semantic segmentation and motion prediction within a point cloud at the voxel level. The network is trained to simultaneously outputs the voxel level class and predicted motion by learning directly from a sequence of point cloud datasets. The proposed backbone includes both a temporal attention module (TAM) and a spatial attention module (SAM) to learn and extract the complex spatiotemporal features. This approach has been evaluated with the nuScenes dataset, and promising performance has been achieved.
Zhensong Wei, Xuewei Qi, Zhengwei Bai, Guoyuan Wu 0001, Saswat Priyadarshi Nayak, Peng Hao 0001, Matthew J. Barth, Yongkang Liu 0005, Kentaro Oguchi 0001
IV8
2022 Aggressive Driving Detection on Other Vehicles
abstract
Aggressive driving became a public safety crisis in the USA. Aggressive drivers tailgate and weave among lanes, which may cause risky events ending up in collisions. Vehicles should be aware of such nearby aggressive driving and notify drivers to keep them away from aggressive ones. In this paper, we focus on that use case. The ego vehicle observes the movements of other nearby cars and detects aggressive driving. We tested the feasibility of the proposed approach through a simulation. Simulation results show that the ego vehicle could identify aggressive driving on other cars with about 94% accuracy.
Tomohiro Matsuda, Seyhan Ucar, Yongkang Liu 0005, Akin Sisbot, Kentaro Oguchi 0001
VTC Fall3
2020 Sensor Fusion of Camera and Cloud Digital Twin Information for Intelligent Vehicles
abstract
With the rapid development of intelligent vehicles and Advanced Driving Assistance Systems (ADAS), a mixed level of human driver engagements is involved in the transportation system. Visual guidance for drivers is essential under this situation to prevent potential risks. To advance the development of visual guidance systems, we introduce a novel sensor fusion methodology, integrating camera image and Digital Twin knowledge from the cloud. Target vehicle bounding box is drawn and matched by combining results of object detector running on ego vehicle and position information from the cloud. The best matching result, with a 79.2% accuracy under 0.7 Intersection over Union (IoU) threshold, is obtained with depth image served as an additional feature source. Game engine-based simulation results also reveal that the visual guidance system could improve driving safety significantly cooperate with the cloud Digital Twin system.
Yongkang Liu 0005, Ziran Wang, Kyungtae Han, Zhenyu Shou, Prashant Tiwari, John H. L. Hansen
IV1
2020 Long-Term Prediction of Lane Change Maneuver Through a Multilayer Perceptron
abstract
Behavior prediction plays an essential role in both autonomous driving systems and Advanced Driver Assistance Systems (ADAS), since it enhances vehicle's awareness of the imminent hazards in the surrounding environment. Many existing lane change prediction models take as input lateral or angle information and make short-term (<; 5 seconds) maneuver predictions. In this study, we propose a longer-term (5~10 seconds) prediction model without any lateral or angle information. Three prediction models are introduced, including a logistic regression model, a multilayer perceptron (MLP) model, and a recurrent neural network (RNN) model, and their performances are compared by using the real-world NGSIM dataset. To properly label the trajectory data, this study proposes a new time-window labeling scheme by adding a time gap between positive and negative samples. Two approaches are also proposed to address the unstable prediction issue, where the aggressive approach propagates each positive prediction for certain seconds, while the conservative approach adopts a roll-window average to smooth the prediction. Evaluation results show that the developed prediction model is able to capture 75% of real lane change maneuvers with an average advanced prediction time of 8.05 seconds.
Zhenyu Shou, Ziran Wang, Kyungtae Han, Yongkang Liu 0005, Prashant Tiwari, Xuan Di
IV4
2017 Navigation-orientated natural spoken language understanding for intelligent vehicle dialogue
abstract
Voice-based human-machine interfaces are becoming a key feature for next generation intelligent vehicles. For the navigation dialogue systems, it is desired to understand a driver's spoken language in a natural way. This study proposes a two-stage framework, which first converts the audio streams into text sentences through Automatic Speech Recognition (ASR), followed by Natural Language Processing (NLP) to retrieve the navigation-associated information. The NLP stage is based on a Deep Neural Network (DNN) framework, which contains sentence-level sentiment analysis and word/phrase-level context extraction. Experiments are conducted using the CU-Move in-vehicle speech corpus. Results indicate that the DNN architecture is effective for navigation dialog language understanding, whereas the NLP performances are affected by ASR errors. Overall, it is expected that the proposed RNN-based NLP approach, with the corresponding reduced vocabulary designed for navigation-oriented tasks, will benefit the development of advanced intelligent vehicle human-machine interfaces.
Yongkang Liu 0005, John H. L. Hansen
Intelligent Vehicles Symposium2