Amr Abdelraouf

dblp:326/4538 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
11since 2021 · last 2026
0000-0001-9068-6664ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Using Intent Communication to Enhance Platooning: Validation with Prototype Vehicles
Ahmadreza Moradipari, Sergei S. Avedisov, Mariam Nour, Shatadal Mishra, Kyungtae Han, Amr Abdelraouf, Takayuki Shimizu, Onur Altintas
INFOCOM7
2026 LLM4AD: Large Language Models for Autonomous Driving - Concept, Review, Benchmark, Experiments, and Future Trends
abstract
With the broader adoption and highly successful development of large language models (LLMs), there has been growing interest and demand for applying LLMs to autonomous driving technology. Driven by their natural language (NL) understanding and reasoning capabilities, LLMs have the potential to enhance various aspects of autonomous driving systems, from perception and scene understanding to interactive decision-making. This article first introduces the novel concept of designing LLMs for autonomous driving (LLM4AD), followed by a review of existing LLM4AD studies. Then, a comprehensive benchmark is proposed for evaluating the instruction-following and reasoning abilities of LLM4AD systems, which includes LaMPilot-Bench, CARLA Leaderboard 1.0 Benchmark in simulation and NuPlanQA for multiview visual question answering (VQA). Furthermore, extensive real-world experiments are conducted on autonomous vehicle platforms, examining both on-cloud and on-edge LLM deployment for personalized decision-making and motion control. Next, the future trends of integrating language diffusion models into autonomous driving are explored, exemplified by the proposed vision-language diffusion (ViLaD) framework. Finally, the main challenges of LLM4AD are discussed, including latency, deployment, security and privacy, safety, trust and transparency, and personalization.
Can Cui 0009, Yunsheng Ma, Sungyeon Park 0001, Zichong Yang, Yupeng Zhou, Peiran Liu 0003, Juanwu Lu, Juntong Peng, Jiaru Zhang, Ruqi Zhang, Lingxi Li 0001, Yaobin Chen, Jitesh H. Panchal, Amr Abdelraouf, Kyungtae Han, Ziran Wang
Proc. IEEE14
2025 Video Token Sparsification for Efficient Multimodal LLMs in Driving Visual Question Answering
abstract
Multimodal large language models (MLLMs) have shown significant potential in enhancing driving scene understanding and visual question answering (VQA) through advanced logical reasoning capabilities. These tasks support driving action generation and explanation, especially in end-to-end autonomous driving applications. However, deploying these models poses a significant challenge due to their substantial parameter sizes and computational demands, which often exceed onboard computational limits. A key limitation stems from the large number of visual tokens needed to capture detailed, long-context visual information, resulting in increased latency and memory use. To address this, we propose Video Token Sparsification (VTS), a novel approach that leverages redundancy in consecutive video frames to reduce visual tokens while preserving critical information. VTS employs a lightweight CNN-based model to identify key frames and prune less informative tokens, mitigating hallucinations and boosting inference throughput without performance loss. Comprehensive experiments on the LingoQA and DRAMA benchmarks show that VTS achieves up to a 33% improvement in inference throughput and a 28% reduction in memory usage compared to baselines, maintaining comparable performance.
Yunsheng Ma, Amr Abdelraouf, Ahmadreza Moradipari, Ziran Wang, Kyungtae Han
IV2
2025 Improved Intent Sharing for Energy-Efficient Vehicle Platooning
abstract
We explore the advantages of using deep learning-based intent sharing for platooning of connected automated vehicles (CAVs). Unlike traditional platooning algorithms that rely on status-sharing - exchanging current position, speed, and acceleration-our approach focuses on intent-sharing, where CAVs share predicted future trajectories with other CAVs. We introduce a deep learning model to generate the intent for each CAV and integrate it into a receding horizon control framework. Our approach aims to minimize spacing errors with the leading vehicle while improving energy efficiency and maintaining string stability. Through microscopic simulations using real-world highway data, we demonstrate that our intent messages significantly enhance energy efficiency compared to conventional status and intent-based platooning algorithms. Moreover, we show that this improvement is particularly pronounced when reducing the frequency of intent message transmission.
Ahmadreza Moradipari, Amr Abdelraouf, Sergei S. Avedisov
IV2
2025 PDB-Eval: An Evaluation of Large Multimodal Models for Description and Explanation of Personalized Driving Behavior
abstract
Understanding a driver's behavior and intentions is important for potential risk assessment and early accident prevention. Safety and driver assistance systems can be tailored to individual drivers' behavior, significantly enhancing their effectiveness. However, existing datasets are limited in describing and explaining general vehicle movements based on external visual evidence. This paper introduces a benchmark, PDB-Eval, for a detailed understanding of Personalized Driver Behavior, and aligning Large Multimodal Models (MLLMs) with driving comprehension and reasoning. Our benchmark consists of two main components, PDB-X and PDBQA. PDB-X can evaluate MLLMs' understanding of temporal driving scenes. Our dataset is designed to find valid visual evidence from the external view to explain the driver's behavior from the internal view. To align MLLMs' reasoning abilities with driving tasks, we propose PDB-QA as a visual explanation question-answering task for MLLM instruction fine-tuning. As a generic learning task for generative models like MLLMs, PDB-QA can bridge the domain gap without harming MLLMs' generalizability. Our evaluation indicates that fine-tuning MLLMs on fine-grained descriptions and explanations can effectively bridge the gap between MLLMs and the driving domain, which improves zero-shot performance on question-answering tasks by up to 73.2%. We further evaluate the MLLMs fine-tuned on PDB-X in Brain4Cars' intention prediction and AIDE's recognition tasks. We observe up to 12.5% performance improvements on the turn intention prediction task in Brain4Cars, and consistent performance improvements up to 11.0% on all tasks in AIDE.
Junda Wu, Jessica Maria Echterhoff, Kyungtae Han, Amr Abdelraouf, Julian J. McAuley
IV4
2024 LaMPilot: An Open Benchmark Dataset for Autonomous Driving with Language Model Programs
abstract
Autonomous driving (AD) has made significant strides in recent years. However, existing frameworks struggle to interpret and execute spontaneous user instructions, such as "overtake the car ahead.” Large Language Models (LLMs) have demonstrated impressive reasoning capabilities showing potential to bridge this gap. In this paper, we present LaMPilot, a novel framework that integrates LLMs into AD systems, enabling them to follow user instructions by generating code that leverages established functional primitives. We also introduce LaMPilot-Bench, the first bench-mark dataset specifically designed to quantitatively evaluate the efficacy of language model programs in AD. Adopting the LaMPilot framework, we conduct extensive experiments to assess the performance of off-the-shelf LLMs on LaMPilot-Bench. Our results demonstrate the potential of LLMs in handling diverse driving scenarios and following user instructions in driving. To facilitate further research in this area, we release our code and data at GitHub.com/PurdueDigitalTwin/LaMPilot.
Yunsheng Ma, Can Cui 0009, Wenqian Ye, Peiran Liu 0003, Juanwu Lu, Amr Abdelraouf, Kyungtae Han, Aniket Bera, James M. Rehg, Ziran Wang
CVPR7
2024 Driving through the Concept Gridlock: Unraveling Explainability Bottlenecks in Automated Driving
abstract
Concept bottleneck models have been successfully used for explainable machine learning by encoding information within the model with a set of human-defined concepts. In the context of human-assisted or autonomous driving, explainability models can help user acceptance and understanding of decisions made by the autonomous vehicle, which can be used to rationalize and explain driver or vehicle behavior. We propose a new approach using concept bottlenecks as visual features for control command predictions and explanations of user and vehicle behavior. We learn a human-understandable concept layer that we use to explain sequential driving scenes while learning vehicle control commands. This approach can then be used to determine whether a change in a preferred gap or steering commands from a human (or autonomous vehicle) is led by an external stimulus or change in preferences. We achieve competitive performance to latent visual features while gaining interpretability within our model setup.1
Jessica Maria Echterhoff, An Yan 0003, Kyungtae Han, Amr Abdelraouf, Julian J. McAuley
WACV4
2023 Real-Time Learning of Driving Gap Preference for Personalized Adaptive Cruise Control
abstract
Advanced Driver Assistance Systems (ADAS) are increasingly important in improving driving safety and comfort, with Adaptive Cruise Control (ACC) being one of the most widely used. However, pre-defined ACC settings may not always align with driver's preferences and habits, leading to discomfort and potential safety issues. Personalized ACC (P-ACC) has been proposed to address this problem, but most existing research uses historical driving data to imitate behaviors that conform to driver preferences, neglecting real-time driver feedback. To bridge this gap, we propose a cloud-vehicle collaborative P-ACC framework that incorporates driver feedback adaptation in real time. The framework is divided into offline and online parts. The offline component records the driver's naturalistic car-following trajectory and uses inverse reinforcement learning (IRL) to train the model on the cloud. In the online component, driver feedback is used to update the driving gap preference in real time. The model is then retrained on the cloud with driver's takeover trajectories, achieving incremental learning to better match driver's preference. Human-in-the-loop (HuiL) simulation experiments demonstrate that our proposed method significantly reduces driver intervention in automatic control systems by up to 62.8%. By incorporating real-time driver feedback, our approach enhances the comfort and safety of P-ACC, providing a personalized and adaptable driving experience.
Zhouqiao Zhao, Xishun Liao, Amr Abdelraouf, Kyungtae Han, Matthew J. Barth, Guoyuan Wu 0001
SMC3
2023 Sequence-to-Sequence Recurrent Graph Convolutional Networks for Traffic Estimation and Prediction Using Connected Probe Vehicle Data
abstract
Traffic estimation is imperative for conducting fundamental transportation engineering tasks such as transportation planning and traffic safety studies. Additionally, traffic prediction is vital for many data-driven intelligent transportation system applications. Most traffic estimation and prediction methods rely on infrastructure-based sensors to collect traffic parameters. However, infrastructure-based data collection can be costly and time consuming to set up and maintain. Additionally, the spatial distribution of the collected traffic data is limited by the location of the deployed hardware sensors. Probe vehicle data can be used to overcome these limitations. Traffic modeling for estimation and prediction is a complex task due to the stochastic nonlinear spatiotemporal dependencies exhibited by traffic parameters. In this paper, a deep learning-based sequence-to-sequence architecture called Seq2seq GCN-LSTM was proposed to estimate and predict network-wide traffic volume and speed. The proposed method utilizes short-term historical traffic data collected from a low-penetration rate probe vehicle fleet to estimate and predict traffic parameters up to 60-minutes ahead. The method utilizes Graph Convolutional Networks for spatial dependency extraction and Long Short-Term Memory networks to model temporal dependencies. The proposed method generated superior traffic results compared to the baseline models. Additionally, the model demonstrated robustness against perturbations caused by the low rate. Furthermore, the probe vehicle penetration rate was varied to test its effect on the proposed method’s modeling capability. The model was able to maintain traffic volume and speed estimation and prediction performance within a reasonable margin of error using penetration rates as low as 1.5% and 0.5%, respectively.
Amr Abdelraouf, Mohamed A. Abdel-Aty, Nada Mahmoud
IEEE Trans. Intell. Transp. Syst.1
2022 Using Vision Transformers for Spatial-Context-Aware Rain and Road Surface Condition Detection on Freeways
abstract
Inclement weather conditions, particularly heavy rain and the consequent wet road surface, have an unfavorable effect on driving conditions, traffic infrastructure, and operational plans. To mitigate the potentially detrimental ramifications of turbulent weather, it must be continuously monitored in real time and with high geospatial granularity. Traditionally, road weather conditions are monitored using weather forecasts or Roadside Weather Information Systems (RWIS). However, these methods are either ill-equipped or too expensive to provide the required fine-grained observations. Alternatively, roadside traffic CCTV cameras are ubiquitously deployed on US freeways and can serve as inexpensive sensors to surveil weather. In this paper, a novel vision-based methodology is proposed to detect rain and road surface conditions from roadside traffic cameras. Vision Transformers were utilized for image-based classification. They demonstrated superior results compared to convolution-based approaches. Furthermore, the geographical distribution of roadside cameras was leveraged to add spatial context awareness to the detection model. A Spatial Self-Attention network was proposed to model the relationship between the detection results of adjacent images as a sequence-to-sequence detection task. The results indicate that the addition of the sequential detection module improved the accuracy of the stand-alone Vision Transformer as measured by the F1-score. The boost in performance enhanced the F1-scores of the stand-alone Vision Transformer by 5.61% and 5.97% for the rain and road surface condition detection tasks, respectively, raising the total F1-score to 96.71% and 98.07%.
Amr Abdelraouf, Mohamed A. Abdel-Aty, Yina Wu
IEEE Trans. Intell. Transp. Syst.1
2022 Utilizing Attention-Based Multi-Encoder-Decoder Neural Networks for Freeway Traffic Speed Prediction
abstract
Speed prediction is a crucial yet complicated task for intelligent transportation systems. The challenge derives from the complex spatiotemporal dependencies of traffic parameters. In the past few years, deep neural networks have achieved the best traffic speed prediction performance. However, most models depend on short-term input sequences to predict short/long-term traffic speed (e.g., predicting speed for the next hour using data from the past hour). These models fail to consider the daily and weekly periodic behavior of traffic. Another problem posed by neural networks is the lack of interpretability as they often operate as “black boxes”. In this paper, an attention-based multi-encoder-decoder (Att-MED) model is proposed to predict traffic speed. The model uses convolutional-LSTMs to capture the spatiotemporal relationship of multiple input sequences, namely short-term, daily and weekly traffic patterns. The model also employs an LSTM to model the output predictions sequentially. Furthermore, attention mechanism is used to weigh the contribution of each traffic sequence towards the output predictions. The proposed network architecture, when trained end-to-end, results in a superior prediction accuracy compared to baseline models. In addition to contributing towards performance, the attention mechanism creates weight values, which when visualized, provide insights into the decision-making process of the neural network, and consequently produce explainable outputs. Att-MED’s extracted attention weights highlight the contribution of daily and weekly periodic input towards speed prediction.
Amr Abdelraouf, Mohamed A. Abdel-Aty, Jinghui Yuan
IEEE Trans. Intell. Transp. Syst.1