Xuan Di

dblp:214/9877 · also Xuan Sharon Di · DBLP profile ↗
← Back
21ranked-venue papers
3as first author
18since 2021 · last 2026
0000-0003-2925-7697ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 1 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 3 since 2021Systems, architecture and hardware · 3 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 AI-Powered CPS-Enabled Vulnerable-User-Aware Urban Transportation Digital Twin: Methods and Applications
abstract
We present methods and applications for the development of digital twins (DT) for urban traffic management. While the majority of studies on the DT focus on its “eyes,” which is the emerging sensing and perception like object detection and tracking, what really distinguishes the DT from a traditional simulator lies in its “brain,” the prediction and decision making capabilities of extracting patterns and making informed decisions from what has been seen and perceived. In order to add value to urban transportation management, DTs need to be powered by artificial intelligence and complement with low-latency high-bandwidth sensing and networking technologies, in other words, cyberphysical systems. This paper can be a pointer to help researchers and practitioners identify challenges and opportunities for the development of DTs; a bridge to initiate conversations across disciplines; and a road map to exploiting potentials of DTs for diverse urban transportation applications.
Yongjie Fu, Mehmet Kerem Türkcan, Mahshid Ghasemi, Zhaobin Mo, Chengbo Zang, Abhishek Adhikari, Zoran Kostic, Gil Zussman, Xuan Di
IEEE Trans. Intell. Transp. Syst.9
2025 Stochastic Semi-Gradient Descent for Learning Mean Field Games with Population-Aware Function Approximation
abstract
Mean field games (MFGs) model interactions in large-population multi-agent systems through population distributions. Traditional learning methods for MFGs are based on fixed-point iteration (FPI), where policy updates and induced population distributions are computed separately and sequentially. However, FPI-type methods may suffer from inefficiency and instability due to potential oscillations caused by this forward-backward procedure. In this work, we propose a novel perspective that treats the policy and population as a unified parameter controlling the game dynamics. By applying stochastic parameter approximation to this unified parameter, we develop SemiSGD, a simple stochastic gradient descent (SGD)-type method, where an agent updates its policy and population estimates simultaneously and fully asynchronously. Building on this perspective, we further apply linear function approximation (LFA) to the unified parameter, resulting in the first population-aware LFA (PA-LFA) for learning MFGs on continuous state-action spaces. A comprehensive finite-time convergence analysis is provided for SemiSGD with PA-LFA, including its convergence to the equilibrium for linear MFGs—a class of MFGs with a linear structure concerning the population—under the standard contractivity condition, and to a neighborhood of the equilibrium under a more practical condition. We also characterize the approximation error for non-linear MFGs. We validate our theoretical findings with six experiments on three MFGs.
Chenyu Zhang 0002, Xu Chen 0033, Xuan Di
ICLR3
2025 Real-Time Video Analytics for Urban Safety: Deployment over Edge and End Devices
abstract
This paper introduces PAVE (Pedestrian Awareness Via Edge analytics), a scalable real-time video analytics system that uses street cameras to enhance pedestrian safety while preserving their privacy. PAVE processes live camera streams on an edge server to track pedestrians and vehicles in real-time, predict vehicles' trajectories, and identify danger zones where pedestrians are present. The coordinates of these zones are sent to pedestrians' mobile devices via a custom iOS app, which locally determines if they are at risk without sharing any data with the edge server, hence preserving privacy. Moreover, anonymized metadata, including real-time location and speed/direction of pedestrians and vehicles, are visualized on a public map. PAVE's effectiveness was validated through deployment on the NSF COSMOS testbed, processing live video from cameras in diverse urban environments. Live field tests show that PAVE can alert at-risk pedestrians ~0.9 s before a vehicle reaches them. Through extensive profiling, we show that optimizing memory/compute configuration per pipeline stage can reduce latency by up to 10× compared to the default operating system configurations.
Mahshid Ghasemi, Yongjie Fu, Peiran Wang, Mehmet Kerem Türkcan, Jhonatan Tavori, Sofia Kleisarchaki, Thomas Calmant, Levent Gürgen, Zoran Kostic, Xuan Di, Gil Zussman, Javad Ghaderi
SEC11
2025 Demo: Real-Time Video Analytics for Urban Safety, Deployment over Edge and End Devices
abstract
We showcase the workflow of PAVE (Pedestrian Awareness Via Edge analytics), a scalable system for real-time video analytics that leverages street cameras to improve pedestrians' safety while maintaining their privacy. PAVE distributes computation across edge servers and end-user mobile devices. Cameras' live streams are processed at the edge to forecast vehicles' trajectories and detect danger zones. Pedestrians' mobile devices then locally determine if the user is inside a danger zone and trigger timely alerts via a custom iOS app. In addition, anonymized metadata, such as pedestrian and vehicle positions, speeds, and directions, are aggregated and displayed on a public map for broader situational awareness. We evaluated PAVE's performance through implementation on the NSF COSMOS testbed's edge server while processing real-time video stream from cameras in diverse urban environments. Live field tests at an intersection in New York City show that PAVE can alert at-risk pedestrians about 0.9 s before a vehicle reaches them. With low-latency cameras, this lead time extends to around 1.6 s which is within the 1–2 s window pedestrians typically need to react.
Mahshid Ghasemi, Yongjie Fu, Peiran Wang, Mehmet Kerem Türkcan, Jhonatan Tavori, Sofia Kleisarchaki, Thomas Calmant, Levent Gürgen, Zoran Kostic, Xuan Di, Gil Zussman, Javad Ghaderi
SEC11
2025 Mean field games for urban mobility: a review
Xuan Di, Zhenhui Xu, Tielong Shen
Sci. China Inf. Sci.1
2024 A Single Online Agent Can Efficiently Learn Mean Field Games
abstract
Mean field games (MFGs) are a promising framework for modeling the behavior of large-population systems. However, solving MFGs can be challenging due to the coupling of forward population evolution and backward agent dynamics. Typically, obtaining mean field Nash equilibria (MFNE) involves an iterative approach where the forward and backward processes are solved alternately, known as fixed-point iteration (FPI). This method requires fully observed population propagation and agent dynamics over the entire spatial domain, which could be impractical in some real-world scenarios. To overcome this limitation, this paper introduces a novel online single-agent model-free learning scheme, which enables a single agent to learn MFNE using online samples, without prior knowledge of the state-action space, reward function, or transition dynamics. Specifically, the agent updates its policy through the value function (Q), while simultaneously evaluating the mean field state (M), using the same batch of observations. We develop two variants of this learning scheme: off-policy and on-policy QM iteration. We prove that they efficiently approximate FPI, and a sample complexity guarantee is provided. The efficacy of our methods is confirmed by numerical experiments.
Chenyu Zhang 0002, Xu Chen 0033, Xuan Di
ECAI3
2024 SLAMuZero: Plan and Learn to Map for Joint SLAM and Navigation
abstract
MuZero has demonstrated remarkable performance in board and video games where Monte Carlo tree search (MCTS) method is utilized to learn and adapt to different game environments. This paper leverages the strength of MuZero to enhance agents’ planning capability for joint active simultaneous localization and mapping (SLAM) and navigation tasks, which require an agent to navigate an unknown environment while simultaneously constructing a map and localizing itself. We propose SLAMuZero, a novel approach for joint SLAM and navigation, which employs a search process that uses an explicit encoder-decoder architecture for mapping, followed by a prediction function to evaluate policy and value based on the generated map. SLAMuZero outperforms the state-of-the-art baseline and significantly reduces training time, underscoring the efficiency of our approach. Additionally, we develop a new open source library for implementing SLAMuZero, which is a flexible and modular toolkit for researchers and practitioners (https://github.com/bwfbowen/SLAMuZero).
Xu Chen 0033, Zhengkun Pan, Xuan Di
ICAPS4
2024 Graphon Mean Field Games with a Representative Player: Analysis and Learning Algorithm
abstract
We propose a discrete time graphon game formulation on continuous state and action spaces using a representative player to study stochastic games with heterogeneous interaction among agents. This formulation admits both conceptual and mathematical advantages, compared to a widely adopted formulation using a continuum of players. We prove the existence and uniqueness of the graphon equilibrium with mild assumptions, and show that this equilibrium can be used to construct an approximate solution for the finite player game, which is challenging to analyze and solve due to curse of dimensionality. An online oracle-free learning algorithm is developed to solve the equilibrium numerically, and sample complexity analysis is provided for its convergence.
Fuzhong Zhou, Chenyu Zhang 0002, Xu Chen 0033, Xuan Di
ICML4
2024 Digital Twin for Pedestrian Safety Warning at a Single Urban Traffic Intersection
abstract
Ensuring the safety of Vulnerable Road Users (VRUs) at intersections is crucial to enhancing urban traffic systems. This paper introduces a novel intelligent warning system specifically designed to increase the safety of VRUs crossing intersections. The proposed system leverages the COSMOS testbed to obtain real time vehicle information and employs Message Queuing Telemetry Transport (MQTT) as a standards-based messaging protocol for device communication and data transmission and utilizes a transformer model and Time To Collision (TTC) method to predict the collision. To validate the effectiveness and reliability of our intelligent alert system, we conducted comprehensive tests using the CARLA simulator, incorporating hardware in the loop simulation approach. The results demonstrate the potential for increased situational awareness and reduced risk factors associated with VRUs at intersections. Our work supports the integration of this intelligent alert system as a viable solution for reducing accidents and enhancing the overall safety of urban intersections in real time.
Yongjie Fu, Mehmet Kerem Türkcan, Vikram Anantha, Zoran Kostic, Gil Zussman, Xuan Di
IV6
2024 Causal Imitation for Markov Decision Processes: a Partial Identification Approach
abstract
Imitation learning enables an agent to learn from expert demonstrations when the performance measure is unknown and the reward signal is not specified. Standard imitation methods do not generally apply when the learner and the expert's sensory capabilities mismatch and demonstrations are contaminated with unobserved confounding bias. To address these challenges, recent advancements in causal imitation learning have been pursued. However, these methods often require access to underlying causal structures that might not always be available, posing practical challenges. In this paper, we investigate robust imitation learning within the framework of canonical Markov Decision Processes (MDPs) using partial identification, allowing the agent to achieve expert performance even when the system dynamics are not uniquely determined from the confounded expert demonstrations. Specifically, first, we theoretically demonstrate that when unobserved confounders (UCs) exist in an MDP, the learner is generally unable to imitate expert performance. We then explore imitation learning in partially identifiable settings --- either transition distribution or reward function is non-identifiable from the available data and knowledge. Augmenting the celebrated GAIL method (Ho \& Ermon, 2016), our analysis leads to two novel causal imitation algorithms that can obtain effective policies guaranteed to achieve expert performance.
Kangrui Ruan, Junzhe Zhang 0001, Xuan Di, Elias Bareinboim
NeurIPS3
2023 Learning Dual Mean Field Games on Graphs
abstract
Reinforcement learning (RL) has been developed for mean field games over graphs (G-MFG) in social media and network economics, in which the transition of agents between a node pair incurs an instantaneous reward. However, agents’ en-route choices on edges are largely neglected that incur an experienced reward depending on agents’ actions and population evolution along edges. Here we focus on a broader class of MFGs, named “dual MFG on graphs” (G-dMFG), which models two interacting MFGs, namely, one on edges and one at nodes over a graph. In this setting, agents select travel speed along edges and next-go-to edge at nodes for a minimum cumulative cost, which arises from the congestion effect when many agents compete for the same resource. This has various implications for autonomous driving navigation, spatial resource allocation, and internet packet routing. We establish formally that G-dMFG is a generic G-MFG, encompassing a more complex cost structure (that is nonseparable between states and actions) and with no need to pre-specify a termination time horizon. RL algorithms are designed to solve mean field equilibria (MFE) on large networks.
Xu Chen 0033, Shuo Liu 0018, Xuan Di
ECAI3
2023 Causal Imitation Learning via Inverse Reinforcement Learning
Kangrui Ruan, Junzhe Zhang 0001, Xuan Di, Elias Bareinboim
ICLR3
2023 Detecting mild cognitive impairment and dementia in older adults using naturalistic driving data and interaction-based classification from influence score
Xuan Di, Yiqiao Yin, Yongjie Fu, Zhaobin Mo, Shaw-Hwa Lo, Carolyn DiGuiseppi, David W. Eby, Linda L. Hill, Thelma J. Mielenz, David Strogatz, Guohua Li
Artif. Intell. Medicine1
2022 Learning Human Driving Behaviors with Sequential Causal Imitation Learning
abstract
Learning human driving behaviors is an efficient approach for self-driving vehicles. Traditional Imitation Learning (IL) methods assume that the expert demonstrations follow Markov Decision Processes (MDPs). However, in reality, this assumption does not always hold true. Spurious correlation may exist through the paths of historical variables because of the existence of unobserved confounders. Accounting for the latent causal relationships from unobserved variables to outcomes, this paper proposes Sequential Causal Imitation Learning (SeqCIL) for imitating driver behaviors. We develop a sequential causal template that generalizes the default MDP settings to one with Unobserved Confounders (MDPUC-HD). Then we develop a sufficient graphical criterion to determine when ignoring causality leads to poor performances in MDPUC-HD. Through the framework of Adversarial Imitation Learning, we develop a procedure to imitate the expert policy by blocking π-backdoor paths at each time step. Our methods are evaluated on a synthetic dataset and a real-world highway driving dataset, both demonstrating that the proposed procedure significantly outperforms non-causal imitation learning methods.
Kangrui Ruan, Xuan Di
AAAI2
2022 Social Learning In Markov Games: Empowering Autonomous Driving
abstract
In a multi-agent system (MAS), a social learning scheme allows independent agents to learn through interactions with agents randomly selected from a pool. Such a scheme is important for autonomous vehicles (AV) to navigate complex traffic environments consisting of many road users. In this paper, we apply the social learning scheme to Markov games and leverage deep reinforcement learning (DRL) to investigate how individual AVs learn policies and form social norms in traffic scenarios. To capture agents’ different attitudes toward traffic environments, a heterogeneous agent pool with cooperative and defective AVs is introduced to the social learning scheme. To solve social norms formed by AVs, we propose a DRL algorithm, and apply them to traffic scenarios: unsignalized intersection and highway platoon. We find that compared to defective AVs, cooperative AVs can easily conform to expected social norms. In addition, cooperative AVs would lead to lower crash rates. We also find that prioritized roads/lanes can make AVs conform to expected social norms.
Xu Chen 0033, Zechu Li, Xuan Di
IV3
2022 TrafficFlowGAN: Physics-Informed Flow Based Generative Adversarial Network for Uncertainty Quantification
Zhaobin Mo, Yongjie Fu, Daran Xu, Xuan Di
ECML/PKDD (3)4
2022 A Physics-Informed Deep Learning Paradigm for Traffic State and Fundamental Diagram Estimation
abstract
Traffic state estimation (TSE) bifurcates into two main categories, model-driven and data-driven (e.g., machine learning, ML) approaches, while each suffers from either deficient physics or small data. To mitigate these limitations, recent studies introduced hybrid methods, such as physics-informed deep learning (PIDL), which contains both model-driven and data-driven components. This paper contributes an improved paradigm, called physics-informed deep learning with a fundamental diagram learner (PIDL + FDL), which integrates ML terms into the model-driven component to learn a functional form of a fundamental diagram (FD), i.e., a mapping from traffic density to flow or velocity. The proposed PIDL + FDL has the advantages of performing the TSE learning, model parameter identification, and FD estimation simultaneously. This paper focuses on highway TSE with observed data from loop detectors, using traffic density or velocity as traffic variables. We demonstrate the use of PIDL + FDL to solve popular first-order and second-order traffic flow models and reconstruct the FD relation as well as model parameters that are outside the FD term. We then evaluate the PIDL + FDL-based TSE using the Next Generation SIMulation (NGSIM) dataset. The experimental results show the superiority of the PIDL + FDL in terms of improved estimation accuracy and data efficiency over advanced baseline TSE methods, and additionally, the capacity to properly learn the unknown underlying FD relation.
Rongye Shi, Zhaobin Mo, Kuang Huang, Xuan Di, Qiang Du 0001
IEEE Trans. Intell. Transp. Syst.4
2021 Physics-Informed Deep Learning for Traffic State Estimation: A Hybrid Paradigm Informed By Second-Order Traffic Models
abstract
Traffic state estimation (TSE) reconstructs the traffic variables (e.g., density or average velocity) on road segments using partially observed data, which is important for traffic managements. Traditional TSE approaches mainly bifurcate into two categories: model-driven and data-driven, and each of them has shortcomings. To mitigate these limitations, hybrid TSE methods, which combine both model-driven and data-driven, are becoming a promising solution. This paper introduces a hybrid framework, physics-informed deep learning (PIDL), to combine second-order traffic flow models and neural networks to solve the TSE problem. PIDL can encode traffic flow models into deep neural networks to regularize the learning process to achieve improved data efficiency and estimation accuracy. We focus on highway TSE with observed data from loop detectors and probe vehicles, using both density and average velocity as the traffic variables. With numerical examples, we show the use of PIDL to solve a popular second-order traffic flow model, i.e., a Greenshields-based Aw-Rascle-Zhang (ARZ) model, and discover the model parameters. We then evaluate the PIDL-based TSE method using the Next Generation SIMulation (NGSIM) dataset. Experimental results demonstrate the proposed PIDL-based approach to outperform advanced baseline methods in terms of data efficiency and estimation accuracy.
Rongye Shi, Zhaobin Mo, Xuan Di
AAAI3
2020 Long-Term Prediction of Lane Change Maneuver Through a Multilayer Perceptron
abstract
Behavior prediction plays an essential role in both autonomous driving systems and Advanced Driver Assistance Systems (ADAS), since it enhances vehicle's awareness of the imminent hazards in the surrounding environment. Many existing lane change prediction models take as input lateral or angle information and make short-term (<; 5 seconds) maneuver predictions. In this study, we propose a longer-term (5~10 seconds) prediction model without any lateral or angle information. Three prediction models are introduced, including a logistic regression model, a multilayer perceptron (MLP) model, and a recurrent neural network (RNN) model, and their performances are compared by using the real-world NGSIM dataset. To properly label the trajectory data, this study proposes a new time-window labeling scheme by adding a time gap between positive and negative samples. Two approaches are also proposed to address the unstable prediction issue, where the aggressive approach propagates each positive prediction for certain seconds, while the conservative approach adopts a roll-window average to smooth the prediction. Evaluation results show that the developed prediction model is able to capture 75% of real lane change maneuvers with an average advanced prediction time of 8.05 seconds.
Zhenyu Shou, Ziran Wang, Kyungtae Han, Yongkang Liu 0005, Prashant Tiwari, Xuan Di
IV6
2019 A Similitude Theory for Modeling Traffic Flow Dynamics
abstract
Similitude theory, particularly dimension analysis, is a common tool for testing scaled-down engineering models and is widely used in vehicle dynamics and other engineering fields. However, it is barely employed in scaling traffic flow dynamics. In this paper, dimension analysis is adopted to scale car-following dynamics. Under the guidance of the similitude theory, a scaled downvehicle test bed is built where seven cars are running on a circular track at a maximum speed initially and congestion emerges after a period of time. In other words, a phantom traffic jam appears in our similitude test bed without any bottlenecks. The fundamental diagrams drawn from the experimental results show that our test bed has the capability of generating traffic hysteresis that is commonly observed in the field. Therefore, the design of this test bed can be used to simulate traffic dynamics to a certain degree and will pave the way for scaled-down connected and automated vehicle systems development.
Xuan Di, Yan Zhao 0011, Shihong Ed Huang, Henry X. Liu
IEEE Trans. Intell. Transp. Syst.1
2018 Large-scale short-term urban taxi demand forecasting using deep learning
abstract
The world has seen in recent years great successes in applying deep learning (DL) for many application domains. Though powerful, DL is not easy to be used well. In this invited paper, we study an urban taxi demand forecast problem using DL, and we show a number of key insights in modeling a domain problem as a suitable DL task. We also conduct a systematic comparison of two recent deep neural networks (DNNs) for taxi demand prediction, i.s., the ST-ResNet and FLC-Net, on New York city taxi record dataset. Our experimental results show DNNs indeed outperform most traditional machine learning techniques, but such superior results can only be achieved with proper design of the right DNN architecture, where domain knowledge plays a key role.
Siyu Liao, Liutong Zhou, Xuan Di, Bo Yuan 0001, Jinjun Xiong
ASP-DAC3