Xiaoyu Mo

dblp:205/4218 · DBLP profile ↗
← Back
16ranked-venue papers
6as first author
14since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 4 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
YearPublicationVenuePosition
2025 Sequence Accumulation and Beyond: Infinite Context Length on Single GPU and Large Clusters
abstract
Linear sequence modeling methods, such as linear attention, state space modeling, and linear RNNs, have recently been recognized as potential alternatives to softmax attention thanks to their linear complexity and competitive performance. However, although their linear-memory advantage during training enables dealing with long sequences, it is still hard to handle extremely long sequences with very limited computational resources. In this paper, we propose Sequence Accumulation (SA) which leverages the common recurrence feature of linear sequence modeling methods to manage infinite context length even on a single GPU. Specifically, SA divides long input sequences into fixed-length sub-sequences and accumulates intermediate states sequentially, which reaches only constant-memory consumption. Additionally, we further propose Sequence Accumulation with Pipeline Parallelism (SAPP), to train large models with infinite context length, without incurring any additional synchronization costs in the sequence dimension. Extensive experiments with a wide range of context lengths are conducted to validate the effectiveness of SA and SAPP on both single and multiple GPUs. Results show that SA and SAPP enable the training of infinite context length on even very limited resources, and are well compatible with the out-of-the-box distributed training techniques.
Weigao Sun, Yongtuo Liu, Xiaqiang Tang, Xiaoyu Mo
AAAI4
2025 Dynamic Prompting Improves Turn-taking in Embodied Spoken Dialogue Systems
abstract
The ability to coordinate turn taking during spoken dialogue is crucial for an embodied spoken dialogue system (SDS), e.g., in a humanoid robot. The SDS needs to model transitions in the conversational floor, which describes each party’s stance (either speaking or listening). Further, the SDS needs to signal its perception of the floor to the human, so that they can coordinate floor transitions and resolve conflicts. Conventional SDS employ standalone modules to control floor transitions but do not produce timely and appropriate responses. Recent end-to-end audio LLMs generate responses quickly, but do not coordinate floor transitions as accurately. In this work, we propose an SDS architecture that dynamically adjusts its prompts to an end-to-end audio LLM based upon its perception of the conversational floor state. The LLM output determines not only the audio output, but also the perceived floor state. This enables the system to signal its stance to the human, both when listening and when speaking. We conducted an experiment where a humanoid robot administered a semi-structured interview with human subjects. Results show that, compared with baseline systems using static prompts, dynamic prompting enables the LLM to model floor transitions more accurately, to generate more appropriate signalling, and to interrupt less, leading to smoother turn-taking in dialogue.
Dingdong Liu, Xiaoyu Mo, Fugee Tsung, Xiaojuan Ma, Bertram E. Shi
RO-MAN3
2025 Toward AI-driven UI transition intuitiveness inspection for smartphone apps
Xiaozhu Hu, Xiaoyu Mo, Xiaofu Jin, Yongquan Hu, Mingming Fan 0001, Tristan Braud
Int. J. Hum. Comput. Stud.2
2025 Hybrid-Prediction Integrated Planning for Autonomous Driving
abstract
Autonomous driving systems require a comprehensive understanding and accurate prediction of the surrounding environment to facilitate informed decision-making in complex scenarios. Recent advances in learning-based systems have highlighted the importance of integrating prediction and planning. However, this integration poses significant alignment challenges through consistency between prediction patterns, to interaction between future prediction and planning. To address these challenges, we introduce a Hybrid-Prediction integrated Planning (HPP) framework, which operates through three novel modules collaboratively. First, we introduce marginal-conditioned occupancy prediction to align joint occupancy with agent-specific motion forecasting. Our proposed MS-OccFormer module achieves spatial-temporal alignment with motion predictions across multiple granularities. Second, we propose a game-theoretic motion predictor, GTFormer, to model the interactive dynamics among agents based on their joint predictive awareness. Third, hybrid prediction patterns are concurrently integrated into the Ego Planner and optimized by prediction guidance. The HPP framework establishes state-of-the-art performance on the nuScenes dataset, demonstrating superior accuracy and safety in end-to-end configurations. Moreover, HPP's interactive open-loop and closed-loop planning performance are demonstrated on the Waymo Open Motion Dataset (WOMD) and CARLA benchmark, outperforming existing integrated pipelines by achieving enhanced consistency between prediction and planning.
Zhiyu Huang, Wenhui Huang 0001, Haohan Yang, Xiaoyu Mo, Chen Lv 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2025 Multimodal Multi-Agent Joint Traffic Simulation With Collision and Run-Off-Road Mitigation
abstract
Simulation environments featuring data-driven traffic agents have become essential tools for training and assessing autonomous driving (AD) systems, offering a safer and more cost-effective alternative to real-world testing. However, prevalent data-driven simulations often focus on modeling individual traffic participants or indirectly representing their joint future motion distribution in a latent space. These approaches may encounter challenges due to the combinatorial explosion in the number of agents or may lack explainability. To tackle these issues, we introduce a novel learning-based traffic simulation framework aimed at directly modeling joint behaviors in the 2D plane. Our framework utilizes a graph-based scene representation to accommodate an arbitrary number of agents and road elements. It directly generates multiple joint goal sets for all agents in a given scenario and subsequently produces reference lines for each goal set. To enhance interaction with the ego, we design a predictive step simulation module to forecast the ego’s short-term motion and prevent collisions between the step rollout and the prediction during reference line tracking. Experimental results on the large-scale Waymo Open Motion Dataset validate that the proposed model can directly generate multiple socially consistent goal sets, showcasing enhancements in collision avoidance, road compliance, and overall realism. This approach shows promise in advancing the development and deployment of AD technology in a safer and more efficient manner.
Xiaoyu Mo, Jintian Ge, Weigao Sun, Chen Lv 0001
IEEE Trans. Intell. Transp. Syst.1
2025 Traffic Scene Representation and Encoding With Graph Structure Learning and Exploration
abstract
Effective scene representation is critical for trajectory generation tasks in autonomous driving. Existing attention-based methods often rely on fixed-length input elements, limiting their ability to adapt to dynamic traffic environments, and frequently overlook lane connectivity. Methods that consider lane connectivity often segment lanes into smaller pieces, which consequently creates a more complex graph structure with an increased number of nodes and edges. Additionally, many current approaches select agent interactions based on fixed distance thresholds, which may miss important long-range or indirect interactions and ignore the fact that proximity does not indicate interaction. In this paper, we propose SceneGNN, a novel framework for learning traffic scenario representation through interaction graph learning and lane graph exploration. SceneGNN constructs a single heterogeneous graph that integrates both agents and lanes, leveraging high-definition maps without over-segmentation. To model lane connectivity more effectively, we introduce LaneGNN, which combines an Omnidirectional Lane Aggregator (OLA) and Directional Lane Explorer (DLE) to explore lane-to-lane interdependencies. Additionally, instead of relying on proximity-based heuristics, our interaction graph learner dynamically constructs inter-agent edges based on learned features, allowing the model to capture meaningful interactions beyond mere distance. We evaluate SceneGNN on the Waymo Open Motion Dataset, where it achieves competitive performance. Our results demonstrate that by efficiently capturing both agent-lane relationships and inter-agent interactions, SceneGNN improves the accuracy and scalability of multi-agent trajectory prediction for real-world autonomous driving applications.
Xiaoyu Mo, Baichuan Lou, Zhiqi Mao, Weigao Sun, Yafei Wang 0001, Chen Lv 0001
IEEE Trans. Intell. Transp. Syst.1
2024 Exploring the Opportunity of Augmented Reality (AR) in Supporting Older Adults to Explore and Learn Smartphone Applications
abstract
The global aging trend compels older adults to navigate the evolving digital landscape, presenting a substantial challenge in mastering smartphone applications. While Augmented Reality (AR) holds promise for enhancing learning and user experience, its role in aiding older adults’ smartphone app exploration remains insufficiently explored. Therefore, we conducted a two-phase study: (1) a workshop with 18 older adults to identify app exploration challenges and potential AR interventions, and (2) tech-probe participatory design sessions with 15 participants to co-create AR support tools. Our research highlights AR’s effectiveness in reducing physical and cognitive strain among older adults during app exploration, especially during multi-app usage and the trial-and-error learning process. We also examined their interactional experiences with AR, yielding design considerations on tailoring AR tools for smartphone app exploration. Ultimately, our study unveils the prospective landscape of AR in supporting the older demographic, both presently and in future scenarios.
Xiaofu Jin, Wai Tong, Xiaoying Wei, Emily Kuang, Xiaoyu Mo, Huamin Qu, Mingming Fan 0001
CHI6
2024 Scalable Traffic Simulation for Autonomous Driving via Multi-Agent Goal Assignment and Autoregressive Goal-Directed Planning
abstract
Simulation provides a fast, cost-effective, and secure environment for developing autonomous driving systems. However, mitigating the gap between simulation and reality is a challenging task as it demands a behavior simulation method that is human-like, diverse, controllable, socially consistent, and scalable. This work proposes a data-driven traffic agent simulation method to address the aforementioned challenges. Our approach centers around a graph-based scene representation and an encoding method, dividing the simulation into two stages: Multi-Agent Goal assignment (MAG) and Goal-Directed Planning (GDP). Firstly, we create joint goal sets for all agents involved in the scenario. Subsequently, we assign target centerlines (TCLs) to each agent based on their predicted goals. To account for any potential mismatch between the predicted joint goal sets and the road structure, we further align the goals of each agent with their respective assigned TCLs. These on-TCL goals serve as inputs for our interactive autoregressive Goal-Directed Planner (AR-GDP), constituting the second stage of our method that generates roll-outs for simulations. Evaluation results on the leaderboard of the Waymo Open Sim Agents Challenge (WOSAC) 2023 show the competitiveness of the proposed method.
Xiaoyu Mo, Zhiyu Huang, Jianwu Fang, Jianru Xue, Chen Lv 0001
IV1
2024 Map-Adaptive Multimodal Trajectory Prediction via Intention-Aware Unimodal Trajectory Predictors
abstract
Autonomous vehicles necessitate prediction of future motions of surrounding traffic participants for safe navigation. However, prediction is challenging due to complex road structures and the multimodality of driving behaviors. Recent approaches usually output a fixed number of predicted trajectories, thus can hardly generalize to situations with more options and match modalities in different situations. They also rely on complex loss design and time-consuming training. This work proposes a novel map-adaptive multimodal trajectory predictor that links driving modalities, driver’s intentions, and a vehicle’s candidate centerlines (CCLs) together, rendering the predictor map-adaptive and the multimodality explainable. The predictor is derived by training an intention-aware unimodal trajectory predictor, which consists of aCCL-based goal predictorand agoal-directed trajectory completer, and aCCL scorerfor estimating the possibilities of a target vehicle (TV) choosing a CCL to follow. This decomposed approach simplifies the training process and reduces the computational resources required, thereby rendering it a faster and more cost-effective alternative. Additionally, the proposed predictor has demonstrated comparable or even superior performance to traditional multimodal predictors in specific applications. Overall, the unimodal predictor presents a promising approach for practical machine learning applications, particularly when computational resources are limited.
Xiaoyu Mo, Zhiyu Huang, Xiuxian Li, Chen Lv 0001
IEEE Trans. Intell. Transp. Syst.1
2023 Designing Loving-Kindness Meditation in Virtual Reality for Long-Distance Romantic Relationships
abstract
Loving-kindness meditation (LKM) is used in clinical psychology for couples' relationship therapy, but physical isolation can make the relationship more strained and inaccessible to LKM. Virtual reality (VR) can provide immersive LKM activities for long-distance couples. However, no suitable commercial VR applications for couples exist to engage in LKM activities of long-distance. This paper organized a series of workshops with couples to build a prototype of a couple-preferred LKM app. Through analysis of participants' design works and semi-structured interviews, we derived design considerations for such VR apps and created a prototype for couples to experience. We conducted a study with couples to understand their experiences of performing LKM using the VR prototype and a traditional video conferencing tool. Results show that LKM session utilizing both tools has a positive effect on the intimate relationship and the VR prototype is a more preferable tool for long-term use. We believe our experience can inform future researchers.
Xiaoyu Mo, Lik-Hang Lee, Xiaoying Wei, Xiaofu Jin, Mingming Fan 0001, Pan Hui 0001
ACM Multimedia2
2022 Multi-modal Motion Prediction with Transformer-based Neural Network for Autonomous Driving
abstract
Predicting the behaviors of other agents on the road is critical for autonomous driving to ensure safety and efficiency. However, the challenging part is how to represent the social interactions between agents and output different possible trajectories with interpretability. In this paper, we introduce a neural prediction framework based on the Transformer structure to model the relationship among the interacting agents and extract the attention of the target agent on the map waypoints. Specifically, we organize the interacting agents into a graph and utilize the multi-head attention Transformer encoder to extract the relations between them. To address the multi-modality of motion prediction, we propose a multi-modal attention Transformer encoder, which modifies the multi-head attention mechanism to multi-modal attention, and each predicted trajectory is conditioned on an independent attention mode. The proposed model is validated on the Argoverse motion forecasting dataset and shows state-of-the-art prediction accuracy while maintaining a small model size and a simple training process. We also demonstrate that the multi-modal attention module can automatically identify different modes of the target agent's attention on the map, which improves the interpretability of the model.
Zhiyu Huang, Xiaoyu Mo
ICRA2
2022 Learning to Compose and Reason with Language Tree Structures for Visual Grounding
abstract
Grounding natural language in images, such as localizing "the black dog on the left of the tree", is one of the core problems in artificial intelligence, as it needs to comprehend the fine-grained language compositions. However, existing solutions merely rely on the association between the holistic language features and visual features, while neglect the nature of composite reasoning implied in the language. In this paper, we propose a natural language grounding model that can automatically compose a binary tree structure for parsing the language and then perform visual reasoning along the tree in a bottom-up fashion. We call our model RvG-Tree: Recursive Grounding Tree, which is inspired by the intuition that any language expression can be recursively decomposed into two constituent parts, and the grounding confidence score can be recursively accumulated by calculating their grounding scores returned by the two sub-trees.RvG-Tree can be trained end-to-end by using the Straight-Through Gumbel-Softmax estimator that allows the gradients from the continuous score functions passing through the discrete tree construction. Experiments on several benchmarks show that our model achieves the state-of-the-art performance with more explainable reasoning.
Richang Hong, Daqing Liu, Xiaoyu Mo, Xiangnan He 0001, Hanwang Zhang
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 Multi-Agent Trajectory Prediction With Heterogeneous Edge-Enhanced Graph Attention Network
abstract
Simultaneous trajectory prediction for multiple heterogeneous traffic participants is essential for safe and efficient operation of connected automated vehicles under complex driving situations. Two main challenges for this task are to handle the varying number of heterogeneous target agents and jointly consider multiple factors that would affect their future motions. This is because different kinds of agents have different motion patterns, and their behaviors are jointly affected by their individual dynamics, their interactions with surrounding agents, as well as the traffic infrastructures. A trajectory prediction method handling these challenges will benefit the downstream decision-making and planning modules of autonomous vehicles. To meet these challenges, we propose a three-channel framework together with a novel Heterogeneous Edge-enhanced graph ATtention network (HEAT). Our framework is able to deal with the heterogeneity of the target agents and traffic participants involved. Specifically, agents’ dynamics are extracted from their historical states using type-specific encoders. The inter-agent interactions are represented with a directed edge-featured heterogeneous graph and processed by the designed HEAT network to extract interaction features. Besides, the map features are shared across all agents by introducing a selective gate-mechanism. And finally, the trajectories of multiple agents are predicted simultaneously. Validations using both urban and highway driving datasets show that the proposed model can realize simultaneous trajectory predictions for multiple agents under complex traffic situations, and achieve state-of-the-art performance with respect to prediction accuracy. The achieved final displacement error (FDE@3sec) is 0.66 meter under urban driving, demonstrating the feasibility and effectiveness of the proposed approach.
Xiaoyu Mo, Zhiyu Huang, Yang Xing 0002, Chen Lv 0001
IEEE Trans. Intell. Transp. Syst.1
2021 Toward Safe and Smart Mobility: Energy-Aware Deep Learning for Driving Behavior Analysis and Prediction of Connected Vehicles
abstract
Connected automated driving technologies have shown tremendous improvement in recent years. However, it is still not clear how driving behaviors and energy consumption correlate with each other and to what extent these factors related to connected vehicles can influence the motion prediction performance. The precise recognition of driving behaviors and prediction of the vehicle motion is critical to the driving safety for connected automated vehicles (CAVs). Hence, in this study, an energy-aware driving pattern analysis and motion prediction system are proposed for CAVs using a deep learning-based time-series modeling approach. First, energy-aware longitudinal acceleration and deceleration behaviors and lateral lane-change behaviors are statistically analyzed. Then, a sliding standard deviation (SSD) test is applied to evaluate the smoothness of the trajectory and velocity signals considering different energy consumption levels. An energy-aware personalized joint time-series modeling (PJTSM) approach based on a deep recurrent neural network (RNN) and long short-term memory (LSTM) cell are proposed for accurate motion (trajectory and velocity) prediction of the leading vehicle. Finally, the differences in the prediction performance regarding different energy consumption levels are compared and discussed. It is shown that due to the higher randomness of the driving behaviors, the prediction accuracy for heavy energy users is the lowest among the three categories, which means it is harder to anticipate the driving behaviors of cars exhibiting heavy energy consumption. The personalized estimation of driving behaviors of CAVs will contribute to safer automated driving and transportation systems.
Yang Xing 0002, Chen Lv 0001, Xiaoyu Mo, Zhongxu Hu, Chao Huang 0006, Peng Hang
IEEE Trans. Intell. Transp. Syst.3
2020 Interaction-Aware Trajectory Prediction of Connected Vehicles using CNN-LSTM Networks
abstract
Predicting the future trajectory of a surrounding vehicle in congested traffic is one of the necessary abilities of an autonomous vehicle. In congestion, a vehicle's future movement is the result of its interaction with surrounding vehicles. A vehicle in congestion may have many neighbors in a relatively short distance, while only a small part of neighbors affect its future trajectory mostly. In this work, An interaction-aware method that predicts the future trajectory of an ego vehicle considering its interaction with eight surrounding vehicles is proposed. The dynamics of vehicles are encoded by LSTMs with shared weights, and the interaction is extracted with a simple CNN. The proposed model is trained and tested on trajectories extracted from the publicly accessible NGSIM US-101 dataset. Quantitative experimental results show that the proposed model outperforms previous models in root-mean-square error (RMSE). Results visualization shows that the model is able to predict future trajectory induced by lane change before the vehicle operates noticeable lateral movement to initiate lane changing.
Xiaoyu Mo, Yang Xing 0002, Chen Lv 0001
IECON1
2019 Secure Pose Estimation for Autonomous Vehicles under Cyber Attacks
abstract
In this paper, we address the problem of secure pose estimation of an autonomous vehicle (AV) under cyber attacks. An extended Kalman filter (EKF) is used to fuse measurements from multiple sensors including GPS, LIDAR, and IMU. To deal with the possible sensor attacks, we design a cumulative sum (CUSUM) detector to monitor the inconsistency between the predicted pose via mathematical model and the sensor measurement. An EKF reconfiguration scheme is proposed to mitigate the influence of sensor attacks once the compromised sensor is identified. The feasibility and effectiveness of the proposed secure pose estimation method are validated using a simulation platform built on Autoware and Gazebo.
Qipeng Liu 0002, Yilin Mo, Xiaoyu Mo, Chen Lv 0001, Ehsan Mihankhah, Danwei Wang
IV3