Xiantong Luo

dblp:336/1614 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
12since 2021 · last 2026
0000-0003-4667-0619ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 2 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Theory of computation · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Flexible Zero-Copy IPC for Processing Chains in ROS 2
abstract
As ROS 2 becomes increasingly adopted in safety-critical real-time systems, the performance of its communication layer, especially inter-process communication (IPC), has emerged as a key bottleneck. While intra-process communication benefits from zero-copy transmission, IPC suffers from significant latency due to serialization and memory copying. Existing shared memory approaches offer limited support for ROS 2 applications, as they impose strict constraints on message formats (e.g., requiring statically sized, POD-compatible types) and overlook end-to-end communication across multi-stage pipelines. In this work, we propose a novel and flexible architecture for enabling zero-copy IPC in ROS 2. Our design supports dynamically structured and non-POD message types, integrates seamlessly with the existing communication framework, and requires no modification to application logic. It consists of a Mini Memory Management System (MMS) for shared memory handling and a Message Propagation Adapter (MPA) that ensures compatibility with the ROS 2 communication framework. Our experimental results show that our method significantly reduces communication latency and supports efficient end-to-end message propagation.
Xiantong Luo, Xu Jiang 0004, Haochun Liang, Yue Tang 0001, Nan Guan, Wang Yi 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2025 On the Scalability and Efficiency of Intra-Process Communication in Ros 2
abstract
The Robot Operating System 2 (ROS 2) has become a widely adopted middleware framework for building modular and distributed robotic systems. Its intra-process communication mechanism is designed to reduce latency by avoiding serialization and memory copying, which is often treated as a negligible or constant-cost operation in both system design and performance analysis. However, this assumption oversimplifies the underlying behavior and may lead to inaccurate performance models and misleading conclusions, especially in latency-sensitive applications. In this paper, we present a comprehensive analysis of intraprocess communication in ROS 2, revealing that its performance is highly sensitive to message configuration, workload structure, and message usage strategies. We identify a scalability risk caused by misaligned communication configurations and propose a guideline to ensure efficient and predictable intra-process communication across varying execution patterns. In addition, we uncover a performance bottleneck in the default ROS 2 implementation stemming from repeated message creation. To address this, we propose a novel message pooling mechanism that reuses message objects to exploit temporal locality and eliminate redundant allocations. Our design is fully compatible with existing ROS 2 APIs and requires no modifications to application-level code. Experimental evaluations using synthetic benchmarks and real-world case studies demonstrate substantial improvements in communication latency, validating the practicality of our design.
Xiantong Luo, Xu Jiang 0004, Nan Guan, Yue Tang 0001, Shaoshuai Zhang
RTSS1
2025 Analysis and optimization of communication delay in multi-subscriber environments of ROS 2
Xiantong Luo, Xu Jiang 0004, Yue Tang 0001, Haochun Liang, Nan Guan, Wang Yi 0001
J. Syst. Archit.1
2025 New Scheduling Algorithm and Analysis for Partitioned Periodic DAG Tasks on Multiprocessors
abstract
Real-time systems are increasingly shifting from single processors to multiprocessors, where software must be parallelized to fully exploit the additional computational power. While the scheduling of real-time parallel tasks modeled as directed acyclic graphs (DAGs) has been extensively studied in the context of global scheduling, the scheduling and analysis of real-time DAG tasks under partitioned scheduling remain far less developed compared to the traditional scheduling of sequential tasks. Existing approaches primarily target plain fixed-priority partitioned scheduling and often rely on self-suspension–based analysis, which limits opportunities for further optimization. In particular, such methods fail to fully leverage fine-grained scheduling management that could improve schedulability. In this paper, we propose a novel approach for scheduling periodic DAG tasks, in which each DAG task is transformed into a set of real-time transactions by incorporating mechanisms for enforcing release offsets and intra-task priority assignments. We further develop corresponding analysis techniques and partitioning algorithms. Through comprehensive experiments, we evaluate the real-time performance of the proposed methods against state-of-the-art scheduling and analysis techniques. The results demonstrate that our approach consistently outperforms existing methods for scheduling periodic DAG tasks across a wide range of parameter settings.
Haochun Liang, Xu Jiang 0004, Xiantong Luo, Songran Liu, Nan Guan, Wang Yi 0001
IEEE Trans. Parallel Distributed Syst.4
2024 Timing Analysis of Cause-Effect Chains for External Events with Finite Validity Intervals
Xiantong Luo, Haochun Liang, Yue Tang 0001, Xu Jiang 0004, Nan Guan, Wang Yi 0001
SETTA1
2024 Timing analysis of processing chains with data refreshing in ROS 2
Yue Tang 0001, Xu Jiang 0004, Nan Guan, Xiantong Luo, Maolin Yang 0004, Wang Yi 0001
J. Syst. Archit.4
2023 Analysis and Optimization of Worst-Case Time Disparity in Cause-Effect Chains
abstract
In automotive systems, an important timing requirement is that the time disparity (the maximum difference among the timestamps of all raw data produced by sensors that an output originates from) must be bounded in a certain range, so that information from different sensors can be correctly synchronized and fused. In this paper, we study the problem of analyzing the worst-case time disparity in cause-effect chains. In particular, we present two bounds, where the first one assumes all chains are independent from each other and the second one takes the fork-join structures into consideration to perform more precise analysis. Moreover, we propose a solution to cut down the worst-case time disparity for a task by designing buffers with proper sizes. Experiments are conducted to show the correctness and effectiveness of both our analysis and optimization methods.
Xu Jiang 0004, Xiantong Luo, Nan Guan, Zheng Dong 0002, Shaoshan Liu, Wang Yi 0001
DATE2
2023 Real-Time Performance Analysis of Processing Systems on ROS 2 Executors
abstract
ROS (Robot Operating System) is one of the most popular robotic software development frameworks. Robotic systems in safety-critical domains are usually subject to hard realtime constraints, so timing behaviors must be formally modeled and analyzed to guarantee that real-time constraints are always honored at run-time. Although a series of analysis techniques has been proposed to analyze the timing performance of ROS 2, the state-of-the-art still generates pessimistic results for ROS 2 systems modeled as DAG (Directed Acyclic Graph). This paper focuses on the analysis of such systems, and proposes techniques to analyze the timing performance in a more precise manner. Experiments with both randomly generated workload and a case study are conducted to evaluate and demonstrate our results.
Yue Tang 0001, Nan Guan, Xu Jiang 0004, Xiantong Luo, Wang Yi 0001
RTAS4
2023 Optimizing End-to-End Latency of Sporadic Cause-Effect Chains Using Priority Inheritance
abstract
Analysis and optimization of end-to-end latency in cause-effect chains is an important problem in real-time systems. Under task-level fixed-priority scheduling, the end-to-end latency largely relies on the relative priority of the tasks in the chain, so previous work has tried to improve the latency via priority assignment. However, the improvement of static priority assignment is limited due to the conflict between schedulability of individual tasks and end-to-end latency of the chain, i.e., a priority assignment leading to good end-to-end latency may make the task set unschedulable. This work proposes a novel method named Dynamic Priority Inheritance Protocol (DPI) to optimize the end-to-end latency of sporadic cause-effect chains. Under DPI, the propagation delay between two communicating jobs is independent of the task relative priority. So the optimization can work on any priority assignment, and no longer conflicts with task schedulability. Moreover, we propose DPI-B, a combination of DPI and a Buffer Manipulation Protocol, for cause-effect chains that also need to meet the determinism requirement. We conduct experiments with both automotive benchmarks and randomly generated workload. The results show the effectiveness of our method in comparison with the state-of-the-art.
Yue Tang 0001, Xu Jiang 0004, Nan Guan, Songran Liu, Xiantong Luo, Wang Yi 0001
RTSS5
2023 Modeling and Analysis of Inter-Process Communication Delay in ROS 2
abstract
ROS 2, the second-generation ROS, is a popular development framework for real-time robotic software. To ensure the timing correctness of applications based on ROS 2, one must model the time delay incurred by two aspects: computation and communication. While significant work has been conducted on computing delay, formal modeling and analysis of communication delay in ROS 2 is still an open issue. In this paper, we first present a formal description on the timing behavior of inter-process communication in ROS 2 with two typical communication policies, namely the InterestTree policy and FIFO policy, and then develop analysis techniques to upper-bound the incurred delay. We conduct experiments to validate the correctness and evaluate the efficacy of our method with case studies on realistic platform.
Xiantong Luo, Xu Jiang 0004, Nan Guan, Haochun Liang, Songran Liu, Wang Yi 0001
RTSS1
2023 Efficient CUDA stream management for multi-DNN real-time inference on embedded GPUs
Weiguang Pang, Xiantong Luo, Kailun Chen, Dong Ji, Lei Qiao 0002, Wang Yi 0001
J. Syst. Archit.2
2023 Comparing Communication Paradigms in Cause-Effect Chains
abstract
A cause-effect chain is a sequence of multi-rate real-time tasks with data dependency. Cause-effect chains are generally subject to end-to-end timing constraints, especially in safety-critical systems. Communication paradigms greatly affect the end-to-end latency of cause-effect chains. This paper compares different communication paradigms (implicit communication, LET, DBP) with regards to the end-to-end latency of cause-effect chains using them, and proposes priority assignment strategies to optimize the end-to-end latency with specific communication paradigm. Experiments with synthesized data based on an automotive benchmark and randomly generated parameters are conducted to evaluate our results.
Yue Tang 0001, Xu Jiang 0004, Nan Guan, Dong Ji, Xiantong Luo, Wang Yi 0001
IEEE Trans. Computers5