Zelin Tong

dblp:252/3776 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
8since 2021 · last 2025
0000-0002-4697-8747ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2025 Asymptotically Optimal Multiprocessor Real-Time Locking for non-JLFP Scheduling
abstract
In prior work, a number of asymptotically optimal suspension-based real-time locking protocols have been presented for job-level fixed priority (JLFP) schedulers, where job priorities do not change. However, the optimality proofs for these locking protocols break down under non-JLFP scheduling, where job priorities can vary. In fact, the problem of designing an asymptotically optimal real-time locking protocol for general non-JLFP scheduling has remained open. This paper closes this problem by presenting the non-JLFP locking protocol (NJLP), the first asymptotically optimal suspension-based real-time locking protocol for non-JLFP schedulers.
Zelin Tong, Syed W. Ali, James H. Anderson
RTAS1
2025 Ros ${ }^{\text {RT }}$: Enabling Flexible Scheduling in Ros 2
abstract
The Robot Operating System 2 (ROS) is heavily used in autonomous systems due to its large ecosystem and modular design. However, ROS remains problematic for real-time applications despite prior efforts to improve its real-time capabilities. Many of these problems are rooted in the implementation of the ROS executor, which does not support preemption or userspecified priorities. These properties are fundamental to real-time scheduling and desirable for general systems to reduce latency. ROS variants in prior work have separately supported the preemption and prioritization of callbacks but place restrictions on the application, preventing their adoption in a real-world workload. This paper addresses these deficiencies with a novel executor framework,$\operatorname{ROS}^{\text{RT}}$, which is compatible with any type of ROS application while supporting preemptive, priority-driven scheduling. Additionally, to support flexible EDF scheduling in ROS${ }^{\text {RT }}$, a custom EDF scheduler implementation is proposed using the new Linux scheduling class SCHED_EXT. ROS${ }^{\text {RT }}$s flexibility and real-time compatibility do not come at the cost of overheads: it achieves a significant decrease in publisher-tosubscriber overhead from the native ROS executor. Finally, this paper concludes with a case study performed on the Autoware Reference System, which simulates the execution of the LiDAR module in an autonomous driving application, demonstrating the ability of$\operatorname{ROS}^{\text{RT}}$at real-time scheduling in real-world scenarios.
Sizhe Liu, Rohan Wagle, Shareef Ahmed, Zelin Tong, James H. Anderson
RTSS4
2024 Predictable GPU Sharing in Component-Based Real-Time Systems
Syed W. Ali, Zelin Tong, Joseph Goh, James H. Anderson
ECRTS2
2023 Want Predictable GPU Execution? Beware SMIs!
abstract
It is common practice today to design complex safety-critical systems by repurposing hardware and software components originally designed for other contexts and using such components in a "black-box" fashion. However, if a black box’s inner workings are not fully understood, then this can be unsafe. This paper reports on an investigation pertaining to a black box that is important for autonomous systems, namely NVIDIA’s CUDA GPU framework. This investigation was motivated by certain timing glitches in CUDA kernels reported in the literature. After extensive tracing and testing efforts, the culprit causing these glitches was surprisingly found to be not CUDA-related at all, but rather delays due to system management interrupts (SMIs), a known source of timing unpredictability on x86 machines that is rarely if ever mentioned in work on real-time GPU usage. The effects of these SMIs are invisible to the operating system and can cause all cores on an x86 machine to become unavailable for over 20ms! This paper describes the methods used to uncover this timing-glitch source. It also discusses some lessons learned when trying to validate the timing behavior of black-box components.
Rohan Wagle, Zelin Tong, Richard L. Sites, James H. Anderson
ICPADS2
2023 Holistically Budgeting Processing Graphs
abstract
To certify the schedulability of a system, valid per-task worst-case execution-time (WCET) estimates are almost always required. Unfortunately, on multicore machines, deriving WCET estimates through static analysis that is not highly pessimistic may never be a practical reality. The alternative is to determine WCETs via a measurement process, but such a process cannot correctly produce accurate WCET estimates with certainty. This lack of certainty necessitates the use of overrun-handling mechanisms, such as budget-enforcement techniques, to preserve temporal correctness at runtime. In many systems of interest today, tasks are interconnected to form processing graphs, which can be quite large. The simplest (and perhaps most common) approach to budget enforcement in this case is to abort an entire graph invocation whenever any node (task) overruns its budget. However, such an approach can result in a high abort rate at the graph level even when the per-node abort rate is low. To remedy this situation, this paper presents a holistic budget-management strategy for directed acyclic graphs (DAGs) that involves reallocating per-node budgets to overrunning nodes to avoid DAG-Ievel aborts. To enable the effects of aborts to be studied analytically, a probabilistic analysis is presented to derive a DAG's abort rate under the proposed budget-management strategy. Experimental results are also presented to demonstrate the utility of budgeting graphs holistically.
Zelin Tong, Shareef Ahmed, James H. Anderson
RTSS1
2022 Overrun-Resilient Multiprocessor Real-Time Locking
Zelin Tong, Shareef Ahmed, James H. Anderson
ECRTS1
2021 TimeWall: Enabling Time Partitioning for Real-Time Multicore+Accelerator Platforms
abstract
Across a range of safety-critical domains, an evolution is underway to endow embedded systems with "thinking" capabilities by using artificial-intelligence (AI) techniques. This evolution is being fueled by the availability of high-performance embedded hardware, typically multicore machines augmented with accelerators. Unfortunately, existing software certification processes rely on time partitioning to isolate system components, and this sense of isolation can be broken by accelerator usage. To address this issue, this paper presents TimeWall, a time-partitioning framework for multicore+accelerator platforms. When applied alongside existing methods for alleviating spatial interference, TimeWall can help enable component-wise certification on multicore+accelerator platforms. The challenges in realizing a TimeWall implementation are discussed in detail in this paper. Additionally, the temporal isolation TimeWall affords is examined experimentally, including via a case study of a computer-vision perception application, on a real platform.
Tanya Amert, Zelin Tong, Sergey Voronov, Joshua Bakita, F. Donelson Smith, James H. Anderson
RTSS2
2021 Statically optimal dynamic soft real-time semi-partitioned scheduling
Clara Hobbs, Zelin Tong, Joshua Bakita, James H. Anderson
Real Time Syst.2