VLDB 2026 Research / reviewers in the wild / expert
Xiaotian Dai 0001
dblp:199/5323
· DBLP profile ↗
22ranked-venue papers
5as first author
14since 2021 · last 2026
0000-0002-6669-5234ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 4 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CAFT-RS: Fault-Tolerant Resource Sharing Protocols With Diverse Preemption SchemesabstractEmerging real-time applications increasingly rely on multicore embedded systems, where tasks must coordinate access to shared local and global resources. Such accesses are protected by critical sections and managed by resource-sharing protocols to ensure mutual exclusion and timing predictability. However, transient faults occurring inside critical sections can corrupt execution and propagate errors across tasks, while directly com- bining conventional locking with fault-tolerance mechanisms can significantly increase blocking. Recent fault-tolerant resource- sharing approaches improve recovery through parallel replica execution, but still suffer from sequential global access and coordination overhead. In previous work, we proposed the Lock- frEe Fault-Tolerant Resource Sharing (LEFT-RS) protocol, which improves fault-tolerant global resource access by allowing con- current critical-section execution. However, LEFT-RS enforces non-preemptive global resource access, which can cause excessive arrival blocking for high-priority tasks and limit schedulability. This paper introduces the CAFT-RS (Ceiling-based Access for Fault-Tolerant Resource Sharing) protocol, which applies a priority-ceiling mechanism to both local and global resource ac- cesses. CAFT-RS allows higher-priority tasks to preempt ongoing global accesses while preserving correctness through dedicated post-preemption rules. We develop a worst-case response-time analysis that accounts for both the reduction in arrival blocking and the additional preemption overhead. Extensive evaluation results show that CAFT-RS improves schedulability by up to 188.5% on average over LEFT-RS. Xiaotian Dai 0001, Tong Cheng, Alan Burns 0001, Iain Bate, Shuai Zhao 0004 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2025 | LEFT-RS: A Lock-Free Fault-Tolerant Resource Sharing Protocol for Multicore Real-Time SystemsabstractEmerging real-time applications have driven the transition to multicore embedded systems, where tasks must share resources due to functional demands and limited availability. These resources, whether local or global, are protected within critical sections to prevent race conditions, with locking protocols ensuring both exclusive access and timing requirements. However, transient faults occurring within critical sections can disrupt execution and propagate errors across multiple tasks. Conventional locking protocols fail to address such faults, and integrating traditional fault tolerance techniques often increases blocking. Recent approaches improve fault recovery through parallel replica execution; however, challenges remain due to sequential accessing, coordination overhead, and susceptibility to common-mode faults. In this paper, we propose a Lock-frEe Fault-Tolerant Resource Sharing (LEFT-RS) protocol for multicore real-time systems. LEFT-RS allows tasks to concurrently access and read global resources while entering their critical sections in parallel. Each task can complete its access earlier upon successful execution if other tasks experience faults, thereby improving the efficiency of resource usage. Our design also limits the overhead and enhances fault resilience. We present a comprehensive worst-case response time analysis to ensure timing guarantees. Extensive evaluation results demonstrate that our method significantly outperforms existing approaches, achieving up to an 84.5% improvement in schedulability on average. Xiaotian Dai 0001, Tong Cheng, Alan Burns 0001, Iain Bate, Shuai Zhao 0004 |
RTSS | 2 |
| 2025 | A cache-aware DAG scheduling method on multicores: Exploiting node affinity and deferred executions
Huixuan Yi, Yuanhai Zhang, Zhiyang Lin, Yiyang Gao, Xiaotian Dai 0001, Shuai Zhao 0004 |
J. Syst. Archit. | 6 |
| 2025 | A Hybrid Approach to Refine WCRT Bounds for DAG Scheduling Using Anomaly ClassificationabstractMotivated by the performance demands and stringent timing requirements of safety-critical systems like avionics and autonomous vehicles, research has focused on providing timing guarantees for the scheduling of Directed Acyclic Graph (DAG) tasks in multicore systems. The structural complexity and timing anomalies make this problem challenging. Existing methods bound the Worst-Case Response Time (WCRT) of tasks through static analysis, but these bounds are complicated, difficult to validate, and often remain pessimistic for many scheduling scenarios. Runtime intervention can be effective in eliminating timing anomalies and providing timing guarantees; however, it is ineffective for anomaly-free scheduling scenarios, leads to non-work-conserving schedules, and incurs additional overhead. This paper proposes a hybrid approach to identify timing anomalies in DAG scheduling scenarios within a system, providing tighter WCRT solutions. The static analysis first offers a sufficient anomaly test to directly identify some anomaly-free DAG scheduling scenarios. Leveraging a wide range of scheduling data collected from the running system or its simulator, we then apply a machine learning approach to train a binary classification model, achieving an accuracy of 99.5%. Identifying the anomaly status enables the application of more precise WCRT bounds for different scheduling scenarios, leading to improved system performance. Specifically, we shorten the WCRT bounds for anomaly-free DAG scheduling by an average of up to 21.58%, with a maximum reduction of up to 55.47% compared to the state-of-the-art method. Xiaotian Dai 0001, Alan Burns 0001, Iain Bate |
IEEE Trans. Computers | 2 |
| 2024 | Special edition on resource partitioning for modern multicore systems
Ian Gray, Xiaotian Dai 0001 |
Real Time Syst. | 2 |
| 2024 | Context-Aware Graceful Degradation for Mixed-Criticality Scheduling in Autonomous SystemsabstractAutonomous systems are of high complexity and often regarded as mixed-criticality systems (MCSs) in which functions are allocated criticality levels according to risk assessment based on safety standards. Typically, tasks have different real-time requirements across criticality levels, and the estimated worst-case execution times (WCETs) are distinct. Further, limitations in computational resources increase the difficulty of integrating tasks onto one shared hardware platform. Conventionally, all nonsafety critical tasks must be discarded or suspended to guarantee the execution of safety-critical tasks when facing a timing fault. This typically leads to a considerable decrease in the system’s Quality-of-Service (QoS). Achieving more graceful degradation is critical to minimizing QoS reduction. This work focuses on tackling timing faults and proposes a novel graceful degradation strategy for use in a mixed-criticality context. Thus, when a system has multiple operational modes depending on the environment or an operational task, our approach can give an effective way of managing degradation to maximize QoS, which is currently not sufficiently recognized in MCS. Furthermore, the proposed causality analysis-based degradation process “bridges the gap” so functional dependencies are considered in scheduling design and thus leads to a graceful degradation that is both feasible and reasonable in functional and nonfunctional terms. The evaluations show that QoS can be better preserved using the proposed context-aware degradation process when compared with more conventional MCS scheduling approaches. Jie Zou 0009, Xiaotian Dai 0001, John A. McDermid |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2023 | A High-Resilience Imprecise Computing Architecture for Mixed-Criticality SystemsabstractConventional mixed-criticality systems (MCS)s are designed to terminate the execution of less critical tasks in exceptional situations so that the timing properties of more critical tasks can be preserved. Such a strategy can be controversial and has proven difficult to implement in practice, as it can lead to hazards and reduced functionality due to the absence of the discarded tasks. To mitigate this issue, the imprecise mixed-critically system model (IMCS) has been proposed. In such a model, instead of completely dropping less-critical tasks, these tasks are executed as much as possible through the use of decreased computation precision. Although IMCS could effectively improve the survivability of the less-critical tasks, it also introduces three key drawbacks - run-time computation errors, real-time performance degradation, and lack of flexibility. In this paper, we present a novel IMCS framework, which can (i) mitigate the computation errors caused by imprecise computation; (ii) achieve real-time performance near to that of a conventional MCS; (iii) enhance system-level throughput; and (iv) provide flexibility for run-time configuration. We describe the design details ofHIART-MCS, and then present the corresponding theoretical analysis and optimisation method for its run-time configuration. Finally,HIART-MCS is evaluated against other MCS frameworks using a variety of experimental metrics. Zhe Jiang 0004, Xiaotian Dai 0001, Alan Burns 0001, Neil C. Audsley, Zonghua Gu 0001, Ian Gray |
IEEE Trans. Computers | 2 |
| 2023 | NPRC-I/O: An NoC-Based Real-Time I/O System With Reduced Contention and Enhanced PredictabilityabstractAll systems rely on inputs and outputs (I/Os) to perceive and interact with their surroundings. In safety-critical systems, it is important to guarantee both the performance and time-predictability of I/O operations. However, with the continued growth of architectural complexity in modern safety-critical systems, satisfying such real-time requirements has become increasingly challenging due to complex I/O transaction paths and extensive hardware contention. In this article, we present a new Network-on-Chip (NoC)-based Predictable I/O system framework (NPRC-I/O) which reduces this contention and ensures the performance and time-predictability of I/O operations. Specifically, NPRC-I/O contains a programmable I/O command controller (NPRC-CC) and a run-time reconfigurable NoC ($\text{R}^{2}$NoC), which provides the capability to adjust I/O transaction paths at run time. Using this flexibility, we construct an end-to-end transmission latency analysis and an optimization engine that produces configurations for NPRC-I/O and the I/O traffic in a given system. The constructed analysis and optimization engine guarantee the timing of all hard real-time traffic while reducing the deadline misses of soft real-time traffic and overall transmission latency. Zhe Jiang 0004, Xiaotian Dai 0001, Ian Gray, Zonghua Gu 0001, Qingling Zhao, Shuai Zhao 0004 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2023 | reTSN: Resilient and Efficient Time-Sensitive Network for Automotive In-Vehicle CommunicationabstractTime-sensitive networking (TSN) is being widely investigated to provide Ethernet capabilities for in-vehicle backbone communication. However, the gate control list (GCL), as a simple mechanism for achieving timing determinism for safety-critical traffic (ST) frames with hard deadlines, is too rigid to handle the intrinsic timing uncertainty of automated driving systems (ADSs). Due to the complexity and the unpredictable operating environment, there can be delayed ST frames that can disrupt the solidly fixed timing behavior. In this work, we present one novel approach to effectively use bandwidth resources to deal with delayed ST frames, which cannot be handled by a traditional fixed GCL, and the discarding of them should be reduced to improve the system’s integrity. An acceptance test is implemented to report to the application layer when an ST frame will miss its deadline and hence be rejected, i.e., prevented from entering the switch. To further improve the efficiency of bandwidth usage, we investigate how to improve the performance of the more important Class A frames in an audio-video-bridging (AVB) switch, which adopts a credit-based shaper mechanism, and we propose a constant bandwidth server to replace the credit-based shaper while taking fairness into consideration. Evaluation with extensive experiments shows that both resilience and efficiency of the TSN are significantly enhanced compared with a credit-based shaper, especially when the traffic load is relatively high. For the delayed ST frames and event-triggered traffic, our approach is able to schedule more than the solution using the AVB switch even with a high network load. Jie Zou 0009, Xiaotian Dai 0001, John A. McDermid |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2022 | Using Digital Twins in the Development of Complex Dependable Real-Time Embedded Systems
Xiaotian Dai 0001, Shuai Zhao 0004, Benjamin Lesage, Iain Bate |
ISoLA (4) | 1 |
| 2022 | Toward an Analysable, Scalable, Energy-Efficient I/O Virtualization for Mixed-Criticality SystemsabstractIn mixed-criticality systems (MCSs), timely handling of I/O operations is a key for the system being successfully implemented and appropriately functioned. The I/O system for an MCS must simultaneously enable different features, including isolation/separation, timing-predictability, performance, scalability, and energy-efficiency. Moreover, such an I/O system also requires to manage I/O resource in an adaptive manner to facilitate efficient yet safe resource sharing among components of different criticality levels. Existing approaches cannot achieve all of these requirements simultaneously. This article presents a mixed-criticality I/O management framework, termed MCS-IOV. MCS-IOV is based on hardware-assisted virtualization, which provides temporal and spatial isolation and prohibits fault propagation with limited extra overhead. MCS-IOV extends a real-time I/O virtualization system, by supporting the concept of mixed criticalities and customized interfaces for schedulers, which offers good timing predictability and scalability. Finally, we introduce an energy management framework for MCS-IOV, ensuring the power-efficiency of the design. The MCS-IOV is the first systematical solution that fulfills all the requirements as a mixed-criticality I/O system. Zhe Jiang 0004, Xiaotian Dai 0001, Pan Dong, Neil C. Audsley, Nan Guan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2022 | DAG Scheduling and Analysis on Multi-Core Systems by Modelling Parallelism and DependencyabstractWith ever more complex functionalities being implemented in emerging real-time applications, multi-core systems are demanded for high performance, with directed acyclic graphs (DAG) being used to model functional dependencies. For a single DAG task, our previous work presented a concurrent provider and consumer (CPC) model that captures the node-level dependency and parallelism, which are the two key factors of a DAG. Based on the CPC, scheduling and analysis methods were constructed to reduce makespan and tighten the analytical bound of the task. However, the CPC-based methods cannot support multi-DAGs as the interference between DAGs (i.e., inter-task interference) is not taken into account. To address this limitation, this article proposes a novel multi-DAG scheduling approach which specifies the number of cores a DAG can utilise so that it does not incur the inter-task interference. This is achieved by modelling and understanding the workload distribution of the DAG and the system. By avoiding the inter-task interference, the constructed schedule provides full compatibility for the CPC-based methods to be applied on each DAG and reduces the pessimism of the existing analysis. Experimental results show that the proposed multi-DAG method achieves an improvement up to 80% in schedulability against the original work that it extends, and outperforms the existing multi-DAG methods by up to 60% for tightening the interference. Shuai Zhao 0004, Xiaotian Dai 0001, Iain Bate |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2021 | Brief Industry Paper: Digital Twin for Dependable Multi-Core Real-Time Systems - Requirements and Open ChallengesabstractDevelopment of dependable multi-/many-core systems requires assurance that the system is operable in a range of conditions, subjected to both functional and non-functional requirements. To achieve this, tools need to be implemented that can enable exploration of design options and be able to detect deficiencies earlier to avoid costly system re-design. In this work we discuss the challenges of design of multi-core realtime systems with timing assurance and discuss what are the requirements for modelling, testing and analysis tools. Digital Twin-based predictive modelling and fast design space evaluation are studied that work toward addressing these challenges. Xiaotian Dai 0001, Shuai Zhao 0004, Iain Bate, Alan Burns 0001, Wanli Chang 0001 |
RTAS | 1 |
| 2021 | HIART-MCS: High Resilience and Approximated Computing Architecture for Imprecise Mixed-Criticality SystemsabstractIn mixed-criticality systems (MCSs), less-critical tasks are often terminated to ensure the correct execution of high critical tasks. This strategy could however lead to safety hazards, and largely reduce system functionality due to the absence of the discarded tasks. To overcome this problem, we introduce a high resilience and approximated computing framework for MCS, i.e., HIART-MCS. HIART-MCS introduces a novel processor which supports approximation at the hardware level. Associated with this, we also introduce a new intermediate system mode which allows less-critical tasks to be executed with reduced precision instead of being directly dropped out. Corresponding to the HIART-MCS, we further present a new theoretical model and schedulability analysis providing a timing guarantee for the system, followed by optimisations of the mode switch strategy. As demonstrated in both the theoretical and practical evaluations, HIART-MCS effectively improves the survivability of less-critical tasks with limited sacrifice of the critical tasks and negligible extra overhead. It is notable that HIART-MCS is the first practical framework for imprecise MCSs. Zhe Jiang 0004, Xiaotian Dai 0001, Neil C. Audsley |
RTSS | 2 |
| 2020 | Timing-Accurate General-Purpose I/O for Multi- and Many-Core Systems: Scheduling and Hardware SupportabstractGeneral-purpose I/O widely exists on multi- and many-core systems. For real-time applications, I/O operations are often required to be timing-predictable, i.e., bounded in the worst case, and timing-accurate, i.e., occur at (or near) an exact desired time instant. Unfortunately, both timing requirements of I/O operations are hard to achieve from the system level, especially for many-core architectures, due to various latency and contention factors presented in the path of instigating an I/O request. This paper considers a dedicated I/O co-processing unit, and proposes two scheduling methods, with the necessary hardware support implemented. It is the first work that guarantees timing predictability and maximises timing accuracy of I/O tasks in the multi-and many-core systems. Shuai Zhao 0004, Zhe Jiang 0004, Xiaotian Dai 0001, Iain Bate, Ibrahim Habli, Wanli Chang 0001 |
DAC | 3 |
| 2020 | All In One Network for Driver Attention MonitoringabstractNowadays, driver drowsiness and driver distraction is considered as a major risk for fatal road accidents around the world. As a result, driver monitoring identifying is emerging as an essential function of automotive safety systems. Its basic features include head pose, gaze direction, yawning and eye state analysis. However, existing work has investigated algorithms to detect these tasks separately and was usually conducted under laboratory environments. To address this problem, we propose a multi-task learning CNN framework which simultaneously solve these tasks. The network is implemented by sharing common features and parameters of highly related tasks. Moreover, we propose Dual-Loss Block to decompose the pose estimation task into pose classification and coarse-to-fine regression and Objectcentric Aware Block to reduce orientation estimation errors. Thus, with such novel designs, our model not only achieves SOA results but also reduces the complexity of integrating into automotive safety systems. It runs at 10 fps on vehicle embedded systems which marks a momentous step for this field. More importantly, to facilitate other researchers, we publish our dataset FDUDrivers which contains 20000 images of 100 different drivers and covers various real driving environments. FDUDrivers might be the first comprehensive dataset regarding driver attention monitoring. Xiaotian Dai 0001, Lizhe Qi, Zhe Jiang 0004 |
ICASSP | 3 |
| 2020 | Fixed-Priority Scheduling and Controller Co-Design for Time-Sensitive NetworksabstractTime-sensitive networking (TSN) is a set of standardised communication protocols developed under the IEEE 802.1 working group. TSN aims to support deterministic communication based on network schedules that are distributively configured. It is widely considered as the future in-vehicle network solution for highly automated driving, where the requirement on timing guarantee is alongside the demand of high communication bandwidth. In this work, we study a setting of periodic control and non-control packets, with implicit and arbitrary deadlines, respectively. As the FIFO (first-in, first-out) queues in the 802.1Qbv switch incur long delay in the worst case, which prevents the control tasks from achieving short sampling periods and thus impedes control performance optimisation, we propose the first fixed-priority scheduling (FPS) approach for TSN by leveraging its gate control features. In this context, we develop a finer-grained frame-level response time analysis, which provides a tighter bound than the conventional packet-level analysis. Building upon FPS and the above analysis, we formulate a co-design optimisation problem to decide the sampling periods and poles of real-time controllers with settling time as the objective to minimise, whilst satisfying the schedulability constraint. Xiaotian Dai 0001, Shuai Zhao 0004, Yu Jiang 0001, Xun Jiao 0002, Xiaobo Sharon Hu, Wanli Chang 0001 |
ICCAD | 1 |
| 2020 | DAG Scheduling and Analysis on Multiprocessor Systems: Exploitation of Parallelism and DependencyabstractWith ever more complex functionalities being implemented in emerging real-time applications, multiprocessor systems are demanded for high performance, and directed acyclic graphs (DAGs) are used to model functional dependencies. In this work, we study a single periodic non-preemptive DAG running on a homogeneous multiprocessor platform, which is a common setup in many domains, such as automotive, robotics, and industrial automation. Aiming to reduce the makespan of the DAG and provide a tight yet safe bound, our contributions involve the exploitation of node-level parallelism and inter-node dependency, which are the two key factors of a DAG topology. First, we introduce a concurrent provider and consumer (CPC) model that precisely captures the above two factors, and can be recursively applied when parsing a DAG. Building upon CPC, we propose a novel scheduling method focused on reducing the makespan that orders the nodes in the following sequence: (i) the critical path, (ii) early predecessor paths of the critical path, and (iii) longer paths. Secondly, new response time analysis is presented, which provides a generic bound for any execution order of the non-critical nodes and a specific (tighter) bound for a fixed such order. Comprehensive evaluation demonstrates that our scheduling approach and analysis outperforms the state-of-the-art methods. Shuai Zhao 0004, Xiaotian Dai 0001, Iain Bate, Alan Burns 0001, Wanli Chang 0001 |
RTSS | 2 |
| 2020 | Period adaptation of real-time control tasks with fixed-priority scheduling in cyber-physical systemsabstractLong-lived, non-stop cyber-physical systems (CPS) are subject to evolutionary changes that can undermine the guarantees of schedulability that were verified at the time of deployment. At the same time, knowledge gleamed from extended periods of execution can be exploited to reduce the uncertainties that were inevitably presented in the system models that are used to define the temporal behaviours of the control tasks. In this paper we utilise this knowledge and present an adaptation method that actively extends the period of control tasks at run-time based on historical measurements. This can lead to lower power consumption or to the accommodation of increased computation resource demands from other components of the CPS. The method relies on online monitoring and model-based prediction to degrade control performance while having a minimal and acceptable impact on ongoing operations. Cloud-based computing is used to facilitate decision making and offload the local computation. We evaluate the effectiveness of the proposed method through control-scheduling co-simulation. Xiaotian Dai 0001, Alan Burns 0001 |
J. Syst. Archit. | 1 |
| 2019 | MCS-IOV: Real-Time I/O Virtualization for Mixed-Criticality SystemsabstractIn mixed-criticality systems, timely handling of I/O is a key for the system being successfully implemented and functioning appropriately. The criticality levels of functions and sometimes the whole system are often dependent on the state of the I/O. An I/O system for a MCS must provide simultaneously isolation/separation, performance/efficiency and timing-predictability, as well as being able to manage I/O resource in an adaptive manner to facilitate efficient yet safe resource sharing among components of different criticality levels. Existing approaches cannot achieve all of these requirements simultaneously. This paper presents a MCS I/O management framework, termed MCS-IOV. MCS-IOV is based on hardware assisted virtualisation, which provides temporal and spatial isolation and prohibits fault propagation with small extra overhead in performance. MCS-IOV extends a real-time I/O virtualisation system, by supporting the concept of mixed criticalities and customised interfaces for schedulers, which offers good timing-preditability. MCS-IOV supports I/O driven criticality mode switch (the mode switch can be triggered by detection of unexpected I/O behaviors, e.g., a higher I/O utilization than expected) and timely I/O resource reconfiguration up on that. Finally, We evaluated and demonstrate MCS-IOV in different aspects. Zhe Jiang 0004, Neil C. Audsley, Pan Dong, Nan Guan, Xiaotian Dai 0001, Lifeng Wei |
RTSS | 5 |
| 2019 | Model based system assurance using the structured assurance case metamodel
Tim Kelly, Xiaotian Dai 0001, Shuai Zhao 0004, Richard Hawkins 0001 |
J. Syst. Softw. | 3 |
| 2019 | A Dual-Mode Strategy for Performance-Maximisation and Resource-Efficient CPS DesignabstractThe emerging scenarios of cyber-physical systems (CPS), such as autonomous vehicles, require implementing complex functionality with limited resources, as well as high performances. This paper considers a common setup in which multiple control and non-control tasks share one processor, and proposes a dual-mode strategy. The control task switches between two sampling periods when rejecting (coping with) a disturbance. We create an optimisation framework looking for the switching sampling periods and time instants that maximise the control performance (indexed by settling time) and resource efficiency (indexed by the number of tasks that are schedulable on the processor). The latter objective is enabled with schedulability analysis tailored for the dual-mode model. Experimental results show that (i) given a set of tasks, the proposed strategy improves the control performances whilst retaining schedulability; and (ii) given requirements on the control performances, the proposed strategy is able to schedule more tasks. Xiaotian Dai 0001, Wanli Chang 0001, Shuai Zhao 0004, Alan Burns 0001 |
ACM Trans. Embed. Comput. Syst. | 1 |