Shuai Zhao 0004

dblp:116/8682-4 · DBLP profile ↗
← Back
51ranked-venue papers
11as first author
37since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 32 · 4 first-author · 25 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 4 first-author · 7 since 2021Software engineering, systems software and programming languages · 7 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 An Environment-Aware Verification Framework for LLM-Generated Robot Control Programs
abstract
Large language models (LLMs) are increasingly used in robotics to translate natural language instructions into executable control programs via task-specific prompts. However, existing approaches often lack correctness guarantees for LLM-generated programs, leading to compilation errors and runtime failures. While some methods consider a verification mechanism, they typically assume complete prior knowledge of the environment, making them unsuitable for complex environments where such knowledge is unavailable. This paper introduces VeBot, an environment-aware verification framework designed to ensure the correctness of robot control programs generated by LLMs. Specifically, VeBot introduces: (i) an LLM-friendly robot control language (RCL) that facilitates the program generation by abstracting away the complex Python code details, (ii) a compiler that translates LLM-generated RCL programs into a control flow graph (CFG) while verifying the lexical, syntactic, and semantic correctness, and (iii) a runtime verification mechanism that checks the CFG and compiles the verified segments into executable Python code, avoiding collisions or planning failures during execution. We illustrate the VeBot framework using a household scenario, and the evaluation shows that it consistently outperforms existing methods across a range of LLMs and tasks, achieving high success rates even with lightweight LLMs.
ZhanShang Nie, Xuanming Liu, Kai Huang 0001, Shuai Zhao 0004
DATE8
2026 K-STAR: Knowledge-Guided Submap Tracking and Recovery for Monocular SLAM in Degraded Visual Conditions
Boyang Li 0009, Shuai Zhao 0004, Kai Huang 0001
KSEM (6)4
2026 CAFT-RS: Fault-Tolerant Resource Sharing Protocols With Diverse Preemption Schemes
abstract
Emerging real-time applications increasingly rely on multicore embedded systems, where tasks must coordinate access to shared local and global resources. Such accesses are protected by critical sections and managed by resource-sharing protocols to ensure mutual exclusion and timing predictability. However, transient faults occurring inside critical sections can corrupt execution and propagate errors across tasks, while directly com- bining conventional locking with fault-tolerance mechanisms can significantly increase blocking. Recent fault-tolerant resource- sharing approaches improve recovery through parallel replica execution, but still suffer from sequential global access and coordination overhead. In previous work, we proposed the Lock- frEe Fault-Tolerant Resource Sharing (LEFT-RS) protocol, which improves fault-tolerant global resource access by allowing con- current critical-section execution. However, LEFT-RS enforces non-preemptive global resource access, which can cause excessive arrival blocking for high-priority tasks and limit schedulability. This paper introduces the CAFT-RS (Ceiling-based Access for Fault-Tolerant Resource Sharing) protocol, which applies a priority-ceiling mechanism to both local and global resource ac- cesses. CAFT-RS allows higher-priority tasks to preempt ongoing global accesses while preserving correctness through dedicated post-preemption rules. We develop a worst-case response-time analysis that accounts for both the reduction in arrival blocking and the additional preemption overhead. Extensive evaluation results show that CAFT-RS improves schedulability by up to 188.5% on average over LEFT-RS.
Xiaotian Dai 0001, Tong Cheng, Alan Burns 0001, Iain Bate, Shuai Zhao 0004
IEEE Trans. Parallel Distributed Syst.6
2025 Unlocking a New Rust Programming Experience: Fast and Slow Thinking with LLMs to Conquer Undefined Behaviors
abstract
To provide flexibility and low-level interaction capabilities, the “unsafe” tag in Rust is essential, but undermines memory safety and introduces Undefined Behaviors (UBs) that reduce safety. Eliminating UBs requires a deep understanding of Rust’s safety rules and strong typing. Traditional methods require depth analysis of code, which is laborious and depends on knowledge design. The powerful semantic understanding capabilities of LLM offer new opportunities to solve this problem. Although existing large model debugging frameworks excel in semantic tasks, limited by fixed processes and lack adaptive and dynamic adjustment capabilities. Inspired by the dual process theory of decision-making (“Fast and Slow Thinking”), we present a LLM-based framework called RustBrain that automatically and flexibly minimizes UBs in Rust projects. Fast thinking extracts features to generate solutions, while slow thinking decomposes, verifies, and generalizes them abstractly. To apply verification and generalization results to solution generation, enabling dynamic adjustments and precise outputs, RustBrain integrates two thinking through a feedback mechanism. Experimental results on Miri dataset show a 94.3% pass rate and 80.4% execution rate, improving flexibility and Rust projects safety.
Renshuang Jiang, Pan Dong, Zhenling Duan, Xiaoxiang Fang, Jun Ma 0015, Shuai Zhao 0004, Zhe Jiang 0004
DAC8
2025 Insights from Rights and Wrongs: A Large Language Model for Solving Assertion Failures in RTL Design
abstract
SystemVerilog Assertions (SVAs) are essential for verifying Register Transfer Level (RTL) designs, as they can be embedded into key functional paths to detect unintended behaviours. During simulation, assertion failures occur when the design’s behaviour deviates from expectations. Solving these failures, i.e., identifying and fixing the issues causing the deviation, requires analysing complex logical and timing relationships between multiple signals. This process heavily relies on human expertise, and there is currently no automatic tool available to assist with it. Here, we present AssertSolver, an opensource Large Language Model (LLM) specifically designed for solving assertion failures. By leveraging synthetic training data and learning from error responses to challenging cases, AssertSolver achieves a bug-fixing pass@1 metric of 88.54% on our testbench, significantly outperforming OpenAI’s o1-preview by up to $\mathbf{1 1. 9 7 \%}$. We release our model and testbench for public access to encourage further research: https://github.com/SEU-ACAL/reproduce-AssertSolver-DAC-25.
Jie Zhou 0001, Youshu Ji, Ning Wang 0071, Xinyao Jiao, Bingkun Yao, Xinwei Fang, Shuai Zhao 0004, Nan Guan, Zhe Jiang 0004
DAC8
2025 Insights from Rights and Wrongs: A Large Language Model for Solving Assertion Failures in RTL Design
abstract
SystemVerilog Assertions (SVAs) are essential for verifying Register Transfer Level (RTL) designs, as they can be embedded into key functional paths to detect unintended behaviours. During simulation, assertion failures occur when the design’s behaviour deviates from expectations. Solving these failures, i.e., identifying and fixing the issues causing the deviation, requires analysing complex logical and timing relationships between multiple signals. This process heavily relies on human expertise, and there is currently no automatic tool available to assist with it. Here, we present AssertSolver, an opensource Large Language Model (LLM) specifically designed for solving assertion failures. By leveraging synthetic training data and learning from error responses to challenging cases, AssertSolver achieves a bug-fixing pass@1 metric of $88.54 \%$ on our testbench, significantly outperforming OpenAI’s o1-preview by up to $\mathbf{1 1. 9 7 \%}$. We release our model and testbench for public access to encourage further research: https://github.com/SEU-ACAL/reproduce-AssertSolver-DAC-25.
Jie Zhou 0001, Youshu Ji, Ning Wang 0071, Xinyao Jiao, Bingkun Yao, Xinwei Fang, Shuai Zhao 0004, Nan Guan, Zhe Jiang 0004
DAC8
2025 From Concept to Practice: an Automated LLM-aided UVM Machine for RTL Verification
abstract
Verification presents a major bottleneck in Integrated Circuit (IC) development, consuming nearly 70% of the total development effort. While the Universal Verification Methodology (UVM) is widely used in industry to improve verification efficiency through structured and reusable testbenches, constructing these testbenches and generating sufficient stimuli remain challenging. These challenges arise from the considerable manual coding effort required, repetitive manual execution of multiple EDA tools, and the need for in-depth domain expertise to navigate complex designs. Here, we present UVM2, an automated verification framework that leverages Large Language Models (LLMs) to generate UVM testbenches and iteratively refine them using coverage feedback, significantly reducing manual effort while maintaining rigorous verification standards. To evaluate UVM2, we introduce a benchmark suite comprising Register Transfer Level (RTL) designs of up to 1.6K lines of code. The results show that UVM2reduces testbench setup time by up to 38.82× compared to experienced engineers, and achieve average code and function coverage of 87.44% and 89.58%, outperforming state- of-the-art solutions by 20.96% and 23.51%, respectively.
Junhao Ye, Dingrong Pan, Qichun Chen, Jie Zhou 0001, Shuai Zhao 0004, Xinwei Fang, Xi Wang 0009, Nan Guan, Zhe Jiang 0004
ICCAD7
2025 An LLM-powered Natural-to-Robotic Language Translation Framework with Correctness Guarantees
abstract
The Large Language Models (LLM) are increasingly being deployed in robotics to generate robot control programs for specific user tasks, enabling embodied intelligence. Existing methods primarily focus on LLM training and prompt design that utilize LLMs to generate executable programs directly from user tasks in natural language. However, due to the inconsistency of the LLMs and the high complexity of the tasks, such best-effort approaches often lead to tremendous programming errors in the generated code, which significantly undermines the effectiveness especially when the light-weight LLMs are applied. This paper introduces a natural-robotic language translation framework that (i) provides correctness verification for generated control programs and (ii) enhances the performance of LLMs in program generation via feedback-based fine-tuning for the programs. To achieve this, a Robot Skill Language (RSL) is proposed to abstract away from the intricate details of the control programs, bridging the natural language tasks with the underlying robot skills. Then, the RSL compiler and debugger are constructed to verify RSL programs generated by the LLM and provide error feedback to the LLM for refining the outputs until being verified by the compiler. This provides correctness guarantees for the LLM-generated programs before being offloaded to the robots for execution, significantly enhancing the effectiveness of LLM-powered robotic applications. Experiments demonstrate NRTrans outperforms the existing method under a range of LLMs and tasks, and achieves a high success rate for light-weight LLMs.
ZhenDong Chen, ZhanShang Nie, ShiXing Wan, JunYi Li, YongTian Cheng, Shuai Zhao 0004
IJCNN6
2025 LEFT-RS: A Lock-Free Fault-Tolerant Resource Sharing Protocol for Multicore Real-Time Systems
abstract
Emerging real-time applications have driven the transition to multicore embedded systems, where tasks must share resources due to functional demands and limited availability. These resources, whether local or global, are protected within critical sections to prevent race conditions, with locking protocols ensuring both exclusive access and timing requirements. However, transient faults occurring within critical sections can disrupt execution and propagate errors across multiple tasks. Conventional locking protocols fail to address such faults, and integrating traditional fault tolerance techniques often increases blocking. Recent approaches improve fault recovery through parallel replica execution; however, challenges remain due to sequential accessing, coordination overhead, and susceptibility to common-mode faults. In this paper, we propose a Lock-frEe Fault-Tolerant Resource Sharing (LEFT-RS) protocol for multicore real-time systems. LEFT-RS allows tasks to concurrently access and read global resources while entering their critical sections in parallel. Each task can complete its access earlier upon successful execution if other tasks experience faults, thereby improving the efficiency of resource usage. Our design also limits the overhead and enhances fault resilience. We present a comprehensive worst-case response time analysis to ensure timing guarantees. Extensive evaluation results demonstrate that our method significantly outperforms existing approaches, achieving up to an 84.5% improvement in schedulability on average.
Xiaotian Dai 0001, Tong Cheng, Alan Burns 0001, Iain Bate, Shuai Zhao 0004
RTSS6
2025 Response Time Analysis for Probabilistic Dag Tasks in Multicore Real-Time Systems
abstract
Parallel real-time systems often contain functionalities with complex dependencies and execution uncertainties, leading to significant timing variability which can be represented as a probabilistic distribution. However, existing timing analysis either produces a single conservative bound or incurs high computational costs due to the exhaustive enumeration of every execution scenario. This significantly hinders the exploitation of the probabilistic timing behaviours during system design, leading to sub-optimal design solutions. Modelling the system as a probabilistic directed acyclic graph ($p$-DAG), this paper presents a probabilistic response time analysis based on different longest paths of the$p$-DAG across all execution scenarios, enhancing the capability of the analysis by eliminating the need for enumeration. We first identify every longest path candidate based on the structure of$\boldsymbol{p}$-DAG and compute the probability of its occurrence, where each candidate is the longest under certain execution scenarios. Then, the worst-case interfering workload is computed for each longest path candidate, forming a complete probabilistic response time distribution with correctness guarantees. Experiments show that compared to the enumeration-based approach, the proposed analysis reduces the computation cost by six orders of magnitude while maintaining a low deviation ($\mathbf{1. 0 4 \%}$on average and below$\mathbf{5 \%}$for most$\boldsymbol{p}$-DAGs).
Shuai Zhao 0004, Yiyang Gao, Zhiyang Lin, Boyang Li 0009, Xinwei Fang, Zhe Jiang 0004, Nan Guan
RTSS1
2025 Tight Cache Contention Analysis for WCET Estimation on Multicore Systems
abstract
WCET (Worst-Case Execution Time) estimation on multicore architecture is particularly challenging mainly due to the complex accesses over cache shared by multiple cores. Existing analysis identifies possible contentions between parallel tasks by leveraging the partial order of the tasks or their program regions. Unfortunately, they overestimate the number of cache misses caused by a remote block access without considering the actual cache state and the number of accesses. This paper reports a new analysis for inter-core cache contention. Based on the order of program regions in a task, we first identify memory references that could be affected if a remote access occurs in a region. Afterwards, a fine-grained contention analysis is constructed that computes the number of cache misses based on the access quantity of local and remote blocks. We demonstrate that the overall inter-core cache interference of a task can be obtained via dynamic programming. Experiments show that compared to existing methods, the proposed analysis reduces inter-core cache interference and WCET estimations by$\mathbf{5 2. 3 1 \%}$and$\mathbf{8. 9 4 \%}$on average, without significantly increasing computation overhead.
Shuai Zhao 0004, Jieyu Jiang, Shenlin Cai, Yaowei Liang, Chen Jie, Yinjie Fang, Wei Zhang 0173, Guoquan Zhang, Yaoyao Gu, Ouyang Ouyang, Wanli Chang 0001
RTSS1
2025 FedPillarNet: Unifying personalized and global features for federated 3D LiDAR object detection
Boyang Li 0009, Siheng Ren, Shuai Zhao 0004, Mingyue Cui, Kai Huang 0001
J. Syst. Archit.3
2025 A cache-aware DAG scheduling method on multicores: Exploiting node affinity and deferred executions
Huixuan Yi, Yuanhai Zhang, Zhiyang Lin, Yiyang Gao, Xiaotian Dai 0001, Shuai Zhao 0004
J. Syst. Archit.7
2025 Energy Efficient Scheduling for Position Reconfiguration of Swarm Drones
abstract
Enhancing the energy efficiency of drones, particularly in extending the flight lifetime, has emerged as a crucial area. Position reconfiguration has been explored as a mechanism to achieve this goal for swarm drones. Building on this concept, we investigate how position reconfiguration can be applied within urban wind environments to further extend the lifetime of drone swarms. Despite its potential, efficiently implementing position reconfiguration remains challenging. To address it, we propose an efficient position reconfiguration scheme that reduces the energy consumption imbalance of the swarm and prolongs the lifetime. The scheme includes: (1) a MIP (mixed integer programming)-based optimization method. (2) an approximation algorithm that runs in pseudo-polynomial time and without the need for an optimization solver. The scheme provides a complete position reconfiguration solution that determines (i) the number of position reconfiguration; (ii) when to perform reconfiguration; (iii) who to change positions. Simulation and experimental results demonstrate the effectiveness of our scheme. Note to Practitioners—In urban environments, the significant variation in wind speeds leads to an energy imbalance among swarm drones performing tasks. This paper addresses the practical issue of extending the lifetime of drones in such environments by optimizing position reconfiguration. Specifically, drones operating in high wind speed areas require more energy to maintain hovering, resulting in faster battery depletion. By allowing drones with more remaining energy to exchange positions with those experiencing higher energy consumption, the overall energy usage can be balanced, thus extending the mission duration. We propose an energy-efficient scheduling scheme to determine when and which drones should reconfigure their positions. The scheme strikes a balance between the benefits of reconfiguration and the associated energy costs, preventing unnecessary movement that could waste energy while ensuring drones do not deplete their batteries prematurely. This solution is particularly suited for drone swarms operating in urban environments. Future research could further explore the integration of this scheme into real-time drone fleet management systems.
Mingxin Wei, Shuai Zhao 0004, Hui Cheng 0002, Kai Huang 0001
IEEE Trans Autom. Sci. Eng.3
2025 Multipath Bound for DAG Tasks
abstract
This article studies the response time bound of a directed acyclic graph (DAG) task. Recently, the idea of using multiple paths to bound the response time of a DAG task, instead of using a single longest path in previous results, was proposed and led to the so-called multipath bound. Multipath bounds can greatly reduce the response time bound and significantly improve the schedulability of DAG tasks. This article derives a new multipath bound and proposes an optimal algorithm to compute this bound. We further present a systematic analysis on the dominance and the sustainability of three existing multipath bounds and the proposed multipath bound. Our bound theoretically dominates and empirically outperforms all existing multipath bounds. What is more, the proposed bound is the only multipath bound that is proved to be self-sustainable.
Qingqiang He, Nan Guan, Shuai Zhao 0004, Mingsong Lv
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2025 MERE: Hardware-Software Co-Design for Masking Cache Miss Latency in Embedded Processors
abstract
Runahead execution is a technique to mask memory latency caused by irregular memory accesses. By pre-executing the application code during occurrences of long-latency operations and prefetching anticipated cache-missed data into the cache hierarchy, runahead effectively masks memory latency for subsequent cache misses and achieves high prefetching accuracy; however, this technique has been limited to superscalar out-of-order and superscalar in-order cores. For implementation in scalar in-order cores, the challenges of area-/energy-constraint and severe cache contention remain. Here, we build the first full-stack system featuring runahead, MERE , from SoC and a dedicated ISA to the OS and programming model. Through this deployment, we show that enabling runahead in scalar in-order cores is possible, with minimal area and power overheads, while still achieving high performance. By re-constructing the sequential runahead employing a hardware/software co-design approach, the system can be implemented on a mature processor and SoC. Building on this, an adaptive runahead mechanism is proposed to mitigate the severe cache contention in scalar in-order cores. Combining this, we provide a comprehensive solution for embedded processors managing irregular workloads. Our evaluation demonstrates that the proposed MERE attains 93.5% of a 2-wide out-of-order core’s performance while constraining area and power overheads below 5%, with the adaptive runahead mechanism delivering an additional 20.1% performance gain through mitigating the severe cache contention issues.
Dean You, Jieyu Jiang, Yushu Du, Zhihang Tan, Hui Wang 0166, Jiapeng Guan, Shuai Zhao 0004, Zhe Jiang 0004
ACM Trans. Embed. Comput. Syst.10
2024 Fault-tolerant DAG Scheduling with Runtime Reconfiguration on Multicore Real-Time Systems
abstract
Fault tolerance and real-time performance are two essential goals for directed acyclic graph (DAG) scheduling. However, the redundant tasks to compensate for faults can significantly prolong the completion time of a DAG, i.e., the makespan. In addition, the unpredictable runtime failure status, i.e., a fault may or may not occur during task execution, results in a huge difference between the offline schedule and the actual execution. Existing list scheduling methods can support fault-tolerant execution with all redundant tasks taken into account before runtime. However, such methods cannot effectively reduce the actual makespan as conditional execution of redundant tasks is not considered during scheduling. To address the above issues, this paper proposes a fault-tolerant DAG scheduling method with a runtime reconfiguration facility to optimize the actual makespan. First, the fault-aware makespan is optimized by a fine-grained offline scheduling method considering the worst-case scenario and the runtime flexibility. Then, the runtime reconfiguration mechanism safely moves the influential nodes ahead using the additional time interval to minimize the actual makespan. The experimental results indicate that the proposed method outperforms the state-of-the-art (SOTA) methods in terms of schedulability and actual makespan.
Yuanhai Zhang, Shuai Zhao 0004, Gang Chen 0023, Kai Huang 0001
ASAP2
2024 A Cache/Algorithm Co-design for Parallel Real-Time Systems with Data Dependency on Multi/Many-core System-on-Chips
abstract
Parallel real-time systems rely on a shared cache for dependent data transmission. A conventional shared cache suffers from intensive interference, yet existing cache management techniques only ensure determinism for single-threaded tasks. This paper introduces a virtual indexed, physically tagged, selectively-inclusive, non-exclusive L1.5 Cache, offering way-level control and fine-grained sharing capabilities. Focusing on DAG tasks, we construct a scheduling method that exploits the L1.5 Cache to reduce data transmission, hence, the makespan. As a systematical solution, we built a real system, from the SoC and the ISA to the programming model. Experiments show that our solution significantly improves the timing performance of DAG tasks with negligible overheads.
Zhe Jiang 0004, Shuai Zhao 0004, Yiyang Gao, Jing Li 0025
DAC2
2024 A Du-Octree based Cross-Attention Model for LiDAR Geometry Compression
abstract
Point cloud compression is an essential technology for efficient storage and transmission of 3D data. Previous methods usually use hierarchical tree data structures for encoding the spatial sparseness of point clouds. However, the node context within the tree is not fully discovered since the feature space among nodes varies significantly. To address this problem, we innovatively represent the LiDAR points in a two-octree structure instead of using traditional single-octree coding, and then design the cross-attention model to capture the hierarchical features between different octrees, of which each octree incorporates a transformer-based deep entropy model and an arithmetic encoder. Besides, we introduce the untied cross-aware position encoding with principal component analysis and different projection matrices, which enhances the correlations over two octrees’ attention feature embeddings. Experimental results show that our method outperforms the previous state-of-the-art works, achieving up to 8.2% Bpp savings on point cloud benchmark datasets with different lasers.
Mingyue Cui, Mingjian Feng, Junhua Long, Daosong Hu, Shuai Zhao 0004, Kai Huang 0001
ICRA5
2024 Resource-Aware Task Allocation on Mixed-Criticality Systems: a Task-Splitting Approach
abstract
An important trend of real-time systems is to integrate applications with different criticality levels on a single multicore platform, enabling resource sharing among applications . However, the existing task allocation schemes suffer from the issue of severe resource contention between cores for accessing mutually exclusively shared resources, resulting in significant blocking time. This jeopardizes the system schedulability and leads to the application difficulty of mixed-criticality systems (MCS) in real-world systems. To tackle this issue, this paper proposes a resource-aware task allocation (RATA) algorithm for multicore mixed-criticality systems. The proposed allocation takes the resource usage of tasks into account and aims to localize the most frequently accessed resources by allocating the requesting tasks on the same core, which effectively reduces the inter-core resource contention, hence, improving system schedulability. In addition, a specialized allocation process is constructed for different execution modes in MCS with a task-splitting mechanism, enabling direct support of MCS with improved resource utilization. The experimental results show that RATA outperforms existing methods by 71.18% on average (up to 217.85%) in terms of system schedulability.
Ruoxian Su, Hanzhi Xu, Jieyu Jiang, Shuai Zhao 0004
Internetware4
2024 An Efficient Position Reconfiguration Approach for Maximizing Lifetime of Fixed-wing Swarm Drones
abstract
With the development and application of swarm drones, some researchers have tried to replicating the migration patterns of geese in drones swarm formation to extend their lifetime. However, the problem of performing appropriate position reconfiguration based on the battery energy still remains an unsolved issue. This paper proposes an efficient position reconfiguration approach that reduces the energy consumption imbalance of the swarm and prolongs the lifetime. The approach includes: (1) a two-step MIP (mixed-integer programming)-based optimization method. (2) a two-step heuristic algorithm that can run in pseudo-polynomial time and without the need for an optimization solver. The approach provides a complete position reconfiguration solution that determines (i) the number of position reconfiguration; (ii) which drones need to exchange positions in every position reconfiguration; (iii) the length of time to maintain each position before next reconfiguration. Finally, the approach is compared with other three methods in experiments which demonstrate the effectiveness of it.
Mingyue Cui, Yunxiao Shan, Shuai Zhao 0004, Kai Huang 0001
IROS5
2024 ROTA-I/O: Hardware/Algorithm Co-design for Real-Time I/O Control with Improved Timing Accuracy and Robustness
abstract
In safety-critical systems, timing accuracy is the key to achieving precise I/O control. To meet such strict timing requirements, dedicated hardware assistance has recently been investigated and developed. However, these solutions are often fragile, due to unforeseen timing defects. In this paper, we propose a robust and timing-accurate I/O co-processor, which manages I/O tasks using Execution Time Servers (ETSs) and a two-level scheduler. The ETSs limit the impact of timing defects between tasks, and the scheduler prioritises ETSs based on their importance, offering a robust and configurable scheduling infrastructure. Based on the hardware design, we present an ETS-based timing-accurate I/O schedule, with the ETS parameters configured to further enhance robustness against timing defects. Experiments show the proposed I/O control method outperforms the state-of-the-art method in terms of timing accuracy and robustness without introducing significant overhead.
Zhe Jiang 0004, Shuai Zhao 0004, Xin Si, Gang Chen 0023, Nan Guan
RTSS2
2024 FRAP: A Flexible Resource Accessing Protocol for Multiprocessor Real-Time Systems
abstract
Fully-partitioned fixed-priority scheduling (FP-FPS) multiprocessor systems are widely found in real-time applications, where spin-based protocols are often deployed to manage the mutually exclusive access of shared resources. Unfortunately, existing approaches either enforce rigid spin priority rules for resource accessing or carry significant pessimism in the schedulability analysis, imposing substantial blocking time regardless of task execution urgency or resource over-provisioning. This paper proposes FRAP, a spin-based flexible resource accessing protocol for FP-FPS systems. A task under FRAP can spin at any priority within a range for accessing a resource, allowing flexible and finegrained resource control with predictable worst-case behaviour. Under flexible spinning, we demonstrate that the existing analysis techniques can lead to incorrect timing bounds and present a novel MCMF (minimum cost maximum flow)-based blocking analysis, providing predictability guarantee for FRAP. A spin priority assignment is reported that fully exploits flexible spinning to reduce the blocking time of tasks with high urgency, enhancing the performance of FRAP. Experimental results show that FRAP outperforms the existing spin-based protocols in schedulability by $\mathbf{1 5. 2 0 \%} \mathbf{- 3 2. 7 3 \%}$ on average, up to $\mathbf{6 5. 8 5 \%}$.
Shuai Zhao 0004, Hanzhi Xu, Ruoxian Su, Wanli Chang 0001
RTSS1
2024 Timing-accurate scheduling and allocation for parallel I/O operations in real-time systems
Yuanhai Zhang, Shuai Zhao 0004, Gang Chen 0023, Haoyu Luo, Kai Huang 0001
J. Syst. Archit.2
2023 A Universal Method for Task Allocation on FP-FPS Multiprocessor Systems with Spin Locks
abstract
Many complex real-time systems, such as increasingly automated vehicles and 5G wireless base stations, contain a large amount of shared resources that must be accessed in a mutually exclusive fashion. This leads to significant contention especially when resources are shared across processors. To reduce the contention, various resource-aware task allocation methods have been developed to localize the shared resources. Unfortunately, these existing methods either are tailored for specific scheduling and analysis approaches, or introduce runtime overhead that undermines their applicability. In this paper, we present a task allocation method for a mainstream type of real-time systems in practice: FP-FPS (fully-partitioned fixed-priority scheduling) multiprocessor systems with spin locks managing shared resources. Instead of relying on timing bounds as guidance, we utilize a model to approximate the degree of resource contention between tasks. The model is decoupled from priority assignment algorithms, resource sharing protocols and schedulability tests. Hence, our task allocation method can be applied without detailed knowledge of the underlying system, which is particularly useful during the initial design phase of the system. More detailed information about the system in the later phases of design will push up the approximation accuracy and further enhance the performance. Experimental results show that the proposed method outperforms the state-of-the-art by 13.6% on average (up to 24.2%) in system schedulability with a much less (57x on average) computation cost and negligible runtime overhead.
Shuai Zhao 0004, Yinjie Fang, Wanli Chang 0001
DAC1
2023 Precise Response Time Analysis for Multiple DAG Tasks with Intra-task Priority Assignment
abstract
In many real-time application domains, there are execution dependencies, such tasks may be formulated as multiple Directed Acyclic Graphs (DAGs) and scheduled with intra-task (i.e., intra-DAG) priority assignment. The worst-case completion time of a DAG must be bounded and schedulability analysis must be conducted during the design phase to estimate the required hardware resources. Typical examples include automotive systems and Ultra-Reliable Low Latency Communications (URLLC), which is the “to-business” protocol in 5G technologies, deployed in industrial automation for instance. To bound the execution time of multiple DAGs, there are two key factors to analyze: the intra-task interference for a single DAG and the inter-task interference between DAGs. While extensive efforts have been invested, the existing methods either still contain a large degree of pessimism or are even erroneous due to errors in the derived analysis. In this paper, we first provide an indepth analysis of the limitation and defects of the existing methods. Inspired by these observations, we construct novel response time analysis for multiple DAG tasks with arbitrary intra-task priority assignment. Our analysis precisely accounts for both the intra- and inter-task interference by fully exploring the node parallelism in each DAG as well as between DAGs. Extensive experimental results show that the proposed analysis obtains tighter bounds and improves the system scheduability by at least 300 % compared to state-of-the-art approaches. This improvement is even larger when the scheduling pressure is relatively high, up to 100 % versus 0 % in many cases. This work notably advances the use of response time analysis in industry. Practitioners have to resort to either potentially unsafe measurement results or significant resource over-provisioning when precise analysis is unavailable.
Shuai Zhao 0004, Ian Gray, Alan Burns 0001, Siyuan Ji, Wanli Chang 0001
RTAS2
2023 Brief Industry Paper: A DAG Generator with Full Topology Coverage
abstract
The increasing computational demand promotes the application of parallel tasks with complex execution dependencies in industrial applications. The Directed Acyclic Graph (DAG) task model is widely applied with dedicated scheduling algorithms to understand and manage the execution of such systems. In order to validate the effectiveness of different DAG scheduling algorithms, DAG generators are often applied to produce synthesized DAGs for performance evaluation. However, existing DAG generators either fail to provide sufficient topology coverage or suffer from severe scalability issues, leading to biased and incomplete evaluation results. This paper proposes a novel DAG generator that provides full topology coverage under the given DAG structural parameters while eliminating isomorphic DAGs as well as redundant edges in each DAG. In addition, a verification method is constructed that enables topology coverage, isomorphic DAG identification, and constraint satisfaction of the generated DAGs. The experimental results show that compared to existing generators, the proposed DAG generator achieves full topology coverage and significantly reduces the number of DAGs being produced. The DAG generator proposed in this work provides a complete solution for synthesised DAG generation, enabling fair and comprehensive evaluation of DAG systems.
Yinjie Fang, Shuai Zhao 0004, Yili Guo, Wanli Chang 0001
RTSS2
2023 FTSC: Fault-tolerant scheduling and control co-design for distributed real-time system
Yuanhai Zhang, Zijin Xu, Nan Guan, Shuai Zhao 0004, Gang Chen 0023, Kai Huang 0001
J. Syst. Archit.5
2023 NPRC-I/O: An NoC-Based Real-Time I/O System With Reduced Contention and Enhanced Predictability
abstract
All systems rely on inputs and outputs (I/Os) to perceive and interact with their surroundings. In safety-critical systems, it is important to guarantee both the performance and time-predictability of I/O operations. However, with the continued growth of architectural complexity in modern safety-critical systems, satisfying such real-time requirements has become increasingly challenging due to complex I/O transaction paths and extensive hardware contention. In this article, we present a new Network-on-Chip (NoC)-based Predictable I/O system framework (NPRC-I/O) which reduces this contention and ensures the performance and time-predictability of I/O operations. Specifically, NPRC-I/O contains a programmable I/O command controller (NPRC-CC) and a run-time reconfigurable NoC ($\text{R}^{2}$NoC), which provides the capability to adjust I/O transaction paths at run time. Using this flexibility, we construct an end-to-end transmission latency analysis and an optimization engine that produces configurations for NPRC-I/O and the I/O traffic in a given system. The constructed analysis and optimization engine guarantee the timing of all hard real-time traffic while reducing the deadline misses of soft real-time traffic and overall transmission latency.
Zhe Jiang 0004, Xiaotian Dai 0001, Ian Gray, Zonghua Gu 0001, Qingling Zhao, Shuai Zhao 0004
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.7
2022 Power Line Detection Based on Feature Fusion Deep Learning Network
Kuansheng Zou, Zhenbang Jiang, Shuai Zhao 0004
CGI3
2022 Using Digital Twins in the Development of Complex Dependable Real-Time Embedded Systems
Xiaotian Dai 0001, Shuai Zhao 0004, Benjamin Lesage, Iain Bate
ISoLA (4)2
2022 MSRP-FT: Reliable Resource Sharing on Multiprocessor Mixed-Criticality Systems
abstract
Driven by applications such as autonomous vehicles, spacecrafts, robotics, and industrial automation, real-time systems are required to implement ever more complex functionalities with high performance, while maintaining conventional timing predictability, reliability, and cost efficiency. Necessarily, large-scale resource sharing on multiprocessor architectures has to be deployed. Unfortunately, existing protocols that manage shared resources and bound blocking delay have not considered reliability, i.e. how to handle faults. Contention over shared resources may be seriously aggravated by re-executions that are essential to satisfy a system’s reliability requirements. Hence, there exists a significant barrier to applying resource sharing in the mission-critical sector. This paper fills that gap between reliability and resource sharing. Focusing on mixed-criticality systems (MCS), which widely exist in practice and make the problem more challenging, we propose a fault-tolerance solution which includes the first fault-tolerance multiprocessor resource sharing protocol (namely MSRP-FT) and a system execution model that supports the application of MSRP-FT in MCS. Our aim is to minimize blocking time while satisfying reliability requirements. A schedulability analysis is reported which can guarantee that timing constraints are respected. Compared to the state-of-the-art method, developed for fault-tolerant MCS without resource sharing, we improve the system schedulability by an average of $ 1.28\times$ in stable modes and $ 1.1\times$ during the mode switch.
Shuai Zhao 0004, Ian Gray, Alan Burns 0001, Siyuan Ji, Wanli Chang 0001
RTAS2
2022 Bridging the Pragmatic Gaps for Mixed-Criticality Systems in the Automotive Industry
abstract
An increasingly important trend in the design of safety-critical systems is the integration of components with different levels of criticality onto a common hardware platform. Mixed-criticality systems (MCSs) have been well researched in academia, but can be difficult to implement in industrial scenarios as the theoretical models underpinning the research do not sufficiently consider industrial safety practice and safety standards. In this article, we make the first attempt toward the implementation of the MCS theoretical model in industrial settings. To this end, we identify the pragmatic gaps between theory and practice, and then propose a generic industrial MCS architecture, termedP-MCS(Practical-MCS).P-MCSis built upon the conventional theoretical MCS model with additional considerations of industrial safety requirements: 1) runtime safety analysis, determining preserved applications in each system mode and 2) correct partitioning and isolation of different critical elements. We introduce three implementing methods forP-MCS. Corresponding to the new system architecture, we present a theoretical model and schedulability analysis (with consideration of shared resources) to ensure system predictability. Finally, we evaluate and demonstrateP-MCSin terms of system schedulability, overheads, throughput, and predictability, along with a real-world case study. As shown in the evaluation, the considerations of industrial requirements lead to extra overheads and performance reduction inP-MCS. Such weaknesses can be considerably mitigated by hardware assistance and acceleration.
Zhe Jiang 0004, Shuai Zhao 0004, Richard Paterson, Nan Guan, Yan Zhuang 0013, Neil C. Audsley
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2022 DAG Scheduling and Analysis on Multi-Core Systems by Modelling Parallelism and Dependency
abstract
With ever more complex functionalities being implemented in emerging real-time applications, multi-core systems are demanded for high performance, with directed acyclic graphs (DAG) being used to model functional dependencies. For a single DAG task, our previous work presented a concurrent provider and consumer (CPC) model that captures the node-level dependency and parallelism, which are the two key factors of a DAG. Based on the CPC, scheduling and analysis methods were constructed to reduce makespan and tighten the analytical bound of the task. However, the CPC-based methods cannot support multi-DAGs as the interference between DAGs (i.e., inter-task interference) is not taken into account. To address this limitation, this article proposes a novel multi-DAG scheduling approach which specifies the number of cores a DAG can utilise so that it does not incur the inter-task interference. This is achieved by modelling and understanding the workload distribution of the DAG and the system. By avoiding the inter-task interference, the constructed schedule provides full compatibility for the CPC-based methods to be applied on each DAG and reduces the pessimism of the existing analysis. Experimental results show that the proposed multi-DAG method achieves an improvement up to 80% in schedulability against the original work that it extends, and outperforms the existing multi-DAG methods by up to 60% for tightening the interference.
Shuai Zhao 0004, Xiaotian Dai 0001, Iain Bate
IEEE Trans. Parallel Distributed Syst.1
2021 Invited: Hardware/Software Co-Synthesis and Co-Optimization for Autonomous Systems
abstract
With ever more complicated functionalities being integrated in modern autonomous systems, traditional design methods may not remain sufficient to deliver trusted and high-performance systems with stringent temporal, safety and cost efficiency requirements. In this paper, we discuss the limitations of the traditional design methods with the above requirements enforced, in which hardware and software design are often considered separately. To tackle these limitations, this paper presents a novel design solution that synthesizes both software-level and hardware-level design. First, we highlight and analyze the interconnections between software-level methods (e.g. priority assignment and task allocation) and hardware design (e.g. cache and memory management), in terms of the resulting system performance, e.g. latency. Second, by applying the identified interconnections, we propose an optimization framework to produce high-quality synthesized solutions of both software and hardware design based on a set of candidate design methods. In addition, we describe potential research directions derived from the work and major challenges that can be investigated jointly by engineers and researchers from embedded systems, system safety and programming languages communities.
Wanli Chang 0001, Shuai Zhao 0004, Simon Burton 0001, Haitong Wang, Ting Chen 0002, Neil C. Audsley
DAC2
2021 Brief Industry Paper: Digital Twin for Dependable Multi-Core Real-Time Systems - Requirements and Open Challenges
abstract
Development of dependable multi-/many-core systems requires assurance that the system is operable in a range of conditions, subjected to both functional and non-functional requirements. To achieve this, tools need to be implemented that can enable exploration of design options and be able to detect deficiencies earlier to avoid costly system re-design. In this work we discuss the challenges of design of multi-core realtime systems with timing assurance and discuss what are the requirements for modelling, testing and analysis tools. Digital Twin-based predictive modelling and fast design space evaluation are studied that work toward addressing these challenges.
Xiaotian Dai 0001, Shuai Zhao 0004, Iain Bate, Alan Burns 0001, Wanli Chang 0001
RTAS2
2021 Priority Assignment on Partitioned Multiprocessor Systems With Shared Resources
abstract
Driven by industry demand, there is an increasing need to develop real-time multiprocessor systems which contain shared resources. The Multiprocessor Stack Resource Policy (MSRP) and Multiprocessor resource sharing Protocol (MrsP) are two major protocols that manage access to shared resources. Both of them can be applied to Fixed-Priority Preemptive Scheduling (FPPS), which is enforced by most commercial real-time systems regulations, and which requires task priorities to be assigned before deployment. Along with MSRP and MrsP, there exist two forms of schedulability tests that bound the worst-case blocking time due to resource accesses: the traditional ones being more widely adopted and the more recently developed holistic ones which deliver tighter analysis. On uniprocessor systems, there are several well-established optimal priority assignment algorithms. Unfortunately, on multiprocessor systems with shared resources, the issue of priority assignment has not been adequately understood. In this article, we investigate three mainstream priority assignment algorithms-Deadline Monotonic Priority Ordering (DMPO), Audsley's Optimal Priority Assignment (OPA), and Robust Priority Assignment (RPA), in the context of partitioned multiprocessor systems with shared resources. Our contributions are multifold: First, we prove that DMPO is optimal with the traditional schedulability tests. Second, two counter examples are given as evidence that DMPO is not optimal with the tighter holistic schedulability tests. Third, we then analyze the pessimism arising from the adoption of OPA and RPA with the holistic tests. Lastly, we propose a Slack-based Priority Ordering (SPO) algorithm that minimises such pessimism, and has polynomial time complexity. Comprehensive experiments show that SPO outperforms (i.e., results in a larger number of schedulable systems) DMPO, OPA, and RPA in general with the holistic schedulability tests, by up to 15 percent. With the theoretical contributions, this paper is a useful guide to priority assignment in real-time partitioned multiprocessor systems with shared resources.
Shuai Zhao 0004, Wanli Chang 0001, Weichen Liu 0001, Nan Guan, Alan Burns 0001, Andy J. Wellings
IEEE Trans. Computers1
2020 Timing-Accurate General-Purpose I/O for Multi- and Many-Core Systems: Scheduling and Hardware Support
abstract
General-purpose I/O widely exists on multi- and many-core systems. For real-time applications, I/O operations are often required to be timing-predictable, i.e., bounded in the worst case, and timing-accurate, i.e., occur at (or near) an exact desired time instant. Unfortunately, both timing requirements of I/O operations are hard to achieve from the system level, especially for many-core architectures, due to various latency and contention factors presented in the path of instigating an I/O request. This paper considers a dedicated I/O co-processing unit, and proposes two scheduling methods, with the necessary hardware support implemented. It is the first work that guarantees timing predictability and maximises timing accuracy of I/O tasks in the multi-and many-core systems.
Shuai Zhao 0004, Zhe Jiang 0004, Xiaotian Dai 0001, Iain Bate, Ibrahim Habli, Wanli Chang 0001
DAC1
2020 CPS-oriented Modeling and Control of Traffic Signals Using Adaptive Back Pressure
abstract
Modeling and design of automotive systems from a cyber-physical system (CPS) perspective have lately attracted extensive attention. As the trend towards automated driving and connectivity accelerates, strong interactions between vehicles and the infrastructure are expected. This requires modeling and control of the traffic network in a similarly formal manner. Modeling of such networks involves a tradeoff between expressivity of the appropriate features and tractability of the control problem. Back-pressure control of traffic signals is gaining ground due to its decentralized implementation, low computational complexity, and no requirements on prior traffic information. It guarantees maximum stability under idealistic assumptions. However, when deployed in real traffic intersections, the existing back-pressure control algorithms may result in poor junction utilization due to (i) fixed-length control phases; (ii) stability as the only objective; and (iii) obliviousness to finite road capacities and empty roads. In this paper, we propose a CPS-oriented model of traffic intersections and control of traffic signals, aiming to address the utilization issue of the back-pressure algorithms. We consider a more realistic model with transition phases and dedicated turning lanes, the latter influencing computation of the pressure and subsequently the utilization. The main technical contribution is an adaptive controller that enables varying-length control phases and considers both stability and utilization, while taking both cases of full roads and empty roads into account. We implement a mechanism to prevent frequent changes of control phases and thus limit the number of transition phases, which have negative impact on the junction utilization. Microscopic simulation results with SUMO on a 3×3 traffic network under various traffic patterns show that the proposed algorithm is at least about 13% better in performance than the existing fixed-length backpressure control algorithms reported in previous works. This is a significant improvement in the context of traffic signal control.
Wanli Chang 0001, Debayan Roy, Shuai Zhao 0004, Anuradha M. Annaswamy, Samarjit Chakraborty
DATE3
2020 Fixed-Priority Scheduling and Controller Co-Design for Time-Sensitive Networks
abstract
Time-sensitive networking (TSN) is a set of standardised communication protocols developed under the IEEE 802.1 working group. TSN aims to support deterministic communication based on network schedules that are distributively configured. It is widely considered as the future in-vehicle network solution for highly automated driving, where the requirement on timing guarantee is alongside the demand of high communication bandwidth. In this work, we study a setting of periodic control and non-control packets, with implicit and arbitrary deadlines, respectively. As the FIFO (first-in, first-out) queues in the 802.1Qbv switch incur long delay in the worst case, which prevents the control tasks from achieving short sampling periods and thus impedes control performance optimisation, we propose the first fixed-priority scheduling (FPS) approach for TSN by leveraging its gate control features. In this context, we develop a finer-grained frame-level response time analysis, which provides a tighter bound than the conventional packet-level analysis. Building upon FPS and the above analysis, we formulate a co-design optimisation problem to decide the sampling periods and poles of real-time controllers with settling time as the objective to minimise, whilst satisfying the schedulability constraint.
Xiaotian Dai 0001, Shuai Zhao 0004, Yu Jiang 0001, Xun Jiao 0002, Xiaobo Sharon Hu, Wanli Chang 0001
ICCAD2
2020 Re-Thinking Mixed-Criticality Architecture for Automotive Industry
abstract
Mixed-Criticality System (MCS) has been considered widely within academic literature, but is proving difficulty to implement in industry as the theoretical models underpinning the research do not always consider industrial safety standards and practice (e.g., DO-178C, ISO26262, and EN50128). This paper analyses and formalises the mismatches between theoretical models and industrial standards, and presents a generic industrial MCS architecture, termed as Z-MCS. Z-MCS is built upon the conventional theoretical MCS model (i.e., Adaptive Mixed-Criticality), but with additional satisfaction on the industrial safety requirements: i). run-time safety analysis, which determines preserved applications in each system mode; ii). correct partitioning and isolation of different critical elements with temporal, spatial and fault isolation. Furthermore, three implementing methods of Z-MCS are proposed, with a generic schedulability analysis for timing guarantee. Finally, we evaluate and demonstrate Z-MCS in terms of system schedulability and overheads, along with a real-world case study. In addition, this paper is the first attempt for connecting the theoretical MCS model with the industrial context.
Zhe Jiang 0004, Shuai Zhao 0004, Pan Dong, Nan Guan, Neil C. Audsley
ICCD2
2020 DAG Scheduling and Analysis on Multiprocessor Systems: Exploitation of Parallelism and Dependency
abstract
With ever more complex functionalities being implemented in emerging real-time applications, multiprocessor systems are demanded for high performance, and directed acyclic graphs (DAGs) are used to model functional dependencies. In this work, we study a single periodic non-preemptive DAG running on a homogeneous multiprocessor platform, which is a common setup in many domains, such as automotive, robotics, and industrial automation. Aiming to reduce the makespan of the DAG and provide a tight yet safe bound, our contributions involve the exploitation of node-level parallelism and inter-node dependency, which are the two key factors of a DAG topology. First, we introduce a concurrent provider and consumer (CPC) model that precisely captures the above two factors, and can be recursively applied when parsing a DAG. Building upon CPC, we propose a novel scheduling method focused on reducing the makespan that orders the nodes in the following sequence: (i) the critical path, (ii) early predecessor paths of the critical path, and (iii) longer paths. Secondly, new response time analysis is presented, which provides a generic bound for any execution order of the non-critical nodes and a specific (tighter) bound for a fixed such order. Comprehensive evaluation demonstrates that our scheduling approach and analysis outperforms the state-of-the-art methods.
Shuai Zhao 0004, Xiaotian Dai 0001, Iain Bate, Alan Burns 0001, Wanli Chang 0001
RTSS1
2020 A complete run-time overhead-aware schedulability analysis for MrsP under nested resources
Shuai Zhao 0004, Jorge Garrido, Alan Burns 0001, Andy J. Wellings, Juan Antonio de la Puente
J. Syst. Softw.1
2020 Development Automation of Real-Time Java: Model-Driven Transformation and Synthesis
abstract
Many applications in emerging scenarios, such as autonomous vehicles, intelligent robots, and industrial automation, are safety-critical with strict timing requirements. However, the development of real-time systems is error prone and highly dependent on sophisticated domain expertise, making it a costly process. This article utilises the principles of model-driven engineering (MDE) and proposes two methodologies to automate the development of real-time Java applications. The first one automatically converts standard time-sharing Java applications to real-time Java applications, using a series of transformations. It is in line with the observed industrial trend, such as for the big data technology, of redeveloping existing software without the real-time notion to realise the real-time features. The second one allows users to automatically generate real-time Java application templates with a lightweight modelling language, which can be used to define the real-time properties—essentially a synthesis process. This article opens up a new research direction on development automation of real-time programming languages and inspires many research questions that can be jointly investigated by the embedded systems, programming languages as well as MDE communities.
Wanli Chang 0001, Shuai Zhao 0004, Andy J. Wellings, Jim Woodcock 0001, Alan Burns 0001
ACM Trans. Embed. Comput. Syst.3
2019 Solving the Multi-objective Flexible Job-Shop Scheduling Problem with Alternative Recipes for a Chemical Production Process
Piotr Dziurzanski, Shuai Zhao 0004, Jerry Swan, Leandro Soares Indrusiak, Sebastian Scholze, Karl Krone
EvoApplications2
2019 Cloud-based dynamic distributed optimisation of integrated process planning and scheduling in smart factories
abstract
In smart factories, process planning and scheduling need to be performed every time a new manufacturing order is received or a factory state change has been detected. A new plan and schedule need to be determined quickly to increase the responsiveness of the factory and enlarge its profit. Simultaneous optimisation of manufacturing process planning and scheduling leads to better results than a traditional sequential approach but is computationally more expensive and thus difficult to be applied to real-world manufacturing scenarios. In this paper, a working approach for cloud-based distributed optimisation of process planning and scheduling is presented. It executes a multi-objective genetic algorithm on multiple subpopulations (islands). The number of islands is automatically decided based on the current optimisation state. A number of test cases based on two real-world manufacturing scenarios are used to show the applicability of the proposed solution.
Shuai Zhao 0004, Piotr Dziurzanski, Michal Przewozniczek, Marcin Komarnicki, Leandro Soares Indrusiak
GECCO1
2019 From Java to real-time Java: a model-driven methodology with automated toolchain (invited paper)
abstract
Real-time systems are receiving increasing attention with the emerging application scenarios that are safety-critical, complex in functionality, high on timing-related performance requirements, and cost-sensitive, such as autonomous vehicles. Development of real-time systems is error-prone and highly dependent on the sophisticated domain expertise, making it a costly process. There is a trend of the existing software without the real-time notion being re-developed to realise real-time features, e.g., in the big data technology. This paper utilises the principles of model-driven engineering (MDE) and proposes the first methodology that automatically converts standard time-sharing Java applications to real-time Java applications. It opens up a new research direction on development automation of real-time programming languages and inspires many research questions that can be jointly investigated by the embedded systems, programming languages as well as MDE communities.
Wanli Chang 0001, Shuai Zhao 0004, Andy J. Wellings, Alan Burns 0001
LCTES2
2019 Model based system assurance using the structured assurance case metamodel
Tim Kelly, Xiaotian Dai 0001, Shuai Zhao 0004, Richard Hawkins 0001
J. Syst. Softw.4
2019 A Dual-Mode Strategy for Performance-Maximisation and Resource-Efficient CPS Design
abstract
The emerging scenarios of cyber-physical systems (CPS), such as autonomous vehicles, require implementing complex functionality with limited resources, as well as high performances. This paper considers a common setup in which multiple control and non-control tasks share one processor, and proposes a dual-mode strategy. The control task switches between two sampling periods when rejecting (coping with) a disturbance. We create an optimisation framework looking for the switching sampling periods and time instants that maximise the control performance (indexed by settling time) and resource efficiency (indexed by the number of tasks that are schedulable on the processor). The latter objective is enabled with schedulability analysis tailored for the dual-mode model. Experimental results show that (i) given a set of tasks, the proposed strategy improves the control performances whilst retaining schedulability; and (ii) given requirements on the control performances, the proposed strategy is able to schedule more tasks.
Xiaotian Dai 0001, Wanli Chang 0001, Shuai Zhao 0004, Alan Burns 0001
ACM Trans. Embed. Comput. Syst.3
2017 New schedulability analysis for MrsP
abstract
In this paper we consider a spin-based multi-processor locking protocol, named the Multiprocessor resource sharing Protocol (MrsP). MrsP adopts a helping-mechanism where the preempted resource holder can migrate. The original schedulability analysis of MrsP carries considerable pessimism as it has been developed assuming limited knowledge of the resource usage for each remote task. In this paper new MrsP schedulability analysis is developed that takes into account such knowledge to provide a less pessimistic analysis than that of the original analysis. Our experiments show that, theoretically, the new analysis offers better (at least identical) schedulability than the FIFO non-preemptive protocol, and can outperform FIFO preemptive spin locks under systems with either intensive resource contention or long critical sections. The paper also develops analysis to include the overhead of MrsP's helping mechanism. Although MrsP's helping mechanism theoretically increases schedulability, our evaluation shows that this increase may be negated when the overheads of migrations are taken into account. To mitigate this, we have modified the MrsP protocol to introduce a short non-preemptive section following migration. Our experiments demonstrate that with migration cost, MrsP may not be favourable for short critical sections but provides a better schedulability than other FIFO spin-based protocols when long critical sections are applied.
Shuai Zhao 0004, Jorge Garrido, Alan Burns 0001, Andy J. Wellings
RTCSA1
2017 Safety-critical Java for embedded systems
abstract
Summary This paper presents the motivation for and outcomes of an engineering research project on certifiable Java for embedded systems. The project supports the upcoming standard for safety‐critical Java, which defines a subset of Java and libraries aiming for development of high criticality systems. The outcome of this project include prototype safety‐critical Java implementations, a time‐predictable Java processor, analysis tools for memory safety, and example applications to explore the usability of safety‐critical Java for this application area. The text summarizes developments and key contributions and concludes with the lessons learned. Copyright © 2016 John Wiley & Sons, Ltd.
Martin Schoeberl, Andreas Engelbredt Dalsgaard, René Rydhof Hansen, Stephan Korsholm, Anders P. Ravn, Juan Ricardo Rios, Tórur Biskopstø Strøm, Hans Søndergaard, Andy J. Wellings, Shuai Zhao 0004
Concurr. Comput. Pract. Exp.10