Zonghua Gu 0001

dblp:93/1025 · DBLP profile ↗
← Back
93ranked-venue papers
15as first author
32since 2021 · last 2026
0000-0003-4228-2774ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 55 · 8 first-author · 20 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 5 since 2021Software engineering, systems software and programming languages · 5 · 1 first-author · 2 since 2021Computer networks · 4 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Rotation-invariant representation learning by sector convolution neural networks
Wenwei Lin, Xunpei Sun, Chonghao Zhong, Haitao Meng, Gang Chen 0023, Bingxian Zhang, Zonghua Gu 0001
Pattern Recognit.7
2026 Introduction to the Special Issue on Large Language and Vision Models on the Edge
Zonghua Gu 0001, Shaohua Wan 0001, Zehui Xiong, Chun Jason Xue
ACM Trans. Embed. Comput. Syst.1
2025 Dynamic Neural Fortresses: An Adaptive Shield for Model Extraction Defense
abstract
Model extraction aims to acquire a pre-trained black-box model concealed behind a black-box API. Existing defense strategies against model extraction primarily concentrate on preventing the unauthorized extraction of API functionality. However, two significant challenges still need to be solved: (i) Neural network architecture of the API constitutes a form of intellectual property that also requires protection; (ii) The current practice of allocating the same network architecture to both attack and benign queries results in substantial resource wastage. To address these challenges, we propose a novel \textit{Dynamic Neural Fortresses} (DNF) defense method, employing a dynamic Early-Exit neural network, deviating from the conventional fixed architecture. Firstly, we facilitate the random exit of attack queries from the network at earlier layers. This strategic exit point selection significantly reduces the computational cost for attack queries. Furthermore, the random exit of attack queries from earlier layers introduces increased uncertainty for attackers attempting to discern the exact architecture, thereby enhancing architectural protection. On the contrary, we aim to facilitate benign queries to exit at later layers, preserving model utility, as these layers typically yield meaningful information. Extensive experiments on defending against various model extraction scenarios and datasets demonstrate the effectiveness of DNF, achieving a notable 2$\times$ improvement in efficiency and an impressive reduction of up to 12\% in clone model accuracy compared to SOTA defense methods. Additionally, DNF provides strong protection against neural architecture theft, effectively safeguarding network architecture from being stolen.
Siyu Luan, Zhenyi Wang 0001, Li Shen 0008, Zonghua Gu 0001, Dacheng Tao
ICLR4
2025 Training multi-bit Spiking Neural Network with Virtual Neurons
Zonghua Gu 0001, Ruimin Sun, De Ma
Neurocomputing2
2025 Partitioned Scheduling With Shared Resources on Imprecise Mixed-Criticality Multiprocessor Systems
abstract
Both resource access protocols and real-time scheduling algorithms have been extensively studied in classic embedded real-time systems. However, there has been relatively little attention given to the resource access protocol and real-time scheduling algorithms in mixed-criticality systems. In this article, we pay attention to the problem of scheduling an imprecise mixed-criticality (IMC) taskset on a multiprocessor platform with shared resources. First, we propose an IMC with MSRP (IMC-MSRP) resource access protocol, which ensures mutually exclusive access to the shared resources for the tasks. Second, we propose the schedulability test based on the IMC-multiprocessor stack resource policy (MSRP) for a given task-to-processor mapping method. Third, we propose a feasible task-to-processor mapping algorithm called resource-aware criticality-unaware worst-fit decreasing (RA-CU-WFD), which first assigns tasks sharing the same resources to the same processor to reduce the global waiting time of the tasks and thus improve the schedulability ratio of the system. And then assigns tasks based on the criticality-unaware worst-fit decreasing (CU-WFD) algorithm. Finally, we conduct experiments using the synthetic tasksets, and the experimental results show that the RA-CU-WFD outperforms the other approaches in terms of the schedulability ratio.
Yiwen Zhang 0002, Jin-Peng Ma, Zonghua Gu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2024 Real-time Stereo-based 3D Object Detection for Streaming Perception
abstract
The ability to promptly respond to environmental changes is crucial for the perception system of autonomous driving. Recently, a new task called streaming perception was proposed. It jointly evaluate the latency and accuracy into a single metric for video online perception. In this work, we introduce StreamDSGN, the first real-time stereo-based 3D object detection framework designed for streaming perception. StreamDSGN is an end-to-end framework that directly predicts the 3D properties of objects in the next moment by leveraging historical information, thereby alleviating the accuracy degradation of streaming perception. Further, StreamDSGN applies three strategies to enhance the perception accuracy: (1) A feature-flow-based fusion method, which generates a pseudo-next feature at the current moment to address the misalignment issue between feature and ground truth. (2) An extra regression loss for explicit supervision of object motion consistency in consecutive frames. (3) A large kernel backbone with a large receptive field for effectively capturing long-range spatial contextual features caused by changes in object positions. Experiments on the KITTI Tracking dataset show that, compared with the strong baseline, StreamDSGN significantly improves the streaming average precision by up to 4.33%. Our code is available at https://github.com/weiyangdaren/streamDSGN-pytorch.
Changcai Li, Zonghua Gu 0001, Gang Chen 0023, Libo Huang 0002, Wei Zhang 0092
NeurIPS2
2024 On the Scheduling of Fault-Tolerant Time-Sensitive Networking With IEEE 802.1CB
abstract
Time-Sensitive Networking (TSN) has become the most popular technique in modern safety-critical Automotive and Industrial Automation Networks by providing deterministic transmission policies. However, the data of TSN messages may be affected by transient faults. IEEE 802.1CB, a reliability standard in TSN, protects against such faults by providing disjoint redundant routes for each stream. However, the unique assumption may present a new challenge, i.e., an inadequate number of redundant routes that may negatively impact stream scheduling. This paper presents an offline fault-tolerant TSN scheduling approach that considers such impacts for real-time streams (such as Time-Trigger (TT) and Audio Video Bridging (AVB) streams). Specifically, we intend to calculate the minimum upper bound number of disjoint routes required for each stream to meet the reliability requirements, subsequently enhancing the network’s schedulability. We also propose a service degradation function for AVB streams when the network is under heavy load caused by redundant transmissions of TT streams. This function will maintain schedulability and reliability for AVB streams. Experiments with small-and large-scale synthetic networks show the efficiency.
Chaoquan Wu, Qingxu Deng, Yuhan Lin 0004, Shichang Gao, Zonghua Gu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2024 Criticality-Aware EDF Scheduling for Constrained-Deadline Imprecise Mixed-Criticality Systems
abstract
EDF-VD first focuses on the classic mixed-criticality task model in which all low-criticality (LO) tasks are abandoned in the high-criticality mode, which is an effective dynamic priority scheduling algorithm for mixed-criticality systems. However, it has low schedulability for the imprecise mixed-criticality (IMC) task model with constrained deadlines, in which LO tasks are provided graceful degradation services instead of being abandoned. In this article, we study how to improve schedulability for the IMC tasks model. First, we propose a novel criticalityaware EDF scheduling algorithm (CA-EDF) that tries to delay the LO task execution to improve schedulability. Second, we derive sufficient conditions of schedulability for CA-EDF based on the Demand Bound Function. Finally, we evaluate CA-EDF through extensive simulation. The experimental results indicate that CA-EDF can improve the schedulability ratio by about 13.10% compared to the existing algorithms.
Yiwen Zhang 0002, Jin-Peng Ma, Zonghua Gu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2024 EDF-Based Energy-Efficient Semi-Clairvoyant Scheduling With Graceful Degradation
abstract
Recent works introduce a semi-clairvoyant model, in which the system mode transition is revealed on the arrival of high-criticality jobs. To solve the problem of inconsistency between the correctness criterion for mixed-criticality systems (MCSs) with a semi-clairvoyant and the actual situation, we study the problem of schedulability and energy in MCS with the semi-clairvoyant model in this article. First, we propose a new correctness criterion for MCS with semi-clairvoyant and graceful degradation and develop the schedulability test based on demand bound function methods denoted as SCS-GD. Second, we propose an energy-efficient semi-clairvoyant scheduling algorithm based on SCS-GD denoted as EE-SCS-GD. Finally, we conduct an experimental evaluation of SCS-GD and EE-SCS-GD by synthetically generated task sets. The experimental results show that SCS-GD can improve the schedulability ratio by 5.98% compared to existing algorithms while EE-SCS-GD can save 56.17% energy compared to SCS-GD.
Yiwen Zhang 0002, Zonghua Gu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2024 Energy Management for Fault-tolerant (m,k)-constrained Real-time Systems That Use Standby-Sparing
abstract
Fault tolerance, energy management, and quality of service (QoS) are essential aspects for the design of real-time embedded systems. In this work, we focus on exploring methods that can simultaneously address the above three critical issues under standby-sparing. The standby-sparing mechanism adopts a dual-processor architecture in which each processor plays the role of the backup for the other one dynamically. In this way, it can provide fault tolerance subject to both permanent and transient faults. Due to its duplicate executions of the real-time jobs/tasks, the energy consumption of a standby-sparing system could be quite high. With the purpose of reducing energy under standby-sparing, we proposed three novel scheduling schemes: The first one is for (1, 1)-constrained tasks, and the second one and the third one (which can be combined into an integrated approach to maximize the overall energy reduction) are for general ( m,k )-constrained tasks that require that among any k consecutive jobs of a task no more than ( k - m ) out of them could miss their deadlines. Through extensive evaluations and performance analysis, our results demonstrate that compared with the existing research, the proposed techniques can reduce energy by up to 11% for (1, 1)-constrained tasks and 25% for general ( m,k )-constrained tasks while assuring ( m,k )-constraints and fault tolerance as well as providing better user perceived QoS levels under standby-sparing.
Linwei Niu, Danda B. Rawat, Dakai Zhu 0001, Jonathan Musselwhite, Zonghua Gu 0001, Qingxu Deng
ACM Trans. Embed. Comput. Syst.5
2024 Energy-Aware Adaptive Mixed-Criticality Scheduling with Semi-Clairvoyance and Graceful Degradation
abstract
The classic Mixed-Criticality System (MCS) task model is a non-clairvoyance model in which the change of the system behavior is based on the completion of high-criticality tasks while dropping low-criticality tasks in high-criticality mode. In this paper, we simultaneously consider graceful degradation and semi-clairvoyance in MCS. We first propose the analysis for adaptive mixed-criticality with semi-clairvoyance denoted as C-AMC-sem. The so-called semi-clairvoyance refers to the system’s behavior change being revealed at the time that jobs are released. Moreover, we propose a new algorithm based on C-AMC-sem to reduce energy consumption. Finally, we verify the performance of the proposed algorithms via experiments upon synthetically generated tasksets. The experimental results indicate that the proposed algorithms significantly outperform the existing algorithms.
Yiwen Zhang 0002, Zonghua Gu 0001
ACM Trans. Embed. Comput. Syst.3
2024 Energy-Constrained Scheduling for Weakly Hard Real-Time Systems Using Standby-Sparing
abstract
For real-time embedded systems, QoS (Quality of Service), fault tolerance, and energy budget constraint are among the primary design concerns. In this research, we investigate the problem of energy constrained standby-sparing for both periodic and aperiodic tasks in a weakly hard real-time environment. The standby-sparing systems adopt a primary processor and a spare processor to provide fault tolerance for both permanent and transient faults. For such kind of systems, we firstly propose several novel standby-sparing schemes for the periodic tasks which can ensure the system feasibility under tighter energy budget constraint than the traditional ones. Then based on them integrated approachs for both periodic and aperiodic tasks are proposed to minimize the aperiodic response time whilst achieving better energy and QoS performance under the given energy budget constraint. The evaluation results demonstrated that the proposed techniques significantly outperformed the existing state-of-the-art approaches in terms of feasibility and system performance while ensuring QoS and fault tolerance under the given energy budget constraint.
Linwei Niu, Danda B. Rawat, Jonathan Musselwhite, Zonghua Gu 0001, Qingxu Deng
ACM Trans. Design Autom. Electr. Syst.4
2023 On the Degree of Parallelism in Real-Time Scheduling of DAG Tasks
abstract
Real-time scheduling and analysis of parallel tasks modeled as directed acyclic graphs (DAG) have been intensively studied in recent years. The degree of parallelism of DAG tasks is an important characterization in scheduling. This paper revisits the definition and the computing algorithms for the degree of parallelism of DAG tasks, and clarifies some misunderstandings regarding the degree of parallelism which exist in real-time literature. Based on the degree of the parallelism, we propose a real-time scheduling approach for DAG tasks, which is quite simple but rather effective and outperforms the state-of-the-art by a considerable margin.
Qingqiang He, Nan Guan, Mingsong Lv, Zonghua Gu 0001
DATE4
2023 Towards Effective Training of Robust Spiking Recurrent Neural Networks Under General Input Noise via Provable Analysis
abstract
Recently, bio-inspired spiking neural networks (SNN) with recurrent structures (SRNN) have received increasingly more attention due to their appealing properties for energy-efficiently solving time-domain classification tasks. SRNN s are often executed in noisy environments on resource-constrained devices which can however greatly compromise its accuracy. Thus, one fundamental question that remains unanswered is whether a formal analysis under the general input noise disturbances can be obtained to guarantee the robustness of SRNNs. Several studies have shown great promises by optimizing the bound over adverse perturbations based on Lipschitz continuity theorem, but most of these theoretical analysis are confined to convolutional neural networks (CNN). In this work, we take a further step towards robust SRNN training via provable robustness analysis over input noise perturbations. We show that it is feasible to establish bound analysis for evaluating noise sensitivity for SRNN by using the relation between the input current and the membrane potential change magnitude across a time window. Inspired by the theoretical analysis, we next propose a targeted penalty term in the objective function for training robust SRNN. Experimental results show that our solution outperforms the more complicated state-of-the-art methods on the commonly tested Fashion MNIST and CIFAR-IO image classification datasets.
Wendong Zheng, Yu Zhou 0029, Gang Chen 0023, Zonghua Gu 0001
ICCAD4
2023 ER3D: An Efficient Real-time 3D Object Detection Framework for Autonomous Driving
abstract
3D object detection is a vital computer vision task in mobile robotics and autonomous driving. However, most existing methods have exclusively focused on achieving high accuracy, leading to complex and bulky systems that can not be deployed in a real-time manner. In this paper, we propose the ER3D (Efficient and Real-time 3D) object detection framework, which takes stereo images as input and predicts 3D bounding boxes. Instead of using the complex network architecture, we leverage a fast-but-inaccurate method of semi-global matching (SGM) for depth estimation. To eliminate the accuracy degradation in 3D detection caused by inaccurate depth estimation, we introduce decoupled regression head and 3D distance-consistency loU loss to boost the accuracy performance of the 3D detector with a small computing overhead. ER3D achieves both high-precision and real-time performance to enable practical applications of 3D object detection systems on robotic systems. Extensive experiments with the comparison of the state of the arts demonstrate the superior practicability of ER3D, which achieves comparable detection accuracy with significant leadership on inference efficiency.
Haitao Meng, Changcai Li, Gang Chen 0023, Zonghua Gu 0001, Alois C. Knoll
ICPADS4
2023 Rotation-Invariant Descriptors Learned with Circulant Convolution Neural Networks
abstract
Extracting local features for accurate correspondences between image pairs is an essential basis for various computer vision tasks. Recent works have shown that deep neural networks (DNNs) have demonstrated promising performance in challenging environments. However, these state-of-the-art DNN-based approaches are not well suited for the scenario with geometry rotations due to their intrinsic deficiencies of square kernel structure. That is, square kernel structures in standard DNNs cannot fully identify the essentials of the rotations in geometry. To address this problem, we present RICNN, a novel deep learning framework that encodes invariance against the rotations in geometry explicitly into convolutional neural networks. Rather than using the square-shaped kernel structure, RICNN adopts sectorshaped convolutional kernels to achieve encoding invariance in all rotations. With the explicitness of such rotation encoding, RICNN enables the transfer of perspective DNN models to obtain rotation-invariant descriptions. Furthermore, we propose a novel multi-level hinge triplet loss function to strengthen the matching constraints against geometry rotations. Comprehensive experiments demonstrate the strong generalization ability of the RICNN descriptor on the HPatches dataset. Toward the rotation invariance evaluation, our method shows state-of-the-art results.
Wenwei Lin, Chonghao Zhong, Xunpei Sun, Haitao Meng, Gang Chen 0023, Biao Hu 0001, Zonghua Gu 0001
ICTAI7
2023 LRP-based network pruning and policy distillation of robust and non-robust DRL agents for embedded systems
abstract
Summary Reinforcement learning (RL) is an effective approach to developing control policies by maximizing the agent's reward. Deep reinforcement learning uses deep neural networks (DNNs) for function approximation in RL, and has achieved tremendous success in recent years. Large DNNs often incur significant memory size and computational overheads, which may impede their deployment into resource‐constrained embedded systems. For deployment of a trained RL agent on embedded systems, it is necessary to compress the policy network of the RL agent to improve its memory and computation efficiency. In this article, we perform model compression of the policy network of an RL agent by leveraging the relevance scores computed by layer‐wise relevance propagation (LRP), a technique for Explainable AI (XAI), to rank and prune the convolutional filters in the policy network, combined with fine‐tuning with policy distillation. Performance evaluation based on several Atari games indicates that our proposed approach is effective in reducing model size and inference time of RL agents. We also consider robust RL agents trained with RADIAL‐RL versus standard RL agents, and show that a robust RL agent can achieve better performance (higher average reward) after pruning than a standard RL agent for different attack strengths and pruning rates.
Siyu Luan, Zonghua Gu 0001, Qingling Zhao, Gang Chen 0023
Concurr. Comput. Pract. Exp.2
2023 A High-Resilience Imprecise Computing Architecture for Mixed-Criticality Systems
abstract
Conventional mixed-criticality systems (MCS)s are designed to terminate the execution of less critical tasks in exceptional situations so that the timing properties of more critical tasks can be preserved. Such a strategy can be controversial and has proven difficult to implement in practice, as it can lead to hazards and reduced functionality due to the absence of the discarded tasks. To mitigate this issue, the imprecise mixed-critically system model (IMCS) has been proposed. In such a model, instead of completely dropping less-critical tasks, these tasks are executed as much as possible through the use of decreased computation precision. Although IMCS could effectively improve the survivability of the less-critical tasks, it also introduces three key drawbacks - run-time computation errors, real-time performance degradation, and lack of flexibility. In this paper, we present a novel IMCS framework, which can (i) mitigate the computation errors caused by imprecise computation; (ii) achieve real-time performance near to that of a conventional MCS; (iii) enhance system-level throughput; and (iv) provide flexibility for run-time configuration. We describe the design details ofHIART-MCS, and then present the corresponding theoretical analysis and optimisation method for its run-time configuration. Finally,HIART-MCS is evaluated against other MCS frameworks using a variety of experimental metrics.
Zhe Jiang 0004, Xiaotian Dai 0001, Alan Burns 0001, Neil C. Audsley, Zonghua Gu 0001, Ian Gray
IEEE Trans. Computers5
2023 NPRC-I/O: An NoC-Based Real-Time I/O System With Reduced Contention and Enhanced Predictability
abstract
All systems rely on inputs and outputs (I/Os) to perceive and interact with their surroundings. In safety-critical systems, it is important to guarantee both the performance and time-predictability of I/O operations. However, with the continued growth of architectural complexity in modern safety-critical systems, satisfying such real-time requirements has become increasingly challenging due to complex I/O transaction paths and extensive hardware contention. In this article, we present a new Network-on-Chip (NoC)-based Predictable I/O system framework (NPRC-I/O) which reduces this contention and ensures the performance and time-predictability of I/O operations. Specifically, NPRC-I/O contains a programmable I/O command controller (NPRC-CC) and a run-time reconfigurable NoC ($\text{R}^{2}$NoC), which provides the capability to adjust I/O transaction paths at run time. Using this flexibility, we construct an end-to-end transmission latency analysis and an optimization engine that produces configurations for NPRC-I/O and the I/O traffic in a given system. The constructed analysis and optimization engine guarantee the timing of all hard real-time traffic while reducing the deadline misses of soft real-time traffic and overall transmission latency.
Zhe Jiang 0004, Xiaotian Dai 0001, Ian Gray, Zonghua Gu 0001, Qingling Zhao, Shuai Zhao 0004
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2023 Energy-Aware Partitioned Scheduling of Imprecise Mixed-Criticality Systems
abstract
We consider partitioned scheduling of an imprecise mixed-criticality (IMC) taskset on a uniform multiprocessor platform, with the earliest deadline first-virtual deadline (EDF-VD) as the uniprocessor task scheduling algorithm, and address the optimization problem of finding a feasible task-to-processor assignment and low-criticality (LO) mode processor speed with the objective of minimizing the system’s average energy consumption in LO mode. We propose a task-to-processor assignment algorithm criticality-unaware worst-fit decreasing (CU-WFD) algorithm, which allocates tasks with the worst-fit decreasing (WFD) heuristic method based on utilization values at their respective criticality levels. We determine the energy-efficient speed for each processor based on EDF-VD scheduling, and present our algorithm energy-efficient partitioned scheduling for imprecise mixed-criticality (EEPSIMC) with the CU-WFD heuristic algorithm to minimize system energy consumption. The experimental results show that our proposed algorithm has good performance in terms both schedulability ratio and normalized energy consumption compared to seven comparison baselines.
Yiwen Zhang 0002, Rong-Kun Chen, Zonghua Gu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2023 Edge-AI-Driven Framework with Efficient Mobile Network Design for Facial Expression Recognition
abstract
Facial Expression Recognition (FER) in the wild poses significant challenges due to realistic occlusions, illumination, scale, and head pose variations of the facial images. In this article, we propose an Edge-AI-driven framework for FER. On the algorithms aspect, we propose two attention modules, Arbitrary-oriented Spatial Pooling (ASP) and Scalable Frequency Pooling (SFP), for effective feature extraction to improve classification accuracy. On the systems aspect, we propose an edge-cloud joint inference architecture for FER to achieve low-latency inference, consisting of a lightweight backbone network running on the edge device, and two optional attention modules partially offloaded to the cloud. Performance evaluation demonstrates that our approach achieves a good balance between classification accuracy and inference latency.
Yirui Wu, Lilai Zhang, Zonghua Gu 0001, Hu Lu, Shaohua Wan 0001
ACM Trans. Embed. Comput. Syst.3
2023 Model-Based Reinforcement Learning and Neural-Network-Based Policy Compression for Spacecraft Rendezvous on Resource-Constrained Embedded Systems
abstract
Autonomous spacecraft rendezvous is very challenging in increasingly complex space missions. In this article, we present our approach model-based reinforcement learning for spacecraft rendezvous guidance (MBRL4SRG). We build a Markov decision process model based on the Clohessy-Wiltshire equation of spacecraft dynamics and use dynamic programming to solve it and generate the decision table as the optimal agent policy. Since the onboard computing system of spacecraft is resource constrained in terms of both memory size and processing speed, we train a neural network (NN) as a compact and efficient function approximation to the tabular representation of the decision table. The NN outputs are formally verified using the verification tool ReluVal, and the verification results show that the robustness of the NN is maintained. Experimental results indicate that MBRL4SRG achieves lower computational overhead than the conventional proportional–integral–derivative algorithm and has higher trustworthiness and better computational efficiency during training than the model-free reinforcement learning algorithms.
Zhibin Yang 0005, Linquan Xing, Zonghua Gu 0001, Yingmin Xiao
IEEE Trans. Ind. Informatics3
2023 ECFA: An Efficient Convergent Firefly Algorithm for Solving Task Scheduling Problems in Cloud-Edge Computing
abstract
In cloud-edge computing paradigms, the integration of edge servers and task offloading mechanisms has posed new challenges to developing task scheduling strategies. This paper proposes an efficient convergent firefly algorithm (ECFA) for scheduling security-critical tasks onto edge servers and the cloud datacenter. The proposed ECFA uses a probability-based mapping operator to convert an individual firefly into a scheduling solution, in order to associate the firefly space with the solution space. Distinct from the standard FA, ECFA employs a low-complexity position update strategy to enhance computational efficiency in solution exploration. In addition, we provide a rigorous theoretical analysis to justify that ECFA owns the capability of converging to the global best individual in the firefly space. Furthermore, we introduce the concept of boundary traps for analyzing firefly movement trajectories, and investigate whether ECFA would fall into boundary traps during the evolutionary procedure under different parameter settings. We create various testing instances to evaluate the performance of ECFA in solving the cloud-edge scheduling problem, demonstrating its superiority over FA-based and other competing metaheuristics. Evaluation results also validate that the parameter range derived from the theoretical analysis can prevent our algorithm from falling into boundary traps.
Lu Yin 0005, Jin Sun 0001, Junlong Zhou, Zonghua Gu 0001, Keqin Li 0001
IEEE Trans. Serv. Comput.4
2022 LRP-based Policy Pruning and Distillation of Reinforcement Learning Agents for Embedded Systems
abstract
Reinforcement Learning (RL) is an effective approach to developing control policies by maximizing the agent’s reward. Deep Reinforcement Learning (DRL) uses Deep Neural Networks (DNNs) for function approximation in RL, and has achieved tremendous success in recent years. Large DNNs often incur significant memory size and computational overheads, which greatly impedes their deployment into resource-constrained embedded systems. For deployment of a trained RL agent on embedded systems, it is necessary to compress the Policy Network of the RL agent to improve its memory and computation efficiency. In this paper, we perform model compression of the Policy Network of an RL agent by leveraging the relevance scores computed by Layer-wise Relevance Propagation (LRP), a technique for Explainable AI (XAI), to rank and prune the convolutional filters in the Policy Network, combined with fine-tuning with Policy Distillation. Performance evaluation based on several Atari games indicates that our proposed approach is effective in reducing model size and inference time of RL agents.
Siyu Luan, Zonghua Gu 0001, Qingling Zhao, Gang Chen 0023
ISORC3
2022 Special Issue on Optimization of Cross-layer Collaborative Resource Allocation for Mobile Edge Computing, Caching and Communication
Shaohua Wan 0001, Remigiusz Wisniewski, George C. Alexandropoulos, Zonghua Gu 0001, Pierluigi Siano
Comput. Commun.4
2022 Toward Minimum WCRT Bound for DAG Tasks Under Prioritized List Scheduling Algorithms
abstract
Many modern real-time parallel applications can be modeled as a directed acyclic graph (DAG) task. Recent studies show that the worst-case response time (WCRT) bound of a DAG task can be significantly reduced when the execution order of the vertices is determined by the priority assigned to each vertex of the DAG. How to obtain the optimal vertex priority assignment, and how far from the best-known WCRT bound of a DAG task to the minimum WCRT bound are still open problems. In this article, we aim to construct the optimal vertex priority assignment and derive the minimum WCRT bound for the DAG task. We encode the priority assignment problem into an integer linear programming (ILP) formulation. To solve the ILP model efficiently, we do not involve all variables or constraints. Instead, we solve the ILP model iteratively, i.e., we initially solve the ILP model with only a few primary variables and constraints, and then at each iteration, we increment the ILP model with the variables and constraints which are more likely to derive the optimal priority assignment. Experimental work shows that our method is capable of solving the ILP model optimally without involving too many variables or constraints, e.g., for instances with 50 vertices, we find the optimal priority assignment by involving 12.67% variables on average and within several minutes on average.
Shuangshuang Chang, Ran Bi 0001, Jinghao Sun, Weichen Liu 0001, Qingxu Deng, Zonghua Gu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.7
2022 Online Rerouting and Rescheduling of Time-Triggered Flows for Fault Tolerance in Time-Sensitive Networking
abstract
Time-sensitive networking (TSN) is an industry-standard networking protocol that is widely deployed in safety-critical industrial and automotive networks thanks to its quality-of-service (QoS) mechanisms, esp. deterministic transmission and bounded end-to-end delay for time-triggered (TT) flows. In this article, we focus on TT flows and address the issue of fault tolerance against permanent and transient faults with both spatial and temporal redundancy. We present an efficient heuristic algorithm for online incremental rerouting and rescheduling of disrupted flows, assuming the paths and schedules of existing flows stay fixed. It is complementary to and can be combined with offline routing and scheduling algorithms for achieving fault tolerance based on frame replication and elimination for reliability (FRER) (IEEE 802.1CB). Performance evaluation shows that our approach is able to better recover the system’s degree of redundancy (DoR) and has a higher acceptance rate than related work.
Zonghua Gu 0001, Haichuan Yu, Qingxu Deng, Linwei Niu
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2022 CAN Bus Intrusion Detection Based on Auxiliary Classifier GAN and Out-of-distribution Detection
abstract
The Controller Area Network (CAN) is a ubiquitous bus protocol present in the Electrical/Electronic (E/E) systems of almost all vehicles. It is vulnerable to a range of attacks once the attacker gains access to the bus through the vehicle’s attack surface. We address the problem of Intrusion Detection on the CAN bus and present a series of methods based on two classifiers trained with Auxiliary Classifier Generative Adversarial Network (ACGAN) to detect and assign fine-grained labels to Known Attacks and also detect the Unknown Attack class in a dataset containing a mixture of (Normal + Known Attacks + Unknown Attack) messages. The most effective method is a cascaded two-stage classification architecture, with the multi-class Auxiliary Classifier in the first stage for classification of Normal and Known Attacks, passing Out-of-Distribution (OOD) samples to the binary Real-Fake Classifier in the second stage for detection of the Unknown Attack class. Performance evaluation demonstrates that our method achieves both high classification accuracy and low runtime overhead, making it suitable for deployment in the resource-constrained in-vehicle environment.
Qingling Zhao, Mingqiang Chen, Zonghua Gu 0001, Siyu Luan, Haibo Zeng 0001, Samarjit Chakraborty
ACM Trans. Embed. Comput. Syst.3
2022 Minimizing Stack Memory for Partitioned Mixed-criticality Scheduling on Multiprocessor Platforms
abstract
A Mixed-Criticality System (MCS) features the integration of multiple subsystems that are subject to different levels of safety certification on a shared hardware platform. In cost-sensitive application domains such as automotive E/E systems, it is important to reduce application memory footprint, since such a reduction may enable the adoption of a cheaper microprocessor in the family. Preemption Threshold Scheduling (PTS) is a well-known technique for reducing system stack usage. We consider partitioned multiprocessor scheduling, with Preemption Threshold Adaptive Mixed-Criticality (PT-AMC) as the task scheduling algorithm on each processor and address the optimization problem of finding a feasible task-to-processor mapping with minimum total system stack usage on a resource-constrained multi-processor. We present the Extended Maximal Preemption Threshold Assignment Algorithm (EMPTAA), with dual purposes of improving the taskset’s schedulability if it is not already schedulable, and minimizing system stack usage of the schedulable taskset. We present efficient heuristic algorithms for finding sub-optimal yet high-quality solutions, including Maximum Utilization Difference based Partitioning (MUDP) and MUDP with Backtrack Mapping (MUDP-BM), as well as a Branch-and-Bound (BnB) algorithm for finding the optimal solution. Performance evaluation with synthetic task sets demonstrates the effectiveness and efficiency of the proposed algorithms.
Qingling Zhao, Mengfei Qu, Zonghua Gu 0001, Haibo Zeng 0001
ACM Trans. Embed. Comput. Syst.3
2022 Perceptual Enhancement for Autonomous Vehicles: Restoring Visually Degraded Images for Context Prediction via Adversarial Training
abstract
Realizing autonomous vehicles is one of the ultimate dreams for humans. However, perceptual information collected by sensors in dynamic and complicated environments, in particular, vision information, may exhibit various types of degradation. This may lead to mispredictions of context followed by more severe consequences. Thus, it is necessary to improve degraded images before employing them for context prediction. To this end, we propose a generative adversarial network to restore images from common types of degradation. The proposed model features a novel architecture with an inverse and a reverse module to address additional attributes between image styles. With the supplementary information, the decoding for restoration can be more precise. In addition, we develop a loss function to stabilize the adversarial training with better training efficiency for the proposed model. Compared with several state-of-the-art methods, the proposed method can achieve better restoration performance with high efficiency. It is highly reliable for assisting in context prediction in autonomous vehicles.
Feng Ding 0007, Keping Yu, Zonghua Gu 0001, Yun Q. Shi 0001
IEEE Trans. Intell. Transp. Syst.3
2021 Exploiting augmented intelligence in the modeling of safety-critical autonomous systems
abstract
Abstract Machine learning (ML) is used increasingly in safety-critical systems to provide more complex autonomy to make the system to do decisions by itself in uncertain environments. Using ML to learn system features is fundamentally different from manually implementing them in conventional components written in source code. In this paper, we make a first step towards exploring the architecture modeling of safety-critical autonomous systems which are composed of conventional components and ML components, based on natural language requirements. Firstly, augmented intelligence for restricted natural language requirement modeling is proposed. In that, several AI technologies such as natural language processing and clustering are used to recommend candidate terms to the glossary, as well as machine learning is used to predict the category of requirements. The glossary including data dictionary and domain glossary and the category of requirements will be used in the restricted natural language requirement specification method RNLReq, which is equipped with a set of restriction rules and templates to structure and restrict the way how users document requirements. Secondly, automatic generation of SysML architecture models from the RNLReq requirement specifications is presented. Thirdly, the prototype tool is implemented based on Papyrus. Finally, it presents the evaluation of the proposed approach using an industrial autonomous guidance, navigation and control case study.
Zhibin Yang 0005, Yang Bao 0007, Yongqiang Yang, Jean-Paul Bodeveix, Mamoun Filali, Zonghua Gu 0001
Formal Aspects Comput.7
2021 An Intelligent Video Analysis Method for Abnormal Event Detection in Intelligent Transportation Systems
abstract
Intelligent transportation systems pervasively deploy thousands of video cameras. Analyzing live video streams from these cameras is of significant importance to public safety. As streaming video is increasing, it becomes infeasible to have human operators sitting in front of hundreds of screens to catch suspicious activities or detect objects of interests in real-time. Actually, with millions of traffic surveillance cameras installed, video retrieval is more vital than ever. To that end, this article proposes a long video event retrieval algorithm based on superframe segmentation. By detecting the motion amplitude of the long video, a large number of redundant frames can be effectively removed from the long video, thereby reducing the number of frames that need to be calculated subsequently. Then, by using a superframe segmentation algorithm based on feature fusion, the remaining long video is divided into several Segments of Interest (SOIs) which include the video events. Finally, the trained semantic model is used to match the answer generated by the text question, and the result with the highest matching value is considered as the video segment corresponding to the question. Experimental results demonstrate that our proposed long video event retrieval and description method which significantly improves the efficiency and accuracy of semantic description, and significantly reduces the retrieval time.
Shaohua Wan 0001, Xiaolong Xu 0001, Tian Wang 0001, Zonghua Gu 0001
IEEE Trans. Intell. Transp. Syst.4
2020 Cognitive computing and wireless communications on the edge for healthcare service robots
Shaohua Wan 0001, Zonghua Gu 0001, Qiang Ni
Comput. Commun.2
2020 Deep Learning Models for Real-time Human Activity Recognition with Smartphones
Shaohua Wan 0001, Lianyong Qi, Xiaolong Xu 0001, Zonghua Gu 0001
Mob. Networks Appl.5
2019 Multi-dimensional data indexing and range query processing via Voronoi diagram for internet of things
Shaohua Wan 0001, Yu Zhao 0009, Tian Wang 0001, Zonghua Gu 0001, Qammer H. Abbasi, Kim-Kwang Raymond Choo
Future Gener. Comput. Syst.4
2019 Overfitting remedy by sparsifying regularization on fully-connected layers of CNNs
Qi Xu 0008, Ming Zhang 0018, Zonghua Gu 0001, Gang Pan 0001
Neurocomputing3
2018 EDF-Based Mixed-Criticality Systems with Weakly-Hard Timing Constraints
Zonghua Gu 0001, Nenggan Zheng
GPC2
2018 GA-Based Mapping and Scheduling of HSDF Graphs on Multiprocessor Platforms
Nenggan Zheng, Zonghua Gu 0001
GPC4
2018 Introduction to the special issue on "Embedded Artificial Intelligence and Smart Computing"
Zonghua Gu 0001, Meikang Qiu
J. Syst. Archit.1
2018 A decomposition-based approach to optimization of TTP-based distributed embedded systems
Ming Zhang 0018, Nenggan Zheng, Zonghua Gu 0001
J. Syst. Archit.4
2018 Schedulability analysis and stack size minimization with preemption thresholds and mixed-criticality scheduling
Qingling Zhao, Zonghua Gu 0001, Haibo Zeng 0001, Nenggan Zheng
J. Syst. Archit.2
2017 Darwin: A neuromorphic hardware co-processor based on spiking neural networks
De Ma, Juncheng Shen, Zonghua Gu 0001, Ming Zhang 0018, Xiaoqiang Xu, Qi Xu 0008, Yangjing Shen, Gang Pan 0001
J. Syst. Archit.3
2017 Design optimization for AUTOSAR models with preemption thresholds and mixed-criticality scheduling
Qingling Zhao, Zonghua Gu 0001, Haibo Zeng 0001
J. Syst. Archit.2
2017 Optimization of Real-Time Software Implementing Multi-Rate Synchronous Finite State Machines
abstract
Model-based design using Synchronous Reactive (SR) models is becoming widespread for control software development in industry. However, software synthesis is challenging for multi-rate SR models consisting of blocks modeled with finite state machines, due to the complexity of validating the system’s real-time schedulability. The existing approach uses the simplified periodic task model to allow an efficient schedulability analysis, which leads to pessimistic and suboptimal solutions. Instead, in this paper, we adopt a more accurate but more complex schedulability analysis. We develop several optimization techniques to improve the algorithm’s efficiency. Experimental results on synthetic systems and an industrial case study show that the proposed optimization framework preserves the solution optimality but is much faster (e.g., 1000× for systems with 15 blocks) than the branch-and-bound algorithm, and it generates better control software than the existing approach.
Yecheng Zhao, Haibo Zeng 0001, Zonghua Gu 0001
ACM Trans. Embed. Comput. Syst.4
2017 Optimized Implementation of Multirate Mixed-Criticality Synchronous Reactive Models
abstract
Model-based design using Synchronous Reactive (SR) models enables early design and verification of application functionality in a platform-independent manner, and the implementation on the target platform should guarantee the preservation of application semantic properties. Mixed-Criticality Scheduling (MCS) is an effective approach to addressing diverse certification requirements of safety-critical systems that integrate multiple subsystems with different levels of criticality. This article considers fixed-priority scheduling of mixed-criticality SR models, and considers two scheduling approaches: Adaptive MCS and Elastic MCS. We formulate the optimization problem of minimizing the total system cost of added functional delays in the implementation while guaranteeing schedulability, and present an optimal algorithm based on branch-and-bound search, and an efficient heuristic algorithm.
Qingling Zhao, Zaid Al-bayati, Zonghua Gu 0001, Haibo Zeng 0001
ACM Trans. Design Autom. Electr. Syst.3
2016 Darwin: a neuromorphic hardware co-processor based on Spiking Neural Networks
Juncheng Shen, De Ma, Zonghua Gu 0001, Ming Zhang 0018, Xiaoqiang Xu, Qi Xu 0008, Yangjing Shen, Gang Pan 0001
Sci. China Inf. Sci.3
2016 Special Issue on High Performance Computing, Communication and Embedded Software/Systems
Zonghua Gu 0001, Meikang Qiu, Dakai Zhu 0001
J. Syst. Archit.1
2016 HLC-PCP: A resource synchronization protocol for certifiable mixed criticality scheduling
Qingling Zhao, Zonghua Gu 0001, Haibo Zeng 0001
J. Syst. Archit.2
2016 Cache-Partitioned Preemption Threshold Scheduling
Zonghua Gu 0001, Chao Wang 0097, Haibo Zeng 0001
ACM Trans. Embed. Comput. Syst.1
2016 Minimizing Stack Memory for Hard Real-Time Applications on Multicore Platforms with Partitioned Fixed-Priority or EDF Scheduling
abstract
Multicore processors are increasingly adopted in resource-constrained real-time embedded applications. In the development of such applications, efficient use of RAM memory is as important as the effective scheduling of software tasks. Preemption Threshold Scheduling (PTS) is a well-known technique for controlling the degree of preemption, possibly improving system schedulability, and to reduce system stack usage. In this paper, we consider partitioned multi-processor scheduling on a multicore processor with either Fixed-Priority or Earliest Deadline First scheduling algorithms with PTS and address the design optimization problem of mapping tasks to processor cores and assignment of task priorities and preemption thresholds with the optimization objective of minimizing system stack usage. We present both optimal solution techniques based on Mixed Integer Linear Programming and efficient heuristic algorithms that can achieve high-quality results. We perform extensive performance evaluations using both synthetic tasksets and industrial case studies.
Chao Wang 0097, Chuansheng Dong, Haibo Zeng 0001, Zonghua Gu 0001
ACM Trans. Design Autom. Electr. Syst.4
2016 Security-Aware Mapping and Scheduling with Hardware Co-Processors for FlexRay-Based Distributed Embedded Systems
abstract
Automotive in-vehicle systems are distributed systems consisting of multiple ECUs (Electronic Control Units) interconnected with a broadcast network such as FlexRay. Message authentication is an effective mechanism to prevent attackers from injecting malicious messages into the network. In order to reduce timing interference of message authentication operations on application tasks, hardware coprocessors in the form of either FPGA or ASIC are adopted to offload computation-intensive cryptographic algorithms from the ECU. However, it may not be feasible or desirable to equip every ECU with a hardware coprocessor, as modern vehicles can contain more than one hundred ECUs, and the automotive industry is cost-sensitive. In this paper, we consider the problem of mapping an application task graph onto a FlexRay-based distributed hardware platform, to meet security and deadline requirements while minimizing the number of hardware coprocessors needed in the system. We present a Mixed Integer Linear Programming (MILP) formulation, a divide-and-conquer heuristic algorithm, and a Simulated Annealing algorithm. We evaluate the algorithms with industrial case studies.
Zonghua Gu 0001, Haibo Zeng 0001, Qingling Zhao
IEEE Trans. Parallel Distributed Syst.1
2016 Global Fixed Priority Scheduling with Preemption Threshold: Schedulability Analysis and Stack Size Minimization
abstract
Memory is a limited resource in cost-sensitive, resource-constrained embedded applications. Preemption Threshold Scheduling (PTS) is a well-known technique for reducing the system stack size requirement. We consider Global Fixed Priority Scheduling with Preemption Threshold (gFPPT), as integration of PTS with global Fixed-Priority scheduling on a homogeneous multiprocessor platform, and formulate the optimization problem of minimizing the system stack size requirement while guaranteeing schedulability. We present schedulability analysis, optimization algorithms for priority and preemption threshold assignment, and an ILP formulation for computing system stack size requirement. Performance evaluation shows that the system stack size requirement can be reduced significantly with gFPPT compared to preemptive scheduling.
Chao Wang 0097, Zonghua Gu 0001, Haibo Zeng 0001
IEEE Trans. Parallel Distributed Syst.2
2015 Enhanced partitioned scheduling of Mixed-Criticality Systems on multicore platforms
abstract
Mixed Criticality Systems (MCS) have gained increasing interest in the past few years due to their industrial relevance. When mixed-criticality systems are implemented on multicore architectures, several challenges arise such as the efficient partitioning of these systems. In this paper, we address this issue by presenting a novel mixed-criticality partitioning algorithm, the Dual-Partitioned Mixed-Criticality (DPM) algorithm, that allows limited migration of LO-criticality tasks to enhance the efficiency of the partitioning while maintaining many of the advantages of partitioned systems. Experimental results show that DPM consistently outperforms existing mixed-criticality partitioning algorithms, for example, at utilizations of 0.8 or higher, DPM is able to schedule 17% more systems.
Zaid Al-bayati, Qingling Zhao, Haibo Zeng 0001, Zonghua Gu 0001
ASP-DAC5
2015 Efficient SAT-based application mapping and scheduling on multiprocessor systems for throughput maximization
abstract
Multiprocessor systems are becoming ubiquitous in today's embedded systems design. In this paper, we address the problem of mapping an application represented by a Homogeneous Synchronous Dataflow (HSDF) graph onto a real-time multiprocessor platform with the objective of maximizing total throughput. We propose that the optimal solution to the problem is composed of three components: actor-to-processor mapping, retiming, and actor ordering on each processor. The entire problem is systematically modeled into a SAT problem and solved by a modern SAT solver formally such that the optimal solution can be guaranteed. In order to explore the vast solution space more efficiently, we develop a specific HSDF theory solver based on the special characteristics of the timed HSDF, and integrate it into the general search framework of the SAT solver. The enhanced optimization framework implemented in branch and bound is able to conduct early branch pruning in the search space, and the scalability is thus greatly improved. Extensive performance evaluation on synthetic examples and a case study on the realistic H.264 Video Decoder shows that our technique provides as much as 76.9% throughput improvement, and it is scalable to industry-sized applications.
Weichen Liu 0001, Zonghua Gu 0001, Yaoyao Ye
CASES2
2015 Integration of Cache Partitioning and Preemption Threshold Scheduling to Improve Schedulability of Hard Real-Time Systems
abstract
For preemptive scheduling with shared cache, different tasks may cause interference in the shared cache, leading to Cache-Related Preemption Overhead (CRPD). Cache partitioning is a well-known technique for mitigating unpredictable cache interference in preemptive scheduling, but it reduces cache space available to each task, causing an increase in task execution time. Non-preemptive scheduling algorithms do not incur CRPD, but they generally have poor schedulability. Preemption Threshold Scheduling (PTS) is an effective approach to strike a balance between preemptive and non-preemptive scheduling. We propose integration of cache partitioning and PTS to optimize schedulability on a uniprocessor. We force each subset of tasks assigned the same cache partition to be a non-preemptive group, by assigning the same PT to all tasks in the subset that is equal to or higher than the highest priority of the tasks in that subset. This eliminates CRPD within each cache partition, and helps to improve schedulability. We present an ILP formulation as well as an efficient heuristic algorithm.
Chao Wang 0097, Zonghua Gu 0001, Haibo Zeng 0001
ECRTS2
2015 Endurance-Aware Allocation of Data Variables on NVM-Based Scratchpad Memory in Real-Time Embedded Systems
abstract
Nonvolatile memory (NVM) has many benefits compared to the traditional static RAM, such as improved reliability and reduced power consumption, but it has long write latency and limited write endurance. Scratchpad memory (SPM) is software-managed small on-chip memory for improving system performance and predicability. We consider SPM based on spin-transfer torque RAM, a type of NVM with high performance and good endurance. We present algorithms for allocating data variables to SPM and distribute write activity evenly in the SPM address space, in order to achieve wear-leveling and prolong the lifetime of NVM. We present two optimization algorithms for minimizing system CPU utilization subject to NVM lifetime constraints: 1) an optimal algorithm based on ILP and 2) an efficient heuristic algorithm that can obtain close-to-optimal solutions.
Zonghua Gu 0001, Zili Shao
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2015 Resource Synchronization and Preemption Thresholds Within Mixed-Criticality Scheduling
abstract
In a mixed-criticality system, multiple tasks with different levels of criticality may coexist on the same hardware platform. The scheduling algorithm EDF-VD (Earliest Deadline First with Virtual Deadlines) has been proposed for mixed-criticality systems, which assumes tasks do not share any common resources. We present MC-SRP (Mixed-Criticality Stack Resource Policy), a resource synchronization protocol for EDF-VD, which allows resource sharing among tasks at the same criticality level and guarantees that each task is blocked at most once in each criticality mode. In addition, we present MC-SRPT (MC-SRP with Thresholds) for reducing the application stack size requirement in resource-constrained embedded systems.
Qingling Zhao, Zonghua Gu 0001, Haibo Zeng 0001
ACM Trans. Embed. Comput. Syst.2
2015 WCET-Aware Energy-Efficient Data Allocation on Scratchpad Memory for Real-Time Embedded Systems
abstract
Scratchpad memory (SPM) is a software-managed, small, on-chip form of memory. For real-time embedded systems, worst case execution time (WCET) is more important than average-case performance. We address the problem of allocating program data variables between main memory and SPM to minimize the energy consumption due to data variable accesses, while respecting a given upper bound on a program's WCET. We present an optimal branch-and-bound algorithm; and an efficient heuristic algorithm. Our approach provides a flexible framework for the designer to perform tradeoff analysis between the program WCET and the energy consumption based on application requirements.
Zonghua Gu 0001, Zili Shao
IEEE Trans. Very Large Scale Integr. Syst.2
2014 Partitioned multiprocessor scheduling of mixed-criticality parallel jobs
abstract
Motivated by the increasing trend in embedded systems towards platform integration, there has been an increasing research interest in scheduling mixed-criticality systems. However, most existing efforts have concentrated on scheduling sequential tasks and ignored intra-task parallelism. In this paper, we study the scheduling of mixed-criticality parallel jobs on multiprocessor platforms. We propose a synchronous mixed-criticality job model, where each job consists of segments, each segment having an arbitrary number of parallel threads that synchronize at the end of the segment. A novel MinLoad algorithm is developed to decompose mixed-criticality parallel jobs into mixed-criticality sequential jobs. This decomposition enables us to leverage existing mixed-criticality scheduling algorithms and schedulability analysis to the multiprocessor scheduling of mixed-criticality parallel jobs. In addition, our MinLoad job decomposition algorithm is designed to make the decomposed mixed-criticality sequential tasks easier to schedule, and thus requires smaller-sized multiprocessor platforms for the mixed-criticality systems.
Guangdong Liu, Ying Lu 0002, Shige Wang, Zonghua Gu 0001
RTCSA4
2013 PT-AMC: integrating preemption thresholds into mixed-criticality scheduling
abstract
Mixed-Criticality Scheduling (MCS) is an effective approach to addressing diverse certification requirements of safety-critical systems that integrate multiple subsystems with different levels of criticality. Preemption Threshold Scheduling (PTS) is a well-known technique for controlling the degree of preemption, ranging from fully-preemptive to fully-non-preemptive scheduling. We present schedulability analysis algorithms to enable integration of PTS with MCS, in order to bring the rich benefits of PTS into MCS, including minimizing the application stack space requirement, reducing the number of runtime task preemptions, and improving schedulability.
Qingling Zhao, Zonghua Gu 0001, Haibo Zeng 0001
DATE2
2013 Integration of resource synchronization and preemption-thresholds into EDF-based mixed-criticality scheduling algorithm
abstract
In mixed-criticality systems, multiple subsystems with different levels of criticality may co-exist on the same hardware platform. Many scheduling algorithms have been proposed to achieve certification at multiple levels of criticality. However, current MCS algorithms and analysis techniques generally assume tasks are independent, i.e., they do not share data that need to be protected with synchronization mechanisms like mutexes or semaphores. In this paper, we address this limitation by presenting an extension to the Stack Resource Protocol (SRP), called Mixed-Criticality-SRP (MC-SRP). Moreover, preemption-threshold scheduling is a well-known technique for reducing stack space size and enhance schedulability in resource-constrained embedded systems.We also present the integration of preemption-thresholds into EDF-based mixed-criticality scheduling (MCS) algorithms, and develop the schedulability analysis methods to such systems.
Qingling Zhao, Zonghua Gu 0001, Haibo Zeng 0001
RTCSA2
2012 Online optimization for scheduling preemptable tasks on IaaS cloud systems
Meikang Qiu, Zhong Ming 0001, Gang Quan, Xiao Qin 0001, Zonghua Gu 0001
J. Parallel Distributed Comput.6
2011 Schedulability analysis for non-preemptive fixed-priority multiprocessor scheduling
Nan Guan, Wang Yi 0001, Qingxu Deng, Zonghua Gu 0001, Ge Yu 0001
J. Syst. Archit.4
2011 Satisfiability Modulo Graph Theory for Task Mapping and Scheduling on Multiprocessor Systems
abstract
Task graph scheduling on multiprocessor systems is a representative multiprocessor scheduling problem. A solution to this problem consists of the mapping of tasks to processors and the scheduling of tasks on each processor. Optimal solution can be obtained by exploring the entire design space of all possible mapping and scheduling choices. Since the problem is NP-hard, scalability becomes the main concern in solving the problem optimally. In this paper, a SAT-based optimization framework is proposed to address this problem, in which SAT solver is enhanced by integrating with a scheduling analysis tool in a branch and bound manner to prune the solution space efficiently. Performance evaluation results show that our technique has average performance improvement in more than an order of magnitude compared to state-of-the-art techniques. We further build a cycle-accurate network-on-chip simulator based on SystemC to verify the effectiveness of the proposed technique on realistic multiprocessor systems.
Weichen Liu 0001, Zonghua Gu 0001, Jiang Xu 0001, Xiaowen Wu, Yaoyao Ye
IEEE Trans. Parallel Distributed Syst.2
2010 Task Allocation and Optimization of Distributed Embedded Systems with Simulated Annealing and Geometric Programming
abstract
We consider the task model of periodic tasks running on a network of processor nodes connected by a bus based on the time-triggered protocol, an industry-standard bus protocol designed for safety-critical automotive and avionics distributed embedded systems, and present an integrated optimization framework that jointly considers one or more of the following attributes: task-to-processor allocation, task priority assignment, task period assignment and bus access configuration. We adopt a hierarchical optimization framework, where each possible task allocation and priority assignment is treated as one top-level coarse-grained state, which may contain many lower-level fine-grained states defined by different task period assignments and bus access configurations. Simulated annealing is used to explore the top-level states, which calls a geometric programming solver as a subroutine to explore the lower-level states contained within a given top-level state. Performance evaluation shows that our framework has good performance in terms of solution quality and scalability.
Xiuqiang He 0001, Zonghua Gu 0001, Yongxin Zhu 0001
Comput. J.2
2010 Implementing a Thermal-Aware Scheduler in Linux Kernel on a Multi-Core Processor
abstract
As power dissipation causes thermal issues in cooling costs, lifetime and reliability, thermal management has become an important issue in today's OS and processor design. Early OS-level thermal management schemes were proposed and evaluated mainly with simulators or analytical models. In this paper, we implement a thermal-aware round-robin scheduling algorithm in the Linux kernel, and compare its performance with the ‘Heat-and-Run’ algorithm and the default Linux baseline scheduler on an Intel Core 2 Duo processor using representative benchmarks from SPEC2000, MiBench and NetBench. Our results indicate that the current Linux scheduler can easily be enhanced with thermal-awareness to show improved performance in terms of both the on-chip temperature condition and application throughput.
Yongxin Zhu 0001, Jingwei Ye, Zonghua Gu 0001
Comput. J.5
2010 Online adaptive utilization control for real-time embedded multiprocessor systems
Jianguo Yao 0002, Xue (Steve) Liu, Zonghua Gu 0001, Jian Li 0021
J. Syst. Archit.3
2010 PAUC: Power-Aware Utilization Control in Distributed Real-Time Systems
abstract
CPU utilization control has recently been demonstrated to be an effective way of meeting end-to-end deadlines for distributed real-time systems running in unpredictable environments. However, current research on utilization control focuses exclusively on task rate adaptation, which cannot effectively handle rate saturation and discrete task rates. Since the CPU utilization contributed by a real-time periodic task is determined by both its rate and execution time, CPU frequency scaling can be used to adapt task execution times for power-efficient utilization control. In this paper, we present PAUC, a two-layer coordinated CPU utilization control architecture. The primary control loop uses frequency scaling to locally control the CPU utilization of each processor, while the secondary control loop adopts rate adaptation to control the utilizations of all the processors at the cluster level on a finer timescale. Both the two control loops are designed and coordinated based on well-established control theory for theoretically guaranteed control accuracy and system stability. Empirical results on a physical testbed demonstrate that our control solution outperforms a state-of-the-art utilization control algorithm by having more accurate control and less power consumption. Extensive simulation results also show that our solution can significantly improve the feasibility of utilization control.
Xue (Steve) Liu, Zonghua Gu 0001
IEEE Trans. Ind. Informatics4
2010 Hardware/software partitioning and pipelined scheduling on runtime reconfigurable FPGAs
abstract
FPGAs are widely used in today's embedded systems design due to their low cost, high performance, and reconfigurability. Partially RunTime-Reconfigurable (PRTR) FPGAs, such as Virtex-2 Pro and Virtex-4 from Xilinx, allow part of the FPGA area to be reconfigured while the remainder continues to operate without interruption, so that HW tasks can be placed and removed dynamically at runtime. We address two problems related to HW task scheduling on PRTR FPGAs: (1) HW/SW partitioning. Given an application in the form of a task graph with known execution times on the HW (FPGA) and SW (CPU), and known area sizes on the FPGA, find an valid allocation of tasks to either HW or SW and a static schedule with the optimization objective of minimizing the total schedule length (makespan). (2) Pipelined scheduling. Given an input task graph, construct a pipelined schedule on a PRTR FPGA with the goal of maximizing system throughput while meeting a given end-to-end deadline. Both problems are NP-hard. Satisfiability Modulo Theories (SMT) is an extension to SAT by adding the ability to handle arithmetic and other decidable theories. We use the SMT solver Yices with Linear Integer Arithmetic (LIA) theory as the optimization engine for solving the two scheduling problems. In addition, we present an efficient heuristic algorithm based on kernel recognition for the pipelined scheduling problem, a technique borrowed from SW pipelining, to overcome the scalability problem of the SMT-based optimal solution technique.
Mingxuan Yuan, Zonghua Gu 0001, Xiuqiang He 0001, Xue (Steve) Liu, Lei Jiang 0001
ACM Trans. Design Autom. Electr. Syst.2
2009 Improving scalability of model-checking for minimizing buffer requirements of synchronous dataflow graphs
abstract
Synchronous dataflow (SDF) is a well-known model of computation for dataflow-oriented applications such as embedded systems for signal processing and multimedia. It is important to minimize the buffer size requirements of applications generated from SDF graphs, since memory space is often a scarce resource in these systems due to cost or power consumption constraints. Some authors have proposed to use model-checking for finding the minimum buffer size requirements, but the scalability of model-checking is limited by state space explosion. In this paper, we present several techniques for reducing state space size and improving scalability of model-checking by exploiting problem-specific properties of SDF graphs.
Nan Guan, Zonghua Gu 0001, Wang Yi 0001, Ge Yu 0001
ASP-DAC2
2009 Power-Aware CPU Utilization Control for Distributed Real-Time Systems
abstract
CPU utilization control has recently been demonstrated to be an effective way of meeting end-to-end deadlines for distributed real-time systems running in unpredictable environments. However, current research on utilization control focuses exclusively on task rate adaptation, which cannot effectively handle rate saturation and discrete task rates. Since the CPU utilization contributed by a real-time periodic task is determined by both its rate and execution time, CPU frequency scaling can be used to adapt task execution times for power-efficient utilization control. In this paper, we present a two-layer coordinated CPU utilization control architecture. The primary control loop uses frequency scaling to locally control the CPU utilization of each processor, while the secondary control loop adopts rate adaptation to control the utilizations of all the processors at the cluster level on a finer timescale. Both the two control loops are designed and coordinated based on well-established control theory for theoretically guaranteed control accuracy and system stability. Empirical results on a physical testbed demonstrate that our control solution outperforms a state-of-the-art utilization control algorithm by having more accurate control and less power consumption.
Xue (Steve) Liu, Zonghua Gu 0001
IEEE Real-Time and Embedded Technology and Applications Symposium4
2008 Performance Comparison of Techniques on Static Path Analysis of WCET
abstract
Static path analysis is a key process of Worst Case Execution Time (WCET) estimation, the objective of which is to find the execution path that has the largest execution time. Currently, there is an argument in the research community whether model checking is another good solution for WCET analysis, besides ILP. To our knowledge, no paper so far has addressed this argument with real performance data. In this paper, we implement both ILP and model checking for static path analysis of WCET, and the experiment results show that ILP yields very good performance, while model checking only works well for simple programs, and it is inclined to scalability problems when dealing with programs that have complex structures and large loop counts.
Mingsong Lv, Zonghua Gu 0001, Nan Guan, Qingxu Deng, Ge Yu 0001
EUC (1)2
2008 Schedulability Analysis of Global Fixed-Priority or EDF Multiprocessor Scheduling with Symbolic Model-Checking
abstract
As Moore's law comes to an end, multi-processor (MP) systems are becoming increasingly important in embedded systems design, hence real-time schedulability analysis for MP systems has become an important research topic. In this paper, we present an exact method for schedulability analysis of global multiprocessor scheduling with either fixed-priority (FP) or earliest-deadline-first (EDF) algorithms using the model-checker NuSMV. Compared to safe but pessimistic schedulability tests based on processor utilization bounds, model-checking can provide an exact answer to the schedulability of a taskset, as well as quantitative information on each task's best-case and worst- case response times.
Nan Guan, Zonghua Gu 0001, Mingsong Lv, Qingxu Deng, Ge Yu 0001
ISORC2
2008 Hardware/Software Partitioning and Static Task Scheduling on Runtime Reconfigurable FPGAs using a SMT Solver
abstract
FPGAs are often used together with a CPU as hardware accelerators. A runtime reconfigurable FPGA allows part of the FPGA area to be reconfigured while the remainder continues to operate without interruption, so that hardware tasks can be placed and removed dynamically at runtime. In this paper, we formulate and solve the problem of optimal hardware/software partitioning and static task scheduling for a hybrid FPGA/CPU device, with the optimization objective of minimizing the total schedule length, in the framework of satisfiability modulo theories (SMT) with linear integer arithmetic.
Mingxuan Yuan, Xiuqiang He 0001, Zonghua Gu 0001
IEEE Real-Time and Embedded Technology and Applications Symposium3
2008 New Schedulability Test Conditions for Non-preemptive Scheduling on Multiprocessor Platforms
abstract
We study the schedulability analysis problem for nonpreemptive scheduling algorithms on multiprocessors. To our best knowledge, the only known work on this problem is the test condition proposed by Baruah for non-preemptive EDF scheduling, which will reject a task set with arbitrarily low utilization if it contains a task whose execution time is equal or greater than the minimal relative deadline among all tasks. In this paper, we firstly derive a linear-time test condition which avoids the problem mentioned above, by building upon previous work for preemptive multiprocessor scheduling. This test condition works on not only non-preemptive EDF, but also any other work-conserving non-preemptive scheduling algorithms. Then we improve the analysis and present test conditions of pseudo-polynomial time-complexity for Non-preemptive Earliest Deadline First scheduling and Non-preemptive Fixed Priority scheduling respectively. Experiments with randomly generated task sets show that our proposed test conditions, especially the improved test conditions, have significant performance improvements compared with [BAR-EDFnp].
Nan Guan, Wang Yi 0001, Zonghua Gu 0001, Qingxu Deng, Ge Yu 0001
RTSS3
2008 Efficient SAT-Based Mapping and Scheduling of Homogeneous Synchronous Dataflow Graphs for Throughput Optimization
abstract
As Moore's law comes to an end, multiprocessor systems are becoming ubiquitous in today's embedded systems design. In this paper, we address the problem of mapping a homogeneous synchronous dataflow (HSDF) graph onto a multiprocessor platform with the objective of maximizing system throughput. We present two optimization approaches based on branch-and-bound and SAT-solving to explore the design space of all possible actor-to-processor mappings and static order schedules on each processor. In the logic-based benders decomposition (LBBD) approach, we decompose the problem into a master problem of finding a feasible actor mapping and scheduling, and a sub-problem of deadlock-checking and throughput computation. In the integrated approach, we integrate branch-and-bound search into the SAT engine to achieve more effective search tree pruning and better scalability. Performance evaluation shows that the integrated approach outperforms the LBBD approach by a large margin.
Weichen Liu 0001, Mingxuan Yuan, Xiuqiang He 0001, Zonghua Gu 0001, Xue (Steve) Liu
RTSS4
2008 Optimal Sampling Rate Assignment with Dynamic Route Selection for Real-Time Wireless Sensor Networks
abstract
The allocation of computation and communication resources in a manner that optimizes aggregate system performance is a crucial aspect of system management. Wireless sensor network poses new challenges due to the resource constraints and real-time requirements. Existing work has dealt with the real-time sampling rate assignment problem, under single processor case and network case with static routing environment. For wireless sensor networks, in order to achieve better overall network performance, routing should be considered together with the rate assignments of individual flows. In this paper, we address the problem of optimizing sampling rates with dynamic route selection for wireless sensor networks. We model the problem as a constrained optimization problem and solve it under the network utility maximization framework. Based on the primal-dual method and dual decomposition technique, we design a distributed algorithm that achieves the optimal global network utility considering both dynamic route decision and rate assignment. Extensive simulations have been conducted to demonstrate the efficiency and efficacy of our proposed solutions.
Weihuan Shu, Xue (Steve) Liu, Zonghua Gu 0001, Sathish Gopalakrishnan
RTSS3
2008 A Hierarchical Framework for Design Space Exploration and Optimization of TTP-Based Distributed Embedded Systems
abstract
Time-triggered protocol (TTP) is a time-division multiple access (TDMA)-based bus protocol designed for use in safety-critical avionics and automotive distributed embedded systems. Design space exploration (DSE) for TTP-based distributed embedded system involves searching through a vast design space of possible task-to-CPU mappings, task/message schedules and bus access configurations to achieve certain design objectives. In this paper, we present an efficient two-level hierarchical DSE framework for TTP-based distributed embedded systems, with the objective of minimizing the total bus utilization while meeting an end-to-end deadline constraint. Logic-based Benders decomposition (LBBD) is used to divide the problem into a master problem of mapping tasks to CPU nodes to minimize the total bus utilization, solved with a satisfiability modulo theories (SMT) solver, and a subproblem of finding a feasible solution of bus access configuration and task/message schedule under an end-to-end deadline constraint for a given task-to-CPU mapping, solved with a constraint programming (CP) solver. Performance evaluation results show that our approach is scalable to problems with realistic size.
Xiuqiang He 0001, Mingxuan Yuan, Zonghua Gu 0001
IEEE Trans. Ind. Informatics3
2008 Schedulability analysis of preemptive and nonpreemptive EDF on partial runtime-reconfigurable FPGAs
abstract
Field Programmable Gate Arrays (FPGAs) are very popular in today's embedded systems design, and Partial Runtime-Reconfigurable (PRTR) FPGAs allow HW tasks to be placed and removed dynamically at runtime. Hardware task scheduling on PRTR FPGAs brings many challenging issues to traditional real-time scheduling theory, which have not been adequately addressed by the research community compared to software task scheduling on CPUs. In this article, we consider the schedulability analysis problem of HW task scheduling on PRPR FPGAs. We derive utilization bounds for several variants of global preemptive/nonpreemptive EDF scheduling, and compare the performance of different utilization bound tests.
Nan Guan, Qingxu Deng, Zonghua Gu 0001, Wenyao Xu, Ge Yu 0001
ACM Trans. Design Autom. Electr. Syst.3
2007 Optimization of Static Task and Bus Access Schedules for Time-Triggered Distributed Embedded Systems with Model-Checking
abstract
Time-Triggered Protocol for the bus and static task scheduling for the CPU are widely used in safety-critical distributed embedded systems. Researchers have presented efficient heuristic algorithms to jointly optimize static task and bus access schedules. In this paper, we use the model checker SPIN to provide a flexible and configurable technique for obtaining provably optimal solutions, and evaluate its performance tradeoffs compared to heuristic algorithms.
Zonghua Gu 0001, Xiuqiang He 0001, Mingxuan Yuan
DAC1
2007 An efficient algorithm for online management of 2D area of partially reconfigurable FPGAs
abstract
Partially runtime-reconfigurable (PRTR) FPGAs allow hardware tasks to be placed and removed dynamically at runtime. We present an efficient algorithm for finding the complete set of maximal empty rectangles on a 2D PRTR FPGA, which is useful for online placement and scheduling of HW tasks. The algorithm is incremental and only updates the local region affected by each task addition or removal event. We use simulation experiments to evaluate its performance and compare to related work
Qingxu Deng, Xiuqiang He 0001, Zonghua Gu 0001
DATE4
2007 Teaching embedded systems software: The HKUST experience
abstract
Recent trends in embedded systems indicate a growing importance of the software component of these systems, and the concomitant increase in the importance of formal design and development of embedded software. In this paper we share our experience with designing and offering courses related to embedded software in the Department of Computer Science and Engineering at the Hong Kong University of Science and Technology. We give a detailed overview of the courses, the structure and the teaching/learning methodology followed in our courses, followed by some discussion and reflections on the courses.
Jogesh K. Muppala, Zonghua Gu 0001, Shing-Chi Cheung
ICPADS2
2007 Improved Schedulability Analysis of EDF Scheduling on Reconfigurable Hardware Devices
abstract
Reconfigurable devices, such as field programmable gate arrays (FPGAs), are very popular in today's embedded systems design due to their low-cost, high-performance and flexibility. Partially runtime-reconfigurable (PRTR) FPGAs allow hardware tasks to be placed and removed dynamically at runtime. Hardware task scheduling on PRTR FPGAs brings many challenging issues to traditional real-time scheduling theory, which have not been adequately addressed by the research community compared to software task scheduling on CPUs. In this paper, we consider the schedulability analysis problem of HW task scheduling on PRPR FPGAs. We derive utilization bound tests for two variants of global EDF scheduling, and use synthetic tasksets to compare performance of the tests to existing work and simulation results.
Nan Guan, Zonghua Gu 0001, Qingxu Deng, Weichen Liu 0001, Ge Yu 0001
IPDPS2
2007 An Efficient Algorithm for Online Soft Real-Time Task Placement on Reconfigurable Hardware Devices
abstract
Reconfigurable devices such as field programmable gate arrays (FPGAs) are very popular in today's embedded systems (design due to their low-cost, high-performance and flexibility. Partially runtime-reconfigurable (PRTR) FPGAs allow hardware tasks to be placed and removed dynamically at runtime. Hardware task scheduling on PRTR FPGAs brings many challenging issues to traditional real-time scheduling theory, which have not been adequately addressed by the real-time research community compared to software task scheduling on CPUs. In this paper, we present an efficient online task placement algorithm for minimizing fragmentation on PRTR FPGAs. First, we present a novel 2D area fragmentation metric that takes into account probability distribution of sizes of future task arrivals; second, we take into the time axis to obtain a 3D fragmentation metric. Simulation experiments indicate that our techniques result in low ratio of task rejection and high FPGA utilization compared to existing techniques
Zonghua Gu 0001, Weichen Liu 0001, Qingxu Deng
ISORC2
2007 Optimal Static Task Scheduling on Reconfigurable Hardware Devices Using Model-Checking
abstract
Real-time scheduling for FPGAs presents unique challenges to traditional real-time scheduling theory, since it is similar to, but more general than multi-processor scheduling. In his paper, we address two problems of static task scheduling on a partially runtime reconfigurable FPGA: finding an optimal static schedule for a task graph with the optimization objective of minimizing the total schedule length, and finding a feasible static schedule for a set of periodic tasks within a hyper-period with the objective of meeting all deadlines. We model the multi-tasking system with timed automata and use reachability analysis of the UPPAAL model-checker to explore the design space and find an optimal or feasible schedule
Zonghua Gu 0001, Mingxuan Yuan, Xiuqiang He 0001
IEEE Real-Time and Embedded Technology and Applications Symposium1
2007 Static Scheduling and Software Synthesis for Dataflow Graphs with Symbolic Model-Checking
abstract
In this paper, we address the problem of static scheduling and software synthesis for dataflow graphs with the symbolic model-checker NuSMV using a two-step process: first use model-checking to obtain a static schedule with the objective of minimizing the data buffer size, then synthesize efficient code from the static schedule with the objective of minimizing code size and performance overheads due to runtime dynamic decisions. We show the effectiveness of these techniques using a number of digital signal processing examples.
Zonghua Gu 0001, Mingxuan Yuan, Nan Guan, Mingsong Lv, Xiuqiang He 0001, Qingxu Deng, Ge Yu 0001
RTSS1
2007 QoS-Optimized Integration of Embedded Software Components with Multiple Modes of Execution
Zonghua Gu 0001, Qingxu Deng
SEKE1
2005 Timing Analysis of Distributed End-to-End Task Graphs with Model-Checking
Zonghua Gu 0001
EUC1
2005 Model-Checking of Component-Based Event-Driven Real-Time Embedded Software
abstract
As complexity of real-time embedded software grows, it is desirable to use formal verification techniques to achieve a high level of assurance. We discuss application of model-checking to verify system-level concurrency properties of component-based real-time embedded software based on CORBA event service, using avionics mission computing software as an application example. We use the process algebra FSP to formalize specification of software components and system architecture, previously only available in the form of natural language and prone to misinterpretation and misunderstanding, and use model-checking to verify system-level concurrency properties. We also discuss effective techniques for coping with the state-space explosion problem by exploiting application domain semantics. We have applied our analysis techniques to realistic application scenarios provided by our industry partner to demonstrate their utility and power.
Zonghua Gu 0001, Kang G. Shin
ISORC1
2005 Synthesis of Real-Time Implementations from Component-Based Software Models
abstract
Component-based software development is an effective technique for tackling the increasing complexity of large-scale embedded software systems. After building a logical software model, the designer must make design decisions, including choosing a multi-threading strategy and assigning priorities to threads, to ensure that the final implementation on the target execution platform satisfies non-functional requirements. Code generators for software design tools produce functional code, but typically ignore concurrency and timing issues. In this paper, we describe techniques for real-time scheduling and design-space exploration and optimization, with the goal of helping the designer synthesize efficient real-time implementations from component-based software models. Experimental evaluation shows that our techniques yield high-quality implementations with reasonable running time of the optimization algorithm.
Zonghua Gu 0001, Kang G. Shin
RTSS1
2003 An Integrated Approach to Modeling and Analysis of Embedded Real-Time Systems Based on Timed Petri Net
abstract
In computer-based control systems, embedded software is taking over what mechanical and dedicated electronic systems used to do, that is, to engage and control the physical world, interacting directly with sensors and actuators. Therefore, software running on a digital processor is tightly-coupled with its surrounding physical environment. We propose an integrated approach based on Timed Petri-Nets for modeling and analysis of embedded real-time systems where real-time scheduling behavior of the controller software is explicitly represented at the model-level, together with the physical environment that it interacts with. This enables the designer to have an integrated view of the entire system while analyzing the system and making design decisions. We also describe a syntax-directed, automated translation procedure from Timed Petri-Nets to Timed Automata, thus enabling the use of model checkers such as UPPAAL for analysis purposes. We consider the railroad crossing problem as an application example, and evaluate alternatives for controller implementation on either single-processor or distributed multi-processor platforms based on the integrated approach.
Zonghua Gu 0001, Kang G. Shin
ICDCS1
2003 Algorithms for effective variable bit rate traffic smoothing
abstract
The transfer of prerecorded, compressed variable-bit-rate video requires multimedia services to support large fluctuations in bandwidth requirements on multiple time scales. Bandwidth smoothing techniques can reduce the burstiness of a VBR (variable bit rate) stream by transmitting data at a series of fixed rates, simplifying the allocation of resources in video servers and the communication network. The RCBR (renegotiated constant bit rate) service model seems ideally suited for smoothed VBR traffic which is piece-wise CBR. Zhimei Jiang (see Proc. IEEE Infocomm 98, p.676, 1998) proposed a dynamic programming algorithm to compute the optimal renegotiation schedule given the relative cost of renegotiation and client buffer size. We show that the schedule produced by his algorithm has high peak rates and frequent renegotiations. We propose another algorithm that computes a renegotiation schedule that has a slightly higher cost than the optimal schedule, but has other desirable properties, such as lower peak rate and lower frequency of renegotiations. We also consider proxy-based online smoothing, and propose an adaptive heuristic algorithm to generate renegotiation schedules at runtime without knowledge of future frame size information. We compare the schedule computed by the algorithm to the optimal schedule computed with full knowledge of future frame sizes.
Zonghua Gu 0001, Kang G. Shin
IPCCC1
2003 An End-to-End Tool Chain for Multi-View Modeling and Analysis of Avionics Mission Computing Software
abstract
We present an end-to-end tool-chain for model-based design and analysis of component-based embedded real-time software, with avionics mission computing as an application domain. The tool-chain covers the entire system development lifecycle including modeling, analysis, code generation, and run-time instrumentation. Emphasis is placed on integration of tools developed by multiple institutions via standardized interface format definitions in XML. By capturing all relevant information explicitly in models at the design level, and performing analysis that provides insight into non-functional aspects of the system, we can raise the level of abstraction for the designer, and facilitate rapid system prototyping.
Zonghua Gu 0001, Shige Wang, Sharath Kodase, Kang G. Shin
RTSS1