Cong Liu 0005

dblp:95/6404-5 · DBLP profile ↗
← Back
112ranked-venue papers
13as first author
43since 2021 · last 2026
0000-0003-1190-522XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 29 · 4 first-author · 8 since 2021Systems, architecture and hardware · 23 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 21 · 17 since 2021Computer networks · 16 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 11 since 2021Software engineering, systems software and programming languages · 11 · 6 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Policy Search, Retrieval, and Composition via Task Similarity in Collaborative Agentic Systems
abstract
Agentic AI aims to create systems that set their own goals, adapt proactively to change, and refine behavior through continuous experience. Recent advances suggest that, when facing multiple and unforeseen tasks, agents could benefit from sharing machine-learned knowledge and reusing policies that have already been fully or partially learned by other agents. However, how to query, select, and retrieve policies from a pool of agents, and how to integrate such policies remains a largely unexplored area. This study explores how an agent decides what knowledge to select, from whom, and when and how to integrate it in its own policy in order to accelerate its own learning. The proposed algorithm, Modular Sharing and Composition in Collective Learning (MOSAIC), improves learning in agentic collectives by combining (1) knowledge selection using performance signals and cosine similarity on Wasserstein task embeddings, (2) modular and transferable neural representations via masks, and (3) policy integration, composition and fine-tuning. MOSAIC outperforms isolated learners and global sharing approaches in both learning speed and overall performance, and in some cases solves tasks that isolated agents cannot. The results also demonstrate that selective, goal-driven reuse leads to less susceptibility to task interference. We also observe the emergence of self-organization, where agents solving simpler tasks accelerate the learning of harder ones through shared knowledge.
Saptarshi Nath, Christos Peridis, Eseoghene Benjamin, Soheil Kolouri, Peter Kinnell, Zexin Li 0001, Cong Liu 0005, Shirin Dora, Andrea Soltoggio
AAAI8
2025 Lemix: Unified Scheduling for Llm Training and Inference on Multi-Gpu Systems
abstract
Modern deployment of large language models (LLMs) frequently involves both inference serving and continuous retraining to stay aligned with evolving data and user feedback. Common practices separate these workloads onto distinct servers in isolated phases, causing substantial inefficiencies (e.g., GPU idleness) and delayed adaptation to new data in distributed settings. Our empirical analysis reveals that these inefficiencies stem from dynamic request arrivals during serving and workload heterogeneity in pipeline-parallel training. To address these challenges, we propose LeMix, a system for co-locating and managing concurrent LLM serving and training workloads. LeMix integrates offline profiling, execution prediction mechanisms, and runtime scheduling to dynamically adapt resource allocation based on workload characteristics and system conditions. By understanding task-specific behaviors and co-execution interference across shared nodes, LeMix improves utilization and serving quality without compromising serving responsiveness. Our evaluation shows that LeMix improves throughput by up to$3.53 \times$, reduces inference loss by up to$0.61 \times$, and delivers up to$2.12 \times$higher response time SLO attainment over traditional separate setups. To our knowledge, this is the first work to uncover and exploit the opportunities of joint LLM inference and training, paving the way for more resource-efficient deployment of LLMs in production environments.
Yufei Li 0001, Zexin Li 0001, Yinglun Zhu, Cong Liu 0005
RTSS4
2024 Safety Alignment in NLP Tasks: Weakly Aligned Summarization as an In-Context Attack
abstract
Recent developments in balancing the usefulness and safety of Large Language Models (LLMs) have raised a critical question: Are mainstream NLP tasks adequately aligned with safety consideration?Our study, focusing on safety-sensitive documents obtained through adversarial attacks, reveals significant disparities in the safety alignment of various NLP tasks.For instance, LLMs can effectively summarize malicious long documents but often refuse to translate them.This discrepancy highlights a previously unidentified vulnerability: attacks exploiting tasks with weaker safety alignment, like summarization, can potentially compromise the integrity of tasks traditionally deemed more robust, such as translation and question-answering (QA).Moreover, the concurrent use of multiple NLP tasks with lesser safety alignment increases the risk of LLMs inadvertently processing harmful content.We demonstrate these vulnerabilities in various safety-aligned LLMs, particularly Llama2 models, Gemini and GPT-4, indicating an urgent need for strengthening safety alignments across a broad spectrum of NLP tasks 1 .
Yu Fu 0009, Yufei Li 0001, Cong Liu 0005, Yue Dong 0002
ACL (1)4
2024 GCAPS: GPU Context-Aware Preemptive Priority-Based Scheduling for Real-Time Tasks
abstract
Scheduling real-time tasks that utilize GPUs with analyzable guarantees poses a significant challenge due to the intricate interaction between CPU and GPU resources, as well as the complex GPU hardware and software stack. While much research has been conducted in the real-time research community, several limitations persist, including the absence or limited availability of GPU-level preemption, extended blocking times, and/or the need for extensive modifications to program code. In this paper, we propose GCAPS, a GPU Context-Aware Preemptive Scheduling approach for real-time GPU tasks. Our approach exerts control over GPU context scheduling at the device driver level and enables preemption of GPU execution based on task priorities by simply adding one-line macros to GPU segment boundaries. In addition, we provide a comprehensive response time analysis of GPU-using tasks for both our proposed approach as well as the default Nvidia GPU driver scheduling that follows a work-conserving round-robin policy. Through empirical evaluations and case studies, we demonstrate the effectiveness of the proposed approaches in improving taskset schedulability and response time. The results highlight significant improvements over prior work as well as the default scheduling approach, with up to 40% higher schedulability, while also achieving predictable worst-case behavior on Nvidia Jetson embedded platforms.
Yidi Wang 0001, Cong Liu 0005, Daniel Wong 0001, Hyoseung Kim 0001
ECRTS2
2024 DuoJoule: Accurate On-Device Deep Reinforcement Learning for Energy and Timeliness
abstract
Deep Reinforcement Learning (DRL) is critical for autonomous systems to continuously learn and adapt in dynamic environments. However, frequent retraining in DRL leads to high energy consumption, posing significant challenges for mobile and battery-dependent robotic systems. Co-optimizing energy, latency, and algorithm performance is essential for efficient on-device DRL. Current approaches either focus on traditional DNNs like CNNs or target only two out of the three dimensions, rather than addressing all three simultaneously. This paper introduces DuoJoule, a comprehensive framework designed to address the unique challenges of DRL workloads by meeting latency deadlines and adhering to energy budgets while maximizing algorithm performance through both application and system-level configurations. DuoJoule dynamically coordinates adjustments in DRL algorithm parameters and system frequency settings using Dynamic Voltage and Frequency Scaling (DVFS). A key innovation of DuoJoule is its runtime metric tracker, which assesses system status against target budgets and calculates a universal efficiency score. This enables rapid and adaptive tuning at runtime, balancing energy efficiency, latency, and algorithm performance. Extensive evaluation using benchmarks along with a realistic autonomous driving case study demonstrates DuoJoule’s versatile cross-platform efficiency, practicality in real-world scenarios, adaptivity to varying constraints, and low runtime overhead evaluated on two widely used autonomous embedded platforms. Empirical results show that DuoJoule consistently meets latency and energy targets while maintaining near-optimal performance, showcasing its effectiveness in managing the complex trade-off space of on-device DRL.
Soheil Shirvani, Aritra Samanta, Zexin Li 0001, Cong Liu 0005
RTSS4
2024 BOXR: Body and head motion Optimization framework for eXtended Reality
abstract
The emergence of standalone Extended Reality (XR) systems has enhanced user mobility, accommodating both subtle, frequent head motions and substantial, less frequent body motions. However, the pervasively used Motion-to-Display (M2D) latency metric, which measures the delay between the most recent motion and its corresponding display update, only accounts for head motions. This oversight can leave users prone to motion sickness if significant body motion is involved. Although existing methods optimize M2D latency through asynchronous task scheduling and reprojection methods, they introduce challenges like resource contention between tasks and outdated pose data. These challenges are further complicated by user motion dynamics and scene changes during runtime. To address these issues, we for the first time introduce the Camera-to-Display (C2D) latency metric, which captures the delay caused by body motions, and present BOXR, a framework designed to co-optimize both body and head motion delays within an XR system. BOXR enhances the coordination between M2D and C2D latencies by efficiently scheduling tasks to avoid contentions while maintaining an up-to-date pose in the output frame. Moreover, BOXR incorporates a motion-driven visual inertial odometer to adjust to user motion dynamics and employs scene-dependent foveated rendering to manage changes in the scene effectively. Our evaluations show that BOXR significantly outperforms state-of-the-art solutions in 11 EuRoC MAV datasets across 4 XR applications across 3 hardware platforms. In controlled motion and scene settings, BOXR reduces M2D and C2D latencies by up to $63 \%$ and $27 \%$, respectively and increases frame rate by up to $43 \%$. In practical deployments, BOXR achieves substantial reductions in real-world scenarios-up to $42 \%$ in M2D latency and $31 \%$ in C2D latency-while maintaining remarkably low miss rates of only $1.6 \%$ for M2D requirements and $\mathbf{1. 0 \%}$ for C2D requirements.
Zexin Li 0001, Hyoseung Kim 0001, Cong Liu 0005
RTSS4
2024 Transferable Adversarial Attacks Against ASR
abstract
Given the extensive research and real-world applications of automatic speech recognition (ASR), ensuring the robustness of ASR models against minor input perturbations becomes a crucial consideration for maintaining their effectiveness in real-time scenarios. Previous explorations into ASR model robustness have predominantly revolved around evaluating accuracy on white-box settings with full access to ASR models. Nevertheless, full ASR model details are often not available in real-world applications. Therefore, evaluating the robustness of black-box ASR models is essential for a comprehensive understanding of ASR model resilience. In this regard, we thoroughly study the vulnerability of practical black-box attacks in cutting-edge ASR models and propose to employ two advanced time-domain-based transferable attacks alongside our differentiable feature extractor. We also propose a speech-aware gradient optimization approach (SAGO) for ASR, which forces mistranscription with minimal impact on human imperceptibility through voice activity detection rule and a speech-aware gradient-oriented optimizer. Our comprehensive experimental results reveal performance enhancements compared to baseline approaches across five models on two databases.
Xiaoxue Gao, Zexin Li 0001, Yiming Chen 0010, Cong Liu 0005, Haizhou Li 0001
IEEE Signal Process. Lett.4
2024 MII: A Multifaceted Framework for Intermittence-Aware Inference and Scheduling
abstract
The concurrent execution of deep neural networks (DNNs) inference tasks on the intermittently-powered batteryless devices (IPDs) has recently garnered much attention due to its potential in a broad range of smart sensing applications. While the checkpointing mechanisms (CMs) provided by the state-of-the-art make this possible, scheduling inference tasks on IPDs is still a complex problem due to significant performance variations across the DNN layers and CM choices. This complexity is further accentuated by dynamic environmental conditions and inherent resource constraints of IPDs. To tackle these challenges, we present MII, a framework designed for the intermittence-aware inference and scheduling on IPDs. MII formulates the shutdown and live time functions of an IPD from profiling the data, which our offline intermittence-aware search scheme uses to find the optimal layer-wise CMs for each task. At runtime, MII enhances the job success rates by dynamically making scheduling decisions to mitigate the workload losses from the power interruptions and adjusting these CMs in response to the actual energy patterns. Our evaluation demonstrates the superiority of MII over the state-of-the-art. In controlled environments, MII achieves an average increase of 21% and 39% in successful jobs under the stable and dynamic energy patterns. In the real-world settings, MII achieves 33% and 24% more successful jobs indoors and outdoors.
Cong Liu 0005, Hyoseung Kim 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2024 Automated Testing Linguistic Capabilities of NLP Models
abstract
Natural language processing (NLP) has gained widespread adoption in the development of real-world applications. However, the black-box nature of neural networks in NLP applications poses a challenge when evaluating their performance, let alone ensuring it. Recent research has proposed testing techniques to enhance the trustworthiness of NLP-based applications. However, most existing works use a single, aggregated metric (i.e., accuracy) which is difficult for users to assess NLP model performance on fine-grained aspects, such as LCs. To address this limitation, we present ALiCT, an automated testing technique for validating NLP applications based on their LCs. ALiCT takes user-specified LCs as inputs and produces diverse test suite with test oracles for each of given LC. We evaluate ALiCT on two widely adopted NLP tasks, sentiment analysis and hate speech detection, in terms of diversity, effectiveness, and consistency. Using Self-BLEU and syntactic diversity metrics, our findings reveal that ALiCT generates test cases that are 190% and 2213% more diverse in semantics and syntax, respectively, compared to those produced by state-of-the-art techniques. In addition, ALiCT is capable of producing a larger number of NLP model failures in 22 out of 25 LCs over the two NLP applications.
Austin Mordahl, Cong Liu 0005, Wei Yang 0013, Shiyi Wei
ACM Trans. Softw. Eng. Methodol.4
2023 Dynamic Transformers Provide a False Sense of Efficiency
abstract
Despite much success in natural language processing (NLP), pre-trained language models typically lead to a high computational cost during inference.Multi-exit is a mainstream approach to address this issue by making a tradeoff between efficiency and accuracy, where the saving of computation comes from an early exit.However, whether such saving from earlyexiting is robust remains unknown.Motivated by this, we first show that directly adapting existing adversarial attack approaches targeting model accuracy cannot significantly reduce inference efficiency.To this end, we propose a simple yet effective attacking framework, SAME, a novel slowdown attack framework on multi-exit models, which is specially tailored to reduce the efficiency of the multi-exit models.By leveraging the multi-exit models' design characteristics, we utilize all internal predictions to guide the adversarial sample generation instead of merely considering the final prediction.Experiments on the GLUE benchmark show that SAME can effectively diminish the efficiency gain of various multi-exit models by 80% on average, convincingly validating its effectiveness and generalization ability. 1
Yiming Chen 0010, Zexin Li 0001, Wei Yang 0013, Cong Liu 0005, Robby T. Tan, Haizhou Li 0001
ACL (1)5
2023 White-Box Multi-Objective Adversarial Attack on Dialogue Generation
abstract
Pre-trained transformers are popular in stateof-the-art dialogue generation (DG) systems.Such language models are, however, vulnerable to various adversarial samples as studied in traditional tasks such as text classification, which inspires our curiosity about their robustness in DG systems.One main challenge of attacking DG models is that perturbations on the current sentence can hardly degrade the response accuracy because the unchanged chat histories are also considered for decision-making.Instead of merely pursuing pitfalls of performance metrics such as BLEU, ROUGE, we observe that crafting adversarial samples to force longer generation outputs benefits attack effectiveness-the generated responses are typically irrelevant, lengthy, and repetitive.To this end, we propose a white-box multi-objective attack method called DGSlow.Specifically, DGSlow balances two objectives-generation accuracy and length, via a gradient-based multiobjective optimizer and applies an adaptive searching mechanism to iteratively craft adversarial samples with only a few modifications.Comprehensive experiments 1 on four benchmark datasets demonstrate that DGSlow could significantly degrade state-of-the-art DG models with a higher success rate than traditional accuracy-based methods.Besides, our crafted sentences also exhibit strong transferability in attacking other models.
Yufei Li 0001, Zexin Li 0001, Yingfan Gao, Cong Liu 0005
ACL (1)4
2023 The Dark Side of Dynamic Routing Neural Networks: Towards Efficiency Backdoor Injection
abstract
Recent advancements in deploying deep neural networks (DNNs) on resource-constrained devices have generated interest in input-adaptive dynamic neural networks (DyNNs). DyNNs offer more efficient inferences and enable the deployment of DNNs on devices with limited resources, such as mobile devices. However, we have discovered a new vulnerability in DyNNs that could potentially compromise their efficiency. Specifically, we investigate whether adversaries can manipulate DyNNs' computational costs to create a false sense of efficiency. To address this question, we propose EfficFrog, an adversarial attack that injects universal efficiency backdoors in DyNNs. To inject a backdoor trigger into DyNNs, EfficFrog poisons only a minimal percentage of the DyNNs' training data. During the inference phase, EfficFrog can slow down the backdoored DyNNs and abuse the computational resources of systems running DyNNs by adding the trigger to any input. To evaluate EfficFrog, we tested it on three DNN backbone architectures (based on VGG16, MobileNet, and ResNet56) using two popular datasets (CIFAR-10 and Tiny ImageNet). Our results demonstrate that EfficFrog reduces the efficiency of DyNNs on triggered input samples while keeping the efficiency of clean samples almost the same.
Mirazul Haque, Cong Liu 0005, Wei Yang 0013
CVPR4
2023 Sibling-Attack: Rethinking Transferable Adversarial Attacks against Face Recognition
abstract
A hard challenge in developing practical face recognition (FR) attacks is due to the black-box nature of the target FR model, i.e., inaccessible gradient and parameter information to attackers. While recent research took an important step towards attacking black-box FR models through leveraging transferability, their performance is still limited, especially against online commercial FR systems that can be pessimistic (e.g., a less than 50% ASR-attack success rate on average). Motivated by this, we present Sibling-Attack, a new FR attack technique for the first time explores a novel multi-task perspective (i.e., leveraging extra information from multi-correlated tasks to boost attacking transferability). Intuitively, Sibling-Attack selects a set of tasks correlated with FR and picks the Attribute Recognition (AR) task as the task used in Sibling-Attack based on theoretical and quantitative analysis. Sibling-Attack then develops an optimization framework that fuses adversarial gradient information through (1) constraining the cross-task features to be under the same space, (2) a Jointtask meta optimization framework that enhances the gradient compatibility among tasks, and (3) a cross-task gradient stabilization method which mitigates the oscillation effect during attacking. Extensive experiments demonstrate that SiblingAttack outperforms state-of-the-art FR attack techniques by a non-trivial margin, boosting ASR by 12.61% and 55.77% on average on state-of-the-art pre-trained FR models and two well-known, widely used commercial FR systems.
Zexin Li 0001, Bangjie Yin, Taiping Yao, Shouhong Ding, Cong Liu 0005
CVPR7
2023 PolicyCleanse: Backdoor Detection and Mitigation for Competitive Reinforcement Learning
abstract
While real-world applications of reinforcement learning (RL) are becoming popular, the security and robustness of RL systems are worthy of more attention and exploration. In particular, recent works have revealed that, in a multi-agent RL environment, backdoor trigger actions can be injected into a victim agent (a.k.a. Trojan agent), which can result in a catastrophic failure as soon as it sees the backdoor trigger action. To ensure the security of RL agents against malicious backdoors, in this work, we propose the problem of Backdoor Detection in a multi-agent competitive reinforcement learning system, with the objective of detecting Trojan agents as well as the corresponding potential trigger actions, and further trying to mitigate their Trojan behavior. In order to solve this problem, we propose PolicyCleanse that is based on the property that the activated Trojan agent’s accumulated rewards degrade noticeably after several timesteps. Along with PolicyCleanse, we also design a machine unlearning-based approach that can effectively mitigate the detected backdoor. Extensive experiments demonstrate that the proposed methods can accurately detect Trojan agents, and outperform existing backdoor mitigation baseline approaches by at least 3% in winning rate across various types of agents and environments.
Lixu Wang, Cong Liu 0005
ICCV4
2023 SCALE-UP: An Efficient Black-box Input-level Backdoor Detection via Analyzing Scaled Prediction Consistency
Yiming Li 0004, Hanqing Guo, Lichao Sun 0001, Cong Liu 0005
ICLR6
2023 SlothSpeech: Denial-of-service Attack Against Speech Recognition Models
Mirazul Haque, Rutvij Shah, Berrak Sisman, Cong Liu 0005, Wei Yang 0013
INTERSPEECH5
2023 PIMbot: Policy and Incentive Manipulation for Multi-Robot Reinforcement Learning in Social Dilemmas
abstract
Recent research has demonstrated the potential of reinforcement learning (RL) in enabling effective multi-robot collaboration, particularly in social dilemmas where robots face a trade-off between self-interests and collective benefits. However, environmental factors such as miscommunication and adversarial robots can impact cooperation, making it crucial to explore how multi-robot communication can be manipulated to achieve different outcomes. This paper presents a novel approach, namely PIMbot, to manipulating the reward function in multi-robot collaboration through two distinct forms of manipulation: policy and incentive manipulation. Our work introduces a new angle for manipulation in recent multi-agent RL social dilemmas that utilize a unique reward function for incentivization. By utilizing our proposed PIMbot mechanisms, a robot is able to manipulate the social dilemma environment effectively. PIMbot has the potential for both positive and negative impacts on the task outcome, where positive impacts lead to faster convergence to the global optimum and maximized rewards for any chosen robot. Conversely, negative impacts can have a detrimental effect on the overall task performance. We present comprehensive experimental results that demonstrate the effectiveness of our proposed methods in the Gazebo-simulated multi-robot environment. Our work provides insights into how inter-robot communication can be manipulated and has implications for various robotic applications.
Shahab Nikkhoo, Zexin Li 0001, Aritra Samanta, Yufei Li 0001, Cong Liu 0005
IROS5
2023 DyCL: Dynamic Neural Network Compilation Via Program Rewriting and Graph Optimization
abstract
The deep learning (DL) compiler serves as a vital infrastructure component to enable the deployment of deep neural networks on diverse hardware platforms such as mobile devices and Raspberry Pi. DL compiler’s primary function is to translate DNN programs written in high-level DL frameworks such as PyTorch and TensorFlow into portable executables. These executables can then be flexibly executed by the deployed host programs. However, existing DL compilers rely on a tracing mechanism, which involves feeding a runtime input to a neural network program and tracing the program execution paths to generate the computational graph necessary for compilation. Unfortunately, this mechanism falls short when dealing with modern dynamic neural networks (DyNNs) that possess varying computational graphs depending on the inputs. Consequently, conventional DL compilers struggle to accurately compile DyNNs into executable code. To address this limitation, we propose DyCL, a general approach that enables any existing DL compiler to successfully compile DyNNs. DyCL tackles the dynamic nature of DyNNs by introducing a compilation mechanism that redistributes the control and data flow of the original DNN programs during the compilation process. Specifically, DyCL develops program analysis and program transformation techniques to convert a dynamic neural network into multiple sub-neural networks. Each sub-neural network is devoid of conditional statements and is compiled independently. Furthermore, DyCL synthesizes a host module that models the control flow of the DyNNs and facilitates the invocation of the sub-neural networks. Our evaluation demonstrates the effectiveness of DyCL, achieving a 100% success rate in compiling all dynamic neural networks. Moreover, the compiled executables generated by DyCL exhibit significantly improved performance, running between 1.12× and 20.21× faster than the original DyNNs executed on general-purpose DL frameworks.
Shiyi Wei, Cong Liu 0005, Wei Yang 0013
ISSTA3
2023 Domain Watermark: Effective and Harmless Dataset Copyright Protection is Closed at Hand
abstract
The prosperity of deep neural networks (DNNs) is largely benefited from open-source datasets, based on which users can evaluate and improve their methods. In this paper, we revisit backdoor-based dataset ownership verification (DOV), which is currently the only feasible approach to protect the copyright of open-source datasets. We reveal that these methods are fundamentally harmful given that they could introduce malicious misclassification behaviors to watermarked DNNs by the adversaries. In this paper, we design DOV from another perspective by making watermarked models (trained on the protected dataset) correctly classify some `hard' samples that will be misclassified by the benign model. Our method is inspired by the generalization property of DNNs, where we find a \emph{hardly-generalized domain} for the original dataset (as its \emph{domain watermark}). It can be easily learned with the protected dataset containing modified samples. Specifically, we formulate the domain generation as a bi-level optimization and propose to optimize a set of visually-indistinguishable clean-label modified data with similar effects to domain-watermarked samples from the hardly-generalized domain to ensure watermark stealthiness. We also design a hypothesis-test-guided ownership verification via our domain watermark and provide the theoretical analyses of our method. Extensive experiments on three benchmark datasets are conducted, which verify the effectiveness of our method and its resistance to potential adaptive methods.
Yiming Li 0004, Lixu Wang, Shutao Xia, Heng Huang 0001, Cong Liu 0005, Bo Li 0026
NeurIPS6
2023 $\mathrm{R}^{3}$: On-Device Real-Time Deep Reinforcement Learning for Autonomous Robotics
abstract
Autonomous robotic systems, like autonomous vehicles and robotic search and rescue, require efficient on-device training for continuous adaptation of Deep Reinforcement Learning (DRL) models in dynamic environments. This research is fundamentally motivated by the need to understand and address the challenges of on-device real-time DRL, which involves balancing timing and algorithm performance under memory constraints, as exposed through our extensive empirical studies. This intricate balance requires co-optimizing two pivotal parameters of DRL training - batch size and replay buffer size. Configuring these parameters significantly affects timing and algorithm performance, while both (unfortunately) require substantial memory allocation to achieve near-optimal performance. This paper presents$\mathbf{R}^{3}$, a holistic solution for managing timing, memory, and algorithm performance in on-device real-time DRL training.$\mathbf{R}^{3}$employs (i) a deadline-driven feedback loop with dynamic batch sizing for optimizing timing, (ii) efficient memory management to reduce memory footprint and allow larger replay buffer sizes, and (iii) a runtime coordinator guided by heuristic analysis and a runtime profiler for dynamically adjusting memory resource reservations. These components collaboratively tackle the trade-offs in on-device DRL training, improving timing and algorithm performance while minimizing the risk of out-of-memory (OOM) errors. We implemented and evaluated$\mathbf{R}^{3}$extensively across various DRL frameworks and benchmarks on three hardware platforms commonly adopted by autonomous robotic systems. Additionally, we integrate$\mathbf{R}^{3}$with a popular realistic autonomous car simulator to demonstrate its real-world applicability. Evaluation results show that$\mathbf{R}^{3}$achieves efficacy across diverse platforms, ensuring consistent latency performance and timing predictability with minimal overhead. Moreover,$\mathbf{R}^{3}$showcases versatility by handling varied optimization goals and adapting to fluctuating systems scenarios.
Zexin Li 0001, Aritra Samanta, Yufei Li 0001, Andrea Soltoggio, Hyoseung Kim 0001, Cong Liu 0005
RTSS6
2023 RT-LM: Uncertainty-Aware Resource Management for Real-Time Inference of Language Models
abstract
Recent advancements in language models (LMs) have gained substantial attentions on their capability to generate human-like responses. Though exhibiting a promising future for various applications such as conversation AI, these LMs face deployment challenges on various devices due to their extreme computational cost and unpredictable inference latency. Such varied inference latency, identified as a consequence of uncertainty intrinsic to the nature of language, can lead to computational inefficiency and degrade the overall performance of LMs, especially under high-traffic workloads. Unfortunately, the bandwidth of these uncertainty sources is extensive, complicating the prediction of latency and the effects emanating from such uncertainties. To understand and mitigate the impact of uncertainty on real-time response-demanding systems, we take the first step to comprehend, quantify and optimize these uncertainty-induced latency performance variations in LMs. Specifically, we present RT-LM, an uncertainty-aware resource management ecosystem for real-time inference of LMs. RT-LM innovatively quantifies how specific input uncertainties, recognized within the NLP community, adversely affect latency, often leading to an increased output length. Exploiting these insights, we devise a lightweight yet effective method to dynamically correlate input text uncertainties with output length at runtime. Utilizing this quantification as a latency heuristic, we integrate the uncertainty information into a system-level scheduler which explores several uncertainty-induced optimization opportunities, including uncertainty-aware prioritization, dynamic consolidation, and strategic CPU offloading. Quantitative experiments across five state-of-the-art LMs on two hardware platforms demonstrates that RT-LM can significantly reduce the average response time and improve throughput while incurring a rather small runtime overhead.
Yufei Li 0001, Zexin Li 0001, Wei Yang 0013, Cong Liu 0005
RTSS4
2023 RED: A Systematic Real-Time Scheduling Approach for Robotic Environmental Dynamics
abstract
Intelligent robots are designed to effectively navigate dynamic and unpredictable environments laden with moving mechanical elements and objects. Such environment-induced dynamics, including moving obstacles, can readily alter the computational demand (e.g., the creation of new tasks) and the structure of workloads (e.g., precedence constraints among tasks) during runtime, thereby adversely affecting overall system performance. This challenge is amplified when multi-task inference is expected on robots operating under stringent resource and real-time constraints. To address such a challenge, we introduce RED, a systematic real-time scheduling approach designed to support multi-task deep neural network workloads in resource-limited robotic systems. It is designed to adaptively manage the Robotic Environmental Dynamics (RED) while adhering to real-time constraints. At the core of RED lies a deadline-based scheduler that employs an intermediate deadline assignment policy, effectively managing to change workloads and asynchronous inference prompted by complex, unpredictable environments. This scheduling framework also facilitates the flexible deployment of MIMONet (multi-input multi-output neural networks), which are commonly utilized in multi-tasking robotic systems to circumvent memory bottlenecks. Building on this scheduling framework, RED recognizes and leverages a unique characteristic of MIMONet: its weight-shared architecture. To further accommodate and exploit this feature, RED devises a novel and effective workload refinement and reconstruction process. This process ensures the scheduling framework's compatibility with MIMONet and maximizes efficiency. We have implemented RED on several widely used embedded and mobile platforms, including the NVIDIA Jetson Nano, TX2, Xavier, and Orin platforms. We evaluated its performance using workloads that span a broad range of settings typical in navigation robots. The experimental results demonstrate that RED surpasses existing approaches (often by a significant margin) across critical metrics such as throughput, timing correctness, interference robustness, adaptability, and overhead.
Zexin Li 0001, Xiaoxi He, Cong Liu 0005
RTSS4
2023 Data Fusion in Infrastructure-Augmented Autonomous Driving System: Why? Where? and How?
abstract
This article is the first to provide a thorough system design overview along with the fusion methods selection criteria of a real-world cooperative autonomous driving system enabled by the Internet of Things (IoT), named infrastructure-augmented autonomous driving (IAAD). We present an in-depth introduction to the IAAD hardware and software on both road side and vehicle side. We extensively characterize the IAAD system and observe that the network condition fluctuation along the road is the main roadblock for cooperative autonomous driving. To address this challenge, we propose new fusion methods, dubbed “interframe fusion” and “planning fusion” to complement the state-of-the-art “intraframe fusion.” We demonstrate that each fusion method has its own benefit and constraint. In order to select the best fusion method under varying network conditions, we propose “fusion criteria” to instruct the IAAD system to intelligently make the selection and implement a system framework named adaptive spatial-temporal (S–T) choice to realize the adaptive fusion guided by the “fusion criteria.” Our real-world field data verifies that S–T choice has significantly improved autonomous driving’s safety and reliability by decreasing the fusion miss ratio from 30% to 7% and remain the planning displacement error within the 1.7 m instead of 4 m when the network condition exacerbates.
Bo Yu 0014, Jie Tang 0003, Shuaiwen Song, Cong Liu 0005, Yang Hu 0001
IEEE Internet Things J.6
2023 MemPerf: Profiling Allocator-Induced Performance Slowdowns
abstract
The memory allocator plays a key role in the performance of applications, but none of the existing profilers can pinpoint performance slowdowns caused by a memory allocator. Consequently, programmers may spend time improving application code incorrectly or unnecessarily, achieving low or no performance improvement. This paper designs the first profiler—MemPerf—to identify allocator-induced performance slowdowns without comparing against another allocator. Based on the key observation that an allocator may impact the whole life-cycle of heap objects, including the accesses (or uses) of these objects, MemPerf proposes a life-cycle based detection to identify slowdowns caused by slow memory management operations and slow accesses separately. For the prior one, MemPerf proposes a thread-aware and type-aware performance modeling to identify slow management operations. For slow memory accesses, MemPerf utilizes a top-down approach to identify all possible reasons for slow memory accesses introduced by the allocator, mainly due to cache and TLB misses, and further proposes a unified method to identify them correctly and efficiently. Based on our extensive evaluation, MemPerf reports 98% medium and large allocator-reduced slowdowns (larger than 5%) correctly without reporting any false positives. MemPerf also pinpoints multiple known and unknown design issues in widely-used allocators.
Sam Silvestro, Steven (Jiaxun) Tang, Hanmei Yang, Hongyu Liu 0005, Guangming Zeng, Bo Wu 0002, Cong Liu 0005, Tongping Liu
Proc. ACM Program. Lang.8
2022 Neural Mean Discrepancy for Efficient Out-of-Distribution Detection
abstract
Various approaches have been proposed for out-of-distribution (OOD) detection by augmenting models, input examples, training sets, and optimization objectives. Deviating from existing work, we have a simple hypothesis that standard off-the-shelf models may already contain sufficient information about the training set distribution which can be leveraged for reliable OOD detection. Our empirical study on validating this hypothesis, which measures the model activation's mean for OOD and in-distribution (ID) minibatches, surprisingly finds that activation means of OOD mini-batches consistently deviate more from those of the training data. In addition, training data's activation means can be computed offline efficiently or retrieved from batch normalization layers as a ‘free lunch’. Based upon this observation, we propose a novel metric called Neural Mean Discrepancy (NMD), which compares neural means of the input examples and training data. Leveraging the simplicity of NMD, we propose an efficient OOD detector that computes neural means by a standard forward pass followed by a lightweight classifier. Extensive experiments show that NMD outperforms state-of-the-art OOD approaches across multiple datasets and model architectures in terms of both detection accuracy and computational cost.
Xin Dong 0009, Wei-Te Ting, Cong Liu 0005, H. T. Kung 0001
CVPR5
2022 NICGSlowDown: Evaluating the Efficiency Robustness of Neural Image Caption Generation Models
abstract
Neural image caption generation (NICG) models have received massive attention from the research community due to their excellent performance in visual understanding. Existing work focuses on improving NICG model ac-curacy while efficiency is less explored. However, many real-world applications require real-time feedback, which highly relies on the efficiency of NICG models. Recent re-search observed that the efficiency of NICG models could vary for different inputs. This observation brings in a new attack surface of NICG models, i.e., An adversary might be able to slightly change inputs to cause the NICG mod-els to consume more computational resources. To further understand such efficiency-oriented threats, we propose a new attack approach, NICGSlowDown, to evaluate the ef-ficiency robustness of NICG models. Our experimental re-sults show that NICGSlowDown can generate images with human-unnoticeable perturbations that will increase the NICG model latency up to 483.86%. We hope this research could raise the community's concern about the efficiency robustness of NICG models.
Mirazul Haque, Cong Liu 0005, Wei Yang 0013
CVPR4
2022 Deep Partial Updating: Towards Communication Efficient Updating for On-Device Inference
Zhongnan Qu, Cong Liu 0005, Lothar Thiele
ECCV (11)2
2022 AEVA: Black-box Backdoor Detection Using Adversarial Extreme Value Analysis
Cong Liu 0005
ICLR3
2022 EREBA: Black-box Energy Testing of Adaptive Neural Networks
abstract
Recently, various Deep Neural Network (DNN) models have been proposed for environments like embedded systems with stringent energy constraints. The fundamental problem of determining the robustness of a DNN with respect to its energy consumption (energy robustness) is relatively unexplored compared to accuracy-based robustness. This work investigates the energy robustness of Adaptive Neural Networks (AdNNs), a type of energy-saving DNNs proposed for many energy-sensitive domains and have recently gained traction. We propose EREBA, the first black-box testing method for determining the energy robustness of an AdNN. EREBA explores and infers the relationship between inputs and the energy consumption of AdNNs to generate energy surging samples. Extensive implementation and evaluation using three state-of-the-art AdNNs demonstrate that test inputs generated by EREBA could degrade the performance of the system substantially. The test inputs generated by EREBA can increase the energy consumption of AdNNs by 2,000% compared to the original inputs. Our results also show that test inputs generated via EREBA are valuable in detecting energy surging inputs.
Mirazul Haque, Yaswanth Yadlapalli, Wei Yang 0013, Cong Liu 0005
ICSE4
2022 Learn to Reverse DNNs from AI Programs Automatically
abstract
With the privatization deployment of DNNs on edge devices, the security of on-device DNNs has raised significant concern. To quantify the model leakage risk of on-device DNNs automatically, we propose NNReverse, the first learning-based method which can reverse DNNs from AI programs without domain knowledge. NNReverse trains a representation model to represent the semantics of binary code for DNN layers. By searching the most similar function in our database, NNReverse infers the layer type of a given function’s binary code. To represent assembly instructions semantics precisely, NNReverse proposes a more fine-grained embedding model to represent the textual and structural-semantic of assembly functions.
Hamed Khanpour, Cong Liu 0005, Wei Yang 0013
IJCAI3
2022 DeepPerform: An Efficient Approach for Performance Testing of Resource-Constrained Neural Networks
abstract
Today, an increasing number of Adaptive Deep Neural Networks (AdNNs) are being used on resource-constrained embedded devices. We observe that, similar to traditional software, redundant computation exists in AdNNs, resulting in considerable performance degradation. The performance degradation is dependent on the input and is referred to as input-dependent performance bottlenecks (IDPBs). To ensure an AdNN satisfies the performance requirements of resource-constrained applications, it is essential to conduct performance testing to detect IDPBs in the AdNN. Existing neural network testing methods are primarily concerned with correctness testing, which does not involve performance testing. To fill this gap, we propose DeepPerform, a scalable approach to generate test samples to detect the IDPBs in AdNNs. We first demonstrate how the problem of generating performance test samples detecting IDPBs can be formulated as an optimization problem. Following that, we demonstrate how DeepPerform efficiently handles the optimization problem by learning and estimating the distribution of AdNNs’ computational consumption. We evaluate DeepPerform on three widely used datasets against five popular AdNN models. The results show that DeepPerform generates test samples that cause more severe performance degradation (FLOPs: increase up to 552%). Furthermore, DeepPerform is substantially more efficient than the baseline methods in generating test inputs (runtime overhead: only 6–10 milliseconds).
Mirazul Haque, Cong Liu 0005, Wei Yang 0013
ASE3
2022 Brief Industry Paper: The Necessity of Adaptive Data Fusion in Infrastructure-Augmented Autonomous Driving System
abstract
This paper is the first to provide a thorough system design overview along with the fusion methods selection criteria of a real-world cooperative autonomous driving system, named Infrastructure-Augmented Autonomous Driving or IAAD. We present an in-depth introduction of the IAAD hardware and software on both road-side and vehicle-side computing/communication platforms. We extensively characterize the IAAD system in the context of real-world deployment scenarios and observe that the network condition fluctuates along the road is currently the main technical roadblock for cooperative autonomous driving. To address this challenge, we propose new fusion methods, dubbed “inter-frame fusion” and “planning fusion” to complement the current state-of-the-art “intra-frame fusion”. We demonstrate that each fusion method has its own benefit and constraint. Adaptively choosing the fusion method according to the real-world condition will benefit the SoV without the violation of the SoV's safety requirements.
Shaoshan Liu, Bo Yu 0014, Jie Tang 0003, Shuaiwen Song, Cong Liu 0005, Yang Hu 0001
RTAS9
2022 A Utilization-based Test for Non-preemptive Gang Tasks on Multiprocessors
abstract
Real-time gang task scheduling has received much recent attention due to the emerging trend of applying highly parallel accelerators (e.g., GPU) and parallel programming models (e.g., OpenMP) in many real-time computing domains. However, existing works on gang task scheduling mainly focus on the preemptive scheduling case, which contradicts a bit with the non-preemptive executing nature of applying gang scheduling techniques in practice. In this paper, we present a set of non-trivial techniques that can analyze the schedulability of scheduling a hard real-time sporadic gang task system under non-preemptive GEDF on multiprocessors. A utilization-based schedulability test (first-of-its-kind) is derived, which is shown to be rather effective via experiments. Rather interestingly, for a special case where each gang task becomes an ordinary sporadic task, our developed test is shown by experiments that it improves schedulability by 75% on average upon a state-of-the-art utilization-based test designed for non-preemptive scheduling of ordinary sporadic tasks on multiprocessors.
Zheng Dong 0002, Cong Liu 0005
RTSS2
2022 NMTSloth: understanding and testing efficiency degradation of neural machine translation systems
abstract
Neural Machine Translation (NMT) systems have received much recent attention due to their human-level accuracy. While existing works mostly focus on either improving accuracy or testing accuracy robustness, the computation efficiency of NMT systems, which is of paramount importance due to often vast translation demands and real-time requirements, has surprisingly received little attention. In this paper, we make the first attempt to understand and test potential computation efficiency robustness in state-of-the-art NMT systems. By analyzing the working mechanism and implementation of 1455 public-accessible NMT systems, we observe a fundamental property in NMT systems that could be manipulated in an adversarial manner to reduce computation efficiency significantly. Our interesting observation is that the output length determines the computation efficiency of NMT systems instead of the input, where the output length depends on two factors: an often sufficiently large yet pessimistic pre-configured threshold controlling the max number of iterations and a runtime generated end of sentence (EOS) token. Our key motivation is to generate test inputs that could sufficiently delay the generation of EOS such that NMT systems would have to go through enough iterations to satisfy the pre-configured threshold. We present NMTSloth, which develops a gradient-guided technique that searches for a minimal and unnoticeable perturbation at character-level, token-level, and structure-level, which sufficiently delays the appearance of EOS and forces these inputs to reach the naturally-unreachable threshold. To demonstrate the effectiveness of NMTSloth, we conduct a systematic evaluation on three public-available NMT systems: Google T5, AllenAI WMT14, and Helsinki-NLP translators. Experimental results show that NMTSloth can increase NMT systems' response latency and energy consumption by 85% to 3153% and 86% to 3052%, respectively, by perturbing just one character or token in the input sentence. Our case study shows that inputs generated by NMTSloth significantly affect the battery power in real-world mobile devices (i.e., drain more than 30 times battery power than normal inputs).
Cong Liu 0005, Mirazul Haque, Wei Yang 0013
ESEC/SIGSOFT FSE2
2022 Deep-Learning-Based Wireless Human Motion Tracking for Mobile Ship Environments
abstract
Being able to track passengers’ movement without invasion of their privacy plays an important role in cruise ships; it enables crucial location-based services, such as maritime search and rescue, tourist services, and epidemic prevention. The past few years have witnessed commodity WiFi holding great potential that provides such services available thanks to its ubiquitous in indoor scenarios. However, existing WiFi-based tracking methods suffer from huge performance degradation in sailing ships due to their complex metal structures and dynamic hull deformation caused by engines and waves/payloads pressure. In this article, we present CRLoc, a deep learning-based passive human tracking system that can overcome the practical limitations of traditional WiFi-based localization approaches applied in a multipath-rich and mobile ship environment, and provide decimeter-level tracking accuracy in cruise ships. Specifically, we make two contributions, i.e., we propose a super-resolution parameter estimation algorithm that better characterizes ship indoor environments, and a deep neural network-based end-to-end solution to remove the impact of noise, interference, and mobility in ships. The real-world implementation and extensive experiments in several passenger ships demonstrate that CRLoc tracks human motions with a median error of 92 cm, better than state-of-the-art localization methods. To our knowledge, this is one of the first WiFi-based passive human motion tracking system in a cruise ship environment.
Kezhong Liu, Mozi Chen, Kai Zheng 0022, Xuming Zeng, Shengkai Zhang, Cong Liu 0005
IEEE Internet Things J.7
2022 Schedulability Analysis for Coscheduling Real-Time Tasks on Multiprocessors
abstract
The real-time coscheduling problem, where tasks may have multiple phases executing on different types of processors, is known to be hard. The (already hard) self-suspending task scheduling simplifies the coscheduling problem by assuming that the latency a task may experience on the other type of processors is naturally bounded, which is unfortunately not true in practice. In this article, we present a novel analysis technique, namely, the vertical view analysis, for analyzing the schedulability of coscheduling sporadic tasks under global earliest-deadline-first (GEDF) on a heterogeneous multiprocessor consisting of two types of processors. We derive both hard (no deadline miss) and soft (bounded response times) real-time utilization-based tests. To the best of our knowledge, these results are the first-of-its-kind for the coscheduling problem and may allow real-time schedulability analysis to be carried out on more practical scenarios under heterogeneous computing.
Zheng Dong 0002, Cong Liu 0005
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2022 Collision-Free Dynamic Convergecast in Low-Duty-Cycle Wireless Sensor Networks
abstract
Convergecast is a fundamental operation in wireless sensor networks (WSNs). To support long-term deployment of WSNs, sensor nodes normally operate at low-duty-cycles. However, the low-duty-cycle operation significantly reduces the communication chance between nodes. Consequently, the risk of data collisions significantly increases when multiple senders transmit packets to a receiver during its very short active period. This problem further causes not only wasted packet retransmissions, but also a large delivery latency. Under such conditions, collision-free medium access is more appealing than recovering after collision for low-duty-cycle WSNs. In this work, we propose anincast-collision-free convergecast protocol, named iCore, to address the many-to-one collision problem in low-duty-cycle WSNs. iCore employs the dynamic forwarding technique, establishes a non-conflicting schedule for efficient convergecast, and improves the channel utilization by allowing senders to opportunistically transmit packets once detecting unused slots. Specifically, we design efficient forwarder assignment and forwarding optimization algorithms that ensure low end-to-end latency under diverse data traffic types. Through comprehensive performance evaluations, we demonstrate that, compared with the baseline protocol, iCore effectively minimizes the end-to-end delay by 25% ~ 57% and maintains high delivery ratio and energy efficiency for different many-to-one convergecast scenarios.
Long Cheng 0005, Linghe Kong, Yu Gu 0001, Jianwei Niu 0002, Ting Zhu 0001, Cong Liu 0005, Shahid Mumtaz, Tian He 0001
IEEE Trans. Wirel. Commun.6
2021 gGuard: Enabling Leakage-Resilient Memory Isolation in GPU-accelerated Autonomous Embedded Systems
abstract
Graphics processing units (GPUs) are being widely used as co-processors for performance acceleration in many autonomous embedded systems such as robotics and autonomous vehicles. However, current GPU hardware and systems software, including GPU device drivers, compilers, and operating systems, do not implement proper memory protection mechanisms due to performance and proprietary reasons, causing severe vulnerabilities such as information leakage. In this paper, we present gGuard, a leakage-resilient GPU memory management system with strong isolation. Based on the intrinsic characteristics of information leakage vulnerabilities on GPUs, gGuard develops a set of efficient and accurate data shredding techniques implemented at the compiler, library, and operating system levels, with the core idea of exploring the data access patterns and dependencies for efficient application-aware data shredding. Our implementation and evaluation show that gGuard can provide effective mitigation on GPU data leakage issues through efficient GPU data shredding while introducing less than 6% overhead in all tested scenarios.
Yaswanth Yadlapalli, Husheng Zhou, Yuqun Zhang, Cong Liu 0005
DAC4
2021 Adv-Makeup: A New Imperceptible and Transferable Attack on Face Recognition
abstract
Deep neural networks, particularly face recognition models, have been shown to be vulnerable to both digital and physical adversarial examples. However, existing adversarial examples against face recognition systems either lack transferability to black-box models, or fail to be implemented in practice. In this paper, we propose a unified adversarial face generation method - Adv-Makeup, which can realize imperceptible and transferable attack under the black-box setting. Adv-Makeup develops a task-driven makeup generation method with the blending module to synthesize imperceptible eye shadow over the orbital region on faces. And to achieve transferability, Adv-Makeup implements a fine-grained meta-learning based adversarial attack strategy to learn more vulnerable or sensitive features from various models. Compared to existing techniques, sufficient visualization results demonstrate that Adv-Makeup is capable to generate much more imperceptible attacks under both digital and physical scenarios. Meanwhile, extensive quantitative experiments show that Adv-Makeup can significantly improve the attack success rate under black-box setting, even attacking commercial systems.
Bangjie Yin, Wenxuan Wang 0003, Taiping Yao, Zelun Kong, Shouhong Ding, Cong Liu 0005
IJCAI8
2021 LAG-Based Analysis Techniques for Scheduling Multiprocessor Hard Real-Time Sporadic DAGs
abstract
Global scheduling of real time tasks with precedence constraints has received significant attention recently. Specifically, on multiprocessor systems, the directed acyclic graph (DAG) model well-represents parallelizable workloads with precedence constraints. Hence, many studies have recently been published analyzing DAG structures using decomposition and window-based analysis methods. We identified a set of schedulable DAG taskets that are hard-to-analyze using state-of-the-art window-based schedulability tests. Additionally, we observe that the window-based test solely depends on one structural feature of the DAG taskset, which raises concerns about its pessimism in many settings. In this paper, to address these concerns, for hard real-time sporadic implicit-deadline DAG tasksets, we perform LAG-based schedulability analysis, which offers a more holistic view of the taskset than the window-based analysis. We present a companion utilization-based schedulability test to the state-of-the-art, which considers additional structural features of DAGs. Our results show that by considering such features, our LAG-based test empirically dominates the state-of-the-art test on over 80% of the evaluated DAG tasksets. Moreover, combining our LAG-based test in conjunction with the window-based tests can achieve high schedulability in many cases.
Yaswanth Yadlapalli, Cong Liu 0005
RTSS2
2021 Efficient algorithms for task mapping on heterogeneous CPU/GPU platforms for fast completion time
Zexin Li 0001, Yuqun Zhang, Husheng Zhou, Cong Liu 0005
J. Syst. Archit.5
2021 SWIM: Speed-Aware WiFi-Based Passive Indoor Localization for Mobile Ship Environment
abstract
Accurate and pervasive device-free indoor localization with meter-level resolution is critical for large cruise and passenger ships due to safety-critical rescue and evacuation requirements when accidents occur. However, existing localization techniques would severely suffer on ships because of their unique mobility characteristics. In this paper, we take the first attempt to build a ubiquitous passive localization system using WiFi fingerprints for the mobile ship environment. By conducting extensive experiments and measurements during several cruise trips, we identified a major influence factor on the fingerprints in the mobile environment: varying the ship speeds may significantly change the patterns of fingerprints at runtime. Since it may be too expensive to identify the fingerprints associated with different speeds, we propose an efficient localization method, namely SWIM, which calibrates the fingerprints from only a single-speed scenario to multiple-speed scenarios using a signal reconstruction analysis. SWIM is designed to learn the predictive fingerprint variation introduced by environmental speed changes and reconstruct the original fingerprints to adapt to the runtime speed scenarios. We have implemented and extensively evaluated SWIM on actual cruise ships. Experimental results demonstrate that SWIM improves localization accuracy from 63.2 to 82.9 percent, while reducing the overall system deployment cost by 87 percent.
Mozi Chen, Kezhong Liu, Yu Gu 0001, Zheng Dong 0002, Cong Liu 0005
IEEE Trans. Mob. Comput.6
2021 Tardiness Bounds for Sporadic Gang Tasks Under Preemptive Global EDF Scheduling
abstract
Following the trend of increasing autonomy in cyber-physical systems, parallel embedded architectures have enabled devices to better handle the large streams of data and intensive computation required by such autonomous systems. However, while the explosion of highly-parallel platforms has seen a proportional growth in the number of applications/devices that utilize these platforms, the embedded systems community's understanding of how to build time-predictable, safety-critical systems with parallel platforms has not kept pace. As a well-motivated but challenging parallel scheduling model, gang scheduling requires all parallel threads of each parallel task to simultaneously execute in unison, which is in contrast to traditional, multi-threaded parallel scheduling, where a parallel task may spawn multiple threads, and each thread will be scheduled independently of other threads of the same task. While increasing research efforts on hard real-time (HRT) gang scheduling have recently been seen, the problem of gang scheduling in the context of soft real-time (SRT) systems, where provably bounded deadline tardiness can be tolerated, has hardly been studied yet. In this article, we derive and prove the first tardiness bounds for sporadic gang task systems under preemptive GEDF scheduling. A total utilization bound for SRT-schedulability is required for ensuring such tardiness bounds but it is shown to be tight with respect to the platform capacity and maximum parallelism-induced idleness. Furthermore, we also empirically evaluate the effects of different degrees of task parallelism upon the SRT-schedulability.
Zheng Dong 0002, Kecheng Yang 0001, Nathan Fisher, Cong Liu 0005
IEEE Trans. Parallel Distributed Syst.4
2020 ILFO: Adversarial Attack on Adaptive Neural Networks
abstract
With the increasing number of layers and parameters in neural networks, the energy consumption of neural networks has become a great concern to society, especially to users of handheld or embedded devices. In this paper, we investigate the robustness of neural networks against energy-oriented attacks. Specifically, we propose ILFO (Intermediate Output-Based Loss Function Optimization) attack against a common type of energy-saving neural networks, Adaptive Neural Networks (AdNN). AdNNs save energy consumption by dynamically deactivating part of its model based on the need of the inputs. ILFO leverages intermediate output as a proxy to infer the relation between input and its corresponding energy consumption. ILFO has shown an increase up to 100 % of the FLOPs (floating-point operations per second) reduced by AdNNs with minimum noise added to input images. To our knowledge, this is the first attempt to attack the energy consumption of an AdNN.
Mirazul Haque, Anki Chauhan, Cong Liu 0005, Wei Yang 0013
CVPR3
2020 PhysGAN: Generating Physical-World-Resilient Adversarial Examples for Autonomous Driving
abstract
Although Deep neural networks (DNNs) are being pervasively used in vision-based autonomous driving systems, they are found vulnerable to adversarial attacks where small-magnitude perturbations into the inputs during test time cause dramatic changes to the outputs. While most of the recent attack methods target at digital-world adversarial scenarios, it is unclear how they perform in the physical world, and more importantly, the generated perturbations under such methods would cover a whole driving scene including those fixed background imagery such as the sky, making them inapplicable to physical world implementation. We present PhysGAN, which generates physical-world-resilient adversarial examples for misleading autonomous driving systems in a continuous manner. We show the effectiveness and robustness of PhysGAN via extensive digital- and real-world evaluations. We compare PhysGAN with a set of state-of-the-art baseline methods, which further demonstrate the robustness and efficacy of our approach. We also show that PhysGAN outperforms state-of-the-art baseline methods. To the best of our knowledge, PhysGAN is probably the first technique of generating realistic and physical-world-resilient adversarial examples for attacking common autonomous driving scenarios.
Zelun Kong, Cong Liu 0005
CVPR4
2020 Practical Poisoning Attacks on Neural Networks
Cong Liu 0005
ECCV (27)2
2020 Simulee: detecting CUDA synchronization bugs via memory-access modeling
abstract
While CUDA has become a mainstream parallel computing platform and programming model for general-purpose GPU computing, how to effectively and efficiently detect CUDA synchronization bugs remains a challenging open problem. In this paper, we propose the first lightweight CUDA synchronization bug detection framework, namely Simulee, to model CUDA program execution by interpreting the corresponding LLVM bytecode and collecting the memory-access information for automatically detecting general CUDA synchronization bugs. To evaluate the effectiveness and efficiency of Simulee, we construct a benchmark with 7 popular CUDA-related projects from GitHub, upon which we conduct an extensive set of experiments. The experimental results suggest that Simulee can detect 21 out of the 24 manually identified bugs in our preliminary study and also 24 previously unknown bugs among all projects, 10 of which have already been confirmed by the developers. Furthermore, Simulee significantly outperforms state-of-the-art approaches for CUDA synchronization bug detection.
Mingyuan Wu, Yicheng Ouyang, Husheng Zhou, Lingming Zhang 0001, Cong Liu 0005, Yuqun Zhang
ICSE5
2020 DeepBillboard: systematic physical-world testing of autonomous driving systems
abstract
Deep Neural Networks (DNNs) have been widely applied in autonomous systems such as self-driving vehicles. Recently, DNN testing has been intensively studied to automatically generate adversarial examples, which inject small-magnitude perturbations into inputs to test DNNs under extreme situations. While existing testing techniques prove to be effective, particularly for autonomous driving, they mostly focus on generating digital adversarial perturbations, e.g., changing image pixels, which may never happen in the physical world. Thus, there is a critical missing piece in the literature on autonomous driving testing: understanding and exploiting both digital and physical adversarial perturbation generation for impacting steering decisions. In this paper, we propose a systematic physical-world testing approach, namely DeepBillboard, targeting at a quite common and practical driving scenario: drive-by billboards. DeepBillboard is capable of generating a robust and resilient printable adversarial billboard test, which works under dynamic changing driving conditions including viewing angle, distance, and lighting. The objective is to maximize the possibility, degree, and duration of the steering-angle errors of an autonomous vehicle driving by our generated adversarial billboard. We have extensively evaluated the efficacy and robustness of DeepBillboard by conducting both experiments with digital perturbations and physical-world case studies. The digital experimental results show that DeepBillboard is effective for various steering models and scenes. Furthermore, the physical case studies demonstrate that DeepBillboard is sufficiently robust and resilient for generating physical-world adversarial billboard tests for real-world driving under various weather conditions, being able to mislead the average steering angle error up to 26.44 degrees. To the best of our knowledge, this is the first study demonstrating the possibility of generating realistic and continuous physical-world tests for practical autonomous driving systems; moreover, DeepBillboard can be directly generalized to a variety of other physical entities/surfaces along the curbside, e.g., a graffiti painted on a wall.
Husheng Zhou, Wei Li 0159, Zelun Kong, Yuqun Zhang, Bei Yu 0001, Lingming Zhang 0001, Cong Liu 0005
ICSE8
2020 Co-Optimizing Performance and Memory Footprint Via Integrated CPU/GPU Memory Management, an Implementation on Autonomous Driving Platform
abstract
Cutting-edge embedded system applications, such as self-driving cars and unmanned drone software, are reliant on integrated CPU/GPU platforms for their DNNs-driven workload, such as perception and other highly parallel components. In this work, we set out to explore the hidden performance implication of GPU memory management methods of integrated CPU/GPU architecture. Through a series of experiments on micro-benchmarks and real-world workloads, we find that the performance under different memory management methods may vary according to application characteristics. Based on this observation, we develop a performance model that can predict system overhead for each memory management method based on application characteristics. Guided by the performance model, we further propose a runtime scheduler. By conducting per-task memory management policy switching and kernel overlapping, the scheduler can significantly relieve the system memory pressure and reduce the multitasking co-run response time. We have implemented and extensively evaluated our system prototype on the NVIDIA Jetson TX2, Drive PX2, and Xavier AGX platforms, using both Rodinia benchmark suite and two real-world case studies of drone software and autonomous driving software.
Soroush Bateni, Yuankun Zhu, Yang Hu 0001, Cong Liu 0005
RTAS5
2020 DENAS: automated rule generation by knowledge extraction from neural networks
abstract
Deep neural networks (DNNs) have been widely applied in the software development process to automatically learn patterns from massive data. However, many applications still make decisions based on rules that are manually crafted and verified by domain experts due to safety or security concerns. In this paper, we aim to close the gap between DNNs and rule-based systems by automating the rule generation process via extracting knowledge from well-trained DNNs. Existing techniques with similar purposes either rely on specific DNNs input instances or use inherently unstable random sampling of the input space. Therefore, these approaches either limit the exploration area to a local decision-space of the DNNs or fail to converge to a consistent set of rules. The resulting rules thus lack representativeness and stability.
Soroush Bateni, Sampath Grandhi, Xiaodi Li 0002, Cong Liu 0005, Wei Yang 0013
ESEC/SIGSOFT FSE5
2020 NeuOS: A Latency-Predictable Multi-Dimensional Optimization Framework for DNN-driven Autonomous Systems
Soroush Bateni, Cong Liu 0005
USENIX ATC2
2020 MoLoc: Unsupervised Fingerprint Roaming for Device-Free Indoor Localization in a Mobile Ship Environment
abstract
Device-free indoor localization may play a critical role in improving passengers' safety in large vessels, particularly for scenarios without equipped radios. However, due to dynamic internal and external influences from the sailing ship such as changing sailing speed, the existing localization systems suffer huge accuracy degradation in a mobile ship environment. The challenges are mainly due to rich and arbitrary ship motions and the resulting complicated impacts on the indoor wireless channels. To address the challenges, in this article, we first propose a ship motion descriptor to extract discriminative latent representation from complex ship motions by leveraging deep-learning techniques. Based on this representation, we then design a novel fingerprint roaming model, i.e., MoLoc, to automatically learn the predictive fingerprint variation pattern and transfer the online fingerprint measurement to adapt to dynamic ship motions in real time. Furthermore, an unsupervised learning strategy is proposed to train the fingerprint roaming model using unlabeled onboard collected data which do not incur any labor costs. We have implemented and extensively evaluated MoLoc on real-world cruise ships, where experimental results demonstrate that MoLoc improves localization accuracy from 63.2% to 92.8% compared to the state-of-the-art localization methods, including Pilot, LiFS, SpotFi, and AutoFi while achieving a mean error of 0.68 m.
Mozi Chen, Kezhong Liu, Xuming Zeng, Zheng Dong 0002, Guangmo Tong, Cong Liu 0005
IEEE Internet Things J.7
2020 Enabling Latency-Aware Data Initialization for Integrated CPU/GPU Heterogeneous Platform
abstract
Nowadays, driven by the needs of autonomous driving and edge intelligence, integrated CPU/GPU heterogeneous platform has gained significant attention from both academia and industry. As the representative series, NVIDIA Jetson family perform well in terms of computation capability, power consumption, and mobile size. Even so, the integrated heterogeneous platform only contains one limited physical memory, which is shared by the CPU and GPU cores and can be the performance bottleneck of the mobile/edge applications. On the other hand, with the unified memory (UM) model introduced in GPU programming, not only the memory allocation is significantly reduced, which mitigates the memory bottleneck of the integrated platforms but also the memory management and programming are simplified. However, as a programming legacy, the UM model still follows the conventional copy-then-execute model, initializing data on the CPU side after allocating memory. This legacy programming mode not only causes significant initialization latency but also slows the execution of the following kernel. In this article, we propose a framework to enable the latency-aware data initialization on the integrated heterogeneous platform. The framework not only includes three data initialization modes, the CPU initialization, GPU initialization, and hybrid initialization, but also utilizes an affinity estimation model to wisely decide the best initialization mode for an application such that the initialization latency performance of the application can be optimized. We evaluate our design on NVIDIA TX2 and AGX platforms. The results demonstrate that the framework can accurately select a data initialization mode for a given application to significantly reduce the initialization latency. We envision this latency-aware data initialization framework being adopted in a full-version of autonomous solution (e.g., Autoware) in the future.
Zihang Jiang, Zhen Wang 0019, Xulong Tang, Cong Liu 0005, Shouyi Yin, Yang Hu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2019 Effective Recycling Planning for Dockless Sharing Bikes
abstract
Bike-sharing systems become more and more popular in the urban transportation system, because of their convenience in recent years. However, due to the high daily usage and lack of effective maintenance, the number of bikes in good condition decreases significantly, and vast piles of broken bikes appear in many big cities. As a result, it is more difficult for regular users to get a working bike, which causes problems both economically and environmentally. Therefore, building an effective broken bike prediction and recycling model becomes a crucial task to promote cycling behavior. In this paper, we propose a predictive model to detect the broken bikes and recommend an optimal recycling program based on the large scale real-world sharing bike data. We incorporate the realistic constraints to formulate our problem and introduce a flexible objective function to tune the trade-off between the broken probability and recycled numbers of the bikes. Finally, we provide extensive experimental results and case studies to demonstrate the effectiveness of our approach.
Cong Zhang 0003, Jie Bao 0003, Sijie Ruan, Tianfu He, Hui Lu 0005, Zhihong Tian 0001, Cong Liu 0005, Jianfeng Lin 0004, Xianen Li
SIGSPATIAL/GIS8
2019 Automating CUDA Synchronization via Program Transformation
abstract
While CUDA has been the most popular parallel computing platform and programming model for general purpose GPU computing, CUDA synchronization undergoes significant challenges for GPU programmers due to its intricate parallel computing mechanism and coding practices. In this paper, we propose AuCS, the first general framework to automate synchronization for CUDA kernel functions. AuCS transforms the original LLVM-level CUDA program control flow graph in a semantic-preserving manner for exploring the possible barrier function locations. Accordingly, AuCS develops mechanisms to correctly place barrier functions for automating synchronization in multiple erroneous (challenging-to-be-detected) synchronization scenarios, including data race, barrier divergence, and redundant barrier functions. To evaluate the effectiveness and efficiency of AuCS, we conduct an extensive set of experiments and the results demonstrate that AuCS can automate 20 out of 24 erroneous synchronization scenarios.
Mingyuan Wu, Lingming Zhang 0001, Cong Liu 0005, Shin Hwei Tan, Yuqun Zhang
ASE3
2019 Diagnosing Vehicles with Automotive Batteries
abstract
The automotive industry is increasingly employing software- based solutions to provide value-added features on vehicles, especially with the coming era of electric vehicles and autonomous driving. The ever-increasing cyber components of vehicles (i.e., computation, communication, and control), however, incur new risks of anomalies, as demonstrated by the millions of vehicles recalled by different manufactures. To mitigate these risks, we design B-Diag, a battery-based diagnostics system that guards vehicles against anomalies with a cyber-physical approach, and implement B-Diag as an add-on module of commodity vehicles attached to automotive batteries, thus providing vehicles an additional layer of protection. B-Diag is inspired by the fact that the automotive battery operates in strong dependency with many physical components of the vehicle, which is observable as correlations between battery voltage and the vehicle's corresponding operational parameters, e.g., a faster revolutions-per-minute (RPM) of the engine, in general, leads to a higher battery voltage. B-Diag exploits such physically-induced correlations to diagnose vehicles by cross-validating the vehicle information with battery voltage, based on a set of data-driven norm models constructed online. Such a design of B-Diag is steered by a dataset collected with a prototype system when driving a 2018 Subaru Crosstrek in real-life over 3 months, covering a total mileage of about 1, 400 miles. Besides the Crosstrek, we have also evaluated B-Diag with driving traces of a 2008 Honda Fit, a 2018 Volvo XC60, and a 2017 Volkswagen Passat, showing B-Diag detects vehicle anomalies with >86% (up to 99%) averaged detection rate.
Liang He 0002, Linghe Kong, Yuanchao Shu, Cong Liu 0005
MobiCom5
2019 PIFA: An Intelligent Phase Identification and Frequency Adjustment Framework for Time-Sensitive Mobile Computing
abstract
Due to the limited battery capacity of mobile devices, various CPU power governors and dynamic frequency adjustment schemes have been proposed to reduce CPU energy consumption. However, most such schemes are app-oblivious, ignoring an important fact that real-world applications often exhibit multiple execution phases that perform different functionality and may request different amounts of hardware resources. Having a unified app-level frequency setting for different phases of an application may not be energy efficient enough and may even violate the desirable latency performance required by certain phases. Motivated by this observation, in this paper, we present PIFA, which is an intelligent Phase Identification and Frequency Adjustment framework for energy-efficient and time-sensitive mobile computing. PIFA addresses two major challenges of fully automatically identifying different execution phases of an application and efficiently integrating the phase identification results for runtime frequency adjustment. We have fully implemented PIFA on the Android platform. An extensive set of experiments using real-world Android applications from multiple app categories demonstrate that PIFA achieves closely better performance than the desired latency requirement specified for each phase, while dramatically reducing energy consumption (e.g., >30% energy reduction for most apps) and incurring rather small runtime overhead (e.g., <;5% overhead for most apps).
Xia Zhang 0001, Xusheng Xiao, Liang He 0002, Yun Ma 0002, Yangyang Huang, Xuanzhe Liu, Wenyao Xu, Cong Liu 0005
RTAS8
2019 Predictable Data-Driven Resource Management: an Implementation using Autoware on Autonomous Platforms
abstract
Autonomous embedded systems (AES) are becoming prominent in many application domains such as self-driving cars. However, the conflict between the rather limited memory space in such systems and the data intensive nature of the workloads creates hard challenges on data and memory management, which may easily cause unpredictability in outputting autonomous control decisions. In this paper, we target data-driven AES featuring the integrated architecture by establishing a data-centric system model inspired by Heijunka, a mature production leveling methodology developed by Toyota. Based on this new model, we develop ResCue which contains a dynamic data scheduler and a flexible memory reservation scheme to ensure both temporal and spatial data availability, which shall guarantee predictability in generating outputs in terms of both meeting deadlines and minimizing jitters. We implement and extensively evaluate ResCue under various settings using a popular end-to-end self-driving software Autoware on top of the AES-specific NVIDIA AGX Xavier SoC. Results show that ResCue never misses a deadline and yields a maximum jitter of merely 834 microseconds, while incurring rather small overhead. Moreover, ResCue is able to noticeably reduce memory consumption compared to vanilla Autoware.
Soroush Bateni, Cong Liu 0005
RTSS2
2019 An Efficient Utilization-Based Test for Scheduling Hard Real-Time Sporadic DAG Task Systems on Multiprocessors
abstract
The scheduling and schedulability analysis of real-time directed acyclic graph (DAG) task systems have received much recent attention. The DAG model can accurately represent intra-task parallelism and precedence constraints existing in many application domains. Existing techniques show that analyzing the DAG model is fundamentally more challenging compared to the ordinary sporadic task model, due to the complex intra-DAG precedence constraints which may cause rather pessimistic schedulability loss. However, such increased loss is counter-intuitive because the DAG structure shall better exploit the hardware parallelism provided by the multiprocessor platform. Our key observation is that the intra-DAG precedence constraints, if not carefully considered by the scheduling algorithm, may cause unpredictable execution behaviors of sub-tasks in a DAG and thus pessimistic analysis. In this paper, we present a set of novel scheduling and analysis techniques for better supporting hard real-time sporadic DAG tasks on multiprocessors, through smartly defining and analyzing the execution order of subtasks in each DAG. Combined with a new DAG-specific interval analysis framework, the proposed subtask ordering technique leads to a highly efficient utilization-based schedulability test. Importantly, the developed test becomes identical to the classical density test designed for the sporadic task model, if each DAG in the system has an out-degree of one (i.e., only containing a chain of subtasks). Experiments show the efficiency of the developed test, which improves schedulability upon existing utilization-based tests by over 60% on average and is often able to guarantee schedulability with little utilization loss.
Zheng Dong 0002, Cong Liu 0005
RTSS2
2019 Work-in-Progress: Non-preemptive Scheduling of Sporadic Gang Tasks on Multiprocessors
abstract
Existing works on gang task scheduling mainly focus on the preemptive scheduling case, which contradicts a bit with the non-preemptive executing nature of applying gang scheduling techniques in practice. In this paper, we present a set of non-trivial techniques that can analyze the schedulability of scheduling a hard real-time sporadic gang task system under non-preemptive GEDF on multiprocessors and a utilization-based schedulability test (first-of-its-kind) is derived.
Zheng Dong 0002, Cong Liu 0005
RTSS2
2019 Many suspensions, many problems: a review of self-suspending tasks in real-time systems
abstract
In general computing systems, a job (process/task) may suspend itself whilst it is waiting for some activity to complete, e.g., an accelerator to return data. In real-time systems, such self-suspension can cause substantial performance/schedulability degradation. This observation, first made in 1988, has led to the investigation of the impact of self-suspension on timing predictability, and many relevant results have been published since. Unfortunately, as it has recently come to light, a number of the existing results are flawed. To provide a correct platform on which future research can be built, this paper reviews the state of the art in the design and analysis of scheduling algorithms and schedulability tests for self-suspending tasks in real-time systems. We provide (1) a systematic description of how self-suspending tasks can be handled in both soft and hard real-time systems; (2) an explanation of the existing misconceptions and their potential remedies; (3) an assessment of the influence of such flawed analyses on partitioned multiprocessor fixed-priority scheduling when tasks synchronize access to shared resources; and (4) a discussion of the computational complexity of analyses for different self-suspension task models.
Jian-Jia Chen, Geoffrey Nelissen, Wen-Hung Kevin Huang, Maolin Yang 0004, Björn B. Brandenburg, Konstantinos Bletsas 0001, Cong Liu 0005, Pascal Richard, Frédéric Ridouard, Neil C. Audsley, Ragunathan Rajkumar, Dionisio de Niz, Georg von der Brüggen
Real Time Syst.7
2019 Analysis techniques for supporting hard real-time sporadic gang task systems
Zheng Dong 0002, Cong Liu 0005
Real Time Syst.2
2019 Extending Battery System Operation via Adaptive Reconfiguration
abstract
Large-scale battery packs are commonly used in applications such as electric vehicles (EVs) and smart grids. Traditionally, to provide stable voltage to the loads, voltage regulators are used to convert battery packs’ output voltage to those of the loads’ required levels, causing power loss especially when the difference between the supplied and required voltages is large or when the load is light. In this article, we address this issue via a reconfiguration framework for the battery system. By abstracting the battery system as a cell graph, we develop an adaptive reconfiguration algorithm to identify the desired system configurations based on real-time load requirements. Our design is evaluated via both prototype-based experiments, EV driving trace-based emulations, and large-scale simulations. The results demonstrate an extended system operation time of up to 5×, especially when facing severe cell imbalance.
Liang He 0002, Linghe Kong, Yu Gu 0001, Cong Liu 0005, Tian He 0001, Kang G. Shin
ACM Trans. Sens. Networks4
2019 A General Analysis Framework for Soft Real-Time Tasks
abstract
Much recent work has been conducted on supporting soft real-time tasks on multiprocessors due to the multicore revolution. While most earlier works focus on the traditional sporadic task model with deterministic worst-case specification, several recent works investigate the stochastic nature of many workloads seen in practice, specifying task execution times using average-case provisioning instead of the worst case. Unfortunately, all the existing work on supporting soft real-time workloads ignores a simple practical fact that the job inter-arrival time (or task period) is also stochastic for many real-world applications. Adopting a fixed worst-case period to model all the arriving pattern is rather pessimistic and may result in significant capacity loss in practice. Based on these observations, we present a general soft real-time multiprocessor schedulability analysis framework in this paper for practical sporadic task systems specified by stochastic period and execution demand, following probability distributions. Our analysis can be generally applied to global tunable priority-based schedulers, which allow any job's priority to be changed dynamically at runtime within a priority window of constant length. We have extensively evaluated the analysis framework using a MPEG video decoding case study and simulation-based experiments. Experimental results demonstrate significant advantages of our analysis, which yields over 200 and 50 percent improvements compared to existing analysis assuming worst-case task periods in terms of schedulability and magnitude of the derived tardiness bound, respectively.
Zheng Dong 0002, Cong Liu 0005, Soroush Bateni, Zelun Kong, Liang He 0002, Lingming Zhang 0001, Ravi Prakash 0001, Yuqun Zhang
IEEE Trans. Parallel Distributed Syst.2
2018 GRU: Exploring Computation and Data Redundancy via Partial GPU Computing Result Reuse
abstract
Graphics processing units (GPUs) have been widely adopted by major cloud vendors for better performance and energy efficiency. Recent research has observed a considerable degree of redundancy in managing computation and data in many datacenters, particularly for several important categories of GPU-accelerated applications such as log mining and machine learning. In this paper, we present GRU, an ecosystem that smartly manages and shares GPU resources through exploiting redundancy. GRU transparently interprets GPU-accelerated computing requests and memoizes results for potential future reuse. To enhance reusability, GRU implements a partial result reuse idea, where GPU computation requests even with different input data and functionality may become reusable w.r.t. each other. To guarantee correctness of partial reuse, GRU employs a compiler-assisted approach that analyzes general data parallel patterns that are reliable for the reuse purpose, and is capable of smartly recognizing such reusable data parallel patterns of incoming requests. We have fully implemented GRU and conducted extensive sets of experiments running micro-benchmarks on local machines and real-world applications including Spark-based uses cases in an AWS cluster. Evaluation results show that GRU is effective in identifying and eliminating redundant GPU computations, achieving up to 5x (2.5x) speedup for compute-intensive (data-intensive) benchmarks. In addition, GPU-managed Spark observes a reduction of 25.3% (39.8%) on average w.r.t. turnaround time (GPU occupation time) over state-of-the-art solutions.
Husheng Zhou, Soroush Bateni, Cong Liu 0005
ICS3
2018 DeepRoad: GAN-based metamorphic testing and input validation framework for autonomous driving systems
abstract
While Deep Neural Networks (DNNs) have established the fundamentals of image-based autonomous driving systems, they may exhibit erroneous behaviors and cause fatal accidents. To address the safety issues in autonomous driving systems, a recent set of testing techniques have been designed to automatically generate artificial driving scenes to enrich test suite, e.g., generating new input images transformed from the original ones. However, these techniques are insufficient due to two limitations: first, many such synthetic images often lack diversity of driving scenes, and hence compromise the resulting efficacy and reliability. Second, for machine-learning-based systems, a mismatch between training and application domain can dramatically degrade system accuracy, such that it is necessary to validate inputs for improving system robustness.
Mengshi Zhang, Yuqun Zhang, Lingming Zhang 0001, Cong Liu 0005, Sarfraz Khurshid
ASE4
2018 Shared-Resource-Centric Limited Preemptive Scheduling: A Comprehensive Study of Suspension-Based Partitioning Approaches
abstract
This paper studies the problem of scheduling a set of hard real-time sporadic tasks that may access CPU cores and a shared resource. Motivated by the observation that the CPU resource is often abundant compared to the shared resources in multi-core and many-core systems, we propose to resolve this problem from a counter-intuitive shared-resource-centric perspective, focusing on judiciously prioritizing and scheduling tasks' requests in a limited preemptive manner on the shared resource while viewing the worst-case latency a task may experience on the CPU cores as suspension delays. We develop a rather comprehensive set of task partitioning algorithms that partition tasks onto the shared resource with the objective of guaranteeing schedulability while minimizing the required size of the shared resource, which plays a critical role in reducing the overall cost and complexity of building resource-constrained embedded systems in many application domains. A GPU-based prototype case study and extensive simulation-based experiments have been conducted, which validate both our shared-resource-centric scheduling philosophy and the efficiency of our suspension-based partitioning solutions in practice.
Zheng Dong 0002, Cong Liu 0005, Soroush Bateni, Kuan-Hsun Chen, Jian-Jia Chen, Georg von der Brüggen
RTAS2
2018 S^3DNN: Supervised Streaming and Scheduling for GPU-Accelerated Real-Time DNN Workloads
abstract
Deep Neural Networks (DNNs) are being widely applied in many advanced embedded systems that require autonomous decision making, e.g., autonomous driving and robotics. To handle resource-demanding DNN workloads, graphic processing units (GPUs) have been used as the main acceleration engine. Although much research has been conducted to algorithmically optimize the efficiency of applying DNN to applications such as object recognition, limited attention has been given to optimizing the execution of GPU-accelerated DNN workloads at the system level. In this paper, we propose S^3DNN, a system solution that optimizes the execution of DNN workloads on GPU in a real-time multi-tasking environment, which simultaneously optimizes the two (sometimes) conflicting goals of real-time correctness and throughput. S^3DNN contains a governor that selectively gathers system-wide DNN requests to perform smart data fusion, and a novel supervised streaming and scheduling framework that combines a deadline-aware scheduler with the concurrency-enabled CUDA stream technique. To simultaneously maximize concurrency-induced benefits and real-time performance, S^3DNN explores a rather interesting and unique characteristic of DNN workloads, where multiple layers of a DNN instance often exhibit a gradually decreased GPU resource utilization pattern. We have fully implemented S^3DNN in a GPU-accelerated system and have conducted extensive sets of experiments evaluating the efficacy of S^3DNN under a wide range of system and workload scenarios. The results show that S^3DNN significantly improves upon state-of-the-art GPU-accelerated DNN processing frameworks, e.g., up to 37% and over 40% improvements in real-time performance and throughput, respectively.
Husheng Zhou, Soroush Bateni, Cong Liu 0005
RTAS3
2018 ApNet: Approximation-Aware Real-Time Neural Network
abstract
Modern embedded cyber-physical systems are becoming entangled with the realm of deep neural networks (DNNs) towards increased autonomy. While applying DNNs can significantly improve the accuracy in making autonomous control decisions, a significant challenge is that DNNs are designed and developed on advanced hardware (e.g., GPU clusters), and will not easily meet strict timing requirements if deployed in a resource-constrained embedded computing environment. One interesting characteristic of DNNs is approximation, which can be used to satisfy real-time requirements by reducing DNNs' execution costs with reasonably sacrificed accuracy. In this paper, we propose ApNet, a timing-predictable runtime system that is able to guarantee deadlines of DNN workloads via efficient approximation. Rather than straightforwardly approximating DNNs, ApNet develops a DNN layer-aware approximation approach that smartly explores the trade-off between the approximation degree and the resulting execution reduction on a per-layer basis. To further reduce approximation-induced accuracy loss at runtime, ApNet explores a rather interesting observation that resource sharing and approximation can mutually supplement one another, particularly in a multi-tasking environment. We have implemented and extensively evaluated ApNet on a mix of 8 different DNN configurations on an NVIDIA Jetson TX2. Experimental results show that ApNet can guarantee timing predictability (i.e., meeting all deadlines), while incurring a reasonable accuracy loss. Moreover, accuracy can be improved by up to 8% via a resource sharing increase of 3.5x on average for overlapping DNN layers.
Soroush Bateni, Cong Liu 0005
RTSS2
2018 PredJoule: A Timing-Predictable Energy Optimization Framework for Deep Neural Networks
abstract
The revolution of deep neural networks (DNNs) is enabling dramatically better autonomy in autonomous driving. However, it is not straightforward to simultaneously achieve both timing predictability (i.e., meeting job latency requirements) and energy efficiency that are essential for any DNN-based autonomous driving system, as they represent two (often) conflicting goals. In this paper, we propose PredJoule, a timing-predictable energy optimization framework for running DNN workloads in a GPU-enabled automotive system. PredJoule achieves both latency guarantees and energy efficiency through a layer-aware design that explores specific performance and energy characteristics of different layers within the same neural network. We implement and evaluate PredJoule on the automotive-specific NVIDIA Jetson TX2 platform for five state-of-the-art DNN models with both high and low variance latency requirements. Experiments show that PredJoule rarely violates job deadlines, and can improve energy by 65% on average compared to five existing approaches and 68% compared to an energy-oriented approach.
Soroush Bateni, Husheng Zhou, Yuankun Zhu, Cong Liu 0005
RTSS4
2018 Work-in-Progress: New Analysis Techniques for Supporting Hard Real-Time Sporadic DAG Task Systems on Multiprocessors
abstract
We consider the problem of globally scheduling hard real-time sporadic DAG task systems on multiprocessors. Existing techniques show that analyzing the DAG model is fundamentally more challenging compared to the ordinary sporadic task model, due to the complex intra-DAG precedence constraints which may cause rather pessimistic schedulability loss. However, such increased loss is counterintuitive because the DAG structure shall better exploit the parallelism provided by the multiprocessor platform. In this work, we present a set of novel scheduling and analysis techniques for better supporting hard real-time sporadic DAG tasks on multiprocessors, through smartly defining and analyzing the execution order of subtasks in each DAG. Interestingly, when each DAG task only contains a single subtask, the proposed utilization-based schedulability test becomes identical to the density test.
Zheng Dong 0002, Cong Liu 0005
RTSS2
2018 Low-Overhead WiFi Fingerprinting
abstract
WiFi-fingerprint localization is recognized as a promising indoor localization technique. However, it suffers from high implementation overhead such as heavy initial training and fingerprint map maintenance overtime. In this paper, we present the design, implementation, and evaluation of AP-Sequence. It is a fingerprint-based localization system that achieves extremely low overhead in fingerprint map construction and maintenance. AP-Sequence achieves this by treating a scan from any reference locations as an input to adjust a large portion of the fingerprint map. The power of AP-Sequence comes from dynamic region partitioning mechanism generating a fingerprint based on relative RSS values. AP-Sequence offers several advantages over existing methods with respect to robustness against environment noises, ability to handle dynamic power control, and mobile device heterogeneity. We have implemented AP-Sequence on an Android platform. Experiment results with over one month of evaluation demonstrate that our design achieves an average localization accuracy of 4-7.6 m over an extended time period with low-overhead in fingerprint map construction and maintenance.
Jung-Hyun Jun, Liang He 0002, Yu Gu 0001, Wenchao Jiang, Gaurav Kushwaha, Vipin A, Long Cheng 0005, Cong Liu 0005, Ting Zhu 0001
IEEE Trans. Mob. Comput.8
2017 Optimal Dataflow Scheduling on a Heterogeneous Multiprocessor With Reduced Response Time Bounds
abstract
Heterogeneous computing platforms with multiple types of computing resources have been widely used in many industrial systems to process dataflow tasks with pre-defined affinity of tasks to subgroups of resources. For many dataflow workloads with soft real-time requirements, guaranteeing fast and bounded response times is often the objective. This paper presents a new set of analysis techniques showing that a classical real-time scheduler, namely earliest-deadline first (EDF), is able to support dataflow tasks scheduled on such heterogeneous platforms with provably bounded response times while incurring no resource capacity loss, thus proving EDF to be an optimal solution for this scheduling problem. Experiments using synthetic workloads with widely varied parameters also demonstrate that the magnitude of the response time bounds yielded under the proposed analysis is reasonably small under all scenarios. Compared to the state-of-the-art soft real-time analysis techniques, our test yields a 68% reduction on response time bounds on average. This work demonstrates the potential of applying EDF into practical industrial systems containing dataflow-based workloads that desire guaranteed bounded response times.
Zheng Dong 0002, Cong Liu 0005, Alan Gatherer, Lee McFearin, Peter Yan, James H. Anderson
ECRTS2
2017 An efficient randomized algorithm for rumor blocking in online social networks
abstract
Social networks allow rapid spread of ideas and innovations while the negative information can also propagate widely. When the cascades with different opinions reaching the same user, the cascade arriving first is the most likely to be taken by the user. Therefore, once misinformation or rumor is detected, a natural containment method is to introduce a positive cascade competing against the rumor. Given a budget k, the rumor blocking problem asks for k seed users to trigger the spread of the positive cascade such that the number of the users who are not influenced by rumor can be maximized. The prior works have shown that the rumor blocking problem can be approximated within a factor of (1 - 1/e- δ) by a classic greedy algorithm combined with Monte Carlo simulation with the running time of O(k3mn ln n/δ2), where n and m are the number of users and edges, respectively. Unfortunately, the Monte-Carlo-simulation-based methods are extremely time consuming and the existing algorithms either trade performance guarantees for practical efficiency or vice versa. In this paper, we present a randomized algorithm which runs in O(km ln n/δ2) expected time and provides a (1 - 1/e - δ)-approximation with a high probability. The experimentally results on both the real-world and synthetic social networks have shown that the proposed randomized rumor blocking algorithm is much more efficient than the state-of-the-art method and it is able to find the seed nodes which are effective in limiting the spread of rumor.
Guangmo Tong, Weili Wu 0001, Deying Li 0001, Cong Liu 0005, Bin Liu 0009, Ding-Zhu Du
INFOCOM5
2017 State of the art for scheduling and analyzing self-suspending sporadic real-time tasks
abstract
In computing systems, a job/process/task/thread may suspend itself when it has to wait for some other internal or external activities, such as computation offloading or memory accesses, to finish before it can continue its execution. In the literature, there are two commonly adopted self-suspending sporadic task models in real-time systems: 1) the dynamic self-suspension model and 2) the segmented self-suspension sporadic task model. A dynamic self-suspending sporadic task is specified with an upper bound on the maximum suspension time for a job (task instance), which allows a job to dynamically suspend itself arbitrary often as long as the suspension time upper bound is not violated. By contrast, a segmented self-suspending sporadic task has a predefined execution and suspension pattern in an interleaving manner. The dynamic self-suspension model is very flexible but inaccurate, whilst the segmented self-suspension model is very restrictive but very accurate. The gap between these two widely-adopted self-suspension task models can be potentially filled by the hybrid self-suspension task model. The investigation of the impact of self-suspension on timing predictability has been started in 1988. This survey paper provides a short summary of the state of the art in the design and analysis of scheduling algorithms and schedulability tests for self-suspending tasks in real-time systems.
Jian-Jia Chen, Georg von der Brüggen, Wen-Hung Kevin Huang, Cong Liu 0005
RTCSA4
2017 Fixed-priority scheduling of mixed soft and hare real-time tasks on multiprocessors
abstract
1This paper answers several open questions of practical concerns to schedule soft real-time (SRT) tasks, to guarantee their bounded tardiness, under fixed-priority scheduling in homogeneous multiprocessor systems. We consider both cases with only SRT tasks and with mixed sets of SRT and hard real-time (HRT) tasks. For the case in which the system has only SRT tasks, we show that any fixed priority assignment policy yields a capacity augmentation factor of 2−1/M where M is the number of processors. We prove the optimality of the utilization-monotonic (UM) priority assignment (i.e., assigning higher priorities to high-utilization tasks) under our sufficient test for guaranteeing bounded tardiness. We show that UM priority assignment can yield a utilization bound of M+1/2M, which is shown asymptotically the best possible bound. For the case in which the system has mixed SRT and HRT tasks, we present two new fixed-priority assignment algorithms and their associated schedulability tests. One is a clustering-based greedy priority assignment policy and another is based on Audsley's optimal priority assignment (OPA) approach. We show that the utilization bounds, augmentation factors, and speedup factors are still maintained by the hard real-time cases. Therefore, introducing soft real-time tasks does not create additional problems (at least in those metrics) for scheduling if the priority assignments are properly done. As demonstrated by extensive experiments, these two policies yield reasonably good performance overall and much better performance than the deadline-monotonic priority assignment.
Jian-Jia Chen, Wen-Hung Kevin Huang, Zheng Dong 0002, Cong Liu 0005
RTCSA4
2017 Analysis Techniques for Supporting Hard Real-Time Sporadic Gang Task Systems
abstract
This paper studies the problem of scheduling hard real-time sporadic gang task systems under global earliest-deadline-first, where a gang application's threads need to be concurrently scheduled on distinct processors. A novel approach combining new lag-based reasoning and executing/non-executing gang interval analysis technique is introduced, which is able to characterize the parallelisminduced idleness, as a key challenge of analyzing gang task schedules. To the best of our knowledge, this approach yields the first utilization-based test for hard real-time gang task systems.
Zheng Dong 0002, Cong Liu 0005
RTSS2
2017 REC: Predictable Charging Scheduling for Electric Taxi Fleets
abstract
Due to the energy security concern, our society is witnessing a surge of EV fleet applications, e.g., public EV taxi fleet systems. A major issue impeding an even more widespread adoption of EVs is range anxiety, which is due to several factors including limited battery capacity, limited availability of battery charging stations, and long charging time compared to traditional gasoline vehicles. By analyzing our accessible real-world EV taxi system-wide datasets, we observe that current EV taxi drivers often suffer from unpredictable, long waiting times at charging stations, due to temporally and spatially unbalanced utilization among charging stations. This is mainly because current taxi fleet management system simply rely on taxi drivers to make charging decisions. In this paper, In this paper, we develop REC, a Real-time Ev Charging scheduling framework for EV taxi fleets, which informs each EV taxi driver at runtime when and where to charge the battery. REC is able to analytically guarantee predictable and tightly bounded waiting times for all EVs in the fleet and temporally/spatially balanced utilization among charging stations, if each driver follows the charging decision made by REC. Moreover, REC can further efficiently handle real-life issues, e.g., allowing a taxi driver to charge at its preferred charging station while still guaranteeing balanced charging station utilization.We have extensively evaluated REC using our accessible real-world EV taxi system-wide datasets. Experimental results show that REC is able to address the unpredictability and unbalancing issues existing in current EV taxi fleet systems, yielding predictable and tightly bounded waiting times, and equally important, temporally/spatially balanced charging station utilization.
Zheng Dong 0002, Cong Liu 0005, Jie Bao 0003, Yu Gu 0001, Tian He 0001
RTSS2
2017 Battery-Aware Mobile Data Service
abstract
Significant research has been devoted to reduce the energy consumption of mobile devices, but how to increase their energy supply has received far less attention. Moreover, reducing the energy consumption alone does not always extend the device operation time due to a unique battery property - the capacity it delivers hinges critically upon how it is discharged. In this paper, we propose B-MODS, a novel design of battery-aware mobile data service on mobile devices. B-MODS constructs battery-friendly discharge patterns utilizing the recovery effect so as to increase the capacity delivered from batteries while meeting data service requirements. We implement B-MODS as an application layer library on the Android platform. Our experiments with diverse mobile devices under various application scenarios have shown that B-MODS increases the capacity delivery from the battery by up to 49.5 percent, with which an increase in the user-perceived data service utilities of up to 28.6 percent is observed.
Liang He 0002, Guozhu Meng, Yu Gu 0001, Cong Liu 0005, Jun Sun 0001, Ting Zhu 0001, Yang Liu 0003, Kang G. Shin
IEEE Trans. Mob. Comput.4
2016 Taming collisions for delay reduction in low-duty-cycle wireless sensor networks
abstract
Many-to-one data collection is a fundamental operation in wireless sensor networks (WSNs). To support long-term deployment of WSNs, sensor nodes normally operate at low-duty-cycles. However, the low-duty-cycle operation significantly reduces the communication chance between nodes. Consequently, the risk of data collisions significantly increases when multiple senders transmit packets to a receiver during its very short active period. Data collision not only results in wasted packet transmissions, but also incurs a large delivery latency. Under such conditions, collision-free medium access is more appealing than recovering after collision for low-duty-cycle WSNs. In this work, we propose an incast-collision-free data collection protocol, named iCore, to address the many-to-one collision problem in low-duty-cycle WSNs. iCore employs the dynamic forwarding technique and establishes a non-conflicting schedule for delay reduction. Specifically, we design efficient forwarder assignment and forwarding optimization algorithms that ensure low end-to-end latency under diverse data traffic types. Through comprehensive performance evaluations, we demonstrate that, compared with the state-of-the-art protocol, iCore effectively minimizes the end-to-end delay by 25% ∼ 57% and maintains high delivery ratio and energy efficiency for different many-to-one convergecast scenarios.
Long Cheng 0005, Yu Gu 0001, Jianwei Niu 0002, Ting Zhu 0001, Cong Liu 0005, Tian He 0001
INFOCOM5
2016 Terminal-set-enhanced community detection in social networks
abstract
Community detection aims to reveal the community structure in a social network, which is one of the fundamental problems. In this paper we investigate the community detection problem based on the concept of terminal set. A terminal set is a group of users within which any two users belong to different communities. Although the community detection is hard in general, the terminal set can be very helpful in designing effective community detection algorithms. We first present a 2-approximation algorithm running in polynomial time for the original community detection problem. In the other issue, in order to better support real applications we further consider the case when extra restrictions are imposed on feasible partitions. For such customized community detection problems, we provide two randomized algorithms which are able to find the optimal partition with a high probability. Demonstrated by the experiments performed on benchmark networks the proposed algorithms are able to produce high-quality communities.
Guangmo Tong, Lei Cui 0010, Weili Wu 0001, Cong Liu 0005, Ding-Zhu Du
INFOCOM4
2016 k2Q: A Quadratic-Form Response Time and Schedulability Analysis Framework for Utilization-Based Analysis
abstract
In this paper, we present a general response-time analysis and schedulability-test framework, called k2Q (k to Q). It provides automatic constructions of closed-form quadratic bounds or utilization bounds for a wide range of applications in real-time systems under fixed-priority scheduling. The key of the framework is a k-point schedulability test or a k-point response time analysis that is based on the utilizations and the execution times of k-1 higher-priority tasks. The natural condition of k2Q is a quadratic form for testing the schedulability or analyzing the response time. The response time analysis and the schedulability analysis provided by the framework can be viewed as a "blackbox'' interface that can result in sufficient utilization-based analysis. Since the framework is independent from the task and platform models, it can be applied to a wide range of applications.
Jian-Jia Chen, Wen-Hung Kevin Huang, Cong Liu 0005
RTSS3
2016 Enabling Predictable Wireless Data Collection in Severe Energy Harvesting Environments
abstract
Micro-powered wireless embedded devices are widely used in many application domains. Their efficiency in practice, however, is significantly constrained by the dual limitations of low harvesting rates and tiny energy buffer. Recent research presents a network stack that efficiently fragments a large packet into many smaller packets that can fit within the available energy in the energy buffer of limited size. While this fragmentation technique represents a major step forward in solving the minuscule energy budget problem, it also introduces a tremendous practical challenge where potentially many fragmented packets belonging to different devices may contend for the communication channel. Designing purely heuristic-based packet transmission protocol is undesirable because the resulting per-packet and end-to-end transmission delay are unknown, thus causing unpredictable system performance which is unacceptable for many applications with real-time constraints. In this paper, we first formulate this packet transmission scheduling problem considering physical properties of the charging and transmission processes. We then develop a novel packet prioritization and transmission protocol NERF that yields tight and predictable delay bounds for transmitting packets from multiple micropowered devices to a charger. We have implemented our protoco on top of the WISP 4.1 platform and the SPEEDWAY RFID READER, and conducted validation experiments. Our experiments validate the correctness of our implementation and show that NERF can reduce the total collection delay by 40% when compared to an existing protocol ALOHA. We have also performed extensive data trace-driven simulations. Simulation results demonstrate the effectiveness of our proposed protocol. On average, our protocol yields an over 30%improvement in terms of runtime transmission delay compared to existing methods, while being able to guarantee tight and provable response time bounds.
Zheng Dong 0002, Yu Gu 0001, Jiming Chen 0001, Shaojie Tang 0001, Tian He 0001, Cong Liu 0005
RTSS6
2016 Closing the Loop for the Selective Conversion Approach: A Utilization-Based Test for Hard Real-Time Suspending Task Systems
abstract
This paper studies the problem of scheduling hard real-time sporadic suspending task systems under global earliest-deadline-first. A novel selective suspension-tocomputation conversion approach has been developed, with the fundamental idea of selecting and converting a limited set of jobs' suspensions into computation to eliminate suspension-induced pessimism in the analysis. To the best of our knowledge, this approach yields the first utilization-based test for globally-scheduled suspending task systems, which analytically dominates the suspension-oblivious approach and dramatically improves schedulability upon existing tests by over 50% on average, as shown by experiments. We believe this paper closes the loop on applying the methodology of selective suspension-to-computation conversion to analyze realtime suspending task systems.
Zheng Dong 0002, Cong Liu 0005
RTSS2
2016 Energy Synchronized Task Assignment in Rechargeable Sensor Networks
abstract
Wireless rechargeable sensor networks have recently emerged as a promising platform that can effectively solve the power constraint problem suffered by traditional battery powered systems. The problem of determining the best charging routes for maximizing charging efficiency has been studied extensively. However, the task assignment problem, which plays a crucial role in efficiently utilizing the harvested energy and thus minimize the charging delay, has received rather limited attention. In this paper, we study the problem of assigning a given set of tasks in a wireless rechargeable sensor network while maximizing the charger's velocity to minimize the charging delay. We first propose an online task assignment algorithm, namely Lower Bound assignment (LB), that yields a quantifiable lower bound on the charging velocity while guaranteeing a feasible assignment. This algorithm further enables the transformation of our considered task assignment problem into a variation of the classical multiple knapsack problem. We then present a fully polynomial-time approximation scheme with a (2+ε)-approximation ratio, namely ACT, that is built upon an existing greedy algorithm designed for the original knapsack problem. Extensive experimental results presented herein demonstrate that ACT is able to achieve near-optimal performance in most cases, and can achieve more than 15% performance improvement compared to the baseline algorithms.
Zheng Dong 0002, Cong Liu 0005, Lingkun Fu, Peng Cheng 0001, Liang He 0002, Yu Gu 0001, Wei Gao 0006, Chau Yuen, Tian He 0001
SECON2
2016 Supporting Soft Real-Time Sporadic Task Systems on Uniform Heterogeneous Multiprocessors with No Utilization Loss
abstract
Uniform heterogeneous multicore architectures are becoming increasingly popular due to their potential of achieving high performance and energy efficiency compared to the homogeneous multicore architectures. In such systems, the real-time scheduling problem becomes more challenging because processors have different speeds. Prior research on uniform heterogeneous multiprocessor real-time scheduling has focused on hard real-time systems, where, significant processing capacity may have to be sacrificed in the worst-case to ensure that all deadlines are met. As meeting hard deadlines is overkill for many soft real-time systems in practice, this paper shows that on soft real-time uniform heterogeneous multiprocessors, bounded response times can be ensured for globally-scheduled sporadic task systems with no utilization loss. A GEDF-based scheduling algorithm, named as GEDF-H, is presented and response time bounds are established under both preemptive and non-preemptive GEDF-H scheduling. Extensive experiments show that the magnitude of the derived response time bound is reasonable, often smaller than four task relative deadlines. To the best of our knowledge, this paper is the first to show that soft real-time sporadic task systems can be supported on uniform heterogeneous multiprocessors without utilization loss under global scheduling, and with reasonable predicted response times.
Guangmo Tong, Cong Liu 0005
IEEE Trans. Parallel Distributed Syst.2
2015 PASS: priority assignment of real-time tasks with dynamic suspending behavior under fixed-priority scheduling
abstract
Self-suspension is becoming an increasingly prominent characteristic in real-time systems such as: (i) I/O-intensive systems, where applications interact intensively with I/O devices, (ii) multi-core processors, where tasks running on different cores have to synchronize and communicate with each other, and (iii) computation offloading systems with coprocessors, like Graphics Processing Units (GPUs). In this paper, we show that rate-monotonic (RM), deadline-monotonic (DM) and laxity-monotonic (LM) scheduling will perform rather poor in dynamic self-suspending systems in terms of speed-up factors. On the other hand, the proposed PASS approach is guaranteed to find a feasible priority assignment on a speed-2 uniprocessor, if one exists on a unit-speed processor. We evaluate the feasibility of the proposed approach via a case study implementation. Furthermore, the effectiveness of the proposed approach is also shown via extensive simulation results.
Wen-Hung Kevin Huang, Jian-Jia Chen, Husheng Zhou, Cong Liu 0005
DAC4
2015 A Computation Offloading Framework for Soft Real-Time Embedded Systems
abstract
Recent developments in embedded hardware have empowered human experiences through pervasive computing. While embedded systems are becoming more powerful, they still fall short when faced with users' growing desire for running more resource-demanding applications. To bridge this gap, one solution is to leverage powerful resources residing at remote sites by performing computation offloading. Unfortunately, the state-of-the-art offloading frameworks cannot be applied in many embedded systems supporting applications with soft real-time (SRT) constraints or high delay sensitivity, as they typically optimize response times on a "best-effort" basis using heuristics. This paper establishes a soft real-time offloading framework that optimizes the resource utilization of the embedded system while analytically guaranteeing SRT schedulability. The key idea behind the proposed framework is to view offloading-induced delays as suspensions occurring at the local embedded system side, which allows a task being offloaded to be modelled as a suspending task and thus existing SRT suspension-aware scheduling and analysis techniques to be leveraged. Based on this idea, we propose an offloading algorithm, namely Real-time Offloading Decision-making Algorithm (RODA), to make offloading decisions such that SRT schedulability of the task system can be ensured. The optimality properties of RODA have been proved on both uniprocessors and multiprocessors. We conducted extensive simulations on evaluating schedulability and implemented a case study offloading system on top of real hardware to test runtime response time performance. Results demonstrated that RODA is superior to existing performance-driven offloading algorithms, particularly under heavy workloads.
Yuchuan Liu, Cong Liu 0005, Xia Zhang 0001, Wei Gao 0006, Liang He 0002, Yu Gu 0001
ECRTS2
2015 GPES: a preemptive execution system for GPGPU computing
abstract
Graphics processing units (GPUs) are being widely used as co-processors in many application domains to accelerate general-purpose workloads that are computationally intensive, known as GPGPU computing. Real-time multi-tasking support is a critical requirement for many emerging GPGPU computing domains. However, due to the asynchronous and non-preemptive nature of GPU processing, in multi-tasking environments, tasks with higher priority may be blocked by lower priority tasks for a lengthy duration. This severely harms the system's timing predictability and is a serious impediment limiting the applicability of GPGPU in many real-time and embedded systems. In this paper, we present an efficient GPGPU preemptive execution system (GPES), which combines user-level and driverlevel runtime engines to reduce the pending time of high-priority GPGPU tasks that may be blocked by long-freezing low-priority competing workloads. GPES automatically slices a long-running kernel execution into multiple subkernel launches and splits data transaction into multiple chunks at user-level, then inserts preemption points between subkernel launches and memorycopy operations at driver-level. We implement a prototype of GPES, and use real-world benchmarks and case studies for evaluation. Experimental results demonstrate that GPES is able to reduce the pending time of high-priority tasks in a multitasking environment by up to 90% over the existing GPU driver solutions, while introducing small overheads.
Husheng Zhou, Guangmo Tong, Cong Liu 0005
RTAS3
2015 k2U: A General Framework from k-Point Effective Schedulability Analysis to Utilization-Based Tests
abstract
To deal with a large variety of workloads in different application domains in real-time embedded systems, a number of expressive task models have been developed. For each individual task model, researchers tend to develop different types of techniques for deriving schedulability tests with different computation complexity and performance. In this paper, we present a general schedulability analysis framework, namely the k2U framework, that can be potentially applied to analyze a large set of real-time task models under any fixed-priority scheduling algorithm, on both uniprocessor and multiprocessor scheduling. The key to k2U is a k-point effective schedulability test, which can be viewed as a "blackbox" interface. For any task model, if a corresponding k-point effective schedulability test can be constructed, then a sufficient utilization-based test can be automatically derived. We show the generality of k2U by applying it to different task models, which results in new and improved tests compared to the state-of-the-art.
Jian-Jia Chen, Wen-Hung Kevin Huang, Cong Liu 0005
RTSS3
2015 CDC: Compressive Data Collection for Wireless Sensor Networks
abstract
Data collection is a crucial operation in wireless sensor networks. The design of data collection schemes is challenging due to the limited energy supply and the hot spot problem. Leveraging empirical observations that sensory data possess strong spatiotemporal compressibility, this paper proposes a novel compressive data collection scheme for wireless sensor networks. We adopt a power-law decaying data model verified by real data sets and then propose a random projection-based estimation algorithm for this data model. Our scheme requires fewer compressed measurements, thus greatly reduces the energy consumption. It allows simple routing strategy without much computation and control overheads, which leads to strong robustness in practical applications. Analytically, we prove that it achieves the optimal estimation error bound. Evaluations on real data sets (from the GreenOrbs, IntelLab and NBDC-CTD projects) show that compared with existing approaches, this new scheme prolongs the network lifetime by 1.5X to 2X for estimation error 5-20 percent.
Xiao-Yang Liu, Yanmin Zhu 0006, Linghe Kong, Cong Liu 0005, Yu Gu 0001, Athanasios V. Vasilakos, Min-You Wu
IEEE Trans. Parallel Distributed Syst.4
2014 Analysis Techniques for Supporting Harmonic Real-Time Tasks with Suspensions
abstract
In many real-time systems, tasks may experience suspension delays when they block to access shared resources or interact with external devices such as I/O. It is known that such suspensions delays may negatively impact schedulability. Particularly in hard real-time systems, a few negative results exist on analyzing the schedulability of such systems, even for very restricted suspending task models on a uniprocessor. In this paper, we focus on the particular case of hard real-time suspending task systems with harmonic periods, which is a special case of practical relevance. We propose a new uniprocessor suspension-aware analysis technique for supporting such task systems under rate-monotonic scheduling. Our analysis technique is able to achieve only Theta(1) suspension-related utilization loss on a uniprocessor. Based upon this technique, we further propose a partitioning scheme that supports suspending task systems with harmonic periods on multiprocessors. The resulting schedulability test shows that compared to existing schedulability tests designed for ordinary non-suspending task systems, suspensions only results in Theta(m) additional suspension-related utilization loss, where m is the number of processors. Furthermore, experiments presented herein show that both our uniprocessor and multiprocessor schedulability tests improve upon prior approaches by a significant margin.
Cong Liu 0005, Jian-Jia Chen, Liang He 0002, Yu Gu 0001
ECRTS1
2014 Supporting read/write applications in embedded real-time systems via suspension-aware analysis
abstract
In many embedded real-time systems, applications often interact with I/O devices via read/write operations, which may incur considerable suspension delays. Unfortunately, prior analysis methods for validating timing correctness in embedded systems become quite pessimistic when suspension delays are present. In this paper, we consider the problem of supporting two common types of I/O applications in a multiprocessor system, that is, write-only applications and read-write applications. For the write-only application model, we present a much improved analysis technique that results in only O(m) suspension-related utilization loss, where m is the number of processors. For the second application model, we present a flexible I/O placement strategy and a corresponding new scheduling algorithm, which can completely circumvent the negative impact due to read- and write-induced suspension delays. We illustrate the feasibility of the proposed I/O-placement-based schedule via a case study implementation. Furthermore, experiments presented herein show that the improvement with respect to system utilization over prior methods is often significant.
Guangmo Tong, Cong Liu 0005
EMSOFT2
2014 Task mapping in heterogeneous embedded systems for fast completion time
abstract
Graphics processing units are being widely used in embedded systems as they can achieve high performance and energy efficiency. In such systems, the problem of computation and data mapping for multiple applications while minimizing the completion time is quite challenging due to a large size of the policy space, including heterogeneous application characteristics, complex application structure, data communication costs, and data partitioning. To achieve fast competition time, a fine-grain mapping framework that explores a set of critical factors is needed for heterogeneous embedded systems. In this paper, we consider this mapping problem by presenting a theoretical framework that yields an optimal integer programming solution. Moreover, based upon several interesting measurements-based case studies, we design three practical mapping algorithms with low time complexity, each of which explores a specific set of factors that may affect the completion time performance. We evaluated the proposed algorithms by implementing them on a real heterogeneous system and using a large set of popular benchmarks for evaluation. Experimental results demonstrate that our proposed algorithms can achieve up to 30% faster completion time compared to the state-of-the-art mapping techniques, and can perform consistently well across different workloads.
Husheng Zhou, Cong Liu 0005
EMSOFT2
2014 On Exploiting Dynamic Execution Patterns for Workload Offloading in Mobile Cloud Applications
abstract
Mobile Cloud Computing (MCC) bridges the gap between limited capabilities of mobile devices and the increasing users' demand of mobile multimedia applications, by offloading the computational workloads from local devices to the remote cloud. Current MCC research focuses on making offloading decisions over different methods of a MCC application, but may inappropriately increase the energy consumption if having transmitted a large amount of program states over expensive wireless channels. Limited research has been done on avoiding such energy waste by exploiting the dynamic patterns of applications' run-time execution for workload offloading. In this paper, we adaptively offload the local computational workload with respect to the run-time application dynamics. Our basic idea is to formulate the dynamic executions of user applications using a semi-Markov model, and to further make offloading decisions based on probabilistic estimations of the offloading operation's energy saving. Such estimation is motivated by experimental investigations over practical smart phone applications, and then builds on analytical modeling of methods' execution times and offloading expenses. Systematic evaluations show that our scheme significantly improves the efficiency of workload offloading compared to existing schemes over various smart phone applications.
Wei Gao 0006, Yong Li 0015, Ting Wang 0006, Cong Liu 0005
ICNP5
2014 Mobile-to-mobile energy replenishment in mission-critical robotic sensor networks
abstract
Recently, much research effort has been devoted to employing mobile chargers for energy replenishment of the robots in robotic sensor networks. Observing the discrepancy between the charging latency of robots and charger travel distance, we propose a novel tree-based charging schedule for the charger, which minimizes its travel distance without causing the robot energy depletion. We analytically evaluate its performance and show its closeness to the optimal solutions. Furthermore, through a queue-based approach, we provide theoretical guidance on the setting of the remaining energy threshold at which the robots request energy replenishment. This guided setting guarantees the feasibility of the tree-based schedule to return a depletion-free charging schedule. The performance of the tree-based charging schedule is evaluated through extensive simulations. The results show that the charger travel distance can be reduced by around 20%, when compared with the schedule that only considers the robot charging latency.
Liang He 0002, Peng Cheng 0001, Yu Gu 0001, Jianping Pan 0001, Ting Zhu 0001, Cong Liu 0005
INFOCOM6
2014 Minimizing response times of automotive dataflows on multicore
abstract
Dataflow software architectures are prevalent in prototypes of advanced automotive systems, for both driver-assisted and autonomous driving. Safety constraints of these systems necessitate real-time performance guarantees. Automotive prototypes often ensure such constraints through over-provisioning and dedicated hardware; however, a commercially viable system must utilize as few low-cost multicore processors as possible to meet size, weight, and power constraints. In short, these platforms must do more with less. To this end, we develop cache-aware and overhead-cognizant scheduling techniques that lessen guaranteed response times without unnecessarily constraining platform utilization. We implement these techniques in PGMRT, a portable middleware framework for managing real-time dataflow applications on multicore platforms. The efficacy of our techniques is demonstrated through overhead-aware schedulability experiments and runtime observations. Results for our test platform show that cache-aware clustered scheduling outperforms naïve partitioned and global approaches in terms of schedulability and end-to-end response times of dataflows.
Glenn A. Elliott, Namhoon Kim, Jeremy P. Erickson, Cong Liu 0005, James H. Anderson
RTCSA4
2014 Fixed-Relative-Deadline Scheduling of Hard Real-Time Tasks with Self-Suspensions
abstract
In many real-time systems, tasks may experience self-suspension delays when accessing external devices. The problem of scheduling such self-suspending tasks to meet hard deadlines on a uniprocessor is known to be NP-hard in the strong sense. Current solutions including the common suspension-oblivious approach of treating all suspensions as computation can be quite pessimistic. This paper shows that another category of scheduling algorithms, namely fixed-relative-deadline (FRD) scheduling, may yield better performance than classical schedulers such as EDF and RM, for real-time tasks that may experience one self-suspension during the execution of a task instance. We analyze a simple FRD algorithm, namely EDA, and derive corresponding pseudo-polynomial-time and linear-time schedulability tests. To analyze the quality of EDA and its schedulability tests, we analyze their resource augmentation factors, with respect to the speed-up factor that is needed to ensure the schedulability and feasibility of the resulting schedule. Specifically, the speed-up factor of EDA is 2 and 3, when referring to the optimal FRD scheduling and any feasible arbitrary scheduling, respectively. Moreover, the speed-up factor of the proposed linear-time schedulability test is 2.787 and 4.875, when referring to the optimal FRD scheduling and any feasible arbitrary scheduling, respectively. Furthermore, extensive experiments presented herein show that our proposed linear-time schedulability test improves upon prior approaches by a significant margin. To our best knowledge, for the scheduling of self-suspending tasks, these are the first results of any sort that indicate it might be possible to design good approximation algorithms.
Jian-Jia Chen, Cong Liu 0005
RTSS2
2014 Bursty-Interference Analysis Techniques for Analyzing Complex Real-Time Task Models
abstract
Due to the recent trend towards building complex real-time cyber-physical systems, system designers need to develop and choose expressive formal models for representing such systems, as the model should be adequately expressive such that it can accurately convey the relevant characteristics of the system being modeled. Compared to the classical sporadic task model, there exist a number of real-time task models that are more expressive. However, such models are often complex and thus are rather difficult to be analyzed efficiently. Due to this reason, prior analysis methods for dealing with such complex task models are pessimistic. In this paper, a novel analysis technique, namely the bur sty-interference analysis, is presented for analyzing two common expressive real-time task models, the general self-suspending task model and the deferrable server task model. This technique is used to derive new uniprocessor utilization-based schedulability tests and rate-monotonic utilization bounds for the two considered task models scheduled under rate-monotonic scheduling. Extensive experiments presented herein show that our proposed tests improve upon prior tests in all scenarios, in many cases by a wide margin. To the best of our knowledge, these are the first techniques that can efficiently analyze the general self-suspending and deferrable server task models on uniprocessors.
Cong Liu 0005, Jian-Jia Chen
RTSS1
2014 REPC: Reliable and efficient participatory computing for mobile devices
abstract
Smartphones and mobile devices have greatly penetrated the daily lives of many people. While participatory/pervasive sensing has gained wide adoptions by leveraging various onboard sensors on mobile devices, another powerful resource, the computational power on these mobile devices has been less frequently harnessed by researchers and practitioners. To fill this gap, we propose in this work the modeling, analysis, and implementation of participatory computing. Specifically, we propose REPC, a generic randomized task assignment framework for the participatory computing paradigm, which guarantees the overall system performance with close to minimal workload at individual participating devices. To achieve these design objectives, we model the intrinsic relationship between the workload of individual devices and the probability they complete their assigned tasks. Based on our modeling results, we analyze the maximal system capacity for any given participatory computing system and derive the minimal workload for individual participating devices to achieve the overall system performance requirement. We have fully implemented our design on the Android platform and demonstrated its performance through a representative participatory computing application. Extensive experiments and simulation results demonstrate that our design is able to achieve more than 90% task completion ratios with only 10% system overhead in practice.
Zheng Dong 0002, Linghe Kong, Peng Cheng 0001, Liang He 0002, Yu Gu 0001, Ting Zhu 0001, Cong Liu 0005
SECON8
2014 Supporting soft real-time parallel applications on multiprocessors
Cong Liu 0005, James H. Anderson
J. Syst. Archit.1
2013 Suspension-Aware Analysis for Hard Real-Time Multiprocessor Scheduling
abstract
In many real-time systems, tasks may experience suspension delays when accessing external devices. The problem of analyzing task systems with such suspensions on multiprocessors has been relatively unexplored. The commonly used suspension-oblivious approach of treating all suspensions as computation can be quite pessimistic. As an alternative, this paper presents the first suspension-aware hard real-time multiprocessor schedulability analysis for task systems with suspensions, under both global fixed-priority and global EDF scheduling. In experiments presented herein, the proposed schedulability tests proved to be superior to suspension-oblivious tests. Moreover, when applied to ordinary arbitrary-deadline sporadic task systems with no suspensions, the proposed analysis for fixed-priority scheduling improves upon prior analysis.
Cong Liu 0005, James H. Anderson
ECRTS1
2013 Exploring Adaptive Reconfiguration to Optimize Energy Efficiency in Large-Scale Battery Systems
abstract
Large-scale battery packs with hundreds/thousands of battery cells are commonly adopted in many emerging cyber-physical systems such as electric vehicles and smart micro-grids. For many applications, the load requirements on the battery systems are dynamic and could significantly change over time. How to resolve the discrepancies between the output power supplied by the battery system and the input power required by the loads is key to the development of large-scale battery systems. Traditionally, voltage regulators are often adopted to convert the voltage outputs to match loads' required input power. Unfortunately, the efficiency of utilizing such voltage regulators degrades significantly when the difference between supplied and required voltages becomes large or the load becomes light. In this paper, we propose to address this problem via an adaptive reconfiguration framework for the battery system. By abstracting the battery system into a graph representation, we develop two adaptive reconfiguration algorithms to identify the desired system configurations dynamically in accordance with real-time load requirements. We extensively evaluate our design with empirical experiments on a prototype battery system, electric vehicle driving trace-based emulation, and battery discharge trace-based simulations. The evaluation results demonstrate that, depending on the system states, our proposed adaptive reconfiguration algorithms are able to achieve 1× to 5× performance improvement with regard to the system operation time.
Liang He 0002, Lipeng Gu, Linghe Kong, Yu Gu 0001, Cong Liu 0005, Tian He 0001
RTSS5
2012 Supporting Soft Real-Time Parallel Applications on Multicore Processors
abstract
The prevalence of multicore processors has resulted in the wider applicability of parallel programming models such as Open MP and MapReduce. A common goal of running parallel applications implemented under such models is to guarantee bounded response times while maximizing system utilization. Unfortunately, little previous work has been done that can provide such performance guarantees. In this paper, this problem is addressed by applying soft real-time scheduling analysis techniques. Analysis and conditions are presented for guaranteeing bounded response times for parallel applications under global EDF multiprocessor scheduling.
Cong Liu 0005, James H. Anderson
RTCSA1
2012 An O(m) Analysis Technique for Supporting Real-Time Self-Suspending Task Systems
abstract
In many real-time and embedded systems, suspension delays may occur when tasks block to access shared resources or interact with external devices. Unfortunately, prior analysis methods for dealing with suspensions are quite pessimistic. In this paper, a novel technique is presented for analyzing soft real-time sporadic self-suspending task systems, for which bounded deadline tardiness is required, scheduled under global schedulers such as global EDF on multiprocessors (or EDF on uniprocessors). This technique is used to derive a new schedulability test that results in only O(m) suspension-related utilization loss, where m is the number of processors. The derived test theoretically dominates prior tests with respect to schedulability. Furthermore, experiments presented herein show that the improvement over prior tests is often quite significant.
Cong Liu 0005, James H. Anderson
RTSS1
2011 Supporting Graph-Based Real-Time Applications in Distributed Systems
abstract
The processing graph method (PGM) is a widely used framework for modeling applications with producer/consumer precedence constraints. PGM was originally developed by the U.S. Navy to model signal-processing applications where data communications exist among connected tasks. Prior work has shown how to schedule PGM-specified systems on uniprocessors and globally-scheduled multiprocessors. In this paper, this work is extended to enable such systems to be supported in a distributed collection of multicore machines. In such a context, pure global and partitioned scheduling approaches are problematic. Moreover, data communication costs must be considered. In this paper, a clustered scheduling algorithm is proposed for soft real-time PGM-specified distributed task systems for which bounded deadline tardiness is acceptable. This algorithm is effective in reducing data communication costs with little utilization loss. This is shown both analytically and via experiments conducted to compare it with an optimal integer linear programming solution.
Cong Liu 0005, James H. Anderson
RTCSA (1)1
2010 Scheduling Suspendable, Pipelined Tasks with Non-Preemptive Sections in Soft Real-Time Multiprocessor Systems
abstract
While most prior work on multiprocessor real-time scheduling focuses on independent tasks, dependencies due to non-preemptive sections, suspensions, and pipeline-based precedence constraints are common in practice. In this paper, such complexities are considered in the context of the global earliest-deadline-first scheduling algorithm. It is shown that any periodic task system with such dependencies can be transformed into one with only suspensions in a way that preserves maximum per-task response times. This result enables analysis directed at systems with suspensions to be applied if non-preemptive sections and/or pipelines are present as well.
Cong Liu 0005, James H. Anderson
IEEE Real-Time and Embedded Technology and Applications Symposium1
2010 Improving the Schedulability of Sporadic Self-Suspending Soft Real-Time Multiprocessor Task Systems
abstract
In work on globally-scheduled soft real-time multiprocessor systems, analysis has been presented for dealing with self-suspensions, but this analysis can be pessimistic. In this paper, we present an approach that is designed to improve the schedulability of such systems. In experimental results that are presented, the proposed approach significantly improved schedulability in most considered scenarios.
Cong Liu 0005, James H. Anderson
RTCSA1
2010 Supporting Soft Real-Time DAG-Based Systems on Multiprocessors with No Utilization Loss
abstract
In work on globally-scheduled real-time multiprocessor systems, analysis is lacking for supporting real-time applications developed using general processing graph models. In this paper, it is shown that bounded deadline tardiness can be ensured for such applications on a multiprocessor with no utilization loss. This result is general: it is applicable to periodic, sporadic, and rate-based directed-acyclic-graph (DAG) models and allows sophisticated notions of precedence to be supported (particularly, notions allowed by the processing graph method). This paper is the first to show that bounded tardiness can be ensured for globally-scheduled DAG-based applications without utilization loss.
Cong Liu 0005, James H. Anderson
RTSS1
2009 Supporting Pipelines in Soft Real-Time Multiprocessor Systems
abstract
In work on multiprocessor real-time systems, processing pipelines have received little attention. In this paper, soft real-time periodic task systems are considered that include such pipelines. Conditions are presented for guaranteeing bounded deadline tardiness in such systems under global EDF or FIFO multiprocessor scheduling.
Cong Liu 0005, James H. Anderson
ECRTS1
2009 Supporting Sporadic Pipelined Tasks with Early-Releasing in Soft Real-Time Multiprocessor Systems
abstract
Soft real-time sporadic multiprocessor task systems are considered that include processing pipelines. Conditions are presented for guaranteeing bounded deadline tardiness in such systems under global EDF or FIFO scheduling. "Early-releasing" is applied to make pipeline scheduling work-conserving. This lessens job response times in lightly-loaded systems.
Cong Liu 0005, James H. Anderson
RTCSA1
2009 Task Scheduling with Self-Suspensions in Soft Real-Time Multiprocessor Systems
abstract
In work on multiprocessor real-time systems, task scheduling with self-suspensions is a relatively unexplored topic. In this paper, soft real-time sporadic task systems are considered that include self-suspending tasks. Conditions are presented for guaranteeing bounded deadline tardiness in such systems under global EDF or FIFO multiprocessor scheduling. These conditions enable many soft real-time task systems with self-suspending tasks to be scheduled with little or no utilization loss.
Cong Liu 0005, James H. Anderson
RTSS1