VLDB 2026 Research / reviewers in the wild / expert
Hyoseung Kim 0001
dblp:19/5798-1
· DBLP profile ↗
59ranked-venue papers
16as first author
29since 2021 · last 2026
0000-0002-8553-732XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 22 · 6 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 18 · 4 first-author · 12 since 2021Computer networks · 7 · 5 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EBVC: Electronic Bee Veterinarian - Beyond Monitoring and Onto Control
Shamima Hossain, Meng-Chieh Lee, Christos Faloutsos, Boris Baer, Hyoseung Kim 0001, Vassilis J. Tsotras |
PAKDD (1) | 5 |
| 2025 | ECLIP: Energy-efficient and Practical Co-Location of ML Inference on Spatially Partitioned GPUsabstractAs AI inference becomes mainstream, research has begun to focus on improving the energy consumption of inference servers. Inference kernels commonly underutilize a GPU’s compute resources and waste power from idling components. To improve utilization and energy efficiency, multiple models can co-locate and share the GPU. However, typical GPU spatial partitioning techniques often experience significant overheads when reconfiguring spatial partitions, which can waste additional energy through repartitioning overheads or non-optimal partition configurations. In this paper, we present ECLIP, a framework to enable low-overhead energy-efficient kernel-wise resource partitioning between co-located inference kernels. ECLIP minimizes repartitioning overheads by pre-allocating pools of CU masked streams and assigns optimal CU assignments to groups of kernels through our resource allocation optimizer. Overall, ECLIP achieves an average of 13% improvement to throughput and 25% improvement to energy efficiency. Ryan Quach, Yidi Wang 0001, Ali Jahanshahi, Daniel Wong 0001, Hyoseung Kim 0001 |
ISLPED | 5 |
| 2025 | Poster Abstract: Low-Cost Soil Sensing and Two-Level Classification for Early Stress Detection in Avocado PlantsabstractWe present a systematic evaluation of low-cost soil sensors for early stress and disease detection in avocado plants. Our monitoring system was deployed across 72 plants divided into four treatment categories within a controlled greenhouse environment collecting data over six months. We developed a two-level hierarchical classifier leveraging soil electrical conductivity (EC) and moisture data to improve classification accuracy. The proposed classifier achieved 75--86% accuracy across different avocado genotypes, outperforming conventional machine learning approaches by over 20%. Our findings demonstrate that while low-cost sensors exhibit certain limitations in field conditions, strategic classification techniques can significantly enhance their utility for precision agriculture. Abdulrahman Bukhari, Bullo Mamo, Shamima Hossain, Daniel Enright, Patricia Manosalva, Hyoseung Kim 0001 |
SenSys | 8 |
| 2025 | Low-Cost Sensing and Classification for Early Stress and Disease Detection in Avocado PlantsabstractWith rising demands for efficient disease and salinity management in agriculture, early detection of plant stressors is crucial, particularly for high-value crops like avocados. This paper presents a comprehensive evaluation of low-cost sensors deployed in the field for early stress and disease detection in avocado plants. Our monitoring system was deployed across 72 plants divided into four treatment categories within a greenhouse environment, with data collected over six months. While leaf temperature and conductivity measurements, widely used metrics for controlled settings, were found unreliable in field conditions due to environmental interference and positioning challenges, leaf spectral measurements produced statistically significant results when combined with our machine learning approach. For soil data analysis, we developed a two-level hierarchical classifier that leverages domain knowledge about treatment characteristics, achieving 75-86% accuracy across different avocado genotypes and outperforming conventional machine learning approaches by over 20%. In addition, performance evaluation on an embedded edge device demonstrated the viability of our approach for resource-constrained environments, with reasonable computational efficiency while maintaining high classification accuracy. Our work bridges the gap between theoretical potential and practical application of low-cost sensors in agriculture and offers insights for developing affordable, scalable monitoring systems. Abdulrahman Bukhari, Bullo Mamo, Shamima Hossain, Daniel Enright, Patricia Manosalva, Hyoseung Kim 0001 |
IEEE Internet Things J. | 8 |
| 2025 | Mixed-trust Computing: Safe and Secure Real-time SystemsabstractVerifying complex Cyber-physical Systems (CPSs) is increasingly important given the push to deploy safety-critical autonomous features. Unfortunately, traditional verification methods do not scale to the complexity of these systems and do not provide systematic methods to protect verified properties when not all the components can be verified. To address these challenges, this article proposes a real-time mixed-trust computing framework that combines verification and protection. The framework introduces a new task model, where an application task can have both an untrusted and a trusted part. The untrusted part allows complex computations supported by a full OS with a real-time scheduler running in a VM hosted by a trusted hypervisor. The trusted part is executed by another scheduler within the hypervisor and is thus protected from the untrusted part. If the untrusted part fails to finish by a specific time, the trusted part is activated to preserve safety (e.g., prevent a crash) including its timing guarantees. This framework is the first allowing the use of untrusted components for CPS critical functions while preserving logical and timing guarantees, even in the presence of malicious attackers. We present the framework, its schedulability analysis, and the coordination protocol between the trusted and untrusted parts. Our implementation on a Raspberry Pi 3 is also discussed along with experiments showing the behavior of the system under failures of untrusted components and a drone application to demonstrate its practicality. Dionisio de Niz, Björn Andersson, Mark Klein 0003, John P. Lehoczky, Hyoseung Kim 0001, Gabriel A. Moreno |
ACM Trans. Cyber Phys. Syst. | 5 |
| 2025 | Principled Mining, Forecasting, and Monitoring of Honeybee Time Series with EBV+abstractHoneybees, as natural crop pollinators, play a significant role in biodiversity and food production for human civilization. Bees actively regulate hive temperature (homeostasis) to maintain a colony’s proper functionality. Deviations from usual thermoregulation behavior due to external stressors (e.g., extreme environmental temperature, parasites, pesticide exposure) indicate an impending colony collapse. Anticipating such threats by forecasting hive temperature and finding changes in temperature patterns would allow beekeepers to take early preventive measures and avoid critical issues. In that case, how can we model bees’ thermoregulation behavior for an interpretable and effective hive monitoring system? In this article, we propose the principled Electronic Bee-Veterinarian Plus (EBV+) method based on the thermal diffusion equation and a novel “ sigmoid ” feedback-loop (P) controller for analyzing hive health with the following properties: (i) it is effective on multiple, real-world beehive time sequences (recorded and streaming), (ii) it is explainable with only a few parameters (e.g., hive health factor) that beekeepers can easily quantify and trust, (iii) it issues proactive alerts to beekeepers before any potential issue affecting homeostasis becomes detrimental, and (iv) it is scalable with a time complexity of \(O(t)\) for reconstructing and \(O(t\times m)\) for finding m cuts of a sequence with t time-ticks. Experimental results on multiple real-world time sequences showcase the potential and practical feasibility of EBV+. Our method yields accurate forecasting (up to 72% improvement in RMSE) with up to 600 times fewer parameters compared to baselines (ARX, seasonal ARX, Holt-winters, and DeepAR), as well as detects discontinuities and raises alerts that coincide with domain experts’ opinions. Moreover, EBV+ is scalable and fast, taking less than 1 minute on a stock laptop to reconstruct 2 months of sensor data. Shamima Hossain, Christos Faloutsos, Boris Baer, Hyoseung Kim 0001, Vassilis J. Tsotras |
ACM Trans. Knowl. Discov. Data | 4 |
| 2024 | GCAPS: GPU Context-Aware Preemptive Priority-Based Scheduling for Real-Time TasksabstractScheduling real-time tasks that utilize GPUs with analyzable guarantees poses a significant challenge due to the intricate interaction between CPU and GPU resources, as well as the complex GPU hardware and software stack. While much research has been conducted in the real-time research community, several limitations persist, including the absence or limited availability of GPU-level preemption, extended blocking times, and/or the need for extensive modifications to program code. In this paper, we propose GCAPS, a GPU Context-Aware Preemptive Scheduling approach for real-time GPU tasks. Our approach exerts control over GPU context scheduling at the device driver level and enables preemption of GPU execution based on task priorities by simply adding one-line macros to GPU segment boundaries. In addition, we provide a comprehensive response time analysis of GPU-using tasks for both our proposed approach as well as the default Nvidia GPU driver scheduling that follows a work-conserving round-robin policy. Through empirical evaluations and case studies, we demonstrate the effectiveness of the proposed approaches in improving taskset schedulability and response time. The results highlight significant improvements over prior work as well as the default scheduling approach, with up to 40% higher schedulability, while also achieving predictable worst-case behavior on Nvidia Jetson embedded platforms. Yidi Wang 0001, Cong Liu 0005, Daniel Wong 0001, Hyoseung Kim 0001 |
ECRTS | 4 |
| 2024 | Rapid Hardware/Software Design Space Exploration for Efficient Intermittent SystemsabstractIntermittent computing enables the functioning of computing systems under unstable power conditions. Designing such systems poses significant challenges due to the vast design space, including hardware and software parameters along with harvested energy availability. In this paper, we propose an analytical framework to predict an application's execution time in intermittently powered systems, enabling rapid design space exploration. Our framework accounts for previously unexplored factors such as the effect of Equivalent Series Resistance (ESR) in capacitors and the choice of checkpoint strategies. With only one-time profiling, it estimates timings of different checkpoint techniques under various design configurations, with an average error of 4.7% under stable power and 10.4% under intermittent power. Furthermore, our evaluation reveals that neglecting checkpoint strategy in design can result in a 3.08x average slowdown compared to optimal setups. Hyoseung Kim 0001 |
ISLPED | 2 |
| 2024 | PAAM: A Framework for Coordinated and Priority-Driven Accelerator Management in ROS 2abstractThis paper proposes a Priority-driven Accelerator Access Management (PAAM) framework for multi-process robotic applications built on top of the Robot Operating System (ROS) 2 middleware platform. The framework addresses the issue of predictable execution of time- and safety-critical callback chains that require hardware accelerators such as GPUs and TPUs. PAAM provides a standalone ROS executor that acts as an accelerator resource server, arbitrating accelerator access requests from all other callbacks at the application layer. This approach enables coordinated and priority-driven accelerator access management in multi-process robotic systems. The framework design is directly applicable to all types of accelerators and enables granular control over how specific chains access accelerators, making it possible to achieve predictable real-time support for accelerators used by safety-critical callback chains without making changes to underlying accelerator device drivers. The paper shows that PAAM also offers a theoretical analysis that can upper bound the worst-case response time of safety-critical callback chains that necessitate accelerator access. This paper also demonstrates that complex robotic systems with extensive accelerator usage that are integrated with PAAM may achieve up to a 91% reduction in end-to-end response time of their critical callback chains. Daniel Enright, Yecheng Xiang, Hyunjong Choi, Hyoseung Kim 0001 |
RTAS | 4 |
| 2024 | BOXR: Body and head motion Optimization framework for eXtended RealityabstractThe emergence of standalone Extended Reality (XR) systems has enhanced user mobility, accommodating both subtle, frequent head motions and substantial, less frequent body motions. However, the pervasively used Motion-to-Display (M2D) latency metric, which measures the delay between the most recent motion and its corresponding display update, only accounts for head motions. This oversight can leave users prone to motion sickness if significant body motion is involved. Although existing methods optimize M2D latency through asynchronous task scheduling and reprojection methods, they introduce challenges like resource contention between tasks and outdated pose data. These challenges are further complicated by user motion dynamics and scene changes during runtime. To address these issues, we for the first time introduce the Camera-to-Display (C2D) latency metric, which captures the delay caused by body motions, and present BOXR, a framework designed to co-optimize both body and head motion delays within an XR system. BOXR enhances the coordination between M2D and C2D latencies by efficiently scheduling tasks to avoid contentions while maintaining an up-to-date pose in the output frame. Moreover, BOXR incorporates a motion-driven visual inertial odometer to adjust to user motion dynamics and employs scene-dependent foveated rendering to manage changes in the scene effectively. Our evaluations show that BOXR significantly outperforms state-of-the-art solutions in 11 EuRoC MAV datasets across 4 XR applications across 3 hardware platforms. In controlled motion and scene settings, BOXR reduces M2D and C2D latencies by up to $63 \%$ and $27 \%$, respectively and increases frame rate by up to $43 \%$. In practical deployments, BOXR achieves substantial reductions in real-world scenarios-up to $42 \%$ in M2D latency and $31 \%$ in C2D latency-while maintaining remarkably low miss rates of only $1.6 \%$ for M2D requirements and $\mathbf{1. 0 \%}$ for C2D requirements. Zexin Li 0001, Hyoseung Kim 0001, Cong Liu 0005 |
RTSS | 3 |
| 2024 | EBV: Electronic Bee-Veterinarian for Principled Mining and Forecasting of Honeybee Time SeriesabstractHoneybees are vital for pollination and food production. Among many factors, extreme temperature (e.g., due to climate change) is particularly dangerous for bee health. Anticipating such extremities would allow beekeepers to take early preventive action. Thus, given sensor (temperature) time series data from beehives, how can we find patterns and do forecasting? Forecasting is crucial as it helps spot unexpected behavior and thus issue warnings to the beekeepers. In that case, what are the right models for forecasting? ARIMA, RNNs, or something else? Shamima Hossain, Christos Faloutsos, Boris Baer, Hyoseung Kim 0001, Vassilis J. Tsotras |
SDM | 4 |
| 2024 | OpenSense: An Open-World Sensing Framework for Incremental Learning and Dynamic Sensor Scheduling on Embedded Edge DevicesabstractRecent advances in Internet-of-Things (IoT) technologies have sparked significant interest towards developing learning-based sensing applications on embedded edge devices. These efforts, however, are being challenged by the complexities of adapting to unforeseen conditions in an open-world environment, mainly due to the intensive computational and energy demands exceeding the capabilities of edge devices. In this paper, we propose OpenSense, an open-world time-series sensing framework for making inferences from time-series sensor data and achieving incremental learning on an embedded edge device with limited resources. The proposed framework is able to achieve two essential tasks, inference and incremental learning, eliminating the necessity for powerful cloud servers. In addition, to secure enough time for incremental learning and reduce energy consumption, we need to schedule sensing activities without missing any events in the environment. Therefore, we propose two dynamic sensor scheduling techniques: (i) a class-level period assignment scheduler that finds an appropriate sensing period for each inferred class, and (ii) a Q-learning-based scheduler that dynamically determines the sensing interval for each classification moment by learning the patterns of event classes. With this framework, we discuss the design choices made to ensure satisfactory learning performance and efficient resource usage. Experimental results demonstrate the ability of the system to incrementally adapt to unforeseen conditions and to efficiently schedule to run on a resource-constrained device. Abdulrahman Bukhari, Seyedmehdi Hosseinimotlagh, Hyoseung Kim 0001 |
IEEE Internet Things J. | 3 |
| 2024 | MII: A Multifaceted Framework for Intermittence-Aware Inference and SchedulingabstractThe concurrent execution of deep neural networks (DNNs) inference tasks on the intermittently-powered batteryless devices (IPDs) has recently garnered much attention due to its potential in a broad range of smart sensing applications. While the checkpointing mechanisms (CMs) provided by the state-of-the-art make this possible, scheduling inference tasks on IPDs is still a complex problem due to significant performance variations across the DNN layers and CM choices. This complexity is further accentuated by dynamic environmental conditions and inherent resource constraints of IPDs. To tackle these challenges, we present MII, a framework designed for the intermittence-aware inference and scheduling on IPDs. MII formulates the shutdown and live time functions of an IPD from profiling the data, which our offline intermittence-aware search scheme uses to find the optimal layer-wise CMs for each task. At runtime, MII enhances the job success rates by dynamically making scheduling decisions to mitigate the workload losses from the power interruptions and adjusting these CMs in response to the actual energy patterns. Our evaluation demonstrates the superiority of MII over the state-of-the-art. In controlled environments, MII achieves an average increase of 21% and 39% in successful jobs under the stable and dynamic energy patterns. In the real-world settings, MII achieves 33% and 24% more successful jobs indoors and outdoors. Cong Liu 0005, Hyoseung Kim 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2023 | Timing Analysis and Priority-driven Enhancements of ROS 2 Multi-threaded ExecutorsabstractThe second generation of Robotic Operating System, ROS 2, has gained much attention for its potential to be used for safety-critical robotic applications. The need to provide a solid foundation for timing correctness and scheduling mechanisms is therefore growing rapidly. Although there are some pioneering studies conducted on formally analyzing the response time of processing chains in ROS 2, the focus has been limited to singlethreaded executors, and multi-threaded executors, despite their advantages, have not been studied well. To fill this knowledge gap, in this paper, we propose a comprehensive response-time analysis framework for chains running on ROS 2 multi-threaded executors. We first analyze the timing behavior of the default scheduling scheme in ROS 2 multi-threaded executors, and then present priority-driven scheduling enhancements to address the limitations of the default scheme. Our framework can analyze chains with both arbitrary and constrained deadlines and also the effect of mutually-exclusive callback groups. Evaluation is conducted by a case study on NVIDIA Jetson AGX Xavier and schedulability experiments using randomly-generated chains. The results demonstrate that our analysis framework can safely upper-bound response times under various conditions and the priority-driven scheduling enhancements not only reduce the response time of critical chains but also improve analytical bounds. Hoora Sobhani, Hyunjong Choi, Hyoseung Kim 0001 |
RTAS | 3 |
| 2023 | $\mathrm{R}^{3}$: On-Device Real-Time Deep Reinforcement Learning for Autonomous RoboticsabstractAutonomous robotic systems, like autonomous vehicles and robotic search and rescue, require efficient on-device training for continuous adaptation of Deep Reinforcement Learning (DRL) models in dynamic environments. This research is fundamentally motivated by the need to understand and address the challenges of on-device real-time DRL, which involves balancing timing and algorithm performance under memory constraints, as exposed through our extensive empirical studies. This intricate balance requires co-optimizing two pivotal parameters of DRL training - batch size and replay buffer size. Configuring these parameters significantly affects timing and algorithm performance, while both (unfortunately) require substantial memory allocation to achieve near-optimal performance. This paper presents$\mathbf{R}^{3}$, a holistic solution for managing timing, memory, and algorithm performance in on-device real-time DRL training.$\mathbf{R}^{3}$employs (i) a deadline-driven feedback loop with dynamic batch sizing for optimizing timing, (ii) efficient memory management to reduce memory footprint and allow larger replay buffer sizes, and (iii) a runtime coordinator guided by heuristic analysis and a runtime profiler for dynamically adjusting memory resource reservations. These components collaboratively tackle the trade-offs in on-device DRL training, improving timing and algorithm performance while minimizing the risk of out-of-memory (OOM) errors. We implemented and evaluated$\mathbf{R}^{3}$extensively across various DRL frameworks and benchmarks on three hardware platforms commonly adopted by autonomous robotic systems. Additionally, we integrate$\mathbf{R}^{3}$with a popular realistic autonomous car simulator to demonstrate its real-world applicability. Evaluation results show that$\mathbf{R}^{3}$achieves efficacy across diverse platforms, ensuring consistent latency performance and timing predictability with minimal overhead. Moreover,$\mathbf{R}^{3}$showcases versatility by handling varied optimization goals and adapting to fluctuating systems scenarios. Zexin Li 0001, Aritra Samanta, Yufei Li 0001, Andrea Soltoggio, Hyoseung Kim 0001, Cong Liu 0005 |
RTSS | 5 |
| 2023 | Message from the Program, Track, and General ChairsabstractOn behalf of the IEEE Technical Committee on Real-Time Systems (TCRTS), it is our pleasure to welcome you to the 44th IEEE Real-Time Systems Symposium (RTSS 2023) during December 5 - 8, 2023 in Taipei. Over the past 44 years, RTSS has established itself as the primary forum for research in the broad field of real-time and embedded systems. Insik Shin, Nan Guan, Renato Mancuso 0001, Hyoseung Kim 0001, Jian-Jia Chen |
RTSS | 4 |
| 2022 | Opportunistic Communication with Latency Guarantees for Intermittently-Powered DevicesabstractEnergy-harvesting wireless sensor nodes have found widespread adoption due to their low cost and small form factor. However, uncertainty in the available power supply introduces significant challenges in engineering communications between intermittently-powered nodes. We propose a constraint-based model for energy harvests that together with a hardware model can be used to enable consistent, opportunistic communication with worst-case latency guarantees. We show that greedy approaches that attempt communication whenever energy is available lead to prolonged latencies in real-world environments. Our approach offers bounded worst-case latency while providing a performance improvement over a conservative, offline approach planned around the worst-case energy harvest. Kacper Wardega, Wenchao Li 0001, Hyoseung Kim 0001, Yawen Wu, Zhenge Jia, Jingtong Hu |
DATE | 3 |
| 2022 | An Open-World Time-Series Sensing Framework for Embedded Edge DevicesabstractThe rapid advancement of IoT technologies has generated much interest in the development of learning-based sensing applications on embedded edge devices. However, these efforts are being challenged by the need to adapt to unforeseen conditions in an open-world environment. Updating a learning model suffers from the lack of training data as well as the high computational demand beyond that available on edge devices. In this paper, we propose an open-world time-series sensing framework for making inferences from time-series sensor data and achieving incremental learning on an embedded edge device with limited resources. The proposed framework is able to achieve two essential tasks, inference and learning, without requiring access to a powerful cloud server. We discuss the design choices made to ensure satisfactory learning performance and efficient resource usage. Experimental results demonstrate the ability of the system to incrementally adapt to unforeseen conditions and to effectively run on a resource-constrained device. Abdulrahman Bukhari, Seyedmehdi Hosseinimotlagh, Hyoseung Kim 0001 |
RTCSA | 3 |
| 2022 | Energy-Adaptive Real-time Sensing for Batteryless DevicesabstractThe use of batteryless energy harvesting devices has been recognized as a promising solution for their low maintenance requirements and ability to work in harsh environments. However, these devices have to harvest energy from ambient energy sources and execute real-time sensing tasks periodically while satisfying data freshness constraints, which is especially challenging as the energy sources are often unreliable and intermittent. In this paper, we develop an energy-adaptive real-time sensing framework for batteryless devices. This framework includes a lightweight machine learning-based energy predictor that is capable of running on microcontroller devices and predicting the energy availability and intensity based on energy traces. Using this, the framework adapts the schedule of real-time tasks by effectively taking into account the predicted energy supply and the resulting age of information of each task, in order to achieve continuous sensing operations and satisfy given data freshness requirements. We discuss various design choices for adaptive scheduling and evaluate their performance in the context of batteryless devices. Experimental results show that the proposed adaptive real-time approach outperforms the recent methods based on static and reactive approaches, in both energy utilization and data freshness. Yidi Wang 0001, Hyoseung Kim 0001 |
RTCSA | 3 |
| 2022 | Towards Energy-Efficient Real-Time Scheduling of Heterogeneous Multi-GPU SystemsabstractWith the increasing demand for computational power, research on general-purpose graphics processing units (GPUs) has been active for various real-time systems spanning from autonomous vehicles to real-time clouds. While the use of GPUs can significantly benefit compute-intensive tasks with timing constraints, their high power consumption becomes an important problem given that it is not rare to see multiple GPUs in today's systems. In this paper, we present our study towards energy-efficient real-time scheduling in heterogeneous multi-GPU systems. We first make observations using a custom power monitoring setup that, in a multi-GPU system, conventional task allocation approaches for multiprocessors do not lead to energy efficiency and there is no clear winner. Then we propose a multi-GPU real-time scheduling framework, sBEET-mg, that builds upon prior work on single-GPU systems and makes offline and runtime scheduling decisions to execute a given job on the energy-optimal GPU while exploiting spatial multitasking on each GPU for better concurrency and real-time performance. We implemented the proposed framework on a real multi-GPU system and evaluated it with randomly-generated task sets of benchmark programs. We also experimentally simulated our method in a system containing more GPUs. Experimental results show that sBEET-mg reduces deadline misses by up to 23% and 18% compared to the conventional load distribution and load concentration methods, respectively, while simultaneously achieving lower energy consumption than them. Yidi Wang 0001, Hyoseung Kim 0001 |
RTSS | 3 |
| 2021 | PiCAS: New Design of Priority-Driven Chain-Aware Scheduling for ROS2abstractIn ROS (Robot Operating System), most applications in time- and safety-critical domain are constructed in the form of callback chains with data dependencies. Due to the shortcomings in its real-time support, ROS does not provide a strong timing guarantee and may lead to disastrous results. Although ROS2 claims to enhance the real-time capability, ensuring predictable end-to-end chain latency still remains a challenging problem. In this paper, we propose a new priority-driven chain-aware scheduler for the ROS2 framework and present end-to-end latency analysis for the proposed scheduler. With our scheduler, callbacks are prioritized based on the given timing requirements of the corresponding chains so that the end-to-end latency of critical chains can be improved with a predictable bound. The proposed scheduling design includes priority assignment and resource allocation considering all ROS2 scheduling-related abstractions, e.g., callbacks, nodes, and executors. To the best of our knowledge, this is the first work to address the inherent limitations of ROS2 in end-to-end latency by proposing a new scheduler design. We have implemented our scheduler in ROS2 running on NVIDIA Xavier NX. We have conducted case studies and schedulability experiments. The results show that the proposed scheduler yields a substantial improvement in end-to-end latency over the default ROS2 scheduler and the latest work in real-world scenarios. Hyunjong Choi, Yecheng Xiang, Hyoseung Kim 0001 |
RTAS | 3 |
| 2021 | Data-Driven Structured Thermal Modeling for COTS Multi-core ProcessorsabstractThermal awareness is increasingly important for real-time systems deployed in harsh environments. As high chip temperature can cause frequency throttling or shutdown of processor cores at unexpected times, many real-time scheduling techniques have been developed to ensure continuous, fail-safe operation of safety-critical tasks with stringent timing constraints. However, their practical use remains largely limited due to the fact that it is extremely difficult to obtain a precise thermal model of commercial processors without using special measurement instruments or access to proprietary information, such as the power traces of micro-architectural units and detailed floorplans.In this paper, we propose a data-driven structured thermal modeling scheme that is directly applicable to commercial off-the-shelf multi-core processors used in real-time embedded systems. By using a small number of thermal profiles obtained from on-chip temperature sensors, our scheme can accurately predict the processor operating temperature under dynamic real-time workloads at various CPU frequencies and ambient conditions. The thermal model derived from our scheme is fast to converge and robust against different sources of errors. Our scheme is non-intrusive, meaning that it does not require changes to the software code or the hardware packaging of the target system. Furthermore, our scheme can estimate the relative power consumption of the processor for a given workload and clock frequency level. Experimental results from a multi-core ARM platform indicate that our scheme estimates the operating temperature with a maximum error of 2.5% while the latest prior work results in 23% error. This highly accurate modeling enables us to obtain the maximum achievable processor utilization that does not cause a thermal safety violation. Seyedmehdi Hosseinimotlagh, Daniel Enright, Christian R. Shelton, Hyoseung Kim 0001 |
RTSS | 4 |
| 2021 | Addressing Multi-core Timing Interference using Co-Runner LockingabstractThis paper presents a task synchronization mechanism, called co-runner locking, to address the timing interference problem in multi-core real-time systems. It prevents certain subsets of tasks from executing simultaneously on different cores in order to avoid large performance penalties from inter-core interference. We provide the general properties of the co-runner locking mechanism and discuss the runtime control policies that determine the execution order of tasks in a co-runner-locking relationship. For schedulability analysis, we derive a response-time test that upper-bounds the delay from co-runner locking and the slowdown imposed by permitted co-runners by combining two new analytic approaches: job-oriented and load-oriented. In evaluation, we demonstrate that the co-runner locking mechanism is an effective alternative to address the "one-out-of-m" problem and brings about a significant improvement in real-time taskset schedulability. Hyoseung Kim 0001, Dionisio de Niz, Björn Andersson, Mark Klein 0003, John P. Lehoczky |
RTSS | 1 |
| 2021 | Resilient Mixed-Trust SchedulingabstractIn this paper we present a new scheduling model for resilient real-time mixed trust systems. This model extends the previous Real-Time Mixed-Trust Computing framework RT-MTC to support degradation modes. Management of these modes has been identified in industrial documents as a key requirement for deploying trusted autonomous vehicles for safe autonomy. RT-MTC uses verified components (known as enforcers) to guarantee that the output of a system is safe by replacing it with a verified safe one if this output is deemed unsafe or is not produced on time. In this paper we extend RT-MTC and develop a scheduling model that uses the digraph scheduling model as a baseline but extends it in four critical ways: (1) it creates extensions for the mixed-preemptive scheduling required by RT-MTC, (2) it enables priority bands in order to separate trusted and untrusted components, (3) it uses these bands in order to calculate intermediate deadlines used by the RT-MTC framework for the scheduling of the trusted components, and (4) it defines system mode semantics to obtain two desirable properties of the new schedulability analysis: low pessimism and low time-complexity. This paper evaluates the new schedulability algorithm and shows that it is efficient in that it only needs to analyze one transition at a time. The new model supports the construction of a resilient autonomous system with provable guarantees protected by verified enforcers within the RT-MTC framework and, more importantly, preserves these guarantees even across failure-triggered mode changes. Dionisio de Niz, Björn Andersson, Hyoseung Kim 0001, Mark Klein 0003, John P. Lehoczky |
RTSS | 3 |
| 2021 | Balancing Energy Efficiency and Real-Time Performance in GPU SchedulingabstractGeneral-purpose graphics processing units (GPUs) made available on embedded platforms have gained much interest in real-time cyber-physical systems. Despite the fact that GPUs generally outperform CPUs on many compute-intensive tasks in a multitasking environment, high power consumption remains a challenging problem. In this paper, we first analyze the power and energy consumption of GPU kernels scheduled with spatial multitasking, which is found to be advantageous for schedulability in recent studies, and prove that its use, however, degrades energy efficiency even in the latest commercially available embedded GPUs like NVIDIA Jetson Xavier AGX. Then, based on our observations, we propose sBEET, a real-time energy-efficient GPU scheduling framework that makes scheduling decisions at runtime to optimize the energy consumption while utilizing spatial multitasking to improve real-time performance. We evaluate the performance of the proposed sBEET framework using well-known GPU benchmarks and randomly-generated timing parameters on real hardware. The results indicate that sBEET reduces deadline misses up to 13% when the system is overloaded, and also achieves 15% to 21% lower energy consumption when the tasksets are schedulable compared to the existing works. Yidi Wang 0001, Yecheng Xiang, Hyoseung Kim 0001 |
RTSS | 4 |
| 2021 | AegisDNN: Dependable and Timely Execution of DNN Tasks with SGXabstractWith the rising demand for emerging DNN applications in safety-critical systems, much attention has been given to the reliability and trustworthiness of DNN inference output against malicious attacks. Although prior work has been conducted to improve the privacy of DNN inference by executing the entire DNN model inside Intel SGX enclaves, existing approaches pose severe performance challenges to achieve dependable and timely execution simultaneously. In this paper, we propose AegisDNN, a DNN inference framework to address this problem. AegisDNN leverages secure SGX enclaves for protecting only the critical part of real-time DNN tasks which are vulnerable to potential fault injection attacks. To choose the right set of layers for protection while ensuring the timeliness of task execution, AegisDNN includes a dynamic-programming based algorithm that finds a layer protection configuration for each task to meet the real-time and dependability requirements based on the layer-wise DNN time and SDC (Silent Data Corruption) profiling mechanism. AegisDNN also utilizes a machine-learning based SDC prediction method to significantly reduce the time for estimating SDC rates for all possible layer protection configurations. We implemented AegisDNN on Caffe, PyTorch, and Tensorflow with Eigen BLAS ported into SGX enclaves to comprehensively demonstrate the effectiveness of AegisDNN against state-of-the-art DNN fault-injection attacks. Experiment results indicate that AegisDNN could satisfy both dependability and real-time requirements simultaneously, when none of the other compared approaches could do so. Yecheng Xiang, Yidi Wang 0001, Hyunjong Choi, Hyoseung Kim 0001 |
RTSS | 5 |
| 2021 | Toward Practical Weakly Hard Real-Time Systems: A Job-Class-Level Scheduling ApproachabstractRecent applications of the Internet of Things and cyber-physical systems require the integration of many sensing and control tasks into resource-constrained embedded devices. Such tasks can often tolerate a bounded number of timing violations. The concept of weakly hard real-time systems can effectively improve resource efficiency without sacrificing system safety. However, the existing studies have limitations on their practical use due to the restrictions imposed on the task timing behavior, high analysis complexity, and the lack of multicore support. In this article, we propose a new job-class-level fixed-priority preemptive scheduler and its schedulability analysis framework for weakly hard real-time tasks. Our proposed scheduler employs the meet-oriented classification of jobs of a task in order to reduce the worst-case temporal interference imposed on other tasks. Under this approach, each job is associated with a “job-class” that is determined by the number of deadlines previously met (with a bounded number of consecutively missed deadlines). This approach allows decomposing the complex weakly hard schedulability problem into two subproblems that are easier to solve: 1) analyzing the response time of a job with each job-class, which can be done by an extension of the existing task-level analysis and 2) finding possible job-class patterns, which can be modeled as a simple reachability tree. We also present a semipartitioned task allocation method for multicore platforms, which enhances the schedulability of weakly hard tasks under the proposed scheduling framework. Experimental results indicate that our scheduler outperforms the prior work in terms of task schedulability and analysis time complexity. We have also implemented a prototype of a job-class-level scheduler in the Linux kernel running on Raspberry Pi with acceptably small-runtime overhead. Hyunjong Choi, Hyoseung Kim 0001, Qi Zhu 0002 |
IEEE Internet Things J. | 2 |
| 2021 | Real-Time Task Scheduling on Intermittently Powered Batteryless DevicesabstractIntermittently powered devices (IPDs) have gained much interest in recent years. However, scheduling real-time tasks while supporting data consistency, timekeeping, and schedulability guarantees on these devices still remains a challenge. Many sensing tasks need long indivisible sensor reading operations, but most prior work has limited their focus to the forward progress of computation-only tasks. In this article, we propose a scheduling framework to execute real-time periodic tasks with atomic sensing operations. Our proposed method keeps track of time progress and ensures the periodic execution of sensing tasks while efficiently utilizing intermittent power sources. We provide schedulability analysis to determine if a taskset is schedulable under a given charging condition. As a proof-of-concept, we design a custom programmable RFID tag device, called R'tag, and demonstrate the effectiveness of our framework in a realistic sensing application. Evaluation results show that the proposed method satisfies the real-time task execution requirements on IPDs in terms of task scheduling, timekeeping, and periodic sensing while significantly outperforming prior work. Hyunjong Choi, Yidi Wang 0001, Yecheng Xiang, Hyoseung Kim 0001 |
IEEE Internet Things J. | 5 |
| 2021 | Cross-Layer Adaptation with Safety-Assured Proactive Task Job SkippingabstractDuring the operation of many real-time safety-critical systems, there are often strong needs for adapting to a dynamic environment or evolving mission objectives, e.g., increasing sampling and control frequencies of some functions to improve their performance under certain situations. However, a system's ability to adapt is often limited by tight resource constraints and rigid periodic execution requirements. In this work, we present a cross-layer approach to improve system adaptability by allowing proactive skipping of task executions, so that the resources can be either saved directly or re-allocated to other tasks for their performance improvement. Our approach includes three novel elements: (1) formal methods for deriving the feasible skipping choices of control tasks with safety guarantees at the functional layer, (2) a schedulability analysis method for assessing system feasibility at the architectural layer under allowed task job skippings, and (3) a runtime adaptation algorithm that efficiently explores job skipping choices and task priorities for meeting system adaptation requirements while ensuring system safety and timing correctness. Experiments demonstrate the effectiveness of our approach in meeting system adaptation needs. Zhilu Wang, Chao Huang 0015, Hyoseung Kim 0001, Wenchao Li 0001, Qi Zhu 0002 |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2020 | Know the Unknowns: Addressing Disturbances and Uncertainties in Autonomous Systems : Invited PaperabstractFuture autonomous systems will employ complex sensing, computation, and communication components for their perception, planning, control, and coordination, and could operate in highly dynamic and uncertain environment with safety and security assurance. To realize this vision, we have to better understand and address the challenges from the "unknowns" - the unexpected disturbances from component faults, environmental interference, and malicious attacks, as well as the inherent uncertainties in system inputs, model inaccuracies, and machine learning techniques (particularly those based on neural networks). In this work, we will discuss these challenges, propose our approaches in addressing them, and present some of the initial results. In particular, we will introduce a cross-layer framework for modeling and mitigating execution uncertainties (e.g., timing violations, soft errors) with weakly-hard paradigm, quantitative and formal methods for ensuring safe and time-predictable application of neural networks in both perception and decision making, and safety-assured adaptation strategies in dynamic environment. Qi Zhu 0002, Wenchao Li 0001, Hyoseung Kim 0001, Yecheng Xiang, Kacper Wardega, Zhilu Wang, Yixuan Wang 0001, Hengyi Liang, Chao Huang 0015, Jiameng Fan, Hyunjong Choi |
ICCAD | 3 |
| 2020 | Chain-Based Fixed-Priority Scheduling of Loosely-Dependent TasksabstractMany cyber-physical applications consist of chains of tasks. Such tasks are often loosely dependent, meaning task execution is time-triggered and independent of the update rate of input data. Since meaningful output can be obtained after processing all the intermediate tasks of a chain, the end-to-end latency of the chain is an important metric that can affect the correctness and quality of the system. In this paper, we present a chain-based fixed-priority preemptive scheduler for multicore real-time systems. The scheduler identifies effective chain instances contributing to the generation of updated output, and employs a runtime policy to improve the end-to-end latency of chains. Based on our scheduler, an analysis method is proposed with two parts: (i) bounding the start and finish time of each job, and (ii) analyzing the end-to-end latency of effective chain instances. Experimental results show that our chain-based scheduler achieves up to 83% reduction in end-to-end latency compared to the state-of-the-art and yields a significant benefit in inter-chain distance over chain-unaware schedulers. Furthermore, our analysis method can be easily adapted to chain-unaware schedulers and provides tighter bounds than prior work. Hyunjong Choi, Hyoseung Kim 0001 |
ICCD | 3 |
| 2020 | On Dynamic Thermal Conditions in Mixed-Criticality SystemsabstractThe rising demand for powerful embedded systems to support modern complex real-time applications signifies the on-chip temperature challenges. Heat conduction between CPU cores interferes in the execution time of tasks running on other cores. The violation of thermal constraints causes timing unpredictability to real-time tasks due to transient performance degradation or permanent system failure. Moreover, dynamic ambient temperature affects the operating temperature on multicore systems significantly.In this paper, we propose a thermal-aware server framework to safely upper-bound the maximum operating temperature of multi-core mixed-criticality systems. With the proposed analysis on the impact of ambient temperature, our framework manages mixed-criticality tasks to satisfy both thermal and timing requirements. We present techniques to find the maximum ambient temperature for each criticality level to guarantee the safe operating temperature bound. We also analyze the minimum time required for a criticality mode change from one level to another. The thermal properties of our framework have been evaluated on a commercial embedded platform. A case study with real-world mixed-critical applications demonstrates the effectiveness of our framework in bounding operating temperature under dynamic ambient temperature changes. Seyedmehdi Hosseinimotlagh, Ali Ghahremannezhad, Hyoseung Kim 0001 |
RTAS | 3 |
| 2020 | Work-In-Progress: Toward Precomputation in Real-Time Mixed-Trust SchedulingabstractThe Real-Time Mixed-Trust (RTMT) Framework [2] enables the use of untrusted components in safety-critical CPS functions (e.g., driving a car) by monitoring their actions with verified and trusted components (called enforcers ) that correct unsafe actions to guarantee critical safety properties (e.g., brake to prevent a crash). The enforcers are run within a verified hypervisor that protects them from security attacks or bugs and the untrusted components are run in an unverified virtual machine (VM) on top of the hypervisor. The untrusted and trusted components are executed as a single coordinated sporadic real-time task, called a mixed-trust task , where the untrusted part is known as the guest task (GT, because it runs in the guest VM) and the trusted part running in the hypervisor (HV) is known as the hypertask (HT). The GT is run by a preemptive fixed-priority scheduler in the VM and the HT by a non-preemptive fixed-priority scheduler in the HV. The non-preemptive scheduler prevents interleavings and simplifies the logical verification [4] , [5] . From a timing point of view, the HT monitors that the GT produces a valid output before the deadline, and if not, the HT itself produces a safe output before the deadline elapses. A new set of schedulability equations to evaluate their schedulability were presented in [2] along with a full discussion of the framework. Dionisio de Niz, Björn Andersson, Hyoseung Kim 0001, Mark Klein 0003, John P. Lehoczky |
RTSS | 3 |
| 2019 | Job-Class-Level Fixed Priority Scheduling of Weakly-Hard Real-Time SystemsabstractMany cyber-physical applications including sensing and control operations can tolerate a certain degree of timing violations as long as the number of the violations are predictably bounded. The notion of weakly-hard real-time systems has been studied to capture this effect, but existing work reveals limitations for practical use due the restrictions imposed on timing model and the high complexity of analysis. In this paper, we propose a new job-class-level fixed-priority preemptive scheduler and its schedulability analysis framework for sporadic tasks with weakly-hard real-time constraints. Our proposed scheduler employs the meet-oriented classification of jobs of a task in order to reduce the worst-case temporal interference imposed on other tasks. Under this approach, each job is associated with a "job-class" that is determined by the number of deadlines previously met (with a bounded number of consecutively-missed deadlines). This approach also allows decomposing the complex weakly-hard schedulability problem into two sub-problems that are easier to solve: (1) analyzing the response time of a job with each job-class, which can be done by an extension of the existing task-level analysis, and (2) finding possible job-class patterns, which can be modeled as a simple reachability tree. Experimental results indicate that our scheduler outperforms prior work in terms of task schedulability and analysis time complexity. We have also implemented a prototype of a job-class-level scheduler in the Linux kernel running on Raspberry Pi with acceptably-small runtime overhead. Hyunjong Choi, Hyoseung Kim 0001, Qi Zhu 0002 |
RTAS | 2 |
| 2019 | Thermal-Aware Servers for Real-Time Tasks on Multi-Core GPU-Integrated Embedded SystemsabstractThe recent trend in real-time applications raises the demand for powerful embedded systems with GPU-CPU integrated systems-on-chips (SoCs). This increased performance, however, comes at the cost of power consumption and resulting heat dissipation. Heat conduction interferes the execution time of tasks running on adjacent CPU and GPU cores. The violation of thermal constraints causes timing unpredictability to real-time tasks due transient performance degradation or permanent system failure. In this paper, we propose a thermal-aware server framework to safely upper bound the maximum temperature of GPU-CPU integrated systems running real-time sporadic tasks. Our framework supports variants of real-time server policies for CPU and GPU cores to satisfy both thermal and timing requirements. In addition, the framework incorporates two mechanisms, miscellaneous-operation-time reservation and pre-ordered scheduling of GPU requests, which significantly reduce task response time. We present analysis to design thermal-server budget and to check the schedulability of CPU-only and GPU-using sporadic tasks. The thermal properties of our framework have been evaluated on a commercial embedded platform. Experimental results with randomly-generated tasksets demonstrate the performance characteristics of our framework with different configurations. Seyedmehdi Hosseinimotlagh, Hyoseung Kim 0001 |
RTAS | 2 |
| 2019 | Mixed-Trust Computing for Real-Time SystemsabstractVerifying complex Cyber-Physical Systems (CPS) is increasingly important given the push to deploy safety-critical autonomous features. Unfortunately, traditional verification methods do not scale to the complexity of these systems and do not provide systematic methods to protect verified properties when not all the components can be verified. To address these challenges, this paper proposes a real-time mixed-trust computing framework that combines verification and protection. The framework introduces a new task model, where an application task can have both an untrusted and a trusted part. The untrusted part allows complex computations supported by a full OS with a realtime scheduler running in a VM hosted by a trusted hypervisor. The trusted part is executed by another scheduler within the hypervisor and is thus protected from the untrusted part. If the untrusted part fails to finish by a specific time, the trusted part is activated to preserve safety (e.g., prevent a crash) including its timing guarantees. This framework is the first allowing the use of untrusted components for CPS critical functions while preserving logical and timing guarantees, even in the presence of malicious attackers. We present the framework design and implementation along with the schedulability analysis and the coordination protocol between the trusted and untrusted parts. We also present our Raspberry Pi 3 implementation along with experiments showing the behavior of the system under failures of untrusted components, and a drone application to demonstrate its practicality. Dionisio de Niz, Björn Andersson, Mark Klein 0003, John P. Lehoczky, Amit Vasudevan, Hyoseung Kim 0001, Gabriel A. Moreno |
RTCSA | 6 |
| 2019 | STGM: Spatio-Temporal GPU Management for Real-Time TasksabstractGraphics Processing Units (GPUs) have been considered as a promising technology to address the high computational demands of real-time data-intensive applications. Today's embedded processors already offer on-chip GPUs, the use of which can greatly help satisfy the timing requirements of realtime tasks by accelerating their execution. However, existing GPU management schemes either underutilize the GPU due to strictly serialized execution or introduce non-deterministic delay caused by uncontrolled concurrent execution. In this paper, we present a spatial-temporal GPU management framework that controls the allocation and sharing of GPU's internal execution engines, e.g., streaming multiprocessors in Nvidia architectures, with analytical bounds. This approach allows multiple GPU-using tasks to simultaneously execute on the GPU, thereby improving GPU utilization and reducing the worst-case response time. Also, it can improve temporal isolation by allocating a portion of GPU execution engines to tasks for their exclusive use. We have examined the feasibility of our framework on two Nvidia GPUs: GTX970 and AGX Xavier. Experimental results with randomly-generated tasksets indicate that our framework yields a significant benefit in schedulability compared to the existing real-time GPU management approaches. Sujan Kumar Saha, Yecheng Xiang, Hyoseung Kim 0001 |
RTCSA | 3 |
| 2019 | Work-in-Progress: Understanding the Effect of Kernel Scheduling on GPU Energy ConsumptionabstractGeneral-purpose graphics processing units (GPUs) made available on embedded platforms have gained much interest in real-time cyber-physical systems. Despite the fact that GPUs generally outperform CPUs on many compute-intensive tasks in a multitasking environment, higher power consumption remains a challenging problem. This paper presents our study on the energy consumption characteristics of an NVIDIA AGX Xavier GPU, the latest commercially available embedded hardware, under different concurrency levels and kernel scheduling orders. Our findings pave the way for designing an energy efficient scheduler for GPUs with real-time guarantees. Yidi Wang 0001, Hyoseung Kim 0001 |
RTSS | 2 |
| 2019 | Pipelined Data-Parallel CPU/GPU Scheduling for Multi-DNN Real-Time InferenceabstractDeep neural networks (DNNs) have been showing significant success in various applications, such as autonomous driving, mobile devices, and Internet of Things. Although much research has been conducted to optimize the structure of DNNs, limited attention has been given to their timely execution, specifically on the scheduling of real-time inference requests to various DNN models. For instance, existing DNN frameworks, such as Caffe, TensorFlow and Torch, only provide a single-level priority, one-DNN-per-process execution model and sequential inference interfaces. They can be particularly problematic when used in edge computing and in-vehicle intelligence systems for multiple DNNs, as response time may become unpredictably long in the worst case while leaving system resources underutilized. This paper presents DART, a DNN scheduling framework that offers deterministic response time to real-time tasks and increased throughput to best-effort tasks. DART employs a pipeline-based scheduling architecture with data parallelism, where heterogeneous CPUs and GPUs are arranged into nodes with different parallelism levels. DART also includes pipeline stage design and node configuration schemes, admission control, execution time profiling, and runtime enforcement techniques. We evaluated DART on Intel x86 Xeon and Nvidia ARM platforms with GPUs. Experimental results indicate that DART significantly outperforms the existing approaches, by up to 98.5% shorter worst-case response time for real-time tasks while simultaneously achieving up to 17.9% higher throughput for best-effort tasks. Yecheng Xiang, Hyoseung Kim 0001 |
RTSS | 2 |
| 2018 | Analytical Enhancements and Practical Insights for MPCP with Self-SuspensionsabstractHardware accelerators such as GP-GPUs and DSPs are being increasingly used in computationally-intensive real-time and multimedia systems. System efficiency is often increased when CPU tasks suspend while using these devices. In this paper, we extend the existing Multiprocessor Priority Ceiling Protocol (MPCP) schedulability analysis in this particular context. We present three methods to improve the traditional MPCP analysis that reduces pessimism in analyzing blocking times. Two of these methods, the request-driven and the job-driven approaches, are motivated by prior work and are adapted to MPCP. The third combines these two approaches in a novel way to consistently outperform either on its own. We note that our underlying observations are general, and that such methods can also be used for analyzing other real-time synchronization protocols. Experimental results indicate that our analytical improvements result in a significantly higher schedulability compared to the traditional recursion-based analysis, even when self-suspensions are not considered. Our approach is also competitive with and often outperforms the linear-programming-based FMLP+ analysis, while having a considerably lower runtime complexity. We further substantiate the practical feasibility of suspension-based MPCP and examine its benefits over the busy-waiting approach by presenting a case-study on an NVIDIA TX2 embedded platform using real-world vision applications. Pratyush Patel, Iljoo Baek, Hyoseung Kim 0001, Ragunathan Rajkumar |
RTAS | 3 |
| 2018 | A server-based approach for predictable GPU access with improved analysis
Hyoseung Kim 0001, Pratyush Patel, Shige Wang, Ragunathan Rajkumar |
J. Syst. Archit. | 1 |
| 2018 | Schedulability Analysis of Tasks with Corunner-Dependent Execution TimesabstractConsider fixed-priority preemptive partitioned scheduling of constrained-deadline sporadic tasks on a multiprocessor. A task generates a sequence of jobs and each job has a deadline that must be met. Assume tasks have Corunner-dependent execution times; i.e., the execution time of a job J depends on the set of jobs that happen to execute (on other processors) at instants when J executes. We present a model that describes Corunner-dependent execution times. For this model, we show that exact schedulability testing is co-NP-hard in the strong sense. Facing this complexity, we present a sufficient schedulability test, which has pseudo-polynomial-time complexity if the number of processors is fixed. We ran experiments with synthetic software benchmarks on a quad-core Intel multicore processor with the Linux/RK operating system and found that for each task, its maximum measured response time was bounded by the upper bound computed by our theory. Björn Andersson, Hyoseung Kim 0001, Dionisio de Niz, Mark Klein 0003, Ragunathan Rajkumar, John P. Lehoczky |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2018 | Predictable Shared Cache Management for Multi-Core Real-Time VirtualizationabstractReal-time virtualization has gained much attention for the consolidation of multiple real-time systems onto a single hardware platform while ensuring timing predictability. However, a shared last-level cache (LLC) on modern multi-core platforms can easily hamper the timing predictability of real-time virtualization due to the resulting temporal interference among consolidated workloads. Since such interference caused by the LLC is highly variable and may have not even existed in legacy systems to be consolidated, it poses a significant challenge for real-time virtualization. In this article, we propose a predictable shared cache management framework for multi-core real-time virtualization. Our framework introduces two hypervisor-level techniques, vLLC and vColoring, that enable the cache allocation of individual tasks running in a virtual machine (VM), which is not achievable by the current state of the art. Our framework also provides a cache management scheme that determines cache allocation to tasks, designs VMs in a cache-aware manner, and minimizes the aggregated utilization of VMs to be consolidated. As a proof of concept, we implemented vLLC and vColoring in the KVM hypervisor running on x86 and ARM multi-core platforms. Experimental results with three different guest OSs (i.e., Linux/RK, vanilla Linux, and MS Windows Embedded) show that our techniques can effectively control the cache allocation of tasks in VMs. Our cache management scheme yields a significant utilization benefit compared to other approaches while satisfying timing constraints. Hyoseung Kim 0001, Ragunathan Rajkumar |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2017 | Mixed-criticality processing pipelinesabstractWhile a number of schemes exist for mixed-criticality scheduling in a single processor setting, no solution exists to cover the industry need for end-to-end scheduling across multiple processors in a pipeline. In this paper, we present an end-to-end zero-slack rate-monotonic scheme (ZSRM) based on real-time pipelines, called the ZSRM pipeline scheduler, that addresses this need. Under ZSRM, each task is associated with a parameter called zero-slack instant, and whenever a higher-criticality job has not finished at its zero-slack instant relative to its arrival time, all jobs of lower criticality are suspended to meet the deadline of the higher-criticality job. We develop a new schedulability test and algorithm for computing the zero-slack instants of tasks scheduled across a pipeline. Dionisio de Niz, Björn Andersson, Hyoseung Kim 0001, Mark Klein 0003, Linh T. X. Phan, Ragunathan Rajkumar |
DATE | 3 |
| 2017 | A server-based approach for predictable GPU access controlabstractWe propose a server-based approach to manage a general-purpose graphics processing unit (GPU) in a predictable and efficient manner. Our proposed approach introduces a GPU server task that is dedicated to handling GPU requests from other tasks on their behalf. The GPU server ensures bounded time to access the GPU, and allows other tasks to suspend during their GPU computation to save CPU cycles. By doing so, we address the two major limitations of the existing real-time synchronization-based GPU management approach: busy waiting and long priority inversion. We implemented a prototype of the server-based approach on a real embedded platform. This case study demonstrates the practicality and effectiveness of the server-based approach. Experimental results indicate that the server-based approach yields significant improvements in task schedulability over the existing synchronization-based approach in most practical settings. Although we focus on a GPU in this paper, the server-based approach can also be used for other types of computational accelerators. Hyoseung Kim 0001, Pratyush Patel, Shige Wang, Ragunathan Rajkumar |
RTCSA | 1 |
| 2016 | Real-time cache management for multi-core virtualizationabstractReal-time virtualization techniques have been investigated with the primary goal of consolidating multiple real-time systems onto a single hardware platform while ensuring timing predictability. However, a shared last-level cache (LLC) on recent multi-core platforms can easily hamper timing predictability due to the resulting temporal interference among consolidated workloads. Since such interference caused by the LLC is highly variable and may have not even existed in legacy systems to be consolidated, it poses a significant challenge for real-time virtualization. In this paper, we propose a real-time cache management framework for multi-core virtualization. Our framework introduces two hypervisor-level techniques, vLLC and vColoring, that enable the cache allocation of individual tasks running in a virtual machine (VM), which is not achievable by the current state of the art. Our framework also provides a cache management scheme that determines cache allocation to tasks, designs VMs in a cache-aware manner, and minimizes the aggregated utilization of VMs to be consolidated. As a proof of concept, we implemented vLLC and vColoring in the KVM hypervisor running on x86 and ARM multi-core platforms. Experimental results with three different guest OSs, namely Linux/RK, vanilla Linux and MS Windows Embedded, show that our techniques can effectively control the cache allocation of tasks in VMs. Our cache management scheme yields a significant utilization benefit compared to other approaches. Hyoseung Kim 0001, Ragunathan Rajkumar |
EMSOFT | 1 |
| 2016 | Bounding and reducing memory interference in COTS-based multi-core systems
Hyoseung Kim 0001, Dionisio de Niz, Björn Andersson, Mark Klein 0003, Onur Mutlu, Ragunathan Rajkumar |
Real Time Syst. | 1 |
| 2015 | Responsive and Enforced Interrupt Handling for Real-Time System VirtualizationabstractThe increasing performance of modern processors makes virtualization a viable solution for consolidating real-time systems into a single hardware platform. Although real-time task scheduling in a virtual machine can benefit from hierarchical scheduling, unbounded interrupt handling time and vulnerability to interrupt storms make practitioners hesitant to virtualize interrupt-driven real-time applications. In this paper, we propose vINT, an interrupt handling scheme designed for real-time system virtualization. vINT provides a pseudo-VCPU abstraction dedicated for interrupt handling, which overcomes the limits imposed by the timing parameters of virtual CPUs in an analyzable way. vINT also accounts for and enforces interrupt handling and resulting execution flows within a guest virtual machine. vINT does not require any change to the guest OS code, so it can be used for virtualizing proprietary, closed-source OSs. We analyze interrupt handling time as well as VCPU and task schedulability, with and without vINT. Our experimental results indicate that vINT achieves timely interrupt handling while providing as good task schedulability as when it is not used. Our case study based on a prototype implementation on the KVM hyper visor shows that vINT yields significant benefits in reducing interrupt handling time and in protecting real-time tasks against interrupt storms permeating into the virtual machine. Hyoseung Kim 0001, Shige Wang, Ragunathan Rajkumar |
RTCSA | 1 |
| 2014 | Bounding memory interference delay in COTS-based multi-core systemsabstractIn commercial-off-the-shelf (COTS) multi-core systems, a task running on one core can be delayed by other tasks running simultaneously on other cores due to interference in the shared DRAM main memory. Such memory interference delay can be large and highly variable, thereby posing a significant challenge for the design of predictable real-time systems. In this paper, we present techniques to provide a tight upper bound on the worst-case memory interference in a COTS-based multi-core system. We explicitly model the major resources in the DRAM system, including banks, buses and the memory controller. By considering their timing characteristics, we analyze the worst-case memory interference delay imposed on a task by other tasks running in parallel. To the best of our knowledge, this is the first work bounding the request re-ordering effect of COTS memory controllers. Our work also enables the quantification of the extent by which memory interference can be reduced by partitioning DRAM banks. We evaluate our approach on a commodity multi-core platform running Linux/RK. Experimental results show that our approach provides an upper bound very close to our measured worst-case interference. Hyoseung Kim 0001, Dionisio de Niz, Björn Andersson, Mark Klein 0003, Onur Mutlu, Ragunathan Rajkumar |
RTAS | 1 |
| 2014 | vMPCP: A Synchronization Framework for Multi-core Virtual MachinesabstractThe virtualization of real-time systems has received much attention for its many benefits, such as the consolidation of individually developed real-time applications while maintaining their implementations. However, the current state of the art still lacks properties required for resource sharing among real-time application tasks in a multi-core virtualization environment. In this paper, we propose vMPCP, a synchronization framework for the virtualization of multi-core real-time systems. Vmpcp exposes the executions of critical sections of tasks in a guest virtual machine to the hyper visor. Using this approach, vMPCP reduces and bounds blocking time on accessing resources shared within and across virtual CPUs (VCPUs) assigned on different physical CPU cores. Vmpcp supports periodic server and deferrable server policies for the VCPU budget replenish policy, with an optional budget overrun to reduce blocking times. We provide the VCPU and task schedulability analyses under vMPCP, with different VCPU budget supply policies, with and without overrun. Experimental results indicate that, under vMPCP, deferrable server outperforms periodic server when overrun is used, with as much as 80% more task sets being schedulable. The case study using our hyper visor implementation shows that vMPCP yields significant benefits compared to a virtualization-unaware multi-core synchronization protocol, with 29% shorter response time on average. Hyoseung Kim 0001, Shige Wang, Ragunathan Rajkumar |
RTSS | 1 |
| 2014 | Memory reservation and shared page management for real-time systems
Hyoseung Kim 0001, Ragunathan Rajkumar |
J. Syst. Archit. | 1 |
| 2013 | A Coordinated Approach for Practical OS-Level Cache Management in Multi-core Real-Time SystemsabstractMany modern multi-core processors sport a large shared cache with the primary goal of enhancing the statistic performance of computing workloads. However, due to resulting cache interference among tasks, the uncontrolled use of such a shared cache can significantly hamper the predictability and analyzability of multi-core real-time systems. Software cache partitioning has been considered as an attractive approach to address this issue because it does not require any hardware support beyond that available on many modern processors. However, the state-of-the-art software cache partitioning techniques face two challenges: (1) the memory co-partitioning problem, which results in page swapping or waste of memory, and (2) the availability of a limited number of cache partitions, which causes degraded performance. These are major impediments to the practical adoption of software cache partitioning. In this paper, we propose a practical OS-level cache management scheme for multi-core real-time systems. Our scheme provides predictable cache performance, addresses the aforementioned problems of existing software cache partitioning, and efficiently allocates cache partitions to schedule a given task set. We have implemented and evaluated our scheme in Linux/RK running on the Intel Core i7 quad-core processor. Experimental results indicate that, compared to the traditional approaches, our scheme is up to 39% more memory space efficient and consumes up to 25% less cache partitions while maintaining cache predictability. Our scheme also yields a significant utilization benefit that increases with the number of tasks. Hyoseung Kim 0001, Arvind Kandhalu, Ragunathan Rajkumar |
ECRTS | 1 |
| 2012 | Shared-Page Management for Improving the Temporal Isolation of Memory Reservations in Resource KernelsabstractMemory reservation provides real-time applications with guaranteed memory access to a specified amount of physical memory. However, previous work on memory reservation primarily focused on private pages, and did not pay attention to shared pages, which are widely used in current operating systems. With previous schemes, a real-time application may experience unexpected timing delays from other applications through shared pages that are shared by another process, even though the application has enough free pages in its reservation. In this paper, we describe problems with shared pages in real-time applications, and propose a shared-page management mechanism to enhance the temporal isolation of memory reservations in resource kernels that use resource reservation. The proposed mechanism consists of two techniques, Shared-Page Conservation (SPC) and Shared-Page Eviction Lock (SPEL), each of which prevents timing penalties caused by the seemingly arbitrary eviction of shared pages. The mechanism can manage shared data for inter-process communication and shared libraries, as well as pages shared by the kernel's copy-on-write technique and file caches. We have implemented and evaluated our schemes on the Linux/RK platform, but it can be applied to other operating systems with paged virtual memory. Hyoseung Kim 0001, Ragunathan Rajkumar |
RTCSA | 1 |
| 2010 | A Decentralized Approach for Monitoring Timing Constraints of Event FlowsabstractThis paper presents a run-time monitoring framework to detect end-to-end timing constraint violations of event flows in distributed real-time systems. The framework analyzes every event on possible event flow paths and automatically inserts timing fault checks for run-time detection. When the framework detects a timing violation, it provides users with the event flow's run-time path and the time consumption of each participating software module. In addition, it invokes a timing fault handler according to the timing fault specification, which allows our approach to aid the monitoring and management of the deployed systems. The experimental results show that the framework correctly detects timing constraint with insignificant overhead and provides related diagnostic information. Hyoseung Kim 0001, Shinyoung Yi 0002, Wonwoo Jung, Hojung Cha |
RTSS | 1 |
| 2007 | Multithreading Optimization Techniques for Sensor Network Operating Systems
Hyoseung Kim 0001, Hojung Cha |
EWSN | 1 |
| 2007 | RETOS: resilient, expandable, and threaded operating system for wireless sensor networksabstractThis paper presents the design principles, implementation, and evaluation of the RETOS operating system which is specifically developed for micro sensor nodes. RETOS has four distinct objectives, which are to provide (1) a multithreaded programming interface, (2) system resiliency, (3) kernel extensibility with dynamic reconfiguration, and (4) WSN-oriented network abstraction. RETOS is a multithreaded operating system, hence it provides the commonly used thread model of programming interface to developers. We have used various implementation techniques to optimize the performance and resource usage of multithreading. RETOS also provides software solutions to separate kernel from user applications, and supports their robust execution on MMU-less hardware. The RETOS kernel can be dynamically reconfigured, via loadable kernel framework, so a application-optimized and resource-efficient kernel is constructed. Finally, the networking architecture in RETOS is designed with a layering concept to provide WSN-specific network abstraction. RETOS currently supports Atmel ATmega128, TI MSP430, and Chipcon CC2430 family of microcontrollers. Several real-world WSN applications are developed for RETOS and the overall evaluation of the systems is described in the paper. Hojung Cha, Sukwon Choi, Inuk Jung, Hyoseung Kim 0001, Hyojeong Shin, Jaehyun Yoo, Chanmin Yoon |
IPSN | 4 |
| 2007 | The RETOS operating system: kernel, tools and applicationsabstractThis demonstration shows the programming development suite of the RETOS operating system for sensor networks, which provides a robust and multithreaded programming interface to application programmers. We first demonstrate how to build the RETOS kernel on the TI MSP430, Atmel ATmega 128 and Chipcon CC2430 family of microcontrollers. The application or a kernel module is then compiled and disseminated, via wireless channel, to the target motes. The GUI-based RMon network management tool for RETOS is also demonstrated to monitor the networked sensors, and even to control the system's parameters or applications running on them via a remote shell. The system is demonstrated to run on a mixed set of MSP430, ATmega128 and CC2430-based motes. Overall, our demonstration will convince attendees of the programming convenience of developing sensor network applications using RETOS, which is, indeed, a mature and practical system that can be used to develop real-world applications. Hojung Cha, Sukwon Choi, Inuk Jung, Hyoseung Kim 0001, Hyojeong Shin, Jaehyun Yoo, Chanmin Yoon |
IPSN | 4 |
| 2007 | Dynamic refresh-rate scaling via frame buffer monitoring for power-aware LCD managementabstractAbstract In recent years, there has been wide‐spread use of large and high‐resolution liquid‐crystal displays (LCDs) on handheld devices. The portion of LCD power consumption in the overall system has gradually increased. While most of the previous research on LCD power management has focused on the hardware level, practical mechanisms at the software level are hardly known. This paper presents a power‐aware LCD management mechanism, based on dynamic refresh‐rate scaling and frame buffer monitoring. The proposed mechanism guarantees the display quality of service, which is inherently specified by content types. The mechanism does not require additional hardware or modifications to applications. The experiment results—on a commercial PDA with a $320\times240$ resolution—show that the proposed mechanisms effectively reduce the power consumption by up to 10%, while satisfying the display quality requirements for the LCD screen. Copyright © 2006 John Wiley & Sons, Ltd. Hyoseung Kim 0001, Hojung Cha, Rhan Ha |
Softw. Pract. Exp. | 1 |
| 2006 | Towards a Resilient Operating System for Wireless Sensor Networks
Hyoseung Kim 0001, Hojung Cha |
USENIX ATC, General Track | 1 |