EDBT 2026 Demo / reviewers in the wild / expert
Woosung Kang 0002
dblp:401/7723-2
· DBLP profile ↗
6ranked-venue papers
4as first author
6since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ZeroSwap: Minimizing Swap Overhead for Real-Time Multi-DNN Inference via SSD-based GPU Memory Extension
Woosung Kang 0002, Filippo Muzzini, Gianluca Brilli, Jinkyu Lee 0001, Hoon Sung Chwa |
RTAS | 1 |
| 2025 | Mitigating Resource Contention for Responsive On-device Machine Learning InferencesabstractOn-device machine learning applications are increasingly deployed in dynamic and open system environments, where resource availability fluctuates unpredictably. This variability, coupled with limited computing resources, poses significant challenges in achieving high responsiveness. Existing on-device machine learning frameworks typically rely on static and coarse-grained resource allocation, leading to performance degradation under resource contention. To address this, we propose FlexOn, a novel framework that combines fine-grained model segmentation and dynamic resource selection to rapidly adapt to highly dynamic runtime conditions and effectively mitigate unpredictable resource contention. A prototype built on LiteRT demonstrates significant improvements in both average and tail latencies of up to 54% and 58%, respectively, across three different embedded platforms under dynamically varying resource availability. To the best of our knowledge, this is the first work that addresses the resource contention in open embedded systems for better machine learning inference responsiveness. Seongjin Chou, Whisoo Chung, Inwoo Kim, Woosung Kang 0002, Hyosu Kim, Sangeun Oh, Hoon Sung Chwa, Kilho Lee |
ICCAD | 6 |
| 2025 | Timing guarantees for inference of AI models in embedded systems
Seunghoon Lee 0002, Woosung Kang 0002, Marko Bertogna, Hoon Sung Chwa, Jinkyu Lee 0001 |
Real Time Syst. | 2 |
| 2024 | RT-Swap: Addressing GPU Memory Bottlenecks for Real-Time Multi-DNN InferenceabstractThe increasing complexity and memory demands of Deep Neural Networks (DNNs) for real-time systems pose new significant challenges, one of which is the GPU memory capacity bottleneck, where the limited physical memory inside GPUs impedes the deployment of sophisticated DNN models. This paper presents, to the best of our knowledge, the first study of addressing the GPU memory bottleneck issues, while simultaneously ensuring the timely inference of multiple DNN tasks. We propose RT-Swap, a real-time memory management framework, that enables transparent and efficient swap scheduling of memory objects, employing the relatively larger CPU memory to extend the available GPU memory capacity, without compromising timing guarantees. We have implemented RT-Swap on top of representative machine-learning frameworks, demonstrating its effectiveness in making significantly more DNN task sets schedulable at least 72% over existing approaches even when the task sets demand up to 96.2% more memory than the GPU's physical capacity. Woosung Kang 0002, Jinkyu Lee 0001, Youngmoon Lee, Sangeun Oh, Kilho Lee, Hoon Sung Chwa |
RTAS | 1 |
| 2022 | DNN-SAM: Split-and-Merge DNN Execution for Real-Time Object DetectionabstractAs real-time object detection systems, such as autonomous cars, need to process input images acquired from multiple cameras, they face significant challenges in delivering accurate and timely inferences often based on machine learning (ML). To meet these challenges, we want to provide different levels of object detection accuracy and timeliness to different portions within each input image with different criticality levels. Specifically, we develop DNN-SAM, a dynamic Split-And-Merge Deep Neural Network (DNN) execution and scheduling framework, that enables seamless split-and-merge DNN execution for unmodified DNN models. Instead of processing an entire input image once in a full DNN model, DNN-SAM first splits a DNN inference task into two smaller sub-tasks-a mandatory sub-task dedicated for a safety-critical (cropped) portion of each image and an optional sub-task for processing a down-scaled image–then executes them independently, and finally merges their results into a complete inference. To achieve DNN-SAM’s timely and accurate detection of objects in each image, we also develop two scheduling algorithms that prioritize sub-tasks according to their criticality levels and adaptively adjust the scale of the input image to meet the timing constraints while minimizing the response time of mandatory sub-tasks or maximizing the accuracy of optional sub-tasks. We have implemented and evaluated DNN-SAM on a representative ML framework. Our evaluation shows DNN-SAM to improve detection accuracy in the safety-critical region by $2.0-3.7\times$ and lower average inference latency by $4.8-9.7\times$ over existing approaches without violating any timing constraints. Woosung Kang 0002, Siwoo Chung, Jeremy Yuhyun Kim, Youngmoon Lee, Kilho Lee, Jinkyu Lee 0001, Kang G. Shin, Hoon Sung Chwa |
RTAS | 1 |
| 2021 | LaLaRAND: Flexible Layer-by-Layer CPU/GPU Scheduling for Real-Time DNN TasksabstractDeep neural networks (DNNs) have shown remarkable success in various machine-learning (ML) tasks useful for many safety-critical, real-time embedded systems. The foremost design goal for enabling DNN execution on real-time embedded systems is to provide worst-case timing guarantees with limited computing resources. Yet, the state-of-the-art ML frameworks hardly leverage heterogeneous computing resources (i.e., CPU, GPU) to improve the schedulability of real-time DNN tasks due to several factors, which include a coarse-grained resource allocation model (one-resource-per-task), the asymmetric nature of DNN execution on CPU and GPU, and lack of schedulability-aware CPU/GPU allocation scheme. This paper presents, to the best of our knowledge, the first study of addressing the above three major barriers and examining their cooperative effect on schedulability improvement. In this paper, we propose LaLaRAND, a real-time layer-level DNN scheduling framework, that enables flexible CPU/GPU scheduling of individual DNN layers by tightly coupling CPU-friendly quantization with fine-grained CPU/GPU allocation schemes (one-resource-per-layer) while mitigating accuracy loss without compromising timing guarantees. We have implemented and evaluated LaLaRAND on top of the state-of-the-art ML framework to demonstrate its effectiveness in making more DNN task sets schedulable by 56% and 80% over an existing approach and a baseline (vanilla PyTorch), respectively, with only up to -0.4% of performance (inference accuracy) difference. Woosung Kang 0002, Kilho Lee, Jinkyu Lee 0001, Insik Shin, Hoon Sung Chwa |
RTSS | 1 |