EDBT 2026 Demo / reviewers in the wild / expert
Heon-Chang Yu
dblp:80/2696 · also HeonChang Yu, Heonchang Yu
· DBLP profile ↗
43ranked-venue papers
0as first author
13since 2021 · last 2026
0000-0003-2216-595XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 5 since 2021Human-computer interaction and ubiquitous computing · 9 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 4 since 2021Software engineering, systems software and programming languages · 4 · 3 since 2021Databases, data management, data science and information retrieval · 2Artificial intelligence and machine learning · 1Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Coordinating Traffic at Group Granularity for Tail Latency Control in Service Meshes
Hokun Park, HyungJun Kim, Hwa-Min Lee, Heon-Chang Yu |
ICDCS | 4 |
| 2026 | Vertical auto-scaling mechanism for elastic memory management of containerized applications in Kubernetes
Taeshin Kang, Heon-Chang Yu |
Future Gener. Comput. Syst. | 3 |
| 2026 | REX: Adaptive distributed training with statistically validated batch scaling and spline-guided Bayesian optimization on heterogeneous GPU clusters
Hwa-Min Lee, Heon-Chang Yu |
Future Gener. Comput. Syst. | 3 |
| 2025 | Korel: Mitigating Stragglers via Real-Time Automatic Mixed Precision in Distributed Deep Learning EnvironmentsabstractIn distributed deep learning systems, straggler nodes are a primary factor in delaying gradient synchronization during synchronous training, thereby diminishing overall efficiency. This issue is particularly pronounced in multi-tenant environments, where resource contention from concurrent tasks exacerbates node slowdowns. Existing approaches typically mitigate this problem by excluding slower nodes or reducing their influence during gradient aggregation, which often leads to resource underutilization. In this study, we introduce a novel solution that leverages Automatic Mixed Precision (AMP) to tackle the straggler problem. Our method employs a MAD-based threshold to detect workers impaired by resource contention and dynamically applies AMP to accelerate their computations. Once contention subsides and node performance recovers, AMP is deactivated to maximize efficiency. Experimental evaluations across diverse deep learning tasks reveal that our approach significantly reduces time-to-accuracy in synchronous Distributed Data Parallel (DDP) environments, enabling rapid convergence to high accuracy. In addition, by effectively counteracting the delay imposed by stragglers-quantified by our mitigation metric which indicates that up to 79.4% of the straggler-induced delay is offset-our approach achieves convergence up to 14 % faster than traditional BSP methods under adverse conditions. Hyunseung Jung, HyungJun Kim, Heon-Chang Yu |
CLOUD | 3 |
| 2025 | MOBOS: Co-Optimizing Cost and Execution Time in Serverless Workflow with Multi-Objective Bayesian OptimizationabstractServerless computing has established itself as a prominent paradigm in cloud computing by enabling developers to focus exclusively on application development without the burden of infrastructure management. In the serverless environment, resource configuration of Function-as-a-Service (FaaS) is a critical factor that directly impacts both execution time and cost; however, determining the optimal configuration remains challenging. This challenge is particularly pronounced in workflows where multiple functions execute sequentially, as individual function configurations influence the overall workflow execution time, thereby increasing the complexity of resource configuration and making it more difficult to identify optimal settings. We propose MOBOS, a method that simultaneously optimizes both execution time and cost of serverless workflows using multi-objective optimization techniques. Our proposed approach is based on Bayesian Optimization, which effectively explores the trade-off between two objectives, execution times and cost, to derive Pareto-optimal memory configurations. Experimental results demonstrate that MOBOS outperforms state-of-the-art single-objective optimization methods and achieves up to 33.83 % faster execution time while simultaneously reducing costs by 13.54 % in real-world application workflows. Moreover, MOBOS provides additional Pareto-optimal configurations with equivalent trade-offs, enabling flexible resource allocation based on specific performance requirements. Heon-Chang Yu |
CLOUD | 2 |
| 2025 | ReSACO: A Meta Reinforcement Learning Method for Fast Offloading in Mobile Edge ComputingabstractMobile Edge Computing (MEC) has emerged as a promising paradigm for latency-sensitive and resource-intensive applications. However, due to the heterogeneous nature of applications, fluctuating workloads, and dynamic resource availability in MEC environments, making optimal offloading decisions remains a challenge. Traditional heuristics and methods based on artificial intelligence can be time consuming to train, which poses a challenge in continually evolving edge-cloud environments where rapid adaptation is critical. In this paper, we propose ReSACO, Reptile-based Soft Actor-Critic for Offloading, a meta-reinforcement learning framework that leverages the meta-learning algorithm with Soft Actor-Critic (SAC). By learning a generalizable policy across multiple scenarios, ReSACO can quickly adapt to new conditions with minimal retraining overhead, enabling robust offloading decisions in the face of unpredictable network fluctuations and resource constraints. We conducted extensive simulations using EdgeCloudSim to validate its performance. Experimental results show that ReSACO achieves up to 16% faster service time under heavy load while substantially reducing both network and VM-based failures. These findings demonstrate the effectiveness of combining meta-learning with entropy-regularized reinforcement learning to address the complexity and variability of edge-cloud environments, offering a scalable and efficient solution for ensuring low latency, high reliability, and fast adaptability in rapidly changing MEC deployments. Myeongjun Kim, Heon-Chang Yu |
CLOUD | 2 |
| 2025 | HEART: Heterogeneous-Aware Traffic Allocation in Multi-Replica Deployments on KubernetesabstractTraffic scheduling for microservices in emerging edge-cloud environments is challenged by heterogeneous node performance and varying inter-node latencies. Inefficient scheduling among replicas may lead to replica overload and excessive communication overhead, ultimately degrading Quality of Service (QoS) metrics-specifically, the 99th percentile tail latency(P99 latency). To address these challenges, we introduce HEART, a novel two-stage traffic scheduler that jointly accounts for node heterogeneity and network latency. In Stage 1, HEART computes per-replica traffic proportions using a sliding window and an exponentially weighted moving average of CPU usage and request rates, thereby capturing recent load trends. In Stage 2, HEART applies k-means clustering to inter-node latency measurements to identify and prune high-latency links, which in turn reduces communication delays. A Maximum Flow algorithm is subsequently employed to verify whether the pruned network supports the required traffic flow; if the reduced network proves insufficient, additional links are incrementally reinstated until a feasible configuration is achieved. Finally, a Minimum-Cost Flow algorithm-selected for its proven optimality in network flow allocation-is applied to distribute traffic cost-effectively across the network. Experimental results demonstrate that HEART significantly reduces the P99 latency compared to existing approaches such as the default Kubernetes scheduler, the Least- Request algorithm of Istio, OptTraffic, and LATA, thereby enhancing overall QoS. Hokun Park, HyungJun Kim, Gyujeong Lim, Heon-Chang Yu |
CLOUD | 5 |
| 2025 | SUPLEC:Microservice Scheduling Under Unexpected Peak Load in the Edge-Cloud ContinuumabstractMicroservice architectures have become widely adopted due to their scalability and flexibility. These applications are commonly deployed in edge-cloud environments. However, in such environments, unexpected peak loads on microservices often lead to violations of service level objectives (SLOs) due to limited compute resources and bandwidth. Moreover, resource utilization of nodes and network states are constantly changing, making it essential to reschedule microservices to ensure SLO compliance. To minimize downtime during rescheduling, the microservice temporarily occupies additional resources until the process is complete. Therefore, frequent rescheduling can result in inefficient resource utilization. In this paper, we introduce SUPLEC, a new scheduling technique that optimally schedules container-based microservices and minimizes the number of reschedules to handle unexpected peak loads in Edge-Cloud environments. SUPLEC consists of two components: Placer and Rescheduler. The Placer identifies frequently interacting microservices, clusters them into groups, and places them on nodes using a worst-fit approach. The Rescheduler analyzes the response times of recent requests to determine whether rescheduling should be performed and, if so, it reschedules a microservice to a better node. Experiment results show that the proposed strategy reduces the average response time by up to 68.5 % under dynamic cluster conditions and loads and reduces the number of reschedules by up to 80 %. Taeshin Kang, Joon-Min Gil, Heon-Chang Yu |
CCGrid | 4 |
| 2025 | Mutli-Metric based GPU Scoring Method in Kubernetes Environments
Jae-Hyung Kim, Woosuk Lee, Sang Hyeop Oh, Hyunsu Jeong, Heon-Chang Yu, Joon-Min Gil |
SERA | 5 |
| 2025 | SADDLE: A runtime feedback control architecture for adaptive distributed deep learning in heterogeneous GPU clusters
Eunyoung Lee, Heon-Chang Yu |
J. Syst. Archit. | 3 |
| 2025 | A-BEE-C: Autonomous Bandwidth-Efficient Edge Codecast
Gyujeong Lim, Joon-Min Gil, Heon-Chang Yu |
Pervasive Mob. Comput. | 3 |
| 2024 | LARE-HPA: Co-optimizing Latency and Resource Efficiency for Horizontal Pod Autoscaling in Kubernetes
EunYoung Lee, Heon-Chang Yu |
ICSOC (2) | 4 |
| 2024 | Optimizing Traffic Allocation for Multi-replica Microservice Deployments in Edge Cloud
Hokun Park, Gyujeong Lim, Heon-Chang Yu |
ICSOC (1) | 5 |
| 2020 | Partial migration technique for GPGPU tasks to Prevent GPU Memory Starvation in RPC-based GPU VirtualizationabstractSummary Graphics processing unit (GPU) virtualization technology enables a single GPU to be shared among multiple virtual machines (VMs), thereby allowing multiple VMs to perform GPU operations simultaneously with a single GPU. Because GPUs exhibit lower resource scalability than central processing units (CPUs), memory, and storage, many VMs encounter resource shortages while running GPU operations concurrently, implying that the VM performing the GPU operation must wait to use the GPU. In this paper, we propose a partial migration technique for general‐purpose graphics processing unit (GPGPU) tasks to prevent the GPU resource shortage in a remote procedure call‐based GPU virtualization environment. The proposed method allows a GPGPU task to be migrated to another physical server's GPU based on the available resources of the target's GPU device, thereby reducing the wait time of the VM to use the GPU. With this approach, we prevent resource shortages and minimize performance degradation for GPGPU operations running on multiple VMs. Our proposed method can prevent GPU memory shortage, improve GPGPU task performance by up to 14%, and improve GPU computational performance by up to 82%. In addition, experiments show that the migration of GPGPU tasks minimizes the impact on other VMs. JongBeom Lim, Heon-Chang Yu |
Softw. Pract. Exp. | 3 |
| 2019 | SafeDB: Spark Acceleration on FPGA Clouds with Enclaved Data Processing and Bitstream ProtectionabstractThis paper proposes SafeDB: Spark Acceleration on FPGA Clouds with Enclaved Data Processing and Bitstream Protection. SafeDB provides a comprehensive and systematic hardware-based security framework from the bitstream protection to data confidentiality, especially for the cloud environment. The AES key shared between FPGA and client for the bitstream encryption is generated in hard-wired logic using PKI and ECC. The data security is assured by the enclaved processing with encrypted data, meaning that the encrypted data is processed inside the FPGA fabric. Thus, no one in the system is able to look into clients' data because plaintext data are not exposed to memory and/or memory-mapped space. SafeDB is resistant not only to the side channel attack but to the attacks from malicious insiders. We have constructed an 8-node cluster prototype with Zynq UltraScale+ FPGAs to demonstrate the security, performance, and practicability. Han-Yee Kim, Rohyoung Myung, Boeui Hong, Heon-Chang Yu, Taeweon Suh, Lei Xu 0012, Larry Shi |
CLOUD | 4 |
| 2018 | Evaluation of P2P and cloud computing as platform for exhaustive key search on block ciphers
JunWeon Yoon, Taeyoung Hong, Jang Won Choi 0001, ChanYeol Park, Ki-Bong Kim, Heon-Chang Yu |
Peer-to-Peer Netw. Appl. | 6 |
| 2018 | Byzantine-resilient dual gossip membership management in clouds
JongBeom Lim, Kwang-Sik Chung, Hwa-Min Lee, Kangbin Yim, Heon-Chang Yu |
Soft Comput. | 5 |
| 2016 | Dynamic group-based fault tolerance technique for reliable resource management in mobile cloud computingabstractSummary Researches on utilizing mobile devices as resources in mobile cloud environments have gained attentions recently because of the enhanced computing power of mobile devices such as the advent of quad‐core chips. However, mobile devices have several problems of the availability and the mobility. Especially, the availability and the mobility of mobile devices cause system faults more frequently due to dynamic changes, and system faults prevent applications using mobile devices from being processed reliably. Therefore, we classify mobile devices into groups according to the availability and the mobility in order to manage reliable mobile resource. Because information of mobile devices is constantly changing, grouping should consider the dynamic environment. We also provide dynamic group‐based mobile cloud computing that applies fault tolerance techniques using checkpoints or replication in each group. The experimental result shows that our algorithm performs dynamic grouping, which is more suitable for the dynamic environment of mobile devices. Copyright © 2014 John Wiley & Sons, Ltd. Heon-Chang Yu, Hyongsoon Kim, EunYoung Lee |
Concurr. Comput. Pract. Exp. | 2 |
| 2014 | An Estimation-Based Task Load Balancing Scheduling in Spot Clouds
Daeyong Jung, Heeseok Choi, Dae-Won Lee, Heon-Chang Yu, EunYoung Lee |
NPC | 4 |
| 2014 | Gossip Membership Management with Social Graphs for Byzantine Fault Tolerance in Clouds
JongBeom Lim, Joon-Min Gil, Kwang-Sik Chung, Dae-Won Lee, Heon-Chang Yu |
NPC | 6 |
| 2014 | A scheduling algorithm with dynamic properties in mobile grid
Jong-Hyuk Lee, SungJin Choi, Joon-Min Gil, Taeweon Suh, Heon-Chang Yu |
Frontiers Comput. Sci. | 5 |
| 2012 | A Gossip-Based Mutual Exclusion Algorithm for Cloud Environments
JongBeom Lim, Kwang-Sik Chung, Sung-Ho Chin, Heon-Chang Yu |
GPC | 4 |
| 2012 | A Termination Detection Technique Using Gossip in Cloud Computing Environments
JongBeom Lim, Kwang-Sik Chung, Heon-Chang Yu |
NPC | 3 |
| 2011 | Group-Based Gossip Multicast Protocol for Efficient and Fault Tolerant Message Dissemination in Clouds
JongBeom Lim, Jong-Hyuk Lee, Sung-Ho Chin, Heon-Chang Yu |
GPC | 4 |
| 2011 | An Efficient Checkpointing Scheme Using Price History of Spot Instances in Cloud Computing Environment
Daeyong Jung, Sung-Ho Chin, Kwang-Sik Chung, Heon-Chang Yu, Joon-Min Gil |
NPC | 4 |
| 2010 | An Effective Job Replication Technique Based on Reliability and Performance in Mobile Grids
Daeyong Jung, Sung-Ho Chin, Kwang-Sik Chung, Taeweon Suh, Heon-Chang Yu, Joon-Min Gil |
GPC | 5 |
| 2010 | Monitoring Service Using Markov Chain Model in Mobile Grid Environment
Kwang-Sik Chung, EunYoung Lee, Young-Sik Jeong, Heon-Chang Yu |
GPC | 5 |
| 2010 | Adaptive service scheduling for workflow applications in Service-Oriented Grid
Sung-Ho Chin, Taeweon Suh, Heon-Chang Yu |
J. Supercomput. | 3 |
| 2009 | Performance Evaluation of Scheduling Mechanism with Checkpoint Sharing and Task Duplication in P2P-Based PC Grid Computing
Joon-Min Gil, Ui-Sung Song, Heon-Chang Yu |
GPC | 3 |
| 2009 | Balanced Scheduling Algorithm Considering Availability in Mobile Grid
Jong-Hyuk Lee, SungJin Song, Joon-Min Gil, Kwang-Sik Chung, Taeweon Suh, Heon-Chang Yu |
GPC | 6 |
| 2007 | Adaptive Workflow Scheduling Strategy in Service-Based Grids
Jong-Hyuk Lee, Sung-Ho Chin, Hwa-Min Lee, TaeMyoung Yoon, Kwang-Sik Chung, Heon-Chang Yu |
GPC | 6 |
| 2006 | Reducing Binding Updates in High Speed Movement Environment Based on HMIPv6
Dae-Won Lee, Kwang-Sik Chung, Sung-Ju Roh, KwangHee Choi, Heon-Chang Yu |
GPC | 5 |
| 2005 | Garbage Collection in a Causal Message Logging Protocol
Kwang-Sik Chung, Heon-Chang Yu |
HPCC | 2 |
| 2005 | Mobile Agent Based Adaptive Scheduling Mechanism in Peer to Peer Grid Computing
SungJin Choi, MaengSoon Baik, Chong-Sun Hwang, Joon-Min Gil, Heon-Chang Yu |
ICCSA (4) | 5 |
| 2005 | Group-Based Scheduling Scheme for Result Checking in Global Computing Systems
HongSoo Kim, SungJin Choi, MaengSoon Baik, KwonWoo Yang, Heon-Chang Yu, Chong-Sun Hwang |
ICCSA (3) | 5 |
| 2005 | A resource management and fault tolerance services in grid computing
Hwa-Min Lee, Kwang-Sik Chung, Sung-Ho Chin, Jong-Hyuk Lee, Dae-Won Lee, Heon-Chang Yu |
J. Parallel Distributed Comput. | 7 |
| 2004 | Volunteer Availability based Fault Tolerant Scheduling Mechanism in Desktop Grid Computing EnvironmentabstractFault tolerance is essential to the further development of desktop grid computing system in order to guarantee continuous and reliable execution of tasks in spite of failures. In a desktop grid computing environment, volunteers are often susceptible to volunteer autonomy failures such as volatility failure and interference failure in the middle of execution of tasks because a desktop grid computing maximally respects autonomy of volunteers. The failures result in an independent livelock problem (i.e. the delay and blocking of the entire execution of a job). Therefore, the failures should be considered in a scheduling mechanism. In This work, in order to tolerate volunteer autonomy failures, we propose a new fault tolerant scheduling mechanism. First, we specify a volunteer autonomy failures and an independent livelock problem. Then, we propose a volunteer availability which reflects the degree of volunteer autonomy failures. Finally, we propose a fault tolerant scheduling mechanism based on volunteer availability (which is called VAFTSM). SungJin Choi, MaengSoon Baik, Chong-Sun Hwang, Joon-Min Gil, Heon-Chang Yu |
NCA | 5 |
| 2003 | Managing Fault Tolerance Information in Multi-agents Based Distributed Systems
Dae-Won Lee, Kwang-Sik Chung, Hwa-Min Lee, Sungbin Park, Young-Jun Lee, Heon-Chang Yu, Won-Gyu Lee |
IDEAL | 6 |
| 2002 | A Recovery Technique Using Multi-agent in Distributed Computing Systems
Hwa-Min Lee, Kwang-Sik Chung, Sang-Chul Shin, Dae-Won Lee, Won-Gyu Lee, Heon-Chang Yu |
COORDINATION | 6 |
| 2002 | Efficient Garbage Collection Schemes for Causal Message Logging with Independent Checkpointing
JinHo Ahn, Sung-Gi Min, Chong-Sun Hwang, Heon-Chang Yu |
J. Supercomput. | 4 |
| 2001 | Optimistic Scheduling Algorithm for Mobile Transactions Based on Reordering
SungSuk Kim, Chong-Sun Hwang, Heon-Chang Yu, SangKeun Lee 0001 |
Mobile Data Management | 3 |
| 2001 | Revisiting Transaction Management in Multidatabase Systems
SangKeun Lee 0001, Chong-Sun Hwang, Heon-Chang Yu |
Distributed Parallel Databases | 3 |
| 1997 | Hybrid checkpointing protocol based on selective-sender-based message loggingabstractThis paper presents a hybrid checkpointing protocol-an asynchronous checkpointing protocol using a message sending/receiving state change for reducing the overhead of failure-free operation combined with a selective sender-based message logging protocol for reducing the cascade rollback of asynchronous checkpointing protocol. The selective sender-based message logging protocol records only potential orphan messages when taking a checkpoint. And this paper presents a message dependency tree recording the inter-process message sending/receiving information on a volatile storage for reducing the search time of inter-process information during the failure recovery. Kwang-Sik Chung, Kibom Kim, Chong-Sun Hwang, Jin Gon Shon, Heon-Chang Yu |
ICPADS | 5 |