EDBT 2026 Demo / reviewers in the wild / expert
Younggeun Kim 0001
dblp:294/6642-1 · also Young Geun Kim 0001, Young-geun Kim 0001
· DBLP profile ↗
20ranked-venue papers
10as first author
10since 2021 · last 2025
0000-0003-4713-819XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 17 · 10 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | First Attentions Last: Better Exploiting First Attentions for Efficient Parallel TrainingabstractAs training billion-scale transformers becomes increasingly common, employing multiple distributed GPUs along with parallel training methods has become a standard practice. However, existing transformer designs suffer from significant communication overhead, especially in Tensor Parallelism (TP), where each block’s MHA–MLP connection requires an all-reduce communication. Through our investigation, we show that the MHA-MLP connections can be bypassed for efficiency, while the attention output of the first layer can serve as an alternative signal for the bypassed connection. Motivated by the observations, we propose FAL (First Attentions Last), an efficient transformer architecture that redirects the first MHA output to the MLP inputs of the following layers, eliminating the per-block MHA-MLP connections. This removes the all-reduce communication and enables parallel execution of MHA and MLP on a single GPU. We also introduce FAL+, which adds the normalized first attention output to the MHA outputs of the following layers to augment the MLP input for the model quality. Our evaluation shows that FAL reduces multi-GPU training time by up to 44%, improves single-GPU throughput by up to 1.18×, and achieves better perplexity compared to the baseline GPT. FAL+ achieves even lower perplexity without increasing the training time than the baseline. Codes are available at: https://casl-ku.github.io/FAL/ Gyudong Kim, Hyukju Na, Jin Kyu Kim, Hyunsung Jang, Jaegi Hwang, Namkoo Ha, Seungryong Kim, Younggeun Kim 0001 |
NeurIPS | 9 |
| 2025 | GreenScale: Carbon Optimization for Edge ComputingabstractGiven billions of mobile users, the environmental impact of edge computing is significant. To address this, future applications need to execute computations on a green component which is fueled by renewable energy sources. However, because of the intermittent nature of the renewable energy sources, the carbon intensity of computing components can significantly vary with location and time of use. This poses a new challenge for edge applications – deciding when and where to run computations across consumer devices at the edge and servers in the cloud. Such scheduling decisions become more complicated with the amortization of the rising embodied emissions and stochastic runtime variance. This work proposes GreenScale, an intelligent execution scaling engine that accurately selects the carbon-optimal execution target for edge applications in different runtime environments. Our evaluation with three representative categories of applications (i.e., AI, Game, and AR/VR) demonstrate that the carbon emissions of the applications can be reduced by 35.2%, on average, with GreenScale. Yonglak Son, Udit Gupta 0001, Andrew McCrabb, Younggeun Kim 0001, Valeria Bertacco, David Brooks 0001, Carole-Jean Wu |
IEEE Internet Things J. | 4 |
| 2025 | Energy-Efficient, Delay-Constrained Edge Computing of a Network of DNNsabstractThis paper presents a novel approach for executing the inference of a network of pre-trained deep neural networks (DNNs) on commercial-off-the-shelf devices that are deployed at the edge. The problem is to partition the computation of the DNNs between an energy-constrained and performance-limited edge device$\boldsymbol{\mathcal{E}}$, and an energy-unconstrained, higher performance device$\boldsymbol{\mathcal{C}}$, referred to as thecloudlet, with the objective of minimizing the energy consumption of$\boldsymbol{\mathcal{E}}$subject to a deadline constraint. The proposed partitioning algorithm takes into account the performance profiles of executing DNNs on the devices, the power consumption profiles, and the variability in the delay of the wireless channel. The algorithm is demonstrated on a platform that consists of an NVIDIA Jetson Nano as the edge device$\boldsymbol{\mathcal{E}}$and a Dell workstation with a Titan Xp GPU as the cloudlet. Experimental results show significant improvements both in terms of energy consumption of$\boldsymbol{\mathcal{E}}$and processing delay of the application. Additionally, it is shown how the energy-optimal solution is changed when the deadline constraint is altered. Moreover, the overhead of decision-making for our proposed method is significantly lower than the state-of-the-art Integer Linear Programming (ILP) solutions. Mehdi Ghasemi 0003, Soroush Heidari, Younggeun Kim 0001, Carole-Jean Wu, Sarma B. K. Vrudhula |
IEEE Trans. Computers | 3 |
| 2024 | Elastic Execution of Multi-Tenant DNNs on Heterogeneous Edge MPSoCsabstractThe growing complexity of machine learning (ML) tasks drives the rapid deployment of multi-tenant ML workloads at the edge presenting unique challenges due to the variable computational demands and strict latency requirements. This paper introduces a holistic elastic scheduler, EMERALD, designed to optimize the execution of multi-tenant machine learning (ML) workloads on heterogeneous edge (Multiprocessor System on Chip) MPSoCs under strict runtime constraints. EMERALD employs input resolution scaling to dynamically adjust the computational demands of deep neural networks (DNNs), thereby enhancing the ability to meet stringent latency requirements while maintaining high accuracy. The scheduler consists of two main components: a local greedy scheduler and a global scheduler. The local scheduler actively manipulates input resolution in response to deadline violations, selecting the resolutions that minimally impact accuracy and maximally reduce response time. The global scheduler, an Integer Linear Programming (ILP)based scheduler, fine-tunes the decisions of the local scheduler by considering factors such as DNN dependencies, scene complexity, hardware heterogeneity, and the trade-offs between accuracy and makespan associated with input scaling adjustments. This hierarchical approach allows EMERALD to effectively balance computational efficiency and accuracy, significantly reducing missed deadlines—achieving 11x and 12.3x fewer missed deadlines compared to CAMDNN and HEFT, respectively, in scenarios demanding 30 frames per second. The results underscore the critical role of adaptive input scaling in managing the complexities of edge-based ML deployments. Soroush Heidari, Mehdi Ghasemi 0003, Younggeun Kim 0001, Carole-Jean Wu, Sarma B. K. Vrudhula |
SEC | 3 |
| 2024 | CLOVER: Carbon Optimization of Federated Learning over Heterogeneous ClientsabstractFederated Learning (FL) is a decentralized approach to train a DNN model without sharing the on-device training samples with a cloud server. Although FL is a practical solution to prevent the privacy leakage in DNN training, the environmental impact of FL can be significant given billions of mobile users. However, optimizing carbon emissions of FL is challenging because of its unique features such as heterogeneous carbon intensity, system/data heterogeneity, and network variability. In this paper, we propose a carbon-aware FL algorithm ---CLOVER--- which enables carbon efficient selections of participants and their respective training samples considering the aforementioned features. In our experiments with various combinations of DNN models and datasets, CLOVER improves the FL carbon efficiency by 25.0%, on average, while still guaranteeing the convergence with better accuracy. Chanwoo Cho, Yonglak Son, Younggeun Kim 0001 |
ISLPED | 4 |
| 2022 | CAMDNN: Content-Aware Mapping of a Network of Deep Neural Networks on Edge MPSoCsabstractMachine Learning (ML) workloads are increasingly deployed at the edge. Enabling efficient inference execution while considering model and system heterogeneity remains challenging, especially for ML tasks built with a network of DNNs. The challenge is to maximize the utilization of all available resources on the multiprocessor system on a chip (MPSoC) at the same time. This becomes even more complicated because the optimal mapping for the network of DNNs can vary with input batch sizes and scene complexity. In this paper, a holistic hierarchical scheduling framework is presented to optimize the execution time for a network of DNN models on an edge MPSoC at runtime, considering varying input characteristics. The framework consists of a local and a global scheduler. The local scheduler maps individual DNNs in the inference pipeline to the best-performing hardware unit while the global scheduler customizes an Integer Linear Programming (ILP) solution to instantiate DNN remapping. To minimize scheduler runtime overhead, an imitation learning (IL) based scheduler is used that approximates the ILP solutions. The proposed scheduling framework (CAMDNN) was implemented on a Qualcomm Robotic RB5 platform. CAMDNN resulted in lower execution time of up to 32% than HEFT, and by factors of 6.67X, 5.6X and 2.17X than the CPU-only, GPU-only and Central Queue schedulers. Soroush Heidari, Mehdi Ghasemi 0003, Younggeun Kim 0001, Carole-Jean Wu, Sarma B. K. Vrudhula |
IEEE Trans. Computers | 3 |
| 2021 | Chasing Carbon: The Elusive Environmental Footprint of ComputingabstractGiven recent algorithm, software, and hardware innovation, computing has enabled a plethora of new applications. As computing becomes increasingly ubiquitous, however, so does its environmental impact. This paper brings the issue to the attention of computer-systems researchers. Our analysis, built on industry-reported characterization, quantifies the environmental effects of computing in terms of carbon emissions. Broadly, carbon emissions have two sources: operational energy consumption, and hardware manufacturing and infrastructure. Although carbon emissions from the former are decreasing thanks to algorithmic, software, and hardware innovations that boost performance and power efficiency, the overall carbon footprint of computer systems continues to grow. This work quantifies the carbon output of computer systems to show that most emissions related to modern mobile and data-center equipment come from hardware manufacturing and infrastructure. We therefore outline future directions for minimizing the environmental impact of computing systems. Udit Gupta 0001, Younggeun Kim 0001, Sylvia Lee, Jordan Tse, Hsien-Hsin S. Lee, Gu-Yeon Wei, David Brooks 0001, Carole-Jean Wu |
HPCA | 2 |
| 2021 | AutoFL: Enabling Heterogeneity-Aware Energy Efficient Federated LearningabstractFederated learning enables a cluster of decentralized mobile devices at the edge to collaboratively train a shared machine learning model, while keeping all the raw training samples on device. This decentralized training approach is demonstrated as a practical solution to mitigate the risk of privacy leakage. However, enabling efficient FL deployment at the edge is challenging because of non-IID training data distribution, wide system heterogeneity and stochastic-varying runtime effects in the field. This paper jointly optimizes time-to-convergence and energy efficiency of state-of-the-art FL use cases by taking into account the stochastic nature of edge execution. We propose AutoFL by tailor-designing a reinforcement learning algorithm that learns and determines which K participant devices and per-device execution targets for each FL model aggregation round in the presence of stochastic runtime variance, system and data heterogeneity. By considering the unique characteristics of FL edge deployment judiciously, AutoFL achieves 3.6 times faster model convergence time and 4.7 and 5.2 times higher energy efficiency for local clients and globally over the cluster of K participants, respectively. Younggeun Kim 0001, Carole-Jean Wu |
MICRO | 1 |
| 2021 | Energy-Efficient Mapping for a Network of DNN Models at the EdgeabstractThis paper describes a novel framework for executing a network of trained deep neural network (DNN) models on commercial-off-the-shelf devices that are deployed in an IoT environment. The scenario consists of two devices connected by a wireless network: a user-end device (U), which is a low-end, energy and performance-limited processor, and a cloudlet (C), which is a substantially higher performance and energy-unconstrained processor. The goal is to distribute the computation of the DNN models between U and C to minimize the energy consumption of U while taking into account the variability in the wireless channel delay and the performance overhead of executing models in parallel. The proposed framework was implemented using an NVIDIA Jetson Nano for U and a Dell workstation with Titan Xp GPU as C. Experiments demonstrate significant improvements both in terms of energy consumption of U and processing delay. Mehdi Ghasemi 0003, Soroush Heidari, Younggeun Kim 0001, Aaron Lamb, Carole-Jean Wu, Sarma B. K. Vrudhula |
SMARTCOMP | 3 |
| 2021 | Thermal-aware adaptive VM allocation considering server locations in heterogeneous data centers
Younggeun Kim 0001, Seon Young Kim, Seung Hun Choi, Sung Woo Chung |
J. Syst. Archit. | 1 |
| 2020 | AutoScale: Energy Efficiency Optimization for Stochastic Edge Inference Using Reinforcement LearningabstractDeep learning inference is increasingly run at the edge. As the programming and system stack support becomes mature, it enables acceleration opportunities in a mobile system, where the system performance envelope is scaled up with a plethora of programmable co-processors. Thus, intelligent services designed for mobile users can choose between running inference on the CPU or any of the co-processors in the mobile system, and exploiting connected systems such as the cloud or a nearby, locally connected mobile system. By doing so, these services can scale out the performance and increase the energy efficiency of edge mobile systems. This gives rise to a new challenge—deciding when inference should run where. Such execution scaling decision becomes more complicated with the stochastic nature of mobile-cloud execution environment, where signal strength variation in the wireless networks and resource interference can affect real-time inference performance and system energy efficiency. To enable energy efficient deep learning inference at the edge, this paper proposes AutoScale, an adaptive and lightweight execution scaling engine built on the custom-designed reinforcement learning algorithm. It continuously learns and selects the most energy efficient inference execution target by considering characteristics of neural networks and available systems in the collaborative cloud-edge execution environment while adapting to stochastic runtime variance. Real system implementation and evaluation, considering realistic execution scenarios, demonstrate an average of 9.8x and 1.6x energy efficiency improvement over the baseline mobile CPU and cloud offloading, respectively, while meeting the real-time performance and accuracy requirements. Younggeun Kim 0001, Carole-Jean Wu |
MICRO | 1 |
| 2020 | An Adaptive Thermal Management Framework for Heterogeneous Multi-Core ProcessorsabstractOff-the-shelf embedded systems have adopted heterogeneous multi-core processors which have high-performance big cores and low-power small cores. Though there are two different types of cores in heterogeneous multi-core processors, conventional DVFS (Dynamic Voltage and Frequency Scaling)-based DTM (Dynamic Thermal Management) techniques do not utilize the different types of cores to cool down hot cores. Rather, they primarily reduce the voltage and frequency of the hot cores, leading to performance degradation. In this article, we propose a novel adaptive DTM framework for heterogeneous multi-core processors, which utilizes the big and small cores to prevent performance degradation. Our proposed framework exploits two migration-based DTM techniques: 1) a technique (denoted as Migrationbig↔big) that migrates applications from hot big cores (big cores whose temperature is above a pre-defined threshold) to cold big cores (big cores whose temperature is below the threshold) and 2) a technique (denoted as Migrationbig↔small) that migrates all applications from the big cores to the small cores. In case of thermal emergency of the big cores, our proposed framework checks the number of cold big cores. When there exist available cold big cores, our proposed framework employs Migrationbig↔big to cool down the hot big cores while not reducing the big core frequency. On the other hand, when there does not exist any available cold big core, our proposed framework employs one between Migrationbig↔small and a DVFS-based DTM technique, which is expected to result in better performance. In our experiments on an embedded development board, our proposed framework improves the average performance by 8.9 percent, compared to ARM's DVFS-based IPA (Intelligent Power Allocation), satisfying thermal constraints. Our framework also improves the average performance by 10.4 percent, compared to a state-of-the-art predictive DVFS-based DTM technique. Younggeun Kim 0001, Minyong Kim, Joonho Kong, Sung Woo Chung |
IEEE Trans. Computers | 1 |
| 2020 | Signal Strength-Aware Adaptive Offloading with Local Image Preprocessing for Energy Efficient Mobile DevicesabstractTo prolong battery life of mobile devices, image processing applications often exploit offloading techniques which run some or all of the computations on remote servers. Unfortunately, the existing offloading techniques do not consider the fact that data transmission time and energy consumption of wireless network interfaces exponentially increase when signal strength decreases. In this paper, we propose an adaptive offloading for image processing applications, which considers wireless signal strength. To improve performance and energy efficiency of offloading, we also propose to adaptively exploit local preprocessing (executing image preprocessing on local mobile devices), considering wireless signal strength; the local preprocessing usually reduces the size of transmission image in offloading. Our proposed technique estimates performance and energy consumption of the following three methods, depending on the wireless signal strength: 1) local execution (executing all the computations on the local mobile devices), 2) offloading without local preprocessing, and 3) offloading with local preprocessing. Based on the estimated performance and energy consumption, our technique employs one among the three methods, which is expected to result in the best performance or energy efficiency. In our evaluation on an off-the-shelf smartphone, when a user prefers performance to energy, our proposed technique improves performance by 27.1 percent, compared to the conventional offloading technique that does not consider the signal strength. On the other hand, when a user prefers energy to performance, our proposed technique saves system-wide (not just CPU nor wireless network interface) energy consumption by 26.3 percent, on average, compared to the conventional offloading technique. Younggeun Kim 0001, Young Seo Lee, Sung Woo Chung |
IEEE Trans. Computers | 1 |
| 2019 | A Framework for Distributed Deep Neural Network Training with Heterogeneous Computing PlatformsabstractDeep neural network (DNN) training is generally performed by cloud computing platforms. However, cloud-based training has several problems such as network bottleneck, server management cost, and privacy. To overcome these problems, one of the most promising solutions is distributed DNN model training which trains the model with not only high-performance servers but also low-end power-efficient mobile edge or user devices. However, due to the lack of a framework which can provide an optimal cluster configuration (i.e., determining which computing devices participate in DNN training tasks), it is difficult to perform efficient DNN model training considering DNN service providers' preferences such as training time or energy efficiency. In this paper, we introduce a novel framework for distributed DNN training that determines the best training cluster configuration with available heterogeneous computing resources. Our proposed framework utilizes pre-training with a small number of training steps and estimates training time, power, energy, and energy-delay product (EDP) for each possible training cluster configuration. Based on the estimated metrics, our framework performs DNN training for the remaining steps with the chosen best cluster configurations depending on DNN service providers' preferences. Our framework is implemented in TensorFlow and evaluated with three heterogeneous computing platforms and five widely used DNN models. According to our experimental results, in 76.67% of the cases, our framework chooses the best cluster configuration depending on DNN service providers' preferences with only a small training time overhead. Bontak Gu, Joonho Kong, Arslan Munir, Younggeun Kim 0001 |
ICPADS | 4 |
| 2019 | Temperature-aware Adaptive VM Allocation in Heterogeneous Data CentersabstractVirtualized data centers usually consist of heterogeneous servers which have different specifications (performance). Though there are usually a number of unused servers with different performance in such heterogeneous data centers, conventional DVFS (Dynamic Voltage and Frequency Scaling)-based DTM (Dynamic Thermal Management) techniques do not exploit the unused servers to cool down hot servers. In this paper, we propose a novel DTM technique which adaptively exploits external computing resources (unused servers with different performance) as well as internal computing resources (unused CPU cores in the server) available in heterogeneous data centers. When the temperature of a CPU core in a server exceeds a pre-defined thermal threshold, our proposed technique first identifies memory intensiveness and usage of VMs (Virtual Machines). Depending on the memory intensiveness and usage of VMs, our technique adaptively employs the following three methods: 1) a method that migrates a VM to another server with different performance, 2) a method that migrates VMs among CPU cores in the server, and 3) a DVFS-based method. In our experiments, our proposed technique improves performance by 9.6% and saves system-wide EDP by 12.9%, on average (by up to 17.1% and 24.5%, respectively), compared to a conventional DVFS-based DTM technique, satisfying thermal constraints. Younggeun Kim 0001, Jeong In Kim, Seung Hun Choi, Seon Young Kim, Sung Woo Chung |
ISLPED | 1 |
| 2018 | A Survey on Recent OS-Level Energy Management Techniques for Mobile Processing UnitsabstractTo improve mobile experience of users, recent mobile devices have adopted powerful processing units (CPUs and GPUs). Unfortunately, the processing units often consume a considerable amount of energy, which in turn shortens battery life of mobile devices. For energy reduction of the processing units, mobile devices adopt energy management techniques based on software, especially OS (Operating Systems), as well as hardware. In this survey paper, we summarize recent OS-level energy management techniques for mobile processing units. We categorize the energy management techniques into three parts, according to main operations of the summarized techniques: 1) techniques adjusting power states of processing units, 2) techniques exploiting other computing resources, and 3) techniques considering interactions between displays and processing units. We believe this comprehensive survey paper will be a useful guideline for understanding recent OS-level energy management techniques and developing more advanced OS-level techniques for energy-efficient mobile processing units. Younggeun Kim 0001, Joonho Kong, Sung Woo Chung |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2017 | Signal strength-aware adaptive offloading for energy efficient mobile devicesabstractTo prolong battery life of mobile devices, applications often exploit offloading techniques which run computations on remote servers. Unfortunately, the existing offloading techniques do not consider the fact that data transmission time and energy consumption of wireless network interfaces exponentially increase when signal strength decreases. In this paper, we propose an adaptive offloading technique that considers signal strength. Our technique estimates gain (reduced computation time and energy of mobile devices) and loss (increased data transmission time and energy of network interfaces) of offloading depending on signal strength. Based on the estimated gain and loss, our technique determines whether it offloads computations to a server or not. In evaluation, our proposed technique improves performance by 30.1% and saves system-wide energy consumption by 25.0%, on average, compared to the conventional offloading technique that does not consider signal strength. Younggeun Kim 0001, Sung Woo Chung |
ISLPED | 1 |
| 2017 | Enhancing Energy Efficiency of Multimedia Applications in Heterogeneous Mobile Multi-Core ProcessorsabstractRecent smart devices have adopted heterogeneous multi-core processors which have high-performance big cores and low-power small cores. Unfortunately, the conventional task scheduler for heterogeneous multi-core processors does not provide appropriate amount of CPU resources for multimedia applications (whose QoS is important to users), resulting in energy waste; it often executes multimedia applications and non-multimedia applications on the same core. In this paper, we propose an advanced task scheduler for heterogeneous multi-core processors, which provides appropriate amount of CPU resources for multimedia applications. Our proposed task scheduler isolates multimedia applications from non-multimedia applications at runtime, exploiting the fact that multimedia applications have a specific thread for video/audio playback (to play video/audio, a multimedia application should use a function that generates the specific thread). Since multimedia applications usually require a smaller amount of CPU resources than non-multimedia applications due to dedicated hardware decoders, our proposed task scheduler allocates the former to the small cores and the latter to the big cores. In our experiments on an Android-based development board, our proposed task scheduler saves system-wide (not just CPU) energy consumption by 8.9 percent, on average, compared to the conventional task scheduler, preserving QoS of multimedia applications. In addition, it improves performance of non-multimedia applications by 13.7 percent, on average, compared to the conventional task scheduler. Younggeun Kim 0001, Minyong Kim, Sung Woo Chung |
IEEE Trans. Computers | 1 |
| 2015 | M-DTM: migration-based dynamic thermal management for heterogeneous mobile multi-core processors
Younggeun Kim 0001, Minyong Kim, Jae Min Kim, Sung Woo Chung |
DATE | 1 |
| 2015 | Stabilizing CPU Frequency and Voltage for Temperature-Aware DVFS in Mobile DevicesabstractRecent mobile devices adopt high-performance processors to support various functions. As a side effect, higher performance inevitably leads to power density increase, eventually resulting in thermal problems. In order to alleviate the thermal problems, off-the-shelf mobile devices rely on dynamic voltage-frequency scaling (DVFS)-based dynamic thermal management (DTM) schemes. Unfortunately, in the DVFS-based DTM schemes, an excessive number of DTM operations worsen not only performance but also power efficiency. In this paper, we propose a temperature-aware DVFS scheme for Android-based mobile devices to optimize power or performance depending on the option. We evaluate our scheme in the off-the-shelf mobile device. Our evaluation results show that our scheme saves energy consumption by 12.7%, on average, when we use the power optimizing option. Our scheme also enhances the performance by 6.3%, on average, by using the performance optimizing scheme, still reducing the energy consumption by 6.7%. Jae Min Kim, Younggeun Kim 0001, Sung Woo Chung |
IEEE Trans. Computers | 2 |