Sunwoo Lee 0001

dblp:56/7811-1 · DBLP profile ↗
← Back
24ranked-venue papers
13as first author
16since 2021 · last 2026
0000-0001-6334-3068ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 7 first-author · 8 since 2021Systems, architecture and hardware · 8 · 4 first-author · 4 since 2021Databases, data management, data science and information retrieval · 6 · 4 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 first-author · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 GEM: A Scale-Aware and Distribution-Sensitive Sparse Fine-Tuning Framework for Effective Downstream Adaptation
abstract
Parameter-efficient fine-tuning (PEFT) has become a popular way to adapt large pre-trained models to new tasks. Most PEFT methods update only a small subset of parameters while freezing the rest, avoiding redundant computation. As they maximize the absolute size of the updates without regard to the parameters’ original scale, the resulting changes in model behavior can be minimal. In contrast, we maximize updates relative to each parameter’s scale, yielding more meaningful downstream adaptation. We propose Gradient-to-Weight Ratio and Entropy-guided Masking (GEM), a parameter scale-aware, distribution-sensitive sparse fine-tuning framework. GEM prioritizes parameters whose updates are significant in proportion to their initial pre-trained values. It also adaptively determines how many parameters to tune at each layer based on the entropy of parameter values, thereby making the most effective use of the computational budget in PEFT. Our empirical study demonstrates the efficacy of GEM on both general-domain tasks (GLUE and SuperGLUE) and domain-specific tasks (GSM8k and MBPP), achieving up to a 1.6% improvement in fine-tuning accuracy over full fine-tuning while updating only 0.1% of model parameters.
Sungmin Kang, Jisoo Kim 0007, Amir Salman Avestimehr, Sunwoo Lee 0001
AAAI4
2026 Multi-Metric Client Activation Method for Fast and Accurate Federated Learning
abstract
While unbiased gradient estimators ensure unbiased solutions in empirical risk minimization problems, they can significantly hinder optimization efficiency and generalization performance in strongly non-IID Federated Learning environments. Although recent studies have demonstrated promising applications of biased estimators, they typically focus only on convergence rates while overlooking generalization performance. We propose a novel multi-metric bias concept, quantified using both local loss and local gradient norm, along with a client activation method based on this bias concept. The proposed method prioritizes training on local datasets that better represent the global dataset, leading to faster convergence and improved generalization. Our extensive empirical study demonstrates that carefully injecting bias into client activation accelerates federated optimization, achieving a substantially improved validation accuracy within a given epoch budget. In representative machine learning benchmarks, our method achieves up to \(12.7\%\) higher accuracy than uniform random sampling and \(2.5\%\) higher accuracy than state-of-the-art biased client activation methods.
Jihyun Lim, Sunwoo Lee 0001
ACM Trans. Intell. Syst. Technol.3
2025 Enabling Weak Client Participation via On-Device Knowledge Distillation in Heterogeneous Federated Learning
abstract
Online Knowledge Distillation (KD) is recently highlighted to train large models in Federated Learning (FL) environments. Many existing studies adopt the logit ensemble method to perform KD on the server side. However, they often assume that unlabeled data collected at the edge is centralized on the server. Moreover, the logit ensemble method personalizes local models, which can degrade the quality of soft targets, especially when data is highly non-IID. To address these critical limitations, we propose a novel on-device KD-based heterogeneous FL method. Our approach leverages a small auxiliary model to learn from labeled local data. Subsequently, a subset of clients with strong system resources transfers knowledge to a large model through on-device KD using their unlabeled data. Our extensive experiments demonstrate that our on-device KD-based heterogeneous FL method effectively utilizes the system resources of all edge devices as well as the unlabeled data, resulting in higher accuracy compared to SOTA KD-based FL methods.
Jihyun Lim, Junhyuk Jo, Sunwoo Lee 0001
ECAI4
2025 Layer-wise Update Aggregation with Recycling for Communication-Efficient Federated Learning
abstract
Expensive communication cost is a common performance bottleneck in Federated Learning (FL), which makes it less appealing in real-world applications. Many communication-efficient FL methods focus on discarding a part of model updates mostly based on gradient magnitude. In this study, we find that recycling previous updates, rather than simply dropping them, more effectively reduces the communication cost while maintaining FL performance. We propose *FedLUAR*, a Layer-wise Update Aggregation with Recycling scheme for communication-efficient FL. We first define a useful metric that quantifies the extent to which the aggregated gradients influences the model parameter values in each layer. *FedLUAR* selects a few layers based on the metric and recycles their previous updates on the server side. Our extensive empirical study demonstrates that the update recycling scheme significantly reduces the communication cost while maintaining model accuracy. For example, our method achieves nearly the same AG News accuracy as FedAvg, while reducing the communication cost to just 17%.
Jisoo Kim 0007, Sungmin Kang, Sunwoo Lee 0001
NeurIPS3
2024 Layer-Wise Adaptive Gradient Norm Penalizing Method for Efficient and Accurate Deep Learning
abstract
Sharpness-aware minimization (SAM) is known to improve the generalization performance of neural networks.However, it is not widely used in real-world applications yet due to its expensive model perturbation cost.A few variants of SAM have been proposed to tackle such an issue, but they commonly do not alleviate the cost noticeably.In this paper, we propose a lightweight layerwise gradient norm penalizing method that tackles the expensive computational cost of SAM while maintaining its superior generalization performance.Our study empirically proves that the gradient norm of the whole model can be effectively suppressed by penalizing the gradient norm of only a few critical layers.We also theoretically show that such a partial model perturbation does not harm the convergence rate of SAM, allowing them to be safely adapted in real-world applications.To demonstrate the efficacy of the proposed method, we perform extensive experiments comparing the proposed method to mini-batch SGD and the conventional SAM using representative computer vision and language modeling benchmarks.
Sunwoo Lee 0001
KDD1
2024 Embracing Federated Learning: Enabling Weak Client Participation via Partial Model Training
abstract
In Federated Learning (FL), clients may have weak devices that cannot train the full model or even hold it in their memory space. To implement large-scale FL applications, thus, it is crucial to develop a distributed learning method that enables the participation of such weak clients. We proposeEmbracingFL, a general FL framework that allows all available clients to join the distributed training regardless of their system resource capacity. The framework is built upon a novel form of partial model training method in which each client trains as many consecutive output-side layers as its system resources allow. Our study demonstrates thatEmbracingFLencourages each layer to have similar data representations across clients, improving FL efficiency. The proposed partial model training method guarantees convergence to a neighbor of stationary points for non-convex and smooth problems. We evaluate the efficacy ofEmbracingFLunder a variety of settings with a mixed number of strong, moderate ($\sim 40\%$memory), and weak ($\sim 15\%$memory) clients, datasets (CIFAR-10, FEMNIST, and IMDB), and models (ResNet20, CNN, and LSTM). Our empirical study shows thatEmbracingFLconsistently achieves high accuracy as like all clients are strong, outperforming the state-of-the-art width reduction methods (i.e. HeteroFL and FjORD).
Sunwoo Lee 0001, Saurav Prakash, Yue Niu 0001, Amir Salman Avestimehr
IEEE Trans. Mob. Comput.1
2023 Layer-Wise Adaptive Model Aggregation for Scalable Federated Learning
abstract
In Federated Learning (FL), a common approach for aggregating local solutions across clients is periodic full model averaging. It is, however, known that different layers of neural networks can have a different degree of model discrepancy across the clients. The conventional full aggregation scheme does not consider such a difference and synchronizes the whole model parameters at once, resulting in inefficient network bandwidth consumption. Aggregating the parameters that are similar across the clients does not make meaningful training progress while increasing the communication cost. We propose FedLAMA, a layer-wise adaptive model aggregation scheme for scalable FL. FedLAMA adjusts the aggregation interval in a layer-wise manner, jointly considering the model discrepancy and the communication cost. This fine-grained aggregation strategy enables to reduce the communication cost without significantly harming the model accuracy. Our extensive empirical study shows that, as the aggregation interval increases, FedLAMA shows a remarkably smaller accuracy drop than the periodic full aggregation, while achieving comparable communication efficiency.
Sunwoo Lee 0001, Amir Salman Avestimehr
AAAI1
2023 FedAudio: A Federated Learning Benchmark for Audio Tasks
abstract
Federated learning (FL) has gained substantial attention in recent years due to data privacy concerns related to the pervasiveness of consumer devices that continuously collect data from users. While a number of FL benchmarks have been developed to facilitate FL research, none of them include audio data and audio-related tasks. In this paper, we fill this critical gap by introducing a new FL benchmark for audio tasks which we refer to as FedAudio. FedAudio includes four representative and commonly used audio datasets from three important audio tasks that are well aligned with FL use cases. In particular, a unique contribution of FedAudio is the introduction of data noises and label errors to the datasets to emulate challenges when deploying FL systems in real-world settings. FedAudio also includes the benchmark results of the datasets and a PyTorch library with the objective of facilitating researchers to fairly compare their algorithms. We hope FedAudio could act as a catalyst to inspire new FL research for audio tasks and thus benefit the acoustic and speech research community. The datasets and benchmark results can be accessed at https://github.com/zhang-tuo-pdf/FedAudio.
Tiantian Feng, Samiul Alam, Sunwoo Lee 0001, Mi Zhang 0002, Shri Narayanan, Amir Salman Avestimehr
ICASSP4
2023 Partial model averaging in Federated Learning: Performance guarantees and benefits
Sunwoo Lee 0001, Anit Kumar Sahu, Chaoyang He 0001, Amir Salman Avestimehr
Neurocomputing1
2023 Achieving small-batch accuracy with large-batch scalability via Hessian-aware learning rate adjustment
Sunwoo Lee 0001, Chaoyang He 0001, Amir Salman Avestimehr
Neural Networks1
2022 Using Multi-Resolution Data to Accelerate Neural Network Training in Scientific Applications
abstract
Neural networks are powerful solutions to many scientific applications; however, they usually require long model training time due to large training data sets or large model size. Research has been focused on developing numerical optimization algorithms and parallel processing to reduce the training time. In this work, we propose a multi-resolution strategy that can reduce the training time by training the model with the reduced-resolution data samples at the beginning and later switching to the original resolution data samples. This strategy is motivated by the observation that coarser versions of many applications can be solved faster than their denser counterparts, and the solution to a coarser problem could be used to initialize the solution to the denser problem. When applying the idea to neural network training, coarse data can have a similar effect on the learning curves at the early stage as the dense data but requires less time. Once the curves no longer improve significantly, our strategy switches to using the data in original resolution. The key in this process is the ability to generate multiple resolutions of a problem automatically, which could usually be done with scientific applications with spatial and temporal continuity. We use two real-world scientific applications, CosmoFlow and DeepCAM, to evaluate the proposed mixed-resolution training strategy. Our experiment results demonstrate that the proposed training strategy effectively reduces the end-to-end training time while achieving a comparable accuracy to that of the training only with the original data. While maintaining the same model accuracy, our multi-resolution training strategy reduces the end-to-end training time up to 30% and 23% for CosmoFlow and DeepCAM, respectively.
Kewei Wang 0002, Sunwoo Lee 0001, Jan Balewski, Alex Sim, Peter Nugent, Ankit Agrawal 0001, Alok N. Choudhary, Kesheng Wu, Wei-keng Liao
CCGRID2
2022 Improving scalability of parallel CNN training by adaptively adjusting parameter update frequency
Sunwoo Lee 0001, Qiao Kang, Reda Al-Bahrani, Ankit Agrawal 0001, Alok N. Choudhary, Wei-keng Liao
J. Parallel Distributed Comput.1
2022 A case study on parallel HDF5 dataset concatenation for high energy physics data analysis
Sunwoo Lee 0001, Kaiyuan Hou, Kewei Wang 0002, Saba Sehrish, Marc F. Paterno, Jim Kowalkowski, Quincey Koziol, Robert B. Ross, Ankit Agrawal 0001, Alok N. Choudhary, Wei-keng Liao
Parallel Comput.1
2021 Supporting Data Compression in PnetCDF
abstract
Recently, the dramatic increase of the data amounts drives up the demand for data compression among HPC applications. Although many file systems and I/O middlewares have incorporated compression features, few high-level parallel I/O libraries support data compression due to the challenges of achieving scalable performance on HPC systems. This paper presents the design and implementation of the variable compression feature in the Parallel NetCDF library. Our design employs the same concept of chunking used by the HDF5 library, but we focus on enabling I/O aggregation across multiple requests to address the challenges on performance and scalability. We evaluate our solution using the I/O kernel of real-world scientific applications and analyze the impacts of data compression on parallel I/O performance. Our result suggests that handling multiple requests at once can significantly improve the parallel I/O performance on chunked and compressed data.
Kaiyuan Hou, Qiao Kang, Sunwoo Lee 0001, Ankit Agrawal 0001, Alok N. Choudhary, Wei-keng Liao
IEEE BigData3
2021 Asynchronous I/O Strategy for Large-Scale Deep Learning Applications
abstract
Many scientific applications have started using deep learning methods for their classification or regression problems. However, for data-intensive scientific applications, I/O performance can be the major performance bottleneck. In order to effectively solve important real-world problems using deep learning methods on High-Performance Computing (HPC) systems, it is essential to address the poor I/O performance issue in large-scale neural network training. In this paper, we propose an asynchronous I/O strategy that can be generally applied to deep learning applications. Our I/O strategy employs an I/O -dedicated thread per process, that performs I/O operations independently of the training progress. The I/O thread reads many training samples at once to reduce the total number of I/O operations per epoch. Given the fixed amount of training data, the fewer the I/O operations per epoch, the shorter the overall I/O time. The I/O operations are also overlapped with the computations using the double-buffering method. We evaluate our I/O strategy using two real-world scientific applications, CosmoFlow and Neuron-Inverter. Our experimental results demonstrate that the proposed I/O strategy significantly improves the scaling performance without affecting the regression performance.
Sunwoo Lee 0001, Qiao Kang, Kewei Wang 0002, Jan Balewski, Alex Sim, Ankit Agrawal 0001, Alok N. Choudhary, Peter Nugent, Kesheng Wu, Wei-keng Liao
HiPC1
2021 SIGRNN: Synthetic Minority Instances Generation in Imbalanced Datasets using a Recurrent Neural Network
Reda Al-Bahrani, Dipendra Jha, Qiao Kang, Sunwoo Lee 0001, Zijiang Yang 0008, Wei-keng Liao, Ankit Agrawal 0001, Alok N. Choudhary
ICPRAM4
2020 Communication-Efficient Local Stochastic Gradient Descent for Scalable Deep Learning
abstract
Synchronous Stochastic Gradient Descent (SGD) with data parallelism, the most popular parallel training strategy for deep learning, suffers from expensive gradient communications. Local SGD with periodic model averaging is a promising alternative to synchronous SGD. The algorithm allows each worker to locally update its own model, and periodically averages the model parameters across all the workers. While this algorithm enjoys less frequent communications, the convergence rate is strongly affected by the number of workers. In order to scale up the local SGD training without losing accuracy, the number of workers should be sufficiently small so that the model converges reasonably fast. In this paper, we discuss how to exploit the degree of parallelism in local SGD while maintaining model accuracy. Our training strategy employs multiple groups of processes and each group trains a local model based on data parallelism. The local models are periodically averaged across all the groups. Based on this hierarchical parallelism, we design a model averaging algorithm that has a cheaper communication cost than allreduce-based approach. We also propose a practical metric for finding the maximum number of workers that does not cause a significant accuracy loss. Our experimental results demonstrate that our proposed training strategy provides a significantly improved scalability while achieving a comparable model accuracy to synchronous SGD.
Sunwoo Lee 0001, Qiao Kang, Ankit Agrawal 0001, Alok N. Choudhary, Wei-keng Liao
IEEE BigData1
2020 Predicting Resource Requirement in Intermediate Palomar Transient Factory Workflow
abstract
Quickly identifying astronomical transients from synoptic surveys is critical to many recent astrophysical discoveries. However, each of the data processing pipelines in these surveys contains dozens of stages with highly varying time and space requirements. Properly predicting the resources required to run these pipelines is critical for the allocation of computing resources and reducing the discovery response time. We propose a machine learning strategy for this prediction task and demonstrate its effectiveness using a set of timing measurements from the intermediate Palomar Transient Factory (iPTF) workflow. The proposed model utilizes the spatiotemporal correlation of astronomical images, where nearby patches of the sky (space) are likely to have a similar number of objects of interest and workflows executed in the recent past (time) are likely to use a similar amount of time because the machines and data storage systems are likely to be in similar states. We capture the relationship among these spatial and temporal features in a Bayesian network and study how they impact the prediction accuracy. This Bayesian network helps us to identify the most influential features for predictions. With proper features, our models achieve errors close to the random variance boundary within batches of images taken at the same time, which can be regarded as the intrinsic limit of prediction accuracy.
Qiao Kang, Alex Sim, Peter Nugent, Sunwoo Lee 0001, Wei-keng Liao, Ankit Agrawal 0001, Alok N. Choudhary, Kesheng Wu
CCGRID4
2020 Improving all-to-many personalized communication in two-phase I/O
abstract
As modern parallel computers enter the exascale era, the communication cost for redistributing requests becomes a significant bottleneck in MPIIO routines. The communication kernel for request redistribution, which has an all-to-many personalized communication pattern for application programs with a large number of noncontiguous requests, plays an essential role in the overall performance. This paper explores the available communication kernels for two-phase I/O communication. We generalize the spread-out algorithm to adapt to the all-to-many communication pattern of two-phase I/O by reducing the communication straggler effect. Communication throttling methods that reduce communication contention for asynchronous MPI implementation are adopted to improve communication performance further. Experimental results are presented using different communication kernels running on Cray XC40 Cori and IBM AC922 Summit supercomputers with different I/O patterns. Our study shows that adjusting communication kernel algorithms for different I/O patterns can improve the end-to-end performance up to 10 times compared with default MPI-IO implementations.
Qiao Kang, Robert B. Ross, Robert Latham, Sunwoo Lee 0001, Ankit Agrawal 0001, Alok N. Choudhary, Wei-keng Liao
SC4
2020 Improving MPI Collective I/O for High Volume Non-Contiguous Requests With Intra-Node Aggregation
abstract
Two-phase I/O is a well-known strategy for implementing collective MPI-IO functions. It redistributes I/O requests among the calling processes into a form that minimizes the file access costs. As modern parallel computers continue to grow into the exascale era, the communication cost of such request redistribution can quickly overwhelm collective I/O performance. This effect has been observed from parallel jobs that run on multiple compute nodes with a high count of MPI processes on each node. To reduce the communication cost, we present a new design for collective I/O by adding an extra communication layer that performs request aggregation among processes within the same compute nodes. This approach can significantly reduce inter-node communication contention when redistributing the I/O requests. We evaluate the performance and compare it with the original two-phase I/O on Cray XC40 parallel computers (Theta and Cori) with Intel KNL and Haswell processors. Using I/O patterns from two large-scale production applications and an I/O benchmark, we show our proposed method effectively reduces the communication cost and hence maintains the scalability for a large number of processes.
Qiao Kang, Sunwoo Lee 0001, Kaiyuan Hou, Robert B. Ross, Ankit Agrawal 0001, Alok N. Choudhary, Wei-keng Liao
IEEE Trans. Parallel Distributed Syst.2
2019 Improving Scalability of Parallel CNN Training by Adjusting Mini-Batch Size at Run-Time
abstract
Training Convolutional Neural Network (CNN) is a computationally intensive task, requiring efficient parallelization to shorten the execution time. Considering the ever-increasing size of available training data, the parallelization of CNN training becomes more important. Data-parallelism, a popular parallelization strategy that distributes the input data among compute processes, requires the mini-batch size to be sufficiently large to achieve a high degree of parallelism. However, training with large batch size is known to produce a low convergence accuracy. In image restoration problems, for example, the batch size is typically tuned to a small value between 16 ~ 64, making it challenging to scale up the training. In this paper, we propose a parallel CNN training strategy that gradually increases the mini-batch size and learning rate at run-time. While improving the scalability, this strategy also maintains the accuracy close to that of the training with a fixed small batch size. We evaluate the performance of the proposed parallel CNN training algorithm with image regression and classification applications using various models and datasets.
Sunwoo Lee 0001, Qiao Kang, Sandeep Madireddy, Prasanna Balaprakash, Ankit Agrawal 0001, Alok N. Choudhary, Rick Archibald, Wei-keng Liao
IEEE BigData1
2017 Parallel Deep Convolutional Neural Network Training by Exploiting the Overlapping of Computation and Communication
abstract
Training Convolutional Neural Network (CNN) is a computationally intensive task whose parallelization has become critical in order to complete the training in an acceptable time. However, there are two obstacles to developing a scalable parallel CNN in a distributed-memory computing environment. One is the high degree of data dependency exhibited in the model parameters across every two adjacent minibatches and the other is the large amount of data to be transferred across the communication channel. In this paper, we present a parallelization strategy that maximizes the overlap of inter-process communication with the computation. The overlapping is achieved by using a thread per compute node to initiate communication after the gradients are available. The output data of backpropagation stage is generated at each model layer, and the communication for the data can run concurrently with the computation of other layers. To study the effectiveness of the overlapping and its impact on the scalability, we evaluated various model architectures and hyperparameter settings. When training VGG-A model using ImageNet data sets, we achieve speedups of 62.97× and 77.97× on 128 compute nodes using mini-batch sizes of 256 and 512, respectively.
Sunwoo Lee 0001, Dipendra Jha, Ankit Agrawal 0001, Alok N. Choudhary, Wei-keng Liao
HiPC1
2016 Evaluation of K-means data clustering algorithm on Intel Xeon Phi
abstract
Intel Xeon Phi is a processor based on MIC architecture that contains a large number of compute cores with a high local memory bandwidth and 512-bit vector processing units. To achieve high performance on Xeon Phi, it is important for programmers to explore all the software features provided by the Intel compiler and libraries to fully utilize the new hardware resources. In this paper, we use the K-Means algorithm to study the performance of various Intel software settings available for Xeon Phi and their impacts to the performance of K-means. At first we examine different memory layouts for storing data points using Intel compiler-intrinsic functions. During distance calculation, the computational kernel of K-means, when the size of individual input data points is not vector-friendly, we pad the data points to align with the VPU width. At last, we implement a parallel reduction to increase memory access parallelism and cache hits. These techniques enable us to successfully take advantage of thread-level parallelism and data-level parallelism on Xeon Phi. Experimental results demonstrate large performance gains over the default auto-vectorization approach. The K-Means implemented with the proposed techniques achieves up to 68.65% and 56.14% performance improvements for aligned datasets and unaligned datasets, respectively. For high-dimensional aligned datasets, we achieved up to 53.49% performance improvement on a large-scale parallel computer.
Sunwoo Lee 0001, Wei-keng Liao, Ankit Agrawal 0001, Nikos Hardavellas, Alok N. Choudhary
IEEE BigData1
2009 Extending Component-Based Approaches for Multithreaded Design of Multiprocessor Embedded Software
abstract
Multiprocessor embedded software presents major challenges including the increased complexity and stringent performance requirements raised by parallel processing capability. Although component-based approaches can greatly alleviate the complexity problem, traditional approaches do not provide adequate support for performance requirements on multiprocessors. In this paper, we extend component-based approaches for performance-aware design and analysis of multiprocessor embedded software. The proposed approach begins with a component model that is able to capture concurrent and performance-critical behavior based on UML activity diagrams. It then identifies and exploits potential sources of concurrency within components to produce a multithreaded software design targeted at multiprocessors. We also provide a schedulability analysis that can be used for validating multithreaded software design from a temporal perspective.
Sunwoo Lee 0001, Byung Kwan Jung, Minsoo Ryu
ISORC1