Sangwoo Hong

dblp:48/334 · DBLP profile ↗
← Back
15ranked-venue papers
7as first author
14since 2021 · last 2026
0000-0002-0270-2781ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 4 since 2021Computer networks · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Trustworthy machine learning · 44% Efficient and distributed learning · 24% Language models and text generation · 11%
Computer architecture, parallel and distributed computing, and storage systems
4 papers
Distributed systems · 98% High-performance computing · 2%
Theoretical computer science
3 papers
Coding theory · 78% Algorithms and data structures · 22%

Topics — the 30 heaviest of 34, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
fairness
2.632026
Bias Alleviation Through Network Pruning for Sparse and Debiased Models · IEEE Trans. Image Process. 2026
Constructing Fair Latent Space for Intersection of Fairness and Explainability · AAAI 2025
Mitigating Spurious Correlations via Disagreement Probability · NeurIPS 2024
Distributed systems › distributed data processing
straggler mitigation
1.932024
Group-Wise Verifiable Coded Computing Under Byzantine Attacks and Stragglers · IEEE Trans. Inf. Forensics Secur. 2024
Straggler-Exploiting Fully Private Distributed Matrix Multiplication With Chebyshev Polynomials · IEEE Trans. Commun. 2023
Chebyshev Polynomial Codes: Task Entanglement-based Coding for Distributed Matrix Multiplication · ICML 2021
Distributed systems
coded computation
1.422024
Group-Wise Verifiable Coded Computing Under Byzantine Attacks and Stragglers · IEEE Trans. Inf. Forensics Secur. 2024
Straggler-Exploiting Fully Private Distributed Matrix Multiplication With Chebyshev Polynomials · IEEE Trans. Commun. 2023
Distributed systems
fault tolerance
1.322024
Group-Wise Verifiable Coded Computing Under Byzantine Attacks and Stragglers · IEEE Trans. Inf. Forensics Secur. 2024
Hierarchical Group Testing for Byzantine Attack Identification in Distributed Matrix Multiplication · IEEE J. Sel. Areas Commun. 2022
Natural language and speech › Language models and text generation
chain-of-thought reasoning
1.012026
Efficient Process Reward Modeling via Contrastive Mutual Information · ACL (1) 2026
Machine learning › Efficient and distributed learning
dataset distillation
1.012026
An Adaptive Sampling Framework for Diffusion-based Dataset Distillation with High Fidelity and Diversity · AAAI 2026
Machine learning › Trustworthy machine learning
debiasing
1.012026
Bias Alleviation Through Network Pruning for Sparse and Debiased Models · IEEE Trans. Image Process. 2026
Machine learning › Efficient and distributed learning › dataset distillation
diffusion-based dataset distillation
1.012026
An Adaptive Sampling Framework for Diffusion-based Dataset Distillation with High Fidelity and Diversity · AAAI 2026
Machine learning › Generative modeling
diffusion model
1.012026
An Adaptive Sampling Framework for Diffusion-based Dataset Distillation with High Fidelity and Diversity · AAAI 2026
Machine learning › Efficient and distributed learning
model compression
1.012026
Bias Alleviation Through Network Pruning for Sparse and Debiased Models · IEEE Trans. Image Process. 2026
Machine learning › Reinforcement learning › reinforcement learning from human feedback
process reward model
1.012026
Efficient Process Reward Modeling via Contrastive Mutual Information · ACL (1) 2026
Machine learning › Efficient and distributed learning › model compression
pruning
1.012026
Bias Alleviation Through Network Pruning for Sparse and Debiased Models · IEEE Trans. Image Process. 2026
Machine learning › Probabilistic and Bayesian machine learning
sampling
1.012026
An Adaptive Sampling Framework for Diffusion-based Dataset Distillation with High Fidelity and Diversity · AAAI 2026
Natural language and speech › Language models and text generation › reasoning verification
step-level verification
1.012026
Efficient Process Reward Modeling via Contrastive Mutual Information · ACL (1) 2026
Machine learning › Trustworthy machine learning › interpretability
counterfactual explanation
0.912025
Constructing Fair Latent Space for Intersection of Fairness and Explainability · AAAI 2025
Machine learning › Trustworthy machine learning › fairness
fair representation learning
0.912025
Constructing Fair Latent Space for Intersection of Fairness and Explainability · AAAI 2025
Machine learning › Trustworthy machine learning
interpretability
0.912025
Constructing Fair Latent Space for Intersection of Fairness and Explainability · AAAI 2025
Machine learning › Trustworthy machine learning
robustness
0.812024
Mitigating Spurious Correlations via Disagreement Probability · NeurIPS 2024
Machine learning › Trustworthy machine learning › robustness › spurious correlation
spurious correlation mitigation
0.812024
Mitigating Spurious Correlations via Disagreement Probability · NeurIPS 2024
Distributed systems › fault tolerance
byzantine fault tolerance
0.812024
Group-Wise Verifiable Coded Computing Under Byzantine Attacks and Stragglers · IEEE Trans. Inf. Forensics Secur. 2024
Coding theory › error-correcting codes › block codes › linear code
polynomial codes
0.722023
Chebyshev Polynomial Codes: Task Entanglement-based Coding for Distributed Matrix Multiplication · ICML 2021
Straggler-Exploiting Fully Private Distributed Matrix Multiplication With Chebyshev Polynomials · IEEE Trans. Commun. 2023
Distributed systems › distributed algorithms
distributed matrix multiplication
0.722022
Chebyshev Polynomial Codes: Task Entanglement-based Coding for Distributed Matrix Multiplication · ICML 2021
Hierarchical Group Testing for Byzantine Attack Identification in Distributed Matrix Multiplication · IEEE J. Sel. Areas Commun. 2022
Coding theory › error-correcting codes › coded computation
coded distributed computing
0.612022
Hierarchical Group Testing for Byzantine Attack Identification in Distributed Matrix Multiplication · IEEE J. Sel. Areas Commun. 2022
Algorithms and data structures
group testing
0.612022
Hierarchical Group Testing for Byzantine Attack Identification in Distributed Matrix Multiplication · IEEE J. Sel. Areas Commun. 2022
Coding theory › local testability
locally testable codes
0.612022
Hierarchical Group Testing for Byzantine Attack Identification in Distributed Matrix Multiplication · IEEE J. Sel. Areas Commun. 2022
Machine learning › Efficient and distributed learning › model compression
sparse neural network
0.312026
Bias Alleviation Through Network Pruning for Sparse and Debiased Models · IEEE Trans. Image Process. 2026
Machine learning › Generative modeling
synthetic data generation
0.312026
An Adaptive Sampling Framework for Diffusion-based Dataset Distillation with High Fidelity and Diversity · AAAI 2026
Machine learning › Representation and self-supervised learning › latent space
latent space manipulation
0.312025
Constructing Fair Latent Space for Intersection of Fairness and Explainability · AAAI 2025
Distributed systems › distributed machine learning
distributed training
0.212024
Group-Wise Verifiable Coded Computing Under Byzantine Attacks and Stragglers · IEEE Trans. Inf. Forensics Secur. 2024
Coding theory
chebyshev polynomials
0.212023
Straggler-Exploiting Fully Private Distributed Matrix Multiplication With Chebyshev Polynomials · IEEE Trans. Commun. 2023

Methods — techniques the papers use, named apart from their topics

coded computation · 2.0chebyshev polynomials · 1.3group testing · 1.1error-correcting codes · 1.1task entanglement · 1.0repulsion regularization · 1.0polynomial coding · 1.0monte carlo estimation · 1.0debiasing through pruning · 1.0contrastive pointwise mutual information · 1.0adaptive sampling · 1.0accumulated confidence · 1.0generative model fine-tuning · 0.9disentanglement · 0.9group-wise verification · 0.8empirical risk minimization · 0.8disagreement probability resampling · 0.8coding theory · 0.8
YearPublicationVenuePosition
2026 An Adaptive Sampling Framework for Diffusion-based Dataset Distillation with High Fidelity and Diversity
abstract
Dataset distillation (DD) aims to generate a compact synthetic dataset that enables efficient training of neural networks while maintaining performance comparable to that achieved with the original dataset. However, existing methods often suffer from two main limitations. They either rely on computationally intensive iterative optimization procedures or depend heavily on architecture-specific designs. These issues limit their practicality for large-scale datasets and hinder generalization across different model architectures. To overcome these challenges, recent research has explored the use of diffusion models as an architecture-agnostic approach to dataset distillation, offering improved scalability and generalization for large-scale datasets across diverse model architectures. While diffusion-based dataset distillation methods have shown considerable potential, several challenges remain. Notably, certain approaches exhibit a distributional mismatch between the pre-trained diffusion model and the target dataset, which can adversely affect the fidelity and representativeness of the generated samples. Others require substantial fine-tuning to achieve high fidelity, which negates the benefits of architectural flexibility. In this work, we propose a new diffusion-based dataset distillation framework that effectively preserves the characteristics of the original dataset without requiring any fine-tuning. Our method employs adaptive sampling and repulsion regularization to enhance both the fidelity and diversity of generated samples. As a result, the proposed approach outperforms state-of-the-art distillation methods across a wide range of datasets and model architectures.
Sunbeom Jeong, Sehwan Kim, Hyeonggeun Han, Hyungjun Joo, Sangwoo Hong, Jungwoo Lee 0001
AAAI5
2026 Efficient Process Reward Modeling via Contrastive Mutual Information
abstract
Recent research has devoted considerable effort to verifying the intermediate reasoning steps of chain-of-thought (CoT) trajectories using process reward models (PRMs) and other verifier models.However, training a PRM typically requires human annotators to assign reward scores to each reasoning step, which is both costly and time-consuming.Existing automated approaches, such as Monte Carlo (MC) estimation, also demand substantial computational resources due to repeated LLM rollouts.To overcome these limitations, we propose contrastive pointwise mutual information (CPMI), a novel automatic reward labeling method that leverages the model's internal probability to infer step-level supervision while significantly reducing the computational burden of annotating dataset.CPMI quantifies how much a reasoning step increases the mutual information between the step and the correct target answer relative to hard-negative alternatives.This contrastive signal serves as a proxy for the step's contribution to the final solution and yields a reliable reward.The experimental results show that CPMI-based labeling reduces dataset construction time by 84% and token generation by 98% compared to MC estimation, while achieving higher accuracy on process-level evaluations and mathematical reasoning benchmarks.The code is available at https://github.com/nakyungLee20/CPMI.
Nakyung Lee, Sangwoo Hong, Jungwoo Lee 0001
ACL (1)2
2026 Bias Alleviation Through Network Pruning for Sparse and Debiased Models
abstract
Pruning is a highly effective method for reducing the size of neural networks with negligible impact on their average performance. However, recent studies have revealed that pruning actually amplifies the bias in the models, leading to decreased performance for underrepresented groups. To address this issue, we first analyze the impact of pruning on the confidence of each sample and introduce Accumulated Confidence (AC). AC is a proxy that facilitates the identification of bias-conflicting and bias-aligned samples without relying on group annotations. We then propose a debiasing algorithm, which is called DEbiasing Network through Pruning (DENP). DENP utilizes AC to mitigate bias within the network. Even without bias information, DENP exhibits remarkable debiasing performance on varying levels of sparsity, effectively mitigating the bias-exacerbating property of pruning and resulting in both sparse and debiased neural networks. Moreover, even when compared with state-of-the-art debiasing baselines under identical conditions, the DENP still achieves the best performance on multiple benchmark datasets, demonstrating its superior debiasing capabilities.
Sangwoo Hong, Sehwan Kim, Hyungjun Joo, Hyeonggeun Han, Jiyoon Shin, Yoav Wald, Jungwoo Lee 0001
IEEE Trans. Image Process.1
2025 Constructing Fair Latent Space for Intersection of Fairness and Explainability
abstract
As the use of machine learning models has increased, numerous studies have aimed to enhance fairness. However, research on the intersection of fairness and explainability remains insufficient, leading to potential issues in gaining the trust of actual users. Here, we propose a novel module that constructs a fair latent space, enabling faithful explanation while ensuring fairness. The fair latent space is constructed by disentangling and redistributing labels and sensitive attributes, allowing the generation of counterfactual explanations for each type of information. Our module is attached to a pretrained generative model, transforming its biased latent space into a fair latent space. Additionally, since only the module needs to be trained, there are advantages in terms of time and cost savings, without the need to train the entire generative model. We validate the fair latent space with various fairness metrics and demonstrate that our approach can effectively provide explanations for biased decisions and assurances of fairness.
Hyungjun Joo, Hyeonggeun Han, Sehwan Kim, Sangwoo Hong, Jungwoo Lee 0001
AAAI4
2025 Adjusting Initial Noise to Mitigate Memorization in Text-to-Image Diffusion Models
abstract
Despite their impressive generative capabilities, text-to-image diffusion models often memorize and replicate training data, prompting serious concerns over privacy and copyright. Recent work has attributed this memorization to an attraction basin—a region where applying classifier-free guidance (CFG) steers the denoising trajectory toward memorized outputs—and has proposed deferring CFG application until the denoising trajectory escapes this basin. However, such delays often result in non-memorized images that are poorly aligned with the input prompts, highlighting the need to promote earlier escape so that CFG can be applied sooner in the denoising process. In this work, we show that the initial noise sample plays a crucial role in determining when this escape occurs. We empirically observe that different initial samples lead to varying escape times. Building on this insight, we propose two mitigation strategies that adjust the initial noise—either collectively or individually—to find and utilize initial samples that encourage earlier basin escape. These approaches significantly reduce memorization while preserving image-text alignment.
Hyeonggeun Han, Sehwan Kim, Hyungjun Joo, Sangwoo Hong, Jungwoo Lee 0001
NeurIPS4
2024 Learning Dual Hierarchical Representation for 3D Surface Reconstruction
Jiyoon Shin, Youngwook Kim 0005, Sangwoo Hong, Jungwoo Lee 0001
ACCV (9)3
2024 Mitigating Spurious Correlations via Disagreement Probability
abstract
Models trained with empirical risk minimization (ERM) are prone to be biased towards spurious correlations between target labels and bias attributes, which leads to poor performance on data groups lacking spurious correlations. It is particularly challenging to address this problem when access to bias labels is not permitted. To mitigate the effect of spurious correlations without bias labels, we first introduce a novel training objective designed to robustly enhance model performance across all data samples, irrespective of the presence of spurious correlations. From this objective, we then derive a debiasing method, Disagreement Probability based Resampling for debiasing (DPR), which does not require bias labels. DPR leverages the disagreement between the target label and the prediction of a biased model to identify bias-conflicting samples—those without spurious correlations—and upsamples them according to the disagreement probability. Empirical evaluations on multiple benchmarks demonstrate that DPR achieves state-of-the-art performance over existing baselines that do not use bias labels. Furthermore, we provide a theoretical analysis that details how DPR reduces dependency on spurious correlations.
Hyeonggeun Han, Sehwan Kim, Hyungjun Joo, Sangwoo Hong, Jungwoo Lee 0001
NeurIPS4
2024 Group-Wise Verifiable Coded Computing Under Byzantine Attacks and Stragglers
abstract
Distributed computing has emerged as a promising solution for accelerating machine learning training processes on large-scale datasets by leveraging the parallel processing capabilities of multiple workers. However, there remain two major issues that still need to be addressed: i) Byzantine attacks from malicious workers, and ii) the effect of slow workers, commonly referred to as stragglers. In this paper, we address both issues concurrently by introducing Group-wise Verifiable Coded Computing (GVCC), a novel approach that combines coding techniques and group-wise verification to enhance robustness against Byzantine attacks and resilience to straggler effects in distributed computing. The key idea of GVCC is to verify a group of computation results from workers at a time, while providing resilience to stragglers through encoding tasks assigned to workers with Group-wise Verifiable Codes. We evaluate the performance of GVCC through experiments conducted on Amazon EC2 clouds and the results show that GVCC outperforms the existing methods in terms of overall processing time and verification time while maintaining the verification performance. This study highlights the potential of GVCC as an effective solution for overcoming the challenges of Byzantine attacks and stragglers in distributed computing for executing matrix multiplication.
Sangwoo Hong, Heecheol Yang, Youngseok Yoon, Jungwoo Lee 0001
IEEE Trans. Inf. Forensics Secur.1
2023 Straggler-Exploiting Fully Private Distributed Matrix Multiplication With Chebyshev Polynomials
abstract
In this paper, we consider coded computation for matrix multiplication tasks in distributed computing to mitigate straggler effects. We assume that the stragglers’ computation results can be leveraged at the master by assigning multiple sub-tasks to the workers. We propose a new coded computation scheme, namely Chebyshev coded fully private matrix multiplication (CFP), to preserve the privacy of a master in a scenario where a master wants to obtain a matrix multiplication result from the libraries which are shared by the workers, while concealing both of the two indices of the desired matrices from each worker. The key idea of CFP is to introduce Chebyshev polynomials, which have commutative property, in queries sent to workers to allocate sub-tasks. We also extend CFP to keep the privacy of a master from colluding workers. In conclusion, we show that CFP can preserve the privacy of a master from each worker and efficiently mitigate straggler effects compared to existing schemes.
Sangwoo Hong, Heecheol Yang, Youngseok Yoon, Jungwoo Lee 0001
IEEE Trans. Commun.1
2022 Doppler Analysis and Compensation for Distributed LEO-MIMO Satellite Communications
abstract
In this paper, we propose a Doppler compensation method in LEO-MIMO communication, where two LEO satellites are used as amplify-and-forward (AF) relays, and line-of-sight (LOS) propagation dominates the communication channel. We first analyze Doppler effects in uplink and downlink, and propose a dual-hop AF relay channel model including Doppler effects. Also we show that Doppler effects in our scenario can be easily compensated at the LEO satellites, without the need of intersatellite link (ISL) communication.
Sangwoo Hong, Wonjae Shin, Jungwoo Lee 0001
APCC1
2022 Byzantine Attack Identification in Distributed Matrix Multiplication via locally testable codes
abstract
Coded computing has proved its efficiency in handling a straggler issue in distributed computing framework. However, in a coded distributed computing framework, there may exist Byzantine workers who send the wrong computation results to a master to contaminate the overall computation output. Therefore, it is essential to identify Byzantine workers from their computation results in coded computing. In this paper, we consider Byzantine attack identification problem in coded computing for distributed matrix multiplication tasks. We propose locally testable codes which facilitate the efficient Byzantine attack identification, and suggest a hierarchical group testing method for Byzantine attack identification. We show that our scheme requires smaller number of tests than the conventional group testing methods for the existing coded computing schemes.
Sangwoo Hong, Heecheol Yang, Jungwoo Lee 0001
ISIT1
2022 Hierarchical Group Testing for Byzantine Attack Identification in Distributed Matrix Multiplication
abstract
Coded computing has proved its efficiency in handling a straggler issue in distributed computing framework. It uses error correcting codes to mitigate the effect of the stragglers. However, in a coded distributed computing framework, there may exist Byzantine workers who send the wrong computation results to a master in order to contaminate the overall computation output. Therefore, it is essential to identify Byzantine workers from their computation results in coded computing. In this paper, we consider Byzantine attack identification problem in coded computing for distributed matrix multiplication tasks. We propose a new coding scheme which facilitates the efficient Byzantine attack identification, namely locally testable codes. We also suggest a hierarchical group testing method for Byzantine attack identification. We claim the required number of tests for group testing in our scheme, and show that it requires smaller number of tests than the conventional group testing method for the existing coded computing schemes.
Sangwoo Hong, Heecheol Yang, Jungwoo Lee 0001
IEEE J. Sel. Areas Commun.1
2021 Chebyshev Polynomial Codes: Task Entanglement-based Coding for Distributed Matrix Multiplication
abstract
Distributed computing has been a prominent solution to efficiently process massive datasets in parallel. However, the existence of stragglers is one of the major concerns that slows down the overall speed of distributed computing. To deal with this problem, we consider a distributed matrix multiplication scenario where a master assigns multiple tasks to each worker to exploit stragglers’ computing ability (which is typically wasted in conventional distributed computing). We propose Chebyshev polynomial codes, which can achieve order-wise improvement in encoding complexity at the master and communication load in distributed matrix multiplication using task entanglement. The key idea of task entanglement is to reduce the number of encoded matrices for multiple tasks assigned to each worker by intertwining encoded matrices. We experimentally demonstrate that, in cloud environments, Chebyshev polynomial codes can provide significant reduction in overall processing time in distributed computing for matrix multiplication, which is a key computational component in modern deep learning.
Sangwoo Hong, Heecheol Yang, Youngseok Yoon, Taehyun Cho, Jungwoo Lee 0001
ICML1
2021 Private and Secure Coded Computation in Straggler-Exploiting Distributed Matrix Multiplication
abstract
In this paper, we consider coded computation for matrix multiplication tasks in distributed computing, which can mitigate the effect of slow workers, called stragglers, by a coding approach. We assume that the stragglers' computation results can be leveraged at the master by assigning multiple sub-tasks to the workers. In this scenario, we propose a new coded computation scheme to preserve the data privacy and security from the non-colluding workers. We also prove that the data privacy and security constraints are satisfied in our scheme in an information-theoretic sense.
Heecheol Yang, Sangwoo Hong, Jungwoo Lee 0001
ISIT2
2003 The Dichotomy of Presence Elements: The Where and What
abstract
One of the goals and defining characteristics of virtual reality systems is to create "presence" and fool the user into believing that one is, or is doing something "in" the synthetic environment. Most research and papers on presence to date have been directed toward coming up with the definitions of presence, and based on them, identifying key elements that affect presence. We carried out an elaborate experiment in which presence levels were measured (with subjective questionnaire) in test virtual worlds configured with different combinations of six visual presence elements.
Dongsik Cho, Gerard Jounghyun Kim, Sangwoo Hong, Sung Ho Han, Seungyong Lee 0001
VR4