VLDB 2026 Research / reviewers in the wild / expert
Sangwoo Hong
dblp:48/334
· DBLP profile ↗
15ranked-venue papers
7as first author
14since 2021 · last 2026
0000-0002-0270-2781ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 4 since 2021Computer networks · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Trustworthy machine learning · 44% Efficient and distributed learning · 24% Language models and text generation · 11% | |
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
Distributed systems · 98% High-performance computing · 2% | |
| Theoretical computer science
3 papers |
Coding theory · 78% Algorithms and data structures · 22% |
Topics — the 30 heaviest of 34, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
fairness |
2.6 | 3 | 2026 | Bias Alleviation Through Network Pruning for Sparse and Debiased Models · IEEE Trans. Image Process. 2026 Constructing Fair Latent Space for Intersection of Fairness and Explainability · AAAI 2025 Mitigating Spurious Correlations via Disagreement Probability · NeurIPS 2024 |
Distributed systems › distributed data processing
straggler mitigation |
1.9 | 3 | 2024 | Group-Wise Verifiable Coded Computing Under Byzantine Attacks and Stragglers · IEEE Trans. Inf. Forensics Secur. 2024 Straggler-Exploiting Fully Private Distributed Matrix Multiplication With Chebyshev Polynomials · IEEE Trans. Commun. 2023 Chebyshev Polynomial Codes: Task Entanglement-based Coding for Distributed Matrix Multiplication · ICML 2021 |
Distributed systems
coded computation |
1.4 | 2 | 2024 | Group-Wise Verifiable Coded Computing Under Byzantine Attacks and Stragglers · IEEE Trans. Inf. Forensics Secur. 2024 Straggler-Exploiting Fully Private Distributed Matrix Multiplication With Chebyshev Polynomials · IEEE Trans. Commun. 2023 |
Distributed systems
fault tolerance |
1.3 | 2 | 2024 | Group-Wise Verifiable Coded Computing Under Byzantine Attacks and Stragglers · IEEE Trans. Inf. Forensics Secur. 2024 Hierarchical Group Testing for Byzantine Attack Identification in Distributed Matrix Multiplication · IEEE J. Sel. Areas Commun. 2022 |
Natural language and speech › Language models and text generation
chain-of-thought reasoning |
1.0 | 1 | 2026 | Efficient Process Reward Modeling via Contrastive Mutual Information · ACL (1) 2026 |
Machine learning › Efficient and distributed learning
dataset distillation |
1.0 | 1 | 2026 | An Adaptive Sampling Framework for Diffusion-based Dataset Distillation with High Fidelity and Diversity · AAAI 2026 |
Machine learning › Trustworthy machine learning
debiasing |
1.0 | 1 | 2026 | Bias Alleviation Through Network Pruning for Sparse and Debiased Models · IEEE Trans. Image Process. 2026 |
Machine learning › Efficient and distributed learning › dataset distillation
diffusion-based dataset distillation |
1.0 | 1 | 2026 | An Adaptive Sampling Framework for Diffusion-based Dataset Distillation with High Fidelity and Diversity · AAAI 2026 |
Machine learning › Generative modeling
diffusion model |
1.0 | 1 | 2026 | An Adaptive Sampling Framework for Diffusion-based Dataset Distillation with High Fidelity and Diversity · AAAI 2026 |
Machine learning › Efficient and distributed learning
model compression |
1.0 | 1 | 2026 | Bias Alleviation Through Network Pruning for Sparse and Debiased Models · IEEE Trans. Image Process. 2026 |
Machine learning › Reinforcement learning › reinforcement learning from human feedback
process reward model |
1.0 | 1 | 2026 | Efficient Process Reward Modeling via Contrastive Mutual Information · ACL (1) 2026 |
Machine learning › Efficient and distributed learning › model compression
pruning |
1.0 | 1 | 2026 | Bias Alleviation Through Network Pruning for Sparse and Debiased Models · IEEE Trans. Image Process. 2026 |
Machine learning › Probabilistic and Bayesian machine learning
sampling |
1.0 | 1 | 2026 | An Adaptive Sampling Framework for Diffusion-based Dataset Distillation with High Fidelity and Diversity · AAAI 2026 |
Natural language and speech › Language models and text generation › reasoning verification
step-level verification |
1.0 | 1 | 2026 | Efficient Process Reward Modeling via Contrastive Mutual Information · ACL (1) 2026 |
Machine learning › Trustworthy machine learning › interpretability
counterfactual explanation |
0.9 | 1 | 2025 | Constructing Fair Latent Space for Intersection of Fairness and Explainability · AAAI 2025 |
Machine learning › Trustworthy machine learning › fairness
fair representation learning |
0.9 | 1 | 2025 | Constructing Fair Latent Space for Intersection of Fairness and Explainability · AAAI 2025 |
Machine learning › Trustworthy machine learning
interpretability |
0.9 | 1 | 2025 | Constructing Fair Latent Space for Intersection of Fairness and Explainability · AAAI 2025 |
Machine learning › Trustworthy machine learning
robustness |
0.8 | 1 | 2024 | Mitigating Spurious Correlations via Disagreement Probability · NeurIPS 2024 |
Machine learning › Trustworthy machine learning › robustness › spurious correlation
spurious correlation mitigation |
0.8 | 1 | 2024 | Mitigating Spurious Correlations via Disagreement Probability · NeurIPS 2024 |
Distributed systems › fault tolerance
byzantine fault tolerance |
0.8 | 1 | 2024 | Group-Wise Verifiable Coded Computing Under Byzantine Attacks and Stragglers · IEEE Trans. Inf. Forensics Secur. 2024 |
Coding theory › error-correcting codes › block codes › linear code
polynomial codes |
0.7 | 2 | 2023 | Chebyshev Polynomial Codes: Task Entanglement-based Coding for Distributed Matrix Multiplication · ICML 2021 Straggler-Exploiting Fully Private Distributed Matrix Multiplication With Chebyshev Polynomials · IEEE Trans. Commun. 2023 |
Distributed systems › distributed algorithms
distributed matrix multiplication |
0.7 | 2 | 2022 | Chebyshev Polynomial Codes: Task Entanglement-based Coding for Distributed Matrix Multiplication · ICML 2021 Hierarchical Group Testing for Byzantine Attack Identification in Distributed Matrix Multiplication · IEEE J. Sel. Areas Commun. 2022 |
Coding theory › error-correcting codes › coded computation
coded distributed computing |
0.6 | 1 | 2022 | Hierarchical Group Testing for Byzantine Attack Identification in Distributed Matrix Multiplication · IEEE J. Sel. Areas Commun. 2022 |
Algorithms and data structures
group testing |
0.6 | 1 | 2022 | Hierarchical Group Testing for Byzantine Attack Identification in Distributed Matrix Multiplication · IEEE J. Sel. Areas Commun. 2022 |
Coding theory › local testability
locally testable codes |
0.6 | 1 | 2022 | Hierarchical Group Testing for Byzantine Attack Identification in Distributed Matrix Multiplication · IEEE J. Sel. Areas Commun. 2022 |
Machine learning › Efficient and distributed learning › model compression
sparse neural network |
0.3 | 1 | 2026 | Bias Alleviation Through Network Pruning for Sparse and Debiased Models · IEEE Trans. Image Process. 2026 |
Machine learning › Generative modeling
synthetic data generation |
0.3 | 1 | 2026 | An Adaptive Sampling Framework for Diffusion-based Dataset Distillation with High Fidelity and Diversity · AAAI 2026 |
Machine learning › Representation and self-supervised learning › latent space
latent space manipulation |
0.3 | 1 | 2025 | Constructing Fair Latent Space for Intersection of Fairness and Explainability · AAAI 2025 |
Distributed systems › distributed machine learning
distributed training |
0.2 | 1 | 2024 | Group-Wise Verifiable Coded Computing Under Byzantine Attacks and Stragglers · IEEE Trans. Inf. Forensics Secur. 2024 |
Coding theory
chebyshev polynomials |
0.2 | 1 | 2023 | Straggler-Exploiting Fully Private Distributed Matrix Multiplication With Chebyshev Polynomials · IEEE Trans. Commun. 2023 |
Methods — techniques the papers use, named apart from their topics
coded computation · 2.0chebyshev polynomials · 1.3group testing · 1.1error-correcting codes · 1.1task entanglement · 1.0repulsion regularization · 1.0polynomial coding · 1.0monte carlo estimation · 1.0debiasing through pruning · 1.0contrastive pointwise mutual information · 1.0adaptive sampling · 1.0accumulated confidence · 1.0generative model fine-tuning · 0.9disentanglement · 0.9group-wise verification · 0.8empirical risk minimization · 0.8disagreement probability resampling · 0.8coding theory · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An Adaptive Sampling Framework for Diffusion-based Dataset Distillation with High Fidelity and DiversityabstractDataset distillation (DD) aims to generate a compact synthetic dataset that enables efficient training of neural networks while maintaining performance comparable to that achieved with the original dataset. However, existing methods often suffer from two main limitations. They either rely on computationally intensive iterative optimization procedures or depend heavily on architecture-specific designs. These issues limit their practicality for large-scale datasets and hinder generalization across different model architectures. To overcome these challenges, recent research has explored the use of diffusion models as an architecture-agnostic approach to dataset distillation, offering improved scalability and generalization for large-scale datasets across diverse model architectures. While diffusion-based dataset distillation methods have shown considerable potential, several challenges remain. Notably, certain approaches exhibit a distributional mismatch between the pre-trained diffusion model and the target dataset, which can adversely affect the fidelity and representativeness of the generated samples. Others require substantial fine-tuning to achieve high fidelity, which negates the benefits of architectural flexibility. In this work, we propose a new diffusion-based dataset distillation framework that effectively preserves the characteristics of the original dataset without requiring any fine-tuning. Our method employs adaptive sampling and repulsion regularization to enhance both the fidelity and diversity of generated samples. As a result, the proposed approach outperforms state-of-the-art distillation methods across a wide range of datasets and model architectures. Sunbeom Jeong, Sehwan Kim, Hyeonggeun Han, Hyungjun Joo, Sangwoo Hong, Jungwoo Lee 0001 |
AAAI | 5 |
| 2026 | Efficient Process Reward Modeling via Contrastive Mutual InformationabstractRecent research has devoted considerable effort to verifying the intermediate reasoning steps of chain-of-thought (CoT) trajectories using process reward models (PRMs) and other verifier models.However, training a PRM typically requires human annotators to assign reward scores to each reasoning step, which is both costly and time-consuming.Existing automated approaches, such as Monte Carlo (MC) estimation, also demand substantial computational resources due to repeated LLM rollouts.To overcome these limitations, we propose contrastive pointwise mutual information (CPMI), a novel automatic reward labeling method that leverages the model's internal probability to infer step-level supervision while significantly reducing the computational burden of annotating dataset.CPMI quantifies how much a reasoning step increases the mutual information between the step and the correct target answer relative to hard-negative alternatives.This contrastive signal serves as a proxy for the step's contribution to the final solution and yields a reliable reward.The experimental results show that CPMI-based labeling reduces dataset construction time by 84% and token generation by 98% compared to MC estimation, while achieving higher accuracy on process-level evaluations and mathematical reasoning benchmarks.The code is available at https://github.com/nakyungLee20/CPMI. Nakyung Lee, Sangwoo Hong, Jungwoo Lee 0001 |
ACL (1) | 2 |
| 2026 | Bias Alleviation Through Network Pruning for Sparse and Debiased ModelsabstractPruning is a highly effective method for reducing the size of neural networks with negligible impact on their average performance. However, recent studies have revealed that pruning actually amplifies the bias in the models, leading to decreased performance for underrepresented groups. To address this issue, we first analyze the impact of pruning on the confidence of each sample and introduce Accumulated Confidence (AC). AC is a proxy that facilitates the identification of bias-conflicting and bias-aligned samples without relying on group annotations. We then propose a debiasing algorithm, which is called DEbiasing Network through Pruning (DENP). DENP utilizes AC to mitigate bias within the network. Even without bias information, DENP exhibits remarkable debiasing performance on varying levels of sparsity, effectively mitigating the bias-exacerbating property of pruning and resulting in both sparse and debiased neural networks. Moreover, even when compared with state-of-the-art debiasing baselines under identical conditions, the DENP still achieves the best performance on multiple benchmark datasets, demonstrating its superior debiasing capabilities. Sangwoo Hong, Sehwan Kim, Hyungjun Joo, Hyeonggeun Han, Jiyoon Shin, Yoav Wald, Jungwoo Lee 0001 |
IEEE Trans. Image Process. | 1 |
| 2025 | Constructing Fair Latent Space for Intersection of Fairness and ExplainabilityabstractAs the use of machine learning models has increased, numerous studies have aimed to enhance fairness. However, research on the intersection of fairness and explainability remains insufficient, leading to potential issues in gaining the trust of actual users. Here, we propose a novel module that constructs a fair latent space, enabling faithful explanation while ensuring fairness. The fair latent space is constructed by disentangling and redistributing labels and sensitive attributes, allowing the generation of counterfactual explanations for each type of information. Our module is attached to a pretrained generative model, transforming its biased latent space into a fair latent space. Additionally, since only the module needs to be trained, there are advantages in terms of time and cost savings, without the need to train the entire generative model. We validate the fair latent space with various fairness metrics and demonstrate that our approach can effectively provide explanations for biased decisions and assurances of fairness. Hyungjun Joo, Hyeonggeun Han, Sehwan Kim, Sangwoo Hong, Jungwoo Lee 0001 |
AAAI | 4 |
| 2025 | Adjusting Initial Noise to Mitigate Memorization in Text-to-Image Diffusion ModelsabstractDespite their impressive generative capabilities, text-to-image diffusion models often memorize and replicate training data, prompting serious concerns over privacy and copyright. Recent work has attributed this memorization to an attraction basin—a region where applying classifier-free guidance (CFG) steers the denoising trajectory toward memorized outputs—and has proposed deferring CFG application until the denoising trajectory escapes this basin. However, such delays often result in non-memorized images that are poorly aligned with the input prompts, highlighting the need to promote earlier escape so that CFG can be applied sooner in the denoising process. In this work, we show that the initial noise sample plays a crucial role in determining when this escape occurs. We empirically observe that different initial samples lead to varying escape times. Building on this insight, we propose two mitigation strategies that adjust the initial noise—either collectively or individually—to find and utilize initial samples that encourage earlier basin escape. These approaches significantly reduce memorization while preserving image-text alignment. Hyeonggeun Han, Sehwan Kim, Hyungjun Joo, Sangwoo Hong, Jungwoo Lee 0001 |
NeurIPS | 4 |
| 2024 | Learning Dual Hierarchical Representation for 3D Surface Reconstruction
Jiyoon Shin, Youngwook Kim 0005, Sangwoo Hong, Jungwoo Lee 0001 |
ACCV (9) | 3 |
| 2024 | Mitigating Spurious Correlations via Disagreement ProbabilityabstractModels trained with empirical risk minimization (ERM) are prone to be biased towards spurious correlations between target labels and bias attributes, which leads to poor performance on data groups lacking spurious correlations. It is particularly challenging to address this problem when access to bias labels is not permitted. To mitigate the effect of spurious correlations without bias labels, we first introduce a novel training objective designed to robustly enhance model performance across all data samples, irrespective of the presence of spurious correlations. From this objective, we then derive a debiasing method, Disagreement Probability based Resampling for debiasing (DPR), which does not require bias labels. DPR leverages the disagreement between the target label and the prediction of a biased model to identify bias-conflicting samples—those without spurious correlations—and upsamples them according to the disagreement probability. Empirical evaluations on multiple benchmarks demonstrate that DPR achieves state-of-the-art performance over existing baselines that do not use bias labels. Furthermore, we provide a theoretical analysis that details how DPR reduces dependency on spurious correlations. Hyeonggeun Han, Sehwan Kim, Hyungjun Joo, Sangwoo Hong, Jungwoo Lee 0001 |
NeurIPS | 4 |
| 2024 | Group-Wise Verifiable Coded Computing Under Byzantine Attacks and StragglersabstractDistributed computing has emerged as a promising solution for accelerating machine learning training processes on large-scale datasets by leveraging the parallel processing capabilities of multiple workers. However, there remain two major issues that still need to be addressed: i) Byzantine attacks from malicious workers, and ii) the effect of slow workers, commonly referred to as stragglers. In this paper, we address both issues concurrently by introducing Group-wise Verifiable Coded Computing (GVCC), a novel approach that combines coding techniques and group-wise verification to enhance robustness against Byzantine attacks and resilience to straggler effects in distributed computing. The key idea of GVCC is to verify a group of computation results from workers at a time, while providing resilience to stragglers through encoding tasks assigned to workers with Group-wise Verifiable Codes. We evaluate the performance of GVCC through experiments conducted on Amazon EC2 clouds and the results show that GVCC outperforms the existing methods in terms of overall processing time and verification time while maintaining the verification performance. This study highlights the potential of GVCC as an effective solution for overcoming the challenges of Byzantine attacks and stragglers in distributed computing for executing matrix multiplication. Sangwoo Hong, Heecheol Yang, Youngseok Yoon, Jungwoo Lee 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2023 | Straggler-Exploiting Fully Private Distributed Matrix Multiplication With Chebyshev PolynomialsabstractIn this paper, we consider coded computation for matrix multiplication tasks in distributed computing to mitigate straggler effects. We assume that the stragglers’ computation results can be leveraged at the master by assigning multiple sub-tasks to the workers. We propose a new coded computation scheme, namely Chebyshev coded fully private matrix multiplication (CFP), to preserve the privacy of a master in a scenario where a master wants to obtain a matrix multiplication result from the libraries which are shared by the workers, while concealing both of the two indices of the desired matrices from each worker. The key idea of CFP is to introduce Chebyshev polynomials, which have commutative property, in queries sent to workers to allocate sub-tasks. We also extend CFP to keep the privacy of a master from colluding workers. In conclusion, we show that CFP can preserve the privacy of a master from each worker and efficiently mitigate straggler effects compared to existing schemes. Sangwoo Hong, Heecheol Yang, Youngseok Yoon, Jungwoo Lee 0001 |
IEEE Trans. Commun. | 1 |
| 2022 | Doppler Analysis and Compensation for Distributed LEO-MIMO Satellite CommunicationsabstractIn this paper, we propose a Doppler compensation method in LEO-MIMO communication, where two LEO satellites are used as amplify-and-forward (AF) relays, and line-of-sight (LOS) propagation dominates the communication channel. We first analyze Doppler effects in uplink and downlink, and propose a dual-hop AF relay channel model including Doppler effects. Also we show that Doppler effects in our scenario can be easily compensated at the LEO satellites, without the need of intersatellite link (ISL) communication. Sangwoo Hong, Wonjae Shin, Jungwoo Lee 0001 |
APCC | 1 |
| 2022 | Byzantine Attack Identification in Distributed Matrix Multiplication via locally testable codesabstractCoded computing has proved its efficiency in handling a straggler issue in distributed computing framework. However, in a coded distributed computing framework, there may exist Byzantine workers who send the wrong computation results to a master to contaminate the overall computation output. Therefore, it is essential to identify Byzantine workers from their computation results in coded computing. In this paper, we consider Byzantine attack identification problem in coded computing for distributed matrix multiplication tasks. We propose locally testable codes which facilitate the efficient Byzantine attack identification, and suggest a hierarchical group testing method for Byzantine attack identification. We show that our scheme requires smaller number of tests than the conventional group testing methods for the existing coded computing schemes. Sangwoo Hong, Heecheol Yang, Jungwoo Lee 0001 |
ISIT | 1 |
| 2022 | Hierarchical Group Testing for Byzantine Attack Identification in Distributed Matrix MultiplicationabstractCoded computing has proved its efficiency in handling a straggler issue in distributed computing framework. It uses error correcting codes to mitigate the effect of the stragglers. However, in a coded distributed computing framework, there may exist Byzantine workers who send the wrong computation results to a master in order to contaminate the overall computation output. Therefore, it is essential to identify Byzantine workers from their computation results in coded computing. In this paper, we consider Byzantine attack identification problem in coded computing for distributed matrix multiplication tasks. We propose a new coding scheme which facilitates the efficient Byzantine attack identification, namely locally testable codes. We also suggest a hierarchical group testing method for Byzantine attack identification. We claim the required number of tests for group testing in our scheme, and show that it requires smaller number of tests than the conventional group testing method for the existing coded computing schemes. Sangwoo Hong, Heecheol Yang, Jungwoo Lee 0001 |
IEEE J. Sel. Areas Commun. | 1 |
| 2021 | Chebyshev Polynomial Codes: Task Entanglement-based Coding for Distributed Matrix MultiplicationabstractDistributed computing has been a prominent solution to efficiently process massive datasets in parallel. However, the existence of stragglers is one of the major concerns that slows down the overall speed of distributed computing. To deal with this problem, we consider a distributed matrix multiplication scenario where a master assigns multiple tasks to each worker to exploit stragglers’ computing ability (which is typically wasted in conventional distributed computing). We propose Chebyshev polynomial codes, which can achieve order-wise improvement in encoding complexity at the master and communication load in distributed matrix multiplication using task entanglement. The key idea of task entanglement is to reduce the number of encoded matrices for multiple tasks assigned to each worker by intertwining encoded matrices. We experimentally demonstrate that, in cloud environments, Chebyshev polynomial codes can provide significant reduction in overall processing time in distributed computing for matrix multiplication, which is a key computational component in modern deep learning. Sangwoo Hong, Heecheol Yang, Youngseok Yoon, Taehyun Cho, Jungwoo Lee 0001 |
ICML | 1 |
| 2021 | Private and Secure Coded Computation in Straggler-Exploiting Distributed Matrix MultiplicationabstractIn this paper, we consider coded computation for matrix multiplication tasks in distributed computing, which can mitigate the effect of slow workers, called stragglers, by a coding approach. We assume that the stragglers' computation results can be leveraged at the master by assigning multiple sub-tasks to the workers. In this scenario, we propose a new coded computation scheme to preserve the data privacy and security from the non-colluding workers. We also prove that the data privacy and security constraints are satisfied in our scheme in an information-theoretic sense. Heecheol Yang, Sangwoo Hong, Jungwoo Lee 0001 |
ISIT | 2 |
| 2003 | The Dichotomy of Presence Elements: The Where and WhatabstractOne of the goals and defining characteristics of virtual reality systems is to create "presence" and fool the user into believing that one is, or is doing something "in" the synthetic environment. Most research and papers on presence to date have been directed toward coming up with the definitions of presence, and based on them, identifying key elements that affect presence. We carried out an elaborate experiment in which presence levels were measured (with subjective questionnaire) in test virtual worlds configured with different combinations of six visual presence elements. Dongsik Cho, Gerard Jounghyun Kim, Sangwoo Hong, Sung Ho Han, Seungyong Lee 0001 |
VR | 4 |