EDBT 2026 Demo / reviewers in the wild / expert
Wonpyo Park
dblp:239/5044
· DBLP profile ↗
10ranked-venue papers
2as first author
6since 2021 · last 2025
0000-0003-0675-6362ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Efficient and distributed learning · 48% Face, body and person analysis · 14% Representation and self-supervised learning · 14% |
Topics — the 17 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
model compression |
3.1 | 4 | 2025 | GuidedQuant: Large Language Model Quantization via Exploiting End Loss Guidance · ICML 2025 Prefixing Attention Sinks can Mitigate Activation Outliers for Large Language Model Quantization · EMNLP 2024 Rethinking Pruning Large Language Models: Benefits and Pitfalls of Reconstruction Error Minimization · EMNLP 2024 |
Natural language and speech › Language models and text generation
large language model inference |
1.2 | 3 | 2025 | Breaking ReLU Barrier: Generalized MoEfication for Dense Pretrained Models · EMNLP 2024 GuidedQuant: Large Language Model Quantization via Exploiting End Loss Guidance · ICML 2025 Prefixing Attention Sinks can Mitigate Activation Outliers for Large Language Model Quantization · EMNLP 2024 |
Machine learning › Representation and self-supervised learning › representation learning
metric learning |
0.9 | 2 | 2021 | Multi-level Distance Regularization for Deep Metric Learning · AAAI 2021 Relational Knowledge Distillation · CVPR 2019 |
Computer vision › Face, body and person analysis
face recognition |
0.9 | 2 | 2020 | BroadFace: Looking at Tens of Thousands of People at once for Face Recognition · ECCV (9) 2020 GroupFace: Learning Latent Groups and Constructing Group-Based Representations for Face Recognition · CVPR 2020 |
Machine learning › Efficient and distributed learning › model compression › quantization
post-training quantization |
0.9 | 1 | 2025 | GuidedQuant: Large Language Model Quantization via Exploiting End Loss Guidance · ICML 2025 |
Machine learning › Efficient and distributed learning › model compression › quantization
activation quantization |
0.8 | 1 | 2024 | Prefixing Attention Sinks can Mitigate Activation Outliers for Large Language Model Quantization · EMNLP 2024 |
Machine learning › Efficient and distributed learning › model compression › pruning › DNN pruning
language model pruning |
0.8 | 1 | 2024 | Rethinking Pruning Large Language Models: Benefits and Pitfalls of Reconstruction Error Minimization · EMNLP 2024 |
Machine learning › Trustworthy machine learning
outlier mitigation |
0.8 | 1 | 2024 | Prefixing Attention Sinks can Mitigate Activation Outliers for Large Language Model Quantization · EMNLP 2024 |
Machine learning › Deep learning architectures and training › mixture of experts
sparse expert activation |
0.8 | 1 | 2024 | Breaking ReLU Barrier: Generalized MoEfication for Dense Pretrained Models · EMNLP 2024 |
Machine learning › Representation and self-supervised learning › representation learning › metric learning
deep metric learning |
0.5 | 1 | 2021 | Multi-level Distance Regularization for Deep Metric Learning · AAAI 2021 |
Machine learning › Deep learning architectures and training
regularization |
0.5 | 1 | 2021 | Multi-level Distance Regularization for Deep Metric Learning · AAAI 2021 |
Computer vision › Face, body and person analysis › face recognition
face verification |
0.4 | 1 | 2020 | GroupFace: Learning Latent Groups and Constructing Group-Based Representations for Face Recognition · CVPR 2020 |
Machine learning › Representation and self-supervised learning › equivariance › equivariant representation learning
group representation learning |
0.4 | 1 | 2020 | GroupFace: Learning Latent Groups and Constructing Group-Based Representations for Face Recognition · CVPR 2020 |
Computer vision › Face, body and person analysis › face recognition
large-scale face recognition |
0.4 | 1 | 2020 | BroadFace: Looking at Tens of Thousands of People at once for Face Recognition · ECCV (9) 2020 |
Machine learning › Efficient and distributed learning › model compression
knowledge distillation |
0.4 | 1 | 2019 | Relational Knowledge Distillation · CVPR 2019 |
Machine learning › Efficient and distributed learning › model compression › knowledge distillation
relational knowledge distillation |
0.4 | 1 | 2019 | Relational Knowledge Distillation · CVPR 2019 |
Computer vision › Face, body and person analysis › face recognition
face identification |
0.1 | 1 | 2020 | GroupFace: Learning Latent Groups and Constructing Group-Based Representations for Face Recognition · CVPR 2020 |
Methods — techniques the papers use, named apart from their topics
non-uniform scalar quantization · 0.9end loss gradient · 0.9reconstruction error minimization · 0.8key-value cache prefixing · 0.8greedy token search · 0.8continual pre-training · 0.8calibration data self-generation · 0.8activation sparsity · 0.8triplet loss · 0.5multi-level distance regularization · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | GuidedQuant: Large Language Model Quantization via Exploiting End Loss GuidanceabstractPost-training quantization is a key technique for reducing the memory and inference latency of large language models by quantizing weights and activations without requiring retraining. However, existing methods either (1) fail to account for the varying importance of hidden features to the end loss or, when incorporating end loss, (2) neglect the critical interactions between model weights. To address these limitations, we propose GuidedQuant, a novel quantization approach that integrates gradient information from the end loss into the quantization objective while preserving cross-weight dependencies within output channels. GuidedQuant consistently boosts the performance of state-of-the-art quantization methods across weight-only scalar, weight-only vector, and weight-and-activation quantization. Additionally, we introduce a novel non-uniform scalar quantization algorithm, which is guaranteed to monotonically decrease the quantization objective value, and outperforms existing methods in this category. We release the code at https://github.com/snu-mllab/GuidedQuant. Marwa El Halabi, Wonpyo Park, Clemens JS Schaefer, Deokjae Lee, Yeonhong Park, Jae W. Lee, Hyun Oh Song |
ICML | 3 |
| 2025 | LCIRC: A Recurrent Compression Approach for Efficient Long-form Context and Query Dependent Modeling in LLMsabstractSumin An, Junyoung Sung, Wonpyo Park, Chanjun Park, Paul Hongsuck Seo. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Sumin An, Junyoung Sung, Wonpyo Park, Chanjun Park, Hongsuck Seo |
NAACL (Long Papers) | 3 |
| 2024 | Breaking ReLU Barrier: Generalized MoEfication for Dense Pretrained ModelsabstractAs the scale of language models (LMs) continues to grow, there is a heightened interest in reducing the inference cost associated with these models.Mixture-of-Experts (MoEs) present an efficient alternative to dense models, while the existing methods to convert pretrained dense models to MoEs is limited to ReLU-based models with natural sparsity.This paper introduces G-MoEfication, applicable to arbitrary dense models, where ReLU-based activation sparsity assumptions no longer hold.For generalizations, we encounter the dilemma of needing to zero-out deactivated experts, while also avoiding excessive zeroing-out to retain dense activation information.We publicly release our code 1 and report results conducted with mBERT, SantaCoder-1.1B,Phi-2-2.7B, and Falcon-7B demonstrating the efficacy of our approach in general scenarios: from multitask to multilingual, from fine-tuning to zero-shot evaluation. Jaeseong Lee 0002, Seung-won Hwang, Wonpyo Park, Mingi Ji |
EMNLP | 3 |
| 2024 | Rethinking Pruning Large Language Models: Benefits and Pitfalls of Reconstruction Error MinimizationabstractThis work suggests fundamentally rethinking the current practice of pruning large language models (LLMs).The way it is done is by divide and conquer: split the model into submodels, sequentially prune them, and reconstruct predictions of the dense counterparts on small calibration data one at a time; the final model is obtained simply by putting the resulting sparse submodels together.While this approach enables pruning under memory constraints, it generates high reconstruction errors.In this work, we first present an array of reconstruction techniques that can significantly reduce this error by more than 90%.Unwittingly, however, we discover that minimizing reconstruction error is not always ideal and can overfit the given calibration data, resulting in rather increased language perplexity and poor performance at downstream tasks.We find out that a strategy of self-generating calibration data can mitigate this trade-off between reconstruction and generalization, suggesting new directions in the presence of both benefits and pitfalls of reconstruction for pruning LLMs. 1 Sungbin Shin, Wonpyo Park, Jaeho Lee 0001, Namhoon Lee |
EMNLP | 2 |
| 2024 | Prefixing Attention Sinks can Mitigate Activation Outliers for Large Language Model QuantizationabstractDespite recent advances in LLM quantization, activation quantization remains to be challenging due to the activation outliers.Conventional remedies, e.g., mixing precisions for different channels, introduce extra overhead and reduce the speedup.In this work, we develop a simple yet effective strategy to facilitate per-tensor activation quantization by preventing the generation of problematic tokens.Precisely, we propose a method to find a set of key-value cache, coined CushionCache, which mitigates outliers in subsequent tokens when inserted as a prefix.CushionCache works in two steps: First, we greedily search for a prompt token sequence that minimizes the maximum activation values in subsequent tokens.Then, we further tune the token cache to regularize the activations of subsequent tokens to be more quantization-friendly.The proposed method successfully addresses activation outliers of LLMs, providing a substantial performance boost for per-tensor activation quantization methods.We thoroughly evaluate our method over a wide range of models and benchmarks and find that it significantly surpasses the established baseline of per-tensor W8A8 quantization and can be seamlessly integrated with the recent activation quantization method. Seungwoo Son 0005, Wonpyo Park, Woohyun Han, Kyuyeun Kim, Jaeho Lee 0001 |
EMNLP | 2 |
| 2021 | Multi-level Distance Regularization for Deep Metric LearningabstractWe propose a novel distance-based regularization method for deep metric learning called Multi-level Distance Regularization (MDR). MDR explicitly disturbs a learning procedure by regularizing pairwise distances between embedding vectors into multiple levels that represents a degree of similarity between a pair. In the training stage, the model is trained with both MDR and an existing loss function of deep metric learning, simultaneously; the two losses interfere with the objective of each other, and it makes the learning process difficult. Moreover, MDR prevents some examples from being ignored or overly influenced in the learning process. These allow the parameters of the embedding network to be settle on a local optima with better generalization. Without bells and whistles, MDR with simple Triplet loss achieves the-state-of-the-art performance in various benchmark datasets: CUB-200-2011, Cars-196, Stanford Online Products, and In-Shop Clothes Retrieval. We extensively perform ablation studies on its behaviors to show the effectiveness of MDR. By easily adopting our MDR, the previous approaches can be improved in performance and generalization ability. Wonpyo Park |
AAAI | 2 |
| 2020 | GroupFace: Learning Latent Groups and Constructing Group-Based Representations for Face RecognitionabstractIn the field of face recognition, a model learns to distinguish millions of face images with fewer dimensional embedding features, and such vast information may not be properly encoded in the conventional model with a single branch. We propose a novel face-recognition-specialized architecture called GroupFace that utilizes multiple group-aware representations, simultaneously, to improve the quality of the embedding feature. The proposed method provides self-distributed labels that balance the number of samples belonging to each group without additional human annotations, and learns the group-aware representations that can narrow down the search space of the target identity. We prove the effectiveness of the proposed method by showing extensive ablation studies and visualizations. All the components of the proposed method can be trained in an end-to-end manner with a marginal increase of computational complexity. Finally, the proposed method achieves the state-of-the-art results with significant improvements in 1:1 face verification and 1:N face identification tasks on the following public datasets: LFW, YTF, CALFW, CPLFW, CFP, AgeDB-30, MegaFace, IJB-B and IJB-C. Wonpyo Park, Myung-Cheol Roh, Jongju Shin |
CVPR | 2 |
| 2020 | BroadFace: Looking at Tens of Thousands of People at once for Face Recognition
Wonpyo Park, Jongju Shin |
ECCV (9) | 2 |
| 2019 | Regularizing Neural Networks via Stochastic Branch LayersabstractWe introduce a novel stochastic regularization technique for deep neural networks, which decomposes a layer into multiple branches with different parameters and merges stochastically sampled combinations of the outputs from the branches during training. Since the factorized branches can collapse into a single branch through a linear operation, inference requires no additional complexity compared to the ordinary layers. The proposed regularization method, referred to as StochasticBranch, is applicable to any linear layers such as fully-connected or convolution layers. The proposed regularizer allows the model to explore diverse regions of the model parameter space via multiple combinations of branches to find better local minima. An extensive set of experiments shows that our method effectively regularizes networks and further improves the generalization performance when used together with other existing regularization techniques. Wonpyo Park, Hongsuck Seo, Bohyung Han, Minsu Cho |
ACML | 1 |
| 2019 | Relational Knowledge DistillationabstractKnowledge distillation aims at transferring knowledge acquired in one model (a teacher) to another model (a student) that is typically smaller. Previous approaches can be expressed as a form of training the student to mimic output activations of individual data examples represented by the teacher. We introduce a novel approach, dubbed relational knowledge distillation (RKD), that transfers mutual relations of data examples instead. For concrete realizations of RKD, we propose distance-wise and angle-wise distillation losses that penalize structural differences in relations. Experiments conducted on different tasks show that the proposed method improves educated student models with a significant margin. In particular for metric learning, it allows students to outperform their teachers' performance, achieving the state of the arts on standard benchmark datasets. Wonpyo Park, Dongju Kim, Minsu Cho |
CVPR | 1 |