VLDB 2026 Research / reviewers in the wild / expert
Sungyoon Lee
dblp:151/5889
· DBLP profile ↗
12ranked-venue papers
7as first author
11since 2021 · last 2025
0000-0002-6097-0565ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 7 first-author · 11 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Variance Sensitivity Induces Attention Entropy Collapse and Instability in TransformersabstractAttention-based language models commonly rely on the softmax function to convert attention logits into probability distributions.However, this softmax re-weighting can lead to attention entropy collapse, in which attention disproportionately concentrates on a single token, ultimately causing training instability.In this work, we identify the high variance sensitivity of softmax as a primary cause of this collapse.We show that entropy-stable attention methods, which either control or are insensitive to the variance of attention logits, can prevent entropy collapse and enable more stable training.We provide empirical evidence of this effect in both large language models (LLMs) and a small Transformer model composed solely of self-attention and support our findings with theoretical analysis.Moreover, we identify that the concentration of attention probabilities increases the probability matrix norm, leading to the gradient exploding. Jonghyun Hong, Sungyoon Lee |
EMNLP | 2 |
| 2025 | Prediction Risk and Estimation Risk of the Ridgeless Least Squares Estimator under General Assumptions on Regression ErrorsabstractIn recent years, there has been a significant growth in research focusing on minimum $\ell_2$ norm (ridgeless) interpolation least squares estimators. However, the majority of these analyses have been limited to an unrealistic regression error structure, assuming independent and identically distributed errors with zero mean and common variance. In this paper, we explore prediction risk as well as estimation risk under more general regression error assumptions, highlighting the benefits of overparameterization in a more realistic setting that allows for clustered or serial dependence. Notably, we establish that the estimation difficulties associated with the variance components of both risks can be summarized through the trace of the variance-covariance matrix of the regression errors. Our findings suggest that the benefits of overparameterization can extend to time series, panel and grouped data. Sungyoon Lee, Sokbae Lee |
ICLR | 1 |
| 2025 | Prior Forgetting and In-Context OverfittingabstractIn-context learning (ICL) is one of the key capabilities contributing to the great success of LLMs. At test time, ICL is known to operate in the two modes: task recognition and task learning. In this paper, we investigate the emergence and dynamics of the two modes of ICL during pretraining. To provide an analytical understanding of the learning dynamics of the ICL abilities, we investigate the in-context random linear regression problem with a simple linear-attention-based transformer, and define and disentangle the strengths of the task recognition and task learning abilities stored in the transformer model’s parameters. We show that, during the pretraining phase, the model first learns the task learning and the task recognition abilities together in the beginning, but it (a) gradually forgets the task recognition ability to recall the priorly learned tasks and (b) relies more on the given context in the later phase, which we call (a) \textit{prior forgetting} and (b) \textit{in-context overfitting}, respectively. Sungyoon Lee |
NeurIPS | 1 |
| 2025 | How Classifier Features Transfer to Downstream: An Asymptotic Analysis in a Two-Layer ModelabstractNeural networks learn effective feature representations, which can be transferred to new tasks without additional training.
While larger datasets are known to improve feature transfer, the theoretical conditions for the success of such transfer remain unclear.
This work investigates feature transfer in networks trained for classification to identify the conditions that enable effective clustering in unseen classes.
We first reveal that higher similarity between training and unseen distributions leads to improved Cohesion and Separability.
We then show that feature expressiveness is enhanced when inputs are similar to the training classes, while the features of irrelevant inputs remain indistinguishable.
We validate our analysis on synthetic and benchmark datasets, including CAR, CUB, SOP, ISC, and ImageNet.
Our analysis highlights the importance of the similarity between training classes and the input distribution for successful feature transfer. Hee Bin Yoo, Sungyoon Lee, Cheongjae Jang, Dong-Sig Han, Jaein Kim 0004, Seunghyeon Lim, Byoung-Tak Zhang |
NeurIPS | 2 |
| 2023 | A new characterization of the edge of stability based on a sharpness measure aware of batch gradient distribution
Sungyoon Lee, Cheongjae Jang |
ICLR | 1 |
| 2023 | Implicit Jacobian regularization weighted with impurity of probability outputabstractThe success of deep learning is greatly attributed to stochastic gradient descent (SGD), yet it remains unclear how SGD finds well-generalized models. We demonstrate that SGD has an implicit regularization effect on the logit-weight Jacobian norm of neural networks. This regularization effect is weighted with the *impurity* of the probability output, and thus it is active in a certain phase of training. Moreover, based on these findings, we propose a novel optimization method that explicitly regularizes the Jacobian norm, which leads to similar performance as other state-of-the-art sharpness-aware optimization methods. Sungyoon Lee, Jinseong Park 0001, Jaewook Lee 0001 |
ICML | 1 |
| 2023 | Bridged adversarial training
Hoki Kim, Sungyoon Lee, Jaewook Lee 0001 |
Neural Networks | 3 |
| 2023 | GradDiv: Adversarial Robustness of Randomized Neural Networks via Gradient Diversity RegularizationabstractDeep learning is vulnerable to adversarial examples. Many defenses based on randomized neural networks have been proposed to solve the problem, but fail to achieve robustness against attacks using proxy gradients such as the Expectation over Transformation (EOT) attack. We investigate the effect of the adversarial attacks using proxy gradients on randomized neural networks and demonstrate that it highly relies on the directional distribution of the loss gradients of the randomized neural network. We show in particular that proxy gradients are less effective when the gradients are more scattered. To this end, we propose Gradient Diversity (GradDiv) regularizations that minimize the concentration of the gradients to build a robust randomized neural network. Our experiments on MNIST, CIFAR10, and STL10 show that our proposed GradDiv regularizations improve the adversarial robustness of randomized neural networks against a variety of state-of-the-art attack methods. Moreover, our method efficiently reduces the transferability among sample models of randomized neural networks. Sungyoon Lee, Hoki Kim, Jaewook Lee 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | A Reparametrization-Invariant Sharpness Measure Based on Information GeometryabstractIt has been observed that the generalization performance of neural networks correlates with the sharpness of their loss landscape. Dinh et al. (2017) have observed that existing formulations of sharpness measures fail to be invariant with respect to scaling and reparametrization. While some scale-invariant measures have recently been proposed, reparametrization-invariant measures are still lacking. Moreover, they often do not provide any theoretical insights into generalization performance nor lead to practical use to improve the performance. Based on an information geometric analysis of the neural network parameter space, in this paper we propose a reparametrization-invariant sharpness measure that captures the change in loss with respect to changes in the probability distribution modeled by neural networks, rather than with respect to changes in the parameter values. We reveal some theoretical connections of our measure to generalization performance. In particular, experiments confirm that using our measure as a regularizer in neural network training significantly improves performance. Cheongjae Jang, Sungyoon Lee, Frank C. Park 0001, Yung-Kyun Noh |
NeurIPS | 2 |
| 2022 | Variational cycle-consistent imputation adversarial networks for general missing patterns
Sungyoon Lee, Junyoung Byun, Hoki Kim, Jaewook Lee 0001 |
Pattern Recognit. | 2 |
| 2021 | Towards Better Understanding of Training Certifiably Robust Models against Adversarial ExamplesabstractWe study the problem of training certifiably robust models against adversarial examples. Certifiable training minimizes an upper bound on the worst-case loss over the allowed perturbation, and thus the tightness of the upper bound is an important factor in building certifiably robust models. However, many studies have shown that Interval Bound Propagation (IBP) training uses much looser bounds but outperforms other models that use tighter bounds. We identify another key factor that influences the performance of certifiable training: \textit{smoothness of the loss landscape}. We find significant differences in the loss landscapes across many linear relaxation-based methods, and that the current state-of-the-arts method often has a landscape with favorable optimization properties. Moreover, to test the claim, we design a new certifiable training method with the desired properties. With the tightness and the smoothness, the proposed method achieves a decent performance under a wide range of perturbations, while others with only one of the two factors can perform well only for a specific range of perturbations. Our code is available at \url{https://github.com/sungyoon-lee/LossLandscapeMatters}. Sungyoon Lee, Jinseong Park 0001, Jaewook Lee 0001 |
NeurIPS | 1 |
| 2020 | Lipschitz-Certifiable Training with a Tight Outer BoundabstractVerifiable training is a promising research direction for training a robust network. However, most verifiable training methods are slow or lack scalability. In this study, we propose a fast and scalable certifiable training algorithm based on Lipschitz analysis and interval arithmetic. Our certifiable training algorithm provides a tight propagated outer bound by introducing the box constraint propagation (BCP), and it efficiently computes the worst logit over the outer bound. In the experiments, we show that BCP achieves a tighter outer bound than the global Lipschitz-based outer bound. Moreover, our certifiable training algorithm is over 12 times faster than the state-of-the-art dual relaxation-based method; however, it achieves comparable or better verification performance, improving natural accuracy. Our fast certifiable training algorithm with the tight outer bound can scale to Tiny ImageNet with verification accuracy of 20.1\% ($\ell_2$-perturbation of $\epsilon=36/255$). Our code is available at \url{https://github.com/sungyoon-lee/bcp}. Sungyoon Lee, Jaewook Lee 0001, Saerom Park |
NeurIPS | 1 |