VLDB 2026 Research / reviewers in the wild / expert
Hanseul Cho 0002
dblp:233/5755-2
· DBLP profile ↗
8ranked-venue papers
3as first author
8since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Deep learning architectures and training · 26% Reinforcement learning · 14% Learning theory · 13% | |
| Theoretical computer science
4 papers |
Mathematical optimization · 54% Computational complexity · 46% |
Topics — the 23 heaviest of 26, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computational complexity › circuit complexity
transformer expressivity |
1.6 | 2 | 2025 | Arithmetic Transformers Can Length-Generalize in Both Operand Length and Count · ICLR 2025 Position Coupling: Improving Length Generalization of Arithmetic Transformers Using Task Structure · NeurIPS 2024 |
Machine learning › Deep learning architectures and training › training dynamics
plasticity loss |
1.4 | 2 | 2024 | DASH: Warm-Starting Neural Network Training in Stationary Settings without Loss of Plasticity · NeurIPS 2024 PLASTIC: Improving Input and Label Plasticity for Sample Efficient Reinforcement Learning · NeurIPS 2023 |
Mathematical optimization
minimax optimization |
1.4 | 2 | 2024 | Fundamental Benefit of Alternating Updates in Minimax Optimization · ICML 2024 SGDA with shuffling: faster convergence for nonconvex-PŁ minimax optimization · ICLR 2023 |
Machine learning › Learning paradigms › continual learning
catastrophic forgetting |
0.9 | 1 | 2025 | Convergence and Implicit Bias of Gradient Descent on Continual Linear Classification · ICLR 2025 |
Machine learning › Learning paradigms
continual learning |
0.9 | 1 | 2025 | Convergence and Implicit Bias of Gradient Descent on Continual Linear Classification · ICLR 2025 |
Machine learning › Optimization for machine learning › convergence guarantees
gradient descent convergence |
0.9 | 1 | 2025 | Convergence and Implicit Bias of Gradient Descent on Continual Linear Classification · ICLR 2025 |
Machine learning › Learning theory
implicit bias |
0.9 | 1 | 2025 | Convergence and Implicit Bias of Gradient Descent on Continual Linear Classification · ICLR 2025 |
Natural language and speech › Language models and text generation › compositional generalization
length generalization |
0.9 | 1 | 2025 | Arithmetic Transformers Can Length-Generalize in Both Operand Length and Count · ICLR 2025 |
Machine learning › Deep learning architectures and training
positional encoding |
0.8 | 1 | 2024 | Position Coupling: Improving Length Generalization of Arithmetic Transformers Using Task Structure · NeurIPS 2024 |
Machine learning › Deep learning architectures and training
transformer |
0.8 | 1 | 2024 | Position Coupling: Improving Length Generalization of Arithmetic Transformers Using Task Structure · NeurIPS 2024 |
Robotics › Motion planning and robot control
warm-starting |
0.8 | 1 | 2024 | DASH: Warm-Starting Neural Network Training in Stationary Settings without Loss of Plasticity · NeurIPS 2024 |
Computational complexity
circuit complexity |
0.8 | 1 | 2024 | Position Coupling: Improving Length Generalization of Arithmetic Transformers Using Task Structure · NeurIPS 2024 |
Mathematical optimization › minimax optimization
gradient descent ascent |
0.8 | 1 | 2024 | Fundamental Benefit of Alternating Updates in Minimax Optimization · ICML 2024 |
Machine learning › Trustworthy machine learning
fairness |
0.7 | 1 | 2023 | Fair Streaming Principal Component Analysis: Statistical and Algorithmic Viewpoint · NeurIPS 2023 |
Machine learning › Trustworthy machine learning › fairness › fair unsupervised learning
fair principal component analysis |
0.7 | 1 | 2023 | Fair Streaming Principal Component Analysis: Statistical and Algorithmic Viewpoint · NeurIPS 2023 |
Machine learning › Efficient and distributed learning
memory-efficient training |
0.7 | 1 | 2023 | Fair Streaming Principal Component Analysis: Statistical and Algorithmic Viewpoint · NeurIPS 2023 |
Machine learning › Reinforcement learning
off-policy reinforcement learning |
0.7 | 1 | 2023 | PLASTIC: Improving Input and Label Plasticity for Sample Efficient Reinforcement Learning · NeurIPS 2023 |
Machine learning › Reinforcement learning
plasticity preservation |
0.7 | 1 | 2023 | PLASTIC: Improving Input and Label Plasticity for Sample Efficient Reinforcement Learning · NeurIPS 2023 |
Machine learning › Reinforcement learning
sample efficiency |
0.7 | 1 | 2023 | PLASTIC: Improving Input and Label Plasticity for Sample Efficient Reinforcement Learning · NeurIPS 2023 |
Machine learning › Time series and sequential data
streaming data |
0.7 | 1 | 2023 | Fair Streaming Principal Component Analysis: Statistical and Algorithmic Viewpoint · NeurIPS 2023 |
Mathematical optimization › stochastic optimization › stochastic gradient methods
stochastic gradient descent |
0.7 | 1 | 2023 | SGDA with shuffling: faster convergence for nonconvex-PŁ minimax optimization · ICLR 2023 |
Machine learning › Deep learning architectures and training
loss landscape |
0.2 | 1 | 2023 | PLASTIC: Improving Input and Label Plasticity for Sample Efficient Reinforcement Learning · NeurIPS 2023 |
Machine learning › Learning theory
PAC learning |
0.2 | 1 | 2023 | Fair Streaming Principal Component Analysis: Statistical and Algorithmic Viewpoint · NeurIPS 2023 |
Methods — techniques the papers use, named apart from their topics
scratchpad · 1.7position coupling · 1.7multi-level position coupling · 1.7positional encoding · 1.5length generalization · 1.5non-asymptotic analysis · 0.9gradient descent analysis · 0.9selective forgetting · 0.8iteration complexity · 0.8direction-aware shrinking · 0.8convergence analysis · 0.8gradient propagation · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Arithmetic Transformers Can Length-Generalize in Both Operand Length and CountabstractTransformers often struggle with *length generalization*, meaning they fail to generalize to sequences longer than those encountered during training. While arithmetic tasks are commonly used to study length generalization, certain tasks are considered notoriously difficult, e.g., multi-operand addition (requiring generalization over both the number of operands and their lengths) and multiplication (requiring generalization over both operand lengths). In this work, we achieve approximately 2–3× length generalization on both tasks, which is the first such achievement in arithmetic Transformers. We design task-specific scratchpads enabling the model to focus on a fixed number of tokens per each next-token prediction step, and apply multi-level versions of *Position Coupling* (Cho et al., 2024; McLeish et al., 2024) to let Transformers know the right position to attend to. On the theory side, we prove that a 1-layer Transformer using our method can solve multi-operand addition, up to operand length and operand count that are exponential in embedding dimension. Hanseul Cho 0002, Jaeyoung Cha, Srinadh Bhojanapalli, Chulhee Yun |
ICLR | 1 |
| 2025 | Convergence and Implicit Bias of Gradient Descent on Continual Linear ClassificationabstractWe study continual learning on multiple linear classification tasks by sequentially running gradient descent (GD) for a fixed budget of iterations per each given task. When all tasks are jointly linearly separable and are presented in a cyclic/random order, we show the directional convergence of the trained linear classifier to the joint (offline) max-margin solution. This is surprising because GD training on a single task is implicitly biased towards the individual max-margin solution for the task, and the direction of the joint max-margin solution can be largely different from these individual solutions. Additionally, when tasks are given in a cyclic order, we present a non-asymptotic analysis on cycle-averaged forgetting, revealing that (1) alignment between tasks is indeed closely tied to catastrophic forgetting and backward knowledge transfer and (2) the amount of forgetting vanishes to zero as the cycle repeats. Lastly, we analyze the case where the tasks are no longer jointly separable and show that the model trained in a cyclic order converges to the unique minimum of the joint loss function. Hyunji Jung, Hanseul Cho 0002, Chulhee Yun |
ICLR | 2 |
| 2024 | Fundamental Benefit of Alternating Updates in Minimax OptimizationabstractThe Gradient Descent-Ascent (GDA) algorithm, designed to solve minimax optimization problems, takes the descent and ascent steps either simultaneously (Sim-GDA) or alternately (Alt-GDA). While Alt-GDA is commonly observed to converge faster, the performance gap between the two is not yet well understood theoretically, especially in terms of global convergence rates. To address this theory-practice gap, we present fine-grained convergence analyses of both algorithms for strongly-convex-strongly-concave and Lipschitz-gradient objectives. Our new iteration complexity upper bound of Alt-GDA is strictly smaller than the lower bound of Sim-GDA; i.e., Alt-GDA is provably faster. Moreover, we propose Alternating-Extrapolation GDA (Alex-GDA), a general algorithmic framework that subsumes Sim-GDA and Alt-GDA, for which the main idea is to alternately take gradients from extrapolations of the iterates. We show that Alex-GDA satisfies a smaller iteration complexity bound, identical to that of the Extra-gradient method, while requiring less gradient computations. We also prove that Alex-GDA enjoys linear convergence for bilinear problems, for which both Sim-GDA and Alt-GDA fail to converge at all. Jaewook Lee 0009, Hanseul Cho 0002, Chulhee Yun |
ICML | 2 |
| 2024 | Position Coupling: Improving Length Generalization of Arithmetic Transformers Using Task StructureabstractEven for simple arithmetic tasks like integer addition, it is challenging for Transformers to generalize to longer sequences than those encountered during training. To tackle this problem, we propose *position coupling*, a simple yet effective method that directly embeds the structure of the tasks into the positional encoding of a (decoder-only) Transformer. Taking a departure from the vanilla absolute position mechanism assigning unique position IDs to each of the tokens, we assign the same position IDs to two or more "relevant" tokens; for integer addition tasks, we regard digits of the same significance as in the same position. On the empirical side, we show that with the proposed position coupling, our models trained on 1 to 30-digit additions can generalize up to *200-digit* additions (6.67x of the trained length). On the theoretical side, we prove that a 1-layer Transformer with coupled positions can solve the addition task involving exponentially many digits, whereas any 1-layer Transformer without positional information cannot entirely solve it. We also demonstrate that position coupling can be applied to other algorithmic tasks such as Nx2 multiplication and a two-dimensional task. Our codebase is available at [github.com/HanseulJo/position-coupling](https://github.com/HanseulJo/position-coupling). Hanseul Cho 0002, Jaeyoung Cha, Pranjal Awasthi, Srinadh Bhojanapalli, Anupam Gupta 0001, Chulhee Yun |
NeurIPS | 1 |
| 2024 | DASH: Warm-Starting Neural Network Training in Stationary Settings without Loss of PlasticityabstractWarm-starting neural network training by initializing networks with previously learned weights is appealing, as practical neural networks are often deployed under a continuous influx of new data. However, it often leads to *loss of plasticity*, where the network loses its ability to learn new information, resulting in worse generalization than training from scratch. This occurs even under stationary data distributions, and its underlying mechanism is poorly understood. We develop a framework emulating real-world neural network training and identify noise memorization as the primary cause of plasticity loss when warm-starting on stationary data. Motivated by this, we propose **Direction-Aware SHrinking (DASH)**, a method aiming to mitigate plasticity loss by selectively forgetting memorized noise while preserving learned features. We validate our approach on vision tasks, demonstrating improvements in test accuracy and training efficiency. Baekrok Shin, Junsoo Oh, Hanseul Cho 0002, Chulhee Yun |
NeurIPS | 3 |
| 2023 | SGDA with shuffling: faster convergence for nonconvex-PŁ minimax optimization
Hanseul Cho 0002, Chulhee Yun |
ICLR | 1 |
| 2023 | PLASTIC: Improving Input and Label Plasticity for Sample Efficient Reinforcement LearningabstractIn Reinforcement Learning (RL), enhancing sample efficiency is crucial, particularly in scenarios when data acquisition is costly and risky. In principle, off-policy RL algorithms can improve sample efficiency by allowing multiple updates per environment interaction. However, these multiple updates often lead the model to overfit to earlier interactions, which is referred to as the loss of plasticity. Our study investigates the underlying causes of this phenomenon by dividing plasticity into two aspects. Input plasticity, which denotes the model's adaptability to changing input data, and label plasticity, which denotes the model's adaptability to evolving input-output relationships. Synthetic experiments on the CIFAR-10 dataset reveal that finding smoother minima of loss landscape enhances input plasticity, whereas refined gradient propagation improves label plasticity. Leveraging these findings, we introduce the **PLASTIC** algorithm, which harmoniously combines techniques to address both concerns. With minimal architectural modifications, PLASTIC achieves competitive performance on benchmarks including Atari-100k and Deepmind Control Suite. This result emphasizes the importance of preserving the model's plasticity to elevate the sample efficiency in RL. The code is available at https://github.com/dojeon-ai/plastic. Hanseul Cho 0002, Hyunseung Kim, Daehoon Gwak, Joonkee Kim, Jaegul Choo, Se-Young Yun, Chulhee Yun |
NeurIPS | 2 |
| 2023 | Fair Streaming Principal Component Analysis: Statistical and Algorithmic ViewpointabstractFair Principal Component Analysis (PCA) is a problem setting where we aim to perform PCA while making the resulting representation fair in that the projected distributions, conditional on the sensitive attributes, match one another. However, existing approaches to fair PCA have two main problems: theoretically, there has been no statistical foundation of fair PCA in terms of learnability; practically, limited memory prevents us from using existing approaches, as they explicitly rely on full access to the entire data. On the theoretical side, we rigorously formulate fair PCA using a new notion called probably approximately fair and optimal (PAFO) learnability. On the practical side, motivated by recent advances in streaming algorithms for addressing memory limitation, we propose a new setting called fair streaming PCA along with a memory-efficient algorithm, fair noisy power method (FNPM). We then provide its statistical guarantee in terms of PAFO-learnability, which is the first of its kind in fair PCA literature. We verify our algorithm in the CelebA dataset without any pre-processing; while the existing approaches are inapplicable due to memory limitations, by turning it into a streaming setting, we show that our algorithm performs fair PCA efficiently and effectively. Hanseul Cho 0002, Se-Young Yun, Chulhee Yun |
NeurIPS | 2 |