Hanseul Cho 0002

dblp:233/5755-2 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
8since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Deep learning architectures and training · 26% Reinforcement learning · 14% Learning theory · 13%
Theoretical computer science
4 papers
Mathematical optimization · 54% Computational complexity · 46%

Topics — the 23 heaviest of 26, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computational complexity › circuit complexity
transformer expressivity
1.622025
Arithmetic Transformers Can Length-Generalize in Both Operand Length and Count · ICLR 2025
Position Coupling: Improving Length Generalization of Arithmetic Transformers Using Task Structure · NeurIPS 2024
Machine learning › Deep learning architectures and training › training dynamics
plasticity loss
1.422024
DASH: Warm-Starting Neural Network Training in Stationary Settings without Loss of Plasticity · NeurIPS 2024
PLASTIC: Improving Input and Label Plasticity for Sample Efficient Reinforcement Learning · NeurIPS 2023
Mathematical optimization
minimax optimization
1.422024
Fundamental Benefit of Alternating Updates in Minimax Optimization · ICML 2024
SGDA with shuffling: faster convergence for nonconvex-PŁ minimax optimization · ICLR 2023
Machine learning › Learning paradigms › continual learning
catastrophic forgetting
0.912025
Convergence and Implicit Bias of Gradient Descent on Continual Linear Classification · ICLR 2025
Machine learning › Learning paradigms
continual learning
0.912025
Convergence and Implicit Bias of Gradient Descent on Continual Linear Classification · ICLR 2025
Machine learning › Optimization for machine learning › convergence guarantees
gradient descent convergence
0.912025
Convergence and Implicit Bias of Gradient Descent on Continual Linear Classification · ICLR 2025
Machine learning › Learning theory
implicit bias
0.912025
Convergence and Implicit Bias of Gradient Descent on Continual Linear Classification · ICLR 2025
Natural language and speech › Language models and text generation › compositional generalization
length generalization
0.912025
Arithmetic Transformers Can Length-Generalize in Both Operand Length and Count · ICLR 2025
Machine learning › Deep learning architectures and training
positional encoding
0.812024
Position Coupling: Improving Length Generalization of Arithmetic Transformers Using Task Structure · NeurIPS 2024
Machine learning › Deep learning architectures and training
transformer
0.812024
Position Coupling: Improving Length Generalization of Arithmetic Transformers Using Task Structure · NeurIPS 2024
Robotics › Motion planning and robot control
warm-starting
0.812024
DASH: Warm-Starting Neural Network Training in Stationary Settings without Loss of Plasticity · NeurIPS 2024
Computational complexity
circuit complexity
0.812024
Position Coupling: Improving Length Generalization of Arithmetic Transformers Using Task Structure · NeurIPS 2024
Mathematical optimization › minimax optimization
gradient descent ascent
0.812024
Fundamental Benefit of Alternating Updates in Minimax Optimization · ICML 2024
Machine learning › Trustworthy machine learning
fairness
0.712023
Fair Streaming Principal Component Analysis: Statistical and Algorithmic Viewpoint · NeurIPS 2023
Machine learning › Trustworthy machine learning › fairness › fair unsupervised learning
fair principal component analysis
0.712023
Fair Streaming Principal Component Analysis: Statistical and Algorithmic Viewpoint · NeurIPS 2023
Machine learning › Efficient and distributed learning
memory-efficient training
0.712023
Fair Streaming Principal Component Analysis: Statistical and Algorithmic Viewpoint · NeurIPS 2023
Machine learning › Reinforcement learning
off-policy reinforcement learning
0.712023
PLASTIC: Improving Input and Label Plasticity for Sample Efficient Reinforcement Learning · NeurIPS 2023
Machine learning › Reinforcement learning
plasticity preservation
0.712023
PLASTIC: Improving Input and Label Plasticity for Sample Efficient Reinforcement Learning · NeurIPS 2023
Machine learning › Reinforcement learning
sample efficiency
0.712023
PLASTIC: Improving Input and Label Plasticity for Sample Efficient Reinforcement Learning · NeurIPS 2023
Machine learning › Time series and sequential data
streaming data
0.712023
Fair Streaming Principal Component Analysis: Statistical and Algorithmic Viewpoint · NeurIPS 2023
Mathematical optimization › stochastic optimization › stochastic gradient methods
stochastic gradient descent
0.712023
SGDA with shuffling: faster convergence for nonconvex-PŁ minimax optimization · ICLR 2023
Machine learning › Deep learning architectures and training
loss landscape
0.212023
PLASTIC: Improving Input and Label Plasticity for Sample Efficient Reinforcement Learning · NeurIPS 2023
Machine learning › Learning theory
PAC learning
0.212023
Fair Streaming Principal Component Analysis: Statistical and Algorithmic Viewpoint · NeurIPS 2023

Methods — techniques the papers use, named apart from their topics

scratchpad · 1.7position coupling · 1.7multi-level position coupling · 1.7positional encoding · 1.5length generalization · 1.5non-asymptotic analysis · 0.9gradient descent analysis · 0.9selective forgetting · 0.8iteration complexity · 0.8direction-aware shrinking · 0.8convergence analysis · 0.8gradient propagation · 0.7
YearPublicationVenuePosition
2025 Arithmetic Transformers Can Length-Generalize in Both Operand Length and Count
abstract
Transformers often struggle with *length generalization*, meaning they fail to generalize to sequences longer than those encountered during training. While arithmetic tasks are commonly used to study length generalization, certain tasks are considered notoriously difficult, e.g., multi-operand addition (requiring generalization over both the number of operands and their lengths) and multiplication (requiring generalization over both operand lengths). In this work, we achieve approximately 2–3× length generalization on both tasks, which is the first such achievement in arithmetic Transformers. We design task-specific scratchpads enabling the model to focus on a fixed number of tokens per each next-token prediction step, and apply multi-level versions of *Position Coupling* (Cho et al., 2024; McLeish et al., 2024) to let Transformers know the right position to attend to. On the theory side, we prove that a 1-layer Transformer using our method can solve multi-operand addition, up to operand length and operand count that are exponential in embedding dimension.
Hanseul Cho 0002, Jaeyoung Cha, Srinadh Bhojanapalli, Chulhee Yun
ICLR1
2025 Convergence and Implicit Bias of Gradient Descent on Continual Linear Classification
abstract
We study continual learning on multiple linear classification tasks by sequentially running gradient descent (GD) for a fixed budget of iterations per each given task. When all tasks are jointly linearly separable and are presented in a cyclic/random order, we show the directional convergence of the trained linear classifier to the joint (offline) max-margin solution. This is surprising because GD training on a single task is implicitly biased towards the individual max-margin solution for the task, and the direction of the joint max-margin solution can be largely different from these individual solutions. Additionally, when tasks are given in a cyclic order, we present a non-asymptotic analysis on cycle-averaged forgetting, revealing that (1) alignment between tasks is indeed closely tied to catastrophic forgetting and backward knowledge transfer and (2) the amount of forgetting vanishes to zero as the cycle repeats. Lastly, we analyze the case where the tasks are no longer jointly separable and show that the model trained in a cyclic order converges to the unique minimum of the joint loss function.
Hyunji Jung, Hanseul Cho 0002, Chulhee Yun
ICLR2
2024 Fundamental Benefit of Alternating Updates in Minimax Optimization
abstract
The Gradient Descent-Ascent (GDA) algorithm, designed to solve minimax optimization problems, takes the descent and ascent steps either simultaneously (Sim-GDA) or alternately (Alt-GDA). While Alt-GDA is commonly observed to converge faster, the performance gap between the two is not yet well understood theoretically, especially in terms of global convergence rates. To address this theory-practice gap, we present fine-grained convergence analyses of both algorithms for strongly-convex-strongly-concave and Lipschitz-gradient objectives. Our new iteration complexity upper bound of Alt-GDA is strictly smaller than the lower bound of Sim-GDA; i.e., Alt-GDA is provably faster. Moreover, we propose Alternating-Extrapolation GDA (Alex-GDA), a general algorithmic framework that subsumes Sim-GDA and Alt-GDA, for which the main idea is to alternately take gradients from extrapolations of the iterates. We show that Alex-GDA satisfies a smaller iteration complexity bound, identical to that of the Extra-gradient method, while requiring less gradient computations. We also prove that Alex-GDA enjoys linear convergence for bilinear problems, for which both Sim-GDA and Alt-GDA fail to converge at all.
Jaewook Lee 0009, Hanseul Cho 0002, Chulhee Yun
ICML2
2024 Position Coupling: Improving Length Generalization of Arithmetic Transformers Using Task Structure
abstract
Even for simple arithmetic tasks like integer addition, it is challenging for Transformers to generalize to longer sequences than those encountered during training. To tackle this problem, we propose *position coupling*, a simple yet effective method that directly embeds the structure of the tasks into the positional encoding of a (decoder-only) Transformer. Taking a departure from the vanilla absolute position mechanism assigning unique position IDs to each of the tokens, we assign the same position IDs to two or more "relevant" tokens; for integer addition tasks, we regard digits of the same significance as in the same position. On the empirical side, we show that with the proposed position coupling, our models trained on 1 to 30-digit additions can generalize up to *200-digit* additions (6.67x of the trained length). On the theoretical side, we prove that a 1-layer Transformer with coupled positions can solve the addition task involving exponentially many digits, whereas any 1-layer Transformer without positional information cannot entirely solve it. We also demonstrate that position coupling can be applied to other algorithmic tasks such as Nx2 multiplication and a two-dimensional task. Our codebase is available at [github.com/HanseulJo/position-coupling](https://github.com/HanseulJo/position-coupling).
Hanseul Cho 0002, Jaeyoung Cha, Pranjal Awasthi, Srinadh Bhojanapalli, Anupam Gupta 0001, Chulhee Yun
NeurIPS1
2024 DASH: Warm-Starting Neural Network Training in Stationary Settings without Loss of Plasticity
abstract
Warm-starting neural network training by initializing networks with previously learned weights is appealing, as practical neural networks are often deployed under a continuous influx of new data. However, it often leads to *loss of plasticity*, where the network loses its ability to learn new information, resulting in worse generalization than training from scratch. This occurs even under stationary data distributions, and its underlying mechanism is poorly understood. We develop a framework emulating real-world neural network training and identify noise memorization as the primary cause of plasticity loss when warm-starting on stationary data. Motivated by this, we propose **Direction-Aware SHrinking (DASH)**, a method aiming to mitigate plasticity loss by selectively forgetting memorized noise while preserving learned features. We validate our approach on vision tasks, demonstrating improvements in test accuracy and training efficiency.
Baekrok Shin, Junsoo Oh, Hanseul Cho 0002, Chulhee Yun
NeurIPS3
2023 SGDA with shuffling: faster convergence for nonconvex-PŁ minimax optimization
Hanseul Cho 0002, Chulhee Yun
ICLR1
2023 PLASTIC: Improving Input and Label Plasticity for Sample Efficient Reinforcement Learning
abstract
In Reinforcement Learning (RL), enhancing sample efficiency is crucial, particularly in scenarios when data acquisition is costly and risky. In principle, off-policy RL algorithms can improve sample efficiency by allowing multiple updates per environment interaction. However, these multiple updates often lead the model to overfit to earlier interactions, which is referred to as the loss of plasticity. Our study investigates the underlying causes of this phenomenon by dividing plasticity into two aspects. Input plasticity, which denotes the model's adaptability to changing input data, and label plasticity, which denotes the model's adaptability to evolving input-output relationships. Synthetic experiments on the CIFAR-10 dataset reveal that finding smoother minima of loss landscape enhances input plasticity, whereas refined gradient propagation improves label plasticity. Leveraging these findings, we introduce the **PLASTIC** algorithm, which harmoniously combines techniques to address both concerns. With minimal architectural modifications, PLASTIC achieves competitive performance on benchmarks including Atari-100k and Deepmind Control Suite. This result emphasizes the importance of preserving the model's plasticity to elevate the sample efficiency in RL. The code is available at https://github.com/dojeon-ai/plastic.
Hanseul Cho 0002, Hyunseung Kim, Daehoon Gwak, Joonkee Kim, Jaegul Choo, Se-Young Yun, Chulhee Yun
NeurIPS2
2023 Fair Streaming Principal Component Analysis: Statistical and Algorithmic Viewpoint
abstract
Fair Principal Component Analysis (PCA) is a problem setting where we aim to perform PCA while making the resulting representation fair in that the projected distributions, conditional on the sensitive attributes, match one another. However, existing approaches to fair PCA have two main problems: theoretically, there has been no statistical foundation of fair PCA in terms of learnability; practically, limited memory prevents us from using existing approaches, as they explicitly rely on full access to the entire data. On the theoretical side, we rigorously formulate fair PCA using a new notion called probably approximately fair and optimal (PAFO) learnability. On the practical side, motivated by recent advances in streaming algorithms for addressing memory limitation, we propose a new setting called fair streaming PCA along with a memory-efficient algorithm, fair noisy power method (FNPM). We then provide its statistical guarantee in terms of PAFO-learnability, which is the first of its kind in fair PCA literature. We verify our algorithm in the CelebA dataset without any pre-processing; while the existing approaches are inapplicable due to memory limitations, by turning it into a streaming setting, we show that our algorithm performs fair PCA efficiently and effectively.
Hanseul Cho 0002, Se-Young Yun, Chulhee Yun
NeurIPS2