VLDB 2026 Research / reviewers in the wild / expert
Yujin Song
dblp:33/7664
· DBLP profile ↗
9ranked-venue papers
2as first author
7since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Theory of computation · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Nonlinear transformers can perform inference-time feature learningabstractPretrained transformers have demonstrated the ability to implement various algorithms at inference time without parameter updates. While theoretical works have established this capability through constructions and approximation guarantees, the optimization and statistical efficiency aspects remain understudied. In this work, we investigate how transformers learn features in-context – a key mechanism underlying their inference-time adaptivity. We focus on the in-context learning of single-index models $y=\sigma_*(⟨\\boldsymbol{x},\\boldsymbol{\beta}⟩)$, which are low-dimensional nonlinear functions parameterized by feature vector $\\boldsymbol\beta$. We prove that transformers pretrained by gradient-based optimization can perform inference-time feature learning, i.e., extract information of the target features $\\boldsymbol{\beta}$ solely from test prompts (despite $\\boldsymbol {\beta}$ varying across different prompts), hence achieving an in-context statistical efficiency that surpasses any non-adaptive (fixed-basis) algorithms such as kernel methods. Moreover, we show that the inference-time sample complexity surpasses the Correlational Statistical Query (CSQ) lower bound, owing to nonlinear label transformations naturally induced by the Softmax self-attention mechanism. Naoki Nishikawa, Yujin Song, Kazusato Oko, Denny Wu, Taiji Suzuki |
ICML | 2 |
| 2025 | How Does Label Noise Gradient Descent Improve Generalization in the Low SNR Regime?abstractThe capacity of deep learning models is often large enough to both learn the underlying statistical signal and overfit to noise in the training set. This noise memorization can be harmful especially for data with a low signal-to-noise ratio (SNR), leading to poor generalization. Inspired by prior observations that label noise provides implicit regularization that improves generalization, in this work, we investigate whether introducing label noise to the gradient updates can enhance the test performance of neural network (NN) in the low SNR regime. Specifically, we consider training a two-layer NN with a simple label noise gradient descent (GD) algorithm, in an idealized signal-noise data setting. We prove that adding label noise during training suppresses noise memorization, preventing it from dominating the learning process; consequently, label noise GD enjoys rapid signal growth while the overfitting remains controlled, thereby achieving good generalization despite the low SNR. In contrast, we also show that NN trained with standard GD tends to overfit to noise in the same low SNR setting and establish a non-vanishing lower bound on its test error, thus demonstrating the benefit of introducing label noise in gradient-based training. Wei Huang 0034, Andi Han, Yujin Song, Yilan Chen 0002, Denny Wu, Difan Zou, Taiji Suzuki |
NeurIPS | 3 |
| 2025 | From Shortcut to Induction Head: How Data Diversity Shapes Algorithm Selection in TransformersabstractTransformers can implement both generalizable algorithms (e.g., induction heads) and simple positional shortcuts (e.g., memorizing fixed output positions). In this work, we study how the choice of pretraining data distribution steers a shallow transformer toward one behavior or the other. Focusing on a minimal trigger-output prediction task -- copying the token immediately following a special trigger upon its second occurrence -- we present a rigorous analysis of gradient-based training of a single-layer transformer.
In both the infinite and finite sample regimes, we prove a transition in the learned mechanism: if input sequences exhibit sufficient diversity, measured by a low “max-sum” ratio of trigger-to-trigger distances, the trained model implements an induction head and generalizes to unseen contexts; by contrast, when this ratio is large, the model resorts to a positional shortcut and fails to generalize out-of-distribution (OOD).
We also reveal a trade-off between the pretraining context length and OOD generalization, and derive the optimal pretraining distribution that minimizes computational cost per sample.
Finally, we validate our theoretical predictions with controlled synthetic experiments, demonstrating that broadening context distributions robustly induces induction heads and enables OOD generalization.
Our results shed light on the algorithmic biases of pretrained transformers and offer conceptual guidelines for data-driven control of their learned behaviors. Ryotaro Kawata, Yujin Song, Alberto Bietti, Naoki Nishikawa, Taiji Suzuki, Samuel Vaiter, Denny Wu |
NeurIPS | 2 |
| 2024 | Learning sum of diverse features: computational hardness and efficient gradient-based training for ridge combinationsabstractWe study the statistical and computational complexity of learning a target function $f_*:\R^d\to\R$ with \textit{additive structure}, that is, $f_*(x) = \frac{1}{\sqrt{M}}\sum_{m=1}^M f_m(⟨x, v_m⟩)$, where $f_1,f_2,...,f_M:\R\to\R$ are nonlinear link functions of single-index models (ridge functions) with diverse and near-orthogonal index features $\{v_m\}_{m=1}^M$, and the number of additive tasks $M$ grows with the dimensionality $M\asymp d^\gamma$ for $\gamma\ge 0$. This problem setting is motivated by the classical additive model literature, the recent representation learning theory of two-layer neural network, and large-scale pretraining where the model simultaneously acquires a large number of “skills” that are often \textit{localized} in distinct parts of the trained network. We prove that a large subset of polynomial $f_*$ can be efficiently learned by gradient descent training of a two-layer neural network, with a polynomial statistical and computational complexity that depends on the number of tasks $M$ and the \textit{information exponent} of $f_m$, despite the unknown link function and $M$ growing with the dimensionality. We complement this learnability guarantee with computational hardness result by establishing statistical query (SQ) lower bounds for both the correlational SQ and full SQ algorithms. Kazusato Oko, Yujin Song, Taiji Suzuki, Denny Wu |
COLT | 2 |
| 2024 | Structural Preprocessing Method for Nonlinear Differential-Algebraic Equations Using Linear Symbolic MatricesabstractDifferential-algebraic equations (DAEs) have been used in modeling various dynamical systems in science and engineering. There are several preprocessing methods that are needed before performing numerical simulations for DAEs, such as consistent initialization and index reduction. Preprocessing methods that use structural information on DAEs run fast and are widely used. Unfortunately, structural preprocessing methods may fail when the system Jacobian, which is a functional matrix, derived from the DAE is singular. Taihei Oki, Yujin Song |
ISSAC | 2 |
| 2024 | Pretrained Transformer Efficiently Learns Low-Dimensional Target Functions In-ContextabstractTransformers can efficiently learn in-context from example demonstrations. Most existing theoretical analyses studied the in-context learning (ICL) ability of transformers for linear function classes, where it is typically shown that the minimizer of the pretraining loss implements one gradient descent step on the least squares objective. However, this simplified linear setting arguably does not demonstrate the statistical efficiency of ICL, since the pretrained transformer does not outperform directly solving linear regression on the test prompt.
In this paper, we study ICL of a nonlinear function class via transformer with nonlinear MLP layer: given a class of \textit{single-index} target functions $f_*(\boldsymbol{x}) = \sigma_*(\langle\boldsymbol{x},\boldsymbol{\beta}\rangle)$, where the index features $\boldsymbol{\beta}\in\mathbb{R}^d$ are drawn from a $r$-dimensional subspace, we show that a nonlinear transformer optimized by gradient descent (with a pretraining sample complexity that depends on the \textit{information exponent} of the link functions $\sigma_*$) learns $f_*$ in-context with a prompt length that only depends on the dimension of the distribution of target functions $r$; in contrast, any algorithm that directly learns $f_*$ on test prompt yields a statistical complexity that scales with the ambient dimension $d$. Our result highlights the adaptivity of the pretrained transformer to low-dimensional structures of the function class, which enables sample-efficient ICL that outperforms estimators that only have access to the in-context data. Kazusato Oko, Yujin Song, Taiji Suzuki, Denny Wu |
NeurIPS | 2 |
| 2021 | Searching for Mathematical Formulas Based on Graph Representation Learning
Yujin Song |
CICM | 1 |
| 2020 | A Model Predictive Control Method for Two Induction Motor Drives Supplied by Four-leg InverterabstractThe four-leg inverter is potential for non-redundant fault-tolerant control to solve multi-leg faults in two-motor drives. This paper proposes a non-redundant fault-tolerant control scheme based on the four-leg inverter. When there are open-circuit faults in both inverters, the fault phases of both motors are connected to the neutral point of DC-link capacitors. Obviously, this will cause more severe capacitor voltages offset and fluctuation than only one motor phase is connected to the neutral point. This paper proposes a fault-tolerant control algorithm based on predictive torque control method to solve this shortcoming. The balanced three-phase current and stable output torque and stator flux are guaranteed. The capacitor voltages offset is suppressed, improving the DC-link voltage utilization. Besides, the algorithm independently calculates the cost function for each motor, eliminating the coupling problem in voltage suppression. It reduces the calculation time and can be applied to multi-motor systems. The effectiveness of the algorithm is verified by experiment results. Yujin Song |
IECON | 1 |
| 2012 | Digital Current Sharing Method for Parallel Interleaved DC-DC Converters Using Input Ripple VoltageabstractThis paper describes a new digital current sharing method for parallel interleaved dc-dc converters. The difference between sensed inductor current values can cause unequal current distribution in the parallel dc-dc converters even though they are controlled by the same average current mode controller with a common current reference. To overcome this problem, a simple digital current sharing algorithm using input voltage ripple difference is proposed. The digital current sharing algorithm calculates the required current reference adjustment value using only the input ripple voltage difference information. No additional input current sensing circuit is required. The proposed algorithm achieves equal current distribution between phases of the parallel converters with high parameter insensitivity. For the experimental verification, a 500 W two phase interleaved synchronous buck dc-dc converter which charges a 24 V Li-ion battery is implemented. Suyong Chae, Yujin Song, Sukin Park, Hakgeun Jeong |
IEEE Trans. Ind. Informatics | 2 |