Yu-Han Wu

dblp:182/6365 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Optimization for machine learning · 35% Generative modeling · 26% Deep learning architectures and training · 22%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Optimization for machine learning
implicit regularization
1.622025
Taking a Big Step: Large Learning Rates in Denoising Score Matching Prevent Memorization · COLT 2025
Implicit regularization of deep residual networks towards neural ODEs · ICLR 2024
Machine learning › Generative modeling › score matching
denoising score matching
0.912025
Taking a Big Step: Large Learning Rates in Denoising Score Matching Prevent Memorization · COLT 2025
Machine learning › Generative modeling
diffusion model
0.912025
Taking a Big Step: Large Learning Rates in Denoising Score Matching Prevent Memorization · COLT 2025
Natural language and speech › Language models and text generation › large language model › knowledge in language models
memorization
0.912025
Taking a Big Step: Large Learning Rates in Denoising Score Matching Prevent Memorization · COLT 2025
Machine learning › Optimization for machine learning
gradient flow
0.812024
Implicit regularization of deep residual networks towards neural ODEs · ICLR 2024
Machine learning › Deep learning architectures and training › neural differential equations
neural ordinary differential equations
0.812024
Implicit regularization of deep residual networks towards neural ODEs · ICLR 2024
Machine learning › Deep learning architectures and training › convolutional neural network
residual network
0.812024
Implicit regularization of deep residual networks towards neural ODEs · ICLR 2024
Machine learning › Learning theory
over-parameterization
0.212024
Implicit regularization of deep residual networks towards neural ODEs · ICLR 2024

Methods — techniques the papers use, named apart from their topics

stochastic gradient descent · 0.9neural network analysis · 0.9polyak-łojasiewicz condition · 0.8gradient flow · 0.8
YearPublicationVenuePosition
2025 Taking a Big Step: Large Learning Rates in Denoising Score Matching Prevent Memorization
abstract
Denoising score matching plays a pivotal role in the performance of diffusion-based generative models. However, the empirical optimal score–the exact solution to the denoising score matching–leads to memorization, where generated samples replicate the training data. Yet, in practice, only a moderate degree of memorization is observed, even without explicit regularization. In this paper, we investigate this phenomenon by uncovering an implicit regularization mechanism driven by large learning rates. Specifically, we show that in the small-noise regime, the empirical optimal score exhibits high irregularity. We then prove that, when trained by stochastic gradient descent with a large enough learning rate, neural networks cannot stably converge to a local minimum with arbitrarily small excess risk. Consequently, the learned score cannot be arbitrarily close to the empirical optimal score, thereby mitigating memorization. To make the analysis tractable, we consider one-dimensional data and two-layer neural networks. Experiments validate the crucial role of the learning rate in preventing memorization, even beyond the one-dimensional setting.
Yu-Han Wu, Pierre Marion, Gérard Biau, Claire Boyer
COLT1
2024 Implicit regularization of deep residual networks towards neural ODEs
abstract
Residual neural networks are state-of-the-art deep learning models. Their continuous-depth analog, neural ordinary differential equations (ODEs), are also widely used. Despite their success, the link between the discrete and continuous models still lacks a solid mathematical foundation. In this article, we take a step in this direction by establishing an implicit regularization of deep residual networks towards neural ODEs, for nonlinear networks trained with gradient flow. We prove that if the network is initialized as a discretization of a neural ODE, then such a discretization holds throughout training. Our results are valid for a finite training time, and also as the training time tends to infinity provided that the network satisfies a Polyak-Łojasiewicz condition. Importantly, this condition holds for a family of residual networks where the residuals are two-layer perceptrons with an overparameterization in width that is only linear, and implies the convergence of gradient flow to a global minimum. Numerical experiments illustrate our results.
Pierre Marion, Yu-Han Wu, Michael E. Sander, Gérard Biau
ICLR2