Shi Chen 0003

dblp:24/2311-3 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2025
0000-0003-0556-3542ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Optimization for machine learning · 40% Deep learning architectures and training · 23% Graph learning · 13%
Theoretical computer science
1 paper
Mathematical optimization · 100%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%

Topics — the 11 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Mathematical optimization › continuous optimization › convex optimization › first-order methods › gradient-based optimization
accelerated gradient methods
0.912025
Accelerating optimization over the space of probability measures · J. Mach. Learn. Res. 2025
Mathematical optimization › continuous optimization › convex optimization › first-order methods
gradient-based optimization
0.912025
Accelerating optimization over the space of probability measures · J. Mach. Learn. Res. 2025
Machine learning › Graph learning
molecular representation learning
0.712023
Learning Harmonic Molecular Representations on Riemannian Manifold · ICLR 2023
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction › manifold learning
riemannian manifold learning
0.712023
Learning Harmonic Molecular Representations on Riemannian Manifold · ICLR 2023
Bioinformatics and computational biology
geometric deep learning
0.712023
Learning Harmonic Molecular Representations on Riemannian Manifold · ICLR 2023
Bioinformatics and computational biology
molecular property prediction
0.712023
Learning Harmonic Molecular Representations on Riemannian Manifold · ICLR 2023
Machine learning › Optimization for machine learning
gradient flow
0.612022
Overparameterization of Deep ResNet: Zero Loss and Mean-field Analysis · J. Mach. Learn. Res. 2022
Machine learning › Learning theory › statistical learning theory › statistical physics of learning
mean-field analysis
0.612022
Overparameterization of Deep ResNet: Zero Loss and Mean-field Analysis · J. Mach. Learn. Res. 2022
Machine learning › Optimization for machine learning
non-convex optimization
0.612022
Overparameterization of Deep ResNet: Zero Loss and Mean-field Analysis · J. Mach. Learn. Res. 2022
Machine learning › Deep learning architectures and training
overparameterized neural network
0.612022
Overparameterization of Deep ResNet: Zero Loss and Mean-field Analysis · J. Mach. Learn. Res. 2022
Machine learning › Deep learning architectures and training › convolutional neural network
residual network
0.612022
Overparameterization of Deep ResNet: Zero Loss and Mean-field Analysis · J. Mach. Learn. Res. 2022

Methods — techniques the papers use, named apart from their topics

momentum methods · 1.7hamiltonian flow · 1.7riemannian geometry · 1.3harmonic analysis · 1.3partial differential equation analysis · 0.6mean-field analysis · 0.6gradient descent · 0.6
YearPublicationVenuePosition
2025 Accelerating optimization over the space of probability measures
abstract
The acceleration of gradient-based optimization methods is a subject of significant practical and theoretical importance, particularly within machine learning applications. While much attention has been directed towards optimizing within Euclidean space, the need to optimize over spaces of probability measures in machine learning motivates the exploration of accelerated gradient methods in this context, too. To this end, we introduce a Hamiltonian-flow approach analogous to momentum-based approaches in Euclidean space. We demonstrate that, in the continuous-time setting, algorithms based on this approach can achieve convergence rates of arbitrarily high order. We complement our findings with numerical examples.
Shi Chen 0003, Qin Li 0007, Oliver Tse, Stephen J. Wright 0001
J. Mach. Learn. Res.1
2023 Learning Harmonic Molecular Representations on Riemannian Manifold
Yuning Shen, Shi Chen 0003
ICLR3
2023 High-Frequency Limit of the Inverse Scattering Problem: Asymptotic Convergence from Inverse Helmholtz to Inverse Liouville
abstract
Abstract. We investigate the asymptotic relation between the inverse problems relying on the Helmholtz equation and the radiative transfer equation (RTE) as physical models in the high-frequency limit. In particular, we evaluate the asymptotic convergence of a generalized version of the inverse scattering problem based on the Helmholtz equation, to the inverse scattering problem of the Liouville equation (a simplified version of RTE). The two inverse problems are connected through the Wigner transform that translates the wave-type description on the physical space to the kinetic-type description on the phase space, and the Husimi transform that models data localized both in location and direction. The finding suggests that impinging tightly concentrated monochromatic beams can indeed provide stable reconstruction of the medium, asymptotically in the high-frequency regime. This fact stands in contrast with the unstable reconstruction for the classical inverse scattering problem when the probing signals are plane waves.
Shi Chen 0003, Zhiyan Ding, Qin Li 0007, Leonardo Zepeda-Núñez
SIAM J. Imaging Sci.1
2022 Overparameterization of Deep ResNet: Zero Loss and Mean-field Analysis
abstract
Finding parameters in a deep neural network (NN) that fit training data is a nonconvex optimization problem, but a basic first-order optimization method (gradient descent) finds a global optimizer with perfect fit (zero-loss) in many practical situations. We examine this phenomenon for the case of Residual Neural Networks (ResNet) with smooth activation functions in a limiting regime in which both the number of layers (depth) and the number of weights in each layer (width) go to infinity. First, we use a mean-field-limit argument to prove that the gradient descent for parameter training becomes a gradient flow for a probability distribution that is characterized by a partial differential equation (PDE) in the large-NN limit. Next, we show that under certain assumptions, the solution to the PDE converges in the training time to a zero-loss solution. Together, these results suggest that the training of the ResNet gives a near-zero loss if the ResNet is large enough. We give estimates of the depth and width needed to reduce the loss below a given threshold, with high probability.
Zhiyan Ding, Shi Chen 0003, Qin Li 0007, Stephen J. Wright 0001
J. Mach. Learn. Res.2