Yuling Yan

dblp:67/6209 · DBLP profile ↗
← Back
13ranked-venue papers
3as first author
7since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 1 since 2021Systems, architecture and hardware · 2Human-computer interaction and ubiquitous computing · 1Theory of computation · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Generative modeling · 41% Reinforcement learning · 37% Learning theory · 15%
Human-computer interaction and pervasive computing
1 paper
Wearable and physiological sensing · 100%

Topics — the 16 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
2.532025
O(d/T) Convergence Theory for Diffusion Probabilistic Models under Minimal Assumptions · J. Mach. Learn. Res. 2025
O(d/T) Convergence Theory for Diffusion Probabilistic Models under Minimal Assumptions · ICLR 2025
Adapting to Unknown Low-Dimensional Structures in Score-Based Diffusion Models · NeurIPS 2024
Machine learning › Generative modeling › diffusion model › score-based generative model
denoising diffusion probabilistic model
1.722025
O(d/T) Convergence Theory for Diffusion Probabilistic Models under Minimal Assumptions · J. Mach. Learn. Res. 2025
O(d/T) Convergence Theory for Diffusion Probabilistic Models under Minimal Assumptions · ICLR 2025
Machine learning › Generative modeling › diffusion model
score-based generative model
1.622025
O(d/T) Convergence Theory for Diffusion Probabilistic Models under Minimal Assumptions · J. Mach. Learn. Res. 2025
Adapting to Unknown Low-Dimensional Structures in Score-Based Diffusion Models · NeurIPS 2024
Machine learning › Learning theory
sample complexity
1.532024
Minimax-optimal reward-agnostic exploration in reinforcement learning · COLT 2024
Sample-Efficient Reinforcement Learning for Linearly-Parameterized MDPs with a Generative Model · NeurIPS 2021
The Efficacy of Pessimism in Asynchronous Q-Learning · IEEE Trans. Inf. Theory 2023
Machine learning › Reinforcement learning › value-based reinforcement learning
q-learning
1.222023
The Efficacy of Pessimism in Asynchronous Q-Learning · IEEE Trans. Inf. Theory 2023
Sample-Efficient Reinforcement Learning for Linearly-Parameterized MDPs with a Generative Model · NeurIPS 2021
Machine learning › Optimization for machine learning
convergence analysis
0.912025
O(d/T) Convergence Theory for Diffusion Probabilistic Models under Minimal Assumptions · ICLR 2025
Machine learning › Reinforcement learning › reinforcement learning theory
convergence theory
0.912025
O(d/T) Convergence Theory for Diffusion Probabilistic Models under Minimal Assumptions · J. Mach. Learn. Res. 2025
Machine learning › Reinforcement learning
exploration
0.812024
Minimax-optimal reward-agnostic exploration in reinforcement learning · COLT 2024
Machine learning › Reinforcement learning › exploration › exploration in markov decision processes
reward-free exploration
0.812024
Minimax-optimal reward-agnostic exploration in reinforcement learning · COLT 2024
Machine learning › Reinforcement learning › value-based reinforcement learning › q-learning
asynchronous q-learning
0.712023
The Efficacy of Pessimism in Asynchronous Q-Learning · IEEE Trans. Inf. Theory 2023
Machine learning › Reinforcement learning › offline reinforcement learning
pessimistic q-learning
0.712023
The Efficacy of Pessimism in Asynchronous Q-Learning · IEEE Trans. Inf. Theory 2023
Machine learning › Reinforcement learning
markov decision process
0.512021
Sample-Efficient Reinforcement Learning for Linearly-Parameterized MDPs with a Generative Model · NeurIPS 2021
Data mining
clustering
0.412020
Efficient Clustering for Stretched Mixtures: Landscape and Optimality · NeurIPS 2020
Wearable and physiological sensing
brain-computer interface
0.312017
Estimation of EMG signal for shoulder joint based on EEG signals for the control of upper-limb power assistance devices · ICRA 2017
Machine learning › Learning theory › sample complexity
optimal sample complexity
0.212023
The Efficacy of Pessimism in Asynchronous Q-Learning · IEEE Trans. Inf. Theory 2023
Machine learning › Learning theory
statistical guarantees
0.112020
Efficient Clustering for Stretched Mixtures: Landscape and Optimality · NeurIPS 2020

Methods — techniques the papers use, named apart from their topics

total variation distance · 1.7score function estimation · 0.9score estimation · 0.9first-order algorithm · 0.9reward-free exploration · 0.8offline RL insights · 0.8denoising diffusion probabilistic model · 0.8coefficient design · 0.8variance reduction · 0.7stochastic approximation · 0.7lower confidence bound · 0.7nonconvex optimization · 0.4non-convex optimization · 0.4principal component analysis · 0.3linear model · 0.3EEG feature extraction · 0.3
YearPublicationVenuePosition
2025 O(d/T) Convergence Theory for Diffusion Probabilistic Models under Minimal Assumptions
abstract
Score-based diffusion models, which generate new data by learning to reverse a diffusion process that perturbs data from the target distribution into noise, have achieved remarkable success across various generative tasks. Despite their superior empirical performance, existing theoretical guarantees are often constrained by stringent assumptions or suboptimal convergence rates. In this paper, we establish a fast convergence theory for the denoising diffusion probabilistic model (DDPM), a widely used SDE-based sampler, under minimal assumptions. Our analysis shows that, provided $\ell_{2}$-accurate estimates of the score functions, the total variation distance between the target and generated distributions is upper bounded by $O(d/T)$ (ignoring logarithmic factors), where $d$ is the data dimensionality and $T$ is the number of steps. This result holds for any target distribution with finite first-order moment. To our knowledge, this improves upon existing convergence theory for the DDPM sampler, while imposing minimal assumptions on the target data distribution and score estimates. This is achieved through a novel set of analytical tools that provides a fine-grained characterization of how the error propagates at each step of the reverse process.
Gen Li 0005, Yuling Yan
ICLR2
2025 O(d/T) Convergence Theory for Diffusion Probabilistic Models under Minimal Assumptions
abstract
Score-based diffusion models, which generate new data by learning to reverse a diffusion process that perturbs data from the target distribution into noise, have achieved remarkable success across various generative tasks. Despite their superior empirical performance, existing theoretical guarantees are often constrained by stringent assumptions or suboptimal convergence rates. In this paper, we establish a fast convergence theory for the denoising diffusion probabilistic model (DDPM), a widely used SDE-based sampler, under minimal assumptions. Our analysis shows that, provided $\ell_{2}$-accurate estimates of the score functions, the total variation distance between the target and generated distributions is upper bounded by $O(d/T)$ (ignoring logarithmic factors), where $d$ is the data dimensionality and $T$ is the number of steps. This result holds for any target distribution with finite first-order moment. Moreover, we show that with careful coefficient design, the convergence rate improves to $O(k/T)$, where $k$ is the intrinsic dimension of the target data distribution. This highlights the ability of DDPM to automatically adapt to unknown low-dimensional structures, a common feature of natural image distributions. These results are achieved through a novel set of analytical tools that provides a fine-grained characterization of how the error propagates at each step of the reverse process.
Gen Li 0005, Yuling Yan
J. Mach. Learn. Res.2
2024 Minimax-optimal reward-agnostic exploration in reinforcement learning
abstract
This paper studies reward-agnostic exploration in reinforcement learning (RL) — a scenario where the learner is unware of the reward functions during the exploration stage — and designs an algorithm that improves over the state of the art. More precisely, consider a finite-horizon inhomogeneous Markov decision process with $S$ states, $A$ actions, and horizon length $H$, and suppose that there are no more than a polynomial number of given reward functions of interest. By collecting an order of $\frac{SAH^3}{\varepsilon^2}$ sample episodes (up to log factor) without guidance of the reward information, our algorithm is able to find $\varepsilon$-optimal policies for all these reward functions, provided that $\varepsilon$ is sufficiently small. This forms the first reward-agnostic exploration scheme in this context that achieves provable minimax optimality. Furthermore, once the sample size exceeds $\frac{S^2AH^3}{\varepsilon^2}$ episodes (up to log factor), our algorithm is able to yield $\varepsilon$ accuracy for arbitrarily many reward functions (even when they are adversarially designed), a task commonly dubbed as “reward-free exploration.” The novelty of our algorithm design draws on insights from offline RL: the exploration scheme attempts to maximize a critical reward-agnostic quantity that dictates the performance of offline RL, while the policy learning paradigm leverages ideas from sample-optimal offline RL paradigms.
Gen Li 0005, Yuling Yan, Yuxin Chen 0002, Jianqing Fan
COLT2
2024 Adapting to Unknown Low-Dimensional Structures in Score-Based Diffusion Models
abstract
This paper investigates score-based diffusion models when the underlying target distribution is concentrated on or near low-dimensional manifolds within the higher-dimensional space in which they formally reside, a common characteristic of natural image distributions. Despite previous efforts to understand the data generation process of diffusion models, existing theoretical support remains highly suboptimal in the presence of low-dimensional structure, which we strengthen in this paper. For the popular Denoising Diffusion Probabilistic Model (DDPM), we find that the dependency of the error incurred within each denoising step on the ambient dimension $d$ is in general unavoidable. We further identify a unique design of coefficients that yields a converges rate at the order of $O(k^{2}/\sqrt{T})$ (up to log factors), where $k$ is the intrinsic dimension of the target distribution and $T$ is the number of steps. This represents the first theoretical demonstration that the DDPM sampler can adapt to unknown low-dimensional structures in the target distribution, highlighting the critical importance of coefficient design. All of this is achieved by a novel set of analysis tools that characterize the algorithmic dynamics in a more deterministic manner.
Gen Li 0005, Yuling Yan
NeurIPS2
2023 The Efficacy of Pessimism in Asynchronous Q-Learning
abstract
This paper is concerned with the asynchronous form of Q-learning, which applies a stochastic approximation scheme to Markovian data samples. Motivated by the recent advances in offline reinforcement learning, we develop an algorithmic framework that incorporates the principle of pessimism into asynchronous Q-learning, which penalizes infrequently-visited state-action pairs based on suitable lower confidence bounds (LCBs). This framework leads to, among other things, improved sample efficiency and enhanced adaptivity in the presence of near-expert data. Our approach permits the observed data in some important scenarios to cover only partial state-action space, which is in stark contrast to prior theory that requires uniform coverage of all state-action pairs. When coupled with the idea of variance reduction, asynchronous Q-learning with LCB penalization achieves near-optimal sample complexity, provided that the target accuracy level is small enough. In comparison, prior works were suboptimal in terms of the dependency on the effective horizon even when i.i.d. sampling is permitted. Our results deliver the first theoretical support for the use of pessimism principle in the presence of Markovian non-i.i.d. data.
Yuling Yan, Gen Li 0005, Yuxin Chen 0002, Jianqing Fan
IEEE Trans. Inf. Theory1
2022 AGRMTS: A virtual aircraft maintenance training system using gesture recognition based on PSO-BPNN model
abstract
Abstract The quality and efficiency of aircraft maintenance are the key to ensure flight safety and on‐time rate, which mainly depend on the techniques and experience of maintenance engineer. Generally, exercises on physical prototypes are used to improve the maintenance capability of engineers, but this will waste a lot of consumables and easily cause safety accidents. With the development of computer technology, maintenance training in a virtual environment has become an advanced and reliable solution. In this paper, a virtual training system of aircraft maintenance based on gesture recognition interaction is established. Leap Motion is used as a sensor to construct a hybrid machine learning gesture recognition model, so as to obtain natural human–computer interaction experience. In the recognition model, the initial weight matrix and the number of hidden layer nodes in the back propagation neural network are jointly optimized by the Particle Swarm Optimization algorithm with self‐adaption inertial weight. This optimization algorithm achieved a recognition rate of 81.26% in the dynamic gesture database constructed in this paper, which is higher than other available algorithms. A preliminary usability evaluation in university classrooms shows that the teaching system in this paper can achieve a better interactive experience.
Yuling Yan, Minye Chen
Comput. Animat. Virtual Worlds1
2021 Sample-Efficient Reinforcement Learning for Linearly-Parameterized MDPs with a Generative Model
abstract
The curse of dimensionality is a widely known issue in reinforcement learning (RL). In the tabular setting where the state space $\mathcal{S}$ and the action space $\mathcal{A}$ are both finite, to obtain a near optimal policy with sampling access to a generative model, the minimax optimal sample complexity scales linearly with $|\mathcal{S}|\times|\mathcal{A}|$, which can be prohibitively large when $\mathcal{S}$ or $\mathcal{A}$ is large. This paper considers a Markov decision process (MDP) that admits a set of state-action features, which can linearly express (or approximate) its probability transition kernel. We show that a model-based approach (resp.$~$Q-learning) provably learns an $\varepsilon$-optimal policy (resp.$~$Q-function) with high probability as soon as the sample size exceeds the order of $\frac{K}{(1-\gamma)^{3}\varepsilon^{2}}$ (resp.$~$$\frac{K}{(1-\gamma)^{4}\varepsilon^{2}}$), up to some logarithmic factor. Here $K$ is the feature dimension and $\gamma\in(0,1)$ is the discount factor of the MDP. Both sample complexity bounds are provably tight, and our result for the model-based approach matches the minimax lower bound. Our results show that for arbitrarily large-scale MDP, both the model-based approach and Q-learning are sample-efficient when $K$ is relatively small, and hence the title of this paper.
Yuling Yan, Jianqing Fan
NeurIPS2
2020 Efficient Clustering for Stretched Mixtures: Landscape and Optimality
abstract
This paper considers a canonical clustering problem where one receives unlabeled samples drawn from a balanced mixture of two elliptical distributions and aims for a classifier to estimate the labels. Many popular methods including PCA and k-means require individual components of the mixture to be somewhat spherical, and perform poorly when they are stretched. To overcome this issue, we propose a non-convex program seeking for an affine transform to turn the data into a one-dimensional point cloud concentrating around -1 and 1, after which clustering becomes easy. Our theoretical contributions are two-fold: (1) we show that the non-convex loss function exhibits desirable geometric properties when the sample size exceeds some constant multiple of the dimension, and (2) we leverage this to prove that an efficient first-order algorithm achieves near-optimal statistical precision without good initialization. We also propose a general methodology for clustering with flexible choices of feature transforms and loss objectives.
Yuling Yan, Mateo Díaz
NeurIPS2
2017 Estimation of EMG signal for shoulder joint based on EEG signals for the control of upper-limb power assistance devices
abstract
Brain-Machine Interface (BMI) has emerged as a powerful tool for assisting disabled people and for augmenting human performance. Up so far, no studies have succeeded in the power augmentation for the multi-DOFs robot based on EEG signals, especially for the complex shoulder joint. In this work, we propose an electromyography (EMG) estimation method based on electroencephalography (EEG) signals to realize the power assistance. The positions of the electrodes where the motion information of shoulder joint is effectively and exactly extracted are discussed, and a linear model that correlates the EMG to the EEG signal is constructed utilizing motion-related features extracted from multi-location EEG measurements. The constructed model is used to estimate the human muscular activity of shoulder joint from EEG using Principal Component Analysis (PCA) method. The proposed approach is experimentally verified, and an average correlation coefficients are as high as about 0.90 for different subjects are obtained between the estimated and the actually measured EMG signal. Our results suggest that the estimation of EMG based on EEG is feasible. This demonstrates the potential of using EEG signals to support human activities via brain-machine interface.
Hongbo Liang, Chi Zhu 0001, Masataka Yoshioka, Naoya Ueda, Yu Iwata, Haoyong Yu, Feng Duan 0006, Yuling Yan
ICRA9
2012 Snake based automatic tracing of vocal-fold motion from high-speed digital images
abstract
High-speed digital imaging (HSDI) of the larynx provides important information on the vocal fold vibrations that are closely associated with voice condition. We present an active contour (snake)-based algorithm for the automatic delineation of the glottis within image sequences captured from an HSDI system. The algorithm has three steps: first, a rough segmentation is performed by global thresholding, and followed by the detection of an ellipse-shaped region that approximates the glottal geometry, secondly, parameters of the ellipse are estimated using the principal component analysis (PCA) method and thirdly, the snake-method is applied using the estimated ellipse as an initial contour. The performance of the proposed approach is demonstrated through the use of clinical samples of the HSDI recordings obtained from subjects having both normal and pathological voice conditions. Finally the proposed method is compared with existing snake-based methods in terms of efficiency and segmentation accuracy.
Yuling Yan, Gan Du, Chi Zhu 0001, Gerard Marriott
ICASSP1
2010 A new type of omnidirectional wheelchair robot for walking support and power assistance
abstract
Up to now, many robotic aids for the elderly's walking support or the disabled's walking rehabilitation are reported, and numerous electrical-powered wheelchairs are developed. In this paper, a new kind of omnidirectional wheelchair typed robot is developed. The robot not only can accomplish the walking support or walking rehabilitation as the elderly or the disabled walk, but also can realize the power assistance for a caregiver when he/she pushes the robot to move. The basic structure of the robot is described, and the omnidirectional mobility of the robot is analyzed. Further, an admittance based human-machine interaction controller is introduced for power assistance. Experiments are implemented, and the experimental results show that the pushing force can be reduced and well controlled arbitrarily as designed. The development purposes of the robot for walking support and power assistance are achieved.
Chi Zhu 0001, Masashi Oda, Masayuki Suzuki, Xiang Luo 0001, Hideomi Watanabe, Yuling Yan
IROS6
2010 Admittance based control of wheelchair typed omnidirectional robot for walking support and power assistance
abstract
In this paper, a new type of omnidirectional mobile robot is developed, that not only can be used for the elderly's walking support, the disabled's walking rehabilitation, but also can be used for a caregiver's power assistance when he/she pushes the robot to move while the elderly or the disabled is sitting in the seat. The omnidirectional mobility of the robot is analyzed, and an admittance based human-machine interaction controller is introduced for power assistance. The experimental results show that the pushing force is reduced and well controlled as we planned. The purpose of walking support and power assistance is achieved.
Masashi Oda, Chi Zhu 0001, Masayuki Suzuki, Xiang Luo 0001, Hideomi Watanabe, Yuling Yan
RO-MAN6
2009 Acoustic and high-speed digital imaging based analysis of pathological voice contributes to better understanding and differential diagnosis of neurological dysphonias and of mimicking phonatory disorders
Krzysztof Izdebski, Yuling Yan, Melda Kunduk
INTERSPEECH2