Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Ya-Xiang Yuan

dblp:45/2834 · also Ya-xiang Yuan, Yaxiang Yuan · DBLP profile ↗
← Back
7ranked-venue papers
0as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Theory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Theoretical computer science
2 papers
Mathematical optimization · 100%
Artificial intelligence
1 paper

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Mathematical optimization
continuous optimization
1.522025
LancBiO: Dynamic Lanczos-aided Bilevel Optimization via Krylov Subspace · ICLR 2025
On the Convergence of Stochastic Gradient Descent with Bandwidth-based Step Size · J. Mach. Learn. Res. 2023
Mathematical optimization
bilevel optimization
0.912025
LancBiO: Dynamic Lanczos-aided Bilevel Optimization via Krylov Subspace · ICLR 2025
Mathematical optimization › iterative methods
krylov subspace methods
0.912025
LancBiO: Dynamic Lanczos-aided Bilevel Optimization via Krylov Subspace · ICLR 2025
Mathematical optimization
convergence analysis
0.712023
On the Convergence of Stochastic Gradient Descent with Bandwidth-based Step Size · J. Mach. Learn. Res. 2023
Mathematical optimization
step-size selection
0.712023
On the Convergence of Stochastic Gradient Descent with Bandwidth-based Step Size · J. Mach. Learn. Res. 2023
Mathematical optimization › stochastic optimization › stochastic gradient methods
stochastic gradient descent
0.712023
On the Convergence of Stochastic Gradient Descent with Bandwidth-based Step Size · J. Mach. Learn. Res. 2023

Methods — techniques the papers use, named apart from their topics

non-monotonic step size · 1.3cyclical step size · 1.3cosine with restart · 1.3lanczos process · 0.9gradient-based optimization · 0.9
YearPublicationVenuePosition
2025 Bilevel Reinforcement Learning via the Development of Hyper-gradient without Lower-Level Convexity
abstract
Bilevel reinforcement learning (RL), which features intertwined two-level problems, has attracted growing interest recently. The inherent non-convexity of the lower-level RL problem is, however, to be an impediment to developing bilevel optimization methods. By employing the fixed point equation associated with the regularized RL, we characterize the hyper-gradient via fully first-order information, thus circumventing the assumption of lower-level convexity. This, remarkably, distinguishes our development of hyper-gradient from the general AID-based bilevel frameworks since we take advantage of the specific structure of RL problems. Moreover, we design both model-based and model-free bilevel reinforcement learning algorithms, facilitated by access to the fully first-order hyper-gradient. Both algorithms enjoy the convergence rate $\mathcal{O}\left(\epsilon^{-1}\right)$. To extend the applicability, a stochastic version of the model-free algorithm is proposed, along with results on its convergence rate and sampling complexity. In addition, numerical experiments demonstrate that the hyper-gradient indeed serves as an integration of exploitation and exploration.
Yan Yang 0012, Bin Gao 0007, Ya-Xiang Yuan
AISTATS3
2025 LancBiO: Dynamic Lanczos-aided Bilevel Optimization via Krylov Subspace
abstract
Bilevel optimization, with broad applications in machine learning, has an intricate hierarchical structure. Gradient-based methods have emerged as a common approach to large-scale bilevel problems. However, the computation of the hyper-gradient, which involves a Hessian inverse vector product, confines the efficiency and is regarded as a bottleneck. To circumvent the inverse, we construct a sequence of low-dimensional approximate Krylov subspaces with the aid of the Lanczos process. As a result, the constructed subspace is able to dynamically and incrementally approximate the Hessian inverse vector product with less effort and thus leads to a favorable estimate of the hyper-gradient. Moreover, we propose a provable subspace-based framework for bilevel problems where one central step is to solve a small-size tridiagonal linear system. To the best of our knowledge, this is the first time that subspace techniques are incorporated into bilevel optimization. This successful trial not only enjoys $\mathcal{O}(\epsilon^{-1})$ convergence rate but also demonstrates efficiency in a synthetic problem and two deep learning tasks.
Yan Yang 0012, Bin Gao 0007, Ya-Xiang Yuan
ICLR3
2025 Revisiting Nesterov's acceleration via high-resolution differential equations
Bin Shi 0005, Ya-Xiang Yuan
J. Glob. Optim.3
2025 A momentum accelerated stochastic method and its application on policy search problems
Boou Jiang, Ya-Xiang Yuan
Neural Comput. Appl.2
2023 On the Convergence of Stochastic Gradient Descent with Bandwidth-based Step Size
abstract
We first propose a general step-size framework for the stochastic gradient descent(SGD) method: bandwidth-based step sizes that are allowed to vary within a banded region. The framework provides efficient and flexible step size selection in optimization, including cyclical and non-monotonic step sizes (e.g., triangular policy and cosine with restart), for which theoretical guarantees are rare. We provide state-of-the-art convergence guarantees for SGD under mild conditions and allow a large constant step size at the beginning of training. Moreover, we investigate the error bounds of SGD under the bandwidth step size where the boundary functions are in the same order and different orders, respectively. Finally, we propose a $1/t$ up-down policy and design novel non-monotonic step sizes. Numerical experiments demonstrate these bandwidth-based step sizes' efficiency and significant potential in training regularized logistic regression and several large-scale neural network tasks.
Xiaoyu Wang 0008, Ya-Xiang Yuan
J. Mach. Learn. Res.2
2011 Fast algorithm for beamforming problems in distributed communication of relay networks
abstract
A sequential quadratic programming (SQP) method is pro posed to solve the distributed beamforming problem in multiple relay networks. The problem is formulated as the minimization of the total relay transmit power, subject to individual signal-to-interference-and-noise ratio constraints at each receiver, which is a nonconvex quadratic constraint quadratic programming. Rather than solving its semi-definite programming (SDP) relaxation, we apply the SQP method to solve its tightened form to replace its inequality constraints with equalities. Its global convergence is guaranteed. Simulations show that it not only runs much faster, but also performs as good as SDP for calculation results.
Cong Sun 0002, Ya-Xiang Yuan
ICASSP2
2010 Modeling and Algorithms of GPS Data Reduction for the Qinghai-Tibet Railway
abstract
Satellites are currently being used to track the positions of trains. Positioning systems using satellites can help reduce the cost of installing and maintaining trackside equipment. This paper develops a nonlinear combinatorial data reduction model for a large amount of railway Global Positioning System (GPS) data to decrease the memory space and, thus, speed up train positioning. Three algorithms are proposed by employing the concept of looking ahead, using the dichotomy idea, or adopting the breadth-first strategy after changing the problem into a shortest path problem to obtain an optimal solution. Two techniques are developed to substantially cut down the computing time for the optimal algorithm. The surveyed GPS data of the Qinghai–Tibet railway (QTR) are used to compare the performance of the algorithms. Results show that the algorithms can extract a few data points from the large amount of GPS data points, thus enabling a simpler representation of the train tracks. Furthermore, these proposed algorithms show a tradeoff between the solution quality and computation time of the algorithms.
Dewang Chen, Yun-Shan Fu, Baigen Cai, Ya-Xiang Yuan
IEEE Trans. Intell. Transp. Syst.4