VLDB 2026 Research / reviewers in the wild / expert
Ya-Xiang Yuan
dblp:45/2834 · also Ya-xiang Yuan, Yaxiang Yuan
· DBLP profile ↗
7ranked-venue papers
0as first author
5since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Theory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Theoretical computer science
2 papers |
Mathematical optimization · 100% | |
| Artificial intelligence
1 paper |
Topics — the 6 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Mathematical optimization
continuous optimization |
1.5 | 2 | 2025 | LancBiO: Dynamic Lanczos-aided Bilevel Optimization via Krylov Subspace · ICLR 2025 On the Convergence of Stochastic Gradient Descent with Bandwidth-based Step Size · J. Mach. Learn. Res. 2023 |
Mathematical optimization
bilevel optimization |
0.9 | 1 | 2025 | LancBiO: Dynamic Lanczos-aided Bilevel Optimization via Krylov Subspace · ICLR 2025 |
Mathematical optimization › iterative methods
krylov subspace methods |
0.9 | 1 | 2025 | LancBiO: Dynamic Lanczos-aided Bilevel Optimization via Krylov Subspace · ICLR 2025 |
Mathematical optimization
convergence analysis |
0.7 | 1 | 2023 | On the Convergence of Stochastic Gradient Descent with Bandwidth-based Step Size · J. Mach. Learn. Res. 2023 |
Mathematical optimization
step-size selection |
0.7 | 1 | 2023 | On the Convergence of Stochastic Gradient Descent with Bandwidth-based Step Size · J. Mach. Learn. Res. 2023 |
Mathematical optimization › stochastic optimization › stochastic gradient methods
stochastic gradient descent |
0.7 | 1 | 2023 | On the Convergence of Stochastic Gradient Descent with Bandwidth-based Step Size · J. Mach. Learn. Res. 2023 |
Methods — techniques the papers use, named apart from their topics
non-monotonic step size · 1.3cyclical step size · 1.3cosine with restart · 1.3lanczos process · 0.9gradient-based optimization · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Bilevel Reinforcement Learning via the Development of Hyper-gradient without Lower-Level ConvexityabstractBilevel reinforcement learning (RL), which features intertwined two-level problems, has attracted growing interest recently. The inherent non-convexity of the lower-level RL problem is, however, to be an impediment to developing bilevel optimization methods. By employing the fixed point equation associated with the regularized RL, we characterize the hyper-gradient via fully first-order information, thus circumventing the assumption of lower-level convexity. This, remarkably, distinguishes our development of hyper-gradient from the general AID-based bilevel frameworks since we take advantage of the specific structure of RL problems. Moreover, we design both model-based and model-free bilevel reinforcement learning algorithms, facilitated by access to the fully first-order hyper-gradient. Both algorithms enjoy the convergence rate $\mathcal{O}\left(\epsilon^{-1}\right)$. To extend the applicability, a stochastic version of the model-free algorithm is proposed, along with results on its convergence rate and sampling complexity. In addition, numerical experiments demonstrate that the hyper-gradient indeed serves as an integration of exploitation and exploration. Yan Yang 0012, Bin Gao 0007, Ya-Xiang Yuan |
AISTATS | 3 |
| 2025 | LancBiO: Dynamic Lanczos-aided Bilevel Optimization via Krylov SubspaceabstractBilevel optimization, with broad applications in machine learning, has an intricate hierarchical structure. Gradient-based methods have emerged as a common approach to large-scale bilevel problems. However, the computation of the hyper-gradient, which involves a Hessian inverse vector product, confines the efficiency and is regarded as a bottleneck. To circumvent the inverse, we construct a sequence of low-dimensional approximate Krylov subspaces with the aid of the Lanczos process. As a result, the constructed subspace is able to dynamically and incrementally approximate the Hessian inverse vector product with less effort and thus leads to a favorable estimate of the hyper-gradient. Moreover, we propose a provable subspace-based framework for bilevel problems where one central step is to solve a small-size tridiagonal linear system. To the best of our knowledge, this is the first time that subspace techniques are incorporated into bilevel optimization. This successful trial not only enjoys $\mathcal{O}(\epsilon^{-1})$ convergence rate but also demonstrates efficiency in a synthetic problem and two deep learning tasks. Yan Yang 0012, Bin Gao 0007, Ya-Xiang Yuan |
ICLR | 3 |
| 2025 | Revisiting Nesterov's acceleration via high-resolution differential equations
Bin Shi 0005, Ya-Xiang Yuan |
J. Glob. Optim. | 3 |
| 2025 | A momentum accelerated stochastic method and its application on policy search problems
Boou Jiang, Ya-Xiang Yuan |
Neural Comput. Appl. | 2 |
| 2023 | On the Convergence of Stochastic Gradient Descent with Bandwidth-based Step SizeabstractWe first propose a general step-size framework for the stochastic gradient descent(SGD) method: bandwidth-based step sizes that are allowed to vary within a banded region. The framework provides efficient and flexible step size selection in optimization, including cyclical and non-monotonic step sizes (e.g., triangular policy and cosine with restart), for which theoretical guarantees are rare. We provide state-of-the-art convergence guarantees for SGD under mild conditions and allow a large constant step size at the beginning of training. Moreover, we investigate the error bounds of SGD under the bandwidth step size where the boundary functions are in the same order and different orders, respectively. Finally, we propose a $1/t$ up-down policy and design novel non-monotonic step sizes. Numerical experiments demonstrate these bandwidth-based step sizes' efficiency and significant potential in training regularized logistic regression and several large-scale neural network tasks. Xiaoyu Wang 0008, Ya-Xiang Yuan |
J. Mach. Learn. Res. | 2 |
| 2011 | Fast algorithm for beamforming problems in distributed communication of relay networksabstractA sequential quadratic programming (SQP) method is pro posed to solve the distributed beamforming problem in multiple relay networks. The problem is formulated as the minimization of the total relay transmit power, subject to individual signal-to-interference-and-noise ratio constraints at each receiver, which is a nonconvex quadratic constraint quadratic programming. Rather than solving its semi-definite programming (SDP) relaxation, we apply the SQP method to solve its tightened form to replace its inequality constraints with equalities. Its global convergence is guaranteed. Simulations show that it not only runs much faster, but also performs as good as SDP for calculation results. Cong Sun 0002, Ya-Xiang Yuan |
ICASSP | 2 |
| 2010 | Modeling and Algorithms of GPS Data Reduction for the Qinghai-Tibet RailwayabstractSatellites are currently being used to track the positions of trains. Positioning systems using satellites can help reduce the cost of installing and maintaining trackside equipment. This paper develops a nonlinear combinatorial data reduction model for a large amount of railway Global Positioning System (GPS) data to decrease the memory space and, thus, speed up train positioning. Three algorithms are proposed by employing the concept of looking ahead, using the dichotomy idea, or adopting the breadth-first strategy after changing the problem into a shortest path problem to obtain an optimal solution. Two techniques are developed to substantially cut down the computing time for the optimal algorithm. The surveyed GPS data of the Qinghai–Tibet railway (QTR) are used to compare the performance of the algorithms. Results show that the algorithms can extract a few data points from the large amount of GPS data points, thus enabling a simpler representation of the train tracks. Furthermore, these proposed algorithms show a tradeoff between the solution quality and computation time of the algorithms. Dewang Chen, Yun-Shan Fu, Baigen Cai, Ya-Xiang Yuan |
IEEE Trans. Intell. Transp. Syst. | 4 |