Yijie Peng

dblp:127/7641 · DBLP profile ↗
← Back
14ranked-venue papers
3as first author
12since 2021 · last 2025
0000-0003-2584-8131ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Theory of computation · 9 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2025 FLOPS: Forward Learning with OPtimal Sampling
abstract
Given the limitations of backpropagation, perturbation-based gradient computation methods have recently gained focus for learning with only forward passes, also referred to as queries. Conventional forward learning consumes enormous queries on each data point for accurate gradient estimation through Monte Carlo sampling, which hinders the scalability of those algorithms. However, not all data points deserve equal queries for gradient estimation. In this paper, we study the problem of improving the forward learning efficiency from a novel perspective: how to reduce the gradient estimation variance with minimum cost? For this, we allocate the optimal number of queries within a set budget during training to balance estimation accuracy and computational efficiency. Specifically, with a simplified proxy objective and a reparameterization technique, we derive a novel plug-and-play query allocator with minimal parameters. Theoretical results are carried out to verify its optimality. We conduct extensive experiments for fine-tuning Vision Transformers on various datasets and further deploy the allocator to two black-box applications: prompt tuning and multimodal alignment for foundation models. All findings demonstrate that our proposed allocator significantly enhances the scalability of forward-learning algorithms, paving the way for real-world applications. The implementation is available at https://github.com/RTkenny/FLOPS-Forward-Learning-with-OPtimal-Sampling.
Tao Ren 0006, Zishi Zhang, Jinyang Jiang 0001, Zeliang Zhang 0001, Mingqian Feng, Yijie Peng
ICLR7
2025 Exploring and Exploiting Model Uncertainty in Bayesian Optimization
abstract
In this work, we consider the problem of Bayesian Optimization (BO) under reward model uncertainty—that is, when the underlying distribution type of the reward is unknown and potentially intractable to specify. This challenge is particularly evident in many modern applications, where the reward distribution is highly ill-behaved, often non-stationary, multi-modal, or heavy-tailed. In such settings, classical Gaussian Process (GP)-based BO methods often fail due to their strong modeling assumptions. To address this challenge, we propose a novel surrogate model, the infinity-Gaussian Process ($\infty$-GP), which represents a sequential spatial Dirichlet Process mixture with a GP baseline. The $\infty$-GP quantifies both value uncertainty and model uncertainty, enabling more flexible modeling of complex reward structures. Combined with Thompson Sampling, the $\infty$-GP facilitates principled exploration and exploitation in the distributional space of reward models. Theoretically, we prove that the $\infty$-GP surrogate model can approximate a broad class of reward distributions by effectively exploring the distribution space, achieving near-minimax-optimal posterior contraction rates. Empirically, our method outperforms state-of-the-art approaches in various challenging scenarios, including highly non-stationary and heavy-tailed reward settings where classical GP-based BO often fails.
Zishi Zhang, Tao Ren 0006, Yijie Peng
NeurIPS3
2025 An Efficient Node Selection Policy for Monte Carlo Tree Search with Neural Networks: INFORMS Journal on Computing Meritorious Paper Awardee
abstract
Monte Carlo tree search (MCTS) has been gaining increasing popularity, and the success of AlphaGo has prompted a new trend of incorporating a value network and a policy network constructed with neural networks into MCTS, namely, NN-MCTS. In this work, motivated by the shortcomings of the widely used upper confidence bounds applied to trees (UCT) policy, we formulate the node selection problem in NN-MCTS as a multistage ranking and selection (R&S) problem and propose a node selection policy that efficiently allocates a limited search budget to maximize the probability of correctly selecting the best action at the root state. The value and policy networks in NN-MCTS further improve the performance of the proposed node selection policy by providing prior knowledge and guiding the selection of the final action, respectively. Numerical experiments on two board games and an OpenAI task demonstrate that the proposed method outperforms the UCT policy used in AlphaGo Zero and MuZero, implying the potential of constructing node selection policies in NN-MCTS with R&S procedures. History: Accepted by Bruno Tuffin, Area Editor for Simulation. Funding: This work was supported by the National Natural Science Foundation of China [Grants 72325007, 72250065, and 72022001], and a PKU-Boya Postdoctoral Fellowship 2406396158. Supplemental Material: The software that supports the findings of this study is available within the paper and its Supplemental Information ( https://pubsonline.informs.org/doi/suppl/10.1287/ijoc.2023.0307 ) as well as from the IJOC GitHub software repository ( https://github.com/INFORMSJoC/2023.0307 ). The complete IJOC Software and Data Repository is available at https://informsjoc.github.io/ .
Xiaotian Liu, Yijie Peng, Ruihan Zhou
INFORMS J. Comput.2
2025 Noise Optimization in Artificial Neural Networks
abstract
Artificial neural network (ANN) has been widely used in automation. However, the vulnerability of ANN under certain attacks poses a security threat to critical automation systems. Previous research has shown that adding noise to ANNs can enhance robustness. Nonetheless, striking a balance between robustness and task performance remains challenging, as excessive noise improves robustness but hampers performance, while low noise offers minor robustness improvement. In this work, we propose to learn the distribution of optimal injected noise, which improves the robustness as well as maintains the performance. Specifically, we compute the pathwise stochastic gradient estimate with respect to the standard deviation of the Gaussian noise added to each neuron of the ANN and optimize both the noise distribution and model parameters during training with negligible additional computational cost. In numerical experiments, our proposed method can achieve significant performance improvement on the robustness of several popular ANN structures under both black box and white box attacks. We also evaluate the proposed technique on two automation tasks: the classic reinforcement learning task of the cart pole game and a fault detection problem. Our results showed that the proposed technique outperforms a conventional neural network in terms of performance, robustness, and visual explainability.Note to Practitioners—The robustness of artificial neural networks is a critical consideration in automation applications as real-world data is often subject to unforeseen perturbations from the environment, potentially causing AI systems to behave unpredictably and unstably. For example, object detection is a widely employed AI technique in automation applications. However, current object detection systems are vulnerable to noise perturbation. Even small, imperceptible noise can lead the model to malfunction. Our work focuses on improving the robustness of neural networks. We propose a novel technique that can be added to any layer of existing neural networks to enhance robustness. Extensive experiments conducted in various scenarios have verified the effectiveness of the proposed method in enhancing both performance and robustness.
Li Xiao 0005, Zeliang Zhang 0001, Kuihua Huang, Jinyang Jiang 0001, Yijie Peng
IEEE Trans Autom. Sci. Eng.5
2025 Efficient Learning for Selecting Top-m Context-Dependent Designs
abstract
We consider a simulation optimization problem for context-dependent decision-making, which aims to determine the top-$m$designs for all contexts. Under a Bayesian framework, we formulate the optimal dynamic sampling decision as a stochastic dynamic programming problem and develop a sequential sampling policy to efficiently learn the performance of each design under each context. The asymptotically optimal sampling ratios are derived to attain the optimal large deviations rate of the worst-case probability of false selection. The proposed sampling policy is proved to be consistent, and its asymptotic sampling ratios are shown to be asymptotically optimal. Numerical experiments demonstrate that the proposed method improves the efficiency for selecting top-$m$context-dependent designs.Note to Practitioners—The performance of a given design may vary across different contexts, and better decision-making is possible by refining the available contextual information. We consider a context-dependent ranking and selection problem, which allows optimal selection to depend on contextual information obtained prior to decision-making. We develop a dynamic sampling scheme to efficiently learn and select the top-$m$designs in all contexts. Numerical experiments, including a honeypot deception game problem and a medical resource allocation problem, demonstrate that the proposed sampling scheme significantly improves the efficiency of context-dependent ranking and selection for the top-$m$designs.
Sihua Chen, Kuihua Huang, Yijie Peng
IEEE Trans Autom. Sci. Eng.4
2024 One Forward is Enough for Neural Network Training via Likelihood Ratio Method
abstract
While backpropagation (BP) is the mainstream approach for gradient computation in neural network training, its heavy reliance on the chain rule of differentiation constrains the designing flexibility of network architecture and training pipelines. We avoid the recursive computation in BP and develop a unified likelihood ratio (ULR) method for gradient estimation with only one forward propagation. Not only can ULR be extended to train a wide variety of neural network architectures, but the computation flow in BP can also be rearranged by ULR for better device adaptation. Moreover, we propose several variance reduction techniques to further accelerate the training process. Our experiments offer numerical results across diverse aspects, including various neural network training scenarios, computation flow rearrangement, and fine-tuning of pre-trained models. All findings demonstrate that ULR effectively enhances the flexibility of neural network training by permitting localized module training without compromising the global objective and significantly boosts the network robustness.
Jinyang Jiang 0001, Zeliang Zhang 0001, Chenliang Xu, Zhaofei Yu, Yijie Peng
ICLR5
2023 Asymptotically Optimal Sampling Policy for Selecting Top-m Alternatives
abstract
We consider selecting the top-m alternatives from a finite number of alternatives via Monte Carlo simulation. Under a Bayesian framework, we formulate the sampling decision as a stochastic dynamic programming problem and develop a sequential sampling policy that maximizes a value function approximation one-step look ahead. To show the asymptotic optimality of the proposed procedure, the asymptotically optimal sampling ratios that optimize the large deviations rate of the probability of false selection for selecting the top-m alternatives have been rigorously defined. The proposed sampling policy is not only proved to be consistent but also achieve the asymptotically optimal sampling ratios. Numerical experiments demonstrate superiority of the proposed allocation procedure over existing ones. History: Accepted by Bruno Tuffin, Area Editor for Simulation. Funding: This work was supported by the National Natural Science Foundation of China [Grants 72250065, 72293582, 72022001, and 71901003], and the National Science Foundation [Grant DMS-2053489], the major project of the National Natural Science Foundation of China [Grant 72293582], and the China Scholarship Council [Grant CSC202206010152]. Supplemental Material: The software that supports the findings of this study is available within the paper and its Supplemental Information ( https://pubsonline.informs.org/doi/suppl/10.1287/ijoc.2021.0333 ) as well as from the IJOC GitHub software repository ( https://github.com/INFORMSJoC/2021.0333 ). The complete IJOC Software and Data Repository is available at https://informsjoc.github.io/ .
Yijie Peng, Jianghua Zhang, Enlu Zhou
INFORMS J. Comput.2
2022 A Stochastic Approximation Method for Simulation-Based Quantile Optimization
abstract
We present a gradient-based algorithm for solving a class of simulation optimization problems in which the objective function is the quantile of a simulation output random variable. In contrast with existing quantile (quantile derivative) estimation techniques, which aim to eliminate the estimator bias by gradually increasing the simulation sample size, our algorithm incorporates a novel recursive procedure that only requires a single simulation sample at each step to simultaneously obtain quantile and quantile derivative estimators that are asymptotically unbiased. We show that these estimators, when coupled with the standard gradient descent method, lead to a multitime-scale stochastic approximation type of algorithm that converges to an optimal quantile value with probability one. In our numerical experiments, the proposed algorithm is applied to optimal investment portfolio problems, resulting in new solutions that complement those obtained under the classical Markowitz mean-variance framework. History: Accepted by Alice E. Smith, Editor-in-Chief; Bruno Tuffin, Area Editor for Simulation. Funding: The work of Y. Peng was supported in part by the National Natural Science Foundation of China (NSFC) [Grants 72022001, 92146003, and 71901003], and by the Key Research and Development Programof Beijing Municipal Science and Technology Commission. Supplemental Material: The e-companion is available at https://doi.org/10.1287/ijoc.2022.1214 .
Jiaqiao Hu, Yijie Peng
INFORMS J. Comput.2
2022 A New Likelihood Ratio Method for Training Artificial Neural Networks
abstract
We investigate a new approach to compute the gradients of artificial neural networks (ANNs), based on the so-called push-out likelihood ratio method. Unlike the widely used backpropagation (BP) method that requires continuity of the loss function and the activation function, our approach bypasses this requirement by injecting artificial noises into the signals passed along the neurons. We show how this approach has a similar computational complexity as BP, and moreover is more advantageous in terms of removing the backward recursion and eliciting transparent formulas. We also formalize the connection between BP, a pivotal technique for training ANNs, and infinitesimal perturbation analysis, a classic path-wise derivative estimation approach, so that both our new proposed methods and BP can be better understood in the context of stochastic gradient estimation. Our approach allows efficient training for ANNs with more flexibility on the loss and activation functions, and shows empirical improvements on the robustness of ANNs under adversarial attacks and corruptions of natural noises. Summary of Contribution: Stochastic gradient estimation has been studied actively in simulation for decades and becomes more important in the era of machine learning and artificial intelligence. The stochastic gradient descent is a standard technique for training the artificial neural networks (ANNs), a pivotal problem in deep learning. The most popular stochastic gradient estimation technique is the backpropagation method. We find that the backpropagation method lies in the family of infinitesimal perturbation analysis, a path-wise gradient estimation technique in simulation. Moreover, we develop a new likelihood ratio-based method, another popular family of gradient estimation technique in simulation, for training more general ANNs, and demonstrate that the new training method can improve the robustness of the ANN.
Yijie Peng, Li Xiao 0005, Bernd Heidergott, L. Jeff Hong, Henry Lam
INFORMS J. Comput.1
2022 Dynamic Sampling Allocation Under Finite Simulation Budget for Feasibility Determination
abstract
Monte Carlo simulation is a commonly used tool for evaluating the performance of complex stochastic systems. In practice, simulation can be expensive, especially when comparing a large number of alternatives, thus motivating the need to intelligently allocate simulation replications. Given a finite set of alternatives whose means are estimated via simulation, we consider the problem of determining the subset of alternatives that have means smaller than a fixed threshold. A dynamic sampling procedure that possesses not only asymptotic optimality, but also desirable finite-sample properties is proposed. Theoretical results show that there is a significant difference between finite-sample optimality and asymptotic optimality. Numerical experiments substantiate the effectiveness of the new method. Summary of Contribution: Simulation is an important tool to estimate the performance of complex stochastic systems. We consider a feasibility determination problem of identifying all those among a finite set of alternatives with mean smaller than a given threshold, in which the means are unknown but can be estimated by sampling replications via stochastic simulation. This problem appears widely in many applications, including call center design and hospital resource allocation. Our work considers how to intelligently allocate simulation replications to different alternatives for efficiently finding the feasible alternatives. Previous work focuses on the asymptotic properties of the sampling allocation procedures, whereas our contribution lies in developing a finite-budget allocation rule that possesses both asymptotic optimality and desirable finite-budget properties.
Zhongshun Shi, Yijie Peng, Leyuan Shi, Chun-Hung Chen, Michael C. Fu 0001
INFORMS J. Comput.2
2021 Computing Sensitivities for Distortion Risk Measures
abstract
Distortion risk measure, defined by an integral of a distorted tail probability, has been widely used in behavioral economics and risk management as an alternative to expected utility. The sensitivity of the distortion risk measure is a functional of certain distribution sensitivities. We propose a new sensitivity estimator for the distortion risk measure that uses generalized likelihood ratio estimators for distribution sensitivities as input and establish a central limit theorem for the new estimator. The proposed estimator can handle discontinuous sample paths and distortion functions.
Peter W. Glynn, Yijie Peng, Michael C. Fu 0001, Jian-Qiang Hu
INFORMS J. Comput.2
2021 Efficient Sampling Allocation Procedures for Optimal Quantile Selection
abstract
We propose a dynamic sampling allocation and selection paradigm for finding the alternative with the optimal quantile in a Bayesian framework. Myopic allocation policies (MAPs), analogous to existing methods in classic ranking and selection for selecting the alternative with the optimal mean, and computationally efficient selection policies are derived for selecting the alternative with the optimal quantile. Under certain conditions, we prove that the proposed MAPs and selection procedures are consistent, which means that the best quantile would be eventually correctly selected as the sample size goes to infinity. Numerical experiments demonstrate that the proposed schemes can significantly improve the performance.
Yijie Peng, Chun-Hung Chen, Michael C. Fu 0001, Jian-Qiang Hu, Ilya O. Ryzhov
INFORMS J. Comput.1
2020 On the Variance of Single-Run Unbiased Stochastic Derivative Estimators
Zhenyu Cui, Michael C. Fu 0001, Jian-Qiang Hu, Yanchu Liu, Yijie Peng, Lingjiong Zhu
INFORMS J. Comput.5
2016 Dynamic Sampling Allocation and Design Selection
abstract
We formulate the statistical selection problem in a general dynamic framework comprising fully sequential sampling allocation and optimal design selection. Because the traditional probability of correct selection measure is not sufficient to capture both aspects in this more general framework, we introduce the integrated probability of correct selection to better characterize the objective. As a result, the usual selection policy of choosing the design with the largest sample mean as the estimate of the best is no longer necessarily optimal. Rather, the optimal selection policy is to choose the design that maximizes the posterior integrated probability of correct selection, which is a function of the posterior mean and the correlation structure induced by the posterior variance. Because determining the optimal selection policy is generally intractable, we also devise an approximation scheme to efficiently approximate the optimal selection policy. For the allocation policy, we study an asymptotic policy called general Bayesian budget allocation, which is comprised of a sampling statistic and a sequential rule. The optimal computing budget allocation algorithm can be interpreted as a special case of the asymptotical sampling statistics. Numerical examples are provided to illustrate the potential performance improvements, especially in small sample behavior.
Yijie Peng, Chun-Hung Chen, Michael C. Fu 0001, Jian-Qiang Hu
INFORMS J. Comput.1