VLDB 2026 Research / reviewers in the wild / expert
Xinyu Zhang 0024
dblp:58/4582-24
· DBLP profile ↗
7ranked-venue papers
0as first author
7since 2021 · last 2025
0000-0002-9044-1410ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Theory of computation · 4 · 4 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Ranking Model Averaging: Ranking Based on Model AveragingabstractRanking problems are commonly encountered in practical applications, including order priority ranking, wine quality ranking, and piston slap noise performance ranking. The responses of these ranking applications are often considered as continuous responses, and there is uncertainty on which scoring function is used to model the responses. In this paper, we address the scoring function uncertainty of continuous response ranking problems by proposing a ranking model averaging (RMA) method. With a set of candidate models varied by scoring functions, RMA assigns weights for each model determined by a K-fold crossvalidation criterion based on pairwise loss. We provide two main theoretical properties for RMA. First, we prove that the averaging ranking predictions of RMA are asymptotically optimal in achieving the lowest possible ranking risk. Second, we provide a bound on the difference between the empirical RMA weights and theoretical optimal ones, and we show that RMA weights are consistent. Simulation results validate RMA superiority over competing methods in reducing ranking risk. Moreover, when applied to empirical examples—order priority, wine quality, and piston slap noise—RMA shows its effectiveness in building accurate ranking systems. History: Accepted by Ram Ramesh, Area Editor for Data Science & Machine Learning. Funding: This research was supported by the Beijing Municipal Natural Science Foundation [Grant 1222002], the National Natural Science Foundation of China [Grants 12071457, 12201018, 12301364, 71925007, 72091212, and 72273120], National Quality Infrastructure [Grant 2022YFF0609903], the Chinese Academy of Sciences Project for Young Scientists in Basic Research [Grant YSBR-008], and the Natural Science Foundation of Anhui Province [Grant 2308085QA09]. Supplemental Material: The software that supports the findings of this study is available within the paper and its Supplemental Information ( https://pubsonline.informs.org/doi/suppl/10.1287/ijoc.2023.0257 ) as well as from the IJOC GitHub software repository ( https://github.com/INFORMSJoC/2023.0257 ). The complete IJOC Software and Data Repository is available at https://informsjoc.github.io/ . Ziheng Feng, Baihua He, Tianfa Xie, Xinyu Zhang 0024, Xianpeng Zong |
INFORMS J. Comput. | 4 |
| 2025 | Model Averaging Under Flexible Loss FunctionsabstractTo address model uncertainty under flexible loss functions in prediction problems, we propose a model averaging method that accommodates various loss functions, including asymmetric linear and quadratic loss functions as well as many other asymmetric/symmetric loss functions as special cases. The flexible loss function allows the proposed method to average a large range of models such as the quantile and expectile regression models. To determine the weights of the candidate models, we establish a J-fold cross-validation criterion. Asymptotic optimality and weight convergence are proved for the proposed method. Simulations and an empirical application show the superior performance of the proposed method compared with other methods of model selection and averaging. History: Accepted by Ram Ramesh, Area Editor for Data Science & Machine Learning. Funding: This work was supported by the Beijing Natural Science Foundation [Grant Z240004], Japan Society for the Promotion of Science (KAKENHI) [Grant 22H00833 to Q. Liu], the CAS Project for Young Scientists in Basic Research [Grant YSBR-008], and the National Natural Science Foundation of China [Grants 71925007, 72091212 and 72495124]. Supplemental Material: The software that supports the findings of this study is available within the paper and its Supplemental Information ( https://pubsonline.informs.org/doi/suppl/10.1287/ijoc.2023.0291 ) as well as from the IJOC GitHub software repository ( https://github.com/INFORMSJoC/2023.0291 ). The complete IJOC Software and Data Repository is available at https://informsjoc.github.io/ . Dieqi Gu, Qingfeng Liu, Xinyu Zhang 0024 |
INFORMS J. Comput. | 3 |
| 2025 | Frequentist Model Averaging for Global Fréchet RegressionabstractTo consider model uncertainty in global Fréchet regression and improve density response prediction, we propose a frequentist model averaging method. The weights are chosen by minimizing a cross-validation criterion based on Wasserstein distance. In the cases where all candidate models are misspecified, we prove that the corresponding model averaging estimator has asymptotic optimality, achieving the lowest possible Wasserstein distance. When there are correctly specified candidate models, we prove that our method asymptotically assigns all weights to the correctly specified models. Numerical results of extensive simulations and a real data analysis on intracerebral hemorrhage data strongly favour our method. Xinyu Zhang 0024, Peng Zhao 0012 |
IEEE Trans. Inf. Theory | 2 |
| 2024 | Optimal Weighted Random ForestsabstractThe random forest (RF) algorithm has become a very popular prediction method for its great flexibility and promising accuracy. In RF, it is conventional to put equal weights on all the base learners (trees) to aggregate their predictions. However, the predictive performance of different trees within the forest can vary significantly due to the randomization of the embedded bootstrap sampling and feature selection. In this paper, we focus on RF for regression and propose two optimal weighting algorithms, namely the 1 Step Optimal Weighted RF (1step-WRF$_\mathrm{opt}$) and 2 Steps Optimal Weighted RF (2steps-WRF$_\mathrm{opt}$), that combine the base learners through the weights determined by weight choice criteria. Under some regularity conditions, we show that these algorithms are asymptotically optimal in the sense that the resulting squared loss and risk are asymptotically identical to those of the infeasible but best possible weighted RF. Numerical studies conducted on real-world data sets and semi-synthetic data sets indicate that these algorithms outperform the equal-weight forest and two other weighted RFs proposed in the existing literature in most cases. Dalei Yu, Xinyu Zhang 0024 |
J. Mach. Learn. Res. | 3 |
| 2024 | Reliability Estimation of $k$-Out-of-$n$: G System With Model UncertaintyabstractImproving the reliability estimation of the$k$-out-of-$n$: G system provides support for its proper operation. Taking the example of a supply chain with the$k$-out-of-$n$: G system, we use the copula function to model the dependence structure among suppliers. This article proposes a model averaging method based on Kullback–Leibler loss to estimate the reliability of the supply chain with a$k$-out-of-$n$: G system. The proposed estimator accommodates the uncertainty in the dependence structures among suppliers and is the first to use the optimal model averaging for developing the reliability estimation. We prove the asymptotic optimality of the proposed estimator and the consistency of weights. Simulation studies and an example in analyzing a real dataset demonstrate the effectiveness of the proposed method. Ziwen Gao, Dalei Yu, Xinyu Zhang 0024 |
IEEE Trans. Reliab. | 3 |
| 2023 | Optimal Parameter-Transfer Learning by Semiparametric Model AveragingabstractIn this article, we focus on prediction of a target model by transferring the information of source models. To be flexible, we use semiparametric additive frameworks for the target and source models. Inheriting the spirit of parameter-transfer learning, we assume that different models possibly share common knowledge across parametric components that is helpful for the target predictive task. Unlike existing parameter-transfer approaches, which need to construct auxiliary source models by parameter similarity with the target model and then adopt a regularization procedure, we propose a frequentist model averaging strategy with a $J$-fold cross-validation criterion so that auxiliary parameter information from different models can be adaptively transferred through data-driven weight assignments. The asymptotic optimality and weight convergence of our proposed method are built under some regularity conditions. Extensive numerical results demonstrate the superiority of the proposed method over competitive methods. Xiaonan Hu, Xinyu Zhang 0024 |
J. Mach. Learn. Res. | 2 |
| 2021 | Reducing Simulation Input-Model Risk via Input Model AveragingabstractInput uncertainty is an aspect of simulation model risk that arises when the driving input distributions are derived or “fit” to real-world, historical data. Although there has been significant progress on quantifying and hedging against input uncertainty, there has been no direct attempt to reduce it via better input modeling. The meaning of “better” depends on the context and the objective: Our context is when (a) there are one or more families of parametric distributions that are plausible choices; (b) the real-world historical data are not expected to perfectly conform to any of them; and (c) our primary goal is to obtain higher-fidelity simulation output rather than to discover the “true” distribution. In this paper, we show that frequentist model averaging can be an effective way to create input models that better represent the true, unknown input distribution, thereby reducing model risk. Input model averaging builds from standard input modeling practice, is not computationally burdensome, requires no change in how the simulation is executed nor any follow-up experiments, and is available on the Comprehensive R Archive Network (CRAN). We provide theoretical and empirical support for our approach. Barry L. Nelson, Alan T. K. Wan, Guohua Zou, Xinyu Zhang 0024, Xi Jiang 0003 |
INFORMS J. Comput. | 4 |