Yang Li 0175

dblp:37/4190-175 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2025
0000-0002-1202-1082ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Theory of computation · 2 · 2 since 2021
YearPublicationVenuePosition
2025 Reproducible Feature Selection for High-Dimensional Measurement Error Models
abstract
The literature has witnessed an upsurge of interest in dealing with corrupted data in diverse operations research and optimization applications. Despite the substantial progress of feature selection, how to control the false discovery rate (FDR) under measurement errors remains largely unexplored, especially in the knockoffs framework. In this paper, we extend the recently developed knockoff procedures designed for clean data sets to deal with corrupted data. To be specific, we propose a new method called the double projection knockoff filter (DP-knockoff) for reproducible feature selection under additive measurement errors in the high-dimensional setup. Our key contribution is to show that the FDR of the proposed DP-knockoff can be asymptotically controlled within a user-specified level. This is nontrivial because there is no way to obtain the exact knockoff copies due to the unobservable measurement errors. We address this issue by resorting to certain bias-corrected test statistics. Our numerical studies and real data analysis demonstrate the effectiveness of the proposed procedure. History: Accepted by Ram Ramesh, Area Editor for Data Science & Machine Learning. Funding: Financial support from the National Key Research and Development Program of China [Grant 2022YFA1008000], the Natural Science Foundation of China [Grants 11671374, 12101584, 71731010, 71921001, and 72071187], the Fundamental Research Funds for the Central Universities [Grants WK3470000017 and WK2040000047], the Doctoral Research Start-up Funds Projects of Anhui University [Grant S020318033/005], and the University Natural Science Research Project of Anhui Province [Grant 2023AH050101] is gratefully acknowledged. Supplemental Material: The software that supports the findings of this study is available within the paper and its Supplemental Information ( https://pubsonline.informs.org/doi/suppl/10.1287/ijoc.2023.0282 ), as well as from the IJOC GitHub software repository ( https://github.com/INFORMSJoC/2023.0282 ). The complete IJOC Software and Data Repository is available at https://informsjoc.github.io/ .
Yang Li 0175, Zemin Zheng, Jie Wu 0004
INFORMS J. Comput.2
2022 L0-Regularized Learning for High-Dimensional Additive Hazards Regression
abstract
Sparse learning in high-dimensional survival analysis is of great practical importance, as exemplified by modern applications in credit risk analysis and high-throughput genomic data analysis. In this article, we consider the L0-regularized learning for simultaneous variable selection and estimation under the framework of additive hazards models and utilize the idea of primal dual active sets to develop an algorithm targeted at solving the traditionally nonpolynomial time optimization problem. Under interpretable conditions, comprehensive statistical properties, including model selection consistency, oracle inequalities under various estimation losses, and the oracle property, are established for the global optimizer of the proposed approach. Moreover, our theoretical analysis for the algorithmic solution reveals that the proposed L0-regularized learning can be more efficient than other regularization methods in that it requests a smaller sample size as well as a lower minimum signal strength to identify the significant features. The effectiveness of the proposed method is evidenced by simulation studies and real-data analysis. Summary of Contribution: Feature selection is a fundamental statistical learning technique under high dimensions and routinely encountered in various areas, including operations research and computing. This paper focuses on the L0-regularized learning for feature selection in high-dimensional additive hazards regression. The matching algorithm for solving the nonconvex L0-constrained problem is scalable and enjoys comprehensive theoretical properties.
Zemin Zheng, Yang Li 0175
INFORMS J. Comput.3