EDBT 2026 Demo / reviewers in the wild / expert
Zemin Zheng
dblp:203/7089
· DBLP profile ↗
6ranked-venue papers
2as first author
4since 2021 · last 2025
0000-0002-0240-9411ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 1 first-author · 1 since 2021Theory of computation · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Robust Parallel Pursuit for Large-Scale Association Network LearningabstractSparse reduced-rank regression is an important tool to uncover the large-scale response-predictor association network, as exemplified by modern applications such as the diffusion networks, and recommendation systems. However, the association networks recovered by existing methods are either sensitive to outliers or not scalable under the big data setup. In this paper, we propose a new statistical learning method called robust parallel pursuit (ROP) for joint estimation and outlier detection in large-scale response-predictor association network analysis. The proposed method is scalable in that it transforms the original large-scale network learning problem into a set of sparse unit-rank estimations via factor analysis, thus facilitating an effective parallel pursuit algorithm. Furthermore, we provide comprehensive theoretical guarantees including consistency in parameter estimation, rank selection, and outlier detection, and we conduct an inference procedure to quantify the uncertainty of existence of outliers. Extensive simulation studies and two real-data analyses demonstrate the effectiveness and the scalability of the suggested approach. History: Accepted by Ram Ramesh, Area Editor/Data Science & Machine Learning. Funding: This work was supported by the National Key R&D Program of China [Grant 2022YFA1008000], Natural Science Foundation of China [Grants 72071187, 72091212, 71731010, and 71921001], China Postdoctoral Science Foundation [Grant 2023M733402], and Fundamental Research Funds for the Central Universities [Grants WK3470000017, WK2040000027, and WK2040000079]. Supplemental Material: The software that supports the findings of this study is available within the paper and its Supplemental Information ( https://pubsonline.informs.org/doi/suppl/10.1287/ijoc.2022.0181 ) as well as from the IJOC GitHub software repository ( https://github.com/INFORMSJoC/2022.0181 ). The complete IJOC Software and Data Repository is available at https://informsjoc.github.io/ . Ruipeng Dong, Zemin Zheng |
INFORMS J. Comput. | 4 |
| 2025 | Reproducible Feature Selection for High-Dimensional Measurement Error ModelsabstractThe literature has witnessed an upsurge of interest in dealing with corrupted data in diverse operations research and optimization applications. Despite the substantial progress of feature selection, how to control the false discovery rate (FDR) under measurement errors remains largely unexplored, especially in the knockoffs framework. In this paper, we extend the recently developed knockoff procedures designed for clean data sets to deal with corrupted data. To be specific, we propose a new method called the double projection knockoff filter (DP-knockoff) for reproducible feature selection under additive measurement errors in the high-dimensional setup. Our key contribution is to show that the FDR of the proposed DP-knockoff can be asymptotically controlled within a user-specified level. This is nontrivial because there is no way to obtain the exact knockoff copies due to the unobservable measurement errors. We address this issue by resorting to certain bias-corrected test statistics. Our numerical studies and real data analysis demonstrate the effectiveness of the proposed procedure. History: Accepted by Ram Ramesh, Area Editor for Data Science & Machine Learning. Funding: Financial support from the National Key Research and Development Program of China [Grant 2022YFA1008000], the Natural Science Foundation of China [Grants 11671374, 12101584, 71731010, 71921001, and 72071187], the Fundamental Research Funds for the Central Universities [Grants WK3470000017 and WK2040000047], the Doctoral Research Start-up Funds Projects of Anhui University [Grant S020318033/005], and the University Natural Science Research Project of Anhui Province [Grant 2023AH050101] is gratefully acknowledged. Supplemental Material: The software that supports the findings of this study is available within the paper and its Supplemental Information ( https://pubsonline.informs.org/doi/suppl/10.1287/ijoc.2023.0282 ), as well as from the IJOC GitHub software repository ( https://github.com/INFORMSJoC/2023.0282 ). The complete IJOC Software and Data Repository is available at https://informsjoc.github.io/ . Yang Li 0175, Zemin Zheng, Jie Wu 0004 |
INFORMS J. Comput. | 3 |
| 2022 | L0-Regularized Learning for High-Dimensional Additive Hazards RegressionabstractSparse learning in high-dimensional survival analysis is of great practical importance, as exemplified by modern applications in credit risk analysis and high-throughput genomic data analysis. In this article, we consider the L0-regularized learning for simultaneous variable selection and estimation under the framework of additive hazards models and utilize the idea of primal dual active sets to develop an algorithm targeted at solving the traditionally nonpolynomial time optimization problem. Under interpretable conditions, comprehensive statistical properties, including model selection consistency, oracle inequalities under various estimation losses, and the oracle property, are established for the global optimizer of the proposed approach. Moreover, our theoretical analysis for the algorithmic solution reveals that the proposed L0-regularized learning can be more efficient than other regularization methods in that it requests a smaller sample size as well as a lower minimum signal strength to identify the significant features. The effectiveness of the proposed method is evidenced by simulation studies and real-data analysis. Summary of Contribution: Feature selection is a fundamental statistical learning technique under high dimensions and routinely encountered in various areas, including operations research and computing. This paper focuses on the L0-regularized learning for feature selection in high-dimensional additive hazards regression. The matching algorithm for solving the nonconvex L0-constrained problem is scalable and enjoys comprehensive theoretical properties. Zemin Zheng, Yang Li 0175 |
INFORMS J. Comput. | 1 |
| 2022 | Fast Stagewise Sparse Factor RegressionabstractSparse factorization of a large matrix is fundamental in modern statistical learning. In particular, the sparse singular value decomposition has been utilized in many multivariate regression methods. The appeal of this factorization is owing to its power in discovering a highly-interpretable latent association network. However, many existing methods are either ad hoc without a general performance guarantee, or are computationally intensive. We formulate the statistical problem as a sparse factor regression and tackle it with a two-stage “deflation + stagewise learning” approach. In the first stage, we consider both sequential and parallel approaches for simplifying the task into a set of co-sparse unit-rank estimation (CURE) problems, and establish the statistical underpinnings of these commonly-adopted and yet poorly understood deflation methods. In the second stage, we innovate a contended stagewise learning technique, consisting of a sequence of simple incremental updates, to efficiently trace out the whole solution paths of CURE. Our algorithm achieves a much lower computational complexity than alternating convex search, and it enables a flexible and principled tradeoff between statistical accuracy and computational efficiency. Our work is among the first to enable stagewise learning for non-convex problems, and the idea can be applicable in many multi-convex problems. Extensive simulation studies and an application in genetics demonstrate the effectiveness and scalability of our approach. Kun Chen 0002, Ruipeng Dong, Wanwan Xu, Zemin Zheng |
J. Mach. Learn. Res. | 4 |
| 2019 | Scalable Interpretable Multi-Response Regression via SEEDabstractSparse reduced-rank regression is an important tool for uncovering meaningful dependence structure between large numbers of predictors and responses in many big data applications such as genome-wide association studies and social media analysis. Despite the recent theoretical and algorithmic advances, scalable estimation of sparse reduced-rank regression remains largely unexplored. In this paper, we suggest a scalable procedure called sequential estimation with eigen-decomposition (SEED) which needs only a single top-$r$ sparse singular value decomposition from a generalized eigenvalue problem to find the optimal low-rank and sparse matrix estimate. Our suggested method is not only scalable but also performs simultaneous dimensionality reduction and variable selection. Under some mild regularity conditions, we show that SEED enjoys nice sampling properties including consistency in estimation, rank selection, prediction, and model selection. Moreover, SEED employs only basic matrix operations that can be efficiently parallelized in high performance computing devices. Numerical studies on synthetic and real data sets show that SEED outperforms the state-of-the-art approaches for large-scale matrix estimation problem. Zemin Zheng, Mohammad Taha Bahadori, Yan Liu 0002, Jinchi Lv |
J. Mach. Learn. Res. | 1 |
| 2016 | The Constrained Dantzig Selector with Enhanced ConsistencyabstractThe Dantzig selector has received popularity for many applications such as compressed sensing and sparse modeling, thanks to its computational efficiency as a linear programming problem and its nice sampling properties. Existing results show that it can recover sparse signals mimicking the accuracy of the ideal procedure, up to a logarithmic factor of the dimensionality. Such a factor has been shown to hold for many regularization methods. An important question is whether this factor can be reduced to a logarithmic factor of the sample size in ultra-high dimensions under mild regularity conditions. To provide an affirmative answer, in this paper we suggest the constrained Dantzig selector, which has more flexible constraints and parameter space. We prove that the suggested method can achieve convergence rates within a logarithmic factor of the sample size of the oracle rates and improved sparsity, under a fairly weak assumption on the signal strength. Such improvement is significant in ultra-high dimensions. This method can be implemented efficiently through sequential linear programming. Numerical studies confirm that the sample size needed for a certain level of accuracy in these problems can be much reduced. Yinfei Kong, Zemin Zheng, Jinchi Lv |
J. Mach. Learn. Res. | 2 |