Jinchi Lv

dblp:96/9709 · DBLP profile ↗
← Back
10ranked-venue papers
0as first author
3since 2021 · last 2026
0000-0002-5881-9591ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 1 since 2021Theory of computation · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
YearPublicationVenuePosition
2026 Asymptotic Theory of Eigenvectors for Latent Embeddings With Generalized Laplacian Matrices
abstract
Laplacian matrices are widely used in practice to capture latent structural information in data, ranging from graphs to manifolds. Their normalization forms naturally induce dependencies among matrix entries, and such dependencies are known to pose significant challenges for advances in random matrix theory (RMT). Motivated by this, we introduce a general class of generalized Laplacian matrices, which includes both the standard Laplacian and random adjacency matrices as special cases, and we develop a new framework—Asymptotic Theory of Eigenvectors for latent embeddings with Generalized Laplacian matrices (ATE-GL)—for studying their spectral properties. Our theory is driven by two key ingredients: the use of generalized quadratic vector equations to handle dependency in RMT, and refined high-order asymptotic expansions for empirical spiked eigenvectors and eigenvalues based on local laws. These results lead to asymptotic normality for both spiked eigenvectors and eigenvalues, enabling precise statistical inference and uncertainty quantification for a broad class of applications involving generalized Laplacian matrices. We also discuss two motivating applications of the ATE-GL framework and demonstrate its effectiveness through numerical examples.
Jianqing Fan, Jinchi Lv, Fan Yang 0106, Diwen Yu
IEEE Trans. Inf. Theory3
2025 Precise Asymptotics and Refined Regret of Variance-Aware UCB
abstract
In this paper, we study the behavior of the Upper Confidence Bound-Variance (UCB-V) algorithm for the Multi-Armed Bandit (MAB) problems, a variant of the canonical Upper Confidence Bound (UCB) algorithm that incorporates variance estimates into its decision-making process. More precisely, we provide an asymptotic characterization of the arm-pulling rates for UCB-V, extending recent results for the canonical UCB in Kalvit and Zeevi (2021) and Khamaru and Zhang (2024). In an interesting contrast to the canonical UCB, our analysis reveals that the behavior of UCB-V can exhibit instability, meaning that the arm-pulling rates may not always be asymptotically deterministic. Besides the asymptotic characterization, we also provide non-asymptotic bounds for the arm-pulling rates in the high probability regime, offering insights into the regret analysis. As an application of this high probability result, we establish that UCB-V can achieve a more refined regret bound, previously unknown even for more complicate and advanced variance-aware online decision-making algorithms. A matching regret lower bound is also established, demonstrating the optimality of our result.
Jinchi Lv, Xiaocong Xu, Zhengyuan Zhou
NeurIPS3
2025 DeepDeconUQ estimates malignant cell fraction prediction intervals in bulk RNA-seq tissue
abstract
Accurate estimation of malignant cell fractions in tissues plays a critical role in cancer diagnosis, prognosis, and subsequent treatment decisions. However, most currently available methods provide only point estimates, neglecting the quantification of uncertainties, which is essential for both clinical and research applications. This study introduces DeepDeconUQ, a deep neural network model developed to estimate prediction intervals for malignant cell fractions based on bulk RNA-seq data. This approach addresses limitations in current malignant cell fraction estimation methods by integrating uncertainty quantification into predictions of cancer cell fractions. DeepDeconUQ leverages single-cell RNA sequencing (scRNA-seq) data in conjunction with conformalized quantile regression to produce reliable prediction intervals. The model trains a quantile regression neural network to establish upper and lower bounds for cancer cell proportions, followed by a calibration step that refines these intervals to ensure both statistical validity (coverage probability) and discrimination (narrow intervals). Benchmark analyses indicate that DeepDeconUQ consistently surpasses existing methods, achieving high coverage accuracy with tight prediction intervals across simulated and real cancer datasets. The robustness of DeepDeconUQ is further demonstrated by its resilience to various gene expression perturbations. The DeepDeconUQ method is publicly accessible at https://github.com/jiaweih14/DeepDeconUQ.
Kevin R. Kelly, Jinchi Lv, Jiang F. Zhong, Fengzhu Sun
PLoS Comput. Biol.4
2019 Nonuniformity of P-values Can Occur Early in Diverging Dimensions
abstract
Evaluating the joint significance of covariates is of fundamental importance in a wide range of applications. To this end, p-values are frequently employed and produced by algorithms that are powered by classical large-sample asymptotic theory. It is well known that the conventional p-values in Gaussian linear model are valid even when the dimensionality is a non-vanishing fraction of the sample size, but can break down when the design matrix becomes singular in higher dimensions or when the error distribution deviates from Gaussianity. A natural question is when the conventional p-values in generalized linear models become invalid in diverging dimensions. We establish that such a breakdown can occur early in nonlinear models. Our theoretical characterizations are confirmed by simulation studies.
Emre Demirkaya, Jinchi Lv
J. Mach. Learn. Res.3
2019 Scalable Interpretable Multi-Response Regression via SEED
abstract
Sparse reduced-rank regression is an important tool for uncovering meaningful dependence structure between large numbers of predictors and responses in many big data applications such as genome-wide association studies and social media analysis. Despite the recent theoretical and algorithmic advances, scalable estimation of sparse reduced-rank regression remains largely unexplored. In this paper, we suggest a scalable procedure called sequential estimation with eigen-decomposition (SEED) which needs only a single top-$r$ sparse singular value decomposition from a generalized eigenvalue problem to find the optimal low-rank and sparse matrix estimate. Our suggested method is not only scalable but also performs simultaneous dimensionality reduction and variable selection. Under some mild regularity conditions, we show that SEED enjoys nice sampling properties including consistency in estimation, rank selection, prediction, and model selection. Moreover, SEED employs only basic matrix operations that can be efficiently parallelized in high performance computing devices. Numerical studies on synthetic and real data sets show that SEED outperforms the state-of-the-art approaches for large-scale matrix estimation problem.
Zemin Zheng, Mohammad Taha Bahadori, Yan Liu 0002, Jinchi Lv
J. Mach. Learn. Res.4
2019 SOFAR: Large-Scale Association Network Learning
abstract
Many modern big data applications feature large scale in both numbers of responses and predictors. Better statistical efficiency and scientific insights can be enabled by understanding the large-scale response-predictor association network structures via layers of sparse latent factors ranked by importance. Yet sparsity and orthogonality have been two largely incompatible goals. To accommodate both features, in this paper we suggest the method of sparse orthogonal factor regression (SOFAR) via the sparse singular value decomposition with orthogonality constrained optimization to learn the underlying association networks, with broad applications to both unsupervised and supervised learning tasks such as biclustering with sparse singular value decomposition, sparse principal component analysis, sparse factor analysis, and spare vector autoregression analysis. Exploiting the framework of convexity-assisted nonconvex optimization, we derive nonasymptotic error bounds for the suggested procedure characterizing the theoretical advantages. The statistical guarantees are powered by an efficient SOFAR algorithm with convergence property. Both computational and theoretical advantages of our procedure are demonstrated with several simulations and real data examples.
Yoshimasa Uematsu, Kun Chen 0002, Jinchi Lv, Wei Lin 0020
IEEE Trans. Inf. Theory4
2018 DeepPINK: reproducible feature selection in deep neural networks
abstract
Deep learning has become increasingly popular in both supervised and unsupervised machine learning thanks to its outstanding empirical performance. However, because of their intrinsic complexity, most deep learning methods are largely treated as black box tools with little interpretability. Even though recent attempts have been made to facilitate the interpretability of deep neural networks (DNNs), existing methods are susceptible to noise and lack of robustness. Therefore, scientists are justifiably cautious about the reproducibility of the discoveries, which is often related to the interpretability of the underlying statistical models. In this paper, we describe a method to increase the interpretability and reproducibility of DNNs by incorporating the idea of feature selection with controlled error rate. By designing a new DNN architecture and integrating it with the recently proposed knockoffs framework, we perform feature selection with a controlled error rate, while maintaining high power. This new method, DeepPINK (Deep feature selection using Paired-Input Nonlinear Knockoffs), is applied to both simulated and real data sets to demonstrate its empirical utility.
Yang Young Lu, Jinchi Lv, William Stafford Noble
NeurIPS3
2016 Estimating and testing high-dimensional mediation effects in epigenetic studies
abstract
MOTIVATION: High-dimensional DNA methylation markers may mediate pathways linking environmental exposures with health outcomes. However, there is a lack of analytical methods to identify significant mediators for high-dimensional mediation analysis. RESULTS: Based on sure independent screening and minimax concave penalty techniques, we use a joint significance test for mediation effect. We demonstrate its practical performance using Monte Carlo simulation studies and apply this method to investigate the extent to which DNA methylation markers mediate the causal pathway from smoking to reduced lung function in the Normative Aging Study. We identify 2 CpGs with significant mediation effects. AVAILABILITY AND IMPLEMENTATION: R package, source code, and simulation study are available at https://github.com/YinanZheng/HIMA CONTACT: [email protected].
Yinan Zheng, Zhou Zhang 0002, Brian Joyce, Grace Yoon, Wei Zhang 0253, Joel Schwartz, Allan C. Just, Elena Colicino, Pantel Vokonas, Lihui Zhao, Jinchi Lv, Andrea A. Baccarelli, Lifang Hou, Lei Liu 0004
Bioinform.13
2016 The Constrained Dantzig Selector with Enhanced Consistency
abstract
The Dantzig selector has received popularity for many applications such as compressed sensing and sparse modeling, thanks to its computational efficiency as a linear programming problem and its nice sampling properties. Existing results show that it can recover sparse signals mimicking the accuracy of the ideal procedure, up to a logarithmic factor of the dimensionality. Such a factor has been shown to hold for many regularization methods. An important question is whether this factor can be reduced to a logarithmic factor of the sample size in ultra-high dimensions under mild regularity conditions. To provide an affirmative answer, in this paper we suggest the constrained Dantzig selector, which has more flexible constraints and parameter space. We prove that the suggested method can achieve convergence rates within a logarithmic factor of the sample size of the oracle rates and improved sparsity, under a fairly weak assumption on the signal strength. Such improvement is significant in ultra-high dimensions. This method can be implemented efficiently through sequential linear programming. Numerical studies confirm that the sample size needed for a certain level of accuracy in these problems can be much reduced.
Yinfei Kong, Zemin Zheng, Jinchi Lv
J. Mach. Learn. Res.3
2011 Nonconcave Penalized Likelihood With NP-Dimensionality
abstract
Penalized likelihood methods are fundamental to ultra-high dimensional variable selection. How high dimensionality such methods can handle remains largely unknown. In this paper, we show that in the context of generalized linear models, such methods possess model selection consistency with oracle properties even for dimensionality of Non-Polynomial (NP) order of sample size, for a class of penalized likelihood approaches using folded-concave penalty functions, which were introduced to ameliorate the bias problems of convex penalty functions. This fills a long-standing gap in the literature where the dimensionality is allowed to grow slowly with the sample size. Our results are also applicable to penalized likelihood with the L(1)-penalty, which is a convex function at the boundary of the class of folded-concave penalty functions under consideration. The coordinate optimization is implemented for finding the solution paths, whose performance is evaluated by a few simulation examples and the real data analysis.
Jianqing Fan, Jinchi Lv
IEEE Trans. Inf. Theory2