Qizhai Li

dblp:81/9004 · DBLP profile ↗
← Back
10ranked-venue papers
0as first author
7since 2021 · last 2026
0000-0002-3325-8265ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 5 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
YearPublicationVenuePosition
2026 ZA-Net: A universal zero-annotation nuclei segmentation network for pathology images via vision-language pre-trained model
Fuqiang Chen, Kun Ru, Miaoxia He, Qizhai Li, Yao Pu, Jing Cai 0001, Wenjian Qin
Pattern Recognit. Lett.6
2025 Composition-on-composition regression analysis for multi-omics integration of metagenomic data
abstract
MOTIVATION: Compositional data are frequently encountered in many disciplines, such as in next-generation sequencing experiments widely used in biomedical studies. Regression analysis with compositional data as either responses or predictors has been well studied. However, when both responses and predictors are compositional, the inventory of analysis tools is surprisingly limited, especially in the high-dimensional setting. Among the few existing methods, most of them rely on a log-ratio transformation to move compositional data from the simplex to real numbers. Yet, a serious weakness of these methods is their failure to handle the substantial fraction of zeroes observed in data collected from next-generation sequencing experiments. RESULTS: To investigate associations between two high-dimensional multi-omics compositions, we propose a composition-on-composition (COC) regression analysis method which does not require log-ratio transformations and hence can handle zeroes in the data. To account for high dimensionality, we estimate regression coefficients using a penalized estimation equation approach. Finally, inference procedures for COC regression are also proposed. Superior performance of COC is demonstrated through both comprehensive numerical simulations and case studies. AVAILABILITY AND IMPLEMENTATION: Source R codes to implement COC method is available at https://github.com/nrios4/COC.
Nicholas Rios, Yuke Shi, Jun Chen 0040, Xiang Zhan, Lingzhou Xue, Qizhai Li
Bioinform.6
2025 The sequence kernel association test for the proportional odds model
abstract
MOTIVATION: The Sequence Kernel Association Test (SKAT) and its extensions are the most popular methods for studying the association between phenotypes and a set of single nucleotide polymorphisms. Their practical application is very wide, but most of these methods are designed for continuous and binary phenotypes. Ordered categorical phenotypes are also very common in practice, so there is an urgent need to develop SKAT-type tests for proportional odds model. RESULTS: To accommodate ordered categorical phenotypes, we propose a test named the Sequence Kernel Association Test for the Proportional Odds Model (POM-SKAT). It constructs a score test for the variance of the coefficients of interest using a quasi-likelihood and the P-value is evaluated by approximating the asymptotic distribution of the test statistic with the Pearson Type III distribution. Simulation studies demonstrate that our method performs well and achieves high power in detecting gene-phenotype associations. We apply POM-SKAT to rheumatoid arthritis data provided by Genetic Analysis Workshop 16, identifying multiple relevant gene variants. AVAILABILITY AND IMPLEMENTATION: Code is available at GitHub (https://github.com/amss-stat/POM-SKAT).
Jingxin Yan, Shuying Wang, Jinjuan Wang, Qizhai Li
Bioinform.5
2024 A Bayesian Approach Toward Robust Multidimensional Ellipsoid-Specific Fitting
abstract
This work presents a novel and effective method for fitting multidimensional ellipsoids (i.e., ellipsoids embedded in [Formula: see text]) to scattered data in the contamination of noise and outliers. Unlike conventional algebraic or geometric fitting paradigms that assume each measurement point is a noisy version of its nearest point on the ellipsoid, we approach the problem as a Bayesian parameter estimate process and maximize the posterior probability of a certain ellipsoidal solution given the data. We establish a more robust correlation between these points based on the predictive distribution within the Bayesian framework, i.e., considering each model point as a potential source for generating each measurement. Concretely, we incorporate a uniform prior distribution to constrain the search for primitive parameters within an ellipsoidal domain, ensuring ellipsoid-specific results regardless of inputs. We then establish the connection between measurement point and model data via Bayes' rule to enhance the method's robustness against noise. Due to independent of spatial dimensions, the proposed method not only delivers high-quality fittings to challenging elongated ellipsoids but also generalizes well to multidimensional spaces. To address outlier disturbances, often overlooked by previous approaches, we further introduce a uniform distribution on top of the predictive distribution to significantly enhance the algorithm's robustness against outliers. Thanks to the uniform prior, our maximum a posterior probability coincides with a more tractable maximum likelihood estimation problem, which is subsequently solved by a numerically stable Expectation Maximization (EM) framework. Moreover, we introduce an ε-accelerated technique to expedite the convergence of EM considerably. We also investigate the relationship between our algorithm and conventional least-squares-based ones, during which we theoretically prove our method's superior robustness. To the best of our knowledge, this is the first comprehensive method capable of performing multidimensional ellipsoid-specific fitting within the Bayesian optimization paradigm under diverse disturbances. We evaluate it across lower and higher dimensional spaces in the presence of heavy noise, outliers, and substantial variations in axis ratios. Also, we apply it to a wide range of practical applications such as microscopy cell counting, 3D reconstruction, geometric shape approximation, and magnetometer calibration tasks. In all these test contexts, our method consistently delivers flexible, robust, ellipsoid-specific performance, and achieves the state-of-the-art results.
Mingyang Zhao 0001, Xiaohong Jia 0001, Lei Ma 0008, Yuke Shi, Jingen Jiang 0001, Qizhai Li, Dong-Ming Yan 0001, Tiejun Huang 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2023 Summary statistics-based association test for identifying the pleiotropic effects with set of genetic variants
abstract
MOTIVATION: Traditional genome-wide association study focuses on testing one-to-one relationship between genetic variants and complex human diseases or traits. While its success in the past decade, this one-to-one paradigm lacks efficiency because it does not utilize the information of intrinsic genetic structure and pleiotropic effects. Due to privacy reasons, only summary statistics of current genome-wide association study data are publicly available. Existing summary statistics-based association tests do not consider covariates for regression model, while adjusting for covariates including population stratification factors is a routine issue. RESULTS: In this work, we first derive the correlation coefficients between summary Wald statistics obtained from linear regression model with covariates. Then, a new test is proposed by integrating three-level information including the intrinsic genetic structure, pleiotropy, and the potential information combinations. Extensive simulations demonstrate that the proposed test outperforms three other existing methods under most of the considered scenarios. Real data analysis of polyunsaturated fatty acids further shows that the proposed test can identify more genes than the compared existing methods. AVAILABILITY AND IMPLEMENTATION: Code is available at https://github.com/bschilder/ThreeWayTest.
Deliang Bu, Qizhai Li
Bioinform.3
2023 A maximum kernel-based association test to detect the pleiotropic genetic effects on multiple phenotypes
abstract
MOTIVATION: Testing the association between multiple phenotypes with a set of genetic variants simultaneously, rather than analyzing one trait at a time, is receiving increasing attention for its high statistical power and easy explanation on pleiotropic effects. The kernel-based association test (KAT), being free of data dimensions and structures, has proven to be a good alternative method for genetic association analysis with multiple phenotypes. However, KAT suffers from substantial power loss when multiple phenotypes have moderate to strong correlations. To handle this issue, we propose a maximum KAT (MaxKAT) and suggest using the generalized extreme value distribution to calculate its statistical significance under the null hypothesis. RESULTS: We show that MaxKAT reduces computational intensity greatly while maintaining high accuracy. Extensive simulations demonstrate that MaxKAT can properly control type I error rates and obtain remarkably higher power than KAT under most of the considered scenarios. Application to a porcine dataset used in biomedical experiments of human disease further illustrates its practical utility. AVAILABILITY AND IMPLEMENTATION: The R package MaxKAT that implements the proposed method is available on Github https://github.com/WangJJ-xrk/MaxKAT.
Jinjuan Wang, Mingya Long, Qizhai Li
Bioinform.3
2022 An adaptive direction-assisted test for microbiome compositional data
abstract
MOTIVATION: Microbial communities have been shown to be associated with many complex diseases, such as cancers and cardiovascular diseases. The identification of differentially abundant taxa is clinically important. It can help understand the pathology of complex diseases, and potentially provide preventive and therapeutic strategies. Appropriate differential analyses for microbiome data are challenging due to its unique data characteristics including compositional constraint, excessive zeros and high dimensionality. Most existing approaches either ignore these data characteristics or only account for the compositional constraint by using log-ratio transformations with zero observations replaced by a pseudocount. However, there is no consensus on how to choose a pseudocount. More importantly, ignoring the characteristic of excessive zeros may result in poorly powered analyses and therefore yield misleading findings. RESULTS: We develop a novel microbiome-based direction-assisted test for the detection of overall difference in microbial relative abundances between two health conditions, which simultaneously incorporates the characteristics of relative abundance data. The proposed test (i) divides the taxa into two clusters by the directions of mean differences of relative abundances and then combines them at cluster level, in light of the compositional characteristic; and (ii) contains a burden type test, which collapses multiple taxa into a single one to account for excessive zeros. Moreover, the proposed test is an adaptive procedure, which can accommodate high-dimensional settings and yield high power against various alternative hypotheses. We perform extensive simulation studies across a wide range of scenarios to evaluate the proposed test and show its substantial power gain over some existing tests. The superiority of the proposed approach is further demonstrated with real datasets from two microbiome studies. AVAILABILITY AND IMPLEMENTATION: An R package for MiDAT is available at https://github.com/zhangwei0125/MiDAT. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Aiyi Liu, Guanjie Chen, Qizhai Li
Bioinform.5
2016 Group-combined P-values with applications to genetic association studies
abstract
MOTIVATION: In large-scale genetic association studies with tens of hundreds of single nucleotide polymorphisms (SNPs) genotyped, the traditional statistical framework of logistic regression using maximum likelihood estimator (MLE) to infer the odds ratios of SNPs may not work appropriately. This is because a large number of odds ratios need to be estimated, and the MLEs may be not stable when some of the SNPs are in high linkage disequilibrium. Under this situation, the P-value combination procedures seem to provide good alternatives as they are constructed on the basis of single-marker analysis. RESULTS: The commonly used P-value combination methods (such as the Fisher's combined test, the truncated product method, the truncated tail strength and the adaptive rank truncated product) may lose power when the significance level varies across SNPs. To tackle this problem, a group combined P-value method (GCP) is proposed, where the P-values are divided into multiple groups and then are combined at the group level. With this strategy, the significance values are integrated at different levels, and the power is improved. Simulation shows that the GCP can effectively control the type I error rates and have additional power over the existing methods-the power increase can be as high as over 50% under some situations. The proposed GCP method is applied to data from the Genetic Analysis Workshop 16. Among all the methods, only the GCP and ARTP can give the significance to identify a genomic region covering gene DSC3 being associated with rheumatoid arthritis, but the GCP provides smaller P-value. AVAILABILITY AND IMPLEMENTATION: http://www.statsci.amss.ac.cn/yjscy/yjy/lqz/201510/t20151027_313273.html CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Xiaonan Hu, Sanguo Zhang, Shuangge Ma, Qizhai Li
Bioinform.5
2016 A two-phase procedure for non-normal quantitative trait genetic association study
abstract
BACKGROUND: The nonparametric trend test (NPT) is well suitable for identifying the genetic variants associated with quantitative traits when the trait values do not satisfy the normal distribution assumption. If the genetic model, defined according to the mode of inheritance, is known, the NPT derived under the given genetic model is optimal. However, in practice, the genetic model is often unknown beforehand. The NPT derived from an uncorrected model might result in loss of power. When the underlying genetic model is unknown, a robust test is preferred to maintain satisfactory power. RESULTS: We propose a two-phase procedure to handle the uncertainty of the genetic model for non-normal quantitative trait genetic association study. First, a model selection procedure is employed to help choose the genetic model. Then the optimal test derived under the selected model is constructed to test for possible association. To control the type I error rate, we derive the joint distribution of the test statistics developed in the two phases and obtain the proper size. CONCLUSIONS: The proposed method is more robust than existing methods through the simulation results and application to gene DNAH9 from the Genetic Analysis Workshop 16 for associated with Anti-cyclic citrullinated peptide antibody further demonstrate its performance.
Huiyun Li, Zhaohai Li, Qizhai Li
BMC Bioinform.4
2011 Robust joint analysis allowing for model uncertainty in two-stage genetic association studies
abstract
BACKGROUND: The cost efficient two-stage design is often used in genome-wide association studies (GWASs) in searching for genetic loci underlying the susceptibility for complex diseases. Replication-based analysis, which considers data from each stage separately, often suffers from loss of efficiency. Joint test that combines data from both stages has been proposed and widely used to improve efficiency. However, existing joint analyses are based on test statistics derived under an assumed genetic model, and thus might not have robust performance when the assumed genetic model is not appropriate. RESULTS: In this paper, we propose joint analyses based on two robust tests, MERT and MAX3, for GWASs under a two-stage design. We developed computationally efficient procedures and formulas for significant level evaluation and power calculation. The performances of the proposed approaches are investigated through the extensive simulation studies and a real example. Numerical results show that the joint analysis based on the MAX3 test statistic has the best overall performance. CONCLUSIONS: MAX3 joint analysis is the most robust procedure among the considered joint analyses, and we recommend using it in a two-stage genome-wide association study.
Dong-Dong Pan, Qizhai Li, Ningning Jiang, Aiyi Liu, Kai F. Yu
BMC Bioinform.2