EDBT 2026 Demo / reviewers in the wild / expert
Xiang Zhan
dblp:160/3294
· DBLP profile ↗
8ranked-venue papers
2as first author
4since 2021 · last 2025
0000-0001-9650-143XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
6 papers |
Bioinformatics and computational biology · 80% Computational science and engineering · 20% | |
| Theoretical computer science
1 paper |
Mathematical optimization · 100% |
Topics — the 16 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology › computational microbiology
microbiome analysis |
1.4 | 2 | 2025 | TCVS: tree-guided compositional variable selection analysis of microbiome data · Bioinform. 2025 MiRKAT: kernel machine regression-based global association tests for the microbiome · Bioinform. 2021 |
Computational science and engineering
compositional data analysis |
0.9 | 1 | 2025 | Composition-on-composition regression analysis for multi-omics integration of metagenomic data · Bioinform. 2025 |
Bioinformatics and computational biology
metagenomics |
0.9 | 1 | 2025 | Composition-on-composition regression analysis for multi-omics integration of metagenomic data · Bioinform. 2025 |
Bioinformatics and computational biology
multi-omics data integration |
0.9 | 1 | 2025 | Composition-on-composition regression analysis for multi-omics integration of metagenomic data · Bioinform. 2025 |
Computational science and engineering
regression modeling |
0.9 | 1 | 2025 | Composition-on-composition regression analysis for multi-omics integration of metagenomic data · Bioinform. 2025 |
Bioinformatics and computational biology › computational microbiology › microbiome analysis
microbiome association testing |
0.6 | 1 | 2022 | Testing microbiome association using integrated quantile regression models · Bioinform. 2022 |
Bioinformatics and computational biology
quantile regression |
0.6 | 1 | 2022 | Testing microbiome association using integrated quantile regression models · Bioinform. 2022 |
Bioinformatics and computational biology › genomics
genome-wide association study |
0.4 | 1 | 2020 | Prioritizing genetic variants in GWAS with lasso using permutation-assisted tuning · Bioinform. 2020 |
Bioinformatics and computational biology › statistical genetics › variant prioritization
genomic variant prioritization |
0.4 | 1 | 2020 | Prioritizing genetic variants in GWAS with lasso using permutation-assisted tuning · Bioinform. 2020 |
Mathematical optimization
regularization |
0.4 | 1 | 2019 | ET-Lasso: A New Efficient Tuning of Lasso-type Regularization for High-Dimensional Data · KDD 2019 |
Bioinformatics and computational biology › statistical genetics › association analysis
copy number variant association |
0.2 | 1 | 2016 | A novel copy number variants kernel association test with application to autism spectrum disorders studies · Bioinform. 2016 |
Bioinformatics and computational biology › statistical genetics › genetic association study
genetic association testing |
0.2 | 1 | 2016 | A novel copy number variants kernel association test with application to autism spectrum disorders studies · Bioinform. 2016 |
Bioinformatics and computational biology › statistical genetics › association testing
kernel association test |
0.2 | 1 | 2016 | A novel copy number variants kernel association test with application to autism spectrum disorders studies · Bioinform. 2016 |
Bioinformatics and computational biology
statistical genetics |
0.2 | 1 | 2016 | A novel copy number variants kernel association test with application to autism spectrum disorders studies · Bioinform. 2016 |
Data mining › statistical analysis
false discovery rate control |
0.1 | 1 | 2019 | ET-Lasso: A New Efficient Tuning of Lasso-type Regularization for High-Dimensional Data · KDD 2019 |
Data mining › dimensionality reduction
feature selection |
0.1 | 1 | 2019 | ET-Lasso: A New Efficient Tuning of Lasso-type Regularization for High-Dimensional Data · KDD 2019 |
Methods — techniques the papers use, named apart from their topics
kernel machine regression · 1.1tree-guided selection · 0.9penalized estimation equation · 0.9log-ratio transformation · 0.9knockoff filtering · 0.9compositional data analysis · 0.9pseudo-features · 0.8permuted features · 0.8knockoff filter · 0.8cross-validation · 0.8BIC · 0.8integrated quantile regression · 0.6permutation-assisted tuning · 0.4LASSO · 0.4kernel methods · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | TCVS: tree-guided compositional variable selection analysis of microbiome dataabstractMOTIVATION: Studies of microbial communities, represented by the relative abundances of taxa at various taxonomic levels, have underscored the significance of microbiota in numerous aspects of human health and disease. A pivotal challenge in microbiome research lies in pinpointing microbial taxa associated with disease outcomes, which could play crucial roles in prevention, detection, and treatment of various health conditions. Alongside these relative abundance data, taxonomic information sometimes offers a unique lens to explore the impact of shared evolutionary histories on patterns of microbial abundance. RESULTS: In pursuit of this goal, we utilize the tree structure to more flexibly identify taxa associated with disease outcomes. To enhance the accuracy of our selection process, we introduce auxiliary knockoff copies of microbiome features designated as noise. This approach allows for the assessment of false positives in the selection process and aids in refining it towards more precise outcomes. Extensive numerical simulations demonstrate that our methodology outperforms several existing methods in terms of selection accuracy. Furthermore, we demonstrate the practicality of our approach by applying it to a widely used gut microbiome dataset, identifying microbial taxa linked to body mass index. AVAILABILITY AND IMPLEMENTATION: TCVS R code is available at https://github.com/Yicong1225/TCVS. Yicong Mao, Zhiwen Jiang, Tianying Wang, Yijuan Hu, Xiang Zhan |
Bioinform. | 5 |
| 2025 | Composition-on-composition regression analysis for multi-omics integration of metagenomic dataabstractMOTIVATION: Compositional data are frequently encountered in many disciplines, such as in next-generation sequencing experiments widely used in biomedical studies. Regression analysis with compositional data as either responses or predictors has been well studied. However, when both responses and predictors are compositional, the inventory of analysis tools is surprisingly limited, especially in the high-dimensional setting. Among the few existing methods, most of them rely on a log-ratio transformation to move compositional data from the simplex to real numbers. Yet, a serious weakness of these methods is their failure to handle the substantial fraction of zeroes observed in data collected from next-generation sequencing experiments. RESULTS: To investigate associations between two high-dimensional multi-omics compositions, we propose a composition-on-composition (COC) regression analysis method which does not require log-ratio transformations and hence can handle zeroes in the data. To account for high dimensionality, we estimate regression coefficients using a penalized estimation equation approach. Finally, inference procedures for COC regression are also proposed. Superior performance of COC is demonstrated through both comprehensive numerical simulations and case studies. AVAILABILITY AND IMPLEMENTATION: Source R codes to implement COC method is available at https://github.com/nrios4/COC. Nicholas Rios, Yuke Shi, Jun Chen 0040, Xiang Zhan, Lingzhou Xue, Qizhai Li |
Bioinform. | 4 |
| 2022 | Testing microbiome association using integrated quantile regression modelsabstractMOTIVATION: Most existing microbiome association analyses focus on the association between microbiome and conditional mean of health or disease-related outcomes, and within this vein, vast computational tools and methods have been devised for standard binary or continuous outcomes. However, these methods tend to be limited either when the underlying microbiome-outcome association occurs somewhere other than the mean level, or when distribution of the outcome variable is irregular (e.g. zero-inflated or mixtures) such that conditional outcome mean is less meaningful. We address this gap by investigating association analysis between microbiome compositions and conditional outcome quantiles. RESULTS: We introduce a new association analysis tool named MiRKAT-IQ within the Microbiome Regression-based Kernel Association Test framework using Integrated Quantile regression models to examine the association between microbiome and the distribution of outcome. For an individual quantile, we utilize the existing kernel machine regression framework to examine the association between that conditional outcome quantile and a group of microbial features (e.g. microbiome community compositions). Then, the goal of examining microbiome association with the whole outcome distribution is achieved by integrating all outcome conditional quantiles over a process, and thus our new MiRKAT-IQ test is robust to both the location of association signals (e.g. mean, variance, median) and the heterogeneous distribution of the outcome. Extensive numerical simulation studies have been conducted to show the validity of the new MiRKAT-IQ test. We demonstrate the potential usefulness of MiRKAT-IQ with applications to actual biological data collected from a previous microbiome study. AVAILABILITY AND IMPLEMENTATION: R codes to implement the proposed methodology is provided in the MiRKAT package, which is available on CRAN. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Tianying Wang, Wodan Ling, Anna M. Plantinga, Michael C. Wu, Xiang Zhan |
Bioinform. | 5 |
| 2021 | MiRKAT: kernel machine regression-based global association tests for the microbiomeabstractSUMMARY: Distance-based tests of microbiome beta diversity are an integral part of many microbiome analyses. MiRKAT enables distance-based association testing with a wide variety of outcome types, including continuous, binary, censored time-to-event, multivariate, correlated and high-dimensional outcomes. Omnibus tests allow simultaneous consideration of multiple distance and dissimilarity measures, providing higher power across a range of simulation scenarios. Two measures of effect size, a modified R-squared coefficient and a kernel RV coefficient, are incorporated to allow comparison of effect sizes across multiple kernels. AVAILABILITY AND IMPLEMENTATION: MiRKAT is available on CRAN as an R package. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Nehemiah Wilson, Ni Zhao, Xiang Zhan, Hyunwook Koh, Weijia Fu, Hongzhe Li, Michael C. Wu, Anna M. Plantinga |
Bioinform. | 3 |
| 2020 | Prioritizing genetic variants in GWAS with lasso using permutation-assisted tuningabstractMOTIVATION: Large scale genome-wide association studies (GWAS) have resulted in the identification of a wide range of genetic variants related to a host of complex traits and disorders. Despite their success, the individual single-nucleotide polymorphism (SNP) analysis approach adopted in most current GWAS can be limited in that it is usually biologically simple to elucidate a comprehensive genetic architecture of phenotypes and statistically underpowered due to heavy multiple-testing correction burden. On the other hand, multiple-SNP analyses (e.g. gene-based or region-based SNP-set analysis) are usually more powerful to examine the joint effects of a set of SNPs on the phenotype of interest. However, current multiple-SNP approaches can only draw an overall conclusion at the SNP-set level and does not directly inform which SNPs in the SNP-set are driving the overall genotype-phenotype association. RESULTS: In this article, we propose a new permutation-assisted tuning procedure in lasso (plasso) to identify phenotype-associated SNPs in a joint multiple-SNP regression model in GWAS. The tuning parameter of lasso determines the amount of shrinkage and is essential to the performance of variable selection. In the proposed plasso procedure, we first generate permutations as pseudo-SNPs that are not associated with the phenotype. Then, the lasso tuning parameter is delicately chosen to separate true signal SNPs and non-informative pseudo-SNPs. We illustrate plasso using simulations to demonstrate its superior performance over existing methods, and application of plasso to a real GWAS dataset gains new additional insights into the genetic control of complex traits. AVAILABILITY AND IMPLEMENTATION: R codes to implement the proposed methodology is available at https://github.com/xyz5074/plasso. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Songshan Yang, Jiawei Wen, Scott T. Eckert, Yaqun Wang, Dajiang J. Liu, Rongling Wu, Runze Li 0001, Xiang Zhan |
Bioinform. | 8 |
| 2019 | ET-Lasso: A New Efficient Tuning of Lasso-type Regularization for High-Dimensional DataabstractThe $L_1 $ regularization (Lasso) has proven to be a versatile tool to select relevant features and estimate the model coefficients simultaneously and has been widely used in many research areas such as genomes studies, finance, and biomedical imaging. Despite its popularity, it is very challenging to guarantee the feature selection consistency of Lasso especially when the dimension of the data is huge. One way to improve the feature selection consistency is to select an ideal tuning parameter. Traditional tuning criteria mainly focus on minimizing the estimated prediction error or maximizing the posterior model probability, such as cross-validation and BIC, which may either be time-consuming or fail to control the false discovery rate (FDR) when the number of features is extremely large. The other way is to introduce pseudo-features to learn the importance of the original ones. Recently, the Knockoff filter is proposed to control the FDR when performing feature selection. However, its performance is sensitive to the choice of the expected FDR threshold. Motivated by these ideas, we propose a new method using pseudo-features to obtain an ideal tuning parameter. In particular, we present the E fficient T uning of Lasso (ET-Lasso ) to separate active and inactive features by adding permuted features as pseudo-features in linear models. The pseudo-features are constructed to be inactive by nature, which can be used to obtain a cutoff to select the tuning parameter that separates active and inactive features. Experimental studies on both simulations and real-world data applications are provided to show that ET-Lasso can effectively and efficiently select active features under a wide range of scenarios. Songshan Yang, Jiawei Wen, Xiang Zhan, Daniel Kifer |
KDD | 3 |
| 2016 | A novel copy number variants kernel association test with application to autism spectrum disorders studiesabstractMOTIVATION: Copy number variants (CNVs) have been implicated in a variety of neurodevelopmental disorders, including autism spectrum disorders, intellectual disability and schizophrenia. Recent advances in high-throughput genomic technologies have enabled rapid discovery of many genetic variants including CNVs. As a result, there is increasing interest in studying the role of CNVs in the etiology of many complex diseases. Despite the availability of an unprecedented wealth of CNV data, methods for testing association between CNVs and disease-related traits are still under-developed due to the low prevalence and complicated multi-scale features of CNVs. RESULTS: We propose a novel CNV kernel association test (CKAT) in this paper. To address the low prevalence, CNVs are first grouped into CNV regions (CNVR). Then, taking into account the multi-scale features of CNVs, we first design a single-CNV kernel which summarizes the similarity between two CNVs, and next aggregate the single-CNV kernel to a CNVR kernel which summarizes the similarity between two CNVRs. Finally, association between CNVR and disease-related traits is assessed by comparing the kernel-based similarity with the similarity in the trait using a score test for variance components in a random effect model. We illustrate the proposed CKAT using simulations and show that CKAT is more powerful than existing methods, while always being able to control the type I error. We also apply CKAT to a real dataset examining the association between CNV and autism spectrum disorders, which demonstrates the potential usefulness of the proposed method. AVAILABILITY AND IMPLEMENTATION: A R package to implement the proposed CKAT method is available at http://works.bepress.com/debashis_ghosh/ CONTACTS: [email protected] or [email protected] information: Supplementary data are available at Bioinformatics online. Xiang Zhan, Santhosh Girirajan, Ni Zhao, Michael C. Wu, Debashis Ghosh |
Bioinform. | 1 |
| 2015 | Kernel approaches for differential expression analysis of mass spectrometry-based metabolomics dataabstractBACKGROUND: Data generated from metabolomics experiments are different from other types of "-omics" data. For example, a common phenomenon in mass spectrometry (MS)-based metabolomics data is that the data matrix frequently contains missing values, which complicates some quantitative analyses. One way to tackle this problem is to treat them as absent. Hence there are two types of information that are available in metabolomics data: presence/absence of a metabolite and a quantitative value of the abundance level of a metabolite if it is present. Combining these two layers of information poses challenges to the application of traditional statistical approaches in differential expression analysis. RESULTS: In this article, we propose a novel kernel-based score test for the metabolomics differential expression analysis. In order to simultaneously capture both the continuous pattern and discrete pattern in metabolomics data, two new kinds of kernels are designed. One is the distance-based kernel and the other is the stratified kernel. While we initially describe the procedures in the case of single-metabolite analysis, we extend the methods to handle metabolite sets as well. CONCLUSIONS: Evaluation based on both simulated data and real data from a liver cancer metabolomics study indicates that our kernel method has a better performance than some existing alternatives. An implementation of the proposed kernel method in the R statistical computing environment is available at http://works.bepress.com/debashis_ghosh/60/ . Xiang Zhan, Andrew D. Patterson, Debashis Ghosh |
BMC Bioinform. | 1 |