EDBT 2026 Demo / reviewers in the wild / expert
Qing Lu 0004
dblp:62/4337-4
· DBLP profile ↗
10ranked-venue papers
0as first author
8since 2021 · last 2024
0000-0002-7943-966XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 10 · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | AIGen: an artificial intelligence software for complex genetic data analysisabstractThe recent development of artificial intelligence (AI) technology, especially the advance of deep neural network (DNN) technology, has revolutionized many fields. While DNN plays a central role in modern AI technology, it has rarely been used in genetic data analysis due to analytical and computational challenges brought by high-dimensional genetic data and an increasing number of samples. To facilitate the use of AI in genetic data analysis, we developed a C++ package, AIGen, based on two newly developed neural networks (i.e. kernel neural networks and functional neural networks) that are capable of modeling complex genotype-phenotype relationships (e.g. interactions) while providing robust performance against high-dimensional genetic data. Moreover, computationally efficient algorithms (e.g. a minimum norm quadratic unbiased estimation approach and batch training) are implemented in the package to accelerate the computation, making them computationally efficient for analyzing large-scale datasets with thousands or even millions of samples. By applying AIGen to the UK Biobank dataset, we demonstrate that it can efficiently analyze large-scale genetic data, attain improved accuracy, and maintain robust performance. Availability: AIGen is developed in C++ and its source code, along with reference libraries, is publicly accessible on GitHub at https://github.com/TingtHou/AIGen. Tingting Hou, Xiaoxi Shen, Muxuan Liang, Li Chen 0029, Qing Lu 0004 |
Briefings Bioinform. | 6 |
| 2024 | Multimodal functional deep learning for multiomics dataabstractWith rapidly evolving high-throughput technologies and consistently decreasing costs, collecting multimodal omics data in large-scale studies has become feasible. Although studying multiomics provides a new comprehensive approach in understanding the complex biological mechanisms of human diseases, the high dimensionality of omics data and the complexity of the interactions among various omics levels in contributing to disease phenotypes present tremendous analytical challenges. There is a great need of novel analytical methods to address these challenges and to facilitate multiomics analyses. In this paper, we propose a multimodal functional deep learning (MFDL) method for the analysis of high-dimensional multiomics data. The MFDL method models the complex relationships between multiomics variants and disease phenotypes through the hierarchical structure of deep neural networks and handles high-dimensional omics data using the functional data analysis technique. Furthermore, MFDL leverages the structure of the multimodal model to capture interactions between different types of omics data. Through simulation studies and real-data applications, we demonstrate the advantages of MFDL in terms of prediction accuracy and its robustness to the high dimensionality and noise within the data. Pei Geng, Feifei Xiao, Guoshuai Cai, Li Chen 0029, Qing Lu 0004 |
Briefings Bioinform. | 7 |
| 2024 | MPRAVarDB: an online database and web server for exploring regulatory effects of genetic variantsabstractSUMMARY: Massively parallel reporter assay (MPRA) is an important technology for evaluating the impact of genetic variants on gene regulation. Here, we present MPRAVarDB, an online database and web server for exploring regulatory effects of genetic variants. MPRAVarDB harbors 18 MPRA experiments designed to assess the regulatory effects of genetic variants associated with GWAS loci, eQTLs, and genomic features, totaling 242 818 variants tested more than 30 cell lines and 30 human diseases or traits. MPRAVarDB enables users to query MPRA variants by genomic region, disease and cell line, or any combination of these parameters. Notably, MPRAVarDB offers a suite of pretrained machine-learning models tailored to the specific disease and cell line, facilitating the prediction of regulatory variants. The user-friendly interface allows users to receive query and prediction results with just a few clicks. AVAILABILITY AND IMPLEMENTATION: https://mpravardb.rc.ufl.edu. Weijia Jin, Javlon Nizomov, Qing Lu 0004, Li Chen 0029 |
Bioinform. | 6 |
| 2024 | scaDA: A novel statistical method for differential analysis of single-cell chromatin accessibility sequencing dataabstractSingle-cell ATAC-seq sequencing data (scATAC-seq) has been widely used to investigate chromatin accessibility on the single-cell level. One important application of scATAC-seq data analysis is differential chromatin accessibility (DA) analysis. However, the data characteristics of scATAC-seq such as excessive zeros and large variability of chromatin accessibility across cells impose a unique challenge for DA analysis. Existing statistical methods focus on detecting the mean difference of the chromatin accessible regions while overlooking the distribution difference. Motivated by real data exploration that distribution difference exists among cell types, we introduce a novel composite statistical test named "scaDA", which is based on zero-inflated negative binomial model (ZINB), for performing differential distribution analysis of chromatin accessibility by jointly testing the abundance, prevalence and dispersion simultaneously. Benefiting from both dispersion shrinkage and iterative refinement of mean and prevalence parameter estimates, scaDA demonstrates its superiority to both ZINB-based likelihood ratio tests and published methods by achieving the highest power and best FDR control in a comprehensive simulation study. In addition to demonstrating the highest power in three real sc-multiome data analyses, scaDA successfully identifies differentially accessible regions in microglia from sc-multiome data for an Alzheimer's disease (AD) study that are most enriched in GO terms related to neurogenesis and the clinical phenotype of AD, and AD-associated GWAS SNPs. Fengdi Zhao, Xin Ma 0028, Qing Lu 0004, Li Chen 0029 |
PLoS Comput. Biol. | 4 |
| 2024 | Functional Neural Networks for High-Dimensional Genetic Data AnalysisabstractArtificial intelligence (AI) is a thriving research field with many successful applications in areas such as computer vision and speech recognition. Machine learning methods, such as artificial neural networks (ANN), play a central role in modern AI technology. While ANN also holds great promise for human genetic research, the high-dimensional genetic data and complex genetic structure bring tremendous challenges. The vast majority of genetic variants on the genome have small or no effects on diseases, and fitting ANN on a large number of variants without considering the underlying genetic structure (e.g., linkage disequilibrium) could bring a serious overfitting issue. Furthermore, while a single disease phenotype is often studied in a classic genetic study, in emerging research fields (e.g., imaging genetics), researchers need to deal with different types of disease phenotypes. To address these challenges, we propose a functional neural networks (FNN) method. FNN uses a series of basis functions to model high-dimensional genetic data and a variety of phenotype data and further builds a multi-layer functional neural network to capture the complex relationships between genetic variants and disease phenotypes. Through simulations, we demonstrate the advantages of FNN for high-dimensional genetic data analysis in terms of robustness and accuracy. The real data applications also showed that FNN attained higher accuracy than the existing methods. Pei Geng, Qing Lu 0004 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2023 | Expectile Neural Networks for Genetic Data Analysis of Complex DiseasesabstractThe genetic etiologies of common diseases are highly complex and heterogeneous. Classic methods, such as linear regression, have successfully identified numerous variants associated with complex diseases. Nonetheless, for most diseases, the identified variants only account for a small proportion of heritability. Challenges remain to discover additional variants contributing to complex diseases. Expectile regression is a generalization of linear regression and provides complete information on the conditional distribution of a phenotype of interest. While expectile regression has many nice properties, it has rarely been used in genetic research. In this paper, we develop an expectile neural network (ENN) method for genetic data analyses of complex diseases. Similar to expectile regression, ENN provides a comprehensive view of relationships between genetic variants and disease phenotypes, which can be used to discover variants predisposing to sub-populations. We further integrate the idea of neural networks into ENN, making it capable of capturing non-linear and non-additive genetic effects (e.g., gene-gene interactions). Through simulations, we showed that the proposed method outperformed an existing expectile regression when there exist complex genotype-phenotype relationships. We also applied the proposed method to the data from the Study of Addiction: Genetics and Environment (SAGE), investigating the relationships of candidate genes with smoking quantity. Jinghang Lin, Xiaoran Tong, Qing Lu 0004 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2022 | Fast heritability estimation based on MINQUE and batch trainingabstractHeritability, the proportion of phenotypic variance explained by genome-wide single nucleotide polymorphisms (SNPs) in unrelated individuals, is an important measure of the genetic contribution to human diseases and plays a critical role in studying the genetic architecture of human diseases. Linear mixed model (LMM) has been widely used for SNP heritability estimation, where variance component parameters are commonly estimated by using a restricted maximum likelihood (REML) method. REML is an iterative optimization algorithm, which is computationally intensive when applied to large-scale datasets (e.g. UK Biobank). To facilitate the heritability analysis of large-scale genetic datasets, we develop a fast approach, minimum norm quadratic unbiased estimator (MINQUE) with batch training, to estimate variance components from LMM (LMM.MNQ.BCH). In LMM.MNQ.BCH, the parameters are estimated by MINQUE, which has a closed-form solution for fast computation and has no convergence issue. Batch training has also been adopted in LMM.MNQ.BCH to accelerate the computation for large-scale genetic datasets. Through simulations and real data analysis, we demonstrate that LMM.MNQ.BCH is much faster than two existing approaches, GCTA and BOLT-REML. Mingsheng Tang, Tingting Hou, Xiaoran Tong, Xiaoxi Shen, Xuefen Zhang, Tong Wang 0019, Qing Lu 0004 |
Briefings Bioinform. | 7 |
| 2022 | Explainable deep transfer learning model for disease risk prediction using high-dimensional genomic dataabstractBuilding an accurate disease risk prediction model is an essential step in the modern quest for precision medicine. While high-dimensional genomic data provides valuable data resources for the investigations of disease risk, their huge amount of noise and complex relationships between predictors and outcomes have brought tremendous analytical challenges. Deep learning model is the state-of-the-art methods for many prediction tasks, and it is a promising framework for the analysis of genomic data. However, deep learning models generally suffer from the curse of dimensionality and the lack of biological interpretability, both of which have greatly limited their applications. In this work, we have developed a deep neural network (DNN) based prediction modeling framework. We first proposed a group-wise feature importance score for feature selection, where genes harboring genetic variants with both linear and non-linear effects are efficiently detected. We then designed an explainable transfer-learning based DNN method, which can directly incorporate information from feature selection and accurately capture complex predictive effects. The proposed DNN-framework is biologically interpretable, as it is built based on the selected predictive genes. It is also computationally efficient and can be applied to genome-wide data. Through extensive simulations and real data analyses, we have demonstrated that our proposed method can not only efficiently detect predictive features, but also accurately predict disease risk, as compared to many existing methods. Cherry Weng, Qing Lu 0004, Tong Wang 0019, Yalu Wen |
PLoS Comput. Biol. | 4 |
| 2020 | Multi-kernel linear mixed model with adaptive lasso for prediction analysis on high-dimensional multi-omics dataabstractMOTIVATION: The use of human genome discoveries and other established factors to build an accurate risk prediction model is an essential step toward precision medicine. While multi-layer high-dimensional omics data provide unprecedented data resources for prediction studies, their corresponding analytical methods are much less developed. RESULTS: We present a multi-kernel penalized linear mixed model with adaptive lasso (MKpLMM), a predictive modeling framework that extends the standard linear mixed models widely used in genomic risk prediction, for multi-omics data analysis. MKpLMM can capture not only the predictive effects from each layer of omics data but also their interactions via using multiple kernel functions. It adopts a data-driven approach to select predictive regions as well as predictive layers of omics data, and achieves robust selection performance. Through extensive simulation studies, the analyses of PET-imaging outcomes from the Alzheimer's Disease Neuroimaging Initiative study, and the analyses of 64 drug responses, we demonstrate that MKpLMM consistently outperforms competing methods in phenotype prediction. AVAILABILITY AND IMPLEMENTATION: The R-package is available at https://github.com/YaluWen/OmicPred. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Qing Lu 0004, Yalu Wen |
Bioinform. | 2 |
| 2017 | A generalized association test based on U statisticsabstractMOTIVATION: Second generation sequencing technologies are being increasingly used for genetic association studies, where the main research interest is to identify sets of genetic variants that contribute to various phenotypes. The phenotype can be univariate disease status, multivariate responses and even high-dimensional outcomes. Considering the genotype and phenotype as two complex objects, this also poses a general statistical problem of testing association between complex objects. RESULTS: We here proposed a similarity-based test, generalized similarity U (GSU), that can test the association between complex objects. We first studied the theoretical properties of the test in a general setting and then focused on the application of the test to sequencing association studies. Based on theoretical analysis, we proposed to use Laplacian Kernel-based similarity for GSU to boost power and enhance robustness. Through simulation, we found that GSU did have advantages over existing methods in terms of power and robustness. We further performed a whole genome sequencing (WGS) scan for Alzherimer's disease neuroimaging initiative data, identifying three genes, APOE , APOC1 and TOMM40 , associated with imaging phenotype. AVAILABILITY AND IMPLEMENTATION: We developed a C ++ package for analysis of WGS data using GSU. The source codes can be downloaded at https://github.com/changshuaiwei/gsu . CONTACT: [email protected] ; [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Changshuai Wei, Qing Lu 0004 |
Bioinform. | 2 |