Lin Hou 0003

dblp:87/8247-3 · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
6since 2021 · last 2024
0000-0002-4283-8501ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 6 since 2021
YearPublicationVenuePosition
2024 Cofea: correlation-based feature selection for single-cell chromatin accessibility data
abstract
Single-cell chromatin accessibility sequencing (scCAS) technologies have enabled characterizing the epigenomic heterogeneity of individual cells. However, the identification of features of scCAS data that are relevant to underlying biological processes remains a significant gap. Here, we introduce a novel method Cofea, to fill this gap. Through comprehensive experiments on 5 simulated and 54 real datasets, Cofea demonstrates its superiority in capturing cellular heterogeneity and facilitating downstream analysis. Applying this method to identification of cell type-specific peaks and candidate enhancers, as well as pathway enrichment analysis and partitioned heritability analysis, we illustrate the potential of Cofea to uncover functional biological process.
Xiaoyang Chen 0007, Shuang Song 0006, Lin Hou 0003, Shengquan Chen, Rui Jiang 0001
Briefings Bioinform.4
2024 Coupled Epidemic-Information Propagation With Stranding Mechanism on Multiplex Metapopulation Networks
abstract
Acknowledging the significance of information propagation and individual adaptive behavior has been regarded as an indispensable prerequisite for a complete understanding of epidemic spreading. Recent studies have widely considered the metapopulation model, where epidemics spread over a single layer of physical networks via individual mobility. However, these advances neglected the interventions of accompanied information and individual behavior response related to epidemics. In this article, we develop a coupled epidemic-information propagation model on multiplex metapopulation networks leveraging the microscopic Markov chain (MMC) approach, aiming to explore the spatiotemporal characteristics of epidemic spreading process. Taking the individual adaptive behavior into account, the stranding mechanism based on infection level and medical resources is introduced to capture the population size dynamics during individual mobility among different patches. Theoretical epidemic threshold is analytically derived under the improved framework. Extensive numerical simulations are performed to validate our theoretical analysis and further examine the impacts of information propagation and spreading parameters on epidemic threshold and steady-state prevalence. Our results indicate that both the scale of information diffusion and the specific configuration of spreading parameters can significantly suppress the epidemic prevalence. These findings shed a novel light on theoretical research and decision-making of coupled epidemic-information process in the spatiotemporal perspective.
Xuming An 0002, Chen Zhang 0007, Lin Hou 0003, Kaibo Wang
IEEE Trans. Comput. Soc. Syst.3
2022 CITEdb: a manually curated database of cell-cell interactions in human
abstract
MOTIVATION: The interactions among various types of cells play critical roles in cell functions and the maintenance of the entire organism. While cell-cell interactions are traditionally revealed from experimental studies, recent developments in single-cell technologies combined with data mining methods have enabled computational prediction of cell-cell interactions, which have broadened our understanding of how cells work together, and have important implications in therapeutic interventions targeting cell-cell interactions for cancers and other diseases. Despite the importance, to our knowledge, there is no database for systematic documentation of high-quality cell-cell interactions at the cell type level, which hinders the development of computational approaches to identify cell-cell interactions. RESULTS: We develop a publicly accessible database, CITEdb (Cell-cell InTEraction database, https://citedb.cn/), which not only facilitates interactive exploration of cell-cell interactions in specific physiological contexts (e.g. a disease or an organ) but also provides a benchmark dataset to interpret and evaluate computationally derived cell-cell interactions from different tools. CITEdb contains 728 pairs of cell-cell interactions in human that are manually curated. Each interaction is equipped with structured annotations including the physiological context, the ligand-receptor pairs that mediate the interaction, etc. Our database provides a web interface to search, visualize and download cell-cell interactions. Users can search for cell-cell interactions by selecting the physiological context of interest or specific cell types involved. CITEdb is the first attempt to catalogue cell-cell interactions at the cell type level, which is beneficial to both experimental, computational and clinical studies of cell-cell interactions. AVAILABILITY AND IMPLEMENTATION: CITEdb is freely available at https://citedb.cn/ and the R package implementing benchmark is available at https://github.com/shanny01/benchmark. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Nayang Shan, Dongyu Li, Jitong Jiang, Linlin Yan, Jiudong Gao, Xing-Ming Zhao, Lin Hou 0003
Bioinform.10
2022 A data-adaptive Bayesian regression approach for polygenic risk prediction
abstract
MOTIVATION: Polygenic risk score (PRS) has been widely exploited for genetic risk prediction due to its accuracy and conceptual simplicity. We introduce a unified Bayesian regression framework, NeuPred, for PRS construction, which accommodates varying genetic architectures and improves overall prediction accuracy for complex diseases by allowing for a wide class of prior choices. To take full advantage of the framework, we propose a summary-statistics-based cross-validation strategy to automatically select suitable chromosome-level priors, which demonstrates a striking variability of the prior preference of each chromosome, for the same complex disease, and further significantly improves the prediction accuracy. RESULTS: Simulation studies and real data applications with seven disease datasets from the Wellcome Trust Case Control Consortium cohort and eight groups of large-scale genome-wide association studies demonstrate that NeuPred achieves substantial and consistent improvements in terms of predictive r2 over existing methods. In addition, NeuPred has similar or advantageous computational efficiency compared with the state-of-the-art Bayesian methods. AVAILABILITY AND IMPLEMENTATION: The R package implementing NeuPred is available at https://github.com/shuangsong0110/NeuPred. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Shuang Song 0006, Lin Hou 0003, Jun S. Liu
Bioinform.2
2022 Erratum to: A data-adaptive Bayesian regression approach for polygenic risk prediction
abstract
Bioinformatics (2022), https://doi.org/10.1093/bioinformatics/btac024 In the originally published version of this manuscript, there was an address error in affiliations 1, 2, and 3. This error has been corrected.
Shuang Song 0006, Lin Hou 0003, Jun S. Liu
Bioinform.2
2021 Openness weighted association studies: leveraging personal genome information to prioritize non-coding variants
abstract
MOTIVATION: Identification and interpretation of non-coding variations that affect disease risk remain a paramount challenge in genome-wide association studies (GWAS) of complex diseases. Experimental efforts have provided comprehensive annotations of functional elements in the human genome. On the other hand, advances in computational biology, especially machine learning approaches, have facilitated accurate predictions of cell-type-specific functional annotations. Integrating functional annotations with GWAS signals has advanced the understanding of disease mechanisms. In previous studies, functional annotations were treated as static of a genomic region, ignoring potential functional differences imposed by different genotypes across individuals. RESULTS: We develop a computational approach, Openness Weighted Association Studies (OWAS), to leverage and aggregate predictions of chromosome accessibility in personal genomes for prioritizing GWAS signals. The approach relies on an analytical expression we derived for identifying disease associated genomic segments whose effects in the etiology of complex diseases are evaluated. In extensive simulations and real data analysis, OWAS identifies genes/segments that explain more heritability than existing methods, and has a better replication rate in independent cohorts than GWAS. Moreover, the identified genes/segments show tissue-specific patterns and are enriched in disease relevant pathways. We use rheumatic arthritis and asthma as examples to demonstrate how OWAS can be exploited to provide novel insights on complex diseases. AVAILABILITY AND IMPLEMENTATION: The R package OWAS that implements our method is available at https://github.com/shuangsong0110/OWAS. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Shuang Song 0006, Nayang Shan, Xiting Yan, Jun S. Liu, Lin Hou 0003
Bioinform.6
2020 Leveraging effect size distributions to improve polygenic risk scores derived from summary statistics of genome-wide association studies
abstract
Genetic risk prediction is an important problem in human genetics, and accurate prediction can facilitate disease prevention and treatment. Calculating polygenic risk score (PRS) has become widely used due to its simplicity and effectiveness, where only summary statistics from genome-wide association studies are needed in the standard method. Recently, several methods have been proposed to improve standard PRS by utilizing external information, such as linkage disequilibrium and functional annotations. In this paper, we introduce EB-PRS, a novel method that leverages information for effect sizes across all the markers to improve prediction accuracy. Compared to most existing genetic risk prediction methods, our method does not need to tune parameters nor external information. Real data applications on six diseases, including asthma, breast cancer, celiac disease, Crohn's disease, Parkinson's disease and type 2 diabetes show that EB-PRS achieved 307.1%, 42.8%, 25.5%, 3.1%, 74.3% and 49.6% relative improvements in terms of predictive r2 over standard PRS method with optimally tuned parameters. Besides, compared to LDpred that makes use of LD information, EB-PRS also achieved 37.9%, 33.6%, 8.6%, 36.2%, 40.6% and 10.8% relative improvements. We note that our method is not the first method leveraging effect size distributions. Here we first justify our method by presenting theoretical optimal property over existing methods in this class of methods, and substantiate our theoretical result with extensive simulation results. The R-package EBPRS that implements our method is available on CRAN.
Shuang Song 0006, Wei Jiang 0019, Lin Hou 0003, Hongyu Zhao 0003
PLoS Comput. Biol.3
2019 Identification of trans-eQTLs using mediation analysis with multiple mediators
abstract
BACKGROUND: Mapping expression quantitative trait loci (eQTLs) has provided insight into gene regulation. Compared to cis-eQTLs, the regulatory mechanisms of trans-eQTLs are less known. Previous studies suggest that trans-eQTLs may regulate expression of remote genes by altering the expression of nearby genes. Trans-association has been studied in the mediation analysis with a single mediator. However, prior applications with one mediator are prone to model misspecification due to correlations between genes. Motivated from the observation that trans-eQTLs are more likely to associate with more than one cis-gene than randomly selected SNPs in the GTEx dataset, we developed a computational method to identify trans-eQTLs that are mediated by multiple mediators. RESULTS: We proposed two hypothesis tests for testing the total mediation effect (TME) and the component-wise mediation effects (CME), respectively. We demonstrated in simulation studies that the type I error rates were controlled in both tests despite model misspecification. The TME test was more powerful than the CME test when the two mediation effects are in the same direction, while the CME test was more powerful than the TME test when the two mediation effects are in opposite direction. Multiple mediator analysis had increased power to detect mediated trans-eQTLs, especially in large samples. In the HapMap3 data, we identified 11 mediated trans-eQTLs that were not detected by the single mediator analysis in the combined samples of African populations. Moreover, the mediated trans-eQTLs in the HapMap3 samples are more likely to be trait-associated SNPs. In terms of computation, although there is no limit in the number of mediators in our model, analysis takes more time when adding additional mediators. In the analysis of the HapMap3 samples, we included at most 5 cis-gene mediators. Majority of the trios we considered have one or two mediators. CONCLUSIONS: Trans-eQTLs are more likely to associate with multiple cis-genes than randomly selected SNPs. Mediation analysis with multiple mediators improves power of identification of mediated trans-eQTLs, especially in large samples.
Nayang Shan, Zuoheng Wang, Lin Hou 0003
BMC Bioinform.3