EDBT 2026 Demo / reviewers in the wild / expert
Lu Zhang 0061
dblp:82/10609-61
· DBLP profile ↗
20ranked-venue papers
3as first author
17since 2021 · last 2026
0000-0002-2794-7371ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 20 · 3 first-author · 17 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | stDyer-image improves clustering analysis of spatially resolved transcriptomics and proteomics with morphological imagesabstractMOTIVATION: Spatially resolved transcriptomics (SRT) and spatially resolved proteomics (SRP) data enable the study of gene expression and protein abundances within their precise spatial and cellular contexts in tissues. Certain SRT and SRP technologies also capture corresponding morphology images, adding another layer of valuable information. However, few existing methods developed for SRT data effectively leverage these supplementary images to enhance clustering performance. RESULTS: Here, we introduce stDyer-image, an end-to-end deep learning framework designed for clustering for SRT and SRP datasets with images. Unlike existing methods that utilize images to complement gene expression data, stDyer-image directly links image features to cluster labels. This approach draws inspiration from pathologists, who can visually identify specific cell types or tumor regions from morphological images without relying on gene expression or protein abundances. Benchmarks against state-of-the-art tools demonstrate that stDyer-image achieves superior performance in clustering. Moreover, it is capable of handling large-scale datasets across diverse technologies, making it a versatile and powerful tool for spatial omics analysis. AVAILABILITY AND IMPLEMENTATION: The source code of stDyer-image and detailed tutorials are available at https://github.com/ericcombiolab/stDyer-image. Xin Maizie Zhou, Lu Zhang 0061 |
Bioinform. | 3 |
| 2026 | Enhanced Disease Susceptible Variant Identification via Short Identity by Descent SegmentsabstractRare diseases affect millions of individuals worldwide, yet diagnostic yields for them still remain low. Among variant identification approaches, identity by descent (IBD) mapping is used to identify disease susceptible variants originating from a recent common ancestor among affected individuals, but existing IBD detection models struggle to identify these variants in short IBD segments. Here, we introduce SILO, a novel model to detect disease susceptible variants in both short and long IBD segments. SILO employs a two-stage procedure to detect IBD segments. In the first stage, SILO identifies long IBD segments based on common variants. In the second stage, SILO utilizes rare variants to detect short IBD segments using a seed-and-extend algorithm. We evaluated SILO in simulated data and real data from the 1000 Genomes Project. Our results demonstrate that SILO outperforms existing models in detecting disease susceptible variants within short IBD segments, and show comparable performance in detecting these variants within longer IBD segments. These findings highlight the potential of SILO to increase diagnostic yields for rare diseases by enhancing the identification of previously overlooked disease susceptible variants in short IBD segments. Nonetheless, we note that the detection of short IBD segments remains challenging due to limited precision, leaving room for future improvement. Chonghao Wang, Werner Pieter Veldsman, Yufen Huang, Xiaodong Fang, Lu Zhang 0061 |
IEEE Trans. Comput. Biol. Bioinform. | 5 |
| 2025 | IDCLP: A Deep Learning Framework for Predicting Chemical-Induced Gene Expression Profiles Through Multisource Data IntegrationabstractPhenotypic Drug Discovery enables the exploration and identification of new compounds with potential therapeutic value without the need to predefine the molecular targets of drug action or hypothesize their mechanisms in pathology. Chemicalinduced transcriptional profiles offer a comprehensive view of phenotypic responses to drugs and serve as a key tool in phenotype-based compound screening. Hence, it is necessary to develop an algorithm for predicting chemical-induced transcription profiles. Existing works tried to predict the transcription profile based on the structure of a compound, but the results are not yet satisfactory. Transcriptional profiles during the induction process are influenced not only by the structural features of compounds, cellular context, and dosages but also by the complex interactions between compounds and biological entities (e.g. target, diseases, and side effects) and the physicochemical properties of compounds. Here, we propose a deep learning model called IDCLP to predict gene expression profiles induced by de novo compounds. It utilizes a heterogeneous graph attention mechanism for extracting drug embedding and integrates the similarity network fusion algorithm to construct a drug similarity network. Moreover, IDCLP utilizes attention mechanisms to model the associations between drugs and cell lines. Experimental results show that IDCLP outperforms state-of-the-art methods, especially with unseen drugs that are dissimilar from the drugs in the supervised learning. Our implementation of IDCLP is available at https://github.com/sdesignates/IDCLP.git. Guo Mao, Hiu Fung Yip, Lu Zhang 0061 |
BIBM | 3 |
| 2025 | A multimodal framework for early detection and classification of social isolation and loneliness among Chinese older adultsabstractAbstract Background Chinese aging population is experiencing growing social isolation and loneliness (SI/L), causing substantial mental concerns. With tremendous efforts devoted to developing interventions, a lack of effective detection hinders the realization of health management for successful aging. Methods This study proposes a multimodal SI/L detection and classification framework that integrates multimodal (i.e., linguistic, acoustic, visual, and demographic) data collected from semi-structured interviews to predict the SI/L severity scores of older adults. Specifically, we construct a novel multimodal dataset tailored to the Chinese context. Then, fine-tuned Chinese large language models (i.e., DeepSeek-R1 and BERT-wwm), behavioral signal processing (i.e., OpenFace software), and prompt-based symptom extraction using GPT-5 are employed to generate the 15-dimensional features. These features are then input into a Random Forest regressor to predict SI/L severity scores. Findings Preliminary results suggest that multimodal models significantly outperform unimodal and dual-modal approaches in accuracy and robustness, with text-based features playing a dominant role and acoustic/visual cues contributing additional insights. By systematically evaluating single, dual, and multimodal settings, our work highlights the advantages of multimodal integration in improving detection precision. Contributions The study proposes a scalable, automated, and linguistically inclusive framework for early SI/L detection among Chinese older adults, addressing current limitations in data quality, population bias, and model interpretability. Our work sheds light on the research of elderly care and the practice in precision and early detection for Chinese older adults’ successful aging. Jueni Lyu, Lu Zhang 0061, Christy M. K. Cheung |
Briefings Bioinform. | 3 |
| 2025 | Mitigation of multi-scale biases in cell-type deconvolution for spatially resolved transcriptomics using HarmoDeconabstractMOTIVATION: The advent of spatially resolved transcriptomics (SRT) has revolutionized our understanding of tissue molecular microenvironments by enabling the study of gene expression in its spatial context. However, many SRT platforms lack single-cell resolution, necessitating cell-type deconvolution methods to estimate cell-type proportions in SRT spots. Despite advancements in existing tools, these methods have not addressed biases occurring at three scales: individual spots, entire tissue samples, and discrepancies between SRT and reference scRNA-seq datasets. These biases result in overbalanced cell-type proportions for each spot, mismatched cell-type fractions at the sample level, and data distribution shifts across platforms. RESULTS: To mitigate these biases, we introduce HarmoDecon, a novel semi-supervised deep learning model for spatial cell-type deconvolution. HarmoDecon leverages pseudo-spots derived from scRNA-seq data and uses Gaussian Mixture Graph Convolutional Networks to address the aforementioned issues. Through extensive simulations on multi-cell spots from STARmap and osmFISH, HarmoDecon outperformed 11 state-of-the-art methods. Additionally, when applied to legacy SRT platforms and 10x Visium datasets, HarmoDecon achieved the highest accuracy in spatial domain clustering and maintained strong correlations between cancer marker genes and cancer cells in human breast cancer samples. These results highlight the utility of HarmoDecon in advancing spatial transcriptomics analysis. AVAILABILITY AND IMPLEMENTATION: The HarmoDecon scripts, with the detailed tutorials, are available at https://github.com/ericcombiolab/HarmoDecon/tree/main. Lu Zhang 0061 |
Bioinform. | 5 |
| 2025 | TRAFICA: an open chromatin language model to improve transcription factor binding affinity predictionabstractMOTIVATION: In silico transcription factor and DNA (TF-DNA) binding affinity prediction plays a vital role in examining TF binding preferences and understanding gene regulation. The existing tools employ TF-DNA binding profiles from in vitro high-throughput technologies to predict TF-DNA binding affinity. However, TFs tend to bind to sequences in open chromatin regions in vivo, such TF binding preference is seldomly considered by these existing tools. RESULTS: In this study, we developed TRAFICA, an open chromatin language model to predict TF-DNA binding affinity by integrating sequence characteristics of open chromatin regions from ATAC-seq experiments and in vitro TF-DNA binding profiles from high-throughput technologies. We pretrained TRAFICA on over 2.8 million nucleotide sequences in open chromatin regions derived from 197 ATAC-seq experiments (115 cell lines) to learn in vivo TF binding preferences. We further fine-tuned TRAFICA using low-rank adaptation (LoRA) on PBM and HT-SELEX TF-DNA binding profiles to learn intrinsic binding preferences for specific TFs. We systematically evaluated TRAFICA and compared its predictive performance with existing prediction tools and advanced DNA language models. The experimental results demonstrated that TRAFICA significantly outperformed the others in predicting in vitro and in vivo TF-DNA binding affinity, achieving state-of-the-art performance. These findings indicate that considering the sequence characteristics from open chromatin regions could significantly improve TF-DNA binding affinity prediction. AVAILABILITY AND IMPLEMENTATION: The source code of TRAFICA and detailed tutorials are available at https://github.com/ericcombiolab/TRAFICA. Chonghao Wang, Aiping Lyu, Lu Zhang 0061 |
Bioinform. | 6 |
| 2025 | RAPID: Reliable and efficient Automatic generation of submission rePortIng checklists with large language moDelsabstractOBJECTIVE: To evaluate an automated reporting checklist generation tool using large language models and retrieval augmentation generation technology, called RAPID. MATERIALS AND METHODS: This study utilized large language models to develop a retrieval augmentation generation architecture. To assess its performance, a total of 91 published journal articles were collected and manually annotated in accordance with the CONSORT and CONSORT-AI medical reporting guidelines. These articles comprised 50 randomized controlled trials conducted without AI intervention and 41 randomized controlled trials that incorporated AI tools. RESULTS: Fifty RCT articles without the intervention of AI tools and 41 RCT articles with the intervention of AI tools were collected as CONSORT and CONSORT-AI datasets. All of the CONSORT reporting items (37) were included in the tool. RAPID achieved a high average accuracy rate of 92.11% and a content consistency score of 81.14% on the CONSORT dataset. Of the CONSORT-AI reporting items, 11 items related to the intervention of AI tools were included in the tool. RAPID achieved an average accuracy of 83.81% with a content consistency score of 72.51% on the CONSORT-AI dataset. DISCUSSION: RAPID may effectively save time and improve working efficiency for different user groups such as medical authors, researchers, editors, and reviewers. CONCLUSION: RAPID has strong scalability, which can be easily adapted to different medical reporting guidelines without transfer learning on a large dataset. RAPID got state-of-the-art performance on 2 datasets for 2 different checklists compared to other methods. Xufei Luo, Zhenhua Yang, Bingyi Wang, Long Ge, Zhaoxiang Bian, Yaolong Chen, Lu Zhang 0061, Dongrui Peng, Honghao Lai, Minjie Duan, Shilin Tang |
J. Am. Medical Informatics Assoc. | 10 |
| 2025 | DBNX: A Machine Learning Method for Ensembling Polygenic Risk Scores and Non-Genetic FactorsabstractPolygenic risk scoring (PRS) holds promise for improving disease prediction and medical treatments by evaluating an individual's genetic susceptibility through multiple genetic variants. However, current PRS calculation methods often excel only in specific diseases and populations, with no single approach consistently outperforming others across all contexts. Furthermore, these methods frequently overlook non-genetic factors, such as lifestyle, that also impact disease risk.We introduce an unsupervised Deep Belief Network (DBN) to aggregate PRS generated by various methods, achieving performance comparable to the Super Learner method-a supervised ensemble approach that combines predictions from multiple methods to improve outcomes. Unlike supervised methods, the DBN does not require training data and can directly ensemble the available PRS. Remarkably, on small-scale datasets, the DBN outperforms the Super Learner. Additionally, we present the DBNX model, which integrates PRS with non-genetic factors using a combination of DBN and XGBoost. DBNX produces a Composite Risk Score (CRS) that incorporates information from both PRS and non-genetic factors. In our experiments using the U.K. Biobank (UKBB) dataset across four diseases, DBNX demonstrated superior performance compared to other commonly used ensemble methods. Xiangzhe Yuan, Chonghao Wang, Shuqin Zhu, Lu Zhang 0061 |
IEEE Trans. Comput. Biol. Bioinform. | 4 |
| 2023 | Benchmarking multi-platform sequencing technologies for human genome assemblyabstractGenome assembly is a computational technique that involves piecing together deoxyribonucleic acid (DNA) fragments generated by sequencing technologies to create a comprehensive and precise representation of the entire genome. Generating a high-quality human reference genome is a crucial prerequisite for comprehending human biology, and it is also vital for downstream genomic variation analysis. Many efforts have been made over the past few decades to create a complete and gapless reference genome for humans by using a diverse range of advanced sequencing technologies. Several available tools are aimed at enhancing the quality of haploid and diploid human genome assemblies, which include contig assembly, polishing of contig errors, scaffolding and variant phasing. Selecting the appropriate tools and technologies remains a daunting task despite several studies have investigated the pros and cons of different assembly strategies. The goal of this paper was to benchmark various strategies for human genome assembly by combining sequencing technologies and tools on two publicly available samples (NA12878 and NA24385) from Genome in a Bottle. We then compared their performances in terms of continuity, accuracy, completeness, variant calling and phasing. We observed that PacBio HiFi long-reads are the optimal choice for generating an assembly with low base errors. On the other hand, we were able to produce the most continuous contigs with Oxford Nanopore long-reads, but they may require further polishing to improve on quality. We recommend using short-reads rather than long-reads themselves to improve the base accuracy of contigs from Oxford Nanopore long-reads. Hi-C is the best choice for chromosome-level scaffolding because it can capture the longest-range DNA connectedness compared to 10× linked-reads and Bionano optical maps. However, a combination of multiple technologies can be used to further improve the quality and completeness of genome assembly. For diploid assembly, hifiasm is the best tool for human diploid genome assembly using PacBio HiFi and Hi-C data. Looking to the future, we expect that further advancements in human diploid assemblers will leverage the power of PacBio HiFi reads and other technologies with long-range DNA connectedness to enable the generation of high-quality, chromosome-level and haplotype-resolved human genome assemblies. Werner Pieter Veldsman, Xiaodong Fang, Yufen Huang, Xuefeng Xie, Aiping Lyu, Lu Zhang 0061 |
Briefings Bioinform. | 7 |
| 2023 | A comprehensive investigation of statistical and machine learning approaches for predicting complex human diseases on genomic variantsabstractQuantifying an individual's risk for common diseases is an important goal of precision health. The polygenic risk score (PRS), which aggregates multiple risk alleles of candidate diseases, has emerged as a standard approach for identifying high-risk individuals. Although several studies have been performed to benchmark the PRS calculation tools and assess their potential to guide future clinical applications, some issues remain to be further investigated, such as lacking (i) various simulated data with different genetic effects; (ii) evaluation of machine learning models and (iii) evaluation on multiple ancestries studies. In this study, we systematically validated and compared 13 statistical methods, 5 machine learning models and 2 ensemble models using simulated data with additive and genetic interaction models, 22 common diseases with internal training sets, 4 common diseases with external summary statistics and 3 common diseases for trans-ancestry studies in UK Biobank. The statistical methods were better in simulated data from additive models and machine learning models have edges for data that include genetic interactions. Ensemble models are generally the best choice by integrating various statistical methods. LDpred2 outperformed the other standalone tools, whereas PRS-CS, lassosum and DBSLMM showed comparable performance. We also identified that disease heritability strongly affected the predictive performance of all methods. Both the number and effect sizes of risk SNPs are important; and sample size strongly influences the performance of all methods. For the trans-ancestry studies, we found that the performance of most methods became worse when training and testing sets were from different populations. Chonghao Wang, Werner Pieter Veldsman, Lu Zhang 0061 |
Briefings Bioinform. | 5 |
| 2023 | Accurate and interpretable gene expression imputation on scRNA-seq data using IGSimputeabstractSingle-cell ribonucleic acid sequencing (scRNA-seq) enables the quantification of gene expression at the transcriptomic level with single-cell resolution, enhancing our understanding of cellular heterogeneity. However, the excessive missing values present in scRNA-seq data hinder downstream analysis. While numerous imputation methods have been proposed to recover scRNA-seq data, high imputation performance often comes with low or no interpretability. Here, we present IGSimpute, an accurate and interpretable imputation method for recovering missing values in scRNA-seq data with an interpretable instance-wise gene selection layer (GSL). IGSimpute outperforms 12 other state-of-the-art imputation methods on 13 out of 17 datasets from different scRNA-seq technologies with the lowest mean squared error as the chosen benchmark metric. We demonstrate that IGSimpute can give unbiased estimates of the missing values compared to other methods, regardless of whether the average gene expression values are small or large. Clustering results of imputed profiles show that IGSimpute offers statistically significant improvement over other imputation methods. By taking the heart-and-aorta and the limb muscle tissues as examples, we show that IGSimpute can also denoise gene expression profiles by removing outlier entries with unexpectedly high expression values via the instance-wise GSL. We also show that genes selected by the instance-wise GSL could indicate the age of B cells from bladder fat tissue of the Tabula Muris Senis atlas. IGSimpute can impute one million cells using 64 min, and thus applicable to large datasets. Chinwang Cheong, Werner Pieter Veldsman, Aiping Lyu, William Kwok-Wai Cheung, Lu Zhang 0061 |
Briefings Bioinform. | 6 |
| 2023 | Benchmarking genome assembly methods on metagenomic sequencing dataabstractMetagenome assembly is an efficient approach to reconstruct microbial genomes from metagenomic sequencing data. Although short-read sequencing has been widely used for metagenome assembly, linked- and long-read sequencing have shown their advancements in assembly by providing long-range DNA connectedness. Many metagenome assembly tools were developed to simplify the assembly graphs and resolve the repeats in microbial genomes. However, there remains no comprehensive evaluation of metagenomic sequencing technologies, and there is a lack of practical guidance on selecting the appropriate metagenome assembly tools. This paper presents a comprehensive benchmark of 19 commonly used assembly tools applied to metagenomic sequencing datasets obtained from simulation, mock communities or human gut microbiomes. These datasets were generated using mainstream sequencing platforms, such as Illumina and BGISEQ short-read sequencing, 10x Genomics linked-read sequencing, and PacBio and Oxford Nanopore long-read sequencing. The assembly tools were extensively evaluated against many criteria, which revealed that long-read assemblers generated high contig contiguity but failed to reveal some medium- and high-quality metagenome-assembled genomes (MAGs). Linked-read assemblers obtained the highest number of overall near-complete MAGs from the human gut microbiomes. Hybrid assemblers using both short- and long-read sequencing were promising methods to improve both total assembly length and the number of near-complete MAGs. This paper also discussed the running time and peak memory consumption of these assembly tools and provided practical guidance on selecting them. Zhenmiao Zhang, Werner Pieter Veldsman, Xiaodong Fang, Lu Zhang 0061 |
Briefings Bioinform. | 5 |
| 2022 | A machine learning model for disease risk prediction by integrating genetic and non-genetic factorsabstractPolygenic risk score (PRS) has been widely used to identify the high-risk individuals from the general population, which would be helpful for disease prevention and early treatment. Many methods have been developed to calculate PRS by weighting and aggregating the phenotype-associated risk alleles from genome-wide association studies. However, only considering genetic effects may not be sufficient for risk prediction because the disease risk is not only related to genetic factors but also non-genetic factors, e.g., diet, physical exercise et al. But it is still a challenge to integrate these genetic and non-genetic factors into a unified machine learning framework for disease risk prediction. In this paper, we proposed PRSIMD (PRS Integrating Multi-source Data), a machine learning model that applies posterior regularization to integrate genetic and non-genetic factors to improve disease risk prediction. Also, we applied Mendelian Randomization analysis to identify the causal non-genetic risk factors for the selected diseases. We applied PRSIMD to predict type 2 diabetes and coronary artery disease from UK Biobank and observed that PRSIMD was significantly better than the existing methods to calculate PRS. In addition, we observed that PRSIMD achieved the better predictive power than the composite risk score. The codes of PRSIMD are available at: https://github.con ericcombiolab/PRSIMD Chonghao Wang, Yunpeng Cai, Ouzhou Young, Aiping Lyu, Lu Zhang 0061 |
BIBM | 7 |
| 2022 | dynDeepDRIM: a dynamic deep learning model to infer direct regulatory interactions using time-course single-cell gene expression dataabstractTime-course single-cell RNA sequencing (scRNA-seq) data have been widely used to explore dynamic changes in gene expression of transcription factors (TFs) and their target genes. This information is useful to reconstruct cell-type-specific gene regulatory networks (GRNs). However, the existing tools are commonly designed to analyze either time-course bulk gene expression data or static scRNA-seq data via pseudo-time cell ordering. A few methods successfully utilize the information from multiple time points while also considering the characteristics of scRNA-seq data. We proposed dynDeepDRIM, a novel deep learning model to reconstruct GRNs using time-course scRNA-seq data. It represents the joint expression of a gene pair as an image and utilizes the image of the target TF-gene pair and the ones of the potential neighbors to reconstruct GRNs from time-course scRNA-seq data. dynDeepDRIM can effectively remove the transitive TF-gene interactions by considering neighborhood context and model the gene expression dynamics using high-dimensional tensors. We compared dynDeepDRIM with six GRN reconstruction methods on both simulation and four real time-course scRNA-seq data. dynDeepDRIM achieved substantially better performance than the other methods in inferring TF-gene interactions and eliminated the false positives effectively. We also applied dynDeepDRIM to annotate gene functions and found it achieved evidently better performance than the other tools due to considering the neighbor genes. Aiping Lyu, William Kwok-Wai Cheung, Lu Zhang 0061 |
Briefings Bioinform. | 5 |
| 2021 | An ensemble deep learning framework to refine large deletions in linked-readsabstractThe detection of structural variants (SVs) remains challenging due to inconsistencies in detected breakpoints and biological complexity of some rearrangements. Linked-reads have demonstrated their superiority in diploid genome assembly and SV detection. Recently developed tools Aquila and Aquila_stLFR use a reference sequence and linked-reads to generate a high quality diploid genome assembly, using which they then detect and phase personal genetic variations. However, they both produce a substantial proportion of false positive deletion SV calls. To take full advantage of linked-reads, an effective downstream filtering and refinement framework is needed pressingly. In this work, we propose AquilaDeepFilter to filter large deletion SVs from Aquila and Aquila_stLFR. AquilaDeepFilter relies on a deep learning ensemble approach by integrating several state-of-the-art CNN backbones. The filtering of deletion SVs is formulated as a binary classification task on image data that are generated through the extraction of multiple alignment signals, including read depth, split reads and discordant read pairs. Three linked-reads libraries sequenced from the well-studied sample NA24385 and the gold standard of GiaB benchmark were used to perform thorough experiments on our proposed method. The results demonstrated that AquilaDeepFilter could increase the precision rate of Aquila while the recall rate of Aquila decreased only slightly, and the overall F1 improved by 20%. Furthermore, AquilaDeepFilter outperformed another deep learning based method for SV filtering, DeepSVFilter. Even though we designed AquilaDeepFilter for linked-reads, the framework could also be used to improve SV detection on short reads. Yunfei Hu, V. Mangal Sanidhya, Lu Zhang 0061, Zhou Xin |
BIBM | 3 |
| 2021 | DeepDRIM: a deep neural network to reconstruct cell-type-specific gene regulatory network using single-cell RNA-seq dataabstractSingle-cell RNA sequencing has enabled to capture the gene activities at single-cell resolution, thus allowing reconstruction of cell-type-specific gene regulatory networks (GRNs). The available algorithms for reconstructing GRNs are commonly designed for bulk RNA-seq data, and few of them are applicable to analyze scRNA-seq data by dealing with the dropout events and cellular heterogeneity. In this paper, we represent the joint gene expression distribution of a gene pair as an image and propose a novel supervised deep neural network called DeepDRIM which utilizes the image of the target TF-gene pair and the ones of the potential neighbors to reconstruct GRN from scRNA-seq data. Due to the consideration of TF-gene pair's neighborhood context, DeepDRIM can effectively eliminate the false positives caused by transitive gene-gene interactions. We compared DeepDRIM with nine GRN reconstruction algorithms designed for either bulk or single-cell RNA-seq data. It achieves evidently better performance for the scRNA-seq data collected from eight cell lines. The simulated data show that DeepDRIM is robust to the dropout rate, the cell number and the size of the training data. We further applied DeepDRIM to the scRNA-seq gene expression of B cells from the bronchoalveolar lavage fluid of the patients with mild and severe coronavirus disease 2019. We focused on the cell-type-specific GRN alteration and observed targets of TFs that were differentially expressed between the two statuses to be enriched in lysosome, apoptosis, response to decreased oxygen level and microtubule, which had been proved to be associated with coronavirus infection. Chinwang Cheong, Liang Lan, Jiming Liu 0001, Aiping Lyu, William Kwok-Wai Cheung, Lu Zhang 0061 |
Briefings Bioinform. | 8 |
| 2021 | METAMVGL: a multi-view graph-based metagenomic contig binning algorithm by integrating assembly and paired-end graphsabstractBACKGROUND: Due to the complexity of microbial communities, de novo assembly on next generation sequencing data is commonly unable to produce complete microbial genomes. Metagenome assembly binning becomes an essential step that could group the fragmented contigs into clusters to represent microbial genomes based on contigs' nucleotide compositions and read depths. These features work well on the long contigs, but are not stable for the short ones. Contigs can be linked by sequence overlap (assembly graph) or by the paired-end reads aligned to them (PE graph), where the linked contigs have high chance to be derived from the same clusters. RESULTS: We developed METAMVGL, a multi-view graph-based metagenomic contig binning algorithm by integrating both assembly and PE graphs. It could strikingly rescue the short contigs and correct the binning errors from dead ends. METAMVGL learns the two graphs' weights automatically and predicts the contig labels in a uniform multi-view label propagation framework. In experiments, we observed METAMVGL made use of significantly more high-confidence edges from the combined graph and linked dead ends to the main graph. It also outperformed many state-of-the-art contig binning algorithms, including MaxBin2, MetaBAT2, MyCC, CONCOCT, SolidBin and GraphBin on the metagenomic sequencing data from simulation, two mock communities and Sharon infant fecal samples. CONCLUSIONS: Our findings demonstrate METAMVGL outstandingly improves the short contig binning and outperforms the other existing contig binning tools on the metagenomic sequencing data from simulation, mock communities and infant fecal samples. Zhenmiao Zhang, Lu Zhang 0061 |
BMC Bioinform. | 2 |
| 2016 | More accurate models for detecting gene-gene interactions from public expression compendiaabstractThe fast accumulation of public gene expression data-made possible by high-throughput technology-provides us an unprecedented opportunity to identify functionally related genes by analyzing their co-expression patterns. However, these data are typically noisy and highly heterogeneous, complicating their use in constructing co-expression network in large expression compendia. Previous studies suggested that the collective gene expression pattern can be better modeled by Gaussian mixtures. This motivates our present work, which proposes Multimodal Mutual Information (MMI) to reconstruct gene co-expression network from public gene expression data. MMI assumes gene pair following bivariate Gaussian mixture models and categories the samples into unique bins with respect to their expression magnitude. Two kinds of correlations in MMI are computed and aggregated to capture both discretized dependency and the expression correlation for each bin. Through extensive simulations, MMI outperforms other approaches with respect to calculating gene-gene interactions, regardless of the level of noise or strength of interactions. The advance of MMI is further validated by three real problems: 1. Infer novel gene functions by their connections with the well-documented genes. We apply principle component analysis to the correlated matrix generated by MMI and Pearson correlation and construct transcriptional components to evaluate gene-gene interactions. MMI enables 1.7 times more eligible transcriptional components than Pearson correlation, which can be used to predict gene functions. 2. Prioritize candidate genes for an affected pedigree. MMI calculates the interactions between candidate and disease established genes and explores KIF1A as the new causal gene of pure hereditary spastic paraparesis. 3. Detect disease “hot genes”. MMI identifies ANK2 as the “hot gene” for autism spectrum disorders by evaluating its co-expression with other disease susceptible genes derived from trio-based exome sequencing data. Lu Zhang 0061, Shuaicheng Li 0001 |
BIBM | 1 |
| 2015 | Reconstructing directed gene regulatory network by only gene expression dataabstractAccurately identifying gene regulatory network serves an important task in understanding in vivo biological activities. The inference of such network is often accomplished through the use of gene expression data. Some methods further predict the regulatory directions in the network by using the location of eQTL single nucleotide polymorphisms, or through gene knock out/down experiments; regrettably, these additional data are not always available, especially for the samples deriving from human tissues. In this paper, we propose Context Based Dependency Network (CBDN), a method that is able to infer gene regulatory networks, complete with the regulatory directions, from only gene expression data. CBDN applies directed data processing inequality (DDPI) to distinguish between direct and transitive relationship between genes. In our experiments with simulated and real data, CBDN outperforms the current state-of-the-art approaches. When used to identify important regulators in a network, CBDN 1. correctly identified TYROBP in the network related to Alzheimer's disease; 2. predicted potential important regulators ZNF329 and RB1 for human brain tumors. Lu Zhang 0061, Yen Kaow Ng, Shuaicheng Li 0001 |
BIBM | 1 |
| 2013 | PriVar: a toolkit for prioritizing SNVs and indels from next-generation sequencing dataabstractUNLABELLED: Next-generation sequencing has become a valuable tool for detecting mutations involved in Mendelian diseases. However, it is a challenge to identify the small subset of functionally important mutations from tens of thousands of rare variants in a whole exome/genome. Therefore, we developed a toolkit called PriVar, a systematic prioritization pipeline that takes into consideration calling quality of the variants, their predicted functional impact, known connection of the gene to the disease and the number of mutations in a gene, and inference from linkage analysis. AVAILABILITY: Executable jar package is available at http://paed.hku.hk/uploadarea/yangwl/html/software.html. Lu Zhang 0061, Dingge Ying, Yu-Lung Lau, Wanling Yang |
Bioinform. | 1 |