EDBT 2026 Demo / reviewers in the wild / expert
Guohua Wang 0001
dblp:74/4514-1
· DBLP profile ↗
95ranked-venue papers
0as first author
85since 2021 · last 2026
0000-0001-7381-2374ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 89 · 79 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Systems, architecture and hardware · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Identification and characterization of lncRNA-stemness-immune regulatory patternsabstractLong noncoding RNAs (lncRNAs) play critical roles in regulating stemness signature genes (SSGs) and tumor immunity, thereby shaping the tumor microenvironment and antitumor immune responses. Increasing evidence suggests that cancer stem cell traits are closely associated with immune evasion and therapeutic resistance, underscoring the need to systematically characterize the pan-cancer interplay among SSGs, lncRNAs, and tumor immunity. Here, we developed an integrative analytical framework that combines network-based modeling with Bayesian network inference to identify core regulatory triplets (STEM-LncCRTs), each consisting of an lncRNA, an SSG, and an immune gene. We demonstrate that specific stemness-related lncRNAs can distinguish cancer subtypes, and that common stemness-related lncRNAs correlate significantly with immune cell infiltration. Notably, the ATAD5/PRR11-AS1/SKP2 triplet exhibits favorable prognostic potential across multiple cancers and consistently outperforms individual gene markers in predicting 1-, 3-, and 5-year overall survival. Furthermore, using four machine learning algorithms across three independent immunotherapy cohorts, we validate the predictive value of STEM-LncCRTs for immune checkpoint inhibitor response. Importantly, integrating STEM-LncCRTs with tumor mutation burden further improves predictive accuracy. Collectively, this study provides a systems-level view of stemness-related lncRNA regulation in tumor immunity and offers practical biomarkers for predicting immunotherapy efficacy. Zhipeng Qian, Chunlong Zhang, Guohua Wang 0001, Chunyu Wang 0002, Yang Li 0130 |
Briefings Bioinform. | 4 |
| 2026 | GFSeeker: a splicing-graph-based approach for accurate gene fusion detection from long-read RNA sequencing dataabstractGene fusions are critical oncogenic drivers and therapeutic targets in diverse cancers. Long-read ribonucleic acid sequencing (RNA-seq) offers an unprecedented opportunity to resolve the full-length structure of fusion isoforms, but its high intrinsic error rates pose significant challenges to the precise identification of true fusion events. Here, we developed GFSeeker, an innovative splicing-graph-based computational framework for accurate gene fusion detection from long-read RNA-seq. GFSeeker employs a unique pipeline based on a splicing graph reference and a dual re-alignment validation to effectively overcome data noise from high error rates. Benchmarking across simulated, non-tumor, and cancer cell line datasets demonstrated GFSeeker's state-of-the-art performance, achieving 6%-15% higher F1 score compared to existing methods. Notably, GFSeeker successfully identified the known fusion event, MATN2-POP1, in the MCF-7 cancer cell line, missed by other tools, highlighting its superior sensitivity in resolving complex fusion events. These results validate GFSeeker as a powerful and reliable tool for gene fusion discovery, heralding its significant potential to advance cancer research and precision diagnostics. Heng Hu, Runtian Gao, Guohua Wang 0001, Tao Jiang 0021 |
Briefings Bioinform. | 4 |
| 2026 | Image-guided spatial omics enhancement reveals hidden spatial microstructures
Gongning Luo, Qiaoming Liu, Suyu Dong, Guohua Wang 0001 |
Bioinform. | 5 |
| 2026 | GMHAN: a heterogeneous graph attention framework for prioritizing coding and non-coding driver genesabstractMOTIVATION: Cancer, a disease of high complexity. Identifying cancer driver genes is fundamental for elucidating oncogenesis and promoting precision medicine. Currently, most approaches mainly focus on homogeneous gene networks and single-omics data, thereby mainly identifying coding driver genes while ignoring non-coding driver genes. RESULT: Thus, we introduced GMHAN, a novel framework based on HAN. Firstly, we integrated the three types of omics data of genes and PPI network topology feature, together with the multi-dimensional features of miRNAs. Afterwards, we used heterogeneous graph attention networks to obtain deep feature embeddings of genes and miRNAs. Finally, the deep feature embeddings are input multilayer perceptron to obtain the probability that genes and miRNAs being cancer drivers. In a comparative evaluation against seven methods, GMHAN demonstrates better performance across both pan-cancer and cancer-specific datasets, achieving higher scores in AUC and AUPR. It has confirmed its effectiveness in identifying carcinogenic drivers. AVAILABILITY AND IMPLEMENTATION: The source code of GMHAN is available at: https://github.com/mping315/GMHAN and https://doi.org/10.5281/zenodo.20154736. Ping Meng, Guohua Wang 0001 |
Bioinform. | 3 |
| 2026 | spAttClu: a spatial domain clustering model leveraging spatially weighted graph attention and contrastive learningabstractMOTIVATION: The rapid growth of spatial transcriptomics data holds potential for deep understanding of spatial specificity and tissue heterogeneity. Recognizing spatial domains is a fundamental step for deciphering tissue functional architecture and dissecting tissue heterogeneity. However, existing models typically define adjacency relations using static weights, which cannot dynamically adjust neighbor importance based on expression context, thereby limiting the accuracy and robustness of spatial domain recognition. RESULTS: We propose spAttClu, a clustering model integrating spatially weighted graph attention with contrastive learning. It adaptively learns neighbor contributions in varying contexts through a distance-weighted graph attention mechanism and enhances embedding discriminability via multi-level contrastive learning. spAttClu demonstrates superior clustering performance on the DLPFC dataset. Moreover, it shows cross-platform adaptability and enables vertical/horizontal inte-gration of multiple tissue slices. Ruolan Zhang, Zhongqian Zhao, Shenghe Li, Yucai Jiang, Binyang Wei, Guohua Wang 0001 |
Bioinform. | 9 |
| 2026 | scTACL: a multitask topology-aware contrastive learning approach for single-cell transcriptomics analysisabstractMOTIVATION: The advent of single-cell RNA sequencing (scRNA-seq) technology has allowed researchers to measure gene expression profiles at the single-cell level, providing valuable insights into cellular heterogeneity. However, due to the limitations of current sequencing platforms, scRNA-seq data often contain significant noise, particularly severe dropout events, which pose major challenges for subsequent analyses. RESULTS: In this study, we developed a new method called topology-aware contrastive learning (scTACL). This approach uses contrastive learning between a cell similarity graph and a cell embedding similarity graph, employing a zero-inflated negative binomial (ZINB) distribution to model the reconstructed data. This alignment helps the processed data better reflect true biological signals. It delivers superior results in key tasks such as data imputation, clustering, batch effect correction, and cell-cell interaction. Additionally, scTACL successfully identified two distinct subtypes of epithelial cells in lung adenocarcinoma tissues, further demonstrating its effectiveness and usefulness in complex biological settings. Notably, without relying on spatial location information, scTACL still effectively distinguished the epithelial and mesenchymal regions in the spatial transcriptome data of liver cancer and identified the COLLAGEN signaling pathway, which plays a crucial role in the epithelial-mesenchymal transition process through intercellular communication analysis. Murong Zhou, Yingjian Liang, Alfred Wei Chieh Kow, Guohua Wang 0001, Qiaoming Liu |
Bioinform. | 5 |
| 2026 | PLATO: ProbabiListic hierArchical mulTi-head mOdel for plug-and-play ambiguous medical image segmentation
Xiangyu Li 0004, Fanding Li, Yongfeng Yuan, Suyu Dong, Kuanquan Wang, Yi Shen 0001, Guohua Wang 0001, Gongning Luo, Shuo Li 0001 |
Knowl. Based Syst. | 7 |
| 2026 | CA-CAE: A deep learning-based multi-omics model for pan-cancer subtype classification and prognosis predictionabstractIn cancer research, identifying cancer subtypes and evaluating prognosis are crucial for personalized diagnosis and treatment of cancer. With the advancement of high-throughput sequencing technologies, multi-omics data has become essential for cancer classification and prognostic analysis. By integrating deep learning techniques, it is possible to more accurately identify cancer subtypes, providing a robust basis for personalized treatment of cancer patients. In this study, we propose a convolutional autoencoder prognostic model incorporating a channel attention mechanism (CA-CAE). The model utilizes multi-omics data to predict survival-associated cancer subtypes and identify prognostic genes. We applied CA-CAE to multiple cancer types, successfully identifying subtypes in 15 distinct cancer types and revealing significant survival differences among these subtypes. Moreover, compared to traditional statistical methods and other deep learning approaches, CA-CAE demonstrated superior performance in predicting survival outcomes. Shumei Zhang, Yicheng Lu, Peixian Li, Junxuan Wu, Guohua Wang 0001 |
PLoS Comput. Biol. | 5 |
| 2026 | Compact Fuzzy-Rule Decision-Level Fusion for Ovarian Cancer Survival Prediction With Controlled Modality ExtensionabstractAccurate survival risk stratification in epithelial ovarian cancer remains challenging because prognostic information is distributed across heterogeneous clinical, histopathological, radiological, and molecular scales, while modality availability is often incomplete across cohorts. We present a compact fuzzy-rule decision-level fusion framework centered on a primary clinical-histopathology survival model and extended through controlled modality-extension analyses. The primary model operates on calibrated unimodal risk scores and integrates fuzzy membership embedding, rule screening, and compact rule distillation to produce a frozen survival score for downstream use. On the clinical-histopathology complete-case subsets, the compact model achieved C-indices of 0.6771 in TCGA-OV and 0.6085 in the independent Memorial Sloan Kettering Cancer Center cohort, and yielded the strongest external discrimination among the evaluated two-modality late-fusion comparators. Paired bootstrap analysis showed significant gains over the clinical unimodal baseline and QMF, while the remaining pairwise comparisons were directionally favorable but not uniformly significant. CT radiomics, evaluated as an auxiliary modality under incomplete availability, provided only modest local refinement and did not redefine the primary model. In the matched molecular subset, transcriptomics provided substantial complementary value beyond the frozen primary score, whereas reverse incremental analysis showed that the primary model retained nonredundant prognostic information beyond the molecular score. Exploratory biological analyses linked the joint molecular extension score to attenuation of immune- and module-related programs and to enrichment of extracellular-matrix and migratory pathways in high-risk tumors. These findings support a compact, interpretable, and deployment-oriented decision-level fusion strategy for ovarian cancer survival modeling. Jianmei Zhao, Yixin Liu 0005, Guohua Wang 0001, Murong Zhou |
IEEE Trans. Fuzzy Syst. | 3 |
| 2026 | Causality-Adjusted Data Augmentation for Domain Continual Medical Image SegmentationabstractIn domain continual medical image segmentation, distillation-based methods mitigate catastrophic forgetting by continuously reviewing old knowledge. However, these approaches often exhibit biases towards both new and old knowledge simultaneously due to confounding factors, which can undermine segmentation performance. To address these biases, we propose the Causality-Adjusted Data Augmentation (CauAug) framework, introducing a novel causal intervention strategy called the Texture-Domain Adjustment Hybrid-Scheme (TDAHS) alongside two causality-targeted data augmentation approaches: the Cross Kernel Network (CKNet) and the Fourier Transformer Generator (FTGen). (1) TDAHS establishes a domain-continual causal model that accounts for two types of knowledge biases by identifying irrelevant local textures (L) and domain-specific features (D) as confounders. It introduces a hybrid causal intervention that combines traditional confounder elimination with a proposed replacement approach to better adapt to domain shifts, thereby promoting causal segmentation. (2) CKNet eliminates confounder L to reduce biases in new knowledge absorption. It decreases reliance on local textures in input images, forcing the model to focus on relevant anatomical structures and thus improving generalization. (3) FTGen causally intervenes on confounder D by selectively replacing it to alleviate biases that impact old knowledge retention. It restores domain-specific features in images, aiding in the comprehensive distillation of old knowledge. Our experiments show that CauAug significantly mitigates catastrophic forgetting and surpasses existing methods in various medical image segmentation tasks. Zhanshi Zhu, Gongning Luo, Wei Wang 0169, Suyu Dong, Kuanquan Wang, Guohua Wang 0001, Shuo Li 0001 |
IEEE J. Biomed. Health Informatics | 8 |
| 2025 | RVC: A Real-Time Variant Calling Framework for Short-Read Sequencing DataabstractAccurate detection of single-nucleotide variants (SNVs) and small insertions/deletions (indels) from second-generation sequencing (NGS) data is essential for clinical applications such as cancer diagnostics, infectious disease monitoring, and rapid genetic screening. However, conventional variant calling pipelines, such as GATK, decouple analysis from sequencing, deferring detection until sequencing is fully completed. We introduce RVC, a real-time variant calling framework tailored for cycle-based NGS workflows. RVC incrementally processes partially sequenced reads and continuously updates variant evidence using a scanline-based alignment algorithm and a lightweight binomial scoring model. This design enables progressive, low-latency SNVs and indels detection during sequencing, without disrupting the sequencing pipeline. In benchmark experiments using the HG002 dataset, RVC completed variant calling within tens of minutes after sequencing, significantly outperforming GATK in runtime. By tightly integrating analysis with sequencing output, RVC bridges the gap between sequencing speed and clinical responsiveness, offering a scalable and practical solution for real-time genomic diagnostics. Miao Cui 0005, Tao Jiang 0021, Yadong Wang 0001, Bo Liu 0023, Guohua Wang 0001, Yadong Liu 0001 |
BIBM | 6 |
| 2025 | Cancer Drug Response Prediction Via Cross-Modal Multilevel Homogeneous and Heterogeneous Feature LearningabstractAccurately predicting cancer drug response (CDR) is crucial for personalized cancer therapies and drug repositioning. Efficient CDR prediction requires to integrate multimodal data including sequences, structures, multilevel omics, and diverse biological networks of drugs and cell lines, to capture intricate underlying patterns. So we proposed TransGCDR, a novel method that integrates cross-modal multilevel homogeneous and heterogeneous features to derive highly discriminative embeddings, thereby enhancing CDR prediction. TransGCDR first learns structural representations of drugs using a Transformer from fingerprint substructures and Graph Convolutional Networks for molecular graphs. Meanwhile, it utilizes a Graph Attention Network to extract cell line representations from graphs integrating multilevel omics data, including gene expression, copy number variation, and somatic mutations. Next, it generates homogeneous embeddings for drugs and cell lines from these cross-modal representations while extracting drug- and cell linecentric heterogeneous contextual embeddings from prior CDRs. These embeddings are then integrated using a Residual Attention Graph Convolutional Network to obtain comprehensive representations. Finally, an MLP-based model is trained on the learned embeddings to predict CDRs. Extensive evaluations demonstrate that TransGCDR outperforms state-of-the-art methods in CDR prediction across various metrics. Ablation studies confirm that every module enhances the precise CDR characterization from the multimodal and multilevel data, which is key to the success of TransGCDR. Moreover, case studies highlight the significant clinical potential of TransGCDR to predict CDRs. Zhixia Teng, Mingxin Yin, Guohua Wang 0001 |
BIBM | 4 |
| 2025 | HGANTLDA: A Hybrid Framework Integrating Sequence Language Modeling and Heterogeneous Graph Attention for lncRNA-Disease Association PredictionabstractLncRNAs have been confirmed by various studies to play an important role in the generation of a variety of diseases, making the accurate prediction of IncRNA-disease associations crucial for diagnosis and therapy. However, experimentally validated associations remain scarce. To address the limitations of existing computational methods regarding biological information coverage and model robustness, we developed HGANTLDA, a novel multimodal information fusion framework. We made three key contributions: 1) integrating the Nucleotide Transformer, a gene language pre-training model, to systematically encode IncRNA sequence semantic information, enhancing sequencelevel feature representation; 2) incorporating miRNA regulatory mechanisms by constructing a heterogeneous graph of IncRNA, miRNA, and disease nodes, and employing a HAN to learn representations from complex semantic paths among these nodes; 3) developing a function similarity-based negative sample selection strategy that significantly reduced pseudo-negative sample interference, effectively improving prediction stability. Extensive experimental results demonstrate that HGANTLDA achieves superior performance, with an AUC of 98.93% and an F1-score of 97.29%, representing a 2.7% improvement over the secondranked model. We have also developed an accessible online service system (http://62.234.10.79:80/) that allows users to perform predictions, customize models, and visualize IncRNA-disease association results interactively. SiCheng Xiang, Yuhai Zhao, Qiaoming Liu, Benzhi Dong, Guohua Wang 0001 |
BIBM | 6 |
| 2025 | Relational similarity-based graph contrastive learning for DTI predictionabstractAs part of the drug repurposing process, it is imperative to predict the interactions between drugs and target proteins in an accurate and efficient manner. With the introduction of contrastive learning into drug-target prediction, the accuracy of drug repurposing will be further improved. However, a large part of DTI prediction methods based on deep learning either focus only on the structural features of proteins and drugs extracted using GNN or CNN, or focus only on their relational features extracted using heterogeneous graph neural networks on a DTI heterogeneous graph. Since the structural and relational features of proteins and drugs describe their attribute information from different perspectives, their combination can improve DTI prediction performance. We propose a relational similarity-based graph contrastive learning for DTI prediction (RSGCL-DTI), which combines the structural and relational features of drugs and proteins to enhance the accuracy of DTI predictions. In our proposed method, the inter-protein relational features and inter-drug relational features are extracted from the heterogeneous drug-protein interaction network through graph contrastive learning, respectively. The results demonstrate that combining the relational features obtained by graph contrastive learning with the structural ones extracted by D-MPNN and CNN enhances feature representation ability, thereby improving DTI prediction performance. Our proposed RSGCL-DTI outperforms eight SOTA baseline models on the four benchmark datasets, performs well on the imbalanced dataset, and also shows excellent generalization ability on unseen drug-protein pairs. Jilong Bian, Limin Wei, Yang Li 0130, Guohua Wang 0001 |
Briefings Bioinform. | 5 |
| 2025 | SVHunter: long-read-based structural variation detection through the transformer modelabstractStructural variations (SVs) are genomic rearrangements larger than 50 bp, that are widely present in the human genome and are associated with various complex diseases. Existing long-read-based SV detection tools often rely on fixed rules or heuristic algorithms, which can oversimplify the complexity of SV signatures. Therefore, these methods usually lack flexibility and cannot fully capture SV signals, leading to reduced accuracy and robustness. To address these issues, we propose SVHunter, a transformer-based method for long-read SV detection. SVHunter combines convolutional neural networks and transformers to capture both local and global SV signatures, enabling accurate identification of SVs. Additionally, SVHunter employs the mean shift clustering algorithm, which dynamically adjusts bandwidth parameters to accommodate different types of SVs without requiring a preset number of clusters, thus allowing precise breakpoint clustering. Validation across multiple sequencing platforms and datasets demonstrates that SVHunter excels at detecting various types of SVs, with a notable reduction in the false discovery rate. This highlights considerable strong potential for both research and clinical applications. Runtian Gao, Heng Hu, Zhongjun Jiang, Shuqi Cao, Guohua Wang 0001, Tao Jiang 0021 |
Briefings Bioinform. | 5 |
| 2025 | TCRdesign: an antigen-specific generative language model for de novo design of T-cell receptorsabstractT-cell receptors (TCR), which are heterodimers of $\alpha $ and $\beta $ chains that recognize foreign antigens, are of great significance to current immunotherapy. Although artificial intelligence (AI) has explosively accelerated de novo protein design, the challenge of therapeutic TCR design has been overlooked by most researchers. Existing TCR engineering relies heavily on isolating antigen-specific TCRs from tumor tissues, which requires a large amount of labor resources and wet experimental verification. To mitigate this issue, we present TCRdesign, a pretrained generative protein language model (PLM) for the de novo design of artificial TCR $\beta $-chain complementarity-determining region 3 sequences conditioned on antigen-binding specificity (BS). In parallel, we develop a high-accuracy binding predictor (TCRBinder) that couples paired $\alpha $/$\beta $ chain information with antigen sequences to assess BS. Our in silico comparisons demonstrate that (i) TCRdesign surpasses state-of-the-art baselines in generating antigen-specific TCR sequences. The model leverages paired-chain coherence to refine amino-acid level interaction patterns. (ii) TCRdesign-generated TCR sequences exhibit better antigen binding capability to diverse oncogenic hotspots compared with natural counterparts. (iii) TCRdesign inherits the intrinsic properties of large PLMs, enabling effectively identify the determinant residues in TCR-antigen binding, which enhances its interpretability. These results highlight the significant capability of TCRdesign in understanding and generating TCR sequences with an antigen-specific interaction pattern, charting a versatile path toward AI-driven T-cell engineering for precision immunotherapy. Xiaokun Li, Qiang Yang 0015, Weihe Dong, Kuanquan Wang, Suyu Dong, Wei Wang 0169, Gongning Luo, Xianyu Zhang 0004, Tiansong Yang, Xin Gao 0001, Guohua Wang 0001 |
Briefings Bioinform. | 12 |
| 2025 | CKG-TPI: integrating collaborative knowledge graph with sequence interactions for TCR-peptide binding specificityabstractAccurately identifying interactions between T-cell receptors (TCRs) and peptides is a fundamental challenge in immunology, with significant implications for vaccine design and immunotherapy. While computational methods offer efficient alternatives to labor-intensive experimental screening, achieving robust and accurate TCR-peptide binding prediction remains a challenging task. To address this, we propose collaborative knowledge graph (CKG-TPI), a novel prediction framework based on graph neural networks that integrates both interaction patterns between TCR and peptide sequences and their higher-order biological context through a constructed collaborative knowledge graph. Experimental results on multiple publicly available independent datasets demonstrate that CKG-TPI consistently outperforms state-of-the-art models. Specifically, it achieves a 9.89% improvement in area under the ROC curve compared to the strongest baseline model UnifyImmun, and a 23.93% increase in area under the precision-recall curve over the leading baseline method. Moreover, attention weight visualization and peptide-specific TCR screening validate the model's effectiveness, underscoring its potential as a powerful tool for immunological research and therapeutic discovery. Yue Liu 0034, Haoyan Wang, Guohua Wang 0001, Yadong Liu 0001, Tao Jiang 0021, Yadong Wang 0001 |
Briefings Bioinform. | 3 |
| 2025 | Graph-based deep learning for integrating single-cell and bulk transcriptomic data to identify clinical cancer subtypesabstractThe integration of single-cell RNA sequencing (scRNA-seq) and bulk transcriptomic data has become essential for deciphering the complex heterogeneity of cancer and identifying clinical cancer subtypes. However, the inherent challenges posed by the high dimensionality, sparsity, and noise characteristics of scRNA-seq data have significantly hindered its widespread clinical translation. To address these limitations, we introduce single-cell and bulk transcriptomic graph deep learning, a graph-based deep learning method that synergistically integrates scRNA-seq and bulk transcriptomic data to precisely identify cancer subtypes and predict clinical outcomes. scBGDL constructs sample-specific gene graphs modeling complex gene-gene interactions and cellular relationships. The architecture employs Graph Attention Networks for feature aggregation, MinCutPool layers for dimensionality reduction, and Transformer modules to capture high-order biological dependencies. Independently validated in each of 16 distinct The Cancer Genome Atlas cancer types, scBGDL significantly outperformed existing methods in prognostic accuracy (mean C-index: 0.7060 versus 0.6709 max competitor), demonstrating robustness and generalizability to diverse transcriptional architectures. To demonstrate clinical versatility, we further evaluated scBGDL in three therapeutic contexts using multicenter cohorts: lung adenocarcinoma survival prediction (n = 1099), epithelial ovarian cancer platinum-based chemotherapy response (n = 762), skin cutaneous melanoma immunotherapy outcome (n = 305). scBGDL consistently delivered robust risk stratification (log-rank P < 0.05 across cohorts), identified key driver edges, and uncovered clinically relevant biological interpretations. By enabling multimodal data integration and interpretable biological insights, scBGDL advances precision oncology for prognosis prediction, therapy optimization, and biomarker discovery. The source code for scBGDL model is available online (https://github.com/NEFLab/scBGDL). Yixin Liu 0005, Guohua Wang 0001 |
Briefings Bioinform. | 5 |
| 2025 | SLGCA: spatial cross-level graph contrastive autoencoder for multislice spatial domain identification and microenvironment explorationabstractThe development of spatial transcriptomics (ST) technologies has enabled researchers to better understand cells' spatial organization and functional heterogeneity within their native tissue context. Spatial domain identification plays a crucial role in ST data analysis. However, most existing spatial domain identification methods do not fully exploit spatial information, and often fail to adequately integrate both local and global features, resulting in suboptimal spatial domain identification. We propose SLGCA, a novel method based on cross-level graph contrastive learning to address these challenges. SLGCA adopts a dual-channel learning mechanism, combining local-level contrastive learning based on spatial neighborhood information and global information contrastive learning across views, thereby significantly enhancing the accuracy of spatial domain identification. SLGCA can integrate multiple tissue sections without needing pre-alignment or external tools, eliminating batch effects and accurately identifying spatial domains across multiple slices. Experimental results show that SLGCA significantly outperforms the benchmark methods in spatial domain identification accuracy on ST data generated by multiple techniques. Moreover, SLGCA enables accurate dissection of tumor heterogeneity in human breast cancer datasets and effectively uncovers the heterogeneous tumor microenvironment in liver cancer, revealing two distinct fibroblast subtypes. Murong Zhou, Guohua Wang 0001, Qiaoming Liu |
Briefings Bioinform. | 3 |
| 2025 | Cross-RNA transferable sequence representation learning for lncRNA m6A site detection via novel deep domain separation networksabstractN6-methyladenosine (m6A) is a key epitranscriptomic marker enriched in long noncoding RNAs (lncRNAs) that is closely involved in complex disease mechanisms. Although accurate detection of m6A sites in lncRNAs is essential for understanding disease mechanisms, the development of effective computational predictors remains challenging due to the limited number of annotated sites. Moreover, most existing predictors are specifically designed for messenger RNAs (mRNAs) based on abundant mRNA-specific knowledge, yet they exhibit limited generalizability to lncRNAs. Given the similarities between mRNAs and lncRNAs, a transferable framework capable of leveraging their shared features is critical for advancing m6A site prediction in lncRNAs. To address this challenge, we propose DSNm6A, a deep learning framework that learns cross-RNA transferable sequence representations for effective lncRNA m6A site detection. To comprehensively capture patterns and signals of m6A sites, lncRNA and mRNA sequences are first encoded from complementary multiple facets, including One-Hot encoding, nucleotide physicochemical properties and cumulative frequency, and position-specific propensity. Based on these sequence encodings, a domain separation network integrating CNN, Bi-LSTM, and BERT modules is then employed to explicitly disentangle domain-invariant features shared between mRNAs and lncRNAs from their domain-specific counterparts. The shared features are finally fed into a fully connected layer for accurate lncRNA m6A sites prediction. Cross-validation and independent test results demonstrate that DSNm6A consistently outperforms existing methods across nearly all performance metrics, attributed to its superior capacity to learn transferable m6A-related features across RNA types. In addition, DSNm6A exhibits strong robustness and generalization across species. Zhixia Teng, Chunyu Wang 0002, Guohua Wang 0001 |
Briefings Bioinform. | 5 |
| 2025 | Long Noncoding RNA function prediction via multiview cross-contrastive learning combined with multiscale semantic adaptive optimizationabstractUnderstanding long noncoding RNA (lncRNA) function is essential for revealing molecular mechanisms and developing effective therapies for complex diseases, as lncRNAs play important regulatory roles in many disease-related biological processes. However, existing lncRNA function predictors struggle to extract discriminative features from multimodal omics data and to model the semantic and topological structure of the gene ontology (GO), which severely limits their ability to achieve biologically meaningful and functionally informative predictions. To address these challenges, we propose a novel framework for lncRNA function prediction, namely MiCLSAO. Firstly, MiCLSAO utilizes multiview cross-contrastive learning with attention mechanisms to extract highly discriminative lncRNA features from diverse omics similarity networks. Secondly, graph convolutional networks are applied to learn initial features of GO terms, while multiscale topological and semantic relationships are incorporated to adaptively refine term representations. Finally, an lncRNA function predictor is developed by dynamically integrating the representations of lncRNAs and GO terms using a Kolmogorov-Arnold network. Extensive experiments demonstrate that MiCLSAO consistently outperforms state-of-the-art methods across multiple metrics, with significant capability to recover known functions and uncover novel ones. Moreover, MiCLSAO demonstrates remarkable practical utility and potential value by providing more functionally informative annotations for lncRNAs. Zhixia Teng, Qingqi Li, Chunyu Wang 0002, Guohua Wang 0001 |
Briefings Bioinform. | 5 |
| 2025 | Enhancing LncRNA-miRNA interaction prediction with multimodal contrastive representation learningabstractInteractions between long non-coding RNAs (lncRNAs) and microRNAs (miRNAs) play an important role in the development of complex human diseases by collaboratively regulating gene transcription and expression. Therefore, identifying lncRNA-miRNA interactions (LMIs) is essential for diagnosing and treating complex human diseases. Because identifying LMIs with wet experiments is time-consuming and labor-intensive, some computational methods have been developed to infer LMIs. However, these approaches excel at utilizing single-modal information but struggle to integrate multimodal data from lncRNAs and miRNAs, which is essential for uncovering complex patterns in LMIs, ultimately limiting their performance. Therefore, this article proposes a novel multimodal contrastive representation learning model (MCRLMI) for LMI predictions. The model fully integrates multi-source similarity information and sequence encodings of lncRNAs and miRNAs. It leverages a graph convolutional network (GCN) and a Transformer to capture local neighborhood structural features and long-distance dependencies, respectively, enabling the collaborative modeling of structural and semantic information. Subsequently, to effectively integrate multimodal characteristics with encoded information, a multichannel attention mechanism and contrastive learning are introduced to fuse the extracted features. Finally, a Kolmogorov-Arnold Network (KAN) is trained with the optimized embeddings to predict LMIs. Extensive experiments show that the proposed MCRLMI consistently outperforms existing methods. Moreover, case studies further validate the potential of MCRLMI to identify novel LMIs in practical applications. Zhixia Teng, Zhaowen Tian, Murong Zhou, Guohua Wang 0001, Zhen Tian 0004 |
Briefings Bioinform. | 4 |
| 2025 | Meta learning for mutant HLA class I epitope immunogenicity prediction to accelerate cancer clinical immunotherapyabstractAccurate prediction of binding between human leukocyte antigen (HLA) class I molecules and antigenic peptide segments is a challenging task and a key bottleneck in personalized immunotherapy for cancer. Although existing prediction tools have demonstrated significant results using established datasets, most can only predict the binding affinity of antigenic peptides to HLA and do not enable the immunogenic interpretation of new antigenic epitopes. This limitation results from the training data for the computational models relying heavily on a large amount of peptide-HLA (pHLA) eluting ligand data, in which most of the candidate epitopes lack immunogenicity. Here, we propose an adaptive immunogenicity prediction model, named MHLAPre, which is trained on the large-scale MS-derived HLA I eluted ligandome (mostly presented by epitopes) that are immunogenic. Allele-specific and pan-allelic prediction models are also provided for endogenous peptide presentation. Using a meta-learning strategy, MHLAPre rapidly assessed HLA class I peptide affinities across the whole pHLA pairs and accurately identified tumor-associated endogenous antigens. During the process of adaptive immune response of T-cells, pHLA-specific binding in the antigen presentation is only a pre-task for CD8+ T-cell recognition. The key factor in activating the immune response is the interaction between pHLA complexes and T-cell receptors (TCRs). Therefore, we performed transfer learning on the pHLA model using the pHLA-TCR dataset. In pHLA binding task, MHLAPre demonstrated significant improvement in identifying neoepitope immunogenicity compared with five state-of-the-art models, proving its effectiveness and robustness. After transfer learning of the pHLA-TCR data, MHLAPre also exhibited relatively superior performance in revealing the mechanism of immunotherapy. MHLAPre is a powerful tool to identify neoepitopes that can interact with TCR and induce immune responses. We believe that the proposed method will greatly contribute to clinical immunotherapy, such as anti-tumor immunity, tumor-specific T-cell engineering, and personalized tumor vaccine. Qiang Yang 0015, Weihe Dong, Xiaokun Li, Kuanquan Wang, Suyu Dong, Xianyu Zhang 0004, Tiansong Yang, Gongning Luo, Xingyu Liao, Xin Gao 0001, Guohua Wang 0001 |
Briefings Bioinform. | 12 |
| 2025 | GAADE: identification spatially variable genes based on adaptive graph attention networkabstractThe rapid advancement of spatial transcriptomics (ST) sequencing technology has made it possible to capture gene expression with spatial coordinate information at the cellular level. Although many methods in ST data analysis can detect spatially variable genes (SVGs), these methods often fail to identify genes with explicit spatial expression patterns due to the lack of consideration for spatial domains. Considering spatial domains is crucial for identifying SVGs as it focuses the analysis of gene expression changes on biologically relevant regions, aiding in the more accurate identification of SVGs associated with specific cell types. Existing methods for identifying SVGs based on spatial domains predefine spot similarity before training, which prevents adaptive learning and limits generalizability across different tissues or samples. This limitation may also lead to inaccurate identification of specific genes at boundary regions. To address these issues, we present GAADE, an unsupervised neural network architecture based on graph-structured data representation learning. GAADE stacks encoder/decoder layers and integrates a self-attention mechanism to reconstruct node attributes and graph structure, effectively capturing spatial domain structures of different sections. Consequently, we confine the identification of SVGs within spatial domains. By performing differential expression analysis on spots within the target spatial domain and their multi-order neighbors, GAADE detects genes with enriched expression patterns within defined domains. Comparative evaluations with five other popular methods on ST datasets across four different species, regions and tissues demonstrate that GAADE exhibits superior performance in detecting SVGs and capturing the extent of spatial gene expression variation. Zhenao Wu, Zhongqian Zhao, Xingjie Zhao, Guohua Wang 0001 |
Briefings Bioinform. | 8 |
| 2025 | KansformerEPI: a deep learning framework integrating KAN and transformer for predicting enhancer-promoter interactionsabstractEnhancer-promoter interaction (EPI) is a critical component of gene regulation. Accurately predicting EPIs across diverse cell types can advance our understanding of the molecular mechanisms behind transcriptional regulation and provide valuable insights into the onset and progression of related diseases. At present, large-scale genome-wide EPI predictions typically rely on computational approaches. However, most of these methods focus on predicting EPIs within a single cell line and lack a global perspective encompassing multiple cell lines. Furthermore, they often fail to fully account for the nonlinear relationships between features, leading to suboptimal prediction accuracy. In this study, we propose KansformerEPI, a global EPI prediction model designed for multiple cell lines. The model is built on Kansformer, an encoder that integrates KAN and Transformer, effectively capturing the nonlinear relationships among various epigenetic and sequence features. We utilized KansformerEPI to achieve cross-tissue prediction of EPIs across different cell types. This approach enhances the model's scalability, eliminating the complexity of designing separate prediction models for individual tissues. As a result, our model is applicable to various tissues, thereby reducing dependency on extensive datasets. Experimental results demonstrate that KansformerEPI surpasses existing methods such as TransEPI, TargetFinder, and SPEID in both accuracy and stability of EPI predictions across datasets including HMEC, IMR90, K562, and NHEK. Saihong Shao, Zhongqian Zhao, Xingjie Zhao, Zhaoxiang Zhang 0001, Guohua Wang 0001 |
Briefings Bioinform. | 8 |
| 2025 | cfDiffusion: diffusion-based efficient generation of high quality scRNA-seq data with classifier-free guidanceabstractSingle-cell RNA sequencing (scRNA-seq) technology provides a powerful means to measure gene expression at the individual cell level, thereby uncovering the intricate cellular heterogeneity that underlies various biological processes, including embryonic development, tumor metastasis, and microbial reproduction. However, the variable amounts of data generated across different cell types within tissues can compromise the accuracy of downstream analyses. Traditional approaches for generating scRNA-seq simulation data often rely on predefined data distributions, which can negatively impact the quality of the simulated data. Furthermore, these methods typically focus on simulating single-attribute cells, necessitating substantial additional data for the simulation of multi-attribute cells, which can lead to increased training times. To address these limitations, we propose cfDiffusion, a novel method grounded in diffusion models that incorporates Classifier-Free Guidance and a high-level feature caching mechanism. By leveraging Classifier-Free Guidance, cfDiffusion significantly reduces the training costs associated with model development compared to traditional Classifier Guidance methods. The integration of a caching mechanism further enhances efficiency by shortening inference times. While the inference duration of cfDiffusion remains longer than that of scDiffusion, it exhibits superior expressiveness and efficiency in generating multi-attribute single-cell data. Evaluated across datasets from multiple sequencing platforms, cfDiffusion consistently outperforms state-of-the-art models across various performance metrics. Additionally, cfDiffusion enables the simulation of single-cell data along a pseudo-time scale, facilitating advanced analyses such as tracking cell differentiation, investigating intercellular communication, and elucidating cellular heterogeneity. Zhongqian Zhao, Jixiang Ren, Guohua Wang 0001 |
Briefings Bioinform. | 6 |
| 2025 | VGAE-CCI: variational graph autoencoder-based construction of 3D spatial cell-cell communication networkabstractCell-cell communication plays a critical role in maintaining normal biological functions, regulating development and differentiation, and controlling immune responses. The rapid development of single-cell RNA sequencing and spatial transcriptomics sequencing (ST-seq) technologies provides essential data support for in-depth and comprehensive analysis of cell-cell communication. However, ST-seq data often contain incomplete data and systematic biases, which may reduce the accuracy and reliability of predicting cell-cell communication. Furthermore, other methods for analyzing cell-cell communication mainly focus on individual tissue sections, neglecting cell-cell communication across multiple tissue layers, and fail to comprehensively elucidate cell-cell communication networks within three-dimensional tissues. To address the aforementioned issues, we propose VGAE-CCI, a deep learning framework based on the Variational Graph Autoencoder, capable of identifying cell-cell communication across multiple tissue layers. Additionally, this model can be applied to spatial transcriptomics data with missing or partially incomplete data and can clustered cells at single-cell resolution based on spatial encoding information within complex tissues, thereby enabling more accurate inference of cell-cell communication. Finally, we tested our method on six datasets and compared it with other state of art methods for predicting cell-cell communication. Our method outperformed other methods across multiple metrics, demonstrating its efficiency and reliability in predicting cell-cell communication. Zhenao Wu, Jixiang Ren, Zhongqian Zhao, Guohua Wang 0001, Tao Wang 0082 |
Briefings Bioinform. | 7 |
| 2025 | scATD: a high-throughput and interpretable framework for single-cell cancer drug resistance prediction and biomarker identificationabstractTransfer learning has been widely applied to drug sensitivity prediction based on single-cell RNA sequencing, leveraging knowledge from large datasets of cancer cell lines or other sources to improve the prediction of drug responses. However, previous studies require model fine-tuning for different patient single-cell datasets, limiting their ability to meet the clinical need for high-throughput rapid prediction. In this research, we introduce single-cell Adaptive Transfer and Distillation model (scATD), a transfer learning framework leveraging large language models for high-throughput drug sensitivity prediction. Based on different large language models (scFoundation and Geneformer) and transfer strategies, scATD includes three distinct sub-models: scATD-sf, scATD-gf, and scATD-sf-dist. scATD-sf and scATD-gf employs an important bidirectional style transfer to enable predictions for new patients without model parameter training. Additionally, scATD-sf-dist uses knowledge distillation from large models to enhance prediction performance, improve efficiency, and reduce resource requirements. Benchmarking across more diverse datasets demonstrates scATD's superior accuracy, generalization and efficiency. Besides, by rigorously selecting reference background samples for feature attribution algorithms, scATD also provides more meaningful insights into the relationship between gene expression and drug resistance mechanisms. Making scATD more interpretability for addressing critical challenges in precision oncology. Murong Zhou, Zeyu Luo, Yu-Hang Yin, Qiaoming Liu, Guohua Wang 0001 |
Briefings Bioinform. | 5 |
| 2025 | MolPrompt: improving multi-modal molecular pre-training with knowledge promptsabstractMOTIVATION: Molecular pre-training has emerged as a foundational approach in computational drug discovery, enabling the extraction of expressive molecular representations from large-scale unlabeled datasets. However, existing methods largely focus on topological or structural features, often neglecting critical physicochemical attributes embedded in molecular systems. RESULT: We present MolPrompt, a knowledge-enhanced multimodal pre-training framework that integrates molecular graphs and textual descriptions via contrastive learning. MolPrompt employs a dual-encoder architecture consisting of Graphormer for graph encoding and BERT for textual encoding, and introduces knowledge prompts, semantic embeddings constructed by converting molecular descriptors into natural language, into the graph encoder to guide structure-aware representation learning. Across tasks including molecular property prediction, toxicity estimation, cross-modal retrieval, and anticancer inhibitor identification, MolPrompt consistently surpasses state-of-the-art baselines. These results highlight the value of embedding domain knowledge into structural learning to improve the depth, interpretability, and transferability of molecular representations. AVAILABILITY AND IMPLEMENTATION: The source code of MolPrompt is available at: https://github.com/catly/MolPrompt. Yang Li 0130, Chang Liu 0082, Xin Gao 0001, Guohua Wang 0001 |
Bioinform. | 4 |
| 2025 | KGCLMDA: a computational model for predicting latent associations of microbial drugs using knowledge graphs and contrastive learningabstractMOTIVATION: Predicting microbe-drug associations (MDgAs) is critical for understanding the role of microbes in drug metabolism, exploring their interactions with host physiology, and advancing personalized therapy. However, traditional methods face challenges in dealing with data sparsity, information imbalance, and the extraction of complex biological knowledge, which limit the accurate prediction of microbe-drug associations. Therefore, developing a computational model that can efficiently integrate multi-source data and address the challenges of data sparsity and information imbalance is essential. RESULTS: The paper proposes a model that integrates knowledge graphs and contrastive learning. By constructing both local and non-local association graphs, the model effectively captures the complex relationships between microbes and drugs. We preprocess and model the embedding representations of microbes and drugs, and design a multi-level interactive contrastive learning mechanism to optimize the information flow both within and outside the graph. Experimental results show that our model significantly outperforms existing methods in metrics such as AUC and AUPR, providing an efficient solution for predicting microbe-drug associations. AVAILABILITY AND IMPLEMENTATION: The source code is available at: https://github.com/SJshujuan/KGCLMDA. The code used in this study is also available on Zenodo: https://doi.org/10.5281/zenodo.16754402. Shujuan Su, Guohua Wang 0001 |
Bioinform. | 3 |
| 2025 | PLiSAGE: enhancing protein-ligand interaction prediction with multimodal surface and geometry encodingabstractMOTIVATION: Accurately predicting protein-ligand interactions is fundamental to elucidating molecular recognition and has far-reaching implications in drug discovery, gene regulation, and signal transduction. Conventional methods predominantly rely on internal structural or sequence-based protein representations. While these approaches have improved predictive performance, their dependence on limited labeled data restricts the capacity to learn expressive features from structural inputs. Moreover, they often neglect the intricate geometric and chemical context encoded on protein surfaces, limiting interpretability, and hindering mechanistic insights into binding interactions. RESULT: Here, we present PLiSAGE, a multimodal framework that integrates 3D structural and surface geometric embeddings to enable accurate prediction of protein-ligand interactions. Central to our approach is the joint pretraining of structural and surface encoders through unsupervised contrastive learning and point cloud reconstruction. Protein surfaces are represented as segmented point cloud patches, allowing the model to capture fine-grained geometric and chemical cues. A Transformer-based encoder further captures both local and global spatial dependencies across patches. The incorporation of spatial topological information during pretraining facilitates the learning of stable, discriminative, and multi-scale protein representations, enhancing the expressive capacity of both modalities. An adaptive fusion module dynamically integrates structural and surface embeddings to yield complete and robust protein representations. PLiSAGE demonstrates superior performance over competitive baselines in binding affinity prediction and interaction classification tasks. Extensive ablation studies underscore the critical contributions of surface features and the pretraining strategy to the model's generalization capabilities. AVAILABILITY AND IMPLEMENTATION: The source code of PLiSAGE is available at: https://github.com/catly/PLiSAGE. Guanyu Qiao, Guohua Wang 0001, Yang Li 0130 |
Bioinform. | 3 |
| 2025 | Multi-omics single-cell data alignment and integration with enhanced contrastive learning and differential attention mechanismabstractMOTIVATION: Identifying cell types that constitute complex tissue components using single-cell sequencing data is a critical issue in the field of biology. With the continuous advancement of sequencing technologies, the recognition of cell types has evolved from analyzing single-omics scRNA-seq data to integrating multi-omics single-cell data. However, existing methods for integrative analysis of high-dimensional multi-omics single-cell sequencing data have several limitations, including reliance on specific distribution assumptions of the data, sensitivity to noise, and clustering accuracy constrained by independent clustering methods. These issues have restricted improvements in the accuracy of cell type identification and hindered the application of such methods to large-scale datasets for cell type recognition. To address these challenges, we propose a novel method for aligning and integrating single-cell multi-omics data-scECDA. RESULTS: The scECDA employs independently designed autoencoders that can autonomously learn the feature distributions of each omics dataset. By incorporating enhanced contrastive learning and differential attention mechanisms, the scECDA effectively reduces the interference of noise during data integration. The model design exhibits high flexibility, enabling adaptation to single-cell omics data generated by different technological platforms. It directly outputs integrated latent features and end-to-end cell clustering results. Through the analysis of the distribution of latent features, the scECDA can effectively identify key biological markers and precisely distinguish cell subtypes, recover cluster-specific motif and infer trajectory. The scECDA was applied to eight paired single-cell multi-omics datasets, covering data generated by 10X Multiome, CITE-seq, and TEA-seq technologies. Compared to eight state-of-the-art methods, scECDA demonstrated higher accuracy in cell clustering. AVAILABILITY AND IMPLEMENTATION: The scECDA code is freely available at https://github.com/SuperheroBetter/scECDA. Zhongqian Zhao, Zhenao Wu, Fang Wang 0028, Guohua Wang 0001 |
Bioinform. | 6 |
| 2025 | Drug repositioning by collaborative learning based on graph convolutional inductive network
Zhixia Teng, Yongliang Li, Zhen Tian 0004, Yingjian Liang, Guohua Wang 0001 |
Future Gener. Comput. Syst. | 5 |
| 2025 | Fine-grained multimodal molecular pretraining via prompt learning
Yang Li 0130, Zhengxin Wei, Chang Liu 0082, Guohua Wang 0001 |
Knowl. Based Syst. | 4 |
| 2025 | MVHGCN: Predicting circRNA-disease associations with multi-view heterogeneous graph convolutional neural networksabstractCircular RNA, a class of RNA molecules gaining widespread attentions, has been widely recognized as a potential biomarker for many diseases. In recent years, significant progress has been made in the study of the associations between circRNA and diseases. However, traditional experimental methods are often inefficient and costly, making computational models an effective alternative. Nevertheless, existing computational methods still face challenges such as data sparsity and the difficulty of confirming negative samples, which limits the accuracy of predictions. To address these challenges, a novel computational method, namely MVHGCN, is proposed based on multi-view and graph convolutional networks to predict potential associations between circRNA and diseases. MVHGCN first constructs a heterogeneous graph and generates feature descriptors by integrating multiple databases. Then it extracts different connection views of circRNA and diseases through meta-paths, maximizing the utilization of known association information, and aggregates deep feature information through graph convolutional networks. Finally, a MLP is used to predict the association scores. The experimental results show that MVHGCN significantly outperforms existing methods on benchmark datasets by 5-fold cross-validation. This research provides an effective new approach to studying the associations between circRNAs and diseases, capable of alleviating the problem of data sparsity and accurately identifying potential associations. Yan Miao, Chunyu Wang 0002, Zhenyuan Sun, Guohua Wang 0001 |
PLoS Comput. Biol. | 5 |
| 2025 | THLANet: A deep learning framework for predicting TCR-pHLA binding in immunotherapy applicationsabstractAdaptive immunity is a targeted immune response that enables the body to identify and eliminate foreign pathogens, playing a critical role in the anti-tumor immune response. Tumor cell expression of antigens forms the foundation for inducing this adaptive response. However, the human leukocyte antigens (HLA)-restricted recognition of antigens by T-cell receptors (TCR) limits their ability to detect all neoantigens, with only a small subset capable of activating T-cells. Accurately predicting neoantigen binding to TCR is, therefore, crucial for assessing their immunogenic potential in clinical settings. We present THLANet, a deep learning model designed to predict the binding specificity of TCR to neoantigens presented by class I HLAs. THLANet employs evolutionary scale modeling-2 (ESM-2), replacing the traditional embedding methods to enhance sequence feature representation. Using scTCR-seq data, we obtained the TCR immune repertoire and constructed a TCR-pHLA binding database to validate THLANet's clinical potential. The model's performance was further evaluated using clinical cancer data across various cancer types. Additionally, by analyzing divided complementarity-determining region (CDR3) sequences and simulating alanine scanning of antigen sequences, we provided new insights into the 3D binding interactions of TCRs and antigens. Predicting TCR-neoantigen pairing remains a significant challenge in immunology, THLANet provides accurate predictions using only the TCR sequence (CDR3β), antigen sequence, and class I HLA, offering novel insights into TCR-antigen interactions. Qiang Yang 0015, Weihe Dong, Xiaokun Li, Kuanquan Wang, Suyu Dong, Gongning Luo, Xianyu Zhang 0004, Tiansong Yang, Xin Gao 0001, Guohua Wang 0001 |
PLoS Comput. Biol. | 11 |
| 2025 | Adjacency-Aware Fuzzy Label Learning for Skin Disease DiagnosisabstractAutomatic acne severity grading is crucial for the accurate diagnosis and effective treatment of skin diseases. However, the acne severity grading process is often ambiguous due to the similar appearance of acne with close severity, making it challenging to achieve reliable acne severity grading. Following the idea of fuzzy logic for handling uncertainty in decision-making, we transforms the acne severity grading task into a fuzzy label learning (FLL) problem, and propose a novel adjacency-aware fuzzy label learning (AFLL) framework to handle uncertainties in this task. The AFLL framework makes four significant contributions, each demonstrated to be highly effective in extensive experiments. First, we introduce a novel adjacency-aware decision sequence generation method that enhances sequence tree construction by reducing bias and improving discriminative power. Second, we present a consistency-guided decision sequence prediction method that mitigates error propagation in hierarchical decision-making through a novel selective masking decision strategy. Third, our proposed sequential conjoint distribution loss innovatively captures the differences for both high and low fuzzy memberships across the entire fuzzy label set while modeling the internal temporal order among different acne severity labels with a cumulative distribution, leading to substantial improvements in FLL. Fourth, to the best of our knowledge, AFLL is the first approach to explicitly address the challenge of distinguishing adjacent categories in acne severity grading tasks. Experimental results on the public ACNE04 dataset demonstrate that AFLL significantly outperforms existing methods, establishing a new state-of-the-art in acne severity grading. Murong Zhou, Baifu Zuo, Guohua Wang 0001, Gongning Luo, Fanding Li, Suyu Dong, Wei Wang 0169, Kuanquan Wang, Xiangyu Li 0004, Lifeng Xu |
IEEE Trans. Fuzzy Syst. | 3 |
| 2024 | scEAGC: an efficient anchor graph clustering for single-cell transcriptomics and proteomics dataabstractSingle-cell multi-omics sequencing allows researchers to simultaneously sequence multiple types of molecular information from the same individual cell, like transcriptomics and proteomics. However, the research of identifying cell types from single-cell multi-omics data is still challenging. In this article, we proposed scEAGC, an efficient anchor graph clustering for single-cell transcriptomics and proteomics data. It first constructs anchor cell graphs for every omics and then integrates separated omics-specific anchor cell graphs on a weighted multi-view clustering model with F-norm and the Orthogonal constraint, finally through the divided iterative optimization method to obtain cluster partition without any extra post-processing. Since scEAGC combines the high-efficiency property of anchor graph clustering, its efficiency is substantially higher than widely used algorithms. Extensive experiments demonstrate that, compared to other state-of-the-art clustering algorithms, scEAGC can boost clustering accuracy and robustness and detect the new cell subtypes in CITE-seq and scRNA-seq data. Qiaoming Liu, Yadong Wang 0001, Guohua Wang 0001 |
BIBM | 3 |
| 2024 | Hierarchical multimodal self-attention-based graph neural network for DTI predictionabstractDrug-target interactions (DTIs) are a key part of drug development process and their accurate and efficient prediction can significantly boost development efficiency and reduce development time. Recent years have witnessed the rapid advancement of deep learning, resulting in an abundance of deep learning-based models for DTI prediction. However, most of these models used a single representation of drugs and proteins, making it difficult to comprehensively represent their characteristics. Multimodal data fusion can effectively compensate for the limitations of single-modal data. However, existing multimodal models for DTI prediction do not take into account both intra- and inter-modal interactions simultaneously, resulting in limited presentation capabilities of fused features and a reduction in DTI prediction accuracy. A hierarchical multimodal self-attention-based graph neural network for DTI prediction, called HMSA-DTI, is proposed to address multimodal feature fusion. Our proposed HMSA-DTI takes drug SMILES, drug molecular graphs, protein sequences and protein 2-mer sequences as inputs, and utilizes a hierarchical multimodal self-attention mechanism to achieve deep fusion of multimodal features of drugs and proteins, enabling the capture of intra- and inter-modal interactions between drugs and proteins. It is demonstrated that our proposed HMSA-DTI has significant advantages over other baseline methods on multiple evaluation metrics across five benchmark datasets. Jilong Bian, Guanghui Dong, Guohua Wang 0001 |
Briefings Bioinform. | 4 |
| 2024 | SVDF: enhancing structural variation detect from long-read sequencing via automatic filtering strategiesabstractStructural variation (SV) is an important form of genomic variation that influences gene function and expression by altering the structure of the genome. Although long-read data have been proven to better characterize SVs, SVs detected from noisy long-read data still include a considerable portion of false-positive calls. To accurately detect SVs in long-read data, we present SVDF, a method that employs a learning-based noise filtering strategy and an SV signature-adaptive clustering algorithm, for effectively reducing the likelihood of false-positive events. Benchmarking results from multiple orthogonal experiments demonstrate that, across different sequencing platforms and depths, SVDF achieves higher calling accuracy for each sample compared to several existing general SV calling tools. We believe that, with its meticulous and sensitive SV detection capability, SVDF can bring new opportunities and advancements to cutting-edge genomic research. Heng Hu, Runtian Gao, Zhongjun Jiang, Murong Zhou, Guohua Wang 0001, Tao Jiang 0021 |
Briefings Bioinform. | 7 |
| 2024 | scLEGA: an attention-based deep clustering method with a tendency for low expression of genes on single-cell RNA-seq dataabstractSingle-cell RNA sequencing (scRNA-seq) enables the exploration of biological heterogeneity among different cell types within tissues at a resolution. Inferring cell types within tissues is foundational for downstream research. Most existing methods for cell type inference based on scRNA-seq data primarily utilize highly variable genes (HVGs) with higher expression levels as clustering features, overlooking the contribution of HVGs with lower expression levels. To address this, we have designed a novel cell type inference method for scRNA-seq data, termed scLEGA. scLEGA employs a novel zero-inflated negative binomial (ZINB) loss function that fully considers the contribution of genes with lower expression levels and combines two distinct scRNA-seq clustering strategies through a multi-head attention mechanism. It utilizes a low-expression optimized denoising autoencoder, based on the novel ZINB model, to extract low-dimensional features and handle dropout events, and a GCN-based graph autoencoder (GAE) that leverages neighbor information to guide dimensionality reduction. The iterative fusion of denoising and topological embedding in scLEGA facilitates the acquisition of cluster-friendly cell representations in the hidden embedding, where similar cells are brought closer together. Compared to 12 state-of-the-art cell type inference methods on 15 scRNA-seq datasets, scLEGA demonstrates superior performance in clustering accuracy, scalability, and stability. Our scLEGA model codes are freely available at https://github.com/Masonze/scLEGA-main. Zhenze Liu, Yingjian Liang, Guohua Wang 0001 |
Briefings Bioinform. | 3 |
| 2024 | DrugReSC: targeting disease-critical cell subpopulations with single-cell transcriptomic data for drug repurposing in cancerabstractThe field of computational drug repurposing aims to uncover novel therapeutic applications for existing drugs through high-throughput data analysis. However, there is a scarcity of drug repurposing methods leveraging the cellular-level information provided by single-cell RNA sequencing data. To address this need, we propose DrugReSC, an innovative approach to drug repurposing utilizing single-cell RNA sequencing data, intending to target specific cell subpopulations critical to disease pathology. DrugReSC constructs a drug-by-cell matrix representing the transcriptional relationships between individual cells and drugs and utilizes permutation-based methods to assess drug contributions to cellular phenotypic changes. We demonstrate DrugReSC's superior performance compared to existing drug repurposing methods based on bulk or single-cell RNA sequencing data across multiple cancer case studies. In summary, DrugReSC offers a novel perspective on the utilization of single-cell sequencing data in drug repurposing methods, contributing to the advancement of precision medicine for cancer. Chonghui Liu, Yingjian Liang, Guohua Wang 0001 |
Briefings Bioinform. | 5 |
| 2024 | DeePhafier: a phage lifestyle classifier using a multilayer self-attention neural network combining protein informationabstractBacteriophages are the viruses that infect bacterial cells. They are the most diverse biological entities on earth and play important roles in microbiome. According to the phage lifestyle, phages can be divided into the virulent phages and the temperate phages. Classifying virulent and temperate phages is crucial for further understanding of the phage-host interactions. Although there are several methods designed for phage lifestyle classification, they merely either consider sequence features or gene features, leading to low accuracy. A new computational method, DeePhafier, is proposed to improve classification performance on phage lifestyle. Built by several multilayer self-attention neural networks, a global self-attention neural network, and being combined by protein features of the Position Specific Scoring Matrix matrix, DeePhafier improves the classification accuracy and outperforms two benchmark methods. The accuracy of DeePhafier on five-fold cross-validation is as high as 87.54% for sequences with length >2000bp. Yan Miao, Zhenyuan Sun, Haoran Gu, Chenjing Ma, Yingjian Liang, Guohua Wang 0001 |
Briefings Bioinform. | 7 |
| 2024 | VirGrapher: a graph-based viral identifier for long sequences from metagenomesabstractViruses are the most abundant biological entities on earth and are important components of microbial communities. A metagenome contains all microorganisms from an environmental sample. Correctly identifying viruses from these mixed sequences is critical in viral analyses. It is common to identify long viral sequences, which has already been passed thought pipelines of assembly and binning. Existing deep learning-based methods divide these long sequences into short subsequences and identify them separately. This makes the relationships between them be omitted, leading to poor performance on identifying long viral sequences. In this paper, VirGrapher is proposed to improve the identification performance of long viral sequences by constructing relationships among short subsequences from long ones. VirGrapher see a long sequence as a graph and uses a Graph Convolutional Network (GCN) model to learn multilayer connections between nodes from sequences after a GCN-based node embedding model. VirGrapher achieves a better AUC value and accuracy on validation set, which is better than three benchmark methods. Yan Miao, Zhenyuan Sun, Chenjing Ma, Guohua Wang 0001, Chunxue Yang |
Briefings Bioinform. | 5 |
| 2024 | Deep learning in template-free de novo biosynthetic pathway design of natural productsabstractNatural products (NPs) are indispensable in drug development, particularly in combating infections, cancer, and neurodegenerative diseases. However, their limited availability poses significant challenges. Template-free de novo biosynthetic pathway design provides a strategic solution for NP production, with deep learning standing out as a powerful tool in this domain. This review delves into state-of-the-art deep learning algorithms in NP biosynthesis pathway design. It provides an in-depth discussion of databases like Kyoto Encyclopedia of Genes and Genomes (KEGG), Reactome, and UniProt, which are essential for model training, along with chemical databases such as Reaxys, SciFinder, and PubChem for transfer learning to expand models' understanding of the broader chemical space. It evaluates the potential and challenges of sequence-to-sequence and graph-to-graph translation models for accurate single-step prediction. Additionally, it discusses search algorithms for multistep prediction and deep learning algorithms for predicting enzyme function. The review also highlights the pivotal role of deep learning in improving catalytic efficiency through enzyme engineering, which is essential for enhancing NP production. Moreover, it examines the application of large language models in pathway design, enzyme discovery, and enzyme engineering. Finally, it addresses the challenges and prospects associated with template-free approaches, offering insights into potential advancements in NP biosynthesis pathway design. Xueying Xie, Lin Gui 0008, Baixue Qiao, Guohua Wang 0001, Shanwen Sun |
Briefings Bioinform. | 4 |
| 2024 | HLAIImaster: a deep learning method with adaptive domain knowledge predicts HLA II neoepitope immunogenic responsesabstractWhile significant strides have been made in predicting neoepitopes that trigger autologous CD4+ T cell responses, accurately identifying the antigen presentation by human leukocyte antigen (HLA) class II molecules remains a challenge. This identification is critical for developing vaccines and cancer immunotherapies. Current prediction methods are limited, primarily due to a lack of high-quality training epitope datasets and algorithmic constraints. To predict the exogenous HLA class II-restricted peptides across most of the human population, we utilized the mass spectrometry data to profile >223 000 eluted ligands over HLA-DR, -DQ, and -DP alleles. Here, by integrating these data with peptide processing and gene expression, we introduce HLAIImaster, an attention-based deep learning framework with adaptive domain knowledge for predicting neoepitope immunogenicity. Leveraging diverse biological characteristics and our enhanced deep learning framework, HLAIImaster is significantly improved against existing tools in terms of positive predictive value across various neoantigen studies. Robust domain knowledge learning accurately identifies neoepitope immunogenicity, bridging the gap between neoantigen biology and the clinical setting and paving the way for future neoantigen-based therapies to provide greater clinical benefit. In summary, we present a comprehensive exploitation of the immunogenic neoepitope repertoire of cancers, facilitating the effective development of "just-in-time" personalized vaccines. Qiang Yang 0015, Weihe Dong, Xiaokun Li, Kuanquan Wang, Suyu Dong, Xianyu Zhang 0004, Tiansong Yang, Feng Jiang 0001, Bin Zhang 0042, Gongning Luo, Xin Gao 0001, Guohua Wang 0001 |
Briefings Bioinform. | 13 |
| 2024 | CPPLS-MLP: a method for constructing cell-cell communication networks and identifying related highly variable genes based on single-cell sequencing and spatial transcriptomics dataabstractIn the growth and development of multicellular organisms, the immune processes of the immune system and the maintenance of the organism's internal environment, cell communication plays a crucial role. It exerts a significant influence on regulating internal cellular states such as gene expression and cell functionality. Currently, the mainstream methods for studying intercellular communication are focused on exploring the ligand-receptor-transcription factor and ligand-receptor-subunit scales. However, there is relatively limited research on the association between intercellular communication and highly variable genes (HVGs). As some HVGs are closely related to cell communication, accurately identifying these HVGs can enhance the accuracy of constructing cell communication networks. The rapid development of single-cell sequencing (scRNA-seq) and spatial transcriptomics technologies provides a data foundation for exploring the relationship between intercellular communication and HVGs. Therefore, we propose CPPLS-MLP, which can identify HVGs closely related to intercellular communication and further analyze the impact of Multiple Input Multiple Output cellular communication on the differential expression of these HVGs. By comparing with the commonly used method CCPLS for constructing intercellular communication networks, we validated the superior performance of our method in identifying cell-type-specific HVGs and effectively analyzing the influence of neighboring cell types on HVG expression regulation. Source codes for the CPPLS_MLP R, python packages and the related scripts are available at 'CPPLS_MLP Github [https://github.com/wuzhenao/CPPLS-MLP]'. Zhenao Wu, Jixiang Ren, Guohua Wang 0001 |
Briefings Bioinform. | 6 |
| 2024 | GTAD: a graph-based approach for cell spatial composition inference from integrated scRNA-seq and ST-seq dataabstractWith the emergence of spatial transcriptome sequencing (ST-seq), research now heavily relies on the joint analysis of ST-seq and single-cell RNA sequencing (scRNA-seq) data to precisely identify cell spatial composition in tissues. However, common methods for combining these datasets often merge data from multiple cells to generate pseudo-ST data, overlooking topological relationships and failing to represent spatial arrangements accurately. We introduce GTAD, a method utilizing the Graph Attention Network for deconvolution of integrated scRNA-seq and ST-seq data. GTAD effectively captures cell spatial relationships and topological structures within tissues using a graph-based approach, enhancing cell-type identification and our understanding of complex tissue cellular landscapes. By integrating scRNA-seq and ST data into a unified graph structure, GTAD outperforms traditional 'pseudo-ST' methods, providing robust and information-rich results. GTAD performs exceptionally well with synthesized spatial data and accurately identifies cell spatial composition in tissues like the mouse cerebral cortex, cerebellum, developing human heart and pancreatic ductal carcinoma. GTAD holds the potential to enhance our understanding of tissue microenvironments and cellular diversity in complex bio-logical systems. The source code is available at https://github.com/zzhjs/GTAD. Benzhi Dong, Guohua Wang 0001 |
Briefings Bioinform. | 5 |
| 2024 | Causal enhanced drug-target interaction prediction based on graph generation and multi-source information fusionabstractMOTIVATION: The prediction of drug-target interaction is a vital task in the biomedical field, aiding in the discovery of potential molecular targets of drugs and the development of targeted therapy methods with higher efficacy and fewer side effects. Although there are various methods for drug-target interaction (DTI) prediction based on heterogeneous information networks, these methods face challenges in capturing the fundamental interaction between drugs and targets and ensuring the interpretability of the model. Moreover, they need to construct meta-paths artificially or a lot of feature engineering (prior knowledge), and graph generation can fuse information more flexibly without meta-path selection. RESULTS: We propose a causal enhanced method for drug-target interaction (CE-DTI) prediction that integrates graph generation and multi-source information fusion. First, we represent drugs and targets by modeling the fusion of their multi-source information through automatic graph generation. Once drugs and targets are combined, a network of drug-target pairs is constructed, transforming the prediction of drug-target interactions into a node classification problem. Specifically, the influence of surrounding nodes on the central node is separated into two groups: causal and non-causal variable nodes. Causal variable nodes significantly impact the central node's classification, while non-causal variable nodes do not. Causal invariance is then used to enhance the contrastive learning of the drug-target pairs network. Our method demonstrates excellent performance compared with other competitive benchmark methods across multiple datasets. At the same time, the experimental results also show that the causal enhancement strategy can explore the potential causal effects between DTPs, and discover new potential targets. Additionally, case studies demonstrate that this method can identify potential drug targets. AVAILABILITY AND IMPLEMENTATION: The source code of AdaDR is available at: https://github.com/catly/CE-DTI. Guanyu Qiao, Guohua Wang 0001, Yang Li 0130 |
Bioinform. | 2 |
| 2024 | Improving ncRNA family prediction using multi-modal contrastive learning of sequence and structureabstractMOTIVATION: Recent advancements in high-throughput sequencing technology have significantly increased the focus on non-coding RNA (ncRNA) research within the life sciences. Despite this, the functions of many ncRNAs remain poorly understood. Research suggests that ncRNAs within the same family typically share similar functions, underlining the importance of understanding their roles. There are two primary methods for predicting ncRNA families: biological and computational. Traditional biological methods are not suitable for large-scale data prediction due to the significant human and resource requirements. Concurrently, most existing computational methods either rely solely on ncRNA sequence data or are exclusively based on the secondary structure of ncRNA molecules. These methods fail to fully utilize the rich multimodal information available from ncRNAs, thereby preventing them from learning more comprehensive and in-depth feature representations. RESULTS: To tackle these problems, we proposed MM-ncRNAFP, a multi-modal contrastive learning framework for ncRNA family prediction. We first used a pre-trained language model to encode the primary sequences of a large mammalian ncRNA dataset. Then, we adopted a contrastive learning framework with an attention mechanism to fuse the secondary structure information obtained by graph neural networks. The MM-ncRNAFP method can effectively fuse multi-modal information. Experimental comparisons with several competitive baselines demonstrated that MM-ncRNAFP can achieve more comprehensive representations of ncRNA features by integrating both sequence and structural information. This integration significantly enhances the performance of ncRNA family prediction. Ablation experiments and qualitative analyses were performed to verify the effectiveness of each component in our model. Moreover, since our model is pre-trained on a large amount of ncRNA data, it has the potential to bring significant improvements to other ncRNA-related tasks. AVAILABILITY AND IMPLEMENTATION: MM-ncRNAFP and the datasets are available at https://github.com/xuruiting2/MM-ncRNAFP. Ruiting Xu, Guohua Wang 0001, Yang Li 0130 |
Bioinform. | 4 |
| 2024 | scDRMAE: integrating masked autoencoder with residual attention networks to leverage omics feature dependencies for accurate cell clusteringabstractMOTIVATION: Cell clustering is foundational for analyzing the heterogeneity of biological tissues using single-cell sequencing data. With the maturation of single-cell multi-omics sequencing technologies, we can integrate multiple omics data to perform cell clustering, thereby overcoming the limitations of insufficient information from single omics data. Existing methods for cell clustering often only consider the differences in data patterns during the analysis of multi-omics data, but the dependencies between omics features of different cell types also significantly influence cell clustering. Moreover, the high dropout rates in scRNA-seq and scATAC-seq data can impact the performance of cell clustering. RESULTS: We propose a cell clustering model based on a masked autoencoder, scDRMAE. Utilizing a masking mechanism, scDRMAE effectively learns the relationships between different features and imputes false zeros caused by dropout events. To differentiate the importance of various omics data in cell clustering, we dynamically adjust the weights of different omics data through an attention mechanism. Finally, we use the K-means algorithm for cluster analysis of the fused multi-omics data. On commonly used sets of 15 multi-omics datasets, our method demonstrates superior cell clustering performance on multiple metrics compared to other computational methods. In addition, when datasets exhibit varying degrees of dropout noise, our method shows better performance and stronger stability on multiple metrics compared to other methods. Moreover, by analyzing the cell clusters classified by scDRMAE, we identified several biologically significant biomarkers that have been validated, further confirming the effectiveness of scDRMAE in cell clustering from a biological perspective. Jixiang Ren, Zhenao Wu, Zhongqian Zhao, Guohua Wang 0001 |
Bioinform. | 6 |
| 2024 | TransGINmer: Identifying viral sequences from metagenomes with self-attention and Graph Isomorphism Network
Zhenyuan Sun, Guohua Wang 0001, Yan Miao |
Future Gener. Comput. Syst. | 3 |
| 2024 | Boosting knowledge diversity, accuracy, and stability via tri-enhanced distillation for domain continual medical image segmentation
Zhanshi Zhu, Xinghua Ma, Wei Wang 0169, Suyu Dong, Kuanquan Wang, Lianming Wu, Gongning Luo, Guohua Wang 0001, Shuo Li 0001 |
Medical Image Anal. | 8 |
| 2024 | Automatically Detecting Anchor Cells and Clustering for scRNA-Seq Data Using scTSNNabstractAdvancing in single-cell RNA sequencing techniques enhances the resolution of cell heterogeneity study. Density-based unsupervised clustering has the potential to detect the representative anchor points and the number of clusters automatically. Meanwhile, discovering the true cell type of scRNA-seq data in the unsupervised scenario is still challenging. To this end, we proposed a tensor shared nearest neighbor anchor clustering for scRNA-seq data, named scTSNN, which first makes use of the tensor affinity learning module to mine the local-global balanced topological structures among cells, next designs density-based shared nearest neighbor measurement method to automatically detect anchor cells, finally partitions the non-anchor cells to obtain the clustering results. Validated on synthetic datasets and scRNA-seq datasets, scTSNN not only exactly detects the complicated structures but also has better performance in accuracy and robustness compared with the state-of-the-art methods. Moreover, case studies on mammalian cells and cervical cancer tumor cells demonstrate the selected anchor cells of scTSNN benefit the cell pseudotime inference and rare cell identification, which show good application and research value of scTSNN. Qiaoming Liu, Dong Wang 0066, Guohua Wang 0001, Yadong Wang 0001 |
IEEE J. Biomed. Health Informatics | 4 |
| 2023 | Co-clustering of single-cell RNA-seq data based on weighted non-negative matrix tri-factorization combined with consensus clusteringabstractSingle-cell RNA sequencing (scRNA-seq) has the ability to accurately identify cell types contained in tissues at single-cell resolution. However, the research of identifying imbalanced cell types using scRNA-seq data is still challenging. In this article, we propose scCO2, a method based on weighted non-negative matrix tri-factorization (NMTF) combined with Kmeans-based consensus clustering for the co-clustering of scRNA-seq data. Compared with several popular methods on six real scRNA-seq data with known cell types, scCO2 can achieve comparable or superior cell clustering performance to selected clustering methods. Through the case study on a real scRNA-seq dataset of the human pancreas, scCO2 could obtain good correspondence between gene clusters and cell clusters. Additionally, scCO2 also shows the ability to identify rare cell types. Moreover, by comparing the gene sets from gene clusters to existing known marker genes, we demonstrate that scCO2 has the potential to identify more underlying cell-type-specific genes, and the weights of genes learned by scCO2 could be used as the indicator of gene importance. Tongtong Ren, Guohua Wang 0001 |
BIBM | 2 |
| 2023 | MCANet: shared-weight-based MultiheadCrossAttention network for drug-target interaction predictionabstractAccurate and effective drug-target interaction (DTI) prediction can greatly shorten the drug development lifecycle and reduce the cost of drug development. In the deep-learning-based paradigm for predicting DTI, robust drug and protein feature representations and their interaction features play a key role in improving the accuracy of DTI prediction. Additionally, the class imbalance problem and the overfitting problem in the drug-target dataset can also affect the prediction accuracy, and reducing the consumption of computational resources and speeding up the training process are also critical considerations. In this paper, we propose shared-weight-based MultiheadCrossAttention, a precise and concise attention mechanism that can establish the association between target and drug, making our models more accurate and faster. Then, we use the cross-attention mechanism to construct two models: MCANet and MCANet-B. In MCANet, the cross-attention mechanism is used to extract the interaction features between drugs and proteins for improving the feature representation ability of drugs and proteins, and the PolyLoss loss function is applied to alleviate the overfitting problem and the class imbalance problem in the drug-target dataset. In MCANet-B, the robustness of the model is improved by combining multiple MCANet models and prediction accuracy further increases. We train and evaluate our proposed methods on six public drug-target datasets and achieve state-of-the-art results. In comparison with other baselines, MCANet saves considerable computational resources while maintaining accuracy in the leading position; however, MCANet-B greatly improves prediction accuracy by combining multiple models while maintaining a balance between computational resource consumption and prediction accuracy. Jilong Bian, Xiying Zhang, Dali Xu, Guohua Wang 0001 |
Briefings Bioinform. | 5 |
| 2023 | Unsupervised construction of gene regulatory network based on single-cell multi-omics data of colorectal cancerabstractIdentifying gene regulatory networks (GRNs) at the resolution of single cells has long been a great challenge, and the advent of single-cell multi-omics data provides unprecedented opportunities to construct GRNs. Here, we propose a novel strategy to integrate omics datasets of single-cell ribonucleic acid sequencing and single-cell Assay for Transposase-Accessible Chromatin using sequencing, and using an unsupervised learning neural network to divide the samples with high copy number variation scores, which are used to infer the GRN in each gene block. Accuracy validation of proposed strategy shows that approximately 80% of transcription factors are directly associated with cancer, colorectal cancer, malignancy and disease by TRRUST; and most transcription factors are prone to produce multiple transcript variants and lead to tumorigenesis by RegNetwork database, respectively. The source code access are available at: https://github.com/Cuily-v/Colorectal_cancer. Lingyu Cui, Jilong Bian, Guohua Wang 0001, Yingjian Liang |
Briefings Bioinform. | 4 |
| 2023 | End-to-end interpretable disease-gene association predictionabstractIdentifying disease-gene associations is a fundamental and critical biomedical task towards understanding molecular mechanisms, the diagnosis and treatment of diseases. It is time-consuming and expensive to experimentally verify causal links between diseases and genes. Recently, deep learning methods have achieved tremendous success in identifying candidate genes for genetic diseases. The gene prediction problem can be modeled as a link prediction problem based on the features of nodes and edges of the gene-disease graph. However, most existing researches either build homogeneous networks based on one single data source or heterogeneous networks based on multi-source data, and artificially define meta-paths, so as to learn the network representation of diseases and genes. The former cannot make use of abundant multi-source heterogeneous information, while the latter needs domain knowledge and experience when defining meta-paths, and the accuracy of the model largely depends on the definition of meta-paths. To address the aforementioned challenges above bottlenecks, we propose an end-to-end disease-gene association prediction model with parallel graph transformer network (DGP-PGTN), which deeply integrates the heterogeneous information of diseases, genes, ontologies and phenotypes. DGP-PGTN can automatically and comprehensively capture the multiple latent interactions between diseases and genes, discover the causal relationship between them and is fully interpretable at the same time. We conduct comprehensive experiments and show that DGP-PGTN outperforms the state-of-the-art methods significantly on the task of disease-gene association prediction. Furthermore, DGP-PGTN can automatically learn the implicit relationship between diseases and genes without manually defining meta paths. Yang Li 0130, Zihou Guo, Xin Gao 0001, Guohua Wang 0001 |
Briefings Bioinform. | 5 |
| 2023 | miProBERT: identification of microRNA promoters based on the pre-trained model BERTabstractAccurate prediction of promoter regions driving miRNA gene expression has become a major challenge due to the lack of annotation information for pri-miRNA transcripts. This defect hinders our understanding of miRNA-mediated regulatory networks. Some algorithms have been designed during the past decade to detect miRNA promoters. However, these methods rely on biosignal data such as CpG islands and still need to be improved. Here, we propose miProBERT, a BERT-based model for predicting promoters directly from gene sequences without using any structural or biological signals. According to our information, it is the first time a BERT-based model has been employed to identify miRNA promoters. We use the pre-trained model DNABERT, fine-tune the pre-trained model on the gene promoter dataset so that the model includes information about the richer biological properties of promoter sequences in its representation, and then systematically scan the upstream regions of each intergenic miRNA using the fine-tuned model. About, 665 miRNA promoters are found. The innovative use of a random substitution strategy to construct a negative dataset improves the discriminative ability of the model and further reduces the false positive rate (FPR) to as low as 0.0421. On independent datasets, miProBERT outperformed other gene promoter prediction methods. With comparison on 33 experimentally validated miRNA promoter datasets, miProBERT significantly outperformed previously developed miRNA promoter prediction programs with 78.13% precision and 75.76% recall. We further verify the predicted promoter regions by analyzing conservation, CpG content and histone marks. The effectiveness and robustness of miProBERT are highlighted. Xin Wang 0124, Xin Gao 0001, Guohua Wang 0001 |
Briefings Bioinform. | 3 |
| 2023 | DeepICSH: a complex deep learning framework for identifying cell-specific silencers and their strength from the human genomeabstractSilencers are noncoding DNA sequence fragments located on the genome that suppress gene expression. The variation of silencers in specific cells is closely related to gene expression and cancer development. Computational approaches that exclusively rely on DNA sequence information for silencer identification fail to account for the cell specificity of silencers, resulting in diminished accuracy. Despite the discovery of several transcription factors and epigenetic modifications associated with silencers on the genome, there is still no definitive biological signal or combination thereof to fully characterize silencers, posing challenges in selecting suitable biological signals for their identification. Therefore, we propose a sophisticated deep learning framework called DeepICSH, which is based on multiple biological data sources. Specifically, DeepICSH leverages a deep convolutional neural network to automatically capture biologically relevant signal combinations strongly associated with silencers, originating from a diverse array of biological signals. Furthermore, the utilization of attention mechanisms facilitates the scoring and visualization of these signal combinations, whereas the employment of skip connections facilitates the fusion of multilevel sequence features and signal combinations, thereby empowering the accurate identification of silencers within specific cells. Extensive experiments on HepG2 and K562 cell line data sets demonstrate that DeepICSH outperforms state-of-the-art methods in silencer identification. Notably, we introduce for the first time a deep learning framework based on multi-omics data for classifying strong and weak silencers, achieving favorable performance. In conclusion, DeepICSH shows great promise for advancing the study and analysis of silencers in complex diseases. The source code is available at https://github.com/lyli1013/DeepICSH. Hailong Sun 0004, Dali Xu, Guohua Wang 0001 |
Briefings Bioinform. | 5 |
| 2023 | MMCL-CDR: enhancing cancer drug response prediction with multi-omics and morphology images contrastive representation learningabstractMOTIVATION: Cancer is a complex disease that results in a significant number of global fatalities. Treatment strategies can vary among patients, even if they have the same type of cancer. The application of precision medicine in cancer shows promise for treating different types of cancer, reducing healthcare expenses, and improving recovery rates. To achieve personalized cancer treatment, machine learning models have been developed to predict drug responses based on tumor and drug characteristics. However, current studies either focus on constructing homogeneous networks from single data source or heterogeneous networks from multiomics data. While multiomics data have shown potential in predicting drug responses in cancer cell lines, there is still a lack of research that effectively utilizes insights from different modalities. Furthermore, effectively utilizing the multimodal knowledge of cancer cell lines poses a challenge due to the heterogeneity inherent in these modalities. RESULTS: To address these challenges, we introduce MMCL-CDR (Multimodal Contrastive Learning for Cancer Drug Responses), a multimodal approach for cancer drug response prediction that integrates copy number variation, gene expression, morphology images of cell lines, and chemical structure of drugs. The objective of MMCL-CDR is to align cancer cell lines across different data modalities by learning cell line representations from omic and image data, and combined with structural drug representations to enhance the prediction of cancer drug responses (CDR). We have carried out comprehensive experiments and show that our model significantly outperforms other state-of-the-art methods in CDR prediction. The experimental results also prove that the model can learn more accurate cell line representation by integrating multiomics and morphological data from cell lines, thereby improving the accuracy of CDR prediction. In addition, the ablation study and qualitative analysis also confirm the effectiveness of each part of our proposed model. Last but not least, MMCL-CDR opens up a new dimension for cancer drug response prediction through multimodal contrastive learning, pioneering a novel approach that integrates multiomics and multimodal drug and cell line modeling. AVAILABILITY AND IMPLEMENTATION: MMCL-CDR is available at https://github.com/catly/MMCL-CDR. Yang Li 0130, Zihou Guo, Xin Gao 0001, Guohua Wang 0001 |
Bioinform. | 4 |
| 2023 | DeepITEH: a deep learning framework for identifying tissue-specific eRNAs from the human genomeabstractMOTIVATION: Enhancers are vital cis-regulatory elements that regulate gene expression. Enhancer RNAs (eRNAs), a type of long noncoding RNAs, are transcribed from enhancer regions in the genome. The tissue-specific expression of eRNAs is crucial in the regulation of gene expression and cancer development. The methods that identify eRNAs based solely on genomic sequence data have high error rates because they do not account for tissue specificity. Specific histone modifications associated with eRNAs offer valuable information for their identification. However, identification of eRNAs using histone modification data requires the use of both RNA-seq and histone modification data. Unfortunately, many public datasets contain only one of these components, which impedes the accurate identification of eRNAs. RESULTS: We introduce DeepITEH, a deep learning framework that leverages RNA-seq data and histone modification data from multiple samples of the same tissue to enhance the accuracy of identifying eRNAs. Specifically, deepITEH initially categorizes eRNAs into two classes, namely, regularly expressed eRNAs and accidental eRNAs, using histone modification data from multiple samples of the same tissue. Thereafter, it integrates both sequence and histone modification features to identify eRNAs in specific tissues. To evaluate the performance of DeepITEH, we compared it with four existing state-of-the-art enhancer prediction methods, SeqPose, iEnhancer-RD, LSTMAtt, and FRL, on four normal tissues and four cancer tissues. Remarkably, seven of these tissues demonstrated a substantially improved specific eRNA prediction performance with DeepITEH, when compared with other methods. Our findings suggest that DeepITEH can effectively predict potential eRNAs on the human genome, providing insights for studying the eRNA function in cancer. AVAILABILITY AND IMPLEMENTATION: The source code and dataset of DeepITEH have been uploaded to https://github.com/lyli1013/DeepITEH. Hailong Sun 0004, Guohua Wang 0001 |
Bioinform. | 4 |
| 2023 | CanMethdb: a database for genome-wide DNA methylation annotation in cancersabstractMOTIVATION: DNA methylation within gene body and promoters in cancer cells is well documented. An increasing number of studies showed that cytosine-phosphate-guanine (CpG) sites falling within other regulatory elements could also regulate target gene activation, mainly by affecting transcription factors (TFs) binding in human cancers. This led to the urgent need for comprehensively and effectively collecting distinct cis-regulatory elements and TF-binding sites (TFBS) to annotate DNA methylation regulation. RESULTS: We developed a database (CanMethdb, http://meth.liclab.net/CanMethdb/) that focused on the upstream and downstream annotations for CpG-genes in cancers. This included upstream cis-regulatory elements, especially those involving distal regions to genes, and TFBS annotations for the CpGs and downstream functional annotations for the target genes, computed through integrating abundant DNA methylation and gene expression profiles in diverse cancers. Users could inquire CpG-target gene pairs for a cancer type through inputting a genomic region, a CpG, a gene name, or select hypo/hypermethylated CpG sets. The current version of CanMethdb documented a total of 38 986 060 CpG-target gene pairs (with 6 769 130 unique pairs), involving 385 217 CpGs and 18 044 target genes, abundant cis-regulatory elements and TFs for 33 TCGA cancer types. CanMethdb might help biologists perform in-depth studies of target gene regulations based on DNA methylations in cancer. AVAILABILITY AND IMPLEMENTATION: The main program is available at https://github.com/chunquanlipathway/CanMethdb. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jianmei Zhao, Fengcui Qian, Xuecang Li, Zhengmin Yu, Yanyu Li, Yongsan Yang, Qi Pan, Qiuyu Wang, Jian Zhang 0084, Guohua Wang 0001, Chunquan Li 0002 |
Bioinform. | 16 |
| 2023 | MTGDC: A Multi-Scale Tensor Graph Diffusion Clustering for Single-Cell RNA Sequencing DataabstractSingle-cell RNA sequencing (scRNA-seq) is a new technology that focuses on the expression levels for each cell to study cell heterogeneity. Thus, new computational methods matching scRNA-seq are designed to detect cell types among various cell groups. Herein, we propose a Multi-scale Tensor Graph Diffusion Clustering (MTGDC) for single-cell RNA sequencing data. It has the following mechanisms: 1) To mine potential similarity distributions among cells, we design a multi-scale affinity learning method to construct a fully connected graph between cells; 2) For each affinity matrix, we propose an efficient tensor graph diffusion learning framework to learn high-order information among multi-scale affinity matrices. First, the tensor graph is explicitly introduced to measure cell-cell edges with local high-order relationship information. To further preserve more global topology structure information in the tensor graph, MTGDC implicitly considers the propagation of information via a data diffusion process by designing a simple and efficient tensor graph diffusion update algorithm. 3) Finally, we mix together the multi-scale tensor graphs to obtain the fusion high-order affinity matrix and apply it to spectral clustering. Experiments and case studies showed that MTGDC had obvious advantages over the state-of-art algorithms in robustness, accuracy, visualization, and speed. Qiaoming Liu, Dong Wang 0066, Jie Li 0055, Guohua Wang 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2023 | A Clustering Ensemble Method for Cell Type Detection by Multiobjective Particle OptimizationabstractSingle-cell RNA sequencing (scRNA-seq) is a new technology different from previous sequencing methods that measure the average expression level for each gene across a large population of cells. Thus, new computational methods are required to reveal cell types among cell populations. We present a clustering ensemble algorithm using optimized multiobjective particle (CEMP). It is featured with several mechanisms: 1) A multi-subspace projection method for mapping the original data to low-dimensional subspaces is applied in order to detect complex data structure at both gene level and sample level. 2) The basic partition module in different subspaces is utilized to generate clustering solutions. 3) A transforming representation between clusters and particles is used to bridge the gap between the discrete clustering ensemble optimization problem and the continuous multiobjective optimization algorithm. 4) We propose a clustering ensemble optimization. To guide the multiobjective ensemble optimization process, three cluster metrics are embedded into CEMP as objective functions in which the final clustering will be dynamically evaluated. Experiments on 9 real scRNA-seq datasets indicated that CEMP had superior performance over several other clustering algorithms in clustering accuracy and robustness. The case study conducted on mouse neuronal cells identified main cell types and cell subtypes successfully. Qiaoming Liu, Xudong Zhao 0002, Guohua Wang 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2023 | MicroRNA Promoter Identification in Human With a Three-level Prediction MethodabstractThe accurate annotation of miRNA promoters is critical for the mechanistic understanding of miRNA gene regulation. Various computational methods have been developed for the prediction of miRNA promoters solely employing a single classifier. Most of these computational methods extract either sequence features or one-sided signal features, and the accuracy and reliability of predictions need to be improved. To address these issues, we present miPTP, a three-level prediction method that combines SVM, RF, and correlation coefficients. It is capable of identifying miRNA promoters based on both DNA sequence and ChIP-Seq data (RPol II). By sequentially integrating these two types of information sources with the three methods selected, miPTP can identify miRNA promoters with higher accuracy and sensitivity compared to specific existing methods. Finally, the reliability of miPTP is validated by examining the conservation, CpG content, and activating histone marks in the identified miRNA promoters. Xin Wang 0124, Jie Li 0055, Guohua Wang 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2022 | Detection of transcription factors binding to methylated DNA by deep recurrent neural networkabstractTranscription factors (TFs) are proteins specifically involved in gene expression regulation. It is generally accepted in epigenetics that methylated nucleotides could prevent the TFs from binding to DNA fragments. However, recent studies have confirmed that some TFs have capability to interact with methylated DNA fragments to further regulate gene expression. Although biochemical experiments could recognize TFs binding to methylated DNA sequences, these wet experimental methods are time-consuming and expensive. Machine learning methods provide a good choice for quickly identifying these TFs without experimental materials. Thus, this study aims to design a robust predictor to detect methylated DNA-bound TFs. We firstly proposed using tripeptide word vector feature to formulate protein samples. Subsequently, based on recurrent neural network with long short-term memory, a two-step computational model was designed. The first step predictor was utilized to discriminate transcription factors from non-transcription factors. Once proteins were predicted as TFs, the second step predictor was employed to judge whether the TFs can bind to methylated DNA. Through the independent dataset test, the accuracies of the first step and the second step are 86.63% and 73.59%, respectively. In addition, the statistical analysis of the distribution of tripeptides in training samples showed that the position and number of some tripeptides in the sequence could affect the binding of TFs to methylated DNA. Finally, on the basis of our model, a free web server was established based on the proposed model, which can be available at https://bioinfor.nefu.edu.cn/TFPM/. Hao Lin 0001, Guohua Wang 0001 |
Briefings Bioinform. | 5 |
| 2022 | Drug-target interaction predication via multi-channel graph neural networksabstractDrug-target interaction (DTI) is an important step in drug discovery. Although there are many methods for predicting drug targets, these methods have limitations in using discrete or manual feature representations. In recent years, deep learning methods have been used to predict DTIs to improve these defects. However, most of the existing deep learning methods lack the fusion of topological structure and semantic information in DPP representation learning process. Besides, when learning the DPP node representation in the DPP network, the different influences between neighboring nodes are ignored. In this paper, a new model DTI-MGNN based on multi-channel graph convolutional network and graph attention is proposed for DTI prediction. We use two independent graph attention networks to learn the different interactions between nodes for the topology graph and feature graph with different strengths. At the same time, we use a graph convolutional network with shared weight matrices to learn the common information of the two graphs. The DTI-MGNN model combines topological structure and semantic features to improve the representation learning ability of DPPs, and obtain the state-of-the-art results on public datasets. Specifically, DTI-MGNN has achieved a high accuracy in identifying DTIs (the area under the receiver operating characteristic curve is 0.9665). Yang Li 0130, Guanyu Qiao, Guohua Wang 0001 |
Briefings Bioinform. | 4 |
| 2022 | scESI: evolutionary sparse imputation for single-cell transcriptomes from nearest neighbor cellsabstractThe ubiquitous dropout problem in single-cell RNA sequencing technology causes a large amount of data noise in the gene expression profile. For this reason, we propose an evolutionary sparse imputation (ESI) algorithm for single-cell transcriptomes, which constructs a sparse representation model based on gene regulation relationships between cells. To solve this model, we design an optimization framework based on nondominated sorting genetics. This framework takes into account the topological relationship between cells and the variety of gene expression to iteratively search the global optimal solution, thereby learning the Pareto optimal cell-cell affinity matrix. Finally, we use the learned sparse relationship model between cells to improve data quality and reduce data noise. In simulated datasets, scESI performed significantly better than benchmark methods with various metrics. By applying scESI to real scRNA-seq datasets, we discovered scESI can not only further classify the cell types and separate cells in visualization successfully but also improve the performance in reconstructing trajectories differentiation and identifying differentially expressed genes. In addition, scESI successfully recovered the expression trends of marker genes in stem cell differentiation and can discover new cell types and putative pathways regulating biological processes. Qiaoming Liu, Ximei Luo, Jie Li 0055, Guohua Wang 0001 |
Briefings Bioinform. | 4 |
| 2022 | A survey on computational methods in discovering protein inhibitors of SARS-CoV-2abstractThe outbreak of acute respiratory disease in 2019, namely Coronavirus Disease-2019 (COVID-19), has become an unprecedented healthcare crisis. To mitigate the pandemic, there are a lot of collective and multidisciplinary efforts in facilitating the rapid discovery of protein inhibitors or drugs against COVID-19. Although many computational methods to predict protein inhibitors have been developed [ 1- 5], few systematic reviews on these methods have been published. Here, we provide a comprehensive overview of the existing methods to discover potential inhibitors of COVID-19 virus, so-called severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2). First, we briefly categorize and describe computational approaches by the basic algorithms involved in. Then we review the related biological datasets used in such predictions. Furthermore, we emphatically discuss current knowledge on SARS-CoV-2 inhibitors with the latest findings and development of computational methods in uncovering protein inhibitors against COVID-19. Qiaoming Liu, Jun Wan 0002, Guohua Wang 0001 |
Briefings Bioinform. | 3 |
| 2022 | CRISPRCasStack: a stacking strategy-based ensemble learning framework for accurate identification of Cas proteinsabstractCRISPR-Cas system is an adaptive immune system widely found in most bacteria and archaea to defend against exogenous gene invasion. One of the most critical steps in the study of exploring and classifying novel CRISPR-Cas systems and their functional diversity is the identification of Cas proteins in CRISPR-Cas systems. The discovery of novel Cas proteins has also laid the foundation for technologies such as CRISPR-Cas-based gene editing and gene therapy. Currently, accurate and efficient screening of Cas proteins from metagenomic sequences and proteomic sequences remains a challenge. For Cas proteins with low sequence conservation, existing tools for Cas protein identification based on homology cannot guarantee identification accuracy and efficiency. In this paper, we have developed a novel stacking-based ensemble learning framework for Cas protein identification, called CRISPRCasStack. In particular, we applied the SHAP (SHapley Additive exPlanations) method to analyze the features used in CRISPRCasStack. Sufficient experimental validation and independent testing have demonstrated that CRISPRCasStack can address the accuracy deficiencies and inefficiencies of the existing state-of-the-art tools. We also provide a toolkit to accurately identify and analyze potential Cas proteins, Cas operons, CRISPR arrays and CRISPR-Cas locus in prokaryotic sequences. The CRISPRCasStack toolkit is available at https://github.com/yrjia1015/CRISPRCasStack. Yuran Jia, Dali Xu, Guohua Wang 0001 |
Briefings Bioinform. | 6 |
| 2022 | Complex genome assembly based on long-read sequencingabstractHigh-quality genome chromosome-scale sequences provide an important basis for genomics downstream analysis, especially the construction of haplotype-resolved and complete genomes, which plays a key role in genome annotation, mutation detection, evolutionary analysis, gene function research, comparative genomics and other aspects. However, genome-wide short-read sequencing is difficult to produce a complete genome in the face of a complex genome with high duplication and multiple heterozygosity. The emergence of long-read sequencing technology has greatly improved the integrity of complex genome assembly. We review a variety of computational methods for complex genome assembly and describe in detail the theories, innovations and shortcomings of collapsed, semi-collapsed and uncollapsed assemblers based on long reads. Among the three methods, uncollapsed assembly is the most correct and complete way to represent genomes. In addition, genome assembly is closely related to haplotype reconstruction, that is uncollapsed assembly realizes haplotype reconstruction, and haplotype reconstruction promotes uncollapsed assembly. We hope that gapless, telomere-to-telomere and accurate assembly of complex genomes can be truly routinely achieved using only a simple process or a single tool in the future. Yuran Jia, Yanan Wei, Guohua Wang 0001 |
Briefings Bioinform. | 6 |
| 2022 | Ensemble classification based signature discovery for cancer diagnosis in RNA expression profiles across different platformsabstractMolecular signatures have been excessively reported for diagnosis of many cancers during the last 20 years. However, false-positive signatures are always found using statistical methods or machine learning approaches, and that makes subsequent biological experiments fail. Therefore, signature discovery has gradually become a non-mainstream work in bioinformatics. Actually, there are three critical weaknesses that make the identified signature unreliable. First of all, a signature is wrongly thought to be a gene set, each component of which keeps differential expressions between or among sample groups. Second, there may be many false-positive genes expressed differentially found, even if samples derived from cancer or normal group can be separated in one-dimensional space. Third, cross-platform validation results of a discovered signature are always poor. In order to solve these problems, we propose a new feature selection framework based on ensemble classification to discover signatures for cancer diagnosis. Meanwhile, a procedure for data transform among different expression profiles across different platforms is also designed. Signatures are found on simulation and real data representing different carcinomas across different platforms. Besides, false positives are suppressed. The experimental results demonstrate the effectiveness of our method. Xudong Zhao 0002, Tong Liu 0013, Guohua Wang 0001 |
Briefings Bioinform. | 3 |
| 2022 | Ensemble classification based feature selection: a case of identification on plant pentatricopeptide repeat proteinsabstractIn order to identify plant pentatricopeptide repeat (PPR) proteins, a framework of variable selection has been proposed. In fact, it is an effective feature selection strategy that focuses on the performance of classification. Random forest has been used as the classifier with certain variables automatically selected for discrimination between PPR functional and non-functional proteins. However, it is found that samples regarded as PPR functional proteins are wrongly classified in a high rate. In this paper, we plan to improve the framework in order to achieve better classification results. Modifications are made on the framework for better identifying PPR functional proteins. Instead of random forest, a hybrid ensemble classifier is built with its base classifiers derived from six different classification methods. Besides, an incremental strategy and a clustering by search in descending order are alternatively used for feature selection, which can effectively select the most representative variables for identification on PPR proteins. In addition, it can be found that different base classifiers alternately play an important role in the ensemble classifier with feature dimension increasing. The experimental results demonstrate the effectiveness of our improvements. Xudong Zhao 0002, Jingwen Zhai, Tong Liu 0013, Guohua Wang 0001 |
Briefings Bioinform. | 4 |
| 2022 | Supervised graph co-contrastive learning for drug-target interaction predictionabstractMOTIVATION: Identification of Drug-Target Interactions (DTIs) is an essential step in drug discovery and repositioning. DTI prediction based on biological experiments is time-consuming and expensive. In recent years, graph learning-based methods have aroused widespread interest and shown certain advantages on this task, where the DTI prediction is often modeled as a binary classification problem of the nodes composed of drug and protein pairs (DPPs). Nevertheless, in many real applications, labeled data are very limited and expensive to obtain. With only a few thousand labeled data, models could hardly recognize comprehensive patterns of DPP node representations, and are unable to capture enough commonsense knowledge, which is required in DTI prediction. Supervised contrastive learning gives an aligned representation of DPP node representations with the same class label. In embedding space, DPP node representations with the same label are pulled together, and those with different labels are pushed apart. RESULTS: We propose an end-to-end supervised graph co-contrastive learning model for DTI prediction directly from heterogeneous networks. By contrasting the topology structures and semantic features of the drug-protein-pair network, as well as the new selection strategy of positive and negative samples, SGCL-DTI generates a contrastive loss to guide the model optimization in a supervised manner. Comprehensive experiments on three public datasets demonstrate that our model outperforms the SOTA methods significantly on the task of DTI prediction, especially in the case of cold start. Furthermore, SGCL-DTI provides a new research perspective of contrastive learning for DTI prediction. AVAILABILITY AND IMPLEMENTATION: The research shows that this method has certain applicability in the discovery of drugs, the identification of drug-target pairs and so on. Yang Li 0130, Guanyu Qiao, Xin Gao 0001, Guohua Wang 0001 |
Bioinform. | 4 |
| 2022 | StackCirRNAPred: computational classification of long circRNA from other lncRNA based on stacking strategyabstractBACKGROUND: CircRNAs are essential for the regulation of post-transcriptional gene expression, including as miRNA sponges, and play an important role in disease development. Some computational tools have been proposed recently to predict circRNA, since only one classifier is used, there is still much that can be done to improve the performance. RESULTS: StackCirRNAPred was proposed, the computational classification of long circRNA from other lncRNA based on stacking strategy. In order to cope with the potential problem that a single feature might not be able to distinguish circRNA well from other lncRNA, we first extracted features from different sources, including nucleic acid composition, sequence spatial features and physicochemical properties, Alu and tandem repeats. We innovatively apply the stacking strategy to integrate the more advantageous classifiers of RF, LightGBM, XGBoost. This allows the model to incorporate these features more flexibly. StackCirRNAPred was found to be significantly better than other tools, with precision, accuracy, F1, recall and MCC of 0.843, 0.833, 0.831, 0.819 and 0.666 respectively. We tested it directly on the mouse dataset. StackCirRNAPred was still significantly better than other methods, with precision, accuracy, F1, recall and MCC of 0.837, 0.839, 0.839, 0.841, 0.677. CONCLUSIONS: We proposed StackCirRNAPred based on stacking strategy to distinguish long circRNAs from other lncRNAs. With the test results demonstrating the validity and robustness of StackCirRNAPred, we hope StackCirRNAPred will complement existing circRNA prediction methods and is helpful in down-stream research. Xin Wang 0124, Yadong Liu 0001, Jie Li 0055, Guohua Wang 0001 |
BMC Bioinform. | 4 |
| 2021 | AlignGraph2: similar genome-assisted reassembly pipeline for PacBio long readsabstractContigs assembled from the third-generation sequencing long reads are usually more complete than the second-generation short reads. However, the current algorithms still have difficulty in assembling the long reads into the ideal complete and accurate genome, or the theoretical best result [1]. To improve the long read contigs and with more and more fully sequenced genomes available, it could still be possible to use the similar genome-assisted reassembly method [2], which was initially proposed for the short reads making use of a closely related genome (similar genome) to the sequencing genome (target genome). The method aligns the contigs and reads to the similar genome, and then extends and refines the aligned contigs with the aligned reads. Here, we introduce AlignGraph2, a similar genome-assisted reassembly pipeline for the PacBio long reads. The AlignGraph2 pipeline is the second version of AlignGraph algorithm proposed by us but completely redesigned, can be inputted with either error-prone or HiFi long reads, and contains four novel algorithms: similarity-aware alignment algorithm and alignment filtration algorithm for alignment of the long reads and preassembled contigs to the similar genome, and reassembly algorithm and weight-adjusted consensus algorithm for extension and refinement of the preassembled contigs. In our performance tests on both error-prone and HiFi long reads, AlignGraph2 can align 5.7-27.2% more long reads and 7.3-56.0% more bases than some current alignment algorithm and is more efficient or comparable to the others. For contigs assembled with various de novo algorithms and aligned to similar genomes (aligned contigs), AlignGraph2 can extend 8.7-94.7% of them (extendable contigs), and obtain contigs of 7.0-249.6% larger N50 value and 5.2-87.7% smaller number of indels per 100 kbp (extended contigs). With genomes of decreased similarities, AlignGraph2 also has relatively stable performance. The AlignGraph2 software can be downloaded for free from this site: https://github.com/huangs001/AlignGraph2. Shien Huang, Guohua Wang 0001, Ergude Bao |
Briefings Bioinform. | 3 |
| 2021 | A deep learning approach for filtering structural variants in short read sequencing dataabstractShort read whole genome sequencing has become widely used to detect structural variants in human genetic studies and clinical practices. However, accurate detection of structural variants is a challenging task. Especially existing structural variant detection approaches produce a large proportion of incorrect calls, so effective structural variant filtering approaches are urgently needed. In this study, we propose a novel deep learning-based approach, DeepSVFilter, for filtering structural variants in short read whole genome sequencing data. DeepSVFilter encodes structural variant signals in the read alignments as images and adopts the transfer learning with pre-trained convolutional neural networks as the classification models, which are trained on the well-characterized samples with known high confidence structural variants. We use two well-characterized samples to demonstrate DeepSVFilter's performance and its filtering effect coupled with commonly used structural variant detection approaches. The software DeepSVFilter is implemented using Python and freely available from the website at https://github.com/yongzhuang/DeepSVFilter. Yongzhuang Liu, Yalin Huang, Guohua Wang 0001, Yadong Wang 0001 |
Briefings Bioinform. | 3 |
| 2021 | CHTKC: a robust and efficient k-mer counting algorithm based on a lock-free chaining hash tableabstractMOTIVATION: Calculating the frequency of occurrence of each substring of length k in DNA sequences is a common task in many bioinformatics applications, including genome assembly, error correction, and sequence alignment. Although the problem is simple, efficient counting of datasets with high sequencing depth or large genome size is a challenge. RESULTS: We propose a robust and efficient method, CHTKC, to solve the k-mer counting problem with a lock-free hash table that uses linked lists to resolve collisions. We also design new mechanisms to optimize memory usage and handle situations where memory is not enough to accommodate all k-mers. CHTKC has been thoroughly tested on seven datasets under multiple memory usage scenarios and compared with Jellyfish2 and KMC3. Our work shows that using a hash-table-based method to effectively solve the k-mer counting problem remains a feasible solution. Guohua Wang 0001 |
Briefings Bioinform. | 4 |
| 2021 | The stacking strategy-based hybrid framework for identifying non-coding RNAsabstractWith the development of next-generation sequencing technology, a large number of transcripts need to be analyzed, and it has been a challenge to distinguish non-coding ribonucleic acid (RNAs) (ncRNAs) from coding RNAs. And for non-model organisms, due to the lack of transcriptional data, many existing methods cannot identify them. Therefore, in addition to using deoxyribonucleic acid-based and RNA-based features, we also proposed a hybrid framework based on the stacking strategy to identify ncRNAs, and we innovatively added eight features based on predicted peptides. The proposed framework was based on stacking two-layer classifier which combined random forest (RF), LightGBM, XGBoost and logistic regression (LR) models. We used this framework to build two types of models. For cross-species ncRNAs identification model, we tested it on six different species: human, mouse, zebrafish, fruit fly, worm and Arabidopsis. Compared with other tools, our model was the best in datasets of Arabidopsis, worm and zebrafish with the accuracy of 98.36%, 99.65% and 94.12%. For performance metrics analysis, the datasets of the six species were considered as a whole set, and the sensitivity, accuracy, precision and F1 values of our model were the best. For the plant-specific ncRNAs identification model, the average values of the six metrics of the two experiments were all greater than 95%, which demonstrated it can be used to identify ncRNAs in plants. The above indicates that the hybrid framework we designed is universal between animals and plants and has significant advantages in the identification of cross-species ncRNAs. Xin Wang 0124, Guohua Wang 0001 |
Briefings Bioinform. | 4 |
| 2021 | Evaluating disease similarity based on gene network reconstruction and representationabstractMOTIVATION: Quantifying the associations between diseases is of great significance in increasing our understanding of disease biology, improving disease diagnosis, re-positioning and developing drugs. Therefore, in recent years, the research of disease similarity has received a lot of attention in the field of bioinformatics. Previous work has shown that the combination of the ontology (such as disease ontology and gene ontology) and disease-gene interactions are worthy to be regarded to elucidate diseases and disease associations. However, most of them are either based on the overlap between disease-related gene sets or distance within the ontology's hierarchy. The diseases in these methods are represented by discrete or sparse feature vectors, which cannot grasp the deep semantic information of diseases. Recently, deep representation learning has been widely studied and gradually applied to various fields of bioinformatics. Based on the hypothesis that disease representation depends on its related gene representations, we propose a disease representation model using two most representative gene resources HumanNet and Gene Ontology to construct a new gene network and learn gene (disease) representations. The similarity between two diseases is computed by the cosine similarity of their corresponding representations. RESULTS: We propose a novel approach to compute disease similarity, which integrates two important factors disease-related genes and gene ontology hierarchy to learn disease representation based on deep representation learning. Under the same experimental settings, the AUC value of our method is 0.8074, which improves the most competitive baseline method by 10.1%. The quantitative and qualitative experimental results show that our model can learn effective disease representations and improve the accuracy of disease similarity computation significantly. AVAILABILITY AND IMPLEMENTATION: The research shows that this method has certain applicability in the prediction of gene-related diseases, the migration of disease treatment methods, drug development and so on. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yang Li 0130, Guohua Wang 0001 |
Bioinform. | 3 |
| 2021 | BP4RNAseq: a babysitter package for retrospective and newly generated RNA-seq data analyses using both alignment-based and alignment-free quantification methodabstractSUMMARY: Processing raw reads of RNA-sequencing (RNA-seq) data, no matter public or newly sequenced data, involves a lot of specialized tools and technical configurations that are often unfamiliar and time-consuming to learn for non-bioinformatics researchers. Here, we develop the R package BP4RNAseq, which integrates the state-of-art tools from both alignment-based and alignment-free quantification workflows. The BP4RNAseq package is a highly automated tool using an optimized pipeline to improve the sensitivity and accuracy of RNA-seq analyses. It can take only two non-technical parameters and output six formatted gene expression quantification at gene and transcript levels. The package applies to both retrospective and newly generated bulk RNA-seq data analyses and is also applicable for single-cell RNA-seq analyses. It, therefore, greatly facilitates the application of RNA-seq. AVAILABILITY AND IMPLEMENTATION: The BP4RNAseq package for R and its documentation are freely available at https://github.com/sunshanwen/BP4RNAseq. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Shanwen Sun, Lei Xu 0047, Quan Zou 0001, Guohua Wang 0001 |
Bioinform. | 4 |
| 2021 | Efficient iterative Hi-C scaffolder based on N-best neighborsabstractBACKGROUND: Efficient and effective genome scaffolding tools are still in high demand for generating reference-quality assemblies. While long read data itself is unlikely to create a chromosome-scale assembly for most eukaryotic species, the inexpensive Hi-C sequencing technology, capable of capturing the chromosomal profile of a genome, is now widely used to complete the task. However, the existing Hi-C based scaffolding tools either require a priori chromosome number as input, or lack the ability to build highly continuous scaffolds. RESULTS: We design and develop a novel Hi-C based scaffolding tool, pin_hic, which takes advantage of contact information from Hi-C reads to construct a scaffolding graph iteratively based on N-best neighbors of contigs. Subsequent to scaffolding, it identifies potential misjoins and breaks them to keep the scaffolding accuracy. Through our tests on three long read based de novo assemblies from three different species, we demonstrate that pin_hic is more efficient than current standard state-of-art tools, and it can generate much more continuous scaffolds, while achieving a higher or comparable accuracy. CONCLUSIONS: Pin_hic is an efficient Hi-C based scaffolding tool, which can be useful for building chromosome-scale assemblies. As many sequencing projects have been launched in the recent years, we believe pin_hic has potential to be applied in these projects and makes a meaningful contribution. Dengfeng Guan, Shane A. McCarthy, Zemin Ning, Guohua Wang 0001, Yadong Wang 0001, Richard Durbin |
BMC Bioinform. | 4 |
| 2021 | Correction to: Efficient iterative Hi-C scaffolder based on N-best neighbors
Dengfeng Guan, Shane A. McCarthy, Zemin Ning, Guohua Wang 0001, Yadong Wang 0001, Richard Durbin |
BMC Bioinform. | 4 |
| 2021 | ReRF-Pred: predicting amyloidogenic regions of proteins based on their pseudo amino acid composition and tripeptide compositionabstractBACKGROUND: Amyloids are insoluble fibrillar aggregates that are highly associated with complex human diseases, such as Alzheimer's disease, Parkinson's disease, and type II diabetes. Recently, many studies reported that some specific regions of amino acid sequences may be responsible for the amyloidosis of proteins. It has become very important for elucidating the mechanism of amyloids that identifying the amyloidogenic regions. Accordingly, several computational methods have been put forward to discover amyloidogenic regions. The majority of these methods predicted amyloidogenic regions based on the physicochemical properties of amino acids. In fact, position, order, and correlation of amino acids may also influence the amyloidosis of proteins, which should be also considered in detecting amyloidogenic regions. RESULTS: To address this problem, we proposed a novel machine-learning approach for predicting amyloidogenic regions, called ReRF-Pred. Firstly, the pseudo amino acid composition (PseAAC) was exploited to characterize physicochemical properties and correlation of amino acids. Secondly, tripeptides composition (TPC) was employed to represent the order and position of amino acids. To improve the distinguishability of TPC, all possible tripeptides were analyzed by the binomial distribution method, and only those which have significantly different distribution between positive and negative samples remained. Finally, all samples were characterized by PseAAC and TPC of their amino acid sequence, and a random forest-based amyloidogenic regions predictor was trained on these samples. It was proved by validation experiments that the feature set consisted of PseAAC and TPC is the most distinguishable one for detecting amyloidosis. Meanwhile, random forest is superior to other concerned classifiers on almost all metrics. To validate the effectiveness of our model, ReRF-Pred is compared with a series of gold-standard methods on two datasets: Pep-251 and Reg33. The results suggested our method has the best overall performance and makes significant improvements in discovering amyloidogenic regions. CONCLUSIONS: The advantages of our method are mainly attributed to that PseAAC and TPC can describe the differences between amyloids and other proteins successfully. The ReRF-Pred server can be accessed at http://106.12.83.135:8080/ReRF-Pred/. Zhixia Teng, Zitong Zhang 0002, Zhen Tian 0004, Yanjuan Li, Guohua Wang 0001 |
BMC Bioinform. | 5 |
| 2020 | AFS-DEA: An automatic feature selection platform for differential expression analysisabstractThe majority of effective information on express profile data remains under-utilized due to the characteristics of high-dimensional features, complex correlations among sample attributes and the lack of skills and resources to analyze and visualize these data. In this paper, we developed an automatic feature selection platform combined with an intuitive and interactive interface for differential expression analysis (abbreviated as AFSDEA), which utilized ensemble learning to automatically select features on expression profile data. Qualitative and quantitative analyses indicating the interpretable and predictive ability of the selected features were provided. Correspondingly, experimental results on simulated and real data proved the effectiveness of the developed platform. The establishment of this platform generates a custom Web browser for automatically selecting features and provides outputs for differential expression analysis, which will facilitate the research on expression profile data. The webserver of AFS-DEA is available at http://bio-nefu.com/afs-dea. Xudong Zhao 0002, Weiqi Su, Hangyu Li 0004, Tong Liu 0013, Denan Kong, Guohua Wang 0001 |
BIBM | 6 |
| 2020 | BYASE: a Python library for estimating gene and isoform level allele-specific expressionabstractSUMMARY: Allele-specific expression (ASE) is involved in many important biological mechanisms. We present a python package BYASE and its graphical user interface (GUI) tool BYASE-GUI for the identification of ASE from single-end and paired-end RNA-seq data based on Bayesian inference, which can simultaneously report differences in gene-level and isoform-level expression. BYASE uses both phased SNPs and non-phased SNPs, and supports polyploid organisms. AVAILABILITY AND IMPLEMENTATION: The source codes of BYASE and BYASE-GUI are freely available at https://github.com/ncjllld/byase and https://github.com/ncjllld/byase_gui. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Guohua Wang 0001 |
Bioinform. | 3 |
| 2020 | Evaluating individual genome similarity with a topic modelabstractMOTIVATION: Evaluating genome similarity among individuals is an essential step in data analysis. Advanced sequencing technology detects more and rarer variants for massive individual genomes, thus enabling individual-level genome similarity evaluation. However, the current methodologies, such as the principal component analysis (PCA), lack the capability to fully leverage rare variants and are also difficult to interpret in terms of population genetics. RESULTS: Here, we introduce a probabilistic topic model, latent Dirichlet allocation, to evaluate individual genome similarity. A total of 2535 individuals from the 1000 Genomes Project (KGP) were used to demonstrate our method. Various aspects of variant choice and model parameter selection were studied. We found that relatively rare (0.001 20 000 bp) variants are more efficient for genome similarity evaluation. At least 100 000 such variants are necessary. In our results, the populations show significantly less mixed and more cohesive visualization than the PCA results. The global similarities among the KGP genomes are consistent with known geographical, historical and cultural factors. AVAILABILITY AND IMPLEMENTATION: The source code and data access are available at: https://github.com/lrjuan/LDA_genome. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Liran Juan, Yongtian Wang, Jingyi Jiang, Guohua Wang 0001, Yadong Wang 0001 |
Bioinform. | 5 |
| 2020 | ECFS-DEA: an ensemble classifier-based feature selection for differential expression analysis on expression profilesabstractBACKGROUND: Various methods for differential expression analysis have been widely used to identify features which best distinguish between different categories of samples. Multiple hypothesis testing may leave out explanatory features, each of which may be composed of individually insignificant variables. Multivariate hypothesis testing holds a non-mainstream position, considering the large computation overhead of large-scale matrix operation. Random forest provides a classification strategy for calculation of variable importance. However, it may be unsuitable for different distributions of samples. RESULTS: Based on the thought of using an ensemble classifier, we develop a feature selection tool for differential expression analysis on expression profiles (i.e., ECFS-DEA for short). Considering the differences in sample distribution, a graphical user interface is designed to allow the selection of different base classifiers. Inspired by random forest, a common measure which is applicable to any base classifier is proposed for calculation of variable importance. After an interactive selection of a feature on sorted individual variables, a projection heatmap is presented using k-means clustering. ROC curve is also provided, both of which can intuitively demonstrate the effectiveness of the selected feature. CONCLUSIONS: Feature selection through ensemble classifiers helps to select important variables and thus is applicable for different sample distributions. Experiments on simulation and realistic data demonstrate the effectiveness of ECFS-DEA for differential expression analysis on expression profiles. The software is available at http://bio-nefu.com/resource/ecfs-dea. Xudong Zhao 0002, Qing Jiao, Hangyu Li 0004, Hanxu Wang, Guohua Wang 0001 |
BMC Bioinform. | 7 |
| 2017 | Assessing the model transferability for prediction of transcription factor binding sites based on chromatin accessibilityabstractBACKGROUND: Computational prediction of transcription factor (TF) binding sites in different cell types is challenging. Recent technology development allows us to determine the genome-wide chromatin accessibility in various cellular and developmental contexts. The chromatin accessibility profiles provide useful information in prediction of TF binding events in various physiological conditions. Furthermore, ChIP-Seq analysis was used to determine genome-wide binding sites for a range of different TFs in multiple cell types. Integration of these two types of genomic information can improve the prediction of TF binding events. RESULTS: We assessed to what extent a model built upon on other TFs and/or other cell types could be used to predict the binding sites of TFs of interest. A random forest model was built using a set of cell type-independent features such as specific sequences recognized by the TFs and evolutionary conservation, as well as cell type-specific features derived from chromatin accessibility data. Our analysis suggested that the models learned from other TFs and/or cell lines performed almost as well as the model learned from the target TF in the cell type of interest. Interestingly, models based on multiple TFs performed better than single-TF models. Finally, we proposed a universal model, BPAC, which was generated using ChIP-Seq data from multiple TFs in various cell types. CONCLUSION: Integrating chromatin accessibility information with sequence information improves prediction of TF binding.The prediction of TF binding is transferable across TFs and/or cell lines suggesting there are a set of universal "rules". A computational tool was developed to predict TF binding sites based on the universal "rules". Cristina Zibetti, Jun Wan 0002, Guohua Wang 0001, Seth Blackshaw |
BMC Bioinform. | 4 |
| 2016 | PrefaceabstractWelcome to the 2016 IEEE International Conference on Bioinformatics and Biomedicine (IEEE BIBM 2016) being held in the Shenzhen, China from December 15–18, 2016. On behalf of the IEEE BIBM 2016 Organizing Team, we would like to thank you for your participation and hope you enjoy the conference. Yadong Wang 0001, Kevin Burrage, Shinichi Morishita, Tianhai Tian, Qinghua Jiang, Jiangning Song, Guohua Wang 0001, Xiaohua Hu 0001 |
BIBM | 8 |
| 2015 | HAlign: Fast multiple similar DNA/RNA sequence alignment based on the centre star strategyabstractAbstract Motivation: Multiple sequence alignment (MSA) is important work, but bottlenecks arise in the massive MSA of homologous DNA or genome sequences. Most of the available state-of-the-art software tools cannot address large-scale datasets, or they run rather slowly. The similarity of homologous DNA sequences is often ignored. Lack of parallelization is still a challenge for MSA research. Results: We developed two software tools to address the DNA MSA problem. The first employed trie trees to accelerate the centre star MSA strategy. The expected time complexity was decreased to linear time from square time. To address large-scale data, parallelism was applied using the hadoop platform. Experiments demonstrated the performance of our proposed methods, including their running time, sum-of-pairs scores and scalability. Moreover, we supplied two massive DNA/RNA MSA datasets for further testing and research. Availability and implementation: The codes, tools and data are accessible free of charge at http://datamining.xmu.edu.cn/software/halign/. Contact: [email protected] or [email protected] Quan Zou 0001, Qinghua Hu, Maozu Guo 0001, Guohua Wang 0001 |
Bioinform. | 4 |
| 2014 | Transcriptional regulation prediction of antiestrogen resistance in breast cancer based on RNA polymerase II binding dataabstractBACKGROUND: Although endocrine therapy impedes estrogen-ER signaling pathway and thus reduces breast cancer mortality, patients remain at continued risk of relapse after tamoxifen or other endocrine therapies. Understanding the mechanisms of endocrine resistance, particularly the role of transcriptional regulation is very important and necessary. METHODS: We propose a two-step workflow based on linear model to investigate the significant differences between MCF7 and OHT cells stimulated by 17β-estradiol (E2) respect to regulatory transcription factors (TFs) and their interactions. We additionally compared predicted regulatory TFs based on RNA polymerase II (PolII) binding quantity data and gene expression data, which were taken from MCF7/MCF7+E2 and OHT/OHT+E2 cell lines following the same analysis workflow. Enrichment analysis concerning diseases and cell functions and regulatory pattern analysis of different motifs of the same TF also were performed. RESULTS: The results showed PolII data could provide more information and predict more recognizably important regulatory TFs. Large differences in TF regulatory mode were found between two cell lines. Through verified through GO annotation, enrichment analysis and related literature regarding these TFs, we found some regulatory TFs such as AP-1, C/EBP, FoxA1, GATA1, Oct-1 and NF-κB, maintained OHT cells through molecular interactions or signaling pathways that were different from the surviving MCF7 cells. From TF regulatory interaction network, we identified E2F, E2F-1 and AP-2 as hub-TFs in MCF7 cells; whereas, in addition to E2F and E2F-1, we identified C/EBP and Oct-1 as hub-TFs in OHT cells. Notably, we found the regulatory patterns of different motifs of the same TF were very different from one another sometimes. CONCLUSIONS: We inferred some regulatory TFs, such as AP-1 and NF-κB, cooperated with ER through both genomic action and non-genomic action. The TFs that were involved in both protein-protein interactions and signaling pathways could be one of the key resistant mechanisms of endocrine therapy and thus also could be new treatment targets for endocrine resistance. Our flexible workflow could be integrated into an existing analytical framework and guide biologists to further determine underlying mechanisms in human diseases. Denan Zhang, Guohua Wang 0001, Yadong Wang 0001 |
BMC Bioinform. | 2 |
| 2010 | Predicting human microRNA-disease associations based on support vector machineabstractThe identification of disease-related microRNAs is vital for understanding the pathogenesis of disease at the molecular level and may lead to the design of specific molecular tools for diagnosis, treatment and prevention. Experimental identification of disease-related microRNAs poses difficulties. Computational prediction of microRNA-disease associations is one of the complementary means. However, one major issue in microRNA studies is the lack of bioinformatics programs to accurately predict microRNA-disease associations. Herein, we present a machine learning-based approach for distinguishing positive microRNA-disease associations from negative microRNA-disease associations. A set of features was extracted for each positive and negative microRNA-disease association, and a support vector machine (SVM) classifier was trained, which achieved the area under the ROC curve of up to 0.8884 in 10-fold cross-validation procedure, indicating that the SVM-based approach described here can be used to predict potential microRNA-disease associations and formulate testable hypotheses to guide future biological experiments. Qinghua Jiang, Guohua Wang 0001, Yadong Wang 0001 |
BIBM | 2 |
| 2008 | Model-based prediction of cis-acting RNA elements regulating tissue-specific alternative splicingabstractHere we describe a model-based approach to predict cis-acting RNA elements which regulate tissue-specific alternative splicing. The model facilitates the identification of cis-acting elements (or CAE) and the estimation of their activities, considering the splicing variants between two different tissues as the combinatorial functions of multiple elements. We implement this model on a set of differentially expressed exons, between heart and liver, derived from Affymetrix GeneChipregHuman Exon 1.0 ST Array sample data. Focusing on hexamers, we select top 15 motifs with greatest cumulative exon inclusion (EIC) scores as the potential cis-acting elements. Eight of the total 15 hexamers are validated based on known exonic splicing regulators (ESRs) and predicted ESRs (PESRs). Permutation test demonstrates that the predicted EIC scores are statistically significant. Based on the prediction, we propose that PTB, hnRNP-B, SRp40, as well as other unknown factors are involved in the tissue-specific alternative splicing between heart and liver. Xin Wang 0008, Guohua Wang 0001, Jeremy Sanford |
BIBE | 3 |