EDBT 2026 Demo / reviewers in the wild / expert
Yuedong Yang
dblp:98/2972 · also Yue-Dong Yang
· DBLP profile ↗
87ranked-venue papers
5as first author
67since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 61 · 2 first-author · 43 since 2021Artificial intelligence and machine learning · 20 · 2 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 1 first-author · 13 since 2021Systems, architecture and hardware · 3 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | De Novo Molecular Generation from Mass Spectra via Many-Body Enhanced DiffusionabstractMolecular structure generation from mass spectrometry is fundamental for understanding cellular metabolism and discovering novel compounds. Although tandem mass spectrometry (MS/MS) enables the high-throughput acquisition of fragment fingerprints, these spectra often reflect higher-order interactions involving the concerted cleavage of multiple atoms and bonds-crucial for resolving complex isomers and non-local fragmentation mechanisms. However, most existing methods adopt atom-centric and pairwise interaction modeling, overlooking higher-order edge interactions and lacking the capacity to systematically capture essential many-body characteristics for structure generation. To overcome these limitations, we present MBGen, a Many-Body enhanced diffusion framework for de novo molecular structure Generation from mass spectra. By integrating a many-body attention mechanism and higher-order edge modeling, MBGen comprehensively leverages the rich structural information encoded in MS/MS spectra, enabling accurate de novo generation and isomer differentiation for novel molecules. Experimental results on the NPLIB1 and MassSpecGym benchmarks demonstrate that MBGen achieves superior performance, with improvements of up to 230% over state-of-the-art methods, highlighting the scientific value and practical utility of many-body modeling for mass spectrometry-based molecular generation. Further analysis and ablation studies show that our approach effectively captures higher-order interactions and exhibits enhanced sensitivity to complex isomeric and non-local fragmentation information. Xichen Sun, Jiahua Rao, Jiancong Xie, Yuedong Yang |
AAAI | 5 |
| 2026 | Informative Subgraph Extraction with Deep Reinforcement Learning for Drug-Drug Interaction PredictionabstractDrug-drug interaction (DDI) prediction is pivotal for drug safety and clinical decision-making. Recently, subgraph-based methods utilizing knowledge graphs (KGs) and domain information have achieved promising results by extracting informative subgraphs for DDI prediction. However, existing subgraph extraction methods are typically coarse-grained and nonspecific, facing two key limitations: First, they are constrained by the vast and noisy nature of real-world KGs, making it challenging to identify the most informative substructures from the massive space of candidate subgraphs. Second, current methods often fail to exploit the molecular structural specificity of drugs to selectively extract relevant subgraphs, lacking effective integration of molecular structure information with knowledge graph context. To address these challenges, we propose RISE-DDI, a novel framework for Reinforced-based Informative Subgraph Extraction approach for drug-drug interaction prediction. Specifically, RISE-DDI formulates the subgraph extraction as a Markov Decision Process (MDP) and leverages a deep reinforcement learning (RL) agent to dynamically and adaptively extract the most informative and context-specific subgraphs for each drug pair. The agent is guided by a learnable structure-aware reward model that considers both the topological context from the knowledge graph and the molecular features of the drug pairs, thereby encouraging the selection of subgraphs that are both structurally relevant and biologically informative. Extensive experiments on DDI benchmark datasets demonstrate that our method outperforms state-of-the-art baselines in both transductive and inductive scenarios, achieving improvements of up to 20%. Furthermore, visualization analyses of the extracted subgraphs highlight the interpretability of our model, providing insights into the underlying mechanisms of drug interactions. Jiancong Xie, Jiahua Rao, Yuedong Yang |
AAAI | 5 |
| 2026 | Advancing Protein Design via Multi-Agent Reinforcement Learning with Pareto-Based Collaborative OptimizationabstractProtein design is revolutionizing biotechnology, yet existing approaches struggle to balance structural foldability with functional performance. Structure-based models excel at generating stable protein backbones but often overlook critical functional properties, while protein language models capture evolutionary and functional signals but frequently predict sequences lacking structural stability. Integrating these complementary approaches remains challenging due to their inherently conflicting objectives. We present MAProt, a multi-agent framework that synergistically combines structure-based and protein language model-based methods for protein design. Each agent specializes in a distinct aspect of the design objective: the structure-based agent (e.g., ProteinMPNN) ensures compatibility with the target backbone, while protein language model-based agents (e.g., ESM, SaProt) capture evolutionary plausibility and functional potential. To reconcile conflicts and achieve optimal trade-offs, we introduce a Pareto-based negotiation module that enables effective multi-objective coordination and consensus among agents. Extensive experiments on benchmark datasets demonstrate that MAProt achieves a remarkable improvement over state-of-the-art baselines, and generalizes robustly across a range of tasks, including thermodynamic folding stability design, functional protein design, and high-affinity antibody design. These results highlight the power of collaborative optimization for advancing rational protein engineering. Mingming Zhu, Jiahua Rao, Qianmu Yuan, Yuedong Yang |
AAAI | 5 |
| 2025 | Advancing Retrosynthesis with Retrieval-Augmented Graph GenerationabstractDiffusion-based molecular graph generative models have achieved significant success in template-free, single-step retrosynthesis prediction. However, these models typically generate reactants from scratch, often overlooking the fact that the scaffold of a product molecule typically remains unchanged during chemical reactions. To leverage this useful observation, we introduce a retrieval-augmented molecular graph generation framework. Our framework comprises three key components: a retrieval component that identifies similar molecules for the given product, an integration component that learns valuable clues from these molecules about which part of the product should remain unchanged, and a base generative model that is prompted by these clues to generate the corresponding reactants. We explore various design choices for critical and under-explored aspects of this framework and instantiate it as the Retrieval-Augmented RetroBridge (RARB). RARB demonstrates state-of-the-art performance on standard benchmarks, achieving a 14.8% relative improvement in top-1 accuracy over its base generative model, highlighting the effectiveness of retrieval augmentation. Additionally, RARB excels in handling out-of-distribution molecules, and its advantages remain significant even with smaller models or fewer denoising steps. These strengths make RARB highly valuable for real-world retrosynthesis applications, where extrapolation to novel molecules and high-throughput prediction are essential. Anjie Qiao, Zhen Wang 0036, Jiahua Rao, Yuedong Yang, Zhewei Wei |
AAAI | 4 |
| 2025 | Cellular-Resolution Reconstruction of Spatial Transcriptomics from Histology Images via Foundation Models and KANsabstractSpatial transcriptomics (ST) enables spatially resolved profiling of gene expression within tissues and is commonly accompanied by paired histology images, such as high-resolution H&E-stained sections. Given that histological morphology often correlates with underlying gene expression patterns, the availability of paired images offers a promising opportunity for enhancing spatial resolution. Here we present Hist2Sr, a novel framework that reconstructs cellular-resolution spatial transcriptomics from histology images. Hist2SR first extracts fine-grained morphological representations for individual cells using a foundation model pre-trained on large-scale pathological image datasets. To bridge the scale gap between cell-level features and spot-level gene expression measurements, we adopt a multiple instance learning (MIL) strategy: cell-wise gene expressions within each spot region are aggregated and supervised using the corresponding spot-level transcriptomic profiles. This allows the model to learn informative cell-level features that align with spatial gene expression signals. Finally, a Kolmogorov-Arnold Network (KAN)-based predictor is employed to capture complex nonlinear mappings from cell morphology to gene expression. Benchmark experiments indicate that Hist2SR achieves state-of-the-art performance across various human tissue types. Ruipeng Huang, Leiming Fang, Wenbing Li, Yuansong Zeng, Yuedong Yang |
BIBM | 5 |
| 2025 | Drug Sensitivity Inference Through Foundation Model-Driven Contrastive Integration of Bulk and Single-Cell OmicsabstractTumor heterogeneity hinders drug response prediction: bulk RNA-seq obscures cell-level resistance, while scRNAseq lacks annotated drug response data. Existing methods transferring knowledge from bulk to single-cell often underuse pretrained models and ignore inter-cell relationships, limiting heterogeneity modeling. We propose COIN (Contrastive Learning-driven Omics Integration Network), integrating labeled bulk RNA-seq with unlabeled scRNA-seq for single-cell drug sensitivity prediction. COIN leverages CellFM, a foundation model pretrained on 100 M cells, and uses a shared feature extractor with contrastive learning to align shared patterns and capture micro-heterogeneity, while an auxiliary reconstruction loss ensures robust representations. Trained solely on bulk data, COIN predicts single-cell sensitivity and outperforms existing methods, with 9% average AUC and 7% AUPR improvements. COIN effectively overcomes the limitations of bulk data for heterogeneity modeling and eliminates the need for costly singlecell drug screening annotations, advancing personalized cancer therapy. Ningyuan Shangguan, Yuansong Zeng, Wenbing Li, Yuedong Yang |
BIBM | 4 |
| 2025 | Interpretably Predicting Chemical Perturbation via Biologically Informed Visible Neural NetworkabstractRecent advances in single-cell transcriptomics have enabled high-resolution characterization of cellular responses to drug perturbations, a critical capability for precision medicine and drug discovery. Deep learning has emerged as a powerful tool for accurate and scalable prediction of single-cell responses to drug perturbations. However, most existing approaches fail to consider gene-level structural priors and biological knowledge. Furthermore, they often operate as black boxes with limited interpretability. To offer reliable drug response prediction in real-world applications, there is an urgent need to develop a model that combines high predictive accuracy with strong interpretability. In this work, we propose a novel framework based on the Visible Neural Network (VNN) to predict cellular responses to drug perturbations in single-cell transcriptomic profiles. In contrast to traditional black-box models, VNN incorporates prior biological knowledge into its architecture. This integration enables interpretable predictions while maintaining biological coherence. We further integrate deep learning models with a knowledge graph of gene-gene relationships to improve the understanding and generalization of transcriptomic responses to unseen drug perturbations by leveraging prior biological knowledge of gene interactions. Our analysis reveals that the hierarchical structure of the VNN aligns well with known biological processes and supports robust feature attribution. In general, our findings underscore the potential of visible architectures to bridge deep learning and systems biology for a scalable and interpretable modeling of cellular drug responses. Jiancong Xie, Yuedong Yang |
BIBM | 3 |
| 2025 | Multi-modal Contrastive Learning with Negative Sampling Calibration for Phenotypic Drug DiscoveryabstractPhenotypic drug discovery presents a promising strategy for identifying first-in-class drugs by bypassing the need for specific drug targets. Recent advances in cell-based phenotypic screening tools, including Cell Painting and the LINCS L1000, provide essential cellular data that capture biological responses to compounds. While the integration of the multi-modal data enhances the use of contrastive learning (CL) methods for molecular phenotypic representation, these approaches treat all negative pairs equally, failing to discriminate molecules with similar phenotypes. To address these challenges, we introduce a foundational framework MINER that dynamically estimates the likelihoods of sample pairs as negative pairs based on uni-modal disentangled representations. In addition, our approach incorporates a mixture fusion strategy to effectively integrate multimodal data, even in cases where certain modalities are missing. Extensive experiments demonstrate that our method enhances both molecular property prediction and molecule-phenotype retrieval accuracy. Moreover, it successfully recommends drug candidates from phenotype for complex diseases documented in the literature. These findings underscore MINER’s potential to advance drug discovery by enabling deeper insights into disease mechanisms and improving drug candidate recommendations. Jiahua Rao, Hanjing Lin, Leyu Chen, Jiancong Xie, Shuangjia Zheng, Yuedong Yang |
CVPR | 6 |
| 2025 | Quadruple Attention in Many-body Systems for Accurate Molecular Property PredictionsabstractWhile Graph Neural Networks and Transformers have shown promise in predicting molecular properties, they struggle with directly modeling complex many-body interactions. Current methods often approximate interactions like three- and four-body terms in message passing, while attention-based models, despite enabling direct atom communication, are typically limited to triplets, making higher-order interactions computationally demanding. To address the limitations, we introduce MABNet, a geometric attention framework designed to model four-body interactions by facilitating direct communication among atomic quartets. This approach bypasses the computational bottlenecks associated with traditional triplet-based attention mechanisms, allowing for the efficient handling of higher-order interactions. MABNet achieves state-of-the-art performance on benchmarks like MD22 and SPICE. These improvements underscore its capability to accurately capture intricate many-body interactions in large molecules. By unifying rigorous many-body physics with computational efficiency, MABNet advances molecular simulations for applications in drug design and materials discovery, while its extensible framework paves the way for modeling higher-order quantum effects. Jiahua Rao, Dahao Xu, Yicong Chen, Mingjun Yang, Yuedong Yang |
ICML | 6 |
| 2025 | Incorporating Retrieval-based Causal Learning with Information Bottlenecks for Interpretable Molecular Graph LearningabstractGraph Neural Networks (GNNs) have gained considerable traction for modeling molecular structures and predicting properties, but their interpretability remains a significant challenge in understanding chemical behaviors. Current interpretation methods often rely on post-hoc explanations, which aim to provide transparency in GNN decisions. However, these approaches struggle with interpreting complex subgraphs and fail to leverage explanations to enhance predictive capabilities. While transparent methods can enhance GNN predictions, they typically compromise on explanation precision. This limitation underscores the need for a new strategy that effectively integrates GNN explanations and predictions. In this study, we have developed a novel interpretable causal GNN framework that combines retrieval-based causal learning with Graph Information Bottleneck (GIB) theory. Our framework semi-parametrically identifies crucial subgraphs through GIB and compresses explanatory subgraphs using a causal module. The framework consistently outperformed state-of-the-art methods, achieving a 32.72% increase in precision for scientific explanation tasks involving diverse substructures. More importantly, the learned explanations were also shown to be able to improve GNN prediction performance. This advancement is particularly vital for molecular graph learning, as it addresses the critical need to interpret how molecular structures influence predicted properties, thereby aiding drug discovery and materials science by providing insights into chemical mechanisms. Jiahua Rao, Hanjing Lin, Jiancong Xie, Zhen Wang 0036, Shuangjia Zheng, Yuedong Yang |
KDD (2) | 6 |
| 2025 | Reinforced Active Learning for Large-Scale Virtual Screening with Learnable Policy ModelabstractVirtual Screening (VS) is vital for drug discovery but struggles with low hit rates and high computational costs. While Active Learning (AL) has shown promise in improving the efficiency of VS, traditional methods rely on inflexible and handcrafted heuristics, limiting adaptability in complex chemical spaces, particularly in balancing molecular diversity and selection accuracy.
To overcome these challenges, we propose GLARE, a reinforced active learning framework that reformulates VS as a Markov Decision Process (MDP). Using Group Relative Policy Optimization (GRPO), GLARE dynamically balances chemical diversity, biological relevance, and computational constraints, eliminating the need for inflexible heuristics.
Experiments show GLARE outperforms state-of-the-art AL methods, with a 64.8% average improvement in Enrichment Factors (EF). Additionally, GLARE enhances the performance of VS foundation models like DrugCLIP, achieving up to an 8-fold improvement in EF$_{0.5\\%}$ with as few as 15 active molecules. These results highlight the transformative potential of GLARE for adaptive and efficient drug discovery. Yicong Chen, Jiahua Rao, Jiancong Xie, Dahao Xu, Zhen Wang 0004, Yuedong Yang |
NeurIPS | 6 |
| 2025 | Accurately Predicting Protein Mutational Effects via a Hierarchical Many-Body Attention NetworkabstractPredicting changes in binding free energy ($\Delta\Delta G$) is essential for understanding protein-protein interactions, which are critical in drug design and protein engineering. However, existing methods often rely on pre-trained knowledge and heuristic features, limiting their ability to accurately model complex mutation effects, particularly higher-order and many-body interactions.
To address these challenges, we propose H3-DDG, a Hypergraph-driven Hierarchical network to capture Higher-order many-body interactions across multiple scales. By introducing a hierarchical communication mechanism, H3-DDG effectively models both local and global mutational effects.
Experimental results demonstrate state-of-the-art performance on multiple benchmarks. On the SKEMPI v2 dataset, H3-DDG achieves a Pearson correlation of 0.75, improving multi-point mutations prediction by 12.10%. On the challenging BindingGYM dataset, it outperforms Prompt-DDG and BA-DDG by 62.61% and 34.26%, respectively.
Ablation and efficiency analyses demonstrate its robustness and scalability, while a case study on SARS-CoV-2 antibodies highlights its practical value in improving binding affinity for therapeutic design. Dahao Xu, Jiahua Rao, Mingming Zhu, Shuangjia Zheng, Yuedong Yang |
NeurIPS | 7 |
| 2025 | A 3D pocket-aware lead optimization model with knowledge guidance and its application for discovery of new glutaminyl cyclase inhibitorsabstractLead optimization, aimed at improving binding affinity or other properties of hit compounds, is a crucial task in drug discovery. Though deep learning-based 3D generative models showed promise in enhancing the efficiency of de novo drug design recently, less research and attention has garnered for structure-based lead optimization. Herein, we propose a 3D pocket-aware diffusion model named Diffleop, which explicitly incorporates the knowledge of protein-ligand binding affinity and information on covalent bonds to guide the denoising sampling process for lead optimization with enhanced binding affinity and rational properties. Specifically, the bond constraint is achieved through diffusion on fully connected molecular graphs, and the determination of atom positions, atom and bond types in each sampling step is guided by the gradient of the binding affinity that is predicted through fitting with an E(3)-equivariant expert network. The comprehensive evaluations indicated that Diffleop outperforms baseline models on lead optimization with higher affinity and more binding interactions, and can generate more drug-like molecules with more rational structures. Diffleop was further applied to optimize 5-methyl-1H-imidazole, our newly discovered lead compound targeting human glutaminyl cyclases (QCs). Three synthesized compounds exhibit substantially improved inhibitory activities against QCs, with the most effective one showing an IC50 value of 8 nM and 3.5-fold better than clinical candidate PQ912. Anjie Qiao, Weifeng Huang, Hao Zhang 0200, Qirui Deng, Jiahua Rao, Ji Deng, Zhen Wang 0004, Mingyuan Xu, Hongming Chen 0001, Jiancong Xie, Shuangjia Zheng, Yuedong Yang, Guo-Bo Li, Jinping Lei |
Briefings Bioinform. | 15 |
| 2024 | Interpretable Drug Response Prediction through Molecule Structure-aware and Knowledge-Guided Visible Neural NetworkabstractPrecise prediction of anti-cancer drug responses has become a crucial obstruction in anti-cancer drug design and clinical applications. In recent years, various deep learning methods have been applied to drug response prediction and become more accurate. However, they are still criticized as being non-transparent. To offer reliable drug response prediction in real-world applications, there is still a pressing demand to develop a model with high predictive performance as well as interpretability. In this study, we propose DrugVNN, an end-to-end interpretable drug response prediction framework, which extracts gene features of cell lines through a knowledge-guided visible neural network (VNN) and learns drug representation through a node-edge communicative message passing network (CMPNN). Additionally, between these two networks, a novel drug-aware gene attention gate is designed to direct the drug representation to VNN to simulate the effects of drugs. By evaluating on the GDSC dataset, DrugVNN achieved state-of-the-art performance. Moreover, DrugVNN can identify active genes and relevant signaling pathways for specific drug-cell line pairs with supporting evidence in the literature, implying the interpretability of our model. Jiancong Xie, Youyou Li, Jiahua Rao, Yuedong Yang |
BIBM | 5 |
| 2024 | GP-nano: a geometric graph network for nanobody polyreactivity predictionabstractNanobodies are emerging therapeutic antibodies with more simple structure, which can target antigen surfaces and tissue types not accessible to conventional antibodies. However, nanobodies exhibit polyreactivity, binding non-specifically to off-target proteins and other biomolecules. This uncertainty can affect the drug development process and pose significant challenges in clinical development. Existing computational polyreactivity prediction methods focus solely on the sequence or fail to fully utilize structural information. In this study, we propose GP-nano, a geometric graph network based model for nanobody polyreactivity using predictive structure. GP-nano starts from sequences, predicts protein structures using ESMfold, and fully utilizes structural geometric information via graph networks. GP-nano can accurately classify the polyreactivity of nanobodies (AUC=0.91). To demonstrate GP-nano’s generalizability, we also trained and tested it on monoclonal antibodies (mAbs). GP-nano outperforms the best methods on both datasets, indicating the contribution of structural information and geometric features to antibody polyreactivity prediction. Qianmu Yuan, Shuangjia Zheng, Yu Wang 0008, Yuedong Yang |
BIBM | 5 |
| 2024 | Instance-level Expert Knowledge and Aggregate Discriminative Attention for Radiology Report GenerationabstractAutomatic radiology report generation can provide sub-stantial advantages to clinical physicians by effectively re-ducing their workload and improving efficiency. Despite the promising potential of current methods, challenges persist in effectively extracting and preventing degradation of prominent features, as well as enhancing attention on piv-otal regions. In this paper, we propose an Instance-level Expert Knowledge and Aggregate Discriminative Attention framework (EKAGen11https://github.com/hnjzbss/EKAGen) for radiology report generation. We convert expert reports into an embedding space and gener-ate comprehensive representations for each disease, which serve as Preliminary Knowledge Support (PKS). To prevent feature disruption, we select the representations in the em-bedding space with the smallest distances to P KS as Rec-tified Knowledge Support (RKS). Then, EKAGen diagnoses the diseases and retrieves knowledge from RKS, creating Instance-level Expert Knowledge (IEK) for each query image, boosting generation. Additionally, we introduce Ag-gregate Discriminative Attention Map (ADM), which uses weak supervision to create maps of discriminative regions that highlight pivotal regions. For training, we propose a Global Information Self-Distillation (GID) strategy, using an iteratively optimized model to distill global knowledge into EKAGen. Extensive experiments and analyses on IU X-Ray and MIMIC-CXR datasets demonstrate that EKAGen outperforms previous state-of-the-art methods. Shenshen Bu, Taiji Li, Yuedong Yang, Zhiming Dai |
CVPR | 3 |
| 2024 | Accurately Deciphering Novel Cell Type in Spatially Resolved Single-Cell Data Through Optimal Transport
Mai Luo, Yuansong Zeng, Jianing Chen 0004, Ningyuan Shangguan, Yuedong Yang |
ISBRA (2) | 6 |
| 2024 | Comprehensive single-cell RNA-seq analysis using deep interpretable generative modeling guided by biological hierarchy knowledgeabstractRecent advances in microfluidics and sequencing technologies allow researchers to explore cellular heterogeneity at single-cell resolution. In recent years, deep learning frameworks, such as generative models, have brought great changes to the analysis of transcriptomic data. Nevertheless, relying on the potential space of these generative models alone is insufficient to generate biological explanations. In addition, most of the previous work based on generative models is limited to shallow neural networks with one to three layers of latent variables, which may limit the capabilities of the models. Here, we propose a deep interpretable generative model called d-scIGM for single-cell data analysis. d-scIGM combines sawtooth connectivity techniques and residual networks, thereby constructing a deep generative framework. In addition, d-scIGM incorporates hierarchical prior knowledge of biological domains to enhance the interpretability of the model. We show that d-scIGM achieves excellent performance in a variety of fundamental tasks, including clustering, visualization, and pseudo-temporal inference. Through topic pathway studies, we found that d-scIGM-learned topics are better enriched for biologically meaningful pathways compared to the baseline models. Furthermore, the analysis of drug response data shows that d-scIGM can capture drug response patterns in large-scale experiments, which provides a promising way to elucidate the underlying biological mechanisms. Lastly, in the melanoma dataset, d-scIGM accurately identified different cell types and revealed multiple melanin-related driver genes and key pathways, which are critical for understanding disease mechanisms and drug development. Hegang Chen, Yuyin Lu, Zhiming Dai, Yuedong Yang, Qing Li 0001, Yanghui Rao |
Briefings Bioinform. | 4 |
| 2024 | Self-supervised learning on millions of primary RNA sequences from 72 vertebrates improves sequence-based RNA splicing predictionabstractLanguage models pretrained by self-supervised learning (SSL) have been widely utilized to study protein sequences, while few models were developed for genomic sequences and were limited to single species. Due to the lack of genomes from different species, these models cannot effectively leverage evolutionary information. In this study, we have developed SpliceBERT, a language model pretrained on primary ribonucleic acids (RNA) sequences from 72 vertebrates by masked language modeling, and applied it to sequence-based modeling of RNA splicing. Pretraining SpliceBERT on diverse species enables effective identification of evolutionarily conserved elements. Meanwhile, the learned hidden states and attention weights can characterize the biological properties of splice sites. As a result, SpliceBERT was shown effective on several downstream tasks: zero-shot prediction of variant effects on splicing, prediction of branchpoints in humans, and cross-species prediction of splice sites. Our study highlighted the importance of pretraining genomic language models on a diverse range of species and suggested that SSL is a promising approach to enhance our understanding of the regulatory logic underlying genomic sequences. Ken Chen 0006, Yue Zhou 0014, Maolin Ding, Yu Wang 0008, Zhixiang Ren, Yuedong Yang |
Briefings Bioinform. | 6 |
| 2024 | Detecting novel cell type in single-cell chromatin accessibility data via open-set domain adaptationabstractRecent advances in single-cell technologies enable the rapid growth of multi-omics data. Cell type annotation is one common task in analyzing single-cell data. It is a challenge that some cell types in the testing set are not present in the training set (i.e. unknown cell types). Most scATAC-seq cell type annotation methods generally assign each cell in the testing set to one known type in the training set but neglect unknown cell types. Here, we present OVAAnno, an automatic cell types annotation method which utilizes open-set domain adaptation to detect unknown cell types in scATAC-seq data. Comprehensive experiments show that OVAAnno successfully identifies known and unknown cell types. Further experiments demonstrate that OVAAnno also performs well on scRNA-seq data. Our codes are available online at https://github.com/lisaber/OVAAnno/tree/master. Yuefan Lin, Zixiang Pan, Yuansong Zeng, Yuedong Yang, Zhiming Dai |
Briefings Bioinform. | 4 |
| 2024 | From intuition to AI: evolution of small molecule representations in drug discoveryabstractWithin drug discovery, the goal of AI scientists and cheminformaticians is to help identify molecular starting points that will develop into safe and efficacious drugs while reducing costs, time and failure rates. To achieve this goal, it is crucial to represent molecules in a digital format that makes them machine-readable and facilitates the accurate prediction of properties that drive decision-making. Over the years, molecular representations have evolved from intuitive and human-readable formats to bespoke numerical descriptors and fingerprints, and now to learned representations that capture patterns and salient features across vast chemical spaces. Among these, sequence-based and graph-based representations of small molecules have become highly popular. However, each approach has strengths and weaknesses across dimensions such as generality, computational cost, inversibility for generative applications and interpretability, which can be critical in informing practitioners' decisions. As the drug discovery landscape evolves, opportunities for innovation continue to emerge. These include the creation of molecular representations for high-value, low-data regimes, the distillation of broader biological and chemical knowledge into novel learned representations and the modeling of up-and-coming therapeutic modalities. Miles McGibbon, Steven R. Shave, Yumiao Gao, Douglas R. Houston, Jiancong Xie, Yuedong Yang, Philippe Schwaller, Vincent Blay |
Briefings Bioinform. | 7 |
| 2024 | Subgraph extraction and graph representation learning for single cell Hi-C imputation and clusteringabstractSingle-cell Hi-C (scHi-C) technology enables the investigation of 3D chromatin structure variability across individual cells. However, the analysis of scHi-C data is challenged by a large number of missing values. Here, we present a scHi-C data imputation model HiC-SGL, based on Subgraph extraction and graph representation learning. HiC-SGL can also learn informative low-dimensional embeddings of cells. We demonstrate that our method surpasses existing methods in terms of imputation accuracy and clustering performance by various metrics. Jiahao Zheng 0005, Yuedong Yang, Zhiming Dai |
Briefings Bioinform. | 2 |
| 2024 | An uncertainty-based interpretable deep learning framework for predicting breast cancer outcomeabstractBACKGROUND: Predicting outcome of breast cancer is important for selecting appropriate treatments and prolonging the survival periods of patients. Recently, different deep learning-based methods have been carefully designed for cancer outcome prediction. However, the application of these methods is still challenged by interpretability. In this study, we proposed a novel multitask deep neural network called UISNet to predict the outcome of breast cancer. The UISNet is able to interpret the importance of features for the prediction model via an uncertainty-based integrated gradients algorithm. UISNet improved the prediction by introducing prior biological pathway knowledge and utilizing patient heterogeneity information. RESULTS: The model was tested in seven public datasets of breast cancer, and showed better performance (average C-index = 0.691) than the state-of-the-art methods (average C-index = 0.650, ranged from 0.619 to 0.677). Importantly, the UISNet identified 20 genes as associated with breast cancer, among which 11 have been proven to be associated with breast cancer by previous studies, and others are novel findings of this study. CONCLUSIONS: Our proposed method is accurate and robust in predicting breast cancer outcomes, and it is an effective way to identify breast cancer-associated genes. The method codes are available at: https://github.com/chh171/UISNet . Siyin Lin, Junqi Lin, Minfan He, Yuedong Yang, Yongzhong OuYang, Huiying Zhao |
BMC Bioinform. | 5 |
| 2023 | SE(3) Equivalent Graph Attention Network as an Energy-Based Model for Protein Side Chain ConformationabstractProtein design energy functions have been developed over decades by leveraging physical forces approximation and knowledge-derived features. However, manual feature engineering and parameter tuning might suffer from knowledge bias. Learning potential energy functions fully from crystal structure data is promising to automatically discover unknown or highorder features contributing to the protein’s energy. Here we propose a novel data-driven energy-based model based on SE(3)-equivariant model for protein conformation, namely GraphEBM. By combining with the graph attention network, GraphEBM improve the massage passing on the chemical bond and capture the interatomic interaction and overlap. GraphEBM was benchmarked on the local rotamer recovery task and found to outperform both Rosetta and the state-of-the-art deep learning based methods. Furthermore, GraphEBM also yielded promising results on combinatorial side chain optimization, improving 13.8% ${\mathcal{X}_1}$ rotamer recovery to the Atom Transformer method on average. Deqin Liu, Shuangjia Zheng, Yuedong Yang |
BIBM | 5 |
| 2023 | Accurately Identifying Muscle-Invasive Bladder Cancer from MRI via Weakly Supervised LearningabstractBladder cancer (BCa) is one of the most common malignancies in the world, which can be categorized into muscleinvasive (MIBC) and non-muscle-invasive (NMIBC). These two types of BCa must be treated differently, and thus it is essential to correctly distinguish MIBC and NMIBC patients preoperatively for adopting different treatment methods accordingly. Currently, the two types can be distinguished through MRI images by radiologists, but manual inspection is time and labor-consuming. Existing machine learning based methods attempt to free radiologists from manual inspection. However, they fail to take full advantage of image features and always require extra laborious refined manual labeling in addition to the classification labels. In this study, we propose a Tumor Staging and Localization Network (TSLNet) to perform preoperative non-invasive assessment of muscle invasion of BCa, which can automatically distinguish MIBC patients from NMIBC patients based on MRI T2-weighted images of BCa. The model adopts the weakly supervised learning method. Specifically, self-produced guidance is used as pixellevel segmentation pseudo labels for auxiliary supervision to extract basic features, and location-recognition based fine-grained image classification technology and inexact consistency labels are used for auxiliary supervision to extract fine-grained features. Moreover, the model can visualize the critical regions of the lesions, which can provide practical reference and a basis for clinicians’ clinical diagnosis. Experimental results show that the model achieves high AUC, accuracy, specificity, sensitivity, and F1-score, which is comparable to experienced clinicians. Fudan Zheng, Yuedong Yang, Tianxin Lin, Shaoxu Wu, Yutong Lu, Zhiguang Chen 0001, Huiying Zhao |
BIBM | 2 |
| 2023 | Efficient On-Device Training via Gradient FilteringabstractDespite its importance for federated learning, continuous learning and many other applications, on-device training remains an open problem for EdgeAI. The problem stems from the large number of operations (e.g., floating point multiplications and additions) and memory consumption required during training by the back-propagation algorithm. Consequently, in this paper, we propose a new gradient filtering approach which enables on-device CNN model training. More precisely, our approach creates a special structure with fewer unique elements in the gradient map, thus significantly reducing the computational complexity and memory consumption of back propagation during training. Extensive experiments on image classification and semantic segmentation with multiple CNN models (e.g., MobileNet, DeepLabV3, UPerNet) and devices (e.g., Raspberry Pi and Jetson Nano) demonstrate the effectiveness and wide applicability of our approach. For example, compared to SOTA, we achieve up to 19× speedup and 77.1% memory savings on ImageNet classification with only 0.1% accuracy loss. Finally, our method is easy to implement and deploy; over 20× speedup and 90% energy savings have been observed compared to highly optimized baselines in MKLDNN and CUDNN on NVIDIA Jetson Nano. Consequently, our approach opens up a new direction of research with a huge potential for on-device training.11Code: https://github.com/SLDGroup/GradientFilter-CVPR23 Yuedong Yang, Guihong Li, Radu Marculescu |
CVPR | 1 |
| 2023 | ZiCo: Zero-shot NAS via inverse Coefficient of Variation on Gradients
Guihong Li, Yuedong Yang, Kartikeya Bhardwaj, Radu Marculescu |
ICLR | 2 |
| 2023 | TIPS: Topologically Important Path Sampling for Anytime Neural NetworksabstractAnytime neural networks (AnytimeNNs) are a promising solution to adaptively adjust the model complexity at runtime under various hardware resource constraints. However, the manually-designed AnytimeNNs are biased by designers’ prior experience and thus provide sub-optimal solutions. To address the limitations of existing hand-crafted approaches, we first model the training process of AnytimeNNs as a discrete-time Markov chain (DTMC) and use it to identify the paths that contribute the most to the training of AnytimeNNs. Based on this new DTMC-based analysis, we further propose TIPS, a framework to automatically design AnytimeNNs under various hardware constraints. Our experimental results show that TIPS can improve the convergence rate and test accuracy of AnytimeNNs. Compared to the existing AnytimeNNs approaches, TIPS improves the accuracy by 2%-6.6% on multiple datasets and achieves SOTA accuracy-FLOPs tradeoffs. Guihong Li, Kartikeya Bhardwaj, Yuedong Yang, Radu Marculescu |
ICML | 3 |
| 2023 | Retrieval-based Knowledge Augmented Vision Language Pre-trainingabstractWith the recent progress in large-scale vision and language representation learning, Vision Language Pre-training (VLP) models have achieved promising improvements on various multi-modal downstream tasks. Albeit powerful, these models have not fully leveraged world knowledge to their advantage. A key challenge of knowledge-augmented VLP is the lack of clear connections between knowledge and multi-modal data. Moreover, not all knowledge present in images/texts is useful, therefore prior approaches often struggle to effectively integrate knowledge, visual, and textual information. In this study, we propose REtrieval-based knowledge Augmented Vision Language (REAVL), a novel knowledge-augmented pre-training framework to address the above issues. For the first time, we introduce a knowledge-aware self-supervised learning scheme that efficiently establishes the correspondence between knowledge and multi-modal data and identifies informative knowledge to improve the modeling of alignment and interactions between visual and textual modalities. By adaptively integrating informative knowledge with visual and textual information, REAVL achieves new state-of-the-art performance uniformly on knowledge-based vision-language understanding and multi-modal entity linking tasks, as well as competitive results on general vision-language tasks while only using 0.2% pre-training data of the best models. Our model shows strong sample efficiency and effective knowledge utilization. Jiahua Rao, Zifei Shan, Longpo Liu, Yao Zhou 0011, Yuedong Yang |
ACM Multimedia | 5 |
| 2023 | Efficient Low-rank Backpropagation for Vision Transformer AdaptationabstractThe increasing scale of vision transformers (ViT) has made the efficient fine-tuning of these large models for specific needs a significant challenge in various applications. This issue originates from the computationally demanding matrix multiplications required during the backpropagation process through linear layers in ViT.
In this paper, we tackle this problem by proposing a new Low-rank BackPropagation via Walsh-Hadamard Transformation (LBP-WHT) method. Intuitively, LBP-WHT projects the gradient into a low-rank space and carries out backpropagation. This approach substantially reduces the computation needed for adapting ViT, as matrix multiplication in the low-rank space is far less resource-intensive. We conduct extensive experiments with different models (ViT, hybrid convolution-ViT model) on multiple datasets to demonstrate the effectiveness of our method. For instance, when adapting an EfficientFormer-L1 model on CIFAR100, our LBP-WHT achieves 10.4\% higher accuracy than the state-of-the-art baseline, while requiring 9 MFLOPs less computation.
As the first work to accelerate ViT adaptation with low-rank backpropagation, our LBP-WHT method is complementary to many prior efforts and can be combined with them for better performance. Yuedong Yang, Hung-Yueh Chiang, Guihong Li, Diana Marculescu, Radu Marculescu |
NeurIPS | 1 |
| 2023 | Fast and accurate protein intrinsic disorder prediction by using a pretrained language modelabstractDetermining intrinsically disordered regions of proteins is essential for elucidating protein biological functions and the mechanisms of their associated diseases. As the gap between the number of experimentally determined protein structures and the number of protein sequences continues to grow exponentially, there is a need for developing an accurate and computationally efficient disorder predictor. However, current single-sequence-based methods are of low accuracy, while evolutionary profile-based methods are computationally intensive. Here, we proposed a fast and accurate protein disorder predictor LMDisorder that employed embedding generated by unsupervised pretrained language models as features. We showed that LMDisorder performs best in all single-sequence-based methods and is comparable or better than another language-model-based technique in four independent test sets, respectively. Furthermore, LMDisorder showed equivalent or even better performance than the state-of-the-art profile-based technique SPOT-Disorder2. In addition, the high computation efficiency of LMDisorder enabled proteome-scale analysis of human, showing that proteins with high predicted disorder content were associated with specific biological functions. The datasets, the source codes, and the trained model are available at https://github.com/biomed-AI/LMDisorder. Yidong Song, Qianmu Yuan, Ken Chen 0006, Yaoqi Zhou, Yuedong Yang |
Briefings Bioinform. | 6 |
| 2023 | Accurately identifying nucleic-acid-binding sites through geometric graph learning on language model predicted structuresabstractThe interactions between nucleic acids and proteins are important in diverse biological processes. The high-quality prediction of nucleic-acid-binding sites continues to pose a significant challenge. Presently, the predictive efficacy of sequence-based methods is constrained by their exclusive consideration of sequence context information, whereas structure-based methods are unsuitable for proteins lacking known tertiary structures. Though protein structures predicted by AlphaFold2 could be used, the extensive computing requirement of AlphaFold2 hinders its use for genome-wide applications. Based on the recent breakthrough of ESMFold for fast prediction of protein structures, we have developed GLMSite, which accurately identifies DNA- and RNA-binding sites using geometric graph learning on ESMFold predicted structures. Here, the predicted protein structures are employed to construct protein structural graph with residues as nodes and spatially neighboring residue pairs for edges. The node representations are further enhanced through the pre-trained language model ProtTrans. The network was trained using a geometric vector perceptron, and the geometric embeddings were subsequently fed into a common network to acquire common binding characteristics. Finally, these characteristics were input into two fully connected layers to predict binding sites with DNA and RNA, respectively. Through comprehensive tests on DNA/RNA benchmark datasets, GLMSite was shown to surpass the latest sequence-based methods and be comparable with structure-based methods. Moreover, the prediction was shown useful for inferring nucleic-acid-binding proteins, demonstrating its potential for protein function discovery. The datasets, codes, and trained models are available at https://github.com/biomed-AI/nucleic-acid-binding. Yidong Song, Qianmu Yuan, Huiying Zhao, Yuedong Yang |
Briefings Bioinform. | 4 |
| 2023 | Fast and accurate protein function prediction from sequence through pretrained language model and homology-based label diffusionabstractProtein function prediction is an essential task in bioinformatics which benefits disease mechanism elucidation and drug target discovery. Due to the explosive growth of proteins in sequence databases and the diversity of their functions, it remains challenging to fast and accurately predict protein functions from sequences alone. Although many methods have integrated protein structures, biological networks or literature information to improve performance, these extra features are often unavailable for most proteins. Here, we propose SPROF-GO, a Sequence-based alignment-free PROtein Function predictor, which leverages a pretrained language model to efficiently extract informative sequence embeddings and employs self-attention pooling to focus on important residues. The prediction is further advanced by exploiting the homology information and accounting for the overlapping communities of proteins with related functions through the label diffusion algorithm. SPROF-GO was shown to surpass state-of-the-art sequence-based and even network-based approaches by more than 14.5, 27.3 and 10.1% in area under the precision-recall curve on the three sub-ontology test sets, respectively. Our method was also demonstrated to generalize well on non-homologous proteins and unseen species. Finally, visualization based on the attention mechanism indicated that SPROF-GO is able to capture sequence domains useful for function prediction. The datasets, source codes and trained models of SPROF-GO are available at https://github.com/biomed-AI/SPROF-GO. The SPROF-GO web server is freely available at http://bio-web1.nscc-gz.cn/app/sprof-go. Qianmu Yuan, Jiancong Xie, Huiying Zhao, Yuedong Yang |
Briefings Bioinform. | 5 |
| 2023 | Identifying spatial domain by adapting transcriptomics with histology through contrastive learningabstractRecent advances in spatial transcriptomics have enabled measurements of gene expression at cell/spot resolution meanwhile retaining both the spatial information and the histology images of the tissues. Accurately identifying the spatial domains of spots is a vital step for various downstream tasks in spatial transcriptomics analysis. To remove noises in gene expression, several methods have been developed to combine histopathological images for data analysis of spatial transcriptomics. However, these methods either use the image only for the spatial relations for spots, or individually learn the embeddings of the gene expression and image without fully coupling the information. Here, we propose a novel method ConGI to accurately exploit spatial domains by adapting gene expression with histopathological images through contrastive learning. Specifically, we designed three contrastive loss functions within and between two modalities (the gene expression and image data) to learn the common representations. The learned representations are then used to cluster the spatial domains on both tumor and normal spatial transcriptomics datasets. ConGI was shown to outperform existing methods for the spatial domain identification. In addition, the learned representations have also been shown powerful for various downstream tasks, including trajectory inference, clustering, and visualization. Yuansong Zeng, Mai Luo, Jianing Chen 0004, Zixiang Pan, Yutong Lu, Weijiang Yu, Yuedong Yang |
Briefings Bioinform. | 8 |
| 2023 | EVlncRNA-Dpred: improved prediction of experimentally validated lncRNAs by deep learningabstractLong non-coding RNAs (lncRNAs) played essential roles in nearly every biological process and disease. Many algorithms were developed to distinguish lncRNAs from mRNAs in transcriptomic data and facilitated discoveries of more than 600 000 of lncRNAs. However, only a tiny fraction (<1%) of lncRNA transcripts (~4000) were further validated by low-throughput experiments (EVlncRNAs). Given the cost and labor-intensive nature of experimental validations, it is necessary to develop computational tools to prioritize those potentially functional lncRNAs because many lncRNAs from high-throughput sequencing (HTlncRNAs) could be resulted from transcriptional noises. Here, we employed deep learning algorithms to separate EVlncRNAs from HTlncRNAs and mRNAs. For overcoming the challenge of small datasets, we employed a three-layer deep-learning neural network (DNN) with a K-mer feature as the input and a small convolutional neural network (CNN) with one-hot encoding as the input. Three separate models were trained for human (h), mouse (m) and plant (p), respectively. The final concatenated models (EVlncRNA-Dpred (h), EVlncRNA-Dpred (m) and EVlncRNA-Dpred (p)) provided substantial improvement over a previous model based on support-vector-machines (EVlncRNA-pred). For example, EVlncRNA-Dpred (h) achieved 0.896 for the area under receiver-operating characteristic curve, compared with 0.582 given by sequence-based EVlncRNA-pred model. The models developed here should be useful for screening lncRNA transcripts for experimental validations. EVlncRNA-Dpred is available as a web server at https://www.sdklab-biophysics-dzu.net/EVlncRNA-Dpred/index.html, and the data and source code can be freely available along with the web server. Bailing Zhou, Maolin Ding, Baohua Ji, Pingping Huang, Junye Zhang, Zanxia Cao, Yuedong Yang, Yaoqi Zhou, Jihua Wang |
Briefings Bioinform. | 9 |
| 2023 | Identifying B-cell epitopes using AlphaFold2 predicted structures and pretrained language modelabstractMOTIVATION: Identifying the B-cell epitopes is an essential step for guiding rational vaccine development and immunotherapies. Since experimental approaches are expensive and time-consuming, many computational methods have been designed to assist B-cell epitope prediction. However, existing sequence-based methods have limited performance since they only use contextual features of the sequential neighbors while neglecting structural information. RESULTS: Based on the recent breakthrough of AlphaFold2 in protein structure prediction, we propose GraphBepi, a novel graph-based model for accurate B-cell epitope prediction. For one protein, the predicted structure from AlphaFold2 is used to construct the protein graph, where the nodes/residues are encoded by ESM-2 learning representations. The graph is input into the edge-enhanced deep graph neural network (EGNN) to capture the spatial information in the predicted 3D structures. In parallel, a bidirectional long short-term memory neural networks (BiLSTM) are employed to capture long-range dependencies in the sequence. The learned low-dimensional representations by EGNN and BiLSTM are then combined into a multilayer perceptron for predicting B-cell epitopes. Through comprehensive tests on the curated epitope dataset, GraphBepi was shown to outperform the state-of-the-art methods by more than 5.5% and 44.0% in terms of AUC and AUPR, respectively. A web server is freely available at http://bio-web1.nscc-gz.cn/app/graphbepi. AVAILABILITY AND IMPLEMENTATION: The datasets, pre-computed features, source codes, and the trained model are available at https://github.com/biomed-AI/GraphBepi. Yuansong Zeng, Zhuoyi Wei, Qianmu Yuan, Weijiang Yu, Yutong Lu, Jianzhao Gao, Yuedong Yang |
Bioinform. | 8 |
| 2023 | SUGAR: Efficient Subgraph-Level Training via Resource-Aware Graph PartitioningabstractGraph Neural Networks (GNNs) have demonstrated a great potential in a variety of graph-based applications, such as recommender systems, drug discovery, and object recognition. Nevertheless, resource-efficient GNN learning is a rarely explored topic despite its many benefits for edge computing and Internet of Things (IoT) applications. To improve this state of affairs, this work proposes efficientsubgraph-level training viaresource-aware graph partitioning (SUGAR). SUGAR first partitions the initial graph into a set of disjoint subgraphs and then performs local training at the subgraph-level We provide a theoretical analysis and conduct extensive experiments on five graph benchmarks to verify its efficacy in practice. Our results across five different hardware platforms demonstrate great runtime speedup and memory reduction of SUGAR on large-scale graphs. We believe SUGAR opens a new research direction towards developing GNN methods that are resource-efficient, hence suitable for IoT deployment. Zihui Xue, Yuedong Yang, Radu Marculescu |
IEEE Trans. Computers | 2 |
| 2023 | A Drug Combination Prediction Framework Based on Graph Convolutional Network and Heterogeneous InformationabstractCombination therapy, which can improve therapeutic efficacy and reduce side effects, plays an important role in the treatment of complex diseases. Yet, a large number of possible combinations among candidate compounds limits our ability to identify effective combinations. Though many studies have focused on predicting potential drug combinations, the existing methods are not entirely satisfactory in terms of performance and scalability. In this study, we propose a new computational pipeline, called DCMGCN, which integrates diverse drug-related information, to predict novel drug combinations. Specifically, DCMGCN first learns low-dimensional representations of drugs from the drug attributes and similarity networks. Then, by quantifying the degree of the nodes in the known drug-drug network and the similarity between connected nodes, we found the drug-drug network has heterophily and sparseness, which may limit the effectiveness of the graph convolutional network (GCN). Therefore, we introduce two designs to modify GCN. Finally, the drug representations are optimized using modified GCN (MGCN) and used to predict drug combinations. The tests on multiple drug combination datasets show that DCMGCN achieved substantial improvements over state-of-the-art methods. Importantly, our model may embed the mechanism of ground-truth drug pairs into the low-dimensional representation of each drug, which may help to further clarify the understanding of mechanisms of drug action. Hegang Chen, Yuyin Lu, Yuedong Yang, Yanghui Rao |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2023 | Onboard Sensors-Based Self-Localization for Autonomous Vehicle With Hierarchical MapabstractLocalization is a fundamental and crucial module for autonomous vehicles. Most of the existing localization methodologies, such as signal-dependent methods (RTK-GPS and Bluetooth), simultaneous localization and mapping (SLAM), and map-based methods, have been utilized in outdoor autonomous driving vehicles and indoor robot positioning. However, they suffer from severe limitations, such as signal-blocked scenes of GPS, computing resource occupation explosion in large-scale scenarios, intolerable time delay, and registration divergence of SLAM/map-based methods. In this article, a self-localization framework, without relying on GPS or any other wireless signals, is proposed. We demonstrate that the proposed homogeneous normal distribution transform algorithm and two-way information interaction mechanism could achieve centimeter-level localization accuracy, which reaches the requirement of autonomous vehicle localization for instantaneity and robustness. In addition, benefitting from hardware and software co-design, the proposed localization approach is extremely light-weighted enough to be operated on an embedded computing system, which is different from other LiDAR localization methods relying on high-performance CPU/GPU. Experiments on a public dataset (Baidu Apollo SouthBay dataset) and real-world verified the effectiveness and advantages of our approach compared with other similar algorithms. Yanqing Shen, Yuedong Yang, Xiaodong Deng, Shi-tao Chen, Jingmin Xin, Nanning Zheng 0001 |
IEEE Trans. Cybern. | 3 |
| 2023 | Subgraph-Aware Few-Shot Inductive Link Prediction Via Meta-LearningabstractLink prediction for knowledge graphs aims to predict missing connections between entities. Prevailing methods are limited to a transductive setting and hard to process unseen entities. The recently proposed subgraph-based models provide alternatives to predict links from the subgraph structure surrounding a candidate triplet. However, these methods require abundant known facts of training triplets and perform poorly on relationships that only have a few triplets. In this paper, we propose Meta-iKG, a novel subgraph-based meta-learner for few-shot inductive relation reasoning. Meta-iKG utilizes local subgraphs to transfer subgraph-specific information and to rapidly learn transferable patterns via meta-gradients. In this way, we find the model can quickly adapt to few-shot relationships using only a handful of known facts with inductive settings. Moreover, we introduce a large-shot relation updating procedure to ensure that our model can generalize well to both few-shot and large-shot relations. We evaluate Meta-iKG on inductive benchmarks sampled from the NELL and Freebase, and the results show that Meta-iKG outperforms the currently state-of-the-art methods in both few-shot scenarios and standard inductive settings. Shuangjia Zheng, Sijie Mai, Haifeng Hu 0001, Yuedong Yang |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2022 | A Meta-learning based Graph-Hierarchical Clustering Method for Single Cell RNA-Seq DataabstractSingle cell sequencing techniques enable researchers view complex bio-tissues from a more precise perspective to identify cell types. However, more and more recent works have been done to find more detailed subtypes within already known cell types. Here, we present MeHi-SCC, a method which utilized meta-learning protocol and brought in multi scRNA-seq datasets’ information in order to assist graph-based hierarchical sub-clustering process. In result, MeHi-SCC outperformed current-prevailing scRNA clustering methods and successfully identified cell subtypes in two large scale cell atlas. Our codes and datasets are available online at https://github.com/biomed-AI/MeHi-SCC Zixiang Pan, Yuansong Zeng, Yuefan Lin, Weijiang Yu, Haokun Zhang, Yuedong Yang |
BIBM | 6 |
| 2022 | Accurately Identifying Coronary Atherosclerotic Heart Disease through Merged Beats of ElectrocardiogramabstractCoronary Atherosclerotic Heart Disease (CAHD) is one kind of severe heart disease that is the dominating cause of death from non-communicable diseases worldwide. CAHD can be early detected through pre-symptomatic health check-ups, and the electrocardiogram (ECG) is common for non-invasive health check diagnoses. Traditionally, ECG signals are utilized to extract clinical features that are then input into machine learning methods for training and prediction. While these extracted features are interpretable, they are difficult to break through known features. On the other hand, ECG can be directly input to deep learning techniques, but such methods are usually limited by small sample sizes. Here, we propose to merge multiple beats of raw signal into one beat, which greatly reduces the complexity while maintaining the raw information. Moreover, we have constructed the largest benchmark dataset for 1113 CAHD patients of 12-lead ECG signals from the UK Biobank database and used the data to train a deep learning model. The results indicated that merged beat signals could achieve the best performance corresponding to an AUC of 0.71 and accuracy of 0.7, which is 4% higher than models using the raw signals and 6% higher than those using the clinical features. Further intuitive interpretation revealed that ST waves in lead II and V3 are the most closely associated with CAHD, consistent with clinical observations. Xinfeng Wang, Mengling Qi, Chengzhi Dong, Yuedong Yang, Huiying Zhao |
BIBM | 5 |
| 2022 | Genetic and phenotypic relationships between coronary atherosclerotic heart disease and electrocardiographic traitsabstractObservational studies have revealed that Coronary Atherosclerotic Heart Disease (CAHD) is associated with abnormal electrocardiogram (ECG) traits. However, it remains unclear whether there are genetic correlations between ECG and CAHD. Here, we explored genetic correlations and putative causal relationships between CAHD and ECG by performing Mendelian randomization (MR) and Polygenic risk score (PRS) analyses on the summary statistics from a large-scale genome-wide association study (GWAS) for CAHD (FinnGen: Ncase 23363, Ncontrol 187840) and ECG traits (UK Biobank: Ncase=1137, Ncontrol=40823). Results showed a causal genetic relationship between CAHD and six ECG traits in the lead V6. These ECG traits combining with age and gender have predicted CAHD risk with an AUC of 0.76. Further summary data-based Mendelian randomization (SMR) analysis identified 11 risk genes associated with the causality between CAHD and ECG. Thus, the revealed putative causal effects of CAHD on ECG traits provide genetic evidence to support the importance of monitoring CAHD risk through the ECG. Xinfeng Wang, Xuehao Xiu, Mengling Qi, Yuedong Yang, Huiying Zhao |
BIBM | 5 |
| 2022 | SCdenoise: a reference-based scRNA-seq denoising method using semi-supervised learningabstractscRNA-seq is a promising technology to perform unbiased, high-throughput, and high-resolution transcriptome analysis at single-cell resolution. The raw data usually suffers from noise and low quality, such as dropout events, which hinder downstream analysis. Thus, it is essential to improve the quality of single-cell data. Although many methods have been developed for denoising scRNA-seq data, the existing methods mainly focus on finding the relationship within the data itself without fully utilizing other datasets with annotated cell labels. Here, we proposed SCdenoise, a semi-supervised denoising method, to denoise unlabeled target data based on annotated cells in the reference datasets, which could utilize biological characteristics hidden in the high-quality reference datasets. Extensive downstream analyses showed that our method outperformed state-of-the-art methods on both simulated and real datasets for single-cell data analyses, including gene expression recovery, differential analysis, and clustering analysis. The source code is available at https://github.com/zhongfqi/SCdenoise-. Fengqi Zhong, Yuansong Zeng, Yuedong Yang |
BIBM | 4 |
| 2022 | Communicative Subgraph Representation Learning for Multi-Relational Inductive Drug-Gene Interaction PredictionabstractIlluminating the interconnections between drugs and genes is an important topic in drug development and precision medicine. Currently, computational predictions of drug-gene interactions mainly focus on the binding interactions without considering other relation types like agonist, antagonist, etc. In addition, existing methods either heavily rely on high-quality domain features or are intrinsically transductive, which limits the capacity of models to generalize to drugs/genes that lack external information or are unseen during the training process. To address these problems, we propose a novel Communicative Subgraph representation learning for Multi-relational Inductive drug-Gene interactions prediction (CoSMIG), where the predictions of drug-gene relations are made through subgraph patterns, and thus are naturally inductive for unseen drugs/genes without retraining or utilizing external domain features. Moreover, the model strengthened the relations on the drug-gene graph through a communicative message passing mechanism. To evaluate our method, we compiled two new benchmark datasets from DrugBank and DGIdb. The comprehensive experiments on the two datasets showed that our method outperformed state-of-the-art baselines in the transductive scenarios and achieved superior performance in the inductive ones. Further experimental analysis including LINCS experimental validation and literature verification also demonstrated the value of our model. Jiahua Rao, Shuangjia Zheng, Sijie Mai, Yuedong Yang |
IJCAI | 4 |
| 2022 | Capturing large genomic contexts for accurately predicting enhancer-promoter interactionsabstractEnhancer-promoter interaction (EPI) is a key mechanism underlying gene regulation. EPI prediction has always been a challenging task because enhancers could regulate promoters of distant target genes. Although many machine learning models have been developed, they leverage only the features in enhancers and promoters, or simply add the average genomic signals in the regions between enhancers and promoters, without utilizing detailed features between or outside enhancers and promoters. Due to a lack of large-scale features, existing methods could achieve only moderate performance, especially for predicting EPIs in different cell types. Here, we present a Transformer-based model, TransEPI, for EPI prediction by capturing large genomic contexts. TransEPI was developed based on EPI datasets derived from Hi-C or ChIA-PET data in six cell lines. To avoid over-fitting, we evaluated the TransEPI model by testing it on independent test datasets where the cell line and chromosome are different from the training data. TransEPI not only achieved consistent performance across the cross-validation and test datasets from different cell types but also outperformed the state-of-the-art machine learning and deep learning models. In addition, we found that the improved performance of TransEPI was attributed to the integration of large genomic contexts. Lastly, TransEPI was extended to study the non-coding mutations associated with brain disorders or neural diseases, and we found that TransEPI was also useful for predicting the target genes of non-coding mutations. Ken Chen 0006, Huiying Zhao, Yuedong Yang |
Briefings Bioinform. | 3 |
| 2022 | AlphaFold2-aware protein-DNA binding site prediction using graph transformerabstractProtein-DNA interactions play crucial roles in the biological systems, and identifying protein-DNA binding sites is the first step for mechanistic understanding of various biological activities (such as transcription and repair) and designing novel drugs. How to accurately identify DNA-binding residues from only protein sequence remains a challenging task. Currently, most existing sequence-based methods only consider contextual features of the sequential neighbors, which are limited to capture spatial information. Based on the recent breakthrough in protein structure prediction by AlphaFold2, we propose an accurate predictor, GraphSite, for identifying DNA-binding residues based on the structural models predicted by AlphaFold2. Here, we convert the binding site prediction problem into a graph node classification task and employ a transformer-based variant model to take the protein structural information into account. By leveraging predicted protein structures and graph transformer, GraphSite substantially improves over the latest sequence-based and structure-based methods. The algorithm is further confirmed on the independent test set of 181 proteins, where GraphSite surpasses the state-of-the-art structure-based method by 16.4% in area under the precision-recall curve and 11.2% in Matthews correlation coefficient, respectively. We provide the datasets, the predicted structures and the source codes along with the pre-trained models of GraphSite at https://github.com/biomed-AI/GraphSite. The GraphSite web server is freely available at https://biomed.nscc-gz.cn/apps/GraphSite. Qianmu Yuan, Jiahua Rao, Shuangjia Zheng, Huiying Zhao, Yuedong Yang |
Briefings Bioinform. | 6 |
| 2022 | Alignment-free metal ion-binding site prediction from protein sequence through pretrained language model and multi-task learningabstractMore than one-third of the proteins contain metal ions in the Protein Data Bank. Correct identification of metal ion-binding residues is important for understanding protein functions and designing novel drugs. Due to the small size and high versatility of metal ions, it remains challenging to computationally predict their binding sites from protein sequence. Existing sequence-based methods are of low accuracy due to the lack of structural information, and time-consuming owing to the usage of multi-sequence alignment. Here, we propose LMetalSite, an alignment-free sequence-based predictor for binding sites of the four most frequently seen metal ions in BioLiP (Zn2+, Ca2+, Mg2+ and Mn2+). LMetalSite leverages the pretrained language model to rapidly generate informative sequence representations and employs transformer to capture long-range dependencies. Multi-task learning is adopted to compensate for the scarcity of training data and capture the intrinsic similarities between different metal ions. LMetalSite was shown to surpass state-of-the-art structure-based methods by more than 19.7, 14.4, 36.8 and 12.6% in area under the precision recall on the four independent tests, respectively. Further analyses indicated that the self-attention modules are effective to learn the structural contexts of residues from protein sequence. We provide the data sets, source codes and trained models of LMetalSite at https://github.com/biomed-AI/LMetalSite. Qianmu Yuan, Yu Wang 0008, Huiying Zhao, Yuedong Yang |
Briefings Bioinform. | 5 |
| 2022 | A robust and scalable graph neural network for accurate single-cell classificationabstractSingle-cell RNA sequencing (scRNA-seq) techniques provide high-resolution data on cellular heterogeneity in diverse tissues, and a critical step for the data analysis is cell type identification. Traditional methods usually cluster the cells and manually identify cell clusters through marker genes, which is time-consuming and subjective. With the launch of several large-scale single-cell projects, millions of sequenced cells have been annotated and it is promising to transfer labels from the annotated datasets to newly generated datasets. One powerful way for the transferring is to learn cell relations through the graph neural network (GNN), but traditional GNNs are difficult to process millions of cells due to the expensive costs of the message-passing procedure at each training epoch. Here, we have developed a robust and scalable GNN-based method for accurate single-cell classification (GraphCS), where the graph is constructed to connect similar cells within and between labelled and unlabeled scRNA-seq datasets for propagation of shared information. To overcome the slow information propagation of GNN at each training epoch, the diffused information is pre-calculated via the approximate Generalized PageRank algorithm, enabling sublinear complexity over cell numbers. Compared with existing methods, GraphCS demonstrates better performance on simulated, cross-platform, cross-species and cross-omics scRNA-seq datasets. More importantly, our model provides a high speed and scalability on large datasets, and can achieve superior performance for 1 million cells within 50 min. Yuansong Zeng, Zhuoyi Wei, Zixiang Pan, Yutong Lu, Yuedong Yang |
Briefings Bioinform. | 5 |
| 2022 | Spatial transcriptomics prediction from histology jointly through Transformer and graph neural networksabstractThe rapid development of spatial transcriptomics allows the measurement of RNA abundance at a high spatial resolution, making it possible to simultaneously profile gene expression, spatial locations of cells or spots, and the corresponding hematoxylin and eosin-stained histology images. It turns promising to predict gene expression from histology images that are relatively easy and cheap to obtain. For this purpose, several methods are devised, but they have not fully captured the internal relations of the 2D vision features or spatial dependency between spots. Here, we developed Hist2ST, a deep learning-based model to predict RNA-seq expression from histology images. Around each sequenced spot, the corresponding histology image is cropped into an image patch and fed into a convolutional module to extract 2D vision features. Meanwhile, the spatial relations with the whole image and neighbored patches are captured through Transformer and graph neural network modules, respectively. These learned features are then used to predict the gene expression by following the zero-inflated negative binomial distribution. To alleviate the impact by the small spatial transcriptomics data, a self-distillation mechanism is employed for efficient learning of the model. By comprehensive tests on cancer and normal datasets, Hist2ST was shown to outperform existing methods in terms of both gene expression prediction and spatial region identification. Further pathway analyses indicated that our model could reserve biological information. Thus, Hist2ST enables generating spatial transcriptomics data from histology images for elucidating molecular signatures of tissues. Yuansong Zeng, Zhuoyi Wei, Weijiang Yu, Yuchen Yuan, Bingling Li, Zhonghui Tang, Yutong Lu, Yuedong Yang |
Briefings Bioinform. | 9 |
| 2022 | A parameter-free deep embedded clustering method for single-cell RNA-seq dataabstractClustering analysis is widely used in single-cell ribonucleic acid (RNA)-sequencing (scRNA-seq) data to discover cell heterogeneity and cell states. While many clustering methods have been developed for scRNA-seq analysis, most of these methods require to provide the number of clusters. However, it is not easy to know the exact number of cell types in advance, and experienced determination is not always reliable. Here, we have developed ADClust, an automatic deep embedding clustering method for scRNA-seq data, which can accurately cluster cells without requiring a predefined number of clusters. Specifically, ADClust first obtains low-dimensional representation through pre-trained autoencoder and uses the representations to cluster cells into initial micro-clusters. The clusters are then compared in between by a statistical test, and similar micro-clusters are merged into larger clusters. According to the clustering, cell representations are updated so that each cell will be pulled toward centers of its assigned cluster and similar clusters, while cells are separated to keep distances between clusters. This is accomplished through jointly optimizing the carefully designed clustering and autoencoder loss functions. This merging process continues until convergence. ADClust was tested on 11 real scRNA-seq datasets and was shown to outperform existing methods in terms of both clustering performance and the accuracy on the number of the determined clusters. More importantly, our model provides high speed and scalability for large datasets. Yuansong Zeng, Zhuoyi Wei, Fengqi Zhong, Zixiang Pan, Yutong Lu, Yuedong Yang |
Briefings Bioinform. | 6 |
| 2022 | A coarse-refine segmentation network for COVID-19 CT imagesabstractThe rapid spread of the novel coronavirus disease 2019 (COVID-19) causes a significant impact on public health. It is critical to diagnose COVID-19 patients so that they can receive reasonable treatments quickly. The doctors can obtain a precise estimate of the infection's progression and decide more effective treatment options by segmenting the CT images of COVID-19 patients. However, it is challenging to segment infected regions in CT slices because the infected regions are multi-scale, and the boundary is not clear due to the low contrast between the infected area and the normal area. In this paper, a coarse-refine segmentation network is proposed to address these challenges. The coarse-refine architecture and hybrid loss is used to guide the model to predict the delicate structures with clear boundaries to address the problem of unclear boundaries. The atrous spatial pyramid pooling module in the network is added to improve the performance in detecting infected regions with different scales. Experimental results show that the model in the segmentation of COVID-19 CT images outperforms other familiar medical segmentation models, enabling the doctor to get a more accurate estimate on the progression of the infection and thus can provide more reasonable treatment options. Ziwang Huang, Xiang Zhang 0012, Huiying Zhao, Yutian Chong, Hejun Wu, Yuedong Yang, Jun Shen 0008, Yunfei Zha |
IET Image Process. | 9 |
| 2022 | Imputing DNA Methylation by Transferred Learning Based Neural Network
Xinfeng Wang, Jiahua Rao, Zhu-Jin Zhang, Yuedong Yang |
J. Comput. Sci. Technol. | 5 |
| 2022 | Dynamic graph dropout for subgraph-based relation prediction
Sijie Mai, Shuangjia Zheng, Yuedong Yang, Haifeng Hu 0001 |
Knowl. Based Syst. | 5 |
| 2022 | To Improve Prediction of Binding Residues With DNA, RNA, Carbohydrate, and Peptide Via Multi-Task Deep Neural NetworksabstractMOTIVATION: The interactions of proteins with DNA, RNA, peptide, and carbohydrate play key roles in various biological processes. The studies of uncharacterized protein-molecules interactions could be aided by accurate predictions of residues that bind with partner molecules. However, the existing methods for predicting binding residues on proteins remain of relatively low accuracies due to the limited number of complex structures in databases. As different types of molecules partially share chemical mechanisms, the predictions for each molecular type should benefit from the binding information with other molecule types. RESULTS: In this study, we employed a multiple task deep learning strategy to develop a new sequence-based method for simultaneously predicting binding residues/sites with multiple important molecule types named MTDsite. By combining four training sets for DNA, RNA, peptide, and carbohydrate-binding proteins, our method yielded accurate and robust predictions with AUC values of 0.852, 0836, 0.758, and 0.776 on their respective independent test sets, which are 0.52 to 6.6% better than other state-of-the-art methods. To my best knowledge, this is the first method using multi-task framework to predict multiple molecular binding sites simultaneously. Shuangjia Zheng, Huiying Zhao, Zhangming Niu, Yutong Lu, Yi Pan 0001, Yuedong Yang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 7 |
| 2021 | Communicative Message Passing for Inductive Relation ReasoningabstractRelation prediction for knowledge graphs aims at predicting missing relationships between entities. Despite the importance of inductive relation prediction, most previous works are limited to a transductive setting and cannot process previously unseen entities. The recent proposed subgraph-based relation reasoning models provided alternatives to predict links from the subgraph structure surrounding a candidate triplet inductively. However, we observe that these methods often neglect the directed nature of the extracted subgraph and weaken the role of relation information in the subgraph modeling. As a result, they fail to effectively handle the asymmetric/anti-symmetric triplets and produce insufficient embeddings for the target triplets. To this end, we introduce a Communicative Message Passing neural network for Inductive reLation rEasoning, CoMPILE, that reasons over local directed subgraph structures and has a vigorous inductive bias to process entity-independent semantic relations. In contrast to existing models, CoMPILE strengthens the message interactions between edges and entitles through a communicative kernel and enables a sufficient flow of relation information. Moreover, we demonstrate that CoMPILE can naturally handle asymmetric/anti-symmetric relations without the need for explosively increasing the number of model parameters by extracting the directed enclosing subgraphs. Extensive experiments show substantial performance gains in comparison to state-of-the-art methods on commonly used benchmark datasets with variant inductive settings. Sijie Mai, Shuangjia Zheng, Yuedong Yang, Haifeng Hu 0001 |
AAAI | 3 |
| 2021 | SEGEM: a Fast and Accurate Automated Protein Backbone Structure Modeling Method for Cryo-EMabstractCryo-electron microscopy (cryo-EM) technique has been widely used in protein structure determination, whereas it remains a challenge to automatically build accurate protein backbone structure from cryo-EM density map. A typical pipeline to automatically build a structure model from cryo-EM map is to first predict $\mathrm{C}\alpha$ sites and then assign them to protein sequence, which is a typical combinatorial optimization task of extremely high computational complexity. Here we propose SEGEM, a fast and accurate automated protein backbone structure modeling method for cryo-EM. We employed 3D Convolutional Neural Networks to predict $\mathrm{C}\alpha$ sites with their amino acid types from cryo-EM, and developed a highly parallel pipeline to assign $\mathrm{C}\alpha$ sites with their predicted amino acid types to protein sequence. We tested SEGEM on three benchmark datasets where it significantly outperformed several state-of-the-art prediction methods including MAINMAST, C-CNN and DeepTracer. In our method plus version SEGEM++, we combined SEGEM with the protein structure prediction algorithm AlphaFold2. SEGEM++is capable to identify whether AlphaFold2 folds a good structure, and rectify the incorrectly folded region through protein threading on cryo-EM map. In our curated dataset of hard targets where AlphaFold2 predicted structures obtained an average RMSD of 7. 87A and GDT-TS score of 0.652 when superimposed to the native structure, SEGEM++ achieved a significantly better RMSD of 2.46A and 0.676 GDT-TS score on average. Furthermore, with our highly parallel pipeline on 30 cores CPU, both SEGEM and SEGEM++ finished structure modeling within 10 minutes on average in our test datasets, indicating their potential in high throughout automated accurate backbone structure modeling for cryo-EM. Xiongjun Li, Yuedong Yang |
BIBM | 5 |
| 2021 | Multi-omics Cancer Prognosis Analysis Based on Graph Convolution NetworkabstractCancer survival analysis is important for patients’ follow-up treatment, and accurate prediction is conductive to improving the survival rate of patients. With the development of gene sequencing technology, multi-omics data is increasingly utilized to predict cancer prognosis. However, it remains challenging to fuse different sources of omics data. To solve this problem, we introduced the interpretable constraint relationship between genes into the model and proposed a new method Graph Survival Network (GraphSurv) by combining the graph convolution network (GCN) with the deep Cox proportional hazard network. Compared with state-of-the-art methods, the average C-index value in 11 cancer datasets was increased by 4.0%. As a case study of liver hepatocellular carcinoma (LIHC), we downloaded GSE54236 and GSE14520 from the GEO database for independent tests. The results showed that our model could significantly distinguish high-risk from low-risk patients (C-index >0.6, log-rank p-value < 0.05). In addition, we performed differential expression analysis based on the divided risk subgroups by GraphSurv Among the top 15 identified differentially expressed genes (DEGs) ranked by log2 change value, 12 (80%) prognostic markers have been confirmed by literature review. The results proved that the method was robust and can identify genes related to the cancer prognosis. Yuedong Yang |
BIBM | 4 |
| 2021 | DGAT-onco: A differential analysis method to detect oncogenes by integrating functional information of mutationsabstractIt is a common strategy to predict oncogenes by differential analysis between somatic mutations and background mutations. Most previous methods only utilize mutations in the cancer population to model its background mutation, which have an obvious bias. A recent method, DiffMut, improves this issue by conducting differential mutational analysis with both mutations in the cancer population and the natural population. However, it assumes the impacts of all mutations are equal, neglecting their functional difference. Thus, we developed a method, DGAT-onco that integrated the functional impacts of mutations to the differential mutational analysis framework of DiffMut. We performed DGAT-onco analysis with 33 cancer types from the Cancer Genome Atlas (TCGA) dataset. Its reliability was further evaluated on an independent test set including 22 cancers from other sources (TS22). Using oncogenes from the Cancer Gene Census (CGC) as the gold standard, our method achieves higher classification performance in oncogene discovery than five alternative methods (i.e., DiffMut, WITER, OncodriveCLUSTL, OncodriveFML, and MutSigCV) with an average AUPRC of 0.197 and 0.187 in TCGA and TS22 respectively. The source code and supplementary materials of DGAT-onco are available at https://github.com/zhanghaoyang0/DGAT-onco. Junkang Wei, Zifeng Liu, Yutian Chong, Yutong Lu, Huiying Zhao, Yuedong Yang |
BIBM | 8 |
| 2021 | DeepANIS: Predicting antibody paratope from concatenated CDR sequences by integrating bidirectional long-short-term memory and transformer neural networksabstractAntibodies are a type of important biomolecules in the humoral immunity system, which can bind tightly to potential antigens with high affinity and specificity. An accurate identification of the paratope, the binding sites with antigens, is crucial for antibody mechanistic research and design. Although many methods have been developed for paratope prediction, further improvement of their accuracy is necessary. In this study, we concatenated the sequences of Complementarity Determining Regions (CDRs) within a single antibody to better capture nonlocal interactions between different CDRs and loop type-specific features for improving paratope prediction. We further integrated BiLSTM and transformer networks to gain the dependencies among the residues within the concatenated CDR sequences and to increase the interpretability of the model. The new method called DeepANIS (Antibody Interacting Site prediction) outperforms other antibody paratope prediction methods compared. The DeepANIS method is freely available as a webserver at https://biomed.nscc-gz.cn/apps/DeepANIS and for download at https://github.com/HideInDust/DeepANIS. Shuangjia Zheng, Yaoqi Zhou, Yuedong Yang |
BIBM | 5 |
| 2021 | Learning Attributed Graph Representation with Communicative Message Passing TransformerabstractConstructing appropriate representations of molecules lies at the core of numerous tasks such as material science, chemistry, and drug designs. Recent researches abstract molecules as attributed graphs and employ graph neural networks (GNN) for molecular representation learning, which have made remarkable achievements in molecular graph modeling. Albeit powerful, current models either are based on local aggregation operations and thus miss higher-order graph properties or focus on only node information without fully using the edge information. For this sake, we propose a Communicative Message Passing Transformer (CoMPT) neural network to improve the molecular graph representation by reinforcing message interactions between nodes and edges based on the Transformer architecture. Unlike the previous transformer-style GNNs that treat molecule as a fully connected graph, we introduce a message diffusion mechanism to leverage the graph connectivity inductive bias and reduce the message enrichment explosion. Extensive experiments demonstrated that the proposed model obtained superior performances (around 4% on average) against state-of-the-art baselines on seven chemical property datasets (graph-level tasks) and two chemical shift datasets (node-level tasks). Further visualization studies also indicated a better representation capacity achieved by our model. Shuangjia Zheng, Jiahua Rao, Yuedong Yang |
IJCAI | 5 |
| 2021 | Integration of Patch Features Through Self-supervised Learning and Transformer for Survival Analysis on Whole Slide Images
Ziwang Huang, Haitao Wang 0026, Yuedong Yang, Hejun Wu |
MICCAI (8) | 5 |
| 2021 | PharmKG: a dedicated knowledge graph benchmark for bomedical data miningabstractBiomedical knowledge graphs (KGs), which can help with the understanding of complex biological systems and pathologies, have begun to play a critical role in medical practice and research. However, challenges remain in their embedding and use due to their complex nature and the specific demands of their construction. Existing studies often suffer from problems such as sparse and noisy datasets, insufficient modeling methods and non-uniform evaluation metrics. In this work, we established a comprehensive KG system for the biomedical field in an attempt to bridge the gap. Here, we introduced PharmKG, a multi-relational, attributed biomedical KG, composed of more than 500 000 individual interconnections between genes, drugs and diseases, with 29 relation types over a vocabulary of ~8000 disambiguated entities. Each entity in PharmKG is attached with heterogeneous, domain-specific information obtained from multi-omics data, i.e. gene expression, chemical structure and disease word embedding, while preserving the semantic and biomedical features. For baselines, we offered nine state-of-the-art KG embedding (KGE) approaches and a new biological, intuitive, graph neural network-based KGE method that uses a combination of both global network structure and heterogeneous domain features. Based on the proposed benchmark, we conducted extensive experiments to assess these KGE models using multiple evaluation metrics. Finally, we discussed our observations across various downstream biological tasks and provide insights and guidelines for how to use a KG in biomedicine. We hope that the unprecedented quality and diversity of PharmKG will lead to advances in biomedical KG construction, embedding and application. Shuangjia Zheng, Jiahua Rao, Xianglu Xiao, Evandro Fei Fang, Yuedong Yang, Zhangming Niu |
Briefings Bioinform. | 7 |
| 2021 | scAdapt: virtual adversarial domain adaptation network for single cell RNA-seq data classification across platforms and speciesabstractIn single cell analyses, cell types are conventionally identified based on expressions of known marker genes, whose identifications are time-consuming and irreproducible. To solve this issue, many supervised approaches have been developed to identify cell types based on the rapid accumulation of public datasets. However, these approaches are sensitive to batch effects or biological variations since the data distributions are different in cross-platforms or species predictions. In this study, we developed scAdapt, a virtual adversarial domain adaptation network, to transfer cell labels between datasets with batch effects. scAdapt used both the labeled source and unlabeled target data to train an enhanced classifier and aligned the labeled source centroids and pseudo-labeled target centroids to generate a joint embedding. The scAdapt was demonstrated to outperform existing methods for classification in simulated, cross-platforms, cross-species, spatial transcriptomic and COVID-19 immune datasets. Further quantitative evaluations and visualizations for the aligned embeddings confirm the superiority in cell mixing and the ability to preserve discriminative cluster structure present in the original datasets. Yuansong Zeng, Huiying Zhao, Yuedong Yang |
Briefings Bioinform. | 5 |
| 2021 | Structure-aware protein-protein interaction site prediction using deep graph convolutional networkabstractMOTIVATION: Protein-protein interactions (PPI) play crucial roles in many biological processes, and identifying PPI sites is an important step for mechanistic understanding of diseases and design of novel drugs. Since experimental approaches for PPI site identification are expensive and time-consuming, many computational methods have been developed as screening tools. However, these methods are mostly based on neighbored features in sequence, and thus limited to capture spatial information. RESULTS: We propose a deep graph-based framework deep Graph convolutional network for Protein-Protein-Interacting Site prediction (GraphPPIS) for PPI site prediction, where the PPI site prediction problem was converted into a graph node classification task and solved by deep learning using the initial residual and identity mapping techniques. We showed that a deeper architecture (up to eight layers) allows significant performance improvement over other sequence-based and structure-based methods by more than 12.5% and 10.5% on AUPRC and MCC, respectively. Further analyses indicated that the predicted interacting sites by GraphPPIS are more spatially clustered and closer to the native ones even when false-positive predictions are made. The results highlight the importance of capturing spatially neighboring residues for interacting site prediction. AVAILABILITY AND IMPLEMENTATION: The datasets, the pre-computed features, and the source codes along with the pre-trained models of GraphPPIS are available at https://github.com/biomed-AI/GraphPPIS. The GraphPPIS web server is freely available at https://biomed.nscc-gz.cn/apps/GraphPPIS. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Qianmu Yuan, Huiying Zhao, Yaoqi Zhou, Yuedong Yang |
Bioinform. | 5 |
| 2021 | Predicting bladder cancer prognosis by integrating multi-omics data through a transfer learning-based Cox proportional hazards network
Yuedong Yang |
CCF Trans. High Perform. Comput. | 4 |
| 2021 | Deep Learning Enables Accurate Diagnosis of Novel Coronavirus (COVID-19) With CT ImagesabstractA novel coronavirus (COVID-19) recently emerged as an acute respiratory syndrome, and has caused a pneumonia outbreak world-widely. As the COVID-19 continues to spread rapidly across the world, computed tomography (CT) has become essentially important for fast diagnoses. Thus, it is urgent to develop an accurate computer-aided method to assist clinicians to identify COVID-19-infected patients by CT images. Here, we have collected chest CT scans of 88 patients diagnosed with COVID-19 from hospitals of two provinces in China, 100 patients infected with bacteria pneumonia, and 86 healthy persons for comparison and modeling. Based on the data, a deep learning-based CT diagnosis system was developed to identify patients with COVID-19. The experimental results showed that our model could accurately discriminate the COVID-19 patients from the bacteria pneumonia patients with an AUC of 0.95, recall (sensitivity) of 0.96, and precision of 0.79. When integrating three types of CT images, our model achieved a recall of 0.93 with precision of 0.86 for discriminating COVID-19 patients from others. Moreover, our model could extract main lesion features, especially the ground-glass opacity (GGO), which are visually helpful for assisted diagnoses by doctors. An online server is available for online diagnoses with CT images by our server (http://biomed.nscc-gz.cn/model.php). Source codes and datasets are available at our GitHub (https://github.com/SY575/COVID19-CT). Shuangjia Zheng, Xiang Zhang 0012, Ziwang Huang, Huiying Zhao, Yutian Chong, Jun Shen 0008, Yunfei Zha, Yuedong Yang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 13 |
| 2020 | An End-to-end Oxford Nanopore Basecaller Using Convolution-augmented TransformerabstractThe following topics are dealt with: learning (artificial intelligence); diseases; medical image processing; molecular biophysics; genetics; medical computing; feature extraction; cancer; genomics; proteins. Xuan Lv, Zhiguang Chen 0001, Yutong Lu, Yuedong Yang |
BIBM | 4 |
| 2020 | Accurately Clustering Single-cell RNA-seq data by Capturing Structural Relations between Cells through Graph Convolutional NetworkabstractRecent advances in single-cell RNA sequencing (scRNA-seq) technologies provide a great opportunity to study gene expression at cellular resolution, and the scRNA-seq data has been routinely conducted to unfold cell heterogeneity and diversity. A critical step for the scRNA-seq analyses is to cluster the same type of cells, and many methods have been developed for cell clustering. However, existing clustering methods are limited to extract the representations from expression data of individual cells, while ignoring the high-order structural relations between cells. Here, we proposed a new method (GraphSCC) to cluster cells based on scRNA-seq data by accounting structural relations between cells through a graph convolutional network. The representation learned from the graph convolutional network, together with another representation output from a denoising autoencoder network, are optimized by a dual self-supervised module for better cell clustering. Extensive experiments indicate that GraphSCC model outperforms state-of-the-art methods in various evaluation metrics on both simulated and real datasets. Yuansong Zeng, Jiahua Rao, Yutong Lu, Yuedong Yang |
BIBM | 5 |
| 2020 | HeLPS: Heterogeneous LiDAR-based Positioning System for Autonomous VehicleabstractLiDAR-based positioning systems are widely used in unmanned systems. However, affected by the high computational complexity of high-precision positioning algorithms, the current positioning system is supported by hardware with low power efficiency and thus hard to integrate into many platforms. In this paper, we analyze features of the positioning system in autonomous driving application, design, and apply Heterogeneous LiDAR-based Positioning System, HeLPS, with software-hardware co-design methodology to achieve better efficiency. Our contributions can be concluded in three aspects. Firstly, we design the CPU-FPGA heterogeneous positioning system accelerating Iterative Closest Point (ICP) algorithm and achieves improvements on both speed and power-efficiency. Secondly, we exploit the spatial locality in the point cloud and design a new compressed data structure for fast neighbor accessing. The experiment reports a significant speedup comparing with other data structures. Lastly, we explore the data access pattern in positioning application and develop a specific cache system and out-of-order execution system reducing memory burden. Our system is deployed on a small and cheap Xilinx Zynq7000 ARM+FPGA platform, which achieves 983.3x speedup compared with Cortex-A9 CPU, and 31.8x speedup compared with i7-7820 CPU, with only 2.37W power consumption. Yuedong Yang, Xiaodong Deng, Yanqing Shen, Shi-tao Chen, Nanning Zheng 0001 |
IECON | 1 |
| 2020 | Communicative Representation Learning on Attributed Molecular GraphsabstractConstructing proper representations of molecules lies at the core of numerous tasks such as molecular property prediction and drug design. Graph neural networks, especially message passing neural network (MPNN) and its variants, have recently made remarkable achievements in molecular graph modeling. Albeit powerful, the one-sided focuses on atom (node) or bond (edge) information of existing MPNN methods lead to the insufficient representations of the attributed molecular graphs. Herein, we propose a Communicative Message Passing Neural Network (CMPNN) to improve the molecular embedding by strengthening the message interactions between nodes and edges through a communicative kernel. In addition, the message generation process is enriched by introducing a new message booster module. Extensive experiments demonstrated that the proposed model obtained superior performances against state-of-the-art baselines on six chemical property datasets. Further visualization also showed better representation capacity of our model. Shuangjia Zheng, Zhangming Niu, Zhang-Hua Fu, Yutong Lu, Yuedong Yang |
IJCAI | 6 |
| 2020 | Accurate prediction of genome-wide RNA secondary structure profile based on extreme gradient boostingabstractMOTIVATION: RNA secondary structure plays a vital role in fundamental cellular processes, and identification of RNA secondary structure is a key step to understand RNA functions. Recently, a few experimental methods were developed to profile genome-wide RNA secondary structure, i.e. the pairing probability of each nucleotide, through high-throughput sequencing techniques. However, these high-throughput methods have low precision and cannot cover all nucleotides due to limited sequencing coverage. RESULTS: Here, we have developed a new method for the prediction of genome-wide RNA secondary structure profile from RNA sequence based on the extreme gradient boosting technique. The method achieves predictions with areas under the receiver operating characteristic curve (AUC) >0.9 on three different datasets, and AUC of 0.888 by another independent test on the recently released Zika virus data. These AUCs are consistently >5% greater than those by the CROSS method recently developed based on a shallow neural network. Further analysis on the 1000 Genome Project data showed that our predicted unpaired probabilities are highly correlated (>0.8) with the minor allele frequencies at synonymous, non-synonymous mutations, and mutations in untranslated regions, which were higher than those generated by RNAplfold. Moreover, the prediction over all human mRNA indicated a consistent result with previous observation that there is a periodic distribution of unpaired probability on codons. The accurate predictions by our method indicate that such model trained on genome-wide experimental data might be an alternative for analytical methods. AVAILABILITY AND IMPLEMENTATION: The GRASP is available for academic use at https://github.com/sysu-yanglab/GRASP. SUPPLEMENTARY INFORMATION: Supplementary data are available online. Yaobin Ke, Jiahua Rao, Huiying Zhao, Yutong Lu, Nong Xiao 0001, Yuedong Yang |
Bioinform. | 6 |
| 2019 | Improving prediction of protein secondary structure, backbone angles, solvent accessibility and contact numbers by using predicted contact maps and an ensemble of recurrent and residual convolutional neural networksabstractMOTIVATION: Sequence-based prediction of one dimensional structural properties of proteins has been a long-standing subproblem of protein structure prediction. Recently, prediction accuracy has been significantly improved due to the rapid expansion of protein sequence and structure libraries and advances in deep learning techniques, such as residual convolutional networks (ResNets) and Long-Short-Term Memory Cells in Bidirectional Recurrent Neural Networks (LSTM-BRNNs). Here we leverage an ensemble of LSTM-BRNN and ResNet models, together with predicted residue-residue contact maps, to continue the push towards the attainable limit of prediction for 3- and 8-state secondary structure, backbone angles (θ, τ, ϕ and ψ), half-sphere exposure, contact numbers and solvent accessible surface area (ASA). RESULTS: The new method, named SPOT-1D, achieves similar, high performance on a large validation set and test set (≈1000 proteins in each set), suggesting robust performance for unseen data. For the large test set, it achieves 87% and 77% in 3- and 8-state secondary structure prediction and 0.82 and 0.86 in correlation coefficients between predicted and measured ASA and contact numbers, respectively. Comparison to current state-of-the-art techniques reveals substantial improvement in secondary structure and backbone angle prediction. In particular, 44% of 40-residue fragment structures constructed from predicted backbone Cα-based θ and τ angles are less than 6 Å root-mean-squared-distance from their native conformations, nearly 20% better than the next best. The method is expected to be useful for advancing protein structure and function prediction. AVAILABILITY AND IMPLEMENTATION: SPOT-1D and its data is available at: http://sparks-lab.org/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jack Hanson, Kuldip K. Paliwal, Thomas Litfin, Yuedong Yang, Yaoqi Zhou |
Bioinform. | 4 |
| 2018 | Sixty-five years of the long march in protein secondary structure prediction: the final stretch?abstractProtein secondary structure prediction began in 1951 when Pauling and Corey predicted helical and sheet conformations for protein polypeptide backbone even before the first protein structure was determined. Sixty-five years later, powerful new methods breathe new life into this field. The highest three-state accuracy without relying on structure templates is now at 82-84%, a number unthinkable just a few years ago. These improvements came from increasingly larger databases of protein sequences and structures for training, the use of template secondary structure information and more powerful deep learning techniques. As we are approaching to the theoretical limit of three-state prediction (88-90%), alternative to secondary structure prediction (prediction of backbone torsion angles and Cα-atom-based angles and torsion angles) not only has more room for further improvement but also allows direct prediction of three-dimensional fragment structures with constantly improved accuracy. About 20% of all 40-residue fragments in a database of 1199 non-redundant proteins have <6 Å root-mean-squared distance from the native conformations by SPIDER2. More powerful deep learning methods with improved capability of capturing long-range interactions begin to emerge as the next generation of techniques for secondary structure prediction. The time has come to finish off the final stretch of the long march towards protein secondary structure prediction. Yuedong Yang, Jianzhao Gao, Jihua Wang, Rhys Heffernan, Jack Hanson, Kuldip K. Paliwal, Yaoqi Zhou |
Briefings Bioinform. | 1 |
| 2018 | Accurate prediction of protein contact maps by coupling residual two-dimensional bidirectional long short-term memory with convolutional neural networksabstractMotivation: Accurate prediction of a protein contact map depends greatly on capturing as much contextual information as possible from surrounding residues for a target residue pair. Recently, ultra-deep residual convolutional networks were found to be state-of-the-art in the latest Critical Assessment of Structure Prediction techniques (CASP12) for protein contact map prediction by attempting to provide a protein-wide context at each residue pair. Recurrent neural networks have seen great success in recent protein residue classification problems due to their ability to propagate information through long protein sequences, especially Long Short-Term Memory (LSTM) cells. Here, we propose a novel protein contact map prediction method by stacking residual convolutional networks with two-dimensional residual bidirectional recurrent LSTM networks, and using both one-dimensional sequence-based and two-dimensional evolutionary coupling-based information. Results: We show that the proposed method achieves a robust performance over validation and independent test sets with the Area Under the receiver operating characteristic Curve (AUC) > 0.95 in all tests. When compared to several state-of-the-art methods for independent testing of 228 proteins, the method yields an AUC value of 0.958, whereas the next-best method obtains an AUC of 0.909. More importantly, the improvement is over contacts at all sequence-position separations. Specifically, a 8.95%, 5.65% and 2.84% increase in precision were observed for the top L∕10 predictions over the next best for short, medium and long-range contacts, respectively. This confirms the usefulness of ResNets to congregate the short-range relations and 2D-BRLSTM to propagate the long-range dependencies throughout the entire protein contact map 'image'. Availability and implementation: SPOT-Contact server url: http://sparks-lab.org/jack/server/SPOT-Contact/. Supplementary information: Supplementary data are available at Bioinformatics online. Jack Hanson, Kuldip K. Paliwal, Thomas Litfin, Yuedong Yang, Yaoqi Zhou |
Bioinform. | 4 |
| 2018 | Structure-based prediction of protein- peptide binding regions using Random ForestabstractMotivation: Protein-peptide interactions are one of the most important biological interactions and play crucial role in many diseases including cancer. Therefore, knowledge of these interactions provides invaluable insights into all cellular processes, functional mechanisms, and drug discovery. Protein-peptide interactions can be analyzed by studying the structures of protein-peptide complexes. However, only a small portion has known complex structures and experimental determination of protein-peptide interaction is costly and inefficient. Thus, predicting peptide-binding sites computationally will be useful to improve efficiency and cost effectiveness of experimental studies. Here, we established a machine learning method called SPRINT-Str (Structure-based prediction of protein-Peptide Residue-level Interaction) to use structural information for predicting protein-peptide binding residues. These predicted binding residues are then employed to infer the peptide-binding site by a clustering algorithm. Results: SPRINT-Str achieves robust and consistent results for prediction of protein-peptide binding regions in terms of residues and sites. Matthews' Correlation Coefficient (MCC) for 10-fold cross validation and independent test set are 0.27 and 0.293, respectively, as well as 0.775 and 0.782, respectively for area under the curve. The prediction outperforms other state-of-the-art methods, including our previously developed sequence-based method. A further spatial neighbor clustering of predicted binding residues leads to prediction of binding sites at 20-116% higher coverage than the next best method at all precision levels in the test set. The application of SPRINT-Str to protein binding with DNA, RNA and carbohydrate confirms the method's capability of separating peptide-binding sites from other functional sites. More importantly, similar performance in prediction of binding residues and sites is obtained when experimentally determined structures are replaced by unbound structures or quality model structures built from homologs, indicating its wide applicability. Availability and implementation: http://sparks-lab.org/server/SPRINT-Str. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Ghazaleh Taherzadeh, Yaoqi Zhou, Alan Wee-Chung Liew, Yuedong Yang |
Bioinform. | 4 |
| 2018 | Grid-based prediction of torsion angle probabilities of protein backbone and its application to discrimination of protein intrinsic disorder regions and selection of model structuresabstractBACKGROUND: bond (τ). Thus, their accurate prediction is useful for structure prediction and model refinement. Early methods predicted torsion angles in a few discrete bins whereas most recent methods have focused on prediction of angles in real, continuous values. Real value prediction, however, is unable to provide the information on probabilities of predicted angles. RESULTS: Here, we propose to predict angles in fine grids of 5° by using deep learning neural networks. We found that this grid-based technique can yield 2-6% higher accuracy in predicting angles in the same 5° bin than existing prediction techniques compared. We further demonstrate the usefulness of predicted probabilities at given angle bins in discrimination of intrinsically disorder regions and in selection of protein models. CONCLUSIONS: The proposed method may be useful for characterizing protein structure and disorder. The method is available at http://sparks-lab.org/server/SPIDER2/ as a part of SPIDER2 package. Jianzhao Gao, Yuedong Yang, Yaoqi Zhou |
BMC Bioinform. | 2 |
| 2017 | Improving protein disorder prediction by deep bidirectional long short-term memory recurrent neural networksabstractMotivation: Capturing long-range interactions between structural but not sequence neighbors of proteins is a long-standing challenging problem in bioinformatics. Recently, long short-term memory (LSTM) networks have significantly improved the accuracy of speech and image classification problems by remembering useful past information in long sequential events. Here, we have implemented deep bidirectional LSTM recurrent neural networks in the problem of protein intrinsic disorder prediction. Results: The new method, named SPOT-Disorder, has steadily improved over a similar method using a traditional, window-based neural network (SPINE-D) in all datasets tested without separate training on short and long disordered regions. Independent tests on four other datasets including the datasets from critical assessment of structure prediction (CASP) techniques and >10 000 annotated proteins from MobiDB, confirmed SPOT-Disorder as one of the best methods in disorder prediction. Moreover, initial studies indicate that the method is more accurate in predicting functional sites in disordered regions. These results highlight the usefulness combining LSTM with deep bidirectional recurrent neural networks in capturing non-local, long-range interactions for bioinformatics applications. Availability and Implementation: SPOT-disorder is available as a web server and as a standalone program at: http://sparks-lab.org/server/SPOT-disorder/index.php . Contact: [email protected] or [email protected] or [email protected]. Supplementary information: Supplementary data is available at Bioinformatics online. Jack Hanson, Yuedong Yang, Kuldip K. Paliwal, Yaoqi Zhou |
Bioinform. | 2 |
| 2017 | Capturing non-local interactions by long short-term memory bidirectional recurrent neural networks for improving prediction of protein secondary structure, backbone angles, contact numbers and solvent accessibilityabstractMOTIVATION: The accuracy of predicting protein local and global structural properties such as secondary structure and solvent accessible surface area has been stagnant for many years because of the challenge of accounting for non-local interactions between amino acid residues that are close in three-dimensional structural space but far from each other in their sequence positions. All existing machine-learning techniques relied on a sliding window of 10-20 amino acid residues to capture some 'short to intermediate' non-local interactions. Here, we employed Long Short-Term Memory (LSTM) Bidirectional Recurrent Neural Networks (BRNNs) which are capable of capturing long range interactions without using a window. RESULTS: We showed that the application of LSTM-BRNN to the prediction of protein structural properties makes the most significant improvement for residues with the most long-range contacts (|i-j| >19) over a previous window-based, deep-learning method SPIDER2. Capturing long-range interactions allows the accuracy of three-state secondary structure prediction to reach 84% and the correlation coefficient between predicted and actual solvent accessible surface areas to reach 0.80, plus a reduction of 5%, 10%, 5% and 10% in the mean absolute error for backbone ϕ , ψ , θ and τ angles, respectively, from SPIDER2. More significantly, 27% of 182724 40-residue models directly constructed from predicted C α atom-based θ and τ have similar structures to their corresponding native structures (6Å RMSD or less), which is 3% better than models built by ϕ and ψ angles. We expect the method to be useful for assisting protein structure and function prediction. AVAILABILITY AND IMPLEMENTATION: The method is available as a SPIDER3 server and standalone package at http://sparks-lab.org . CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Rhys Heffernan, Yuedong Yang, Kuldip K. Paliwal, Yaoqi Zhou |
Bioinform. | 2 |
| 2017 | SPOT-ligand 2: improving structure-based virtual screening by binding-homology search on an expanded structural template libraryabstractMotivation: The high cost of drug discovery motivates the development of accurate virtual screening tools. Binding-homology, which takes advantage of known protein-ligand binding pairs, has emerged as a powerful discrimination technique. In order to exploit all available binding data, modelled structures of ligand-binding sequences may be used to create an expanded structural binding template library. Results: SPOT-Ligand 2 has demonstrated significantly improved screening performance over its previous version by expanding the template library 15 times over the previous one. It also performed better than or similar to other binding-homology approaches on the DUD and DUD-E benchmarks. Availability and Implementation: The server is available online at http://sparks-lab.org . Contacts: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Thomas Litfin, Yaoqi Zhou, Yuedong Yang |
Bioinform. | 3 |
| 2017 | LRFragLib: an effective algorithm to identify fragments for de novo protein structure predictionabstractMotivation: The quality of fragment library determines the efficiency of fragment assembly, an approach that is widely used in most de novo protein-structure prediction algorithms. Conventional fragment libraries are constructed mainly based on the identities of amino acids, sometimes facilitated by predicted information including dihedral angles and secondary structures. However, it remains challenging to identify near-native fragment structures with low sequence homology. Results: We introduce a novel fragment-library-construction algorithm, LRFragLib, to improve the detection of near-native low-homology fragments of 7-10 residues, using a multi-stage, flexible selection protocol. Based on logistic regression scoring models, LRFragLib outperforms existing techniques by achieving a significantly higher precision and a comparable coverage on recent CASP protein sets in sampling near-native structures. The method also has a comparable computational efficiency to the fastest existing techniques with substantially reduced memory usage. Availability and Implementation: The source code is available for download at http://166.111.152.91/Downloads.html. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Tong Wang 0014, Yuedong Yang, Yaoqi Zhou, Haipeng Gong |
Bioinform. | 2 |
| 2016 | Fast and accurate non-sequential protein structure alignment using a new asymmetric linear sum assignment heuristicabstractMOTIVATION: The three dimensional tertiary structure of a protein at near atomic level resolution provides insight alluding to its function and evolution. As protein structure decides its functionality, similarity in structure usually implies similarity in function. As such, structure alignment techniques are often useful in the classifications of protein function. Given the rapidly growing rate of new, experimentally determined structures being made available from repositories such as the Protein Data Bank, fast and accurate computational structure comparison tools are required. This paper presents SPalignNS, a non-sequential protein structure alignment tool using a novel asymmetrical greedy search technique. RESULTS: The performance of SPalignNS was evaluated against existing sequential and non-sequential structure alignment methods by performing trials with commonly used datasets. These benchmark datasets used to gauge alignment accuracy include (i) 9538 pairwise alignments implied by the HOMSTRAD database of homologous proteins; (ii) a subset of 64 difficult alignments from set (i) that have low structure similarity; (iii) 199 pairwise alignments of proteins with similar structure but different topology; and (iv) a subset of 20 pairwise alignments from the RIPC set. SPalignNS is shown to achieve greater alignment accuracy (lower or comparable root-mean squared distance with increased structure overlap coverage) for all datasets, and the highest agreement with reference alignments from the challenging dataset (iv) above, when compared with both sequentially constrained alignments and other non-sequential alignments. AVAILABILITY AND IMPLEMENTATION: SPalignNS was implemented in C++. The source code, binary executable, and a web server version is freely available at: http://sparks-lab.org CONTACT: [email protected]. Wayne J. Pullan, Yuedong Yang, Yaoqi Zhou |
Bioinform. | 3 |
| 2016 | Predicting the errors of predicted local backbone angles and non-local solvent- accessibilities of proteins by deep neural networksabstractMOTIVATION: Backbone structures and solvent accessible surface area of proteins are benefited from continuous real value prediction because it removes the arbitrariness of defining boundary between different secondary-structure and solvent-accessibility states. However, lacking the confidence score for predicted values has limited their applications. Here we investigated whether or not we can make a reasonable prediction of absolute errors for predicted backbone torsion angles, Cα-atom-based angles and torsion angles, solvent accessibility, contact numbers and half-sphere exposures by employing deep neural networks. RESULTS: We found that angle-based errors can be predicted most accurately with Spearman correlation coefficient (SPC) between predicted and actual errors at about 0.6. This is followed by solvent accessibility (SPC∼0.5). The errors on contact-based structural properties are most difficult to predict (SPC between 0.2 and 0.3). We showed that predicted errors are significantly better error indicators than the average errors based on secondary-structure and amino-acid residue types. We further demonstrated the usefulness of predicted errors in model quality assessment. These error or confidence indictors are expected to be useful for prediction, assessment, and refinement of protein structures. AVAILABILITY AND IMPLEMENTATION: The method is available at http://sparks-lab.org as a part of SPIDER2 package. CONTACT: [email protected] or [email protected] information: Supplementary data are available at Bioinformatics online. Jianzhao Gao, Yuedong Yang, Yaoqi Zhou |
Bioinform. | 2 |
| 2016 | Highly accurate sequence-based prediction of half-sphere exposures of amino acid residues in proteinsabstractMOTIVATION: Solvent exposure of amino acid residues of proteins plays an important role in understanding and predicting protein structure, function and interactions. Solvent exposure can be characterized by several measures including solvent accessible surface area (ASA), residue depth (RD) and contact numbers (CN). More recently, an orientation-dependent contact number called half-sphere exposure (HSE) was introduced by separating the contacts within upper and down half spheres defined according to the Cα-Cβ (HSEβ) vector or neighboring Cα-Cα vectors (HSEα). HSEα calculated from protein structures was found to better describe the solvent exposure over ASA, CN and RD in many applications. Thus, a sequence-based prediction is desirable, as most proteins do not have experimentally determined structures. To our best knowledge, there is no method to predict HSEα and only one method to predict HSEβ. RESULTS: This study developed a novel method for predicting both HSEα and HSEβ (SPIDER-HSE) that achieved a consistent performance for 10-fold cross validation and two independent tests. The correlation coefficients between predicted and measured HSEβ (0.73 for upper sphere, 0.69 for down sphere and 0.76 for contact numbers) for the independent test set of 1199 proteins are significantly higher than existing methods. Moreover, predicted HSEα has a higher correlation coefficient (0.46) to the stability change by residue mutants than predicted HSEβ (0.37) and ASA (0.43). The results, together with its easy Cα-atom-based calculation, highlight the potential usefulness of predicted HSEα for protein structure prediction and refinement as well as function prediction. AVAILABILITY AND IMPLEMENTATION: The method is available at http://sparks-lab.org CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Rhys Heffernan, Abdollah Dehzangi, James G. Lyons, Kuldip K. Paliwal, Alok Sharma, Jihua Wang, Abdul Sattar 0001, Yaoqi Zhou, Yuedong Yang |
Bioinform. | 9 |
| 2015 | DDIG-in: detecting disease-causing genetic variations due to frameshifting indels and nonsense mutations employing sequence and structural properties at nucleotide and protein levelsabstractAbstract Motivation: Frameshifting (FS) indels and nonsense (NS) variants disrupt the protein-coding sequence downstream of the mutation site by changing the reading frame or introducing a premature termination codon, respectively. Despite such drastic changes to the protein sequence, FS indels and NS variants have been discovered in healthy individuals. How to discriminate disease-causing from neutral FS indels and NS variants is an understudied problem. Results: We have built a machine learning method called DDIG-in (FS) based on real human genetic variations from the Human Gene Mutation Database (inherited disease-causing) and the 1000 Genomes Project (GP) (putatively neutral). The method incorporates both sequence and predicted structural features and yields a robust performance by 10-fold cross-validation and independent tests on both FS indels and NS variants. We showed that human-derived NS variants and FS indels derived from animal orthologs can be effectively employed for independent testing of our method trained on human-derived FS indels. DDIG-in (FS) achieves a Matthews correlation coefficient (MCC) of 0.59, a sensitivity of 86%, and a specificity of 72% for FS indels. Application of DDIG-in (FS) to NS variants yields essentially the same performance (MCC of 0.43) as a method that was specifically trained for NS variants. DDIG-in (FS) was shown to make a significant improvement over existing techniques. Availability and implementation: The DDIG-in web-server for predicting NS variants, FS indels, and non-frameshifting (NFS) indels is available at http://sparks-lab.org/ddig. Contact: [email protected] Supplementary information: Supplementary data are available at Bioinformatics online. Lukas Folkman, Yuedong Yang, Zhixiu Li, Bela Stantic, Abdul Sattar 0001, Matthew E. Mort, David N. Cooper, Yaoqi Zhou |
Bioinform. | 2 |
| 2011 | Improving protein fold recognition and template-based modeling by employing probabilistic-based matching between predicted one-dimensional structural properties of query and corresponding native properties of templatesabstractMOTIVATION: In recent years, development of a single-method fold-recognition server lags behind consensus and multiple template techniques. However, a good consensus prediction relies on the accuracy of individual methods. This article reports our efforts to further improve a single-method fold recognition technique called SPARKS by changing the alignment scoring function and incorporating the SPINE-X techniques that make improved prediction of secondary structure, backbone torsion angle and solvent accessible surface area. RESULTS: The new method called SPARKS-X was tested with the SALIGN benchmark for alignment accuracy, Lindahl and SCOP benchmarks for fold recognition, and CASP 9 blind test for structure prediction. The method is compared to several state-of-the-art techniques such as HHPRED and BoostThreader. Results show that SPARKS-X is one of the best single-method fold recognition techniques. We further note that incorporating multiple templates and refinement in model building will likely further improve SPARKS-X. AVAILABILITY: The method is available as a SPARKS-X server at http://sparks.informatics.iupui.edu/ Yuedong Yang, Eshel Faraggi, Huiying Zhao, Yaoqi Zhou |
Bioinform. | 1 |
| 2010 | Structure-based prediction of DNA-binding proteins by structural alignment and a volume-fraction corrected DFIRE-based energy functionabstractMOTIVATION: Template-based prediction of DNA binding proteins requires not only structural similarity between target and template structures but also prediction of binding affinity between the target and DNA to ensure binding. Here, we propose to predict protein-DNA binding affinity by introducing a new volume-fraction correction to a statistical energy function based on a distance-scaled, finite, ideal-gas reference (DFIRE) state. RESULTS: We showed that this energy function together with the structural alignment program TM-align achieves the Matthews correlation coefficient (MCC) of 0.76 with an accuracy of 98%, a precision of 93% and a sensitivity of 64%, for predicting DNA binding proteins in a benchmark of 179 DNA binding proteins and 3797 non-binding proteins. The MCC value is substantially higher than the best MCC value of 0.69 given by previous methods. Application of this method to 2235 structural genomics targets uncovered 37 as DNA binding proteins, 27 (73%) of which are putatively DNA binding and only 1 protein whose annotated functions do not contain DNA binding, while the remaining proteins have unknown function. The method provides a highly accurate and sensitive technique for structure-based prediction of DNA binding proteins. AVAILABILITY: The method is implemented as a part of the Structure-based function-Prediction On-line Tools (SPOT) package available at http://sparks.informatics.iupui.edu/spot Huiying Zhao, Yuedong Yang, Yaoqi Zhou |
Bioinform. | 2 |