Tianyi Zang

dblp:00/6500 · DBLP profile ↗
← Back
43ranked-venue papers
3as first author
18since 2021 · last 2025
0000-0003-0195-8731ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 36 · 1 first-author · 15 since 2021Systems, architecture and hardware · 3 · 1 first-authorArtificial intelligence and machine learning · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 first-authorComputer networks · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Mkdban-Tei: a Multi-Level Knowledge Distillation-Based Deep Learning Architecture for Predicting T Cell Receptor-Epitope Binding Specificity
abstract
Understanding the underlying mechanisms of TCR-epitope interactions is crucial for studying the adaptive immune system and promoting the field of immunotherapy. Given the high cost of traditional experimental methods, it is urgent to develop computational methods to predict TCRepitope binding. With the advancement of experimental technology, an increasing number of TCR-epitope binding pairs have been archived in public databases, creating opportunities for the advancement of computational methods. In this study, we propose a novel framework called MKDBAN-TEI for predicting TCR-epitope binding. We encode TCR and epitope sequences using a learnable residue embedding matrix and employ CNN layers to extract features. An interpretable bilinear attention network is then used to capture the interaction patterns between TCR and epitope. To improve the model's performance and generalization capability, we introduce a multi-level knowledge distillation framework: first, we cluster epitopes in the training set based on sequence similarity to define distinct domains; second, for each domain, we integrate three types of protein sequence features (protein language embeddings, physicochemical information, and evolutionary information) to train domain-specific teacher models via internal multi-feature knowledge distillation, capturing domain-specific binding patterns; finally, we distill knowledge from all domain-specific teachers to a universal student model through inter-domain knowledge distillation, enhancing generalization to unseen epitopes. Compared to several state-of-the-art models, MKDBAN-TEI demonstrates superior performance and generalization capability. Further experiments illustrate the effectiveness of the model in realworld scenarios. Visualizing the attention maps learned by MKDBAN-TEI provides new insights into TCR-epitope interactions. The code is available at: https://github.com/X/MKDBAN-TEI
Haoyan Wang, Tianyi Zang
BIBM2
2025 Multi-Channel Learning Framework Based on Relation-Aware Transformer For Drug Synergy Prediction
abstract
Combined therapeutic strategies have demonstrated their importance in addressing complex diseases, particularly among patient populations where monotherapy is less effective. Compared to single drug treatments, the use of drug combinations can reduce the occurrence of drug resistance and enhance the efficacy of cancer treatments. Therefore, developing effective combination therapies through clinical trials holds significant importance for researchers and society as a whole. However, facing a vast library of compounds, conducting high-throughput drug combination screening is not only challenging but also costly. To overcome these difficulties, researchers have developed various computational methods that utilize biomedical data related to drugs to efficiently identify potential drug combinations. This study has developed a novel multi-channel representation learning framework based on knowledge graphs and Transformers models, aimed at predicting drug synergy. Unlike full-graph analysis, we employ a node-centric sampling strategy to extract subgraphs for learning node representations. Furthermore, we utilize a relation-based self-attention mechanism to handle complex relations between nodes. Through the multi-channel network, we input drug molecular structures, cell line information, and biomedical knowledge graphs into different channels, and fuse the feature representations of these channels to enhance the accuracy of drug synergy prediction. Our model benefits from both the structural features of drugs and rich biomedical background information. Extensive experiments on two representative databases have validated the effectiveness of our model.
Kaiyuan Zhang 0006, Tianyi Zang, Yanli Zhao
IJCNN4
2024 MFTEP: A Multimodal Fusion Deep Learning Framework for T Cell Receptor-epitope Interaction Prediction
abstract
Accurately predicting immunogenic peptides recognized by T cell receptors (TCR) is a crucial step toward personalized immunotherapy. However, prediction of TCR-epitope interactions is still a challenging task. Early works, like molecular dynamics simulation-based methods, suffer from slow speed and poor generalization capabilities. It is necessary to develop novel computational methods to predict TCR-epitope interactions precisely. With the development of high-throughput sequencing technologies, more and more TCR-epitope interaction data have been recorded in public databases. With the help of these databases, many in silico predictive methods have shown promising performance. However, current methods still perform poorly on unseen TCRs and epitopes. Moreover, most current models still accept single-modal information about TCRs and epitopes, such as sequences or physicochemical information. Effectively utilizing the multimodal information of TCRs and epitopes, such as molecular graphs and 3D structure, may enhance the model’s prediction performance. To address the above issues, we presented MFTEP, a multimodal fusion method to predict TCR-epitope interactions by fusing the sequence features, molecular graph features, and 3D structure features of TCRs and epitopes. The ablation study highlights the importance of the multi-modal fusion module in enhancing the model’s performance. Several datasets were collected and utilized to evaluate the generality and robustness of the proposed model. According to the experimental results, MFTEP performs better than other state-of-the-art methods, indicating its high predictive power. Overall, the results demonstrate that MFTEP can learn general TCR-epitope interaction patterns and is a powerful prediction tool to apply to real-world scenarios. All the data and code are available at: https://github.com/skybluewhy/MFTEP
Haoyan Wang, Tianyi Zang
BIBM3
2024 TDLM: A Diffusion Language Model for TCR Sequence Exploration and Generation
abstract
The adaptive immune response relies on the ability of T-cell receptors (TCRs) to recognize specific antigens. The vast diversity of TCRs allows T-cells to recognize a broad spectrum of antigens, but this complexity also poses challenges for understanding and predicting TCR-antigen binding specificity. Despite the development of various machine learning and deep learning methods for prediction and clustering, there remains a need for a versatile and effective TCR language framework that can be flexibly applied to various downstream tasks, including sequence generation. Here we present TDLM, a T-cell Receptor (TCR) diffusion language model, designed to decode complex patterns within TCR sequences and apply them across various downstream tasks. Firstly, TDLM can be trained on unlabeled TCR sequence data, enabling it to utilize vast datasets to generate comprehensive embeddings. When compared to other embedding methods, TDLM embeddings enhance TCR-antigen binding prediction accuracy and enable effective TCR sequence clustering and similarity analysis, helping identify TCRs with shared antigen specificity. Furthermore, as a diffusion-based generative model, TDLM can generate highly diverse and specific TCR sequences. This ability is invaluable for the rapid screening and optimization of TCRs with target antigen specificities, offering significant potential in disease diagnosis, personalized immunotherapy, and vaccine research. The code is available at: https://github.com/skybluewhy/TDLM
Haoyan Wang, Tianyi Zang, Yadong Liu 0001
BIBM4
2024 KGE-UNIT: toward the unification of molecular interactions prediction based on knowledge graph and multi-task learning on drug discovery
abstract
The prediction of molecular interactions is vital for drug discovery. Existing methods often focus on individual prediction tasks and overlook the relationships between them. Additionally, certain tasks encounter limitations due to insufficient data availability, resulting in limited performance. To overcome these limitations, we propose KGE-UNIT, a unified framework that combines knowledge graph embedding (KGE) and multi-task learning, for simultaneous prediction of drug-target interactions (DTIs) and drug-drug interactions (DDIs) and enhancing the performance of each task, even when data availability is limited. Via KGE, we extract heterogeneous features from the drug knowledge graph to enhance the structural features of drug and protein nodes, thereby improving the quality of features. Additionally, employing multi-task learning, we introduce an innovative predictor that comprises the task-aware Convolutional Neural Network-based (CNN-based) encoder and the task-aware attention decoder which can fuse better multimodal features, capture the contextual interactions of molecular tasks and enhance task awareness, leading to improved performance. Experiments on two imbalanced datasets for DTIs and DDIs demonstrate the superiority of KGE-UNIT, achieving high area under the receiver operating characteristics curves (AUROCs) (0.942, 0.987) and area under the precision-recall curve ( AUPRs) (0.930, 0.980) for DTIs and high AUROCs (0.975, 0.989) and AUPRs (0.966, 0.988) for DDIs. Notably, on the LUO dataset where the data were more limited, KGE-UNIT exhibited a more pronounced improvement, with increases of 4.32$\%$ in AUROC and 3.56$\%$ in AUPR for DTIs and 6.56$\%$ in AUROC and 8.17$\%$ in AUPR for DDIs. The scalability of KGE-UNIT is demonstrated through its extension to protein-protein interactions prediction, ablation studies and case studies further validate its effectiveness.
Tianyi Zang, Tianyi Zhao 0001
Briefings Bioinform.2
2024 DPDFormer: A Coarse-to-Fine Model for Monocular Depth Estimation
abstract
Monocular depth estimation attracts great attention from computer vision researchers for its convenience in acquiring environment depth information. Recently classification-based MDE methods show its promising performance and begin to act as an essential role in many multi-view applications such as reconstruction and 3D object detection. However, existed classification-based MDE models usually apply fixed depth range discretization strategy across a whole scene. This fixed depth range discretization leads to the imbalance of discretization scale among different depth ranges, resulting in the inexact depth range localization. In this article, to alleviate the imbalanced depth range discretization problem in classification-based monocular depth estimation (MDE) method we follow the coarse-to-fine principle and propose a novel depth range discretization method called depth post-discretization (DPD). Based on a coarse depth anchor roughly indicating the depth range, the DPD generates the depth range discretization adaptively for every position. The depth range discretization with DPD is more fine-grained around the actual depth, which is beneficial for locating the depth range more precisely for each scene position. Besides, to better manage the prediction of the coarse depth anchor and depth probability distribution for calculating the final depth, we design a dual-decoder transformer-based network, i.e., DPDFormer, which is more compatible with our proposed DPD method. We evaluate DPDFormer on popular depth datasets NYU Depth V2 and KITTI. The experimental results prove the superior performance of our proposed method.
Chunpu Liu, Guanglei Yang, Wangmeng Zuo, Tianyi Zang
ACM Trans. Multim. Comput. Commun. Appl.4
2023 TetraCVD: A Temporal-Textual Transformer based Model for Cardiovascular Disease Diagnosis
abstract
Cardiovascular disease (CVD) is one of the leading causes of death globally. There is considerable clinical significance and an emerging need of assisting doctors to diagnose cardiovascular disease and identify the subtype of it, from which doctors can provide different treatments and medications to increase the cure rate. The goal of this paper is to develop a deep learning model to predict cardiovascular disease and classify its subtype, which by handling data from two modalities of time-series vital signs and text report. We propose a temporal-textual transformer based model for cardiovascular disease diagnosis, TetraCVD, to address the challenges of irregular temporal feature extraction and medical long-text feature extraction respectively. TetraCVD is a multimodal deep learning model, consisting of two networks, cvdGNN and cvdHierBERT, as its time-series and language backbones, which leverage knowledge from temporal vital signs and text reports of the individuals respectively. Our results show that TetraCVD achieves promising performance in predicting subtypes of cardiovascular disease using the P18-ECER dataset and obtains state-of-the-art results. This study is among the first efforts that use both time-series vital signs and text report data to predict cardiovascular disease and its subtype. We argue that our approach can be generalized to predict and diagnose other diseases easily, and it can potentially play a significant role in the domain of general disease diagnosis in the future.
Kailong Lu, Penghuan Gu, Haoyan Wang, Tianyi Zang
BIBM5
2023 Relative order constraint for monocular depth estimation
Chunpu Liu, Wangmeng Zuo, Guanglei Yang, Wanlong Li, Hongbo Zhang 0004, Tianyi Zang
Appl. Intell.7
2022 MTOR hypermethylation may associate with the susceptibility and survival of SARS-CoV-2 infections to lung adenocarcinoma patients based on multi-omics data and machine learning
abstract
Recent studies have shown that lung adenocarcinoma (LUAD) patients have a higher risk and worse prognosis of COVID-19 caused by SARS-CoV-2 compared to normal samples. Whereas, in addition to the receptor for SARS-CoV-2, other genes also deserve attention. In our study, we identified 19 differentially methylated genes (DMGs) that were co-upregulated in LUAD and COVID-19 samples. These 19 DMGs mainly regulated the immune-related and multiple viral infection signaling pathways. Gene Ontology and pathway enrichment analysis were applied with these genes. Then, 6 key DMGs (MTOR, ACE, IGF1, PTPRC, C3, and PTGS2) were identified by constructing and analyzing the protein-protein interaction (PPI) network. Besides, MTOR was significantly associated with 5 prognostic markers (CDO1, NEURL4, SMAP1, NPEPPS, IQCK) identified by survival analysis based on machine learning. In total, MTOR hypermethylation may be related to the susceptibility of LUAD patients to SARS-CoV-2 and the prognosis of LUAD patients suffering from COVID-19.
Yang Hu 0008, Tianyi Zang
BIBM4
2022 Comparison of the Nanopore and PacBio sequencing technologies for DNA 5-methylcytosine detection
abstract
DNA methylation provides a pivotal layer of epigenetic regulation in eukaryotes that has significant involvement for numerous biological processes in health and disease. Recent long-read sequencing technology including Oxford Nanopore sequencing and PacBio HiFi sequencing greatly expands the capacity of long-range, single-molecule, and direct DNA modification detection from reads without extra laboratory techniques. A growing number of analytical pipelines including base-calling and 5mC methylation detection have been developed, but there is still a lack of comprehensive evaluations of the two sequencing technologies. Here, we assess the performance of different methylation-calling pipelines based on Nanopore and HiFi sequencing datasets to provide a systematic evaluation to guide researchers on how to select the long-read sequencing technologies in performing human epigenome-wide studies.
Yadong Liu 0001, Zhongyu Liu, Tao Jiang 0021, Tianyi Zang, Yadong Wang 0001
BIBM4
2022 MHCRoBERTa: pan-specific peptide-MHC class I binding prediction through transfer learning with label-agnostic protein sequences
abstract
Predicting the binding of peptide and major histocompatibility complex (MHC) plays a vital role in immunotherapy for cancer. The success of Alphafold of applying natural language processing (NLP) algorithms in protein secondary struction prediction has inspired us to explore the possibility of NLP methods in predicting peptide-MHC class I binding. Based on the above motivations, we propose the MHCRoBERTa method, RoBERTa pre-training approach, for predicting the binding affinity between type I MHC and peptides. Analysis of the results on benchmark dataset demonstrates that MHCRoBERTa can outperform other state-of-art prediction methods with an increase of the Spearman rank correlation coefficient (SRCC) value. Notably, our model gave a significant improvement on IC50 value. Our method has achieved SRCC value and AUC value as 0.785 and 0.817, respectively. Our SRCC value is 14.3% higher than NetMHCpan3.0 (the second highest SRCC value on pan-specific) and is 3% higher than MHCflurry (the second highest SRCC value on all methods). The AUC value is also better than any other pan-specific methods. Moreover, we visualize the multi-head self-attention for the token representation across the layers and heads by this method. Through the analysis of the representation of each layer and head, we can show whether the model has learned the syntax and semantics necessary to perform the prediction task well. All these results demonstrate that our model can accurately predict the peptide-MHC class I binding affinity and that MHCRoBERTa is a powerful tool for screening potential neoantigens for cancer immunotherapy. MHCRoBERTa is available as an open source software at github (https://github.com/FuxuWang/MHCRoBERTa).
Fuxu Wang, Haoyan Wang, Lizhuang Wang, Haoyu Lu, Shizheng Qiu, Tianyi Zang, Xinjun Zhang, Yang Hu 0008
Briefings Bioinform.6
2022 CNN-DDI: a learning-based method for predicting drug-drug interactions using convolution neural networks
abstract
BACKGROUND: Drug-drug interactions (DDIs) are the reactions between drugs. They are compartmentalized into three types: synergistic, antagonistic and no reaction. As a rapidly developing technology, predicting DDIs-associated events is getting more and more attention and application in drug development and disease diagnosis fields. In this work, we study not only whether the two drugs interact, but also specific interaction types. And we propose a learning-based method using convolution neural networks to learn feature representations and predict DDIs. RESULTS: In this paper, we proposed a novel algorithm using a CNN architecture, named CNN-DDI, to predict drug-drug interactions. First, we extract feature interactions from drug categories, targets, pathways and enzymes as feature vectors and employ the Jaccard similarity as the measurement of drugs similarity. Then, based on the representation of features, we build a new convolution neural network as the DDIs' predictor. CONCLUSION: The experimental results indicate that drug categories is effective as a new feature type applied to CNN-DDI method. And using multiple features is more informative and more effective than single feature. It can be concluded that CNN-DDI has more superiority than other existing algorithms on task of predicting DDIs.
Tianyi Zang
BMC Bioinform.3
2022 A multi-network integration approach for measuring disease similarity based on ncRNA regulation and heterogeneous information
abstract
BACKGROUND: Measuring similarity between complex diseases has significant implications for revealing the pathogenesis of diseases and development in the domain of biomedicine. It has been consentaneous that functional associations between disease-related genes and semantic associations can be applied to calculate disease similarity. Currently, more and more studies have demonstrated the profound involvement of non-coding RNA in the regulation of genome organization and gene expression. Thus, taking ncRNA into account can be useful in measuring disease similarities. However, existing methods ignore the regulation functions of ncRNA in biological process. In this study, we proposed a novel deep-learning method to deduce disease similarity. RESULTS: In this article, we proposed a novel method, ImpAESim, a framework integrating multiple networks embedding to learn compact feature representations and disease similarity calculation. We first utilize three different disease-related information networks to build up a heterogeneous network, after a network diffusion process, RWR, a compact feature learning model composed of classic Auto Encoder (AE) and improved AE model is proposed to extract constraints and low-dimensional feature representations. We finally obtain an accurate and low-dimensional feature representation of diseases, then we employed the cosine distance as the measurement of disease similarity. CONCLUSION: ImpAESim focuses on extracting a low-dimensional vector representation of features based on ncRNA regulation, and gene-gene interaction network. Our method can significantly reduce the calculation bias resulted from the sparse disease associations which are derived from semantic associations.
Ningyi Zhang, Tianyi Zang
BMC Bioinform.2
2021 prePathCluster: An novel deep-learning based method for endocrine disease pathway analysis
abstract
Identification and investigation of molecular pathways are important in exploring the underlying mechanism of complex diseases. With the help of molecular and intermolecular information derived from pathways, we developed a disease pathway prediction method prePathCluster, to identify significant pathway sub-clusters that may have not been revealed to be associated with specific diseases currently. Since the interactions between pathways, genes and protein complexes can be regarded as a genome-scale interaction network, prePathCluster could identify disease associated pathways with less characterized roles in previous studies based on pathway features and network topological structure. Unlike most of current databases providing only a few candidate disease pathways, our method finds pathway sub-clusters which are significantly associated with the disease. As a result, we identified 8, 30, 63, 9, 29 pathways associated with Graves’ disease (GD), Insulin-like growth hormone factor I deficiency (IGFI), polycystic ovary syndrome (PCOS), type 1 diabetes mellitus (T1DM) and type 2 diabetes mellitus (T2DM), respectively. By illuminating these subnetworks derived from pathway network, our pathway prediction method provides a novel roadmap to explore new diagnostic and therapeutic opportunities across endocrine diseases.
Ningyi Zhang, Tianyi Zang
BIBM2
2021 Identifying drug-target interactions based on graph convolutional network and deep neural network
abstract
Identification of new drug-target interactions (DTIs) is an important but a time-consuming and costly step in drug discovery. In recent years, to mitigate these drawbacks, researchers have sought to identify DTIs using computational approaches. However, most existing methods construct drug networks and target networks separately, and then predict novel DTIs based on known associations between the drugs and targets without accounting for associations between drug-protein pairs (DPPs). To incorporate the associations between DPPs into DTI modeling, we built a DPP network based on multiple drugs and proteins in which DPPs are the nodes and the associations between DPPs are the edges of the network. We then propose a novel learning-based framework, 'graph convolutional network (GCN)-DTI', for DTI identification. The model first uses a graph convolutional network to learn the features for each DPP. Second, using the feature representation as an input, it uses a deep neural network to predict the final label. The results of our analysis show that the proposed framework outperforms some state-of-the-art approaches by a large margin.
Tianyi Zhao 0001, Yang Hu 0008, Linda R. Valsdottir, Tianyi Zang, Jiajie Peng
Briefings Bioinform.4
2021 Prediction and collection of protein-metabolite interactions
abstract
Interactions between proteins and small molecule metabolites play vital roles in regulating protein functions and controlling various cellular processes. The activities of metabolic enzymes, transcription factors, transporters and membrane receptors can all be mediated through protein-metabolite interactions (PMIs). Compared with the rich knowledge of protein-protein interactions, little is known about PMIs. To the best of our knowledge, no existing database has been developed for collecting PMIs. The recent rapid development of large-scale mass spectrometry analysis of biomolecules has led to the discovery of large amounts of PMIs. Therefore, we developed the PMI-DB to provide a comprehensive and accurate resource of PMIs. A total of 49 785 entries were manually collected in the PMI-DB, corresponding to 23 small molecule metabolites, 9631 proteins and 4 species. Unlike other databases that only provide positive samples, the PMI-DB provides non-interaction between proteins and metabolites, which not only reduces the experimental cost for biological experimenters but also facilitates the construction of more accurate algorithms for researchers using machine learning. To show the convenience of the PMI-DB, we developed a deep learning-based method to predict PMIs in the PMI-DB and compared it with several methods. The experimental results show that the area under the curve and area under the precision-recall curve of our method are 0.88 and 0.95, respectively. Overall, the PMI-DB provides a user-friendly interface for browsing the biological functions of metabolites/proteins of interest, and experimental techniques for identifying PMIs in different species, which provides important support for furthering the understanding of cellular processes. The PMI-DB is freely accessible at http://easybioai.com/PMIDB.
Tianyi Zhao 0001, Tianyi Zang, Jiajie Peng
Briefings Bioinform.6
2021 interacCircos: an R package based on JavaScript libraries for the generation of interactive circos plots
abstract
SUMMARY: JavaScript-based Circos libraries have been widely implemented to generate interactive Circos plots in web applications. However, these libraries require either local installation, which requires the compilation of extra libraries, or extra data processing procedures to prepare input and configuration for each track of plot, which limits the utility and capability of integration with powerful R packages. In this report, we present interacCircos, an R package for creating interactive Circos plots through the integration of JavaScript-based libraries. interacCircos can simply and flexibly implement 14 track-plot functions and 7 auxiliary functions for presenting large-scale genomic data in interactive Circos plots. AVAILABILITY AND IMPLEMENTATION: InteracCircos and its manual are freely available at https://github.com/mrcuizhe/interacCircos under the GPL license. The online documentation is available at https://mrcuizhe.github.io/interacCircos_documentation/index.html. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Ya Cui, Tianyi Zang
Bioinform.3
2021 SKSV: ultrafast structural variation detection from circular consensus sequencing reads
abstract
SUMMARY: Circular consensus sequencing reads are promising for the comprehensive detection of structural variants (SVs). However, alignment-based SV calling pipelines are computationally intensive due to the generation of complete read-alignments and its post-processing. Herein, we propose a SKeleton-based analysis toolkit for Structural Variation detection (SKSV). Benchmarks on real and simulated datasets demonstrate that SKSV has an order of magnitude of faster speed than state-of-the-art SV calling approaches; moreover, it achieves higher F1 scores for various types of SVs. AVAILABILITY AND IMPLEMENTATION: SKSV is available from https://github.com/ydLiu-HIT/SKSV. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yadong Liu 0001, Tao Jiang 0021, Junhao Su, Bo Liu 0023, Tianyi Zang, Yadong Wang 0001
Bioinform.5
2020 Assessment of Machine Learning Methods for Classification in Single Cell ATAC-seq
abstract
Single-cell assay for transposase accessible chromatin using sequencing(scATAC-seq) is rapidly advancing our understanding of the cellular composition of complex tissues and organisms. The similarity of data structure and feature between scRNA-seq and scATAC-seq makes it feasible to identify the cell types in scATAC-seq through traditional supervised machine learning methods. Here, we evaluated 6 popular machine learning methods for classification in scATAC-seq. The performance of the methods is evaluated using 4 public single cell ATAC-seq datasets of different tissues, sizes and technologies. We evaluated these methods using intradatasets experiments of 5-folds cross validation based on accuracy, recall and percentage of correctly predicted cells. We found that these methods may perform well in some types of cells in a single dataset, but the overall results are not as well as in scRNA-seq analysis. For testing the classification ability of machine learning methods across datasets, we applied inter-dataset experiments to test the performance of machine learning methods in realistic scenarios. SVM and NMC are overall the top 2 best-performing methods across all experiments. We recommend researchers to apply SVM and NMC as the underlying classifier when developing an automatic classification method in scATAC-seq.
Bo Liu 0023, Liran Juan, Tianyi Zang, Tao Jiang 0021, Yadong Wang 0001
BIBM4
2020 NCRR: A novel method for measuring disease similarity based on non-coding RNA regulation
abstract
Complex diseases are not simply caused by a single gene, single mRNA transcript or single protein but the effect of their collaborations. Measuring similarity between complex diseases plays an important role in understanding the mechanism of the diseases, which can also support identifying potential therapeutic drugs for diseases. With the rapid development of technology, it has been consentaneous that functional associations between disease-related genes and semantic associations can be applied to calculate disease similarity. Recent years, more and more studies have demonstrated a profound involvement of the non-coding RNA in the regulation of genome organization and gene expression. Non-coding RNA seem to operate at several biological levels such as epigenetic processes that control differentiation and development. Thus, taking noncoding RNA into account can be useful in measuring disease similarities. However, existing methods ignore the regulation functions of non-coding RNA in biological process. In this work, we proposed a novel method, NCRR (non-coding RNA regulated based similarity measurement), to measure disease similarity integrating functional associations between disease-related genes, semantic associations between diseases and similarities between disease-related non-coding RNAs. NCRR employs the Jaccard coefficient to measure the similarity between gene sets regulated by disease-related lncRNAs and disease-related miRNAs.
Ningyi Zhang, Liran Juan, Tianyi Zang
BIBM3
2020 CNN-DDI: A novel deep learning method for predicting drug-drug interactions
abstract
Predicting drug-drug interactions (DDIs) is one of the major concerns in patients' medication, which is crucial for patient safety and public health. Most of studies study whether drugs interact or not. In this study, we focus on 65 categories of drug-drug interaction-associated events and proposed a new method based on convolutional neural network (CNN), named CNN-DDI, for predicting DDIs. First, the categories, targets, pathways and enzymes of drugs were extracted as the features of drugs, which constructed the input of CNN-DDI. Then, these features were as input vectors of our CNN network, and the output is the prediction of drug-drug interaction-associated events' categories. In the computational experiments, the CNN-DDI method achieves an accuracy rate up to 0.8914, an area under the precision-recall curve up to 0.9322. And the experiments also prove using feature combinations outperforms one feature. Compared with other state-of-the-art methods, the CNN-DDI method has better performance in the superiority and the effectiveness for predicting DDI's events.
Tianyi Zang
BIBM2
2020 DRACP: a novel method for identification of anticancer peptides
abstract
BACKGROUND: Millions of people are suffering from cancers, but accurate early diagnosis and effective treatment are still tough for all doctors. Common ways against cancer include surgical operation, radiotherapy and chemotherapy. However, they are all very harmful for patients. Recently, the anticancer peptides (ACPs) have been discovered to be a potential way to treat cancer. Since ACPs are natural biologics, they are safer than other methods. However, the experimental technology is an expensive way to find ACPs so we purpose a new machine learning method to identify the ACPs. RESULTS: Firstly, we extracted the feature of ACPs in two aspects: sequence and chemical characteristics of amino acids. For sequence, average 20 amino acids composition was extracted. For chemical characteristics, we classified amino acids into six groups based on the patterns of hydrophobic and hydrophilic residues. Then, deep belief network has been used to encode the features of ACPs. Finally, we purposed Random Relevance Vector Machines to identify the true ACPs. We call this method 'DRACP' and tested the performance of it on two independent datasets. Its AUC and AUPR are higher than 0.9 in both datasets. CONCLUSION: We developed a novel method named 'DRACP' and compared it with some traditional methods. The cross-validation results showed its effectiveness in identifying ACPs.
Tianyi Zhao 0001, Yang Hu 0008, Tianyi Zang
BMC Bioinform.3
2019 Identification of anticancer peptides based on Random Relevance Vector Machines
abstract
Cancer is the most threat to human's health and life. At present, people have developed several ways to against cancer, such as surgical operation, radiotherapy and chemotherapy. However, cancers still cause highly mortality rate. A main part of the reason is that the traditional methods bring treatment effect as well as the negative effect. Recently, the anticancer peptides (ACPs) have been discovered which can be a new way to treat cancer. Since ACPs are natural biologics, they are safer than other methods. However, the experimental technology is an expensive way to find ACPs so we purpose a new machine learning method to identify the ACPs which named Random Relevance Vector Machines (RRVMs). The cross validations experiments show the high accuracy and stability of this new method.
Tianyi Zhao 0001, Tianyi Zang, Yang Hu 0008
BIBM3
2019 Prioritizing candidate diseases-related metabolites based on literature and functional similarity
abstract
BACKGROUND: As the terminal products of cellular regulatory process, functional related metabolites have a close relationship with complex diseases, and are often associated with the same or similar diseases. Therefore, identification of disease related metabolites play a critical role in understanding comprehensively pathogenesis of disease, aiming at improving the clinical medicine. Considering that a large number of metabolic markers of diseases need to be explored, we propose a computational model to identify potential disease-related metabolites based on functional relationships and scores of referred literatures between metabolites. First, obtaining associations between metabolites and diseases from the Human Metabolome database, we calculate the similarities of metabolites based on modified recommendation strategy of collaborative filtering utilizing the similarities between diseases. Next, a disease-associated metabolite network (DMN) is built with similarities between metabolites as weight. To improve the ability of identifying disease-related metabolites, we introduce scores of text mining from the existing database of chemicals and proteins into DMN and build a new disease-associated metabolite network (FLDMN) by fusing functional associations and scores of literatures. Finally, we utilize random walking with restart (RWR) in this network to predict candidate metabolites related to diseases. RESULTS: We construct the disease-associated metabolite network and its improved network (FLDMN) with 245 diseases, 587 metabolites and 28,715 disease-metabolite associations. Subsequently, we extract training sets and testing sets from two different versions of the Human Metabolome database and assess the performance of DMN and FLDMN on 19 diseases, respectively. As a result, the average AUC (area under the receiver operating characteristic curve) of DMN is 64.35%. As a further improved network, FLDMN is proven to be successful in predicting potential metabolic signatures for 19 diseases with an average AUC value of 76.03%. CONCLUSION: In this paper, a computational model is proposed for exploring metabolite-disease pairs and has good performance in predicting potential metabolites related to diseases through adequate validation. This result suggests that integrating literature and functional associations can be an effective way to construct disease associated metabolite network for prioritizing candidate diseases-related metabolites.
Yongtian Wang, Liran Juan, Jiajie Peng, Tianyi Zang, Yadong Wang 0001
BMC Bioinform.4
2019 LncDisAP: a computation model for LncRNA-disease association prediction based on multiple biological datasets
abstract
BACKGROUND: Over the past decades, a large number of long non-coding RNAs (lncRNAs) have been identified. Growing evidence has indicated that the mutation and dysregulation of lncRNAs play a critical role in the development of many complex human diseases. Consequently, identifying potential disease-related lncRNAs is an effective means to improve the quality of disease diagnostics and treatment, which is the motivation of this work. Here, we propose a computational model (LncDisAP) for potential disease-related lncRNA identification based on multiple biological datasets. First, the associations between lncRNA and different data sources are collected from different databases. With these data sources as dimensions, we calculate the functional associations between lncRNAs by the recommendation strategy of collaborative filtering. Subsequently, a disease-associated lncRNA functional network is built with functional similarities between lncRNAs as the weight. Ultimately, potential disease-related lncRNAs can be identified based on ranked scores derived by random walking with restart (RWR). Then, training sets and testing sets are extracted from two different versions of a disease-lncRNA dataset to assess the performance of LncDisAP on 54 diseases. RESULTS: A lncRNA functional network is built based on the proposed computational model, and it contains 66,060 associations among 364 lncRNAs associated with 182 diseases in total. We extract 218 known disease-lncRNA pairs associated with 54 diseases to assess the network. As a result, the average AUC (area under the receiver operating characteristic curve) of LncDisAP is 78.08%. CONCLUSION: In this article, a computational model integrating multiple lncRNA-related biological datasets is proposed for identifying potential disease-related lncRNAs. The result shows that LncDisAP is successful in predicting novel disease-related lncRNA signatures. In addition, with several common cancers taken as case studies, we found some unknown lncRNAs that could be associated with these diseases through our network. These results suggest that this method can be helpful in improving the quality for disease diagnostics and treatment.
Yongtian Wang, Liran Juan, Jiajie Peng, Tianyi Zang, Yadong Wang 0001
BMC Bioinform.4
2019 Identifying Alzheimer's disease-related proteins by LRRGD
abstract
BACKGROUND: Alzheimer's disease (AD) imposes a heavy burden on society and every family. Therefore, diagnosing AD in advance and discovering new drug targets are crucial, while these could be achieved by identifying AD-related proteins. The time-consuming and money-costing biological experiment makes researchers turn to develop more advanced algorithms to identify AD-related proteins. RESULTS: Firstly, we proposed a hypothesis "similar diseases share similar related proteins". Therefore, five similarity calculation methods are introduced to find out others diseases which are similar to AD. Then, these diseases' related proteins could be obtained by public data set. Finally, these proteins are features of each disease and could be used to map their similarity to AD. We developed a novel method 'LRRGD' which combines Logistic Regression (LR) and Gradient Descent (GD) and borrows the idea of Random Forest (RF). LR is introduced to regress features to similarities. Borrowing the idea of RF, hundreds of LR models have been built by randomly selecting 40 features (proteins) each time. Here, GD is introduced to find out the optimal result. To avoid the drawback of local optimal solution, a good initial value is selected by some known AD-related proteins. Finally, 376 proteins are found to be related to AD. CONCLUSION: Three hundred eight of three hundred seventy-six proteins are the novel proteins. Three case studies are done to prove our method's effectiveness. These 308 proteins could give researchers a basis to do biological experiments to help treatment and diagnostic AD.
Tianyi Zhao 0001, Yang Hu 0008, Tianyi Zang, Liang Cheng 0006
BMC Bioinform.3
2018 A Novel Method for Identifying Alzheimer's Disease-related Proteins
Yang Hu 0008, Jun Zhang 0041, Tianyi Zhao 0001, Liang Cheng 0006, Tianyi Zang
BIBM5
2018 DeepDNA: a hybrid convolutional and recurrent neural network for compressing human mitochondrial genomes
Yan-Shuo Chu, Yongtian Wang, Mingrui Sun, Junyi Li 0004, Tianyi Zang, Yadong Wang 0001
BIBM8
2018 Identifying Candidate Diseases-related Metabolites Based on Disease Similarity
Yongtian Wang, Liran Juan, Chunpu Liu, Tianyi Zang
BIBM4
2018 Predicting candidate disease-related lncRNAs based on network random walk
Yongtian Wang, Liran Juan, Jiajie Peng, Tianyi Zang, Yadong Wang 0001
BIBM4
2018 Analysis for Early Seizure Detection System Based on Deep Learning Algorithm
Fuxu Wang, Mingrui Sun, Tengfei Min, Yueying Wang, Chunpu Liu, Tianyi Zang
BIBM6
2018 Identifying diseases-related metabolites using random walk
abstract
BACKGROUND: Metabolites disrupted by abnormal state of human body are deemed as the effect of diseases. In comparison with the cause of diseases like genes, these markers are easier to be captured for the prevention and diagnosis of metabolic diseases. Currently, a large number of metabolic markers of diseases need to be explored, which drive us to do this work. METHODS: The existing metabolite-disease associations were extracted from Human Metabolome Database (HMDB) using a text mining tool NCBO annotator as priori knowledge. Next we calculated the similarity of a pair-wise metabolites based on the similarity of disease sets of them. Then, all the similarities of metabolite pairs were utilized for constructing a weighted metabolite association network (WMAN). Subsequently, the network was utilized for predicting novel metabolic markers of diseases using random walk. RESULTS: Totally, 604 metabolites and 228 diseases were extracted from HMDB. From 604 metabolites, 453 metabolites are selected to construct the WMAN, where each metabolite is deemed as a node, and the similarity of two metabolites as the weight of the edge linking them. The performance of the network is validated using the leave one out method. As a result, the high area under the receiver operating characteristic curve (AUC) (0.7048) is achieved. The further case studies for identifying novel metabolites of diabetes mellitus were validated in the recent studies. CONCLUSION: In this paper, we presented a novel method for prioritizing metabolite-disease pairs. The superior performance validates its reliability for exploring novel metabolic markers of diseases.
Yang Hu 0008, Tianyi Zhao 0001, Ningyi Zhang, Tianyi Zang, Jun Zhang 0041, Liang Cheng 0006
BMC Bioinform.4
2017 A bucket index correction based method for compression of genomic sequencing data
abstract
As high-throughput sequencing technologies are generating vast amounts of data, there is urgent need to develop efficient algorithms for sequencing data compression. Existing methods usually dispatch the similar sequences into the same bucket based on their same minimizer, that is the lexicographical smallest k-mer within the sequence, for data compression. However, when the sequencing error existed in the minimizer area, it could cause sequences to be distributed into the improper buckets, which could result in a negative effect in the following compression process. In this paper, we propose a novel method BIC, a bucket index correction method for sequencing data compression. BIC is the first method to correct sequencing errors in minimizer area, which dispatches more similar sequences into the same buckets, that could effectively compress sequencing data. Compared with three state-of-the-art methods on five different data sets, BIC could reach more compression rate. The codes of BIC are available at https://github.com/rongjiewang/BIC.
Qianlong Cheng, Tianyi Zang
BIBM4
2017 FNSemSim: An improved disease similarity method based on network fusion
abstract
Discovering similar diseases is very helpful for revealing the pathogenesis of diseases and making direction in drug use. And related diseases are often triggered by disease-related genes. Therefore, function interaction networks structured by disease-related genes are suitable for measurement of disease similarity, and some methods have utilized the advantage of function interaction of disease-related genes. However, all of them were developed by using a single gene functional network, some of them ignoring the effect of non-neighbour nodes in a functional interaction network. In this study, we propose a new method, FNSemSim, for computing relatedness between diseases by fusing two protein networks, which could be utilized fully based on random walk with restart (RWR). And a benchmark set of similar disease pairs are used to assess the performance of FNSemSim. As a result, FNSemSim achieved a very good performance with a high AUC (area under the receiver operating characteristic curve) reached 98.7%. Furthermore, we further studied the impact of different data sources, including function interaction networks and disease-related genes databases. It was found that the quality of the data sources has a greater impact on the performance of disease similarity calculation than the size of the data source, and utilizing function interaction networks and gene-disease association data could improve the performance of FNSemSim.
Yongtian Wang, Liran Juan, Yan-Shuo Chu, Tianyi Zang
BIBM5
2016 DMcompress: Dynamic Markov models for bacterial genome compression
abstract
Genome data increasing exponentially since the last decade, compressing genome with Markov models has been proposed as an effective statistical method. However, existing methods set a static order-k Markov models to compress various genomes. Employing static order-k Markov model could result in a sub-optimal orders on some genomes. In this paper, we propose a compression method that relies on a pre-analysis of the data before compression, with the aim of estimating Markov models order k, yielding improvements over static Markov models. Experimental results on the latest complete bacterial genome data show that our method could effectively compress genome with a better performance than the state-of-the-art method. The codes of DMcompress are available at https://rongjiewang.github.io/DMcompress.
Mingxiang Teng, Tianyi Zang, Yadong Wang 0001
BIBM4
2015 SS4CSHC: A Services System for the Collaboration in Stroke Healthcare Cycle
abstract
Stroke remains the third leading cause of death. The rising of stroke illness and the continuous aging of the global population requires a collaboration of stroke services systems based on relations and exchange of information to address needs of caregiver and patients in the healthcare community. The collaborative stroke services system involves interconnected changes and the development of integrated healthcare information systems and novel healthcare services. We present the SS4CSHC, a services system for the collaboration in stroke healthcare cycle, to contribute to information sharing and cooperation for the stroke surveillance, prevention, emergency therapy, diagnosis, treatment, and rehabilitation. Some of critical and novel stroke healthcare services are developed and deployed, and innovative models and services for delivery and the coordination of stroke healthcare are explored by the stroke healthcare center. The feasibility and efficiency of our services systems are validated in the local medical institutes.
Mingrui Sun, Di Dai, Shiquan Wang, Tianyi Zang, Xiaofei Xu 0001
ICSS5
2015 Family genome browser: visualizing genomes with pedigree information
abstract
MOTIVATION: Families with inherited diseases are widely used in Mendelian/complex disease studies. Owing to the advances in high-throughput sequencing technologies, family genome sequencing becomes more and more prevalent. Visualizing family genomes can greatly facilitate human genetics studies and personalized medicine. However, due to the complex genetic relationships and high similarities among genomes of consanguineous family members, family genomes are difficult to be visualized in traditional genome visualization framework. How to visualize the family genome variants and their functions with integrated pedigree information remains a critical challenge. RESULTS: We developed the Family Genome Browser (FGB) to provide comprehensive analysis and visualization for family genomes. The FGB can visualize family genomes in both individual level and variant level effectively, through integrating genome data with pedigree information. Family genome analysis, including determination of parental origin of the variants, detection of de novo mutations, identification of potential recombination events and identical-by-decent segments, etc., can be performed flexibly. Diverse annotations for the family genome variants, such as dbSNP memberships, linkage disequilibriums, genes, variant effects, potential phenotypes, etc., are illustrated as well. Moreover, the FGB can automatically search de novo mutations and compound heterozygous variants for a selected individual, and guide investigators to find high-risk genes with flexible navigation options. These features enable users to investigate and understand family genomes intuitively and systematically. AVAILABILITY AND IMPLEMENTATION: The FGB is available at http://mlg.hit.edu.cn/FGB/.
Liran Juan, Yongzhuang Liu, Yongtian Wang, Mingxiang Teng, Tianyi Zang, Yadong Wang 0001
Bioinform.5
2010 Cooperative Work Systems for the Security of Digital Computing Infrastructure
abstract
On open digital computing infrastructure, various large-scale and complicated malicious behaviors are increasingly threatening the security of digital computing infrastructure. In this paper, a Cooperative Work Model (CRM) is presented by extending the conceptions of the Universal Turing Machine to deal with the threats. Then the Cooperative Work System Framework (CWSF) is derived from the model. Based on the framework, two practical Cooperative Work Systems (CWSs) are developed to track and analyze the Botnet and DDoS on digital computing infrastructure respectively. The systems collectively use and coordinate various monitoring systems distributed in the back-bone network of the infrastructure. The experimental results of analyzing typical security events show that the framework and systems are efficient and effective to collaboratively use diverse related network systems for monitoring and analyzing the large-scale network events. Currently, the systems are running steadily in the monitoring environment of a large-scale back-bone network.
Tianning Zang, Xiao-chun Yun, Tianyi Zang, Yongzheng Zhang 0002, Chaoguang Men
ICPADS3
2010 Comprehensive Practice on Service Engineering: An Experimental Solution
abstract
To develop a well-designed curriculum of service engineering is a prerequisite for training students get new skills and ability to design and develop globally effective solutions in a service business environment. As service engineering is an application-oriented discipline, sufficient practice is extremely important. In this paper, we show an experimental solution of a course "comprehensive practice on service engineering", which has been explored for three years in Harbin Institute of Technology. Its objective, process, execution program, and detailed design of each phase, are elaborately presented. Lessons learned from the explorations and future improvement on this solution, are briefly discussed.
Zhongjie Wang 0003, Xiaofei Xu 0001, Lanshun Nie, Tianyi Zang
ICSS5
2008 WSRF-Based Modeling of Clinical Trial Information for Collaborative Cancer Research
abstract
The CancerGrid consortium is developing open- standards cancer informatics to address the challenges posed by modern cancer clinical trials. This paper presents the service-oriented software paradigm implemented in CancerGrid to derive clinical trial information management systems for collaborative cancer research across multiple institutions. Our proposal is founded on a combination of a clinical trial (meta)model and WSRF (Web Services Resource Framework), and is currently being evaluated for use in early phase trials. Although primarily targeted at cancer research, our approach is readily applicable to other areas for which a similar information model is available.
Tianyi Zang, Radu Calinescu, Steve Harris, Andrew Tsui, Marta Z. Kwiatkowska, Jeremy Gibbons, Jim Davies, Peter Maccallum, Carlos Caldas
CCGRID1
2008 Metamodel-Based Generation of WSRF-Compliant SOA for Collaborative Cancer Research
abstract
Cancer clinical trials pose significant challenges to the e-Science community. The information technology required to enable this kind of large-scale, collaborative science will need to support easy and rapid development and deployment of reliable and flexible software systems that enable syntactic, semantic and computational interoperability. CancerGrid, an e-Science consortium funded by the UK Medical Research Council, is addressing these challenges through the development of model-driven, service-oriented technology for cancer informatics. This poster presents recent significant efforts in CancerGrid, resulting in the metamodel-based automated generation of WSRF (Web Services Resource Framework) compliant trial management systems. The most important advantages of our approach are discussed.
Tianyi Zang, Radu Calinescu, Steve Harris, Andrew Tsui, Charles Crichton, Marta Z. Kwiatkowska, Jeremy Gibbons, Jim Davies, James D. Brenton, Carlos Caldas
eScience1
2006 GRASG - A Framework for "Gridifying" and Running Applications on Service-Oriented Grids
abstract
The convergence of grid computing technologies and Web services offers many opportunities to utilize resources distributed across the Internet and solves many issues of interoperability. As a result, enabling applications as Web services are required intensively. Hence, a framework for "gridifying" and running applications on service-oriented grids (GRASG) was built to offer developers a flexible and effective tool for "gridifying" applications and making use of distributed resources on grid environment without much effort from the developers. It allows users to quickly enable an application as a Web service and access this service in a simple fashion. Further, in order to make use of distributed resources, GRASG provides a metascheduling mechanism that is able to schedule jobs to grid resources using Web services protocol. These features reduce the time taken for application development and execution.
Quoc-Thuan Ho, Terence Hung, Wei Jie, Hoong-Maeng Chan, Sindhu Emilda, Subramaniam Ganesan, Tianyi Zang, Xiaorong Li
CCGRID7
2004 The Design and Implementation of An OGSA-based Grid Information Service
abstract
The information service is a key component of a grid environment and critical to the operation of a computational grid. In this work, an OGSA (Open Grid Services Architecture) based information service that complies with OGSI (Open Grid Services Infrastructure) is presented. The main functionality of this information service is the provision of information essential for applications running on a computational grid such as resource information, job status, resource workload, service meta-information, and queue status. This OGSI-compliant information service is built on Globus Toolkit MDS-3, and it works with meta-scheduling services and local job scheduling systems to support resource discovery, job scheduling, and execution management. In this paper, the architecture of the Information Service and the models of information data organization are presented. Some implementation issues are discussed as well.
Tianyi Zang, Wei Jie, Terence Hung, Stephen John Turner, Wentong Cai 0001
ICWS1