Taijiao Jiang

dblp:08/5867 · DBLP profile ↗
← Back
31ranked-venue papers
1as first author
16since 2021 · last 2026
0000-0002-6280-6347ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 28 · 1 first-author · 14 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Knowledge from medical ontology can significantly enhance mainstream text embedding models in medical information retrieval
Lizong Deng, Yifan Qi, Chunli Shao, Taijiao Jiang
Inf. Process. Manag.7
2025 TransSSVs: a Transformer-based deep learning model for accurate detection of somatic small variants in paired tumor and normal sequencing data
abstract
Accurate identification of somatic small variants in tumors plays a crucial role in cancer diagnosis. Various somatic mutation callers have been developed; however, existing methods face limitations in modeling mapping information of flanking genomic sites (genomic sites adjacent to a somatic site) that influence the state of the somatic site. Additionally, they are unable to analyze inter-site interactions within the context sequence centered around a somatic site or appropriately weigh the effects of flanking genomic sites on the somatic site. To address these limitations, the Transformer model is utilized to develop TransSSVs for detecting somatic small variants. The core functionality of TransSSVs relies on the multi-head attention mechanism, which generates a reliable representation of interactions between a candidate somatic site and its flanking genomic sites within the context sequence. TransSSVs effectively extract mapping features of various genomic sites in the context sequence to enhance prediction accuracy. Benchmarking experiments demonstrate that TransSSVs exhibit robust performance when compared with state-of-the-art methods on well-characterized real and simulated tumor datasets. Furthermore, the contributions of flanking genomic sites to the detection of somatic sites are assessed, and attention weight patterns for positive and negative somatic sites are analyzed.
Jiangyuan Wang, Jingze Liu, Wenkai Song, Aiping Wu 0002, Taijiao Jiang
Appl. Intell.7
2025 PREDAC-FluB: predicting antigenic clusters of seasonal influenza B viruses with protein language model embedding based convolutional neural network
abstract
Influenza poses a significant global public health threat, with vaccination being the most effective and economical preventive measure. However, these punctuated antigenic changes, particularly in HA, result in escape from the immunity that was induced by prior infection or vaccination. Accurately predicting antigenic variation and understanding the antigenic dynamics of influenza viruses are crucial for selecting appropriate vaccine strains, but no established methods exist for influenza B viruses. Therefore, we present PREDAC-FluB, a hybrid deep learning framework that integrates spatial feature extraction via CNN to model interactions in HA1 sequences, multimodal sequence representation combining ESM-2 embeddings with six physicochemical descriptors and continuous encoding (ESM2-7-features), and UMAP-guided clustering for antigenic cluster identification. Using data from 9036 B/Victoria-lineage and 4520 B/Yamagata-lineage influenza virus pair. PREDAC-FluB demonstrates superior performance over traditional machine learning methods in predicting antigenic variation in influenza viruses, successfully identifying major antigenic clusters. Specifically, PREDAC-FluB classified the B/Victoria lineage into nine antigenic clusters and the B/Yamagata lineage into three antigenic clusters. In five-fold cross-validation for B/Victoria viruses, PREDAC-FluB with ESM2-7-features encoding achieved AUROC values of 0.9961 on the validation set and 0.9856 on the independent test set. In retrospective testing for B/Victoria viruses, PREDAC-FluB achieved AUROC values ranging from 0.83 to 0.97, demonstrating high prediction accuracy and effectively capturing antigenic variation information. In conclusion, PREDAC-FluB is a robust tool for antigenic computation, capable of accurately predicting antigenic variation in influenza B viruses. Its high prediction accuracy makes it a promising auxiliary method for recommending future influenza vaccine strains.
Wenping Xie, Jingze Liu, Jiangyuan Wang, Wenjie Han, Yousong Peng, Xiangjun Du, Kang Ning 0001, Taijiao Jiang
Briefings Bioinform.10
2025 PhenoAlign: A Hybrid Data-Knowledge-Driven Approach for Precisely Aligning Phenotype Information in Medical Texts
abstract
Precisely aligning phenotypic information within medical texts is paramount in advancing intelligent medical applications, such as similar patient case retrieval. However, despite its criticality, an algorithm specifically designed for this task is lacking. We previously introduced a fine-grained semantic information model, the semantic structured unit of phenotypes (PhenoSSU), and an automatic extraction algorithm. This model accurately characterizes and extracts phenotypic information from medical texts. In this study, we explore different PhenoSSU alignment strategies. The results show that the data-knowledge-driven approach best aligns PhenoSSUs. Specifically, employing a BERT-based pre-trained language model (PLMs) to align phrase-type PhenoSSUs and a knowledge-based method for logic-type PhenoSSUs demonstrates efficacy. Moreover, by successfully integrating the PhenoSSU alignment and extraction algorithms, we have developed PhenoAlign, a novel medical text phenotype alignment tool. This tool facilitates precisely aligning phenotypic information by processing two medical texts, generating accurate alignment outcomes. PhenoAlign exhibited satisfactory medical text phenotypic alignment using the expert annotated gold standard test set, with an end-to-end F1 score of 0.820. The F1 scores for phenoSSU extraction and alignment were 0.885 and 0.927, respectively. Our analysis extends to the potential application of ChatGPT in phenotype alignment tasks, and find significant challenges encountered by large language models in this domain. We developed a simple, effective tool for medical text phenotype information alignment. This tool will be valuable to intelligent medical applications, facilitating patient care and medical research advancements.
Lizong Deng, Taijiao Jiang
IEEE J. Biomed. Health Informatics4
2024 PREDAC-CNN: predicting antigenic clusters of seasonal influenza A viruses with convolutional neural network
abstract
Vaccination stands as the most effective and economical strategy for prevention and control of influenza. The primary target of neutralizing antibodies is the surface antigen hemagglutinin (HA). However, ongoing mutations in the HA sequence result in antigenic drift. The success of a vaccine is contingent on its antigenic congruence with circulating strains. Thus, predicting antigenic variants and deducing antigenic clusters of influenza viruses are pivotal for recommendation of vaccine strains. The antigenicity of influenza A viruses is determined by the interplay of amino acids in the HA1 sequence. In this study, we exploit the ability of convolutional neural networks (CNNs) to extract spatial feature representations in the convolutional layers, which can discern interactions between amino acid sites. We introduce PREDAC-CNN, a model designed to track antigenic evolution of seasonal influenza A viruses. Accessible at http://predac-cnn.cloudna.cn, PREDAC-CNN formulates a spatially oriented representation of the HA1 sequence, optimized for the convolutional framework. It effectively probes interactions among amino acid sites in the HA1 sequence. Also, PREDAC-CNN focuses exclusively on physicochemical attributes crucial for the antigenicity of influenza viruses, thereby eliminating unnecessary amino acid embeddings. Together, PREDAC-CNN is adept at capturing interactions of amino acid sites within the HA1 sequence and examining the collective impact of point mutations on antigenic variation. Through 5-fold cross-validation and retrospective testing, PREDAC-CNN has shown superior performance in predicting antigenic variants compared to its counterparts. Additionally, PREDAC-CNN has been instrumental in identifying predominant antigenic clusters for A/H3N2 (1968-2023) and A/H1N1 (1977-2023) viruses, significantly aiding in vaccine strain recommendation.
Jingze Liu, Wenkai Song, Honglei Li 0002, Jiangyuan Wang, Le Zhang 0004, Yousong Peng, Aiping Wu 0002, Taijiao Jiang
Briefings Bioinform.9
2023 TeaBERT: An Efficient Knowledge Infused Cross-Lingual Language Model for Mapping Chinese Medical Entities to the Unified Medical Language System
abstract
Medical entity normalization is an important task for medical information processing. The Unified Medical Language System (UMLS), a well-developed medical terminology system, is crucial for medical entity normalization. However, the UMLS primarily consists of English medical terms. For languages other than English, such as Chinese, a significant challenge for normalizing medical entities is the lack of robust terminology systems. To address this issue, we propose a translation-enhancing training strategy that incorporates the translation and synonym knowledge of the UMLS into a language model using the contrastive learning approach. In this work, we proposed a cross-lingual pre-trained language model called TeaBERT, which can align synonymous Chinese and English medical entities across languages at the concept level. As the evaluation results showed, the TeaBERT language model outperformed previous cross-lingual language models with Acc@5 values of 92.54%, 87.14% and 84.77% on the ICD10-CN, CHPO and RealWorld-v2 datasets, respectively. It also achieved a new state-of-the-art cross-lingual entity mapping performance without fine-tuning. The translation-enhancing strategy is applicable to other languages that face the similar challenge due to the absence of well-developed medical terminology systems.
Yifan Qi, Aiping Wu 0002, Lizong Deng, Taijiao Jiang
IEEE J. Biomed. Health Informatics5
2022 vsRNAfinder: a novel method for identifying high-confidence viral small RNAs from small RNA-Seq data
abstract
Virus-encoded small RNAs (vsRNA) have been reported to play an important role in viral infection. Unfortunately, there is still a lack of an effective method for vsRNA identification. Herein, we presented vsRNAfinder, a de novo method for identifying high-confidence vsRNAs from small RNA-Seq (sRNA-Seq) data based on peak calling and Poisson distribution and is publicly available at https://github.com/ZenaCai/vsRNAfinder. vsRNAfinder outperformed two widely used methods namely miRDeep2 and ShortStack in identifying viral miRNAs with a significantly improved sensitivity. It can also be used to identify sRNAs in animals and plants with similar performance to miRDeep2 and ShortStack. vsRNAfinder would greatly facilitate effective identification of vsRNAs from sRNA-Seq data.
Zena Cai, Ye Qiu, Aiping Wu 0002, Gaihua Zhang, Taijiao Jiang, Xing-Yi Ge, Haizhen Zhu, Yousong Peng
Briefings Bioinform.7
2022 An atlas of human viruses provides new insights into diversity and tissue tropism of human viruses
abstract
MOTIVATION: Viruses continue to threaten human health. Yet, the complete viral species carried by humans and their infection characteristics have not been fully revealed. RESULTS: This study curated an atlas of human viruses from public databases and literature, and built the Human Virus Database (HVD). The HVD contains 1131 virus species of 54 viral families which were more than twice the number of the human-infecting virus species reported in previous studies. These viruses were identified in human samples including 68 human tissues, the excreta and body fluid. The viral diversity in humans was age-dependent with a peak in the infant and a valley in the teenager. The tissue tropism of viruses was found to be associated with several factors including the viral group (DNA, RNA or reverse-transcribing viruses), enveloped or not, viral genome length and GC content, viral receptors and the virus-interacting proteins. Finally, the tissue tropism of DNA viruses was predicted using a random-forest algorithm with a middle performance. Overall, the study not only provides a valuable resource for further studies of human viruses but also deepens our understanding toward the diversity and tissue tropism of human viruses. AVAILABILITY AND IMPLEMENTATION: The HVD is available at http://computationalbiology.cn/humanVirusBase/#/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Sifan Ye, Congyu Lu, Ye Qiu, Heping Zheng, Xingyi Ge, Aiping Wu 0002, Zanxian Xia, Taijiao Jiang, Haizhen Zhu, Yousong Peng
Bioinform.8
2022 PIAT: An Evolutionarily Intelligent System for Deep Phenotyping of Chinese Electronic Health Records
abstract
Electronic health record (EHR) resources are valuable but remain underexplored because most clinical information, especially phenotype information, is buried in the free text of EHRs. An intelligent annotation tool plays an important role in unlocking the full potential of EHRs by transforming free-text phenotype information into a computer-readable form. Deep phenotyping has shown its advantage in representing phenotype information in EHRs with high fidelity; however, most existing annotation tools are not suitable for the deep phenotyping task. Here, we developed an intelligent annotation tool named PIAT with a major focus on the deep phenotyping of Chinese EHRs. PIAT can improve the annotation efficiency for EHR-based deep phenotyping with a simple but effective interactive interface, automatic preannotation support, and a learning mechanism. Specifically, experts can proofread automatic annotation results from the annotation algorithm in the web-based interactive interface, and EHRs reviewed by experts can be used for evolving the underlying annotation algorithm. In this way, the annotation process of deep phenotyping EHRs will become easier. In conclusion, we create a powerful intelligent system for the deep phenotyping of Chinese EHRs. It is hoped that our work will inspire further studies in constructing intelligent systems for deep phenotyping English and non-English EHRs.
Lizong Deng, Taijiao Jiang
IEEE J. Biomed. Health Informatics6
2021 VirusCircBase: a database of virus circular RNAs
abstract
Circular RNAs (circRNAs) are covalently closed long noncoding RNAs critical in diverse cellular activities and multiple human diseases. Several cancer-related viral circRNAs have been identified in double-stranded DNA viruses (dsDNA), yet no systematic study about the viral circRNAs has been reported. Herein, we have performed a systematic survey of 11 924 circRNAs from 23 viral species by computational prediction of viral circRNAs from viral-infection-related RNA sequencing data. Besides the dsDNA viruses, our study has also revealed lots of circRNAs in single-stranded RNA viruses and retro-transcribing viruses, such as the Zika virus, the Influenza A virus, the Zaire ebolavirus, and the Human immunodeficiency virus 1. Most viral circRNAs had reverse complementary sequences or repeated sequences at the flanking sequences of the back-splice sites. Most viral circRNAs only expressed in a specific cell line or tissue in a specific species. Functional enrichment analysis indicated that the viral circRNAs from dsDNA viruses were involved in KEGG pathways associated with cancer. All viral circRNAs presented in the current study were stored and organized in VirusCircBase, which is freely available at http://www.computationalbiology.cn/ViruscircBase/home.html and is the first virus circRNA database. VirusCircBase forms the fundamental atlas for the further exploration and investigation of viral circRNAs in the context of public health.
Zena Cai, Yunshi Fan, Congyu Lu, Zhaozhong Zhu, Taijiao Jiang, Tongling Shan, Yousong Peng
Briefings Bioinform.6
2021 Identification and characterization of circRNAs encoded by MERS-CoV, SARS-CoV-1 and SARS-CoV-2
abstract
The life-threatening coronaviruses MERS-CoV, SARS-CoV-1 and SARS-CoV-2 (SARS-CoV-1/2) have caused and will continue to cause enormous morbidity and mortality to humans. Virus-encoded noncoding RNAs are poorly understood in coronaviruses. Data mining of viral-infection-related RNA-sequencing data has resulted in the identification of 28 754, 720 and 3437 circRNAs encoded by MERS-CoV, SARS-CoV-1 and SARS-CoV-2, respectively. MERS-CoV exhibits much more prominent ability to encode circRNAs in all genomic regions than those of SARS-CoV-1/2. Viral circRNAs typically exhibit low expression levels. Moreover, majority of the viral circRNAs exhibit expressions only in the late stage of viral infection. Analysis of the competitive interactions of viral circRNAs, human miRNAs and mRNAs in MERS-CoV infections reveals that viral circRNAs up-regulated genes related to mRNA splicing and processing in the early stage of viral infection, and regulated genes involved in diverse functions including cancer, metabolism, autophagy, viral infection in the late stage of viral infection. Similar analysis in SARS-CoV-2 infections reveals that its viral circRNAs down-regulated genes associated with metabolic processes of cholesterol, alcohol, fatty acid and up-regulated genes associated with cellular responses to oxidative stress in the late stage of viral infection. A few genes regulated by viral circRNAs from both MERS-CoV and SARS-CoV-2 were enriched in several biological processes such as response to reactive oxygen and centrosome localization. This study provides the first glimpse into viral circRNAs in three deadly coronaviruses and would serve as a valuable resource for further studies of circRNAs in coronaviruses.
Zena Cai, Congyu Lu, Yuanqiang Zou, Zhaozhong Zhu, Xingyi Ge, Aiping Wu 0002, Taijiao Jiang, Heping Zheng, Yousong Peng
Briefings Bioinform.10
2021 Comparative viromes of Culicoides and mosquitoes reveal their consistency and diversity in viral profiles
abstract
The genus Culicoides includes biting midges, some of which are vectors for viruses that cause diseases in humans and animals. Knowledge of the roles of Culicoides in viral ecology is inadequate. We collected ~300 000 samples of Culicoides and mosquitoes in 15 representative regions within Yunnan, China. Using mosquitoes as reference vectors, we designed a comparative virome strategy to study the viral composition, diversity, hosts and spatiotemporal distribution of Culicoides. A map of viromes in Culicoides and mosquitoes in Yunan province, China, was constructed. At the same locations, Culicoides and mosquitoes usually share a similar viral diversity. At least 10 important pathogenic viruses were detected from Culicoides. Many novel viruses were discovered, including 21 segmented viruses of Flaviviridae, 180 viruses of Monjiviricetes and 130 viruses of Bunyavirales. The findings demonstrate that Culicoides is an important part of viral ecology and should be studied and monitored for potentially emerging viruses.
Qin Shen, Yuwen He, Na Han, Xianyue Wang, Jinxin Meng, Yousong Peng, Mei Pan, Yuting Jin, Taijiao Jiang, Wenjie Tan, Jinglin Wang, Aiping Wu 0002
Briefings Bioinform.11
2021 DeepSSV: detecting somatic small variants in paired tumor and normal sequencing data with convolutional neural network
abstract
It is of considerable interest to detect somatic mutations in paired tumor and normal sequencing data. A number of callers that are based on statistical or machine learning approaches have been developed to detect somatic small variants. However, they take into consideration only limited information about the reference and potential variant allele in both tumor and normal samples at a candidate somatic site. Also, they differ in how biological and technological noises are addressed. Hence, they are expected to produce divergent outputs. To overcome the drawbacks of existing somatic callers, we develop a deep learning-based tool called DeepSSV, which employs a convolutional neural network (CNN) model to learn increasingly abstract feature representations from the raw data in higher feature layers. DeepSSV creates a spatially oriented representation of read alignments around the candidate somatic sites adapted for the convolutional architecture, which enables it to expand to effectively gather scattered evidence. Moreover, DeepSSV incorporates the mapping information of both reference allele-supporting and variant allele-supporting reads in the tumor and normal samples at a genomic site that are readily available in the pileup format file. Together, the CNN model can process the whole alignment information. Such representational richness allows the model to capture the dependencies in the sequence and identify context-based sequencing artifacts. We fitted the model on ground truth somatic mutations and did benchmarking experiments on simulated and real tumors. The benchmarking results demonstrate that DeepSSV outperforms its state-of-the-art competitors in overall F1 score.
Brandon Victor, Zhen He 0002, Hongde Liu 0001, Taijiao Jiang
Briefings Bioinform.5
2021 Co-mutation modules capture the evolution and transmission patterns of SARS-CoV-2
abstract
The rapid spread and huge impact of the COVID-19 pandemic caused by the emerging SARS-CoV-2 have driven large efforts for sequencing and analyzing the viral genomes. Mutation analyses have revealed that the virus keeps mutating and shows a certain degree of genetic diversity, which could result in the alteration of its infectivity and pathogenicity. Therefore, appropriate delineation of SARS-CoV-2 genetic variants enables us to understand its evolution and transmission patterns. By focusing on the nucleotides that co-substituted, we first identified 42 co-mutation modules that consist of at least two co-substituted nucleotides during the SARS-CoV-2 evolution. Then based on these co-mutation modules, we classified the SARS-CoV-2 population into 43 groups and further identified the phylogenetic relationships among groups based on the number of inconsistent co-mutation modules, which were validated with phylogenetic trees. Intuitively, we tracked tempo-spatial patterns of the 43 groups, of which 11 groups were geographic-specific. Different epidemic periods showed specific co-circulating groups, where the dominant groups existed and had multiple sub-groups of parallel evolution. Our work enables us to capture the evolution and transmission patterns of SARS-CoV-2, which can contribute to guiding the prevention and control of the COVID-19 pandemic. An interactive website for grouping SARS-CoV-2 genomes and visualizing the spatio-temporal distribution of groups is available at https://www.jianglab.tech/cmm-grouping/.
Luyao Qin, Qingfeng Chen, Taijiao Jiang
Briefings Bioinform.6
2021 Compositional diversity and evolutionary pattern of coronavirus accessory proteins
abstract
Accessory proteins play important roles in the interaction between coronaviruses and their hosts. Accordingly, a comprehensive study of the compositional diversity and evolutionary patterns of accessory proteins is critical to understanding the host adaptation and epidemic variation of coronaviruses. Here, we developed a standardized genome annotation tool for coronavirus (CoroAnnoter) by combining open reading frame prediction, transcription regulatory sequence recognition and homologous alignment. Using CoroAnnoter, we annotated 39 representative coronavirus strains to form a compositional profile for all of the accessary proteins. Large variations were observed in the number of accessory proteins of 1-10 for different coronaviruses, with SARS-CoV-2 and SARS-CoV having the most (9 and 10, respectively). The variation between SARS-CoV and SARS-CoV-2 accessory proteins could be traced back to related coronaviruses in other hosts. The genomic distribution of accessory proteins had significant intra-genus conservation and inter-genus diversity and could be grouped into 1, 4, 2 and 1 types for alpha-, beta-, gamma-, and delta-coronaviruses, respectively. Evolutionary analysis suggested that accessory proteins are more conservative locating before the N-terminal of proteins E and M (E-M), while they are more diverse after these proteins. Furthermore, comparison of virus-host interaction networks of SARS-CoV-2 and SARS-CoV accessory proteins showed that they share multiple antiviral signaling pathways, those involved in the apoptotic process, viral life cycle and response to oxidative stress. In summary, our study provides a tool for coronavirus genome annotation and builds a comprehensive profile for coronavirus accessory proteins covering their composition, classification, evolutionary pattern and host interaction.
Jingzhe Shang, Na Han, Yousong Peng, Hangyu Zhou, Chengyang Ji, Taijiao Jiang, Aiping Wu 0002
Briefings Bioinform.9
2021 Classification and characterization of multigene family proteins of African swine fever viruses
abstract
African swine fever virus (ASFV) poses serious threats to the pig industry. The multigene family (MGF) proteins are extensively distributed in ASFVs and are generally classified into five families, including MGF-100, MGF-110, MGF-300, MGF-360 and MGF-505. Most MGF proteins, however, have not been well characterized and classified within each family. To bridge this gap, this study first classified MGF proteins into 31 groups based on protein sequence homology and network clustering. A web server for classifying MGF proteins was established and kept available for free at http://www.computationalbiology.cn/MGF/home.html. Results showed that MGF groups of the same family were most similar to each other and had conserved sequence motifs; the genetic diversity of MGF groups varied widely, mainly due to the occurrence of indels. In addition, the MGF proteins were predicted to have large structural and functional diversity, and MGF proteins of the same MGF family tended to have similar structure, location and function. Reconstruction of the ancestral states of MGF groups along the ASFV phylogeny showed that most MGF groups experienced either the copy number variations or the gain-or-loss changes, and most of these changes happened within strains of the same genotype. It is found that the copy number decrease and the loss of MGF groups were much larger than the copy number increase and the gain of MGF groups, respectively, suggesting the ASFV tended to lose MGF proteins in the evolution. Overall, the work provides a detailed classification for MGF proteins and would facilitate further research on MGF proteins.
Zhaozhong Zhu, Huiting Chen, Taijiao Jiang, Yuanqiang Zou, Yousong Peng
Briefings Bioinform.5
2020 FluReassort: a database for the study of genomic reassortments among influenza viruses
abstract
Genomic reassortment is an important genetic event in the generation of emerging influenza viruses, which can cause numerous serious flu endemics and epidemics within hosts or even across different hosts. However, there is no dedicated and comprehensive repository for reassortment events among influenza viruses. Here, we present FluReassort, a database for understanding the genomic reassortment events in influenza viruses. Through manual curation of thousands of literature references, the database compiles 204 reassortment events among 56 subtypes of influenza A viruses isolated in 37 different countries. FluReassort provides an interface for the visualization and evolutionary analysis of reassortment events, allowing users to view the events through the phylogenetic analysis with varying parameters. The reassortment networks in FluReassort graphically summarize the correlation and causality between different subtypes of the influenza virus and facilitate the description and interpretation of the reassortment preference among subtypes. We believe FluReassort is a convenient and powerful platform for understanding the evolution of emerging influenza viruses. FluReassort is freely available at https://www.jianglab.tech/FluReassort.
Xuye Yuan, Longfei Mao, Aiping Wu 0002, Taijiao Jiang
Briefings Bioinform.5
2020 FluPhenotype - a one-stop platform for early warnings of the influenza A virus
abstract
MOTIVATION: Newly emerging influenza viruses keep challenging global public health. To evaluate the potential risk of the viruses, it is critical to rapidly determine the phenotypes of the viruses, including the antigenicity, host, virulence and drug resistance. RESULTS: Here, we built FluPhenotype, a one-stop platform to rapidly determinate the phenotypes of the influenza A viruses. The input of FluPhenotype is the complete or partial genomic/protein sequences of the influenza A viruses. The output presents five types of information about the viruses: (i) sequence annotation including the gene and protein names as well as the open reading frames, (ii) potential hosts and human-adaptation-associated amino acid markers, (iii) antigenic and genetic relationships with the vaccine strains of different HA subtypes, (iv) mammalian virulence-related amino acid markers and (v) drug resistance-related amino acid markers. FluPhenotype will be a useful bioinformatic tool for surveillance and early warnings of the newly emerging influenza A viruses. AVAILABILITY AND IMPLEMENTATION: It is publicly available from: http://www.computationalbiology.cn : 18888/IVEW. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Congyu Lu, Zena Cai, Yuanqiang Zou, Lizong Deng, Xiangjun Du, Aiping Wu 0002, Lei Yang 0002, Dayan Wang, Yuelong Shu, Taijiao Jiang, Yousong Peng
Bioinform.12
2020 Phage protein receptors have multiple interaction partners and high expressions
abstract
MOTIVATION: Receptors on host cells play a critical role in viral infection. How phages select receptors is still unknown. RESULTS: Here, we manually curated a high-quality database named phageReceptor, including 427 pairs of phage-host receptor interactions, 341 unique viral species or sub-species and 69 bacterial species. Sugars and proteins were most widely used by phages as receptors. The receptor usage of phages in Gram-positive bacteria was different from that in Gram-negative bacteria. Most protein receptors were located on the outer membrane. The phage protein receptors (PPRs) were highly diverse in their structures, and had little sequence identity and no common protein domain with mammalian virus receptors. Further functional characterization of PPRs in Escherichia coli showed that they had larger node degrees and betweennesses in the protein-protein interaction network, and higher expression levels, than other outer membrane proteins, plasma membrane proteins or other intracellular proteins. These findings were consistent with what observed for mammalian virus receptors reported in previous studies, suggesting that viral protein receptors tend to have multiple interaction partners and high expressions. The study deepens our understanding of virus-host interactions. AVAILABILITY AND IMPLEMENTATION: phageReceptor is publicly available from: http://www.computationalbiology.cn/phageReceptor/index.html. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Fen Yu, Yuanqiang Zou, Ye Qiu, Aiping Wu 0002, Taijiao Jiang, Yousong Peng
Bioinform.6
2020 LATTE: A knowledge-based method to normalize various expressions of laboratory test results in free text of Chinese electronic health records
Chunyan Wu, Longfei Mao, Yongyou Wu, Lizong Deng, Taijiao Jiang
J. Biomed. Informatics8
2019 Cell membrane proteins with high N-glycosylation, high expression and multiple interaction partners are preferred by mammalian viruses as receptors
abstract
MOTIVATION: Receptor mediated entry is the first step for viral infection. However, the question of how viruses select receptors remains unanswered. RESULTS: Here, by manually curating a high-quality database of 268 pairs of mammalian virus-host receptor interaction, which included 128 unique viral species or sub-species and 119 virus receptors, we found the viral receptors are structurally and functionally diverse, yet they had several common features when compared to other cell membrane proteins: more protein domains, higher level of N-glycosylation, higher ratio of self-interaction and more interaction partners, and higher expression in most tissues of the host. This study could deepen our understanding of virus-receptor interaction. AVAILABILITY AND IMPLEMENTATION: The database of mammalian virus-host receptor interaction is available at http://www.computationalbiology.cn: 5000/viralReceptor. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Zhaozhong Zhu, Zena Cai, Beibei Xu, Zhiying Tan, Aiping Wu 0002, Xingyi Ge, Xinhong Guo, Zhongyang Tan, Zanxian Xia, Haizhen Zhu, Taijiao Jiang, Yousong Peng
Bioinform.13
2018 Multiple Visual Fields Cascaded Convolutional Neural Network for Breast Cancer Detection
Haomiao Ni, Hong Liu 0007, Zichao Guo, Taijiao Jiang, Kuansong Wang, Yueliang Qian
PRICAI (1)5
2017 cooccurNet: an R package for co-occurrence network construction and analysis
abstract
MOTIVATION: Previously, we developed a computational model to identify genomic co-occurrence networks that was applied to capture the coevolution patterns within genomes of influenza viruses. To facilitate easy public use of this model, an R package 'cooccurNet' is presented here. RESULTS: 'cooccurNet' includes functionalities of construction and analysis of residues (e.g. nucleotides, amino acids and SNPs) co-occurrence network. In addition, a new method for measuring residues coevolution, defined as residue co-occurrence score (RCOS), is proposed and implemented in 'cooccurNet' based on the co-occurrence network. AVAILABILITY AND IMPLEMENTATION: 'cooccurNet' is publicly available on CRAN repositories under the GPL-3 Open Source License ( http://cran.r-project.org/package=cooccurNet ). CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yuanqiang Zou, Zhiqiang Wu 0001, Lizong Deng, Aiping Wu 0002, Fan Wu 0016, Kenli Li 0001, Taijiao Jiang, Yousong Peng
Bioinform.7
2016 PREDAC-H3: a user-friendly platform for antigenic surveillance of human influenza a(H3N2) virus based on hemagglutinin sequences
abstract
MOTIVATION: Timely surveillance of the antigenic dynamics of the influenza virus is critical for accurate selection of vaccine strains, which is important for effective prevention of viral spread and infection. RESULTS: Here, we provide a computational platform, called PREDAC-H3, for antigenic surveillance of human influenza A(H3N2) virus based on the sequence of surface protein hemagglutinin (HA). PREDAC-H3 not only determines the antigenic variants and antigenic cluster (grouped for similar antigenicity) to which the virus belongs, based on HA sequences, but also allows visualization of the spatial distribution and temporal dynamics of antigenic clusters of viruses isolated from around the world, thus assisting in antigenic surveillance of human influenza A(H3N2) virus. AVAILABILITY AND IMPLEMENTATION: It is publicly available from: http://biocloud.hnu.edu.cn/influ411/html/index.php CONTACTS: : [email protected] or [email protected].
Yousong Peng, Lei Yang 0002, Honglei Li 0002, Yuanqiang Zou, Lizong Deng, Aiping Wu 0002, Xiangjun Du, Dayan Wang, Yuelong Shu, Taijiao Jiang
Bioinform.10
2014 Exploring protein domain organization by recognition of secondary structure packing interfaces
abstract
MOTIVATION: Protein domains are fundamental units of protein structure, function and evolution; thus, it is critical to gain a deep understanding of protein domain organization. Previous works have attempted to identify key residues involved in organization of domain architecture. Because one of the most important characteristics of domain architecture is the arrangement of secondary structure elements (SSEs), here we present a picture of domain organization through an integrated consideration of SSE arrangements and residue contact networks. RESULTS: In this work, by representing SSEs as main-chain scaffolds and side-chain interfaces and through construction of residue contact networks, we have identified the SSE interfaces well packed within protein domains as SSE packing clusters. In total, 17 334 SSE packing clusters were recognized from 9015 Structural Classification of Proteins domains of <40% sequence identity. The similar SSE packing clusters were observed not only among domains of the same folds, but also among domains of different folds, indicating their roles as common scaffolds for organization of protein domains. Further analysis of 14 small single-domain proteins reveals a high correlation between the SSE packing clusters and the folding nuclei. Consistent with their important roles in domain organization, SSE packing clusters were found to be more conserved than other regions within the same proteins. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Lizong Deng, Aiping Wu 0002, Wentao Dai, Tingrui Song, Ya Cui, Taijiao Jiang
Bioinform.6
2011 Improved side-chain modeling by coupling clash-detection guided iterative search with rotamer relaxation
abstract
MOTIVATION: Side-chain modeling has seen wide applications in computational structure biology. Most of the popular side-chain modeling programs explore the conformation space using discrete rigid rotamers for speed and efficiency. However, in the tightly packed environments of protein interiors, these methods will inherently lead to atomic clashes and hinder the prediction accuracy. RESULTS: We present a side-chain modeling method (CIS-RR), which couples a novel clash-detection guided iterative search (CIS) algorithm with continuous torsion space optimization of rotamers (RR). Benchmark testing shows that compared with the existing popular side-chain modeling methods, CIS-RR removes atomic clashes much more effectively and achieves comparable or even better prediction accuracy while having comparable computational cost. We believe that CIS-RR could be a useful method for accurate side-chain modeling. AVAILABILITY: CIS-RR is available to non-commercial users at our website: http://jianglab.ibp.ac.cn/lims/cisrr/cisrr.html.
Zhichao Miao, Liqing Tian, Taijiao Jiang
Bioinform.6
2011 RASP: rapid modeling of protein side chain conformations
abstract
MOTIVATION: Modeling of side chain conformations constitutes an indispensable effort in protein structure modeling, protein-protein docking and protein design. Thanks to an intensive attention to this field, many of the existing programs can achieve reasonably good and comparable prediction accuracy. Moreover, in our previous work on CIS-RR, we argued that the prediction with few atomic clashes can complement the current existing methods for subsequent analysis and refinement of protein structures. However, these recent efforts to enhance the quality of predicted side chains have been accompanied by a significant increase of computational cost. RESULTS: In this study, by mainly focusing on improving the speed of side chain conformation prediction, we present a RApid Side-chain Predictor, called RASP. To achieve a much faster speed with a comparable accuracy to the best existing methods, we not only employ the clash elimination strategy of CIS-RR, but also carefully optimize energy terms and integrate different search algorithms. In comprehensive benchmark testings, RASP is over one order of magnitude faster (~ 40 times over CIS-RR) than the recently developed methods, while achieving comparable or even better accuracy.
Zhichao Miao, Taijiao Jiang
Bioinform.3
2011 NCACO-score: An effective main-chain dependent scoring function for structure modeling
abstract
BACKGROUND: Development of effective scoring functions is a critical component to the success of protein structure modeling. Previously, many efforts have been dedicated to the development of scoring functions. Despite these efforts, development of an effective scoring function that can achieve both good accuracy and fast speed still presents a grand challenge. RESULTS: Based on a coarse-grained representation of a protein structure by using only four main-chain atoms: N, Cα, C and O, we develop a knowledge-based scoring function, called NCACO-score, that integrates different structural information to rapidly model protein structure from sequence. In testing on the Decoys'R'Us sets, we found that NCACO-score can effectively recognize native conformers from their decoys. Furthermore, we demonstrate that NCACO-score can effectively guide fragment assembly for protein structure prediction, which has achieved a good performance in building the structure models for hard targets from CASP8 in terms of both accuracy and speed. CONCLUSIONS: Although NCACO-score is developed based on a coarse-grained model, it is able to discriminate native conformers from decoy conformers with high accuracy. NCACO is a very effective scoring function for structure modeling.
Liqing Tian, Aiping Wu 0002, Xiaoxi Dong, Taijiao Jiang
BMC Bioinform.6
2010 Correlation of Influenza Virus Excess Mortality with Antigenic Variation: Application to Rapid Estimation of Influenza Mortality Burden
abstract
The variants of human influenza virus have caused, and continue to cause, substantial morbidity and mortality. Timely and accurate assessment of their impact on human death is invaluable for influenza planning but presents a substantial challenge, as current approaches rely mostly on intensive and unbiased influenza surveillance. In this study, by proposing a novel host-virus interaction model, we have established a positive correlation between the excess mortalities caused by viral strains of distinct antigenicity and their antigenic distances to their previous strains for each (sub)type of seasonal influenza viruses. Based on this relationship, we further develop a method to rapidly assess the mortality burden of influenza A(H1N1) virus by accurately predicting the antigenic distance between A(H1N1) strains. Rapid estimation of influenza mortality burden for new seasonal strains should help formulate a cost-effective response for influenza control and prevention.
Aiping Wu 0002, Yousong Peng, Xiangjun Du, Yuelong Shu, Taijiao Jiang
PLoS Comput. Biol.5
2006 Paircoil2: improved prediction of coiled coils from sequence
abstract
Abstract Summary: We introduce Paircoil2, a new version of the Paircoil program, which uses pairwise residue probabilities to detect coiled–coil motifs in protein sequence data. Paircoil2 achieves 98% sensitivity and 97% specificity on known coiled coils in leave-family-out cross-validation. It also shows superior performance compared with published methods in tests on proteins of known structure. Availability: Paircoil2 is freely available as a web application and for download at Contact: [email protected]; [email protected] Supplementary information: Available at Bioinformatics online and at the Paircoil website.
A. V. McDonnell, Taijiao Jiang, Amy E. Keating, Bonnie Berger
Bioinform.2
2005 AVID: An integrative framework for discovering functional relationships among proteins
abstract
Abstract Background Determining the functions of uncharacterized proteins is one of the most pressing problems in the post-genomic era. Large scale protein-protein interaction assays, global mRNA expression analyses and systematic protein localization studies provide experimental information that can be used for this purpose. The data from such experiments contain many false positives and false negatives, but can be processed using computational methods to provide reliable information about protein-protein relationships and protein function. An outstanding and important goal is to predict detailed functional annotation for all uncharacterized proteins that is reliable enough to effectively guide experiments. Results We present AVID, a computational method that uses a multi-stage learning framework to integrate experimental results with sequence information, generating networks reflecting functional similarities among proteins. We illustrate use of the networks by making predictions of detailed Gene Ontology (GO) annotations in three categories: molecular function, biological process, and cellular component. Applied to the yeast Saccharomyces cerevisiae , AVID provides 37,451 pair-wise functional linkages between 4,191 proteins. These relationships are ~65–78% accurate, as assessed by cross-validation testing. Assignments of highly detailed functional descriptors to proteins, based on the networks, are estimated to be ~67% accurate for GO categories describing molecular function and cellular component and ~52% accurate for terms describing biological process. The predictions cover 1,490 proteins with no previous annotation in GO and also assign more detailed functions to many proteins annotated only with less descriptive terms. Predictions made by AVID are largely distinct from those made by other methods. Out of 37,451 predicted pair-wise relationships, the greatest number shared in common with another method is 3,413. Conclusion AVID provides three networks reflecting functional associations among proteins. We use these networks to generate new, highly detailed functional predictions for roughly half of the yeast proteome that are reliable enough to drive targeted experimental investigations. The predictions suggest many specific, testable hypotheses. All of the data are available as downloadable files as well as through an interactive website at http://web.mit.edu/biology/keating/AVID . Thus, AVID will be a valuable resource for experimental biologists.
Taijiao Jiang, Amy E. Keating
BMC Bioinform.1