Aiping Wu 0002

dblp:25/1688-2 · DBLP profile ↗
← Back
22ranked-venue papers
1as first author
11since 2021 · last 2025
0000-0002-5869-651XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 21 · 1 first-author · 10 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2025 TransSSVs: a Transformer-based deep learning model for accurate detection of somatic small variants in paired tumor and normal sequencing data
abstract
Accurate identification of somatic small variants in tumors plays a crucial role in cancer diagnosis. Various somatic mutation callers have been developed; however, existing methods face limitations in modeling mapping information of flanking genomic sites (genomic sites adjacent to a somatic site) that influence the state of the somatic site. Additionally, they are unable to analyze inter-site interactions within the context sequence centered around a somatic site or appropriately weigh the effects of flanking genomic sites on the somatic site. To address these limitations, the Transformer model is utilized to develop TransSSVs for detecting somatic small variants. The core functionality of TransSSVs relies on the multi-head attention mechanism, which generates a reliable representation of interactions between a candidate somatic site and its flanking genomic sites within the context sequence. TransSSVs effectively extract mapping features of various genomic sites in the context sequence to enhance prediction accuracy. Benchmarking experiments demonstrate that TransSSVs exhibit robust performance when compared with state-of-the-art methods on well-characterized real and simulated tumor datasets. Furthermore, the contributions of flanking genomic sites to the detection of somatic sites are assessed, and attention weight patterns for positive and negative somatic sites are analyzed.
Jiangyuan Wang, Jingze Liu, Wenkai Song, Aiping Wu 0002, Taijiao Jiang
Appl. Intell.6
2025 Influenza virus reassortment patterns exhibit preference and continuity while uncovering cross-species transmission events
abstract
Genomic reassortment is a key driver of influenza virus evolution and a major factor in pandemic emergence, as reassorted strains can exhibit significantly altered antigenicity. However, due to technical and ethical constraints, research on reassortment patterns (RPs) has been limited, impeding effective surveillance and control strategies. To address this gap, we developed FluRPId, a framework for identifying RPs based on the genetic diversity of influenza viruses. FluRPId integrates principles of reassortment diversity maximization, dominance, and epidemiological likelihood to assess the credibility of detected reassortment events. Applying FluRPId, we constructed a comprehensive reassortment landscape of influenza viruses, encompassing widespread reassortment events with high credibility, which also include most previously reported reassortment events. Our analysis revealed that the NS gene frequently reassorts with PA and NA, while reassortment involving HA, NA, and NS occurs more frequently than expected. Furthermore, we identified specific loci combinations that exhibit strong linkage during reassortment, providing insights into segment association preferences. Additionally, extensive reassortment chains were observed across all subtypes, underscoring the continuity of reassortment in influenza virus evolution. Notably, we identified significant cross-species reassortment events and characterized host adaptation changes in cross-species-transmitted viruses. Our study provides the most comprehensive reassortment landscape of influenza viruses to date, uncovering key patterns, preferences, and evolutionary continuity. These findings bridge a critical gap in macro-scale reassortment studies and offer insights for future research and control efforts.
Yun Ma 0016, Jingze Liu, Luyao Qin, Aiping Wu 0002
Briefings Bioinform.6
2024 PREDAC-CNN: predicting antigenic clusters of seasonal influenza A viruses with convolutional neural network
abstract
Vaccination stands as the most effective and economical strategy for prevention and control of influenza. The primary target of neutralizing antibodies is the surface antigen hemagglutinin (HA). However, ongoing mutations in the HA sequence result in antigenic drift. The success of a vaccine is contingent on its antigenic congruence with circulating strains. Thus, predicting antigenic variants and deducing antigenic clusters of influenza viruses are pivotal for recommendation of vaccine strains. The antigenicity of influenza A viruses is determined by the interplay of amino acids in the HA1 sequence. In this study, we exploit the ability of convolutional neural networks (CNNs) to extract spatial feature representations in the convolutional layers, which can discern interactions between amino acid sites. We introduce PREDAC-CNN, a model designed to track antigenic evolution of seasonal influenza A viruses. Accessible at http://predac-cnn.cloudna.cn, PREDAC-CNN formulates a spatially oriented representation of the HA1 sequence, optimized for the convolutional framework. It effectively probes interactions among amino acid sites in the HA1 sequence. Also, PREDAC-CNN focuses exclusively on physicochemical attributes crucial for the antigenicity of influenza viruses, thereby eliminating unnecessary amino acid embeddings. Together, PREDAC-CNN is adept at capturing interactions of amino acid sites within the HA1 sequence and examining the collective impact of point mutations on antigenic variation. Through 5-fold cross-validation and retrospective testing, PREDAC-CNN has shown superior performance in predicting antigenic variants compared to its counterparts. Additionally, PREDAC-CNN has been instrumental in identifying predominant antigenic clusters for A/H3N2 (1968-2023) and A/H1N1 (1977-2023) viruses, significantly aiding in vaccine strain recommendation.
Jingze Liu, Wenkai Song, Honglei Li 0002, Jiangyuan Wang, Le Zhang 0004, Yousong Peng, Aiping Wu 0002, Taijiao Jiang
Briefings Bioinform.8
2023 TeaBERT: An Efficient Knowledge Infused Cross-Lingual Language Model for Mapping Chinese Medical Entities to the Unified Medical Language System
abstract
Medical entity normalization is an important task for medical information processing. The Unified Medical Language System (UMLS), a well-developed medical terminology system, is crucial for medical entity normalization. However, the UMLS primarily consists of English medical terms. For languages other than English, such as Chinese, a significant challenge for normalizing medical entities is the lack of robust terminology systems. To address this issue, we propose a translation-enhancing training strategy that incorporates the translation and synonym knowledge of the UMLS into a language model using the contrastive learning approach. In this work, we proposed a cross-lingual pre-trained language model called TeaBERT, which can align synonymous Chinese and English medical entities across languages at the concept level. As the evaluation results showed, the TeaBERT language model outperformed previous cross-lingual language models with Acc@5 values of 92.54%, 87.14% and 84.77% on the ICD10-CN, CHPO and RealWorld-v2 datasets, respectively. It also achieved a new state-of-the-art cross-lingual entity mapping performance without fine-tuning. The translation-enhancing strategy is applicable to other languages that face the similar challenge due to the absence of well-developed medical terminology systems.
Yifan Qi, Aiping Wu 0002, Lizong Deng, Taijiao Jiang
IEEE J. Biomed. Health Informatics3
2022 vsRNAfinder: a novel method for identifying high-confidence viral small RNAs from small RNA-Seq data
abstract
Virus-encoded small RNAs (vsRNA) have been reported to play an important role in viral infection. Unfortunately, there is still a lack of an effective method for vsRNA identification. Herein, we presented vsRNAfinder, a de novo method for identifying high-confidence vsRNAs from small RNA-Seq (sRNA-Seq) data based on peak calling and Poisson distribution and is publicly available at https://github.com/ZenaCai/vsRNAfinder. vsRNAfinder outperformed two widely used methods namely miRDeep2 and ShortStack in identifying viral miRNAs with a significantly improved sensitivity. It can also be used to identify sRNAs in animals and plants with similar performance to miRDeep2 and ShortStack. vsRNAfinder would greatly facilitate effective identification of vsRNAs from sRNA-Seq data.
Zena Cai, Ye Qiu, Aiping Wu 0002, Gaihua Zhang, Taijiao Jiang, Xing-Yi Ge, Haizhen Zhu, Yousong Peng
Briefings Bioinform.4
2022 An atlas of human viruses provides new insights into diversity and tissue tropism of human viruses
abstract
MOTIVATION: Viruses continue to threaten human health. Yet, the complete viral species carried by humans and their infection characteristics have not been fully revealed. RESULTS: This study curated an atlas of human viruses from public databases and literature, and built the Human Virus Database (HVD). The HVD contains 1131 virus species of 54 viral families which were more than twice the number of the human-infecting virus species reported in previous studies. These viruses were identified in human samples including 68 human tissues, the excreta and body fluid. The viral diversity in humans was age-dependent with a peak in the infant and a valley in the teenager. The tissue tropism of viruses was found to be associated with several factors including the viral group (DNA, RNA or reverse-transcribing viruses), enveloped or not, viral genome length and GC content, viral receptors and the virus-interacting proteins. Finally, the tissue tropism of DNA viruses was predicted using a random-forest algorithm with a middle performance. Overall, the study not only provides a valuable resource for further studies of human viruses but also deepens our understanding toward the diversity and tissue tropism of human viruses. AVAILABILITY AND IMPLEMENTATION: The HVD is available at http://computationalbiology.cn/humanVirusBase/#/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Sifan Ye, Congyu Lu, Ye Qiu, Heping Zheng, Xingyi Ge, Aiping Wu 0002, Zanxian Xia, Taijiao Jiang, Haizhen Zhu, Yousong Peng
Bioinform.6
2022 sitePath: a visual tool to identify polymorphism clades and help find fixed and parallel mutations
abstract
BACKGROUND: Identifying polymorphism clades on phylogenetic trees could help detect punctual mutations that are associated with viral functions. With visualization tools coloring the tree, it is easy to visually find clades where most sequences have the same polymorphism state. However, with the fast accumulation of viral sequences, a computational tool to automate this process is urgently needed. RESULTS: Here, by implementing a branch-and-bound-like search method, we developed an R package named sitePath to identify polymorphism clades automatically. Based on the identified polymorphism clades, fixed and parallel mutations could be inferred. Furthermore, sitePath also integrated visualization tools to generate figures of the calculated results. In an example with the influenza A virus H3N2 dataset, the detected fixed mutations coincide with antigenic shift mutations. The highly specificity and sensitivity of sitePath in finding fixed mutations were achieved for a range of parameters and different phylogenetic tree inference software. CONCLUSIONS: The result suggests that sitePath can identify polymorphism clades per site. The clustering of sequences on a phylogenetic tree can be used to infer fixed and parallel mutations. High-quality figures of the calculated results could also be generated by sitePath.
Chengyang Ji, Na Han, Yexiao Cheng, Jingzhe Shang, Shenghui Weng, Hangyu Zhou, Aiping Wu 0002
BMC Bioinform.8
2021 Identification and characterization of circRNAs encoded by MERS-CoV, SARS-CoV-1 and SARS-CoV-2
abstract
The life-threatening coronaviruses MERS-CoV, SARS-CoV-1 and SARS-CoV-2 (SARS-CoV-1/2) have caused and will continue to cause enormous morbidity and mortality to humans. Virus-encoded noncoding RNAs are poorly understood in coronaviruses. Data mining of viral-infection-related RNA-sequencing data has resulted in the identification of 28 754, 720 and 3437 circRNAs encoded by MERS-CoV, SARS-CoV-1 and SARS-CoV-2, respectively. MERS-CoV exhibits much more prominent ability to encode circRNAs in all genomic regions than those of SARS-CoV-1/2. Viral circRNAs typically exhibit low expression levels. Moreover, majority of the viral circRNAs exhibit expressions only in the late stage of viral infection. Analysis of the competitive interactions of viral circRNAs, human miRNAs and mRNAs in MERS-CoV infections reveals that viral circRNAs up-regulated genes related to mRNA splicing and processing in the early stage of viral infection, and regulated genes involved in diverse functions including cancer, metabolism, autophagy, viral infection in the late stage of viral infection. Similar analysis in SARS-CoV-2 infections reveals that its viral circRNAs down-regulated genes associated with metabolic processes of cholesterol, alcohol, fatty acid and up-regulated genes associated with cellular responses to oxidative stress in the late stage of viral infection. A few genes regulated by viral circRNAs from both MERS-CoV and SARS-CoV-2 were enriched in several biological processes such as response to reactive oxygen and centrosome localization. This study provides the first glimpse into viral circRNAs in three deadly coronaviruses and would serve as a valuable resource for further studies of circRNAs in coronaviruses.
Zena Cai, Congyu Lu, Yuanqiang Zou, Zhaozhong Zhu, Xingyi Ge, Aiping Wu 0002, Taijiao Jiang, Heping Zheng, Yousong Peng
Briefings Bioinform.9
2021 Progress and challenge for computational quantification of tissue immune cells
abstract
Tissue immune cells have long been recognized as important regulators for the maintenance of balance in the body system. Quantification of the abundance of different immune cells will provide enhanced understanding of the correlation between immune cells and normal or abnormal situations. Currently, computational methods to predict tissue immune cell compositions from bulk transcriptomes have been largely developed. Therefore, summarizing the advantages and disadvantages is appropriate. In addition, an examination of the challenges and possible solutions for these computational models will assist the development of this field. The common hypothesis of these models is that the expression of signature genes for immune cell types might represent the proportion of immune cells that contribute to the tissue transcriptome. In general, we grouped all reported tools into three groups, including reference-free, reference-based scoring and reference-based deconvolution methods. In this review, a summary of all the currently reported computational immune cell quantification tools and their applications, limitations, and perspectives are presented. Furthermore, some critical problems are found that have limited the performance and application of these models, including inadequate immune cell type, the collinearity problem, the impact of the tissue environment on the immune cell expression level, and the deficiency of standard datasets for model validation. To address these issues, tissue specific training datasets that include all known immune cells, a hierarchical computational framework, and benchmark datasets including both tissue expression profiles and the abundances of all the immune cells are proposed to further promote the development of this field.
Aiping Wu 0002
Briefings Bioinform.2
2021 Comparative viromes of Culicoides and mosquitoes reveal their consistency and diversity in viral profiles
abstract
The genus Culicoides includes biting midges, some of which are vectors for viruses that cause diseases in humans and animals. Knowledge of the roles of Culicoides in viral ecology is inadequate. We collected ~300 000 samples of Culicoides and mosquitoes in 15 representative regions within Yunnan, China. Using mosquitoes as reference vectors, we designed a comparative virome strategy to study the viral composition, diversity, hosts and spatiotemporal distribution of Culicoides. A map of viromes in Culicoides and mosquitoes in Yunan province, China, was constructed. At the same locations, Culicoides and mosquitoes usually share a similar viral diversity. At least 10 important pathogenic viruses were detected from Culicoides. Many novel viruses were discovered, including 21 segmented viruses of Flaviviridae, 180 viruses of Monjiviricetes and 130 viruses of Bunyavirales. The findings demonstrate that Culicoides is an important part of viral ecology and should be studied and monitored for potentially emerging viruses.
Qin Shen, Yuwen He, Na Han, Xianyue Wang, Jinxin Meng, Yousong Peng, Mei Pan, Yuting Jin, Taijiao Jiang, Wenjie Tan, Jinglin Wang, Aiping Wu 0002
Briefings Bioinform.14
2021 Compositional diversity and evolutionary pattern of coronavirus accessory proteins
abstract
Accessory proteins play important roles in the interaction between coronaviruses and their hosts. Accordingly, a comprehensive study of the compositional diversity and evolutionary patterns of accessory proteins is critical to understanding the host adaptation and epidemic variation of coronaviruses. Here, we developed a standardized genome annotation tool for coronavirus (CoroAnnoter) by combining open reading frame prediction, transcription regulatory sequence recognition and homologous alignment. Using CoroAnnoter, we annotated 39 representative coronavirus strains to form a compositional profile for all of the accessary proteins. Large variations were observed in the number of accessory proteins of 1-10 for different coronaviruses, with SARS-CoV-2 and SARS-CoV having the most (9 and 10, respectively). The variation between SARS-CoV and SARS-CoV-2 accessory proteins could be traced back to related coronaviruses in other hosts. The genomic distribution of accessory proteins had significant intra-genus conservation and inter-genus diversity and could be grouped into 1, 4, 2 and 1 types for alpha-, beta-, gamma-, and delta-coronaviruses, respectively. Evolutionary analysis suggested that accessory proteins are more conservative locating before the N-terminal of proteins E and M (E-M), while they are more diverse after these proteins. Furthermore, comparison of virus-host interaction networks of SARS-CoV-2 and SARS-CoV accessory proteins showed that they share multiple antiviral signaling pathways, those involved in the apoptotic process, viral life cycle and response to oxidative stress. In summary, our study provides a tool for coronavirus genome annotation and builds a comprehensive profile for coronavirus accessory proteins covering their composition, classification, evolutionary pattern and host interaction.
Jingzhe Shang, Na Han, Yousong Peng, Hangyu Zhou, Chengyang Ji, Taijiao Jiang, Aiping Wu 0002
Briefings Bioinform.10
2020 FluReassort: a database for the study of genomic reassortments among influenza viruses
abstract
Genomic reassortment is an important genetic event in the generation of emerging influenza viruses, which can cause numerous serious flu endemics and epidemics within hosts or even across different hosts. However, there is no dedicated and comprehensive repository for reassortment events among influenza viruses. Here, we present FluReassort, a database for understanding the genomic reassortment events in influenza viruses. Through manual curation of thousands of literature references, the database compiles 204 reassortment events among 56 subtypes of influenza A viruses isolated in 37 different countries. FluReassort provides an interface for the visualization and evolutionary analysis of reassortment events, allowing users to view the events through the phylogenetic analysis with varying parameters. The reassortment networks in FluReassort graphically summarize the correlation and causality between different subtypes of the influenza virus and facilitate the description and interpretation of the reassortment preference among subtypes. We believe FluReassort is a convenient and powerful platform for understanding the evolution of emerging influenza viruses. FluReassort is freely available at https://www.jianglab.tech/FluReassort.
Xuye Yuan, Longfei Mao, Aiping Wu 0002, Taijiao Jiang
Briefings Bioinform.4
2020 Tissue-specific deconvolution of immune cell composition by integrating bulk and single-cell transcriptomes
abstract
MOTIVATION: Many methods have been developed to estimate immune cell composition from tissue transcriptomes. One common characteristic of these methods is that they are trained using a set of general immune cell transcriptomes that ignores tissue specificities. However, as immune cells are localized in different tissues, they may have distinct expression profiles. Hence, calculations that use general signature matrices may hinder the deconvolution accuracy. RESULTS: This study used single cell RNA-sequencing (scRNA-Seq) data from different mouse tissues instead of general signature expression values to generate tissue-specific signature gene matrices that are used as the input of the deconvolution model. First, the transcriptome of immune cells in each tissue was extracted from scRNA-Seq data and used to construct the entire expression matrix of tissue immune cells. Then, after comparing different gene selection strategies, the expressions of 162 seq-ImmuCC derived signature genes in tissue immune cell scRNA-Seq data were regarded as the tissue specific signature matrices. Finally, a modest improvement in performance was observed in multiple tissues that refer to a traditional general signature matrix in the deconvolution model. With the fast accumulation of scRNA-Seq data, the introduction of these data into an estimation of immune cell compositions for different tissues will open a new window for avoiding tissue bias for immune cell expression. AVAILABILITY AND IMPLEMENTATION: The signature matrices were available at https://github.com/wuaipinglab/ImmuCC/tree/master/tissue_immucc/SignatureMatrix). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Chengyang Ji, Qin Shen, F. Xiao-Feng Qin, Aiping Wu 0002
Bioinform.6
2020 FluPhenotype - a one-stop platform for early warnings of the influenza A virus
abstract
MOTIVATION: Newly emerging influenza viruses keep challenging global public health. To evaluate the potential risk of the viruses, it is critical to rapidly determine the phenotypes of the viruses, including the antigenicity, host, virulence and drug resistance. RESULTS: Here, we built FluPhenotype, a one-stop platform to rapidly determinate the phenotypes of the influenza A viruses. The input of FluPhenotype is the complete or partial genomic/protein sequences of the influenza A viruses. The output presents five types of information about the viruses: (i) sequence annotation including the gene and protein names as well as the open reading frames, (ii) potential hosts and human-adaptation-associated amino acid markers, (iii) antigenic and genetic relationships with the vaccine strains of different HA subtypes, (iv) mammalian virulence-related amino acid markers and (v) drug resistance-related amino acid markers. FluPhenotype will be a useful bioinformatic tool for surveillance and early warnings of the newly emerging influenza A viruses. AVAILABILITY AND IMPLEMENTATION: It is publicly available from: http://www.computationalbiology.cn : 18888/IVEW. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Congyu Lu, Zena Cai, Yuanqiang Zou, Lizong Deng, Xiangjun Du, Aiping Wu 0002, Lei Yang 0002, Dayan Wang, Yuelong Shu, Taijiao Jiang, Yousong Peng
Bioinform.8
2020 Phage protein receptors have multiple interaction partners and high expressions
abstract
MOTIVATION: Receptors on host cells play a critical role in viral infection. How phages select receptors is still unknown. RESULTS: Here, we manually curated a high-quality database named phageReceptor, including 427 pairs of phage-host receptor interactions, 341 unique viral species or sub-species and 69 bacterial species. Sugars and proteins were most widely used by phages as receptors. The receptor usage of phages in Gram-positive bacteria was different from that in Gram-negative bacteria. Most protein receptors were located on the outer membrane. The phage protein receptors (PPRs) were highly diverse in their structures, and had little sequence identity and no common protein domain with mammalian virus receptors. Further functional characterization of PPRs in Escherichia coli showed that they had larger node degrees and betweennesses in the protein-protein interaction network, and higher expression levels, than other outer membrane proteins, plasma membrane proteins or other intracellular proteins. These findings were consistent with what observed for mammalian virus receptors reported in previous studies, suggesting that viral protein receptors tend to have multiple interaction partners and high expressions. The study deepens our understanding of virus-host interactions. AVAILABILITY AND IMPLEMENTATION: phageReceptor is publicly available from: http://www.computationalbiology.cn/phageReceptor/index.html. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Fen Yu, Yuanqiang Zou, Ye Qiu, Aiping Wu 0002, Taijiao Jiang, Yousong Peng
Bioinform.5
2019 Cell membrane proteins with high N-glycosylation, high expression and multiple interaction partners are preferred by mammalian viruses as receptors
abstract
MOTIVATION: Receptor mediated entry is the first step for viral infection. However, the question of how viruses select receptors remains unanswered. RESULTS: Here, by manually curating a high-quality database of 268 pairs of mammalian virus-host receptor interaction, which included 128 unique viral species or sub-species and 119 virus receptors, we found the viral receptors are structurally and functionally diverse, yet they had several common features when compared to other cell membrane proteins: more protein domains, higher level of N-glycosylation, higher ratio of self-interaction and more interaction partners, and higher expression in most tissues of the host. This study could deepen our understanding of virus-receptor interaction. AVAILABILITY AND IMPLEMENTATION: The database of mammalian virus-host receptor interaction is available at http://www.computationalbiology.cn: 5000/viralReceptor. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Zhaozhong Zhu, Zena Cai, Beibei Xu, Zhiying Tan, Aiping Wu 0002, Xingyi Ge, Xinhong Guo, Zhongyang Tan, Zanxian Xia, Haizhen Zhu, Taijiao Jiang, Yousong Peng
Bioinform.7
2017 cooccurNet: an R package for co-occurrence network construction and analysis
abstract
MOTIVATION: Previously, we developed a computational model to identify genomic co-occurrence networks that was applied to capture the coevolution patterns within genomes of influenza viruses. To facilitate easy public use of this model, an R package 'cooccurNet' is presented here. RESULTS: 'cooccurNet' includes functionalities of construction and analysis of residues (e.g. nucleotides, amino acids and SNPs) co-occurrence network. In addition, a new method for measuring residues coevolution, defined as residue co-occurrence score (RCOS), is proposed and implemented in 'cooccurNet' based on the co-occurrence network. AVAILABILITY AND IMPLEMENTATION: 'cooccurNet' is publicly available on CRAN repositories under the GPL-3 Open Source License ( http://cran.r-project.org/package=cooccurNet ). CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yuanqiang Zou, Zhiqiang Wu 0001, Lizong Deng, Aiping Wu 0002, Fan Wu 0016, Kenli Li 0001, Taijiao Jiang, Yousong Peng
Bioinform.4
2016 PREDAC-H3: a user-friendly platform for antigenic surveillance of human influenza a(H3N2) virus based on hemagglutinin sequences
abstract
MOTIVATION: Timely surveillance of the antigenic dynamics of the influenza virus is critical for accurate selection of vaccine strains, which is important for effective prevention of viral spread and infection. RESULTS: Here, we provide a computational platform, called PREDAC-H3, for antigenic surveillance of human influenza A(H3N2) virus based on the sequence of surface protein hemagglutinin (HA). PREDAC-H3 not only determines the antigenic variants and antigenic cluster (grouped for similar antigenicity) to which the virus belongs, based on HA sequences, but also allows visualization of the spatial distribution and temporal dynamics of antigenic clusters of viruses isolated from around the world, thus assisting in antigenic surveillance of human influenza A(H3N2) virus. AVAILABILITY AND IMPLEMENTATION: It is publicly available from: http://biocloud.hnu.edu.cn/influ411/html/index.php CONTACTS: : [email protected] or [email protected].
Yousong Peng, Lei Yang 0002, Honglei Li 0002, Yuanqiang Zou, Lizong Deng, Aiping Wu 0002, Xiangjun Du, Dayan Wang, Yuelong Shu, Taijiao Jiang
Bioinform.6
2015 Uncover protein complexes in E.coli network
abstract
Recent advances in proteomic technologies have enabled high-throughput binary data on protein-protein interactions of E. coli to be released into public domain, and many protein complexes have been identified by experimental methods. Although it has a long study history, a large-scale analysis of protein complex in binary PPI network of E. coli is still absent. We used a novel link clustering algorithm named ELPA to infer protein complexes and functional modules in E. coli PPI network. By mapping our results to 276 gold standard protein complexes and protein function annotations offered by EcoCyc, we found that 80.2% of predicted modules mapping well with one or more complexes, while 92.8% of predicted modules tally well with certain GO terms. Furthermore, we compare our results with MCL algorithm, and evaluated our results with several accuracy measures and biological relevance, the result shows that ELPA achieved an average 18.3% improvement over MCL based on the accuracy measures, which means our method will contributes to uncover the complexes of Ecoli.
Wei Liu 0154, Aiping Wu 0002
BIBM2
2014 Exploring protein domain organization by recognition of secondary structure packing interfaces
abstract
MOTIVATION: Protein domains are fundamental units of protein structure, function and evolution; thus, it is critical to gain a deep understanding of protein domain organization. Previous works have attempted to identify key residues involved in organization of domain architecture. Because one of the most important characteristics of domain architecture is the arrangement of secondary structure elements (SSEs), here we present a picture of domain organization through an integrated consideration of SSE arrangements and residue contact networks. RESULTS: In this work, by representing SSEs as main-chain scaffolds and side-chain interfaces and through construction of residue contact networks, we have identified the SSE interfaces well packed within protein domains as SSE packing clusters. In total, 17 334 SSE packing clusters were recognized from 9015 Structural Classification of Proteins domains of <40% sequence identity. The similar SSE packing clusters were observed not only among domains of the same folds, but also among domains of different folds, indicating their roles as common scaffolds for organization of protein domains. Further analysis of 14 small single-domain proteins reveals a high correlation between the SSE packing clusters and the folding nuclei. Consistent with their important roles in domain organization, SSE packing clusters were found to be more conserved than other regions within the same proteins. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Lizong Deng, Aiping Wu 0002, Wentao Dai, Tingrui Song, Ya Cui, Taijiao Jiang
Bioinform.2
2011 NCACO-score: An effective main-chain dependent scoring function for structure modeling
abstract
BACKGROUND: Development of effective scoring functions is a critical component to the success of protein structure modeling. Previously, many efforts have been dedicated to the development of scoring functions. Despite these efforts, development of an effective scoring function that can achieve both good accuracy and fast speed still presents a grand challenge. RESULTS: Based on a coarse-grained representation of a protein structure by using only four main-chain atoms: N, Cα, C and O, we develop a knowledge-based scoring function, called NCACO-score, that integrates different structural information to rapidly model protein structure from sequence. In testing on the Decoys'R'Us sets, we found that NCACO-score can effectively recognize native conformers from their decoys. Furthermore, we demonstrate that NCACO-score can effectively guide fragment assembly for protein structure prediction, which has achieved a good performance in building the structure models for hard targets from CASP8 in terms of both accuracy and speed. CONCLUSIONS: Although NCACO-score is developed based on a coarse-grained model, it is able to discriminate native conformers from decoy conformers with high accuracy. NCACO is a very effective scoring function for structure modeling.
Liqing Tian, Aiping Wu 0002, Xiaoxi Dong, Taijiao Jiang
BMC Bioinform.2
2010 Correlation of Influenza Virus Excess Mortality with Antigenic Variation: Application to Rapid Estimation of Influenza Mortality Burden
abstract
The variants of human influenza virus have caused, and continue to cause, substantial morbidity and mortality. Timely and accurate assessment of their impact on human death is invaluable for influenza planning but presents a substantial challenge, as current approaches rely mostly on intensive and unbiased influenza surveillance. In this study, by proposing a novel host-virus interaction model, we have established a positive correlation between the excess mortalities caused by viral strains of distinct antigenicity and their antigenic distances to their previous strains for each (sub)type of seasonal influenza viruses. Based on this relationship, we further develop a method to rapidly assess the mortality burden of influenza A(H1N1) virus by accurately predicting the antigenic distance between A(H1N1) strains. Rapid estimation of influenza mortality burden for new seasonal strains should help formulate a cost-effective response for influenza control and prevention.
Aiping Wu 0002, Yousong Peng, Xiangjun Du, Yuelong Shu, Taijiao Jiang
PLoS Comput. Biol.1