VLDB 2026 Research / reviewers in the wild / expert
Juan Wang 0011
dblp:74/3634-11
· DBLP profile ↗
22ranked-venue papers
2as first author
20since 2021 · last 2025
0000-0002-8289-1614ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 20 · 2 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SWEC-DTA: Slide-Window Transformer and ECAttention Method for Predicting Drug-Target Binding AffinityabstractPrecise drug-target affinity (DTA) prediction is essential for accelerating drug discovery pipeline. To improve prediction accuracy, in this paper, we utilized two new datasets based on the BindingDB database, which have been updated with newly discovered DTA data. We propose a new method called SWEC-DTA, which based sliding window attention mechanism and efficient channel attention mechanism. It takes the SMILES sequence of the drug and the amino acid sequence of the protein as input. Sliding window transformer (SWIN) is used to obtain 1D sequence features for both the protein and the drug. After converting the drug's SMILES sequence into a molecular graph and molecular fingerprints (MACCS, PubChem, ErG) are combined to enhance the 2D representation. ESM is employed to obtain protein evolutionary and functional features. Finally, ECAttention is used to dynamically learn the interaction relationships between features. SWEC-DTA has been trained and tested on three datasets. Experimental results demonstrate that SWEC-DTA significantly outperforms other baseline methods. Additionally, we conducted ablation experiments and parameter optimization on DB_KD to validate the effectiveness of the method. Yu-Jiang Cheng, Juan Wang 0011, Shang-Jun Yang, Xiang-Zhen Song |
BIBM | 2 |
| 2025 | AFMCDR: An Adaptive Fusion Model to Predict Cancer Drug Responses Based on Hierarchical Feature Enhancement from Multi-Omics of Driver MutationsabstractPrediction of cancer drug response based on deep learning methods has become the basis for personalized medicine. A large amount of multi-omics data on human cell lines and anti-cancer drugs are available for this task. Feature extraction from multi-omics data is essential for this task. Driver mutations affect key signaling pathways, which can significantly alter the sensitivity of tumor cells to drugs and affect the prediction of drug response. However, the contribution of cancer driver mutations to model performance has not been fully investigated. We propose AFMCDR, a novel approach that addresses key challenges in predicting cancer drug response. AFMCDR innovatively utilizes multi-omics data on driver mutations in cell lines and multidimensional drug data, filling a gap in understanding how these data affect drug response mechanisms, a key research area; we propose a hierarchical feature enhancement network that captures both the histological feature level and deep associations at the biological entity level; and we propose a biomedical crossnetwork attention mechanism to capture associations between biomedical entities. In addition, we design a representation learning method that combines drug responses with biomedical information. Finally, we propose a biological network-based adaptive fusion mechanism for dynamically optimizing the fusion of multi-omics features. Experiments on two benchmark datasets show that AFMCDR outperforms existing methods in predicting cancer drug responses. Visualization analysis shows that driver mutation features significantly improves the performance of drug response prediction. These findings confirm that AFMCDR is a powerful tool for predicting cancer drug response. The source code and datasets of AFMCDR are available at https://anonymous.4open.science/r/AFMCDR-DA65/. Weiliang Han, Zhensong Wang, Jiajie Xing, Juan Wang 0011 |
BIBM | 6 |
| 2025 | Stnet: an Efficient Framework for Phylogenetic Network Construction Via Supertree-Based ClusteringabstractPhylogenetic networks are essential for revealing reticulate evolutionary relationships such as hybridization and horizontal gene transfer, but balancing accuracy and scalability remains challenging. Existing tree-input-based methods often suffer from reduced accuracy and inefficiency when processing large datasets due to topological conflicts or high complexity. To address this, we propose STNET, a phylogenetic network construction algorithm based on a three-stage clus-tering-supertree-network framework. First, an improved kmeans++ clustering with RF distance preserves intra-cluster consistency and reduces computational scale. Second, within each cluster, a BCD+ supertree is constructed using source tree ordering and incremental max-flow strategies, ensuring accurate integration of evolutionary information. Finally, supertrees are input into Fast-RF-Net to generate candidate networks, and the optimal one is selected by minimizing cluster distance. Comparative experiments on simulated datasets ($6-20$taxa,$100-1000$trees) and real datasets (placental mammals, marsupials, OMM) show that STNET outperforms mainstream methods (Fast-RFNet, MP, CoalRe, PhyG) in equivalent distance, cluster distance, and runtime. On real datasets, it achieves significantly faster runtime while maintaining comparable accuracy. STNET provides an effective solution to the accuracy-scalability trade-off in large-scale phylogenetic network construction. Yuxin Hao, Juan Wang 0011 |
BIBM | 3 |
| 2025 | HiFusion-Pro: Geometric-Structural Aware Protein Function Prediction via Hierarchical Interaction Fusion and Multi-Task LearningabstractAccurate protein function prediction, such as assigning Enzyme Commission (EC) numbers and Gene Ontology (GO) terms, is a fundamental challenge in bioinformatics. We propose HiFusion-Pro, a deep learning framework that enhances prediction accuracy through two key innovations. First, we introduce a hierarchical fusion architecture that deeply integrates sequence and structure information. This is achieved by initializing a structural encoder (GearNet) with features from a protein language model (ESM) for early information sharing, and by aggregating multi-level intermediate representations from both encoders. This approach overcomes the limitations of conventional late-stage fusion. Second, we employ a multitask learning strategy that co-predicts protein function (primary task) with inter-residue distances (auxiliary task). This strategy regularizes the model, promoting robust feature learning that captures both global functional patterns and local geometric details. Experimental results demonstrate that HiFusion-Pro significantly outperforms state-of-the-art methods across multiple benchmarks. The codes for HiFusion-Pro and datasets are available at https://github.com/YuQing-cs/HiFusion-Pro. Zhongyu Hu, Juan Wang 0011, Guanglai Gao |
BIBM | 2 |
| 2025 | EGSDRP: Enhanced Graph Attention Networks and Drug Substructures for Cancer Drug Response PredictionabstractPredicting cancer drug response is crucial for personalized cancer therapy. While numerous deep learning methods have been developed for predicting cancer drug response, many focus solely on omics data, neglecting the valuable information contained within gene-gene interactions. Even though some studies have employed gene relationship networks, they still have not effectively extracted cell line features. Furthermore, drug feature extraction often overlooks substructure-based learning. In this study, we propose EGSDRP, a novel model designed to address these limitations. EGSDRP incorporates a new graph attention network based on Lipschitz normalization to extract features from three types of cancer cell line omics data: gene expression, copy number, and mutation. Additionally, it employs a substructure feature extraction module that integrates a gated message passing network and a Transformer encoder to dynamically capture drug substructure features. Experimental results demonstrate that EGSDRP achieves a Pearson Correlation Coefficient (PCC) of 0.944 on the GDSC2 dataset, outperforming other benchmark models. Significantly, our model successfully prioritized clinically validated anti-cancer agents, including Docetaxel and Epothilone B, underscoring its potential for effective drug discovery. Shangjun Yang, Juan Wang 0011, Xiangzhen Song, Weiliang Han |
BIBM | 4 |
| 2025 | HyperCausal: A Multi-Task Causal Learning Framework for Cancer Drug Response PredictionabstractAccurate prediction of anti-cancer drug response requires simultaneous assessment of both categorical sensitivity and Individual Treatment Effects (ITE). Current computational approaches predominantly focus on correlations rather than causal mechanisms. We present HyperCausal, a novel multi-task framework integrating hypergraph representation learning with causal inference for robust drug response prediction and personalized treatment selection. HyperCausal constructs bipartite graphs and hypergraphs to capture complex biological relationships. The framework explicitly models network inter-ference effects through hypergraph convolutional networks. To address confounding bias, we employ representation balancing based on Wasserstein distance for robust causal effect estimation. Evaluation on GDSC and CCLE datasets achieves state-of-the-art performance. The framework reaches AUC of 0.9074 and AP of 0.9073 on GDSC dataset. Comprehensive ablation studies confirm the criticality of causal inference components. EGFR case studies demonstrate clinical relevance through correct stratification based on mutation status. This establishes a principled approach for computational drug discovery with mechanistic understanding. Code and data are available https://anonymous.4open.science/r/HyperCausal-2436 for reproducibility. Jiajie Xing, Juan Wang 0011 |
BIBM | 3 |
| 2025 | SAGA-DRP: A Gradient-Guided Alignment and Momentum Prototyping Framework for Stable Cross-Domain Drug Response PredictionabstractDrug response prediction (DRP) is crucial for precision oncology. Current models exhibit poor generalization from preclinical cell lines to patient tumors due to domain shift. While unsupervised domain adaptation offers a solution, existing frameworks suffer from two training instabilities. These include gradient conflicts between competing objectives and prototype instability from high-variance cancer subtype representations. These fundamental issues prevent reliable identification of transferable biomarkers essential for clinical decision-making. We propose SAGA-DRP, a framework with dual stabilization mechanisms grounded in cancer biology. Our gradient-guided annealing identifies parameter regions corresponding to conserved biological pathways. This resolves destructive interference between objectives. Momentum-smoothed prototype updates stabilize molecular signatures of cancer subtypes. This ensures consistent drug response pattern capture across domains. SAGA-DRP represents the first framework to simultaneously address both stability challenges through biologically-motivated solutions. Evaluations on TCGA and PDTC datasets across nine FDA-approved drugs demonstrate significant improvements. SAGA-DRP achieves superior accuracy with AUC of 0.6851, representing a 1.9% improvement. The framework also shows enhanced stability with 73% prototype variance reduction. This provides a robust foundation for clinical translation in precision oncology. Statistical significance testing confirms the reliability of improvements. Pathway analysis validates that identified biomarkers align with known cancer biology. To ensure reproducibility, data and code are accessible via anonymous links at https://anonymous.4open.science/r/SAGA-DRP-34CC. Jiajie Xing, Juan Wang 0011 |
BIBM | 3 |
| 2025 | Disease-Gene Association Prediction via Hybrid Negative Sampling and Contrastive Leaming-Driven Variational Graph Auto-EncoderabstractThe accurate prediction of genedisease associations is essential for understanding pathogenic mechanisms and identifying therapeutic targets. While existing computational methods have integrated diverse biological information, they struggle with the sparsity of known associations and the lack of reliable negative samples. To address these issues, we propose NCVGAE, a novel framework that extends the Variational Graph Auto-Encoder (VGAE) by incorporating a hybrid negative sampling strategy and contrastive learning to enhance representation learning under sparse supervision. We devise a Hybrid Negative Sampling strategy that combines similarity-guided selection-using topological and biological similarity to identify high-confidence negatives-with random sampling to ensure diversity. This helps reduce noise from unreliable negatives and improves the quality of supervision. The encoder module employs a Graph Neural Network (GNN) to effectively learn robust embeddings for genes and diseases. During the training phase, a contrastive learning objective refines embeddings by pulling together true associations and pushing apart negatives. A multilayer perceptron then predicts association scores. Experiments on three public datasets show that NCVGAE consistently outperforms state-of-the-art methods in terms of AUC and AUPR, highlighting its effectiveness and generalizability. Jiajie Xing, Yuwei Sun, Juan Wang 0011, Shangjun Yang, Yujiang Cheng |
BIBM | 4 |
| 2025 | GATOmics: A Novel Multi-Omics Graph Attention Network Model for Cancer Driver Gene DetectionabstractIdentifying cancer driver genes remains challenging due to the complexity of gene interactions in cancer genomics. Existing methods often face difficulties in integrating multidimensional biological data, which limits their ability to capture diverse gene relationships. GATOmics, a novel multi-omics framework, addresses this by integrating four networks, protein-protein interactions, tissue co-expression, pathway co-occurrence, and gene semantic similarity. Utilizing a graph attention network to enhance feature extraction. By employing self-attention mechanisms and convolutional modules, GATOmics captures long-range gene interactions, improving prediction accuracy. The use of semi-supervised learning enables the model to leverage both labeled and unlabeled data, enhancing generalization across different cancer types. Evaluations on pan-cancer datasets demonstrate that GATOmics consistently outperformed state-of-the-art methods, achieving higher AUC and AUPRC scores, highlighting its potential for broader applications in cancer genomics research. The code and data for GATOmics are available at https://github.com/ggkong/GATOmics Ge Kong, Juan Wang 0011 |
ICASSP | 3 |
| 2025 | TriFP-NGram: Integrating Three Complementary Fingerprint and N-Gram Features for Enhanced Drug-Target Affinity PredictionabstractAccurate prediction of drug-target affinity plays a vital role in drug discovery and design. TriFP-NGram integrates multiple fingerprint and n-gram features to predict drug-target binding affinities. It surpasses current methods by leveraging a comprehensive set of molecular and protein features, enhancing predictive performance across six datasets. The innovative model architecture allows for deep feature extraction, including graph convolutional networks and selective receptive fields. Its reliable and efficient affinity predictions offer a significant advancement in drug discovery and design. The source code and datasets of TriFp-NGram are available at https://github.com/991212/TriFP-NGram.git Ge Kong, Juan Wang 0011 |
ICASSP | 3 |
| 2025 | Identifying Disease-Gene Associations by Topological and Biological Feature-based Data Augmentation and Graph Neural NetworksabstractPredicting gene-disease associations is essential for understanding disease pathogenesis and determining therapeutic targets. While prior methods have integrated diverse biological information to make predictions, they still encounter several challenges. First, incomplete and sparse gene-disease association data constrain model performance. Second, integrating heterogeneous data sources is not straightforward. To address these challenges, we propose a novel method, DAVGAE, which combines data augmentation, Variational Graph Auto-Encoders (VGAE), and attention mechanisms. DAVGAE integrates both the biological and topological features of genes and diseases to address challenges such as data sparsity and heterogeneity. By leveraging these features, it calculates cosine similarity scores for gene-disease pairs and applies a novel data augmentation strategy to enhance association data by selecting gene-disease associations with higher similarity scores. Using a four-layer Graph Neural Network (GNN) encoder, DAVGAE effectively learns robust and discriminative representations for genes and diseases within the association network. Finally, an inner product decoder predicts association scores for all gene-disease pairs. Comprehensive experiments on three gene-disease association datasets reveal that DAVGAE outperforms baseline models in predicting gene-disease associations. Juan Wang 0011, Jiajie Xing |
IEEE J. Biomed. Health Informatics | 2 |
| 2024 | CDM-MBT: Context-Dependent Modeling With Memory Bank Traceback for Emotion Recognition in ConversationabstractEmotion Recognition in Conversation (ERC) is a key component in affective dialogue systems, which helps the system understand users’ emotions and generate empathetic responses. The typical approaches for modeling context have been to generate local representation of each utterance in a conversation and then perform contextual modeling of these. However, these approaches are insufficient in understanding the context due to lacking the ability to jointly capture both utterance- and conversation-level contextual information. To better model the context, we propose Context-Dependent Modeling (CDM) to capture context-dependent representation of the query utterance by utilizing the powerful representation capability of pre-trained language models. We input both the query utterance and its previous conversational context into RoBERTa and discard the need to model the context separately after obtaining the representation, as it already provides an effective embedding of such context. Besides, the emotion of the query utterance can be traced back to the stable mood of the speaker, which is inherited from previous emotional states. We design Memory Bank Traceback (MBT) to track and summarize the speaker’s emotional inertia. We treat MBT as an auxiliary information source that enriches the representation of the query utterance. Extensive experiments on four benchmark datasets demonstrate that CDM-MBT outperforms most baseline models. Juan Wang 0011 |
BIBM | 3 |
| 2024 | SLP-T5: A New Method for Predicting Protein Sub-chloroplast Localization Based on Text-to-Text Transfer TransformerabstractChloroplasts convert light energy into chemical energy and are essential for photosynthesis. Proteins in different chloroplast positions have different functions. Thus, precise prediction of sub-chloroplast localization of an unknown protein is crucial for understanding related biochemical processes. The state-of-the-art methods for this task are mainly based on machine learning and deep learning, but they continue to struggle with data imbalance and the lack of training data, both of which result in low accuracy. Here, we proposed a new method for predicting Sub-chloroplast Localization Proteins based on the Text-to-Text Transfer Transformer (T5) model, called SLP-T5, which included a feature extraction module with ProtT5, utilized the Multi-Labeled Data Augmentation module to generate additional training data and address the lack of training data, and employed a multi-label predictor with a Threshold Selection Module to combat data imbalance and improve prediction accuracy. Experimental results showed that SLP-T5 outperformed current methods, increasing the overall actual accuracy by 9% to 92.5%. The SLP-T5 provided further evidence for the localization of predicted chloroplast proteins in the UniProtKB/Swiss-Prot database, with 95% accuracy, and formed a comprehensive documentation. SLP-T5 significantly improved the performance of protein sub-chloroplast localization, providing a robust, accurate, and efficient predictor, showcasing a successful application of large language models in protein function prediction. The codes for SLP-T5 and datasets are available at https://anonymous.4open.science/r/SLP-T5-0809. Ge Kong, Yuanhao Fan, Juan Wang 0011 |
BIBM | 4 |
| 2024 | MLFF-DTA: A Multi-Level Feature Fusion Method for Predicting Drug-Target Binding AffinityabstractAccurate prediction of drug-target binding affinity (DTA) is crucial for drug discovery. In deep learning-based methods, robust feature representations of drugs and targets and their interaction features play a key role in improving the accuracy of DTA prediction. Additionally, biological data on DTA have been significantly updated in recent years, and the ability to predict and identify new data is also an important consideration. In this paper, we have constructed two new datasets to update the newly discovered DTA data. We propose a new method based on Multi-Level Feature Fusion for predicting DTA, called MLFF-DTA. It uses SMILES strings of compounds and amino acid sequences of proteins as inputs and then extracts molecular graph and fingerprint features for compounds and n-gram features for proteins. Here, three kinds of complementary fingerprints, i.e., MACCS, PubChem, and Pharmacophore ErG fingerprints, are fused as a new fingerprint feature, called FP, to describe the physicochemical properties of compounds. The n-gram features describe the types and quantities of amino acid functional groups in protein sequences. MLFF-DTA is trained and tested on three datasets, i.e., a public dataset and two new datasets. MLFF-DTA was also tested for the ability of generalization on four small datasets, in which each protein is not included in training sets. The experimental results demonstrate that the performance of MLFF-DTA is superior to other baseline methods. Furthermore, we apply MLFF-DTA to discover three potential drugs for breast cancer with support from relevant clinical experiments. The source code and datasets of MLFF-DTA are available at https://anonymous.4open.science/r/MLFF-DTA. Ge Kong, Juan Wang 0011 |
BIBM | 3 |
| 2024 | Predicting miRNA-disease association based on Hybrid Graph AutoencoderabstractSince microRNA (miRNA) can participate in the post-transcriptional regulation of gene expression, they can become potential markers of diseases, therapeutic targets and regulatory factors of disease development. So it is vital to predict miRNA-disease association(MDA). However, recent researches lack feature learning based on diverse views, and pay little attention to the sequence features of miRNA. In this study, we extract the initial feature from miRNA sequences and build homogeneous graph and heterogeneous graph for miRNA-disease association. To effectively learn low-dimensional representations of miRNA and disease features from different perspectives, we propose a Hybrid Graph Autoencoder model (HGAMDA). The model learn the embedding of nodes from multiple homogeneous graphs separately, utilizing graph-level attention mechanisms to learn the importance of different homogeneous graphs. Then, it further learns low-dimensional features of nodes in the heterogeneous graph. The experimental results show that the AUC value of HGAMDA in 5-fold cross-validation is 0.9878 on HMDD v4.0 dataset, significantly outperforming other baseline methods. Jiajie Xing, Juan Wang 0011 |
BIBM | 3 |
| 2024 | MDMD: A Computational Model for Predicting Drug-Related Microbes Based on the Aggregated Metapaths from a Heterogeneous NetworkabstractClinical studies have shown that microbes in the human body are closely related to human health. Microbes can influence the activity and toxicity of drugs. So they play an important role in the treatment of diseases. It is critical to research the associations between drugs and microbes for drug development and precision medicine. Recently, there are several computation methods for predicting drug-related microbes. However, these methods ignore the information of diseases because diseases are the bridge between drugs and microbes. Here we introduce a new model (called MDMD) proposed to predict drug-related microbes based on the Metapaths from a heterogeneous network constructed by using the data of Diseases, Microbes, Drugs, the associations of microbe-disease and disease-drug. The MDMD uses an aggregation of the metapath features that can effectively abundance the embedding of the features for different types of nodes and edges in the heterogeneous networks. Then, the MDMD uses the attention mechanism to mark the importance of the metapath vector for each node type which can improve the quality of feature embedding. Experimental results demonstrate that the MDMD improves accuracy by 1.9% compared with other models. The MDMD is also used to predict the microbes of two drugs Lamivudine and Tenofovir which are the antiretroviral drugs used to treat the Acquired Immune Deficiency Syndrome(AIDS). The results show that 90-95% of microbes are reported in the PubMed. In addition, we found that lamivudine may be useful for the treatment of tuberculosis caused by Mycobacterium tuberculosis (Mtb). An online platform of the MDMD is available in https://mdmd2023.bit1024.top/, in which the source code of the MDMD and the data in the work can be downloaded. Jiajie Xing, Juan Wang 0011 |
BIBM | 4 |
| 2023 | Prediction of Oncology Drug Targets Based on Ensemble Learning and Sample Weight UpdatingabstractThe selection of targets is a critical function in the design and development of new drugs. However, finding targets solely through biological experiments is time-consuming and unrealistic due to the complexity of drug-target interactions. Mining potential drug targets based on machine learning methods can significantly shorten the drug development cycle. In the field of oncology, there are only a limited number of proteins that are approved drug targets. Therefore, it is a typical data imbalance problem for predicting oncology drug targets from a large number of proteins. Here, we propose an improved AdaBoost algorithm called HardBoost. The algorithm optimizes the weight update strategy of AdaBoost. First, HardBoost makes the model pay more attention to hard-to-classify samples by adding weight control coefficients. Then, in order to reduce the impact of noisy data on the model, the algorithm evaluates the noise samples. Based on the classification results of the base classifier in the Boosting process, the algorithm then adjusts the weight of these noise samples. Finally, we use the HardBoost algorithm as the base classifier of EasyEnsemble (EE) to get an ensemble model called EE-HardBoost. Experimental results show that EE-HardBoost effectively improves the classification performance of imbalanced data. EE-HardBoost is freely available at https://github.com/imustu/EE-HardBoost. Juan Wang 0011 |
BIBM | 3 |
| 2022 | A new method for predicting plant proteins function based on multi label classification algorithmabstractProtein function annotation is an important content of bioinformatics. It is unrealistic to experimentally annotate a large number of protein functions, so automated prediction of the functions of proteins is required. Studying plant proteins can help us cultivate new plant varieties. However, the existing methods mainly aim at predicting the single function of plant proteins. Most proteins only have sequence information. Therefore, the multi-function prediction method of plant protein based on sequence has become the focus of research. In this study, we propose a method based on sequence to predict the function of plant proteins, called PlantGO. We describe the functional annotation problem as a multi-label classification problem using Gene Ontology (GO) terminology, and predict multiple functions for plant proteins. PlantGO extracts features from three aspects and performs feature fusion respectively. To avoid redundant information in features, PlantGO uses the random forest to select features and obtains three groups of optimal features. PlantGO uses a multi-label learning algorithm MLKNN to predict protein function, which proves the effectiveness and accuracy of the algorithm. Finally, PlantGO adopts the ensemble learning method to integrate all models into a unified model, further improving the prediction performance. Therefore, PlantGO has achieved excellent performance. The performance metrics are Accuracy (0.936), F1 Score (0.944), Ranking loss (0.011), which is better than the previous method in the independent test. We have developed an online server to predict the functions of plant proteins. Juan Wang 0011, Maozu Guo 0001 |
BIBM | 2 |
| 2022 | Sub-Loc: predicting protein sub-mitochondrial localization based on sequence embeddingabstractMitochondria are subcellular organelles existing in most eukaryotic organisms. They have a pivotal role in lots of bio-chemical processes for cells. Proteins in different compartments of mitochondria have their transport routes. Locating proteins in mitochondria can provide a solid foundation for mitochondrial pathologies. So far, there have been several computational methods for solving the issue. However, their accuracy is so low that they do not identify the localization precisely. We develop an unsupervised learning model to represent mitochondrial proteins as n-dimensional vectors, called sequence embedding, which can learn the global and context information from mitochondrial proteins. We design a new model, called Sub-Loc, to predict the sub-mitochondrial localization of proteins using the SVM classifier and the sequence embedding method. The sequence embedding method and the Sub-Loc are tested by experiments. Experimental results show the sequence embedding method remarkably enhances the performance of prediction for sub-mitochondrial localization of proteins compared with other feature representations. The Sub-Loc outperforms other approaches for predicting sub-mitochondrial localization of proteins. Juan Wang 0011, Haodong Bian, Maozu Guo 0001 |
BIBM | 2 |
| 2022 | DeepRCI: Predicting ATP-Binding Proteins Using the Residue-Residue Contact InformationabstractAdenine-5'-triphosphate (ATP) is a direct energy source for various activities of tissues and cells in the body. The release of ATP energies requires the assistance of ATP-binding proteins. Therefore, the identification of ATP-binding proteins is of great significance for the research on organisms. So far, there are several methods for predicting ATP-binding proteins. However, the accuracies of these methods are so low that the predicted proteins are inaccurate. Here, we designed a novel method, called as DeepRCI (based on Deep convolutional neural network and Residue-residue Contact Information), for predicting ATP-binding proteins. In order to maximize the performance of our method, we experimented with different hyperparameters and finally chose a 12-depth-512-filters deep convolutional neural network with an input size of 448*448. By using this model, DeepRCI achieved an accuracy of 93.61% on the test set which means a significant improvement of 11.78% over the state-of-the-art methods. We also compared the performance of residue-residue contact information datasets with different noise levels which are mainly due to gaps in the multiple sequence alignment. Compared with the low-noise dataset, the prediction accuracy on the high-noise dataset is reduced by 6.78%, which affects the performance of DeepRCI to a certain extent. We believe that with the increase of sequence data, this problem will eventually be solved. Finally, we provide a web service of DeepRCI which link can be obtained in Data Availability. Yulan Zhao, Juan Wang 0011, Maozu Guo 0001 |
IEEE J. Biomed. Health Informatics | 3 |
| 2019 | A review of metrics measuring dissimilarity for rooted phylogenetic networksabstractA rooted phylogenetic network is an important structure in the description of evolutionary relationships. Computing the distance (topological dissimilarity) between two rooted phylogenetic networks is a fundamental in phylogenic analysis. During the past few decades, several polynomial-time computable metrics have been described. Here, we give a comprehensive review and analysis on those metrics, including the correlation among metrics and the distribution of distance values computed by each metric. Moreover, we describe the software and website, CDRPN (Computing Distance for Rooted Phylogenetic Networks), for measuring the topological dissimilarity between rooted phylogenetic networks. AVAILABILITY: http://bioinformatics.imu.edu.cn/distance/. CONTACT: [email protected]. Juan Wang 0011, Maozu Guo 0001 |
Briefings Bioinform. | 1 |
| 2013 | Lnetwork: an efficient and effective method for constructing phylogenetic networksabstractMOTIVATION: The evolutionary history of species is traditionally represented with a rooted phylogenetic tree. Each tree comprises a set of clusters, i.e. subsets of the species that are descended from a common ancestor. When rooted phylogenetic trees are built from several different datasets (e.g. from different genes), the clusters are often conflicting. These conflicting clusters cannot be expressed as a simple phylogenetic tree; however, they can be expressed in a phylogenetic network. Phylogenetic networks are a generalization of phylogenetic trees that can account for processes such as hybridization, horizontal gene transfer and recombination, which are difficult to represent in standard tree-like models of evolutionary histories. There is currently a large body of research aimed at developing appropriate methods for constructing phylogenetic networks from cluster sets. The Cass algorithm can construct a much simpler network than other available methods, but is extremely slow for large datasets or for datasets that need lots of reticulate nodes. The networks constructed by Cass are also greatly dependent on the order of input data, i.e. it generally derives different phylogenetic networks for the same dataset when different input orders are used. RESULTS: In this study, we introduce an improved Cass algorithm, Lnetwork, which can construct a phylogenetic network for a given set of clusters. We show that Lnetwork is significantly faster than Cass and effectively weakens the influence of input data order. Moreover, we show that Lnetwork can construct a much simpler network than most of the other available methods. AVAILABILITY: Lnetwork has been built as a Java software package and is freely available at http://nclab.hit.edu.cn/∼wangjuan/Lnetwork/. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Juan Wang 0011, Maozu Guo 0001, Yang Liu 0006, Chunyu Wang 0002, Linlin Xing, Kai Che |
Bioinform. | 1 |