Douglas E. V. Pires

dblp:08/10819 · also Douglas Pires · DBLP profile ↗
← Back
23ranked-venue papers
4as first author
17since 2021 · last 2025
0000-0002-3004-2119ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 23 · 4 first-author · 17 since 2021
YearPublicationVenuePosition
2025 Uncovering digital overdiagnosis - Quantification and mitigation using clinical trajectories: Heparin-induced thrombocytopenia use case
abstract
OBJECTIVE: Overdiagnosis occurs when abnormalities meeting diagnostic criteria would remain asymptomatic if undiagnosed. Cases initially identified through digital diagnostic tools but later recognised as overdiagnosis are referred to as 'digital overdiagnosis'. Data-driven frameworks to quantify and mitigate overdiagnosis remain limited. This study introduces a framework that integrates clinical trajectories to train a machine learning (ML)-based disease classifier, enabling the quantification and mitigation of digital overdiagnosis, using Heparin-Induced Thrombocytopenia (HIT) as a case study. METHODS: A pre-existing HIT classifier identified HIT-positive and HIT-negative cases, with ground truth based on HIT diagnostic criteria. Clinical trajectories for True Positive (TP) and True Negative (TN) patients were clustered using a novel process-models-based approach. Overdiagnosis was detected when TP cases clustered with predominantly TN cases. The classifier was then retrained with an 'updated label' integrating both HIT criteria and the concordant trajectory, to reduce overdiagnosis while maintaining accuracy. RESULTS: 7.2% of TP cases were identified as overdiagnosed. Retraining with the updated labels successfully reclassified 89.5% of overdiagnosed cases as TN, with only a minimal reduction in performance (MCC decreased by 0.03, positive likelihood ratio decreased by 0.49, and negative likelihood ratio increased by 0.05). Clinical outcomes-length of stay, thrombotic events, and mortality-differed significantly between non-overdiagnosed and overdiagnosed cases, and between non-overdiagnosed and TN cases, but not between overdiagnosed and TN cases, confirming that overdiagnosed patients resemble TN patients. CONCLUSION: Incorporating clinical trajectories into ML-based diagnosis enables the quantification of digital overdiagnosis. This approach could refine ML algorithms by prompting a reassessment of criteria-based disease labels in supervised learning.
Prabodi Senevirathna, Douglas E. V. Pires, Daniel Capurro
J. Biomed. Informatics2
2024 Large scale paired antibody language models
abstract
Antibodies are proteins produced by the immune system that can identify and neutralise a wide variety of antigens with high specificity and affinity, and constitute the most successful class of biotherapeutics. With the advent of next-generation sequencing, billions of antibody sequences have been collected in recent years, though their application in the design of better therapeutics has been constrained by the sheer volume and complexity of the data. To address this challenge, we present IgBert and IgT5, the best performing antibody-specific language models developed to date which can consistently handle both paired and unpaired variable region sequences as input. These models are trained comprehensively using the more than two billion unpaired sequences and two million paired sequences of light and heavy chains present in the Observed Antibody Space dataset. We show that our models outperform existing antibody and protein language models on a diverse range of design and regression tasks relevant to antibody engineering. This advancement marks a significant leap forward in leveraging machine learning, large scale data sets and high-performance computing for enhancing antibody design for therapeutic development.
Henry Kenlay, Frédéric A. Dreyer, Aleksandr Kovaltsuk, Dom Miketa, Douglas E. V. Pires, Charlotte M. Deane
PLoS Comput. Biol.5
2023 epitope1D: accurate taxonomy-aware B-cell linear epitope prediction
abstract
The ability to identify B-cell epitopes is an essential step in vaccine design, immunodiagnostic tests and antibody production. Several computational approaches have been proposed to identify, from an antigen protein or peptide sequence, which residues are more likely to be part of an epitope, but have limited performance on relatively homogeneous data sets and lack interpretability, limiting biological insights that could otherwise be obtained. To address these limitations, we have developed epitope1D, an explainable machine learning method capable of accurately identifying linear B-cell epitopes, leveraging two new descriptors: a graph-based signature representation of protein sequences, based on our well-established Cutoff Scanning Matrix algorithm and Organism Ontology information. Our model achieved Areas Under the ROC curve of up to 0.935 on cross-validation and blind tests, demonstrating robust performance. A comprehensive comparison to alternative methods using distinct benchmark data sets was also employed, with our model outperforming state-of-the-art tools. epitope1D represents not only a significant advance in predictive performance, but also allows biologically meaningful features to be combined and used for model interpretation. epitope1D has been made available as a user-friendly web server interface and application programming interface at https://biosig.lab.uq.edu.au/epitope1d/.
Bruna Moreira da Silva, David B. Ascher, Douglas E. V. Pires
Briefings Bioinform.3
2023 LEGO-CSM: a tool for functional characterization of proteins
abstract
MOTIVATION: With the development of sequencing techniques, the discovery of new proteins significantly exceeds the human capacity and resources for experimentally characterizing protein functions. Localization, EC numbers, and GO terms with the structure-based Cutoff Scanning Matrix (LEGO-CSM) is a comprehensive web-based resource that fills this gap by leveraging the well-established and robust graph-based signatures to supervised learning models using both protein sequence and structure information to accurately model protein function in terms of Subcellular Localization, Enzyme Commission (EC) numbers, and Gene Ontology (GO) terms. RESULTS: We show our models perform as well as or better than alternative approaches, achieving area under the receiver operating characteristic curve of up to 0.93 for subcellular localization, up to 0.93 for EC, and up to 0.81 for GO terms on independent blind tests. AVAILABILITY AND IMPLEMENTATION: LEGO-CSM's web server is freely available at https://biosig.lab.uq.edu.au/lego_csm. In addition, all datasets used to train and test LEGO-CSM's models can be downloaded at https://biosig.lab.uq.edu.au/lego_csm/data.
Thanh-Binh Nguyen 0007, Alex G. C. de Sá, Carlos H. M. Rodrigues, Douglas E. V. Pires, David B. Ascher
Bioinform.4
2023 Understanding the complementarity and plasticity of antibody-antigen interfaces
abstract
MOTIVATION: While antibodies have been ground-breaking therapeutic agents, the structural determinants for antibody binding specificity remain to be fully elucidated, which is compounded by the virtually unlimited repertoire of antigens they can recognize. Here, we have explored the structural landscapes of antibody-antigen interfaces to identify the structural determinants driving target recognition by assessing concavity and interatomic interactions. RESULTS: We found that complementarity-determining regions utilized deeper concavity with their longer H3 loops, especially H3 loops of nanobody showing the deepest use of concavity. Of all amino acid residues found in complementarity-determining regions, tryptophan used deeper concavity, especially in nanobodies, making it suitable for leveraging concave antigen surfaces. Similarly, antigens utilized arginine to bind to deeper pockets of the antibody surface. Our findings fill a gap in knowledge about the antibody specificity, binding affinity, and the nature of antibody-antigen interface features, which will lead to a better understanding of how antibodies can be more effective to target druggable sites on antigen surfaces. AVAILABILITY AND IMPLEMENTATION: The data and scripts are available at: https://github.com/YoochanMyung/scripts.
Yoochan Myung, Douglas E. V. Pires, David B. Ascher
Bioinform.2
2023 Developing a deep learning natural language processing algorithm for automated reporting of adverse drug reactions
Christopher McMaster, Julia Chan, David F. L. Liew, Elizabeth Su, Albert G. Frauman, Wendy W. Chapman, Douglas E. V. Pires
J. Biomed. Informatics7
2023 Data-driven overdiagnosis definitions: A scoping review
abstract
INTRODUCTION: Adequate methods to promptly translate digital health innovations for improved patient care are essential. Advances in Artificial Intelligence (AI) and Machine Learning (ML) have been sources of digital innovation and hold the promise to revolutionize the way we treat, manage and diagnose patients. Understanding the benefits but also the potential adverse effects of digital health innovations, particularly when these are made available or applied on healthier segments of the population is essential. One of such adverse effects is overdiagnosis. OBJECTIVE: to comprehensively analyze quantification strategies and data-driven definitions for overdiagnosis reported in the literature. METHODS: we conducted a scoping systematic review of manuscripts describing quantitative methods to estimate the proportion of overdiagnosed patients. RESULTS: we identified 46 studies that met our inclusion criteria. They covered a variety of clinical conditions, primarily breast and prostate cancer. Methods to quantify overdiagnosis included both prospective and retrospective methods including randomized clinical trials, and simulations. CONCLUSION: a variety of methods to quantify overdiagnosis have been published, producing widely diverging results. A standard method to quantify overdiagnosis is needed to allow its mitigation during the rapidly increasing development of new digital diagnostic tools.
Prabodi Senevirathna, Douglas E. V. Pires, Daniel Capurro
J. Biomed. Informatics2
2022 CSM-carbohydrate: protein-carbohydrate binding affinity prediction and docking scoring function
abstract
Protein-carbohydrate interactions are crucial for many cellular processes but can be challenging to biologically characterise. To improve our understanding and ability to model these molecular interactions, we used a carefully curated set of 370 protein-carbohydrate complexes with experimental structural and biophysical data in order to train and validate a new tool, cutoff scanning matrix (CSM)-carbohydrate, using machine learning algorithms to accurately predict their binding affinity and rank docking poses as a scoring function. Information on both protein and carbohydrate complementarity, in terms of shape and chemistry, was captured using graph-based structural signatures. Across both training and independent test sets, we achieved comparable Pearson's correlations of 0.72 under cross-validation [root mean square error (RMSE) of 1.58 Kcal/mol] and 0.67 on the independent test (RMSE of 1.72 Kcal/mol), providing confidence in the generalisability and robustness of the final model. Similar performance was obtained across mono-, di- and oligosaccharides, further highlighting the applicability of this approach to the study of larger complexes. We show CSM-carbohydrate significantly outperformed previous approaches and have implemented our method and make all data freely available through both a user-friendly web interface and application programming interface, to facilitate programmatic access at http://biosig.unimelb.edu.au/csm_carbohydrate/. We believe CSM-carbohydrate will be an invaluable tool for helping assess docking poses and the effects of mutations on protein-carbohydrate affinity, unravelling important aspects that drive binding recognition.
Thanh-Binh Nguyen 0007, Douglas E. V. Pires, David B. Ascher
Briefings Bioinform.2
2022 GASS-Metal: identifying metal-binding sites on protein structures using genetic algorithms
abstract
Metals are present in >30% of proteins found in nature and assist them to perform important biological functions, including storage, transport, signal transduction and enzymatic activity. Traditional and experimental techniques for metal-binding site prediction are usually costly and time-consuming, making computational tools that can assist in these predictions of significant importance. Here we present Genetic Active Site Search (GASS)-Metal, a new method for protein metal-binding site prediction. The method relies on a parallel genetic algorithm to find candidate metal-binding sites that are structurally similar to curated templates from M-CSA and MetalPDB. GASS-Metal was thoroughly validated using homologous proteins and conservative mutations of residues, showing a robust performance. The ability of GASS-Metal to identify metal-binding sites was also compared with state-of-the-art methods, outperforming similar methods and achieving an MCC of up to 0.57 and detecting up to 96.1% of the sites correctly. GASS-Metal is freely available at https://gassmetal.unifei.edu.br. The GASS-Metal source code is available at https://github.com/sandroizidoro/gassmetal-local.
Vinícius de Almeida Paiva, Murillo Ventura Mendonça, Sabrina de Azevedo Silveira, David B. Ascher, Douglas E. V. Pires, Sandro C. Izidoro
Briefings Bioinform.5
2022 Systematic evaluation of computational tools to predict the effects of mutations on protein stability in the absence of experimental structures
abstract
Changes in protein sequence can have dramatic effects on how proteins fold, their stability and dynamics. Over the last 20 years, pioneering methods have been developed to try to estimate the effects of missense mutations on protein stability, leveraging growing availability of protein 3D structures. These, however, have been developed and validated using experimentally derived structures and biophysical measurements. A large proportion of protein structures remain to be experimentally elucidated and, while many studies have based their conclusions on predictions made using homology models, there has been no systematic evaluation of the reliability of these tools in the absence of experimental structural data. We have, therefore, systematically investigated the performance and robustness of ten widely used structural methods when presented with homology models built using templates at a range of sequence identity levels (from 15% to 95%) and contrasted performance with sequence-based tools, as a baseline. We found there is indeed performance deterioration on homology models built using templates with sequence identity below 40%, where sequence-based tools might become preferable. This was most marked for mutations in solvent exposed residues and stabilizing mutations. As structure prediction tools improve, the reliability of these predictors is expected to follow, however we strongly suggest that these factors should be taken into consideration when interpreting results from structure-based predictors of mutation effects on protein stability.
Qisheng Pan, Thanh-Binh Nguyen 0007, David B. Ascher, Douglas E. V. Pires
Briefings Bioinform.4
2022 cropCSM: designing safe and potent herbicides with graph-based signatures
abstract
Herbicides have revolutionised weed management, increased crop yields and improved profitability allowing for an increase in worldwide food security. Their widespread use, however, has also led to a rise in resistance and concerns about their environmental impact. Despite the need for potent and safe herbicidal molecules, no herbicide with a new mode of action has reached the market in 30 years. Although development of computational approaches has proven invaluable to guide rational drug discovery pipelines, leading to higher hit rates and lower attrition due to poor toxicity, little has been done in contrast for herbicide design. To fill this gap, we have developed cropCSM, a computational platform to help identify new, potent, nontoxic and environmentally safe herbicides. By using a knowledge-based approach, we identified physicochemical properties and substructures enriched in safe herbicides. By representing the small molecules as a graph, we leveraged these insights to guide the development of predictive models trained and tested on the largest collected data set of molecules with experimentally characterised herbicidal profiles to date (over 4500 compounds). In addition, we developed six new environmental and human toxicity predictors, spanning five different species to assist in molecule prioritisation. cropCSM was able to correctly identify 97% of herbicides currently available commercially, while predicting toxicity profiles with accuracies of up to 92%. We believe cropCSM will be an essential tool for the enrichment of screening libraries and to guide the development of potent and safe herbicides. We have made the method freely available through a user-friendly webserver at http://biosig.unimelb.edu.au/crop_csm.
Douglas E. V. Pires, Keith A. Stubbs, Joshua S. Mylne, David B. Ascher
Briefings Bioinform.1
2022 Evaluating hierarchical machine learning approaches to classify biological databases
abstract
The rate of biological data generation has increased dramatically in recent years, which has driven the importance of databases as a resource to guide innovation and the generation of biological insights. Given the complexity and scale of these databases, automatic data classification is often required. Biological data sets are often hierarchical in nature, with varying degrees of complexity, imposing different challenges to train, test and validate accurate and generalizable classification models. While some approaches to classify hierarchical data have been proposed, no guidelines regarding their utility, applicability and limitations have been explored or implemented. These include 'Local' approaches considering the hierarchy, building models per level or node, and 'Global' hierarchical classification, using a flat classification approach. To fill this gap, here we have systematically contrasted the performance of 'Local per Level' and 'Local per Node' approaches with a 'Global' approach applied to two different hierarchical datasets: BioLip and CATH. The results show how different components of hierarchical data sets, such as variation coefficient and prediction by depth, can guide the choice of appropriate classification schemes. Finally, we provide guidelines to support this process when embarking on a hierarchical classification task, which will help optimize computational resources and predictive performance.
Pâmela M. Rezende, Joicymara S. Xavier, David B. Ascher, Gabriel R. Fernandes, Douglas E. V. Pires
Briefings Bioinform.5
2022 Structural landscapes of PPI interfaces
abstract
Proteins are capable of highly specific interactions and are responsible for a wide range of functions, making them attractive in the pursuit of new therapeutic options. Previous studies focusing on overall geometry of protein-protein interfaces, however, concluded that PPI interfaces were generally flat. More recently, this idea has been challenged by their structural and thermodynamic characterisation, suggesting the existence of concave binding sites that are closer in character to traditional small-molecule binding sites, rather than exhibiting complete flatness. Here, we present a large-scale analysis of binding geometry and physicochemical properties of all protein-protein interfaces available in the Protein Data Bank. In this review, we provide a comprehensive overview of the protein-protein interface landscape, including evidence that even for overall larger, more flat interfaces that utilize discontinuous interacting regions, small and potentially druggable pockets are utilized at binding sites.
Carlos H. M. Rodrigues, Douglas E. V. Pires, Tom L. Blundell, David B. Ascher
Briefings Bioinform.2
2022 toxCSM: comprehensive prediction of small molecule toxicity profiles
abstract
Drug discovery is a lengthy, costly and high-risk endeavour that is further convoluted by high attrition rates in later development stages. Toxicity has been one of the main causes of failure during clinical trials, increasing drug development time and costs. To facilitate early identification and optimisation of toxicity profiles, several computational tools emerged aiming at improving success rates by timely pre-screening drug candidates. Despite these efforts, there is an increasing demand for platforms capable of assessing both environmental as well as human-based toxicity properties at large scale. Here, we present toxCSM, a comprehensive computational platform for the study and optimisation of toxicity profiles of small molecules. toxCSM leverages on the well-established concepts of graph-based signatures, molecular descriptors and similarity scores to develop 36 models for predicting a range of toxicity properties, which can assist in developing safer drugs and agrochemicals. toxCSM achieved an Area Under the Receiver Operating Characteristic (ROC) Curve (AUC) of up to 0.99 and Pearson's correlation coefficients of up to 0.94 on 10-fold cross-validation, with comparable performance on blind test sets, outperforming all alternative methods. toxCSM is freely available as a user-friendly web server and API at http://biosig.lab.uq.edu.au/toxcsm.
Alex G. C. de Sá, Yangyang Long, Stephanie Portelli, Douglas E. V. Pires, David B. Ascher
Briefings Bioinform.4
2022 epitope3D: a machine learning method for conformational B-cell epitope prediction
abstract
The ability to identify antigenic determinants of pathogens, or epitopes, is fundamental to guide rational vaccine development and immunotherapies, which are particularly relevant for rapid pandemic response. A range of computational tools has been developed over the past two decades to assist in epitope prediction; however, they have presented limited performance and generalization, particularly for the identification of conformational B-cell epitopes. Here, we present epitope3D, a novel scalable machine learning method capable of accurately identifying conformational epitopes trained and evaluated on the largest curated epitope data set to date. Our method uses the concept of graph-based signatures to model epitope and non-epitope regions as graphs and extract distance patterns that are used as evidence to train and test predictive models. We show epitope3D outperforms available alternative approaches, achieving Mathew's Correlation Coefficient and F1-scores of 0.55 and 0.57 on cross-validation and 0.45 and 0.36 during independent blind tests, respectively.
Bruna Moreira da Silva, Yoochan Myung, David B. Ascher, Douglas E. V. Pires
Briefings Bioinform.4
2022 CSM-AB: graph-based antibody-antigen binding affinity prediction and docking scoring function
abstract
MOTIVATION: Understanding antibody-antigen interactions is key to improving their binding affinities and specificities. While experimental approaches are fundamental for developing new therapeutics, computational methods can provide quick assessment of binding landscapes, guiding experimental design. Despite this, little effort has been devoted to accurately predicting the binding affinity between antibodies and antigens and to develop tailored docking scoring functions for this type of interaction. Here, we developed CSM-AB, a machine learning method capable of predicting antibody-antigen binding affinity by modelling interaction interfaces as graph-based signatures. RESULTS: CSM-AB outperformed alternative methods achieving a Pearson's correlation of up to 0.64 on blind tests. We also show CSM-AB can accurately rank near-native poses, working effectively as a docking scoring function. We believe CSM-AB will be an invaluable tool to assist in the development of new immunotherapies. AVAILABILITY AND IMPLEMENTATION: CSM-AB is freely available as a user-friendly web interface and API at http://biosig.unimelb.edu.au/csm_ab/datasets. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yoochan Myung, Douglas E. V. Pires, David B. Ascher
Bioinform.2
2022 Large-scale protein-protein post-translational modification extraction with distant supervision and confidence calibrated BioBERT
abstract
MOTIVATION: Protein-protein interactions (PPIs) are critical to normal cellular function and are related to many disease pathways. A range of protein functions are mediated and regulated by protein interactions through post-translational modifications (PTM). However, only 4% of PPIs are annotated with PTMs in biological knowledge databases such as IntAct, mainly performed through manual curation, which is neither time- nor cost-effective. Here we aim to facilitate annotation by extracting PPIs along with their pairwise PTM from the literature by using distantly supervised training data using deep learning to aid human curation. METHOD: We use the IntAct PPI database to create a distant supervised dataset annotated with interacting protein pairs, their corresponding PTM type, and associated abstracts from the PubMed database. We train an ensemble of BioBERT models-dubbed PPI-BioBERT-x10-to improve confidence calibration. We extend the use of ensemble average confidence approach with confidence variation to counteract the effects of class imbalance to extract high confidence predictions. RESULTS AND CONCLUSION: The PPI-BioBERT-x10 model evaluated on the test set resulted in a modest F1-micro 41.3 (P =5 8.1, R = 32.1). However, by combining high confidence and low variation to identify high quality predictions, tuning the predictions for precision, we retained 19% of the test predictions with 100% precision. We evaluated PPI-BioBERT-x10 on 18 million PubMed abstracts and extracted 1.6 million (546507 unique PTM-PPI triplets) PTM-PPI predictions, and filter [Formula: see text] (4584 unique) high confidence predictions. Of the 5700, human evaluation on a small randomly sampled subset shows that the precision drops to 33.7% despite confidence calibration and highlights the challenges of generalisability beyond the test set even with confidence calibration. We circumvent the problem by only including predictions associated with multiple papers, improving the precision to 58.8%. In this work, we highlight the benefits and challenges of deep learning-based text mining in practice, and the need for increased emphasis on confidence calibration to facilitate human curation efforts.
Aparna Elangovan, Yuan Li 0012, Douglas E. V. Pires, Melissa J. Davis, Karin Verspoor
BMC Bioinform.3
2020 mCSM-AB2: guiding rational antibody design using graph-based signatures
abstract
MOTIVATION: A lack of accurate computational tools to guide rational mutagenesis has made affinity maturation a recurrent challenge in antibody (Ab) development. We previously showed that graph-based signatures can be used to predict the effects of mutations on Ab binding affinity. RESULTS: Here we present an updated and refined version of this approach, mCSM-AB2, capable of accurately modelling the effects of mutations on Ab-antigen binding affinity, through the inclusion of evolutionary and energetic terms. Using a new and expanded database of over 1800 mutations with experimental binding measurements and structural information, mCSM-AB2 achieved a Pearson's correlation of 0.73 and 0.77 across training and blind tests, respectively, outperforming available methods currently used for rational Ab engineering. AVAILABILITY AND IMPLEMENTATION: mCSM-AB2 is available as a user-friendly and freely accessible web server providing rapid analysis of both individual mutations or the entire binding interface to guide rational antibody affinity maturation at http://biosig.unimelb.edu.au/mcsm_ab2. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yoochan Myung, Carlos H. M. Rodrigues, David B. Ascher, Douglas E. V. Pires
Bioinform.4
2020 EasyVS: a user-friendly web-based tool for molecule library selection and structure-based virtual screening
abstract
SUMMARY: EasyVS is a web-based platform built to simplify molecule library selection and virtual screening. With an intuitive interface, the tool allows users to go from selecting a protein target with a known structure and tailoring a purchasable molecule library to performing and visualizing docking in a few clicks. Our system also allows users to filter screening libraries based on molecule properties, cluster molecules by similarity and personalize docking parameters. AVAILABILITY AND IMPLEMENTATION: EasyVS is freely available as an easy-to-use web interface at http://biosig.unimelb.edu.au/easyvs. CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Douglas E. V. Pires, Wandré N. P. Veloso, Yoochan Myung, Carlos H. M. Rodrigues, Michael Silk, Pâmela M. Rezende, Francislon Silva, Joicymara S. Xavier, João P. L. Velloso, Carlos H. Silveira, David B. Ascher
Bioinform.1
2015 PDBest: a user-friendly platform for manipulating and enhancing protein structures
abstract
UNLABELLED: PDBest (PDB Enhanced Structures Toolkit) is a user-friendly, freely available platform for acquiring, manipulating and normalizing protein structures in a high-throughput and seamless fashion. With an intuitive graphical interface it allows users with no programming background to download and manipulate their files. The platform also exports protocols, enabling users to easily share PDB searching and filtering criteria, enhancing analysis reproducibility. AVAILABILITY AND IMPLEMENTATION: PDBest installation packages are freely available for several platforms at http://www.pdbest.dcc.ufmg.br CONTACT: [email protected], [email protected], [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Wellisson R. S. Gonçalves, Valdete M. Gonçalves-Almeida, Aleksander L. Arruda, Wagner Meira Jr., Carlos H. Silveira, Douglas E. V. Pires, Raquel Cardoso de Melo Minardi
Bioinform.6
2014 mCSM: predicting the effects of mutations in proteins using graph-based signatures
abstract
MOTIVATION: Mutations play fundamental roles in evolution by introducing diversity into genomes. Missense mutations in structural genes may become either selectively advantageous or disadvantageous to the organism by affecting protein stability and/or interfering with interactions between partners. Thus, the ability to predict the impact of mutations on protein stability and interactions is of significant value, particularly in understanding the effects of Mendelian and somatic mutations on the progression of disease. Here, we propose a novel approach to the study of missense mutations, called mCSM, which relies on graph-based signatures. These encode distance patterns between atoms and are used to represent the protein residue environment and to train predictive models. To understand the roles of mutations in disease, we have evaluated their impacts not only on protein stability but also on protein-protein and protein-nucleic acid interactions. RESULTS: We show that mCSM performs as well as or better than other methods that are used widely. The mCSM signatures were successfully used in different tasks demonstrating that the impact of a mutation can be correlated with the atomic-distance patterns surrounding an amino acid residue. We showed that mCSM can predict stability changes of a wide range of mutations occurring in the tumour suppressor protein p53, demonstrating the applicability of the proposed method in a challenging disease scenario. AVAILABILITY AND IMPLEMENTATION: A web server is available at http://structure.bioc.cam.ac.uk/mcsm.
Douglas E. V. Pires, David B. Ascher, Tom L. Blundell
Bioinform.1
2013 aCSM: noise-free graph-based signatures to large-scale receptor-based ligand prediction
abstract
MOTIVATION: Receptor-ligand interactions are a central phenomenon in most biological systems. They are characterized by molecular recognition, a complex process mainly driven by physicochemical and structural properties of both receptor and ligand. Understanding and predicting these interactions are major steps towards protein ligand prediction, target identification, lead discovery and drug design. RESULTS: We propose a novel graph-based-binding pocket signature called aCSM, which proved to be efficient and effective in handling large-scale protein ligand prediction tasks. We compare our results with those described in the literature and demonstrate that our algorithm overcomes the competitor's techniques. Finally, we predict novel ligands for proteins from Trypanosoma cruzi, the parasite responsible for Chagas disease, and validate them in silico via a docking protocol, showing the applicability of the method in suggesting ligands for pockets in a real-world scenario. AVAILABILITY AND IMPLEMENTATION: Datasets and the source code are available at http://www.dcc.ufmg.br/∼dpires/acsm. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Douglas E. V. Pires, Raquel Cardoso de Melo Minardi, Carlos H. Silveira, Frederico F. Campos, Wagner Meira Jr.
Bioinform.1
2012 HydroPaCe: understanding and predicting cross-inhibition in serine proteases through hydrophobic patch centroids
abstract
MOTIVATION: Protein-protein interfaces contain important information about molecular recognition. The discovery of conserved patterns is essential for understanding how substrates and inhibitors are bound and for predicting molecular binding. When an inhibitor binds to different enzymes (e.g. dissimilar sequences, structures or mechanisms what we call cross-inhibition), identification of invariants is a difficult task for which traditional methods may fail. RESULTS: To clarify how cross-inhibition happens, we model the problem, propose and evaluate a methodology called HydroPaCe to detect conserved patterns. Interfaces are modeled as graphs of atomic apolar interactions and hydrophobic patches are computed and summarized by centroids (HP-centroids), and their conservation is detected. Despite sequence and structure dissimilarity, our method achieves an appropriate level of abstraction to obtain invariant properties in cross-inhibition. We show examples in which HP-centroids successfully predicted enzymes that could be inhibited by the studied inhibitors according to BRENDA database. AVAILABILITY: www.dcc.ufmg.br/~raquelcm/hydropace CONTACT: [email protected]; [email protected]; [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Valdete M. Gonçalves-Almeida, Douglas E. V. Pires, Raquel Cardoso de Melo Minardi, Carlos H. Silveira, Wagner Meira Jr., Marcelo M. Santoro
Bioinform.2