VLDB 2026 Research / reviewers in the wild / expert
Baldomero Oliva
dblp:o/BaldomeroOliva · also Baldo Oliva
· DBLP profile ↗
24ranked-venue papers
1as first author
5since 2021 · last 2026
0000-0003-0702-0250ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 24 · 1 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ExoFILT: transfer learning for robust and accelerated analysis of exocytosis single-particle tracking dataabstractMOTIVATION: Understanding constitutive exocytosis at the molecular level requires quantitative characterization of protein dynamics during the process. Single-particle tracking allows the measurement of protein dynamics in living cells. However, identifying bona fide exocytic events requires extensive manual annotation, limiting throughput and introducing personal biases that affect reproducibility. RESULTS: We present ExoFILT, a deep learning-based classifier designed to identify exocytic events in single-particle tracking data, using the exocyst complex as a reference. Trained via transfer learning on simulated and experimental data, ExoFILT reduces the time required for manual annotation by ten-fold while improving measurement consistency across researchers. When applied to simultaneous dual-color time-lapse movies, ExoFILT enabled the systematic quantification of temporal relationships between exocytic proteins. The increased throughput uncovered distinct subpopulations of exocytic events with differential molecular composition (e.g. events with and without detectable levels of Sec1), underscoring the potential of ExoFILT to reveal mechanistic insights into exocytosis. AVAILABILITY: All raw data and code used for this article is available in GitHub (https://github.com/GallegoLab/ExoFILT) and Zenodo (https://zenodo.org/records/18962705). Eric Kramer, Laura I Betancur, Sasha Meek, Sébastien Tosi, Carlo Manzo, Baldomero Oliva, Oriol Gallego |
Bioinform. | 6 |
| 2025 | SNPeBoT: a tool for predicting transcription factor allele specific bindingabstractBACKGROUND: Mutations in non-coding regulatory regions of DNA may lead to disease through the disruption of transcription factor binding. However, our understanding of binding patterns of transcription factors and the effects that changes to their binding sites have on their action remains limited. To address this issue we trained a Deep learning model to predict the effects of Single Nucleotide Polymorphisms (SNP) on transcription factor binding. Allele specific binding (ASB) data from Chromatin Immunoprecipitation sequencing (ChIP-seq) experiments were paired with high sequence-identity DNA binding Domains assessed in Protein Binding Microarray (PBM) experiments. For each transcription factor a paired DNA binding Domain was selected from which we derived E-score profiles for reference and alternate DNA sequences of ASB events. A Convolutional Neural Network (CNN) was trained to predict whether these profiles were indicative of ASB gain/loss or no change in binding. 18211 E-score profiles from 113 transcription factors were split into train, validation and test data. We compared the performance of the trained model with other available platforms for predicting the effect of SNP on transcription factor binding. Our model demonstrated increased accuracy and ASB recall in comparison to the best scoring benchmark tools. CONCLUSION: In this paper we present our model SNPeBoT (Single Nucleotide Polymorphism effect on Binding of Transcription Factors) in its standalone and web server form. The increased recovery and prediction accuracy of allele specific binding events could prove useful in discovering non-coding mutations relevant to disease. Patrick Gohl, Baldomero Oliva |
BMC Bioinform. | 2 |
| 2024 | Genopyc: a Python library for investigating the functional effects of genomic variants associated to complex diseasesabstractMOTIVATION: Integrative Biomedicl Informatics, Research Program on Biomedical Informatics (IBI - GRIB), Hospital Del Mar Medical Research Institute (IMIM), Department of Experimental and Health Sciences, Universitat Pompeu Fabra (UPF) C/ del Dr. Aiguader 88 Barcelona 08003 Spain.Understanding the genetic basis of complex diseases is one of the main challenges in modern genomics. However, current tools often lack the versatility to efficiently analyze the intricate relationships between genetic variations and disease outcomes. To address this, we introduce Genopyc, a novel Python library designed for comprehensive investigation of how the variants associated to complex diseases affects downstream pathways. Genopyc offers an extensive suite of functions for heterogeneous data mining and visualization, enabling researchers to delve into and integrate biological information from large-scale genomic datasets. RESULTS: In this work, we present the Genopyc library through application to real-world genome wide association studies variants. Using Genopyc to investigate the functional consequences of variants associated to intervertebral disc degeneration enabled a deeper understanding of the potential dysregulated pathways involved in the disease, which can be explored and visualized by exploiting the functionalities featured in the package. Genopyc emerges as a powerful asset for researchers, facilitating the investigation of complex diseases paving the way for more targeted therapeutic interventions. AVAILABILITY AND IMPLEMENTATION: Genopyc is available on pip https://pypi.org/project/genopyc/.The source code of Genopyc is available at https://github.com/freh-g/genopyc. A tutorial notebook is available at https://github.com/freh-g/genopyc/blob/main/tutorials/Genopyc_tutorial_notebook.ipynb. Finally, a detailed documentation is available at: https://genopyc.readthedocs.io/en/latest/. Francesco Gualdi, Baldomero Oliva, Janet Piñero |
Bioinform. | 2 |
| 2023 | SBILib: a handle for protein modeling and engineeringabstractSUMMARY: The SBILib Python library provides an integrated platform for the analysis of macromolecular structures and interactions. It combines simple 3D file parsing and workup methods with more advanced analytical tools. SBILib includes modules for macromolecular interactions, loops, super-secondary structures, and biological sequences, as well as wrappers for external tools with which to integrate their results and facilitate the comparative analysis of protein structures and their complexes. The library can handle macromolecular complexes formed by proteins and/or nucleic acid molecules (i.e. DNA and RNA). It is uniquely capable of parsing and calculating protein super-secondary structure and loop geometry. We have compiled a list of example scenarios which SBILib may be applied to and provided access to these within the library. AVAILABILITY AND IMPLEMENTATION: SBILib is made available on Github at https://github.com/structuralbioinformatics/SBILib. Patrick Gohl, Jaume Bonet, Oriol Fornes, Joan Planas-Iglesias, Narcis Fernandez-Fuentes, Baldomero Oliva |
Bioinform. | 6 |
| 2021 | SPServer: split-statistical potentials for the analysis of protein structures and protein-protein interactionsabstractBACKGROUND: Statistical potentials, also named knowledge-based potentials, are scoring functions derived from empirical data that can be used to evaluate the quality of protein folds and protein-protein interaction (PPI) structures. In previous works we decomposed the statistical potentials in different terms, named Split-Statistical Potentials, accounting for the type of amino acid pairs, their hydrophobicity, solvent accessibility and type of secondary structure. These potentials have been successfully used to identify near-native structures in protein structure prediction, rank protein docking poses, and predict PPI binding affinities. RESULTS: Here, we present the SPServer, a web server that applies the Split-Statistical Potentials to analyze protein folds and protein interfaces. SPServer provides global scores as well as residue/residue-pair profiles presented as score plots and maps. This level of detail allows users to: (1) identify potentially problematic regions on protein structures; (2) identify disrupting amino acid pairs in protein interfaces; and (3) compare and analyze the quality of tertiary and quaternary structural models. CONCLUSIONS: While there are many web servers that provide scoring functions to assess the quality of either protein folds or PPI structures, SPServer integrates both aspects in a unique easy-to-use web server. Moreover, the server permits to locally assess the quality of the structures and interfaces at a residue level and provides tools to compare the local assessment between structures. SERVER ADDRESS: https://sbi.upf.edu/spserver/ . Joaquim Aguirre-Plans, Alberto Meseguer, Ruben Molina-Fernandez, Manuel Alejandro Marín-López, Gaurav Jumde, Kevin Casanova, Jaume Bonet, Oriol Fornes, Narcis Fernandez-Fuentes, Baldomero Oliva |
BMC Bioinform. | 10 |
| 2018 | On the mechanisms of protein interactions: predicting their affinity from unbound tertiary structuresabstractMotivation: The characterization of the protein-protein association mechanisms is crucial to understanding how biological processes occur. It has been previously shown that the early formation of non-specific encounters enhances the realization of the stereospecific (i.e. native) complex by reducing the dimensionality of the search process. The association rate for the formation of such complex plays a crucial role in the cell biology and depends on how the partners diffuse to be close to each other. Predicting the binding free energy of proteins provides new opportunities to modulate and control protein-protein interactions. However, existing methods require the 3D structure of the complex to predict its affinity, severely limiting their application to interactions with known structures. Results: We present a new approach that relies on the unbound protein structures and protein docking to predict protein-protein binding affinities. Through the study of the docking space (i.e. decoys), the method predicts the binding affinity of the query proteins when the actual structure of the complex itself is unknown. We tested our approach on a set of globular and soluble proteins of the newest affinity benchmark, obtaining accuracy values comparable to other state-of-art methods: a 0.4 correlation coefficient between the experimental and predicted values of ΔG and an error < 3 Kcal/mol. Availability and implementation: The binding affinity predictor is implemented and available at http://sbi.upf.edu/BADock and https://github.com/badocksbi/BADock. Contact: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Manuel Alejandro Marín-López, Joan Planas-Iglesias, Joaquim Aguirre-Plans, Jaume Bonet, Javier García-García 0002, Narcis Fernandez-Fuentes, Baldomero Oliva |
Bioinform. | 7 |
| 2015 | Knowledge-based modeling of peptides at protein interfaces: PiPreDabstractMOTIVATION: Protein-protein interactions (PPIs) underpin virtually all cellular processes both in health and disease. Modulating the interaction between proteins by means of small (chemical) agents is therefore a promising route for future novel therapeutic interventions. In this context, peptides are gaining momentum as emerging agents for the modulation of PPIs. RESULTS: We reported a novel computational, structure and knowledge-based approach to model orthosteric peptides to target PPIs: PiPreD. PiPreD relies on a precompiled and bespoken library of structural motifs, iMotifs, extracted from protein complexes and a fast structural modeling algorithm driven by the location of native chemical groups on the interface of the protein target named anchor residues. PiPreD comprehensive and systematically samples the entire interface deriving peptide conformations best suited for the given region on the protein interface. PiPreD complements the existing technologies and provides new solutions for the disruption of selected interactions. AVAILABILITY AND IMPLEMENTATION: Database and accessory scripts and programs are available upon request to the authors or at http://www.bioinsilico.org/PIPRED. CONTACT: [email protected]. Baldomero Oliva, Narcis Fernandez-Fuentes |
Bioinform. | 1 |
| 2014 | Frag'r'Us: knowledge-based sampling of protein backbone conformations for de novo structure-based protein designabstractMOTIVATION: The remodeling of short fragment(s) of the protein backbone to accommodate new function(s), fine-tune binding specificities or change/create novel protein interactions is a common task in structure-based computational design. Alternative backbone conformations can be generated de novo or by redeploying existing fragments extracted from protein structures i.e. knowledge-based. We present Frag'r'Us, a web server designed to sample alternative protein backbone conformations in loop regions. The method relies on a database of super secondary structural motifs called smotifs. Thus, sampling of conformations reflects structurally feasible fragments compiled from existing protein structures. Availability and implementation Frag'r'Us has been implemented as web application and is available at http://www.bioinsilico.org/FRAGRUS. Jaume Bonet, Joan Segura, Joan Planas-Iglesias, Baldomero Oliva, Narcis Fernandez-Fuentes |
Bioinform. | 4 |
| 2014 | GUILDify: a web server for phenotypic characterization of genes through biological data integration and network-based prioritization algorithmsabstractSUMMARY: Determining genetic factors underlying various phenotypes is hindered by the involvement of multiple genes acting cooperatively. Over the past years, disease-gene prioritization has been central to identify genes implicated in human disorders. Special attention has been paid on using physical interactions between the proteins encoded by the genes to link them with diseases. Such methods exploit the guilt-by-association principle in the protein interaction network to uncover novel disease-gene associations. These methods rely on the proximity of a gene in the network to the genes associated with a phenotype and require a set of initial associations. Here, we present GUILDify, an easy-to-use web server for the phenotypic characterization of genes. GUILDify offers a prioritization approach based on the protein-protein interaction network where the initial phenotype-gene associations are retrieved via free text search on biological databases. GUILDify web server does not restrict the prioritization to any predefined phenotype, supports multiple species and accepts user-specified genes. It also prioritizes drugs based on the ranking of their targets, unleashing opportunities for repurposing drugs for novel therapies. AVAILABILITY AND IMPLEMENTATION: Available online at http://sbi.imim.es/GUILDify.php Emre Guney, Javier García-García 0002, Baldomero Oliva |
Bioinform. | 3 |
| 2013 | iLoops: a protein-protein interaction prediction server based on structural featuresabstractSUMMARY: Protein-protein interactions play a critical role in many biological processes. Despite that, the number of servers that provide an easy and comprehensive method to predict them is still limited. Here, we present iLoops, a web server that predicts whether a pair of proteins can interact using local structural features. The inputs of the server are as follows: (i) the sequences of the query proteins and (ii) the pairs to be tested. Structural features are assigned to the query proteins by sequence similarity. Pairs of structural features (formed by loops or domains) are classified according to their likelihood to favor or disfavor a protein-protein interaction, depending on their observation in known interacting and non-interacting pairs. The server evaluates the putative interaction using a random forest classifier. AVAILABILITY: iLoops is available at http://sbi.imim.es/iLoops.php CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Joan Planas-Iglesias, Manuel Alejandro Marín-López, Jaume Bonet, Javier García-García 0002, Baldomero Oliva |
Bioinform. | 5 |
| 2011 | Toward PWAS: discovering pathways associated with human disordersabstractThe past decade has witnessed dramatic advances in genome sequencing and a substantial shift in the number of genome wide association studies (GWAS). These efforts have expanded considerably our knowledge on the sequential variations in Human DNA and their consequences on the human biology. Nevertheless, complex genetic disorders often involve products of multiple genes acting cooperatively and pinpointing the decisive elements of such disease pathways remains a challenge. Network biology recently proved its use in identifying candidate disease genes based on the simple observation that proteins translated by phenotypically related genes tend to interact, the so called guilt-by-association principle. Emre Guney, Baldomero Oliva |
BMC Bioinform. | 2 |
| 2010 | Biana: a software framework for compiling biological interactions and analyzing networksabstractBACKGROUND: The analysis and usage of biological data is hindered by the spread of information across multiple repositories and the difficulties posed by different nomenclature systems and storage formats. In particular, there is an important need for data unification in the study and use of protein-protein interactions. Without good integration strategies, it is difficult to analyze the whole set of available data and its properties. RESULTS: We introduce BIANA (Biologic Interactions and Network Analysis), a tool for biological information integration and network management. BIANA is a Python framework designed to achieve two major goals: i) the integration of multiple sources of biological information, including biological entities and their relationships, and ii) the management of biological information as a network where entities are nodes and relationships are edges. Moreover, BIANA uses properties of proteins and genes to infer latent biomolecular relationships by transferring edges to entities sharing similar properties. BIANA is also provided as a plugin for Cytoscape, which allows users to visualize and interactively manage the data. A web interface to BIANA providing basic functionalities is also available. The software can be downloaded under GNU GPL license from http://sbi.imim.es/web/BIANA.php. CONCLUSIONS: BIANA's approach to data unification solves many of the nomenclature issues common to systems dealing with biological data. BIANA can easily be extended to handle new specific data repositories and new specific data types. The unification protocol allows BIANA to be a flexible tool suitable for different user requirements: non-expert users can use a suggested unification protocol while expert users can define their own specific unification rules. Javier García-García 0002, Emre Guney, Ramon Aragues, Joan Planas-Iglesias, Baldomero Oliva |
BMC Bioinform. | 5 |
| 2009 | ModLink+: improving fold recognition by using protein-protein interactionsabstractMOTIVATION: Several strategies have been developed to predict the fold of a target protein sequence, most of which are based on aligning the target sequence to other sequences of known structure. Previously, we demonstrated that the consideration of protein-protein interactions significantly increases the accuracy of fold assignment compared with PSI-BLAST sequence comparisons. A drawback of our method was the low number of proteins to which a fold could be assigned. Here, we present an improved version of the method that addresses this limitation. We also compare our method to other state-of-the-art fold assignment methodologies. RESULTS: Our approach (ModLink+) has been tested on 3716 proteins with domain folds classified in the Structural Classification Of Proteins (SCOP) as well as known interacting partners in the Database of Interacting Proteins (DIP). For this test set, the ratio of success [positive predictive value (PPV)] on fold assignment increases from 75% for PSI-BLAST, 83% for HHSearch and 81% for PRC to >90% for ModLink+at the e-value cutoff of 10(-3). Under this e-value, ModLink+can assign a fold to 30-45% of the proteins in the test set, while our previous method could cover <25%. When applied to 6384 proteins with unknown fold in the yeast proteome, ModLink+combined with PSI-BLAST assigns a fold for domains in 3738 proteins, while PSI-BLAST alone covers only 2122 proteins, HHSearch 2969 and PRC 2826 proteins, using a threshold e-value that would represent a PPV >82% for each method in the test set. AVAILABILITY: The ModLink+server is freely accessible in the World Wide Web at http://sbi.imim.es/modlink/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Oriol Fornes, Ramon Aragues, Jordi Espadaler, Marc A. Martí-Renom, Andrej Sali, Baldomero Oliva |
Bioinform. | 6 |
| 2008 | Predicting cancer involvement of genes from heterogeneous dataabstractBACKGROUND: Systematic approaches for identifying proteins involved in different types of cancer are needed. Experimental techniques such as microarrays are being used to characterize cancer, but validating their results can be a laborious task. Computational approaches are used to prioritize between genes putatively involved in cancer, usually based on further analyzing experimental data. RESULTS: We implemented a systematic method using the PIANA software that predicts cancer involvement of genes by integrating heterogeneous datasets. Specifically, we produced lists of genes likely to be involved in cancer by relying on: (i) protein-protein interactions; (ii) differential expression data; and (iii) structural and functional properties of cancer genes. The integrative approach that combines multiple sources of data obtained positive predictive values ranging from 23% (on a list of 811 genes) to 73% (on a list of 22 genes), outperforming the use of any of the data sources alone. We analyze a list of 20 cancer gene predictions, finding that most of them have been recently linked to cancer in literature. CONCLUSION: Our approach to identifying and prioritizing candidate cancer genes can be used to produce lists of genes likely to be involved in cancer. Our results suggest that differential expression studies yielding high numbers of candidate cancer genes can be filtered using protein interaction networks. Ramon Aragues, Chris Sander, Baldomero Oliva |
BMC Bioinform. | 3 |
| 2008 | Prediction of enzyme function by combining sequence similarity and protein interactionsabstractBACKGROUND: A number of studies have used protein interaction data alone for protein function prediction. Here, we introduce a computational approach for annotation of enzymes, based on the observation that similar protein sequences are more likely to perform the same function if they share similar interacting partners. RESULTS: The method has been tested against the PSI-BLAST program using a set of 3,890 protein sequences from which interaction data was available. For protein sequences that align with at least 40% sequence identity to a known enzyme, the specificity of our method in predicting the first three EC digits increased from 80% to 90% at 80% coverage when compared to PSI-BLAST. CONCLUSION: Our method can also be used in proteins for which homologous sequences with known interacting partners can be detected. Thus, our method could increase 10% the specificity of genome-wide enzyme predictions based on sequence matching by PSI-BLAST alone. Jordi Espadaler, Narayanan Eswar, Enrique Querol, Francesc X. Avilés, Andrej Sali, Marc A. Martí-Renom, Baldomero Oliva |
BMC Bioinform. | 7 |
| 2008 | Model Formulation: Training Multidisciplinary Biomedical Informatics Students: Three Years of ExperienceabstractOBJECTIVE: The European INFOBIOMED Network of Excellence recognized that a successful education program in biomedical informatics should include not only traditional teaching activities in the basic sciences but also the development of skills for working in multidisciplinary teams. DESIGN: A carefully developed 3-year training program for biomedical informatics students addressed these educational aspects through the following four activities: (1) an internet course database containing an overview of all Medical Informatics and BioInformatics courses, (2) a BioMedical Informatics Summer School, (3) a mobility program based on a 'brokerage service' which published demands and offers, including funding for research exchange projects, and (4) training challenges aimed at the development of multi-disciplinary skills. MEASUREMENTS: This paper focuses on experiences gained in the development of novel educational activities addressing work in multidisciplinary teams. The training challenges described here were evaluated by asking participants to fill out forms with Likert scale based questions. For the mobility program a needs assessment was carried out. RESULTS: The mobility program supported 20 exchanges which fostered new BMI research, resulted in a number of peer-reviewed publications and demonstrated the feasibility of this multidisciplinary BMI approach within the European Union. Students unanimously indicated that the training challenge experience had contributed to their understanding and appreciation of multidisciplinary teamwork. CONCLUSION: The training activities undertaken in INFOBIOMED have contributed to a multi-disciplinary BMI approach. It is our hope that this work might provide an impetus for training efforts in Europe, and yield a new generation of biomedical informaticians. Erik M. van Mulligen, Montserrat Cases, Kristina M. Hettne, Eva Molero, Marc Weeber, Kevin A. Robertson, Baldomero Oliva, Guillermo de la Calle, Victor Maojo |
J. Am. Medical Informatics Assoc. | 7 |
| 2007 | Structure-based evaluation of in silico predictions of protein-protein interactions using Comparative DockingabstractMOTIVATION: Due to the limitations in experimental methods for determining binary interactions and structure determination of protein complexes, the need exists for computational models to fill the increasing gap between genome sequence information and protein annotation. Here we describe a novel method that uses structural models to reduce a large number of in silico predictions to a high confidence subset that is amenable to experimental validation. RESULTS: A two-stage evaluation procedure was developed, first, a sequence-based method assessed the conservation of protein interface patches used in the original in silico prediction method, both in terms of position within the primary sequence, and in terms of sequence conservation. When applying the most stringent conditions it was found that 20.5% of the data set being assessed passed this test. Secondly, a high-throughput structure-based docking evaluation procedure assessed the soundness of three dimensional models produced for the putative interactions. Of the data set being assessed, 8264 interactions or over 70% could be modelled in this way, and 27% of these can be considered 'valid' by the applied criteria. In all, 6.9% of the interactions passed both the tests and can be considered to be a high confidence set of predicted interactions, several of which are described. AVAILABILITY: http://bioinformatics.leeds.ac.uk/~bmb4sjc. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Simon J. Cockell, Baldomero Oliva, Richard M. Jackson |
Bioinform. | 2 |
| 2007 | Characterization of Protein Hubs by Inferring Interacting Motifs from Protein InteractionsabstractThe characterization of protein interactions is essential for understanding biological systems. While genome-scale methods are available for identifying interacting proteins, they do not pinpoint the interacting motifs (e.g., a domain, sequence segments, a binding site, or a set of residues). Here, we develop and apply a method for delineating the interacting motifs of hub proteins (i.e., highly connected proteins). The method relies on the observation that proteins with common interaction partners tend to interact with these partners through a common interacting motif. The sole input for the method are binary protein interactions; neither sequence nor structure information is needed. The approach is evaluated by comparing the inferred interacting motifs with domain families defined for 368 proteins in the Structural Classification of Proteins (SCOP). The positive predictive value of the method for detecting proteins with common SCOP families is 75% at sensitivity of 10%. Most of the inferred interacting motifs were significantly associated with sequence patterns, which could be responsible for the common interactions. We find that yeast hubs with multiple interacting motifs are more likely to be essential than hubs with one or two interacting motifs, thus rationalizing the previously observed correlation between essentiality and the number of interacting partners of a protein. We also find that yeast hubs with multiple interacting motifs evolve slower than the average protein, contrary to the hubs with one or two interacting motifs. The proposed method will help us discover unknown interacting motifs and provide biological insights about protein hubs and their roles in interaction networks. Ramon Aragues, Andrej Sali, Jaume Bonet, Marc A. Martí-Renom, Baldomero Oliva |
PLoS Comput. Biol. | 5 |
| 2006 | PIANA: protein interactions and network analysisabstractUNLABELLED: We present a software framework and tool called Protein Interactions And Network Analysis (PIANA) that facilitates working with protein interaction networks by (1) integrating data from multiple sources, (2) providing a library that handles graph-related tasks and (3) automating the analysis of protein-protein interaction networks. PIANA can also be used as a stand-alone application to create protein interaction networks and perform tasks such as predicting protein interactions and helping to identify spots in a 2D electrophoresis gel. AVAILABILITY: PIANA is under the GNU GPL. Source code, database and detailed documentation may be freely downloaded from http://sbi.imim.es/piana. Ramon Aragues, Daniel Jaeggi, Baldomero Oliva |
Bioinform. | 3 |
| 2006 | Identification of function-associated loop motifs and application to protein function predictionabstractMOTIVATION: The detection of function-related local 3D-motifs in protein structures can provide insights towards protein function in absence of sequence or fold similarity. Protein loops are known to play important roles in protein function and several loop classifications have been described, but the automated identification of putative functional 3D-motifs in such classifications has not yet been addressed. This identification can be used on sequence annotations. RESULTS: We evaluated three different scoring methods for their ability to identify known motifs from the PROSITE database in ArchDB. More than 500 new putative function-related motifs not reported in PROSITE were identified. Sequence patterns derived from these motifs were especially useful at predicting precise annotations. The number of reliable sequence annotations could be increased up to 100% with respect to standard BLAST. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary Data are available at Bioinformatics online. Jordi Espadaler, Enrique Querol, Francesc X. Avilés, Baldomero Oliva |
Bioinform. | 4 |
| 2005 | Prediction of protein-protein interactions using distant conservation of sequence patterns and structure relationshipsabstractMOTIVATION: Given that association and dissociation of protein molecules is crucial in most biological processes several in silico methods have been recently developed to predict protein-protein interactions. Structural evidence has shown that usually interacting pairs of close homologs (interologs) physically interact in the same way. Moreover, conservation of an interaction depends on the conservation of the interface between interacting partners. In this article we make use of both, structural similarities among domains of known interacting proteins found in the Database of Interacting Proteins (DIP) and conservation of pairs of sequence patches involved in protein-protein interfaces to predict putative protein interaction pairs. RESULTS: We have obtained a large amount of putative protein-protein interaction (approximately 130,000). The list is independent from other techniques both experimental and theoretical. We separated the list of predictions into three sets according to their relationship with known interacting proteins found in DIP. For each set, only a small fraction of the predicted protein pairs could be independently validated by cross checking with the Human Protein Reference Database (HPRD). The fraction of validated protein pairs was always larger than that expected by using random protein pairs. Furthermore, a correlation map of interacting protein pairs was calculated with respect to molecular function, as defined in the Gene Ontology database. It shows good consistency of the predicted interactions with data in the HPRD database. The intersection between the lists of interactions of other methods and ours produces a network of potentially high-confidence interactions. Jordi Espadaler, Oriol Romero-Isart, Richard M. Jackson, Baldomero Oliva |
Bioinform. | 4 |
| 2003 | SNOW: Standard NOmenclature Wizard to help searching for (bio) chemical standardized namesabstractUNLABELLED: When developing bioinformatical tools dealing with enzymatic activity, metabolism or enzymatic networks, the problem of the lack of a clear nomenclature for biochemical compounds often arises. This problem leads us to develop a small web-based tool (SNOW, Standard NOmenclature Wizard) which may help to find recommended and trivial names or the correct closest spelling for a query compound name, if it exists. AVAILABILITY: Web-based interface available at http://ibb.uab.es/snow/ SUPPLEMENTARY INFORMATION: http://ibb.uab.es/snow/snow_moreinfo.html Daniel Aguilar, Baldomero Oliva, Francesc X. Avilés, Enrique Querol |
Bioinform. | 2 |
| 2002 | TranScout: prediction of gene expression regulatory proteins from their sequencesabstractAbstract Motivation: The advent of genomics yields thousands of reading frames in search of function. Identification of conserved functional motifs in protein sequences can be helpful for function prediction. Results: A database and a classification of reported DNA-binding protein motifs has been designed. A program (‘TranScout’) has been developed for the detection and evaluation of conserved motifs in prokaryotic and eukaryotic sequences of proteins with a gene regulatory function. The efficiency of the program is shown in a benchmark against a database obtained from SWISS-PROT without the protein sequences used to train the program. All motifs were detected with a mean average sensitivity of 0.98 and a mean average specificity of 0.92. Availability: The program is freely available for use on the internet at http://luz.uab.es/transcout/. The user can find additional information at this site. Contact: [email protected]; [email protected] * To whom correspondence should be addressed. 2 Present address: Laboratorio de Biologı́a Estructural Computacional (GRIB), Universitat Pompeu Fabra, C/Doctor Aiguader, Barcelona 08003, Spain. Daniel Aguilar, Baldomero Oliva, Francesc X. Avilés, Enrique Querol |
Bioinform. | 2 |
| 1997 | 'TransMem': a neural network implemented in Excel spreadsheets for predicting transmembrane domains of proteinsabstractMOTIVATION: Genomic sequences from different organisms, even prokaryotic, have plenty of orphan ORFs, making necessary methods for the prediction of protein structure and function. The prediction of the presence of hydrophobic transmembrane (HTM) stretches is a valuable clue for this. RESULTS: The program. TransMem, based on a neural network and running on personal computers (either Apple Macintosh or PC, using Excel worksheets), for the prediction and distribution of amino acid residues in transmembrane segments of integral membrane proteins is reported. The percentage of residue predictive accuracy obtained for the set of proteins tested is 93%, ranging from 99.9% for the best to 71.7% for the worst prediction. The segment-based accuracy is 93.6%; 63.6% of the protein set match any of the predicted and observed segment locations. AVAILABILITY: TransMem is available upon request or by anonymous up: IP address: luz.uab.es, directory/pub/ TransMem. It is also placed on the EMBL file server (ftp:/(/)ftp.ebi.ac.uk/pub/software/mac/TransMem ). Patrick Aloy, Juan Cedano, Baldomero Oliva, Francesc X. Avilés, Enrique Querol |
Comput. Appl. Biosci. | 3 |