VLDB 2026 Research / reviewers in the wild / expert
Miguel A. Andrade-Navarro
dblp:a/MiguelAAndradeNavarro · also Miguel A. Andrade
· DBLP profile ↗
41ranked-venue papers
7as first author
5since 2021 · last 2026
0000-0001-6650-1711ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 39 · 6 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Cross-domain transfer learning from peptides to metabolites using a multi-property fine-tuned LLMabstractMOTIVATION: Accurate liquid chromatography retention time (RT) prediction is a critical component of compound identification in metabolomics and lipidomics. However, existing RT prediction approaches are often limited by the scarcity of experimental RT measurements for many molecular classes, restricting model generalization and the construction of comprehensive RT libraries. Transfer learning from data-rich chemical domains offers a potential strategy to overcome these limitations, but its effectiveness for metabolite RT prediction remains insufficiently explored. RESULTS: We developed a transfer learning framework based on ChemBERTa that leverages large peptide datasets to improve metabolite RT prediction under data-sparse conditions. A peptide-pretrained model was trained using a multi-task objective that jointly predicted RT and seven RDKit-derived molecular descriptors. Compared with an RT-only model, the multi-task approach learned more robust chemical representations and demonstrated superior generalization to metabolites, achieving a median test R² of 0.842 versus 0.820. When transferred to metabolite RT prediction, the multi-task pretrained model substantially outperformed models trained from scratch at low-data regimes. Using only 3% of metabolite training data (2129 compounds), transfer learning achieved a median test R² of 0.322 compared with 0.216 for the baseline model, while reducing MAE from 131.7 to 114.9. Significant improvements were also observed at 5% and 10% training fractions, with benefits gradually diminishing as larger metabolite datasets became available. In contrast, a peptide-pretrained single-task RT model showed performance comparable to the baseline, indicating that the observed gains arise primarily from multi-task molecular property learning rather than peptide pretraining alone. These findings demonstrate that multi-task transfer learning provides an effective and scalable strategy for improving RT prediction in metabolomics, particularly when experimental training data are limited. AVAILABILITY: Freely available on https://github.com/uchealex/CHEMBEDDING. Uchenna Alex Anyaegbunam, David Teschner, Thierry Schmidlin, Andreas Hildebrandt 0001, Johannes U. Mayer, Maximilian Sprang, Miguel A. Andrade-Navarro |
Bioinform. | 7 |
| 2025 | Xsurvey: Web Tool to Query the Set of Homorepeats of all Reference ProteomesabstractHomorepeats are low complexity regions in protein sequences composed of repetitions of one specific amino acid residue. There is currently no automatic way to compare the set of homorepeats between two or more species. Here we present Xsurvey, a web tool to query the set of homorepeats of 23,150 completely-sequenced proteomes. The polyX usage values can be easily compared visually, which simplifies the interpretation of the results. Xsurvey is freely available for public use at https://cbdm-01.zdv.uni-mainz.de/∼munoz/xsurvey/. Miguel A. Andrade-Navarro, Pablo Mier |
IEEE Trans. Comput. Biol. Bioinform. | 1 |
| 2022 | The sequence context in poly-alanine regions: structure, function and conservationabstractMOTIVATION: Poly-alanine (polyA) regions are protein stretches mostly composed of alanines. Despite their abundance in eukaryotic proteomes and their association to nine inherited human diseases, the structural and functional roles exerted by polyA stretches remain poorly understood. In this work we study how the amino acid context in which polyA regions are settled in proteins influences their structure and function. RESULTS: We identified glycine and proline as the most abundant amino acids within polyA and in the flanking regions of polyA tracts, in human proteins as well as in 17 additional eukaryotic species. Our analyses indicate that the non-structuring nature of these two amino acids influences the α-helical conformations predicted for polyA, suggesting a relevant role in reducing the inherent aggregation propensity of long polyA. Then, we show how polyA position in protein N-termini relates with their function as transit peptides. PolyA placed just after the initial methionine is often predicted as part of mitochondrial transit peptides, whereas when placed in downstream positions, polyA are part of signal peptides. A few examples from known structures suggest that short polyA can emerge by alanine substitutions in α-helices; but evolution by insertion is observed for longer polyA. Our results showcase the importance of studying the sequence context of homorepeats as a mechanism to shape their structure-function relationships. AVAILABILITY AND IMPLEMENTATION: The datasets used and/or analyzed during the current study are available from the corresponding author onreasonable request. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Pablo Mier, Carlos A. Elena-Real, Juan Cortés, Pau Bernadó, Miguel A. Andrade-Navarro |
Bioinform. | 5 |
| 2022 | Batch effect detection and correction in RNA-seq data using machine-learning-based automated assessment of qualityabstractBACKGROUND: The constant evolving and development of next-generation sequencing techniques lead to high throughput data composed of datasets that include a large number of biological samples. Although a large number of samples are usually experimentally processed by batches, scientific publications are often elusive about this information, which can greatly impact the quality of the samples and confound further statistical analyzes. Because dedicated bioinformatics methods developed to detect unwanted sources of variance in the data can wrongly detect real biological signals, such methods could benefit from using a quality-aware approach. RESULTS: We recently developed statistical guidelines and a machine learning tool to automatically evaluate the quality of a next-generation-sequencing sample. We leveraged this quality assessment to detect and correct batch effects in 12 publicly available RNA-seq datasets with available batch information. We were able to distinguish batches by our quality score and used it to correct for some batch effects in sample clustering. Overall, the correction was evaluated as comparable to or better than the reference method that uses a priori knowledge of the batches (in 10 and 1 datasets of 12, respectively; total = 92%). When coupled to outlier removal, the correction was more often evaluated as better than the reference (comparable or better in 5 and 6 datasets of 12, respectively; total = 92%). CONCLUSIONS: In this work, we show the capabilities of our software to detect batches in public RNA-seq datasets from differences in the predicted quality of their samples. We also use these insights to correct the batch effect and observe the relation of sample quality and batch effect. These observations reinforce our expectation that while batch effects do correlate with differences in quality, batch effects also arise from other artifacts and are more suitably corrected statistically in well-designed experiments. Maximilian Sprang, Miguel A. Andrade-Navarro, Jean-Fred Fontaine |
BMC Bioinform. | 2 |
| 2021 | LipiDisease: associate lipids to diseases using literature miningabstractSUMMARY: Lipids exhibit an essential role in cellular assembly and signaling. Dysregulation of these functions has been linked with many complications including obesity, diabetes, metabolic disorders, cancer and more. Investigating lipid profiles in such conditions can provide insights into cellular functions and possible interventions. Hence the field of lipidomics is expanding in recent years. Even though the role of individual lipids in diseases has been investigated, there is no resource to perform disease enrichment analysis considering the cumulative association of a lipid set. To address this, we have implemented the LipiDisease web server. The tool analyzes millions of records from the PubMed biomedical literature database discussing lipids and diseases, predicts their association and ranks them according to false discovery rates generated by random simulations. The tool takes into account 4270 diseases and 4798 lipids. Since the tool extracts the information from PubMed records, the number of diseases and lipids will be expanded over time as the biomedical literature grows. AVAILABILITY AND IMPLEMENTATION: The LipiDisease webserver can be freely accessed at http://cbdm-01.zdv.uni-mainz.de:3838/piyusmor/LipiDisease/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Piyush More, Laura Bindila, Philipp Wild, Miguel A. Andrade-Navarro, Jean-Fred Fontaine |
Bioinform. | 4 |
| 2020 | Disentangling the complexity of low complexity proteinsabstractThere are multiple definitions for low complexity regions (LCRs) in protein sequences, with all of them broadly considering LCRs as regions with fewer amino acid types compared to an average composition. Following this view, LCRs can also be defined as regions showing composition bias. In this critical review, we focus on the definition of sequence complexity of LCRs and their connection with structure. We present statistics and methodological approaches that measure low complexity (LC) and related sequence properties. Composition bias is often associated with LC and disorder, but repeats, while compositionally biased, might also induce ordered structures. We illustrate this dichotomy, and more generally the overlaps between different properties related to LCRs, using examples. We argue that statistical measures alone cannot capture all structural aspects of LCRs and recommend the combined usage of a variety of predictive tools and measurements. While the methodologies available to study LCRs are already very advanced, we foresee that a more comprehensive annotation of sequences in the databases will enable the improvement of predictions and a better understanding of the evolution and the connection between structure and function of LCRs. This will require the use of standards for the generation and exchange of data describing all aspects of LCRs. SHORT ABSTRACT: There are multiple definitions for low complexity regions (LCRs) in protein sequences. In this critical review, we focus on the definition of sequence complexity of LCRs and their connection with structure. We present statistics and methodological approaches that measure low complexity (LC) and related sequence properties. Composition bias is often associated with LC and disorder, but repeats, while compositionally biased, might also induce ordered structures. We illustrate this dichotomy, plus overlaps between different properties related to LCRs, using examples. Pablo Mier, Lisanna Paladin, Stella Tamana, Sophia Petrosian, Borbála Hajdu-Soltész, Annika Urbanek, Aleksandra Gruca, Dariusz Plewczynski, Marcin Grynberg, Pau Bernadó, Zoltán Gáspári, Christos A. Ouzounis, Vasilis J. Promponas, Andrey V. Kajava, John M. Hancock, Silvio C. E. Tosatto, Zsuzsanna Dosztányi, Miguel A. Andrade-Navarro |
Briefings Bioinform. | 18 |
| 2019 | Toward completion of the Earth's proteome: an update a decade laterabstractProtein databases are steadily growing driven by the spread of new more efficient sequencing techniques. This growth is dominated by an increase in redundancy (homologous proteins with various degrees of sequence similarity) and by the incapability to process and curate sequence entries as fast as they are created. To understand these trends and aid bioinformatic resources that might be compromised by the increasing size of the protein sequence databases, we have created a less-redundant protein data set. In parallel, we analyzed the evolution of protein sequence databases in terms of size and redundancy. While the SwissProt database has decelerated its growth mostly because of a focus on increasing the level of annotation of its sequences, its counterpart TrEMBL, much less limited by curation steps, is still in a phase of accelerated growth. However, we predict that before 2020, almost all entries deposited in UniProtKB will be homologous to known proteins. We propose that new sequencing projects can be made more useful if they are driven to sequencing voids, parts of the tree of life far from already sequenced species or model organisms. We show these voids are present in the Archaea and Eukarya domains of life. The approach to the certainty of the redundancy of new protein sequence entries leads to the consideration that most of the protein diversity on Earth has already been described, which we estimate to be of around 3.75 million proteins, revising down the prediction we did a decade ago. Pablo Mier, Miguel A. Andrade-Navarro |
Briefings Bioinform. | 2 |
| 2019 | Traitpedia: a collaborative effort to gather species traitsabstractSUMMARY: Traitpedia is a collaborative database aimed to collect binary traits in a tabular form for a growing number of species. AVAILABILITY AND IMPLEMENTATION: Traitpedia can be accessed from http://cbdm-01.zdv.uni-mainz.de/~munoz/traitpedia. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Pablo Mier, Miguel A. Andrade-Navarro |
Bioinform. | 2 |
| 2018 | The latent geometry of the human protein interaction networkabstractMotivation: A series of recently introduced algorithms and models advocates for the existence of a hyperbolic geometry underlying the network representation of complex systems. Since the human protein interaction network (hPIN) has a complex architecture, we hypothesized that uncovering its latent geometry could ease challenging problems in systems biology, translating them into measuring distances between proteins. Results: We embedded the hPIN to hyperbolic space and found that the inferred coordinates of nodes capture biologically relevant features, like protein age, function and cellular localization. This means that the representation of the hPIN in the two-dimensional hyperbolic plane offers a novel and informative way to visualize proteins and their interactions. We then used these coordinates to compute hyperbolic distances between proteins, which served as likelihood scores for the prediction of plausible protein interactions. Finally, we observed that proteins can efficiently communicate with each other via a greedy routing process, guided by the latent geometry of the hPIN. We show that these efficient communication channels can be used to determine the core members of signal transduction pathways and to study how system perturbations impact their efficiency. Availability and implementation: An R implementation of our network embedder is available at https://github.com/galanisl/NetHypGeom. Also, a web tool for the geometric analysis of the hPIN accompanies this text at http://cbdm-01.zdv.uni-mainz.de/~galanisl/gapi. Supplementary information: Supplementary data are available at Bioinformatics online. Gregorio Alanis-Lobato, Pablo Mier, Miguel A. Andrade-Navarro |
Bioinform. | 3 |
| 2018 | Automated selection of homologs to track the evolutionary history of proteinsabstractBACKGROUND: The selection of distant homologs of a query protein under study is a usual and useful application of protein sequence databases. Such sets of homologs are often applied to investigate the function of a protein and the degree to which experimental results can be transferred from one organism to another. In particular, a variety of databases facilitates static browsing for orthologs. However, these resources have a limited power when identifying orthologs between taxonomically distant species. In addition, in some situations, for a given query protein, it is advantageous to compare the sets of orthologs from different specific organisms: this recursive step-wise search might give an idea of the evolutionary path of the protein as a series of consecutive steps, for example gaining or losing domains. However, a step-wise orthology search is a time-consuming task if the number of steps is high. RESULTS: To illustrate a solution for this problem, we present the web tool ProteinPathTracker, which allows to track the evolutionary history of a query protein by locating homologs in selected proteomes along several evolutionary paths. Additional functionalities include locking a region of interest to follow its evolution in the discovered homologous sequences and the study of the protein function evolution by analysis of the annotations of the homologs. CONCLUSIONS: ProteinPathTracker is an easy-to-use web tool that automatises the practice of looking for selected homologs in distant species in a straightforward way for non-expert users. Pablo Mier, Antonio J. Pérez-Pulido, Miguel A. Andrade-Navarro |
BMC Bioinform. | 3 |
| 2017 | dAPE: a web server to detect homorepeats and follow their evolutionabstractSummary: Homorepeats are low complexity regions consisting of repetitions of a single amino acid residue. There is no current consensus on the minimum number of residues needed to define a functional homorepeat, nor even if mismatches are allowed. Here we present dAPE, a web server that helps following the evolution of homorepeats based on orthology information, using a sensitive but tunable cutoff to help in the identification of emerging homorepeats. Availability and Implementation: dAPE can be accessed from http://cbdm-01.zdv.uni-mainz.de/∼munoz/polyx . Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Pablo Mier, Miguel A. Andrade-Navarro |
Bioinform. | 2 |
| 2014 | CAFE: an R package for the detection of gross chromosomal abnormalities from gene expression microarray dataabstractSUMMARY: The current methods available to detect chromosomal abnormalities from DNA microarray expression data are cumbersome and inflexible. CAFE has been developed to alleviate these issues. It is implemented as an R package that analyzes Affymetrix *.CEL files and comes with flexible plotting functions, easing visualization of chromosomal abnormalities. AVAILABILITY AND IMPLEMENTATION: CAFE is available from https://bitbucket.org/cob87icW6z/cafe/ as both source and compiled packages for Linux and Windows. It is released under the GPL version 3 license. CAFE will also be freely available from Bioconductor. CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Sander Bollen, Mathias Leddin, Miguel A. Andrade-Navarro, Nancy Mah |
Bioinform. | 3 |
| 2014 | Characterizing Protein Interactions Employing a Genome-Wide siRNA Cellular Phenotyping ScreenabstractCharacterizing the activating and inhibiting effect of protein-protein interactions (PPI) is fundamental to gain insight into the complex signaling system of a human cell. A plethora of methods has been suggested to infer PPI from data on a large scale, but none of them is able to characterize the effect of this interaction. Here, we present a novel computational development that employs mitotic phenotypes of a genome-wide RNAi knockdown screen and enables identifying the activating and inhibiting effects of PPIs. Exemplarily, we applied our technique to a knockdown screen of HeLa cells cultivated at standard conditions. Using a machine learning approach, we obtained high accuracy (82% AUC of the receiver operating characteristics) by cross-validation using 6,870 known activating and inhibiting PPIs as gold standard. We predicted de novo unknown activating and inhibiting effects for 1,954 PPIs in HeLa cells covering the ten major signaling pathways of the Kyoto Encyclopedia of Genes and Genomes, and made these predictions publicly available in a database. We finally demonstrate that the predicted effects can be used to cluster knockdown genes of similar biological processes in coherent subgroups. The characterization of the activating or inhibiting effect of individual PPIs opens up new perspectives for the interpretation of large datasets of PPIs and thus considerably increases the value of PPIs as an integrated resource for studying the detailed function of signaling pathways of the cellular system of interest. Apichat Suratanee, Martin H. Schaefer 0001, Matthew J. Betts, Zita Soons, Heiko A. Mannsperger, Nathalie Harder, Marcus Oswald, Markus Gipp, Ellen Ramminger, Guillermo Marcus Martinez, Reinhard Männer, Karl Rohr, Erich E. Wanker, Robert B. Russell, Miguel A. Andrade-Navarro, Roland Eils, Rainer König |
PLoS Comput. Biol. | 15 |
| 2013 | Using cited references to improve the retrieval of related biomedical documentsabstractBACKGROUND: A popular query from scientists reading a biomedical abstract is to search for topic-related documents in bibliographic databases. Such a query is challenging because the amount of information attached to a single abstract is little, whereas classification-based retrieval algorithms are optimally trained with large sets of relevant documents. As a solution to this problem, we propose a query expansion method that extends the information related to a manuscript using its cited references. RESULTS: Data on cited references and text sections in 249,108 full-text biomedical articles was extracted from the Open Access subset of the PubMed Central® database (PMC-OA). Of the five standard sections of a scientific article, the Introduction and Discussion sections contained most of the citations (mean = 10.2 and 9.9 citations, respectively). A large proportion of articles (98.4%) and their cited references (79.5%) were indexed in the PubMed® database. Using the MedlineRanker abstract classification tool, cited references allowed accurate retrieval of the citing document in a test set of 10,000 documents and also of documents related to six biomedical topics defined by particular MeSH® terms from the entire PMC-OA (p-value<0.01). Classification performance was sensitive to the topic and also to the text sections from which the references were selected. Classifiers trained on the baseline (i.e., only text from the query document and not from the references) were outperformed in almost all the cases. Best performance was often obtained when using all cited references, though using the references from Introduction and Discussion sections led to similarly good results. This query expansion method performed significantly better than pseudo relevance feedback in 4 out of 6 topics. CONCLUSIONS: The retrieval of documents related to a single document can be significantly improved by using the references cited by this document (p-value<0.01). Using references from Introduction and Discussion performs almost as well as using all references, which might be useful for methods that require reduced datasets due to computational limitations. Cited references from particular sections might not be appropriate for all topics. Our method could be a better alternative to pseudo relevance feedback though it is limited by full text availability. Francisco M. Ortuño Guzman, Ignacio Rojas, Miguel A. Andrade-Navarro, Jean-Fred Fontaine |
BMC Bioinform. | 3 |
| 2013 | A novel approach for protein subcellular location prediction using amino acid exposureabstractBACKGROUND: Proteins perform their functions in associated cellular locations. Therefore, the study of protein function can be facilitated by predictions of protein location. Protein location can be predicted either from the sequence of a protein alone by identification of targeting peptide sequences and motifs, or by homology to proteins of known location. A third approach, which is complementary, exploits the differences in amino acid composition of proteins associated to different cellular locations, and can be useful if motif and homology information are missing. Here we expand this approach taking into account amino acid composition at different levels of amino acid exposure. RESULTS: Our method has two stages. For stage one, we trained multiple Support Vector Machines (SVMs) to score eukaryotic protein sequences for membership to each of three categories: nuclear, cytoplasmic and extracellular, plus extra category nucleocytoplasmic, accounting for the fact that a large number of proteins shuttles between those two locations. In stage two we use an artificial neural network (ANN) to propose a category from the scores given to the four locations in stage one. The method reaches an accuracy of 68% when using as input 3D-derived values of amino acid exposure. Calibration of the method using predicted values of amino acid exposure allows classifying proteins without 3D-information with an accuracy of 62% and discerning proteins in different locations even if they shared high levels of identity. CONCLUSIONS: In this study we explored the relationship between residue exposure and protein subcellular location. We developed a new algorithm for subcellular location prediction that uses residue exposure signatures. Our algorithm uses a novel approach to address the multiclass classification problem. The algorithm is implemented as web server 'NYCE' and can be accessed at http://cbdm.mdc-berlin.de/~amer/nyce. Arvind Mer, Miguel A. Andrade-Navarro |
BMC Bioinform. | 2 |
| 2013 | Adding Protein Context to the Human Protein-Protein Interaction Network to Reveal Meaningful InteractionsabstractInteractions of proteins regulate signaling, catalysis, gene expression and many other cellular functions. Therefore, characterizing the entire human interactome is a key effort in current proteomics research. This challenge is complicated by the dynamic nature of protein-protein interactions (PPIs), which are conditional on the cellular context: both interacting proteins must be expressed in the same cell and localized in the same organelle to meet. Additionally, interactions underlie a delicate control of signaling pathways, e.g. by post-translational modifications of the protein partners - hence, many diseases are caused by the perturbation of these mechanisms. Despite the high degree of cell-state specificity of PPIs, many interactions are measured under artificial conditions (e.g. yeast cells are transfected with human genes in yeast two-hybrid assays) or even if detected in a physiological context, this information is missing from the common PPI databases. To overcome these problems, we developed a method that assigns context information to PPIs inferred from various attributes of the interacting proteins: gene expression, functional and disease annotations, and inferred pathways. We demonstrate that context consistency correlates with the experimental reliability of PPIs, which allows us to generate high-confidence tissue- and function-specific subnetworks. We illustrate how these context-filtered networks are enriched in bona fide pathways and disease proteins to prove the ability of context-filters to highlight meaningful interactions with respect to various biological questions. We use this approach to study the lung-specific pathways used by the influenza virus, pointing to IRAK1, BHLHE40 and TOLLIP as potential regulators of influenza virus pathogenicity, and to study the signalling pathways that play a role in Alzheimer's disease, identifying a pathway involving the altered phosphorylation of the Tau protein. Finally, we provide the annotated human PPI network via a web frontend that allows the construction of context-specific networks in several ways. Martin H. Schaefer 0001, Tiago J. S. Lopes, Nancy Mah, Jason E. Shoemaker, Yukiko Matsuoka, Jean-Fred Fontaine, Caroline Louis-Jeune, Amie J. Eisfeld, Gabriele Neumann, Carolina Perez-Iratxeta, Yoshihiro Kawaoka, Hiroaki Kitano, Miguel A. Andrade-Navarro |
PLoS Comput. Biol. | 13 |
| 2012 | Keynote lecturesabstractThese tutorials/keynote speeches discuss the following: the effects of nicotine exposure on the complexity and the genetic patterns of dopamine neurons in VTA; from 6-Ps medicine to cardiovascular health informatics; computer-aided interpretation of vascular images towards valid diagnosis and risk stratification of atherosclerosis; turning data into predictions of gene and protein function; from reading to writing (and rewriting) the code of life: the future of biology - scientific, ethical, legal, civil and social issues. Metin Akay, Yuan-Ting Zhang, Konstantina S. Nikita, Miguel A. Andrade-Navarro, Christos A. Ouzounis |
BIBE | 4 |
| 2011 | PDBpaint, a visualization webservice to tag protein structures with sequence annotationsabstractSUMMARY: Protein features are often displayed along the linear sequence of amino acids that make up that protein, but in reality these features occupy a position in the folded protein's 3D space. Mapping sequence features to known or predicted protein structures is useful when trying to deduce the function of those features and when evaluating sequence or structural predictions. To facilitate this goal, we developed PDBpaint, a simple tool that displays protein sequence features gathered from bioinformatics resources on top of protein structures, which are displayed in an interactive window (using the Jmol Java viewer). PDBpaint can be used either with existing protein structures or with novel structures provided by the user. The current version of PDBpaint allows the visualization of annotations from Pfam, ARD (detection of HEAT-repeats), UniProt, TMHMM2.0 and SignalP. Users can also add other annotations manually. AVAILABILITY AND IMPLEMENTATION: PDBpaint is accessible at http://cbdm.mdc-berlin.de/~pdbpaint. Code is available from http://sourceforge.net/projects/pdbpaint. The website was implemented in Perl, with all major browsers supported. CONTACT: [email protected]. David Fournier, Miguel A. Andrade-Navarro |
Bioinform. | 2 |
| 2011 | Tissue-specific subnetworks and characteristics of publicly available human protein interaction databasesabstractMOTIVATION: Protein-protein interaction (PPI) databases are widely used tools to study cellular pathways and networks; however, there are several databases available that still do not account for cell type-specific differences. Here, we evaluated the characteristics of six interaction databases, incorporated tissue-specific gene expression information and finally, investigated if the most popular proteins of scientific literature are involved in good quality interactions. RESULTS: We found that the evaluated databases are comparable in terms of node connectivity (i.e. proteins with few interaction partners also have few interaction partners in other databases), but may differ in the identity of interaction partners. We also observed that the incorporation of tissue-specific expression information significantly altered the interaction landscape and finally, we demonstrated that many of the most intensively studied proteins are engaged in interactions associated with low confidence scores. In summary, interaction databases are valuable research tools but may lead to different predictions on interactions or pathways. The accuracy of predictions can be improved by incorporating datasets on organ- and cell type-specific gene expression, and by obtaining additional interaction evidence for the most 'popular' proteins. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Tiago J. S. Lopes, Martin H. Schaefer 0001, Jason E. Shoemaker, Yukiko Matsuoka, Jean-Fred Fontaine, Gabriele Neumann, Miguel A. Andrade-Navarro, Yoshihiro Kawaoka, Hiroaki Kitano |
Bioinform. | 7 |
| 2011 | PESCADOR, a web-based tool to assist text-mining of biointeractions extracted from PubMed queriesabstractBACKGROUND: Biological function is greatly dependent on the interactions of proteins with other proteins and genes. Abstracts from the biomedical literature stored in the NCBI's PubMed database can be used for the derivation of interactions between genes and proteins by identifying the co-occurrences of their terms. Often, the amount of interactions obtained through such an approach is large and may mix processes occurring in different contexts. Current tools do not allow studying these data with a focus on concepts of relevance to a user, for example, interactions related to a disease or to a biological mechanism such as protein aggregation. RESULTS: To help the concept-oriented exploration of such data we developed PESCADOR, a web tool that extracts a network of interactions from a set of PubMed abstracts given by a user, and allows filtering the interaction network according to user-defined concepts. We illustrate its use in exploring protein aggregation in neurodegenerative disease and in the expansion of pathways associated to colon cancer. CONCLUSIONS: PESCADOR is a platform independent web resource available at: http://cbdm.mdc-berlin.de/tools/pescador/ Adriano Barbosa-Silva, Jean-Fred Fontaine, Elisa R. Donnard, Fernanda Stussi, José Miguel Ortega, Miguel A. Andrade-Navarro |
BMC Bioinform. | 6 |
| 2010 | LAITOR - Literature Assistant for Identification of Terms co-Occurrences and RelationshipsabstractBACKGROUND: Biological knowledge is represented in scientific literature that often describes the function of genes/proteins (bioentities) in terms of their interactions (biointeractions). Such bioentities are often related to biological concepts of interest that are specific of a determined research field. Therefore, the study of the current literature about a selected topic deposited in public databases, facilitates the generation of novel hypotheses associating a set of bioentities to a common context. RESULTS: We created a text mining system (LAITOR: Literature Assistant for Identification of Terms co-Occurrences and Relationships) that analyses co-occurrences of bioentities, biointeractions, and other biological terms in MEDLINE abstracts. The method accounts for the position of the co-occurring terms within sentences or abstracts. The system detected abstracts mentioning protein-protein interactions in a standard test (BioCreative II IAS test data) with a precision of 0.82-0.89 and a recall of 0.48-0.70. We illustrate the application of LAITOR to the detection of plant response genes in a dataset of 1000 abstracts relevant to the topic. CONCLUSIONS: Text mining tools combining the extraction of interacting bioentities and biological concepts with network displays can be helpful in developing reasonable hypotheses in different scientific backgrounds. Adriano Barbosa-Silva, Theodoros G. Soldatos, Ivan L. F. Magalhães, Georgios A. Pavlopoulos, Jean-Fred Fontaine, Miguel A. Andrade-Navarro, Reinhard Schneider 0002, José Miguel Ortega |
BMC Bioinform. | 6 |
| 2009 | Detection of Alpha-Rod Protein Repeats Using a Neural Network and Application to HuntingtinabstractA growing number of solved protein structures display an elongated structural domain, denoted here as alpha-rod, composed of stacked pairs of anti-parallel alpha-helices. Alpha-rods are flexible and expose a large surface, which makes them suitable for protein interaction. Although most likely originating by tandem duplication of a two-helix unit, their detection using sequence similarity between repeats is poor. Here, we show that alpha-rod repeats can be detected using a neural network. The network detects more repeats than are identified by domain databases using multiple profiles, with a low level of false positives (<10%). We identify alpha-rod repeats in approximately 0.4% of proteins in eukaryotic genomes. We then investigate the results for all human proteins, identifying alpha-rod repeats for the first time in six protein families, including proteins STAG1-3, SERAC1, and PSMD1-2 & 5. We also characterize a short version of these repeats in eight protein families of Archaeal, Bacterial, and Fungal species. Finally, we demonstrate the utility of these predictions in directing experimental work to demarcate three alpha-rods in huntingtin, a protein mutated in Huntington's disease. Using yeast two hybrid analysis and an immunoprecipitation technique, we show that the huntingtin fragments containing alpha-rods associate with each other. This is the first definition of domains in huntingtin and the first validation of predicted interactions between fragments of huntingtin, which sets up directions toward functional characterization of this protein. An implementation of the repeat detection algorithm is available as a Web server with a simple graphical output: http://www.ogic.ca/projects/ard. This can be further visualized using BiasViz, a graphic tool for representation of multiple sequence alignments. Gareth A. Palidwor, Sergey Shcherbinin, Matthew R. Huska, Tamas Rasko, Ulrich Stelzl, Anup Arumughan, Raphaele Foulle, Pablo Porras, Luis Sánchez-Pulido, Erich E. Wanker, Miguel A. Andrade-Navarro |
PLoS Comput. Biol. | 11 |
| 2007 | Evolving research trends in bioinformaticsabstractThe cross-disciplinary nature of bioinformatics entails co-evolution with other biomedical disciplines, whereby some bioinformatics applications become popular in certain disciplines and, in turn, these disciplines influence the focus of future bioinformatics development efforts. We observe here that the growth of computational approaches within various biomedical disciplines is not merely a reflection of a general extended usage of computers and the Internet, but due to the production of useful bioinformatics databases and methods for the rest of the biomedical scientific community. We have used the abstracts stored both in the MEDLINE database of biomedical literature and in NIH-funded project grants, to quantify two effects. First, we examine the biomedical literature as a whole and find that the use of computational methods has become increasingly prevalent across biomedical disciplines over the past three decades, while use of databases and the Internet have been rapidly increasing over the past decade. Second, we study the recent trends in the use of bioinformatics topics. We observe that molecular sequence databases are a widely adopted contribution in biomedicine from the field of bioinformatics, and that microarray analysis is one of the major new topics engaged by the bioinformatics community. Via this analysis, we were able to identify areas of rapid growth in the use of informatics to aid in curriculum planning, development of computational infrastructure and strategies for workforce education and funding. Carolina Perez-Iratxeta, Miguel A. Andrade-Navarro, Jonathan D. Wren |
Briefings Bioinform. | 2 |
| 2007 | BiasViz: visualization of amino acid biased regions in protein alignmentsabstractAbstract Summary: About a third of all protein sequences have at least one composition biased region (CBR). Such regions might act as linkers between protein domains but often confer specific binding to various molecules; therefore, their characterization in terms of their boundaries and over-represented residues is important. Analysis of CBRs in a particular sequence can be time consuming if several types of biases have to be explored and their position visualized. Assessment of the significance of the detected CBRs can be approached by comparison to homologous protein sequences. To assist this procedure, we have developed BiasViz, a tool that allows to graphically studying local amino acid composition in protein sequences of a multiple sequence alignment. Availability: BiasViz java applet and source code can be accessed from http://biasviz.sourceforge.net Contact: [email protected] Matthew R. Huska, Henrik Buschmann, Miguel A. Andrade-Navarro |
Bioinform. | 3 |
| 2006 | Amplification of the Gene Ontology annotation of Affymetrix probe setsabstractBACKGROUND: The annotations of Affymetrix DNA microarray probe sets with Gene Ontology terms are carefully selected for correctness. This results in very accurate but incomplete annotations which is not always desirable for microarray experiment evaluation. RESULTS: Here we present a protocol to amplify the set of Gene Ontology annotations associated to Affymetrix DNA microarray probe sets using information from related databases. CONCLUSION: Predicted novel annotations and the evidence producing them can be accessed at Probe2GO: http://www.ogic.ca/p2g. Scripts are available on demand. Enrique M. Muro, Carolina Perez-Iratxeta, Miguel A. Andrade-Navarro |
BMC Bioinform. | 3 |
| 2006 | Taxonomic colouring of phylogenetic trees of protein sequencesabstractBACKGROUND: Phylogenetic analyses of protein families are used to define the evolutionary relationships between homologous proteins. The interpretation of protein-sequence phylogenetic trees requires the examination of the taxonomic properties of the species associated to those sequences. However, there is no online tool to facilitate this interpretation, for example, by automatically attaching taxonomic information to the nodes of a tree, or by interactively colouring the branches of a tree according to any combination of taxonomic divisions. This is especially problematic if the tree contains on the order of hundreds of sequences, which, given the accelerated increase in the size of the protein sequence databases, is a situation that is becoming common. RESULTS: We have developed PhyloView, a web based tool for colouring phylogenetic trees upon arbitrary taxonomic properties of the species represented in a protein sequence phylogenetic tree. Provided that the tree contains SwissProt, SpTrembl, or GenBank protein identifiers, the tool retrieves the taxonomic information from the corresponding database. A colour picker displays a summary of the findings and allows the user to associate colours to the leaves of the tree according to any number of taxonomic partitions. Then, the colours are propagated to the branches of the tree. CONCLUSION: PhyloView can be used at http://www.ogic.ca/projects/phyloview/. A tutorial, the software with documentation, and GPL licensed source code, can be accessed at the same web address. Gareth A. Palidwor, Emmanuel G. Reynaud, Miguel A. Andrade-Navarro |
BMC Bioinform. | 3 |
| 2005 | Inconsistencies over time in 5% of NetAffx probe-to-gene annotationsabstractBACKGROUND: DNA microarray probes are designed to match particular mRNA transcripts, often based on expressed sequences like ESTs, or cDNAs, many times incomplete. As a result, the relations between probes and genes can change as the sequence data are updated. However, it is frequent that the reported results of microarray analyses are given just as lists of genes without any reference to the underlying probes. RESULTS: We show for a particular commercial microarray design that the number of probes associated to some genes change with time. These changes concern approximately 5% of the probe sets across the history of annotation releases over a two year span. CONCLUSION: We recommend to report probe set identifiers when publishing microarray results, and to submit those analyses to microarray public databases to ensure that the interpretation of the data is updated with the latest set of annotations. Carolina Perez-Iratxeta, Miguel A. Andrade-Navarro |
BMC Bioinform. | 2 |
| 2005 | Ranking the whole MEDLINE database according to a large training set using text indexingabstractBACKGROUND: The MEDLINE database contains over 12 million references to scientific literature, with about 3/4 of recent articles including an abstract of the publication. Retrieval of entries using queries with keywords is useful for human users that need to obtain small selections. However, particular analyses of the literature or database developments may need the complete ranking of all the references in the MEDLINE database as to their relevance to a topic of interest. This report describes a method that does this ranking using the differences in word content between MEDLINE entries related to a topic and the whole of MEDLINE, in a computational time appropriate for an article search query engine. RESULTS: We tested the capabilities of our system to retrieve MEDLINE references which are relevant to the subject of stem cells. We took advantage of the existing annotation of references with terms from the MeSH hierarchical vocabulary (Medical Subject Headings, developed at the National Library of Medicine). A training set of 81,416 references was constructed by selecting entries annotated with the MeSH term stem cells or some child in its sub tree. Frequencies of all nouns, verbs, and adjectives in the training set were computed and the ratios of word frequencies in the training set to those in the entire MEDLINE were used to score references. Self-consistency of the algorithm, benchmarked with a test set containing the training set and an equal number of references randomly selected from MEDLINE was better using nouns (79%) than adjectives (73%) or verbs (70%). The evaluation of the system with 6,923 references not used for training, containing 204 articles relevant to stem cells according to a human expert, indicated a recall of 65% for a precision of 65%. CONCLUSION: This strategy appears to be useful for predicting the relevance of MEDLINE references to a given concept. The method is simple and can be used with any user-defined training set. Choice of the part of speech of the words used for classification has important effects on performance. Lists of words, scripts, and additional information are available from the web address http://www.ogic.ca/projects/ks2004/. Brian P. Suomela, Miguel A. Andrade-Navarro |
BMC Bioinform. | 2 |
| 2004 | Gene annotation from scientific literature using mappings between keyword systemsabstractMOTIVATION: The description of genes in databases by keywords helps the non-specialist to quickly grasp the properties of a gene and increases the efficiency of computational tools that are applied to gene data (e.g. searching a gene database for sequences related to a particular biological process). However, the association of keywords to genes or protein sequences is a difficult process that ultimately implies examination of the literature related to a gene. RESULTS: To support this task, we present a procedure to derive keywords from the set of scientific abstracts related to a gene. Our system is based on the automated extraction of mappings between related terms from different databases using a model of fuzzy associations that can be applied with all generality to any pair of linked databases. We tested the system by annotating genes of the SWISS-PROT database with keywords derived from the abstracts linked to their entries (stored in the MEDLINE database of scientific references). The performance of the annotation procedure was much better for SWISS-PROT keywords (recall of 47%, precision of 68%) than for Gene Ontology terms (recall of 8%, precision of 67%). AVAILABILITY: The algorithm can be publicly accessed and used for the annotation of sequences through a web server at http://www.bork.embl.de/kat Antonio J. Pérez-Pulido, Carolina Perez-Iratxeta, Peer Bork, Guillermo Thode, Miguel A. Andrade-Navarro |
Bioinform. | 5 |
| 2003 | Evaluation of annotation strategies using an entire genome sequenceabstractAbstract Motivation: Genome-wide functional annotation either by manual or automatic means has raised considerable concerns regarding the accuracy of assignments and the reproducibility of methodologies. In addition, a performance evaluation of automated systems that attempt to tackle sequence analyses rapidly and reproducibly is generally missing. In order to quantify the accuracy and reproducibility of function assignments on a genome-wide scale, we have re-annotated the entire genome sequence of Chlamydia trachomatis (serovar D), in a collaborative manner. Results: We have encoded all annotations in a structured format to allow further comparison and data exchange and have used a scale that records the different levels of potential annotation errors according to their propensity to propagate in the database due to transitive function assignments. We conclude that genome annotation may entail a considerable amount of errors, ranging from simple typographical errors to complex sequence analysis problems. The most surprising result of this comparative study is that automatic systems might perform as well as the teams of experts annotating genome sequences. Availability and supplementary information: http://www.ebi.ac.uk/research/cgg/annotation/cteval/ Contact: [email protected] * To whom correspondence should be addressed. † INA-EKETA, GR-57001 Thessaloniki, Greece ‡ Computational Biology Center, Memorial Sloan-Kettering Cancer Center, New York, NY 10021, USA § Aetion Technologies LLC, Worthington, OH 43085, USA ¶ Institut Curie, F-75248 Paris, France ∥ CNRS, UMR6543, F-06108 Nice, France ** Alma Bioinformatics, E-28760 Madrid, Spain †† Cap Gemini Ernst & Young, London SW1X 7LX, UK ‡‡ Univ. of Rome ‘La Sapienza’, I-00185 Rome, Italy §§ MWG-Biotech AG, Ebersberg, D-85560 Berlin, Germany ¶¶ Wellcome Trust Biocentre, Univ. of Dundee, Dundee DD1 5HN, UK Sophia Tsoka, Miguel A. Andrade-Navarro, Anton J. Enright, Mark Carroll, Patrick Poullet, Vasilis J. Promponas, Theodore Liakopoulos, Giorgos Palaios, Claude Pasquier, Stavros J. Hamodrakas, Javier Tamames, Asutosh T. Yagnik, Anna Tramontano, Damien Devos, Christian Blaschke, Alfonso Valencia, David Brett, David M. A. Martin, Christophe Leroy, Isidore Rigoutsos, Chris Sander, Christos A. Ouzounis |
Bioinform. | 3 |
| 2003 | Information extraction from full text scientific articles: Where are the keywords?abstractBACKGROUND: To date, many of the methods for information extraction of biological information from scientific articles are restricted to the abstract of the article. However, full text articles in electronic version, which offer larger sources of data, are currently available. Several questions arise as to whether the effort of scanning full text articles is worthy, or whether the information that can be extracted from the different sections of an article can be relevant. RESULTS: In this work we addressed those questions showing that the keyword content of the different sections of a standard scientific article (abstract, introduction, methods, results, and discussion) is very heterogeneous. CONCLUSIONS: Although the abstract contains the best ratio of keywords per total of words, other sections of the article may be a better source of biologically relevant data. Parantu K. Shah, Carolina Perez-Iratxeta, Peer Bork, Miguel A. Andrade-Navarro |
BMC Bioinform. | 4 |
| 2000 | NAIL-Network Analysis Interface for Linking HMMER resultsabstractAbstract Summary: Network Analysis Interface for Linking HMMER results (NAIL) is a web-based tool for the analysis of results from a HMMER protein database-search. NAIL facilitates the selection of protein hits and the creation of an alignment, which can be used for a new sequence similarity search. Availability: From http://www.bork.embl-heidelberg.de/NAIL/ Contact: [email protected] * To whom correspondence should be addressed. Luis Sánchez-Pulido, Yan P. Yuan, Miguel A. Andrade-Navarro, Peer Bork |
Bioinform. | 3 |
| 1999 | Position-Specific Annotation of Protein Function Based on Multiple Homologs
Miguel A. Andrade-Navarro |
ISMB | 1 |
| 1999 | Automatic Extraction of Biological Information from Scientific Text: Protein-Protein Interactions
Christian Blaschke, Miguel A. Andrade-Navarro, Christos A. Ouzounis, Alfonso Valencia |
ISMB | 2 |
| 1999 | Automated genome sequence analysis and annotationabstractMOTIVATION: Large-scale genome projects generate a rapidly increasing number of sequences, most of them biochemically uncharacterized. Research in bioinformatics contributes to the development of methods for the computational characterization of these sequences. However, the installation and application of these methods require experience and are time consuming. RESULTS: We present here an automatic system for preliminary functional annotation of protein sequences that has been applied to the analysis of sets of sequences from complete genomes, both to refine overall performance and to make new discoveries comparable to those made by human experts. The GeneQuiz system includes a Web-based browser that allows examination of the evidence leading to an automatic annotation and offers additional information, views of the results, and links to biological databases that complement the automatic analysis. System structure and operating principles concerning the use of multiple sequence databases, underlying sequence analysis tools, lexical analyses of database annotations and decision criteria for functional assignments are detailed. The system makes automatic quality assessments of results based on prior experience with the underlying sequence analysis tools; overall error rates in functional assignment are estimated at 2.5-5% for cases annotated with highest reliability ('clear' cases). Sources of over-interpretation of results are discussed with proposals for improvement. A conservative definition for reporting 'new findings' that takes account of database maturity is presented along with examples of possible kinds of discoveries (new function, family and superfamily) made by the system. System performance in relation to sequence database coverage, database dynamics and database search methods is analysed, demonstrating the inherent advantages of an integrated automatic approach using multiple databases and search methods applied in an objective and repeatable manner. AVAILABILITY: The GeneQuiz system is publicly available for analysis of protein sequences through a Web server at http://www.sander.ebi.ac. uk/gqsrv/submit Miguel A. Andrade-Navarro, Nigel P. Brown, Christophe Leroy, S. Hörsch, Antoine de Daruvar, C. Reich, Angelo Franchini, Javier Tamames, Alfonso Valencia, Christos A. Ouzounis, Chris Sander |
Bioinform. | 1 |
| 1998 | Automatic extraction of keywords from scientific text: application to the knowledge domain of protein familiesabstractMOTIVATION: Annotation of the biological function of different protein sequences is a time-consuming process currently performed by human experts. Genome analysis tools encounter great difficulty in performing this task. Database curators, developers of genome analysis tools and biologists in general could benefit from access to tools able to suggest functional annotations and facilitate access to functional information. APPROACH: We present here the first prototype of a system for the automatic annotation of protein function. The system is triggered by collections of s related to a given protein, and it is able to extract biological information directly from scientific literature, i.e. MEDLINE abstracts. Relevant keywords are selected by their relative accumulation in comparison with a domain-specific background distribution. Simultaneously, the most representative sentences and MEDLINE abstracts are selected and presented to the end-user. Evolutionary information is considered as a predominant characteristic in the domain of protein function. Our system consequently extracts domain-specific information from the analysis of a set of protein families. RESULTS: The system has been tested with different protein families, of which three examples are discussed in detail here: 'ataxia-telangiectasia associated protein', 'ran GTPase' and 'carbonic anhydrase'. We found generally good correlation between the amount of information provided to the system and the quality of the annotations. Finally, the current limitations and future developments of the system are discussed. AVAILABILITY: The current system can be considered as a prototype system. As such, it can be accessed as a server at http://columba.ebi.ac. uk:8765/andrade/abx. The system accepts text related to the protein or proteins to be evaluated (optimally, the result of a MEDLINE search by keyword) and the results are returned in the form of Web pages for keywords, sentences and s. SUPPLEMENTARY INFORMATION: Web pages containing full information on the examples mentioned in the text are available at: http://www.cnb.uam.es/ approximately cnbprot/keywords/ CONTACT: [email protected] Miguel A. Andrade-Navarro, Alfonso Valencia |
Bioinform. | 1 |
| 1998 | Computational space reduction and parallelization of a new clustering approach for large groups of sequencesabstractMOTIVATION: The explosive growth of the biological sequences databases stimulated by genome projects has modified the framework of several applications in the biological sequence analysis area. In most cases, this new scenario is characterized by studies on large sets of sequences, suggesting the need for effective and automatic methods for their clustering. A more effective clustering of the database could be followed by the application of common family analysis schemes to the groups so formed. RESULTS: In this work, we present a new strategy to reduce the computational cost associated with the clustering of large sets of sequences which are expected to contain several families. The strategy is based on the grouping of the sequences into families by using a dynamic threshold on a pairwise sequence similarity criterion. Routine clustering of large data sets can now be done very efficiently. The method developed here achieves a computational space reduction of about an order of magnitude over more traditional ones of all-versus-all comparisons. The outcome of this approach produces family groupings that reproduce closely already accepted biological results. Our work includes a parallel implementation for distributed memory multiprocessors with a dynamic scheduling strategy for performance optimization. AVAILABILITY: By anonymous ftp at ftp.ac.uma.es (/pub/ots/pCluster directory), or from our Web site http://www.cnb. uam.es/www/software/software_index.html CONTACT: [email protected] Oswaldo Trelles, Miguel A. Andrade-Navarro, Alfonso Valencia, Emilio L. Zapata, José María Carazo |
Bioinform. | 2 |
| 1997 | Automatic Annotation for Biological Sequences by Etraction of Keywords from MEDLINE Abstracts: Development of a Prototype System
Miguel A. Andrade-Navarro, Alfonso Valencia |
ISMB | 1 |
| 1997 | Sequence analysis of the Methanococcus jannaschii genome and the prediction of protein functionabstractMiguel Andrade, Georg Casari, Antoine de Daruvar, Chris Sander, Reinhard Schneider, Javier Tamames, Alfonso Valencia, Christos Ouzounis; Sequence analysis Miguel A. Andrade-Navarro, Georg Casari, Antoine de Daruvar, Chris Sander, Reinhard Schneider 0002, Javier Tamames, Alfonso Valencia, Christos A. Ouzounis |
Comput. Appl. Biosci. | 1 |
| 1997 | Receptive Field Map Development by Anti-Hebbian Learning
Miguel A. Andrade-Navarro, Federico Morán |
Neural Networks | 1 |
| 1994 | Proteinotopic feature maps
Juan Julián Merelo Guervós, Miguel A. Andrade-Navarro, Alberto Prieto, Federico Morán |
Neurocomputing | 2 |