Juan Garcia Ranea

dblp:11/1500 · also Juan A. G. Ranea, Juan Antonio Garcia-Ranea, Juan Antonio Ranea · DBLP profile ↗
← Back
16ranked-venue papers
2as first author
7since 2021 · last 2026
0000-0003-0327-1837ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 14 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Predicting gene compensation in disease with graph embedding techniques
abstract
Genetic compensation plays a critical role in mitigating the effects of deleterious mutations in genetic diseases. Identifying functionally compensatory gene relationships represents a promising strategy for discovering therapeutic targets in inherited genetic disorders and cancer. We present a novel approach that combines multiplexed network analysis with graph embedding techniques to predict compensatory genes. Our method constructs a multi-source gene similarity network, embedding each source with kernels and node2vec methods, to finally integrate all sources in a comprehensive functional similarity gene network. Our approach demonstrates high predictive performance, successfully prioritizing known compensatory genes in disorders such as Duchenne muscular dystrophy, spinal muscular atrophy, and β-thalassemia. Furthermore, we extend its application to cancer, where genetic compensation mechanisms contribute to treatment resistance. Notable examples include androgen receptor (AR) in prostate cancer and RBL1 suppression compensation by RBL2 in breast cancer. These results demonstrate that embedding network representation is useful for prioritizing compensatory genes.
Federico García-Criado, Jesús Pérez-García, Elena Rojano, Pedro Seoane, Juan Garcia Ranea
Artif. Intell. Medicine5
2025 Advancing edge-based clustering and graph embedding for biological network analysis: a case study in RASopathies
abstract
Understanding and predicting biological processes from protein-protein interaction (PPI) networks requires accurate and efficient representations of their structure. However, many existing methods fail to capture the complex, overlapping modular structure of biological systems. To address this, we propose a network embedding strategy that improves both biological interpretability and predictive power. By transforming networks into a low-dimensional space while preserving key topological properties, embedding enables the discovery of novel functional relationships. Pre-clustering a network before embedding enhances representation quality, i.e. the ability to preserve meaningful structural and functional properties in the embedding space. However, traditional non-overlapping clustering methods can introduce bias by ignoring the overlapping nature of biological communities. We overcome this limitation by integrating the Hierarchical Link Clustering (HLC) algorithm into an embedding workflow tailored for large, weighted, undirected networks. First, we introduce two optimized HLC implementations for Python and R, both outperforming existing methods in clustering accuracy and scalability. Then, by restricting random walks to HLC-defined communities, we improve the representation of biological pathways, as shown using Reactome on the human PPI network. We also apply our full cluster embedding workflow to analyze RASopathies, a group of interrelated disorders with a diverse range of phenotypes, caused by mutations in genes from the RAS/MAPK pathway. This approach was used not only to represent known pathways, but also to identify potential novel gene candidates associated with RASopathies, including Noonan and Costello syndrome. HLC implementations are available in the CDLIB library (https://github.com/GiulioRossetti/cdlib), and at https://github.com/jimrperkins/linkcomm for Python and R, respectively.
Federico García-Criado, Pedro Seoane, Elena Rojano, Juan Garcia Ranea, James Richard Perkins
Briefings Bioinform.4
2024 Exploring miRNA-target gene pair detection in disease with coRmiT
abstract
A wide range of approaches can be used to detect micro RNA (miRNA)-target gene pairs (mTPs) from expression data, differing in the ways the gene and miRNA expression profiles are calculated, combined and correlated. However, there is no clear consensus on which is the best approach across all datasets. Here, we have implemented multiple strategies and applied them to three distinct rare disease datasets that comprise smallRNA-Seq and RNA-Seq data obtained from the same samples, obtaining mTPs related to the disease pathology. All datasets were preprocessed using a standardized, freely available computational workflow, DEG_workflow. This workflow includes coRmiT, a method to compare multiple strategies for mTP detection. We used it to investigate the overlap of the detected mTPs with predicted and validated mTPs from 11 different databases. Results show that there is no clear best strategy for mTP detection applicable to all situations. We therefore propose the integration of the results of the different strategies by selecting the one with the highest odds ratio for each miRNA, as the optimal way to integrate the results. We applied this selection-integration method to the datasets and showed it to be robust to changes in the predicted and validated mTP databases. Our findings have important implications for miRNA analysis. coRmiT is implemented as part of the ExpHunterSuite Bioconductor package available from https://bioconductor.org/packages/ExpHunterSuite.
José Córdoba-Caballero, James Richard Perkins, Federico García-Criado, Diana Gallego, Alicia Navarro-Sánchez, Mireia Moreno-Estellés, Concepción Garcés, Fernando Bonet, Carlos Romá-Mateo, Rocio Toro, Belén Perez, Pascual Sanz, Matthias Kohl, Elena Rojano, Pedro Seoane, Juan Garcia Ranea
Briefings Bioinform.16
2023 Integrating differential expression, co-expression and gene network analysis for the identification of common genes associated with tumor angiogenesis deregulation
abstract
Angiogenesis is essential for tumor growth and cancer metastasis. Identifying the molecular pathways involved in this process is the first step in the rational design of new therapeutic strategies to improve cancer treatment. In recent years, RNA-seq data analysis has helped to determine the genetic and molecular factors associated with different types of cancer. In this work we performed integrative analysis using RNA-seq data from human umbilical vein endothelial cells (HUVEC) and patients with angiogenesis-dependent diseases to find genes that serve as potential candidates to improve the prognosis of tumor angiogenesis deregulation and understand how this process is orchestrated at the genetic and molecular level. We downloaded four RNA-seq datasets (including cellular models of tumor angiogenesis and ischaemic heart disease) from the Sequence Read Archive. Our integrative analysis includes a first step to determine differentially and co-expressed genes. For this, we used the ExpHunter Suite, an R package that performs differential expression, co-expression and functional analysis of RNA-seq data. We used both differentially and co-expressed genes to explore the human gene interaction network and determine which genes were found in the different datasets that may be key for the angiogenesis deregulation. Finally, we performed drug repositioning analysis to find potential targets related to angiogenesis inhibition. We found that that among the transcriptional alterations identified, SEMA3D and IL33 genes are deregulated in all datasets. Microenvironment remodeling, cell cycle, lipid metabolism and vesicular transport are the main molecular pathways affected. In addition to this, interacting genes are involved in intracellular signaling pathways, especially in immune system and semaphorins, respiratory electron transport and fatty acid metabolism. The methodology presented here can be used for finding common transcriptional alterations in other genetically-based diseases.
Beatriz Monterde, Elena Rojano, José Córdoba-Caballero, Pedro Seoane, James Richard Perkins, Miguel Angel Medina, Juan Garcia Ranea
J. Biomed. Informatics7
2022 Deepening the knowledge of rare diseases dependent on angiogenesis through semantic similarity clustering and network analysis
abstract
BACKGROUND: Angiogenesis is regulated by multiple genes whose variants can lead to different disorders. Among them, rare diseases are a heterogeneous group of pathologies, most of them genetic, whose information may be of interest to determine the still unknown genetic and molecular causes of other diseases. In this work, we use the information on rare diseases dependent on angiogenesis to investigate the genes that are associated with this biological process and to determine if there are interactions between the genes involved in its deregulation. RESULTS: We propose a systemic approach supported by the use of pathological phenotypes to group diseases by semantic similarity. We grouped 158 angiogenesis-related rare diseases in 18 clusters based on their phenotypes. Of them, 16 clusters had traceable gene connections in a high-quality interaction network. These disease clusters are associated with 130 different genes. We searched for genes associated with angiogenesis througth ClinVar pathogenic variants. Of the seven retrieved genes, our system confirms six of them. Furthermore, it allowed us to identify common affected functions among these disease clusters. AVAILABILITY: https://github.com/ElenaRojano/angio_cluster. CONTACT: [email protected] and [email protected].
Raquel Pagano-Márquez, José Córdoba-Caballero, Beatriz Martínez-Poveda, Ana R. Quesada, Elena Rojano, Pedro Seoane, Juan Garcia Ranea, Miguel Angel Medina
Briefings Bioinform.7
2022 Assigning protein function from domain-function associations using DomFun
abstract
BACKGROUND: Protein function prediction remains a key challenge. Domain composition affects protein function. Here we present DomFun, a Ruby gem that uses associations between protein domains and functions, calculated using multiple indices based on tripartite network analysis. These domain-function associations are combined at the protein level, to generate protein-function predictions. RESULTS: We analysed 16 tripartite networks connecting homologous superfamily and FunFam domains from CATH-Gene3D with functional annotations from the three Gene Ontology (GO) sub-ontologies, KEGG, and Reactome. We validated the results using the CAFA 3 benchmark platform for GO annotation, finding that out of the multiple association metrics and domain datasets tested, Simpson index for FunFam domain-function associations combined with Stouffer's method leads to the best performance in almost all scenarios. We also found that using FunFams led to better performance than superfamilies, and better results were found for GO molecular function compared to GO biological process terms. DomFun performed as well as the highest-performing method in certain CAFA 3 evaluation procedures in terms of [Formula: see text] and [Formula: see text] We also implemented our own benchmark procedure, Pathway Prediction Performance (PPP), which can be used to validate function prediction for additional annotations sources, such as KEGG and Reactome. Using PPP, we found similar results to those found with CAFA 3 for GO, moreover we found good performance for the other annotation sources. As with CAFA 3, Simpson index with Stouffer's method led to the top performance in almost all scenarios. CONCLUSIONS: DomFun shows competitive performance with other methods evaluated in CAFA 3 when predicting proteins function with GO, although results vary depending on the evaluation procedure. Through our own benchmark procedure, PPP, we have shown it can also make accurate predictions for KEGG and Reactome. It performs best when using FunFams, combining Simpson index derived domain-function associations using Stouffer's method. The tool has been implemented so that it can be easily adapted to incorporate other protein features, such as domain data from other sources, amino acid k-mers and motifs. The DomFun Ruby gem is available from https://rubygems.org/gems/DomFun . Code maintained at https://github.com/ElenaRojano/DomFun . Validation procedure scripts can be found at https://github.com/ElenaRojano/DomFun_project .
Elena Rojano, Fernando Moreno Jabato, James Richard Perkins, José Córdoba-Caballero, Federico García-Criado, Ian Sillitoe, Christine A. Orengo, Juan Garcia Ranea, Pedro Seoane
BMC Bioinform.8
2021 Protein residues determining interaction specificity in paralogous families
abstract
MOTIVATION: Predicting the residues controlling a protein's interaction specificity is important not only to better understand its interactions but also to design mutations aimed at fine-tuning or swapping them as well. RESULTS: In this work, we present a methodology that combines sequence information (in the form of multiple sequence alignments) with interactome information to detect that kind of residues in paralogous families of proteins. The interactome is used to define pairwise similarities of interaction contexts for the proteins in the alignment. The method looks for alignment positions with patterns of amino-acid changes reflecting the similarities/differences in the interaction neighborhoods of the corresponding proteins. We tested this new methodology in a large set of human paralogous families with structurally characterized interactions, and discuss in detail the results for the RasH family. We show that this approach is a better predictor of interfacial residues than both, sequence conservation and an equivalent 'unsupervised' method that does not use interactome information. AVAILABILITY AND IMPLEMENTATION: http://csbg.cnb.csic.es/pazos/Xdet/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Borja Pitarch, Juan Garcia Ranea, Florencio Pazos
Bioinform.2
2019 Regulatory variants: from detection to predicting impact
abstract
Variants within non-coding genomic regions can greatly affect disease. In recent years, increasing focus has been given to these variants, and how they can alter regulatory elements, such as enhancers, transcription factor binding sites and DNA methylation regions. Such variants can be considered regulatory variants. Concurrently, much effort has been put into establishing international consortia to undertake large projects aimed at discovering regulatory elements in different tissues, cell lines and organisms, and probing the effects of genetic variants on regulation by measuring gene expression. Here, we describe methods and techniques for discovering disease-associated non-coding variants using sequencing technologies. We then explain the computational procedures that can be used for annotating these variants using the information from the aforementioned projects, and prediction of their putative effects, including potential pathogenicity, based on rule-based and machine learning approaches. We provide the details of techniques to validate these predictions, by mapping chromatin-chromatin and chromatin-protein interactions, and introduce Clustered Regularly Interspaced Short Palindromic Repeats-Associated Protein 9 (CRISPR-Cas9) technology, which has already been used in this field and is likely to have a big impact on its future evolution. We also give examples of regulatory variants associated with multiple complex diseases. This review is aimed at bioinformaticians interested in the characterization of regulatory variants, molecular biologists and geneticists interested in understanding more about the nature and potential role of such variants from a functional point of views, and clinicians who may wish to learn about variants in non-coding genomic regions associated with a given disease and find out what to do next to uncover how they impact on the underlying mechanisms.
Elena Rojano, Pedro Seoane, Juan Garcia Ranea, James Richard Perkins
Briefings Bioinform.3
2017 How can functional annotations be derived from profiles of phenotypic annotations?
abstract
BACKGROUND: Loss-of-function phenotypes are widely used to infer gene function using the principle that similar phenotypes are indicative of similar functions. However, converting phenotypic to functional annotations requires careful interpretation of phenotypic descriptions and assessment of phenotypic similarity. Understanding how functions and phenotypes are linked will be crucial for the development of methods for the automatic conversion of gene loss-of-function phenotypes to gene functional annotations. RESULTS: We explored the relation between cellular phenotypes from RNAi-based screens in human cells and gene annotations of cellular functions as provided by the Gene Ontology (GO). Comparing different similarity measures, we found that information content-based measures of phenotypic similarity were the best at capturing gene functional similarity. However, phenotypic similarities did not map to the Gene Ontology organization of gene function but to functions defined as groups of GO terms with shared gene annotations. CONCLUSIONS: Our observations have implications for the use and interpretation of phenotypic similarities as a proxy for gene functions both in RNAi screen data analysis and curation and in the prediction of disease genes.
Beatriz Serrano-Solano, Antonio Díaz Ramos, Jean-Karim Hériché, Juan Garcia Ranea
BMC Bioinform.4
2017 Erratum to: How can functional annotations be derived from profiles of phenotypic annotations?
Beatriz Serrano-Solano, Antonio Díaz Ramos, Jean-Karim Hériché, Juan Garcia Ranea
BMC Bioinform.4
2015 FUN-L: gene prioritization for RNAi screens
abstract
MOTIVATION: Most biological processes remain only partially characterized with many components still to be identified. Given that a whole genome can usually not be tested in a functional assay, identifying the genes most likely to be of interest is of critical importance to avoid wasting resources. RESULTS: Given a set of known functionally related genes and using a state-of-the-art approach to data integration and mining, our Functional Lists (FUN-L) method provides a ranked list of candidate genes for testing. Validation of predictions from FUN-L with independent RNAi screens confirms that FUN-L-produced lists are enriched in genes with the expected phenotypes. In this article, we describe a website front end to FUN-L. AVAILABILITY AND IMPLEMENTATION: The website is freely available to use at http://funl.org
Jonathan G. Lees, Jean-Karim Hériché, Ian Morilla, José María Fernández 0001, Priit Adler, Martin Krallinger, Jaak Vilo, Alfonso Valencia, Jan Ellenberg, Juan Garcia Ranea, Christine A. Orengo
Bioinform.10
2014 A Cloud-based GWAS Analysis Pipeline for Clinical Researchers
abstract
The cost of obtaining genome-scale biomedical data continues to drop rapidly, with many hospitals and universities being able to produce large amounts of data. Managing and analysing such ever-growing datasets is becoming a crucial issue. Cloud computing presents a good solution to this problem due to its flexibility in obtaining computational resources. However, it is essential to allow end-users with no experience to take advantage of the cloud computing model of elastic resource provisioning. This paper presents a workflow that allows the end user to perform the core steps of a genome wide association analysis, consisting of 1) uploading raw data files to the cloud, 2) genotype calling, i.e. converting the raw data into information on genome variation 3) quality assessment of the genotype data and subsequent filtering based on user-provided parameters and 4) downloading the filtered genotype data in a standard file format. A number of steps in this process are computationally intensive; moreover, the computational resources involved vary greatly depending on the size of the study, from a few samples to a few thousand. Therefore cloud computing provides an ideal solution to this problem. The paper describes in detail how the pipeline was implemented, focussing on the cloud infrastructure and the different tools and software involved for software management and GUI construction. The key contributions of this paper are: 1a) iIt presents a real world application of cloud-computing to address a critical problem in biomedicine, . 2b) iIt provides a thorough description of how such a pipeline was implemented, in terms of data management and user-interface, such that the end-user does not need to focus on the computational aspect but can instead concentrate on data analysis and biological interpretation of results, and. c3) wWe show how cloud-computing can be used in a more effective way through the parallelisation of the appropriate parts of the pipeline.
Paul Heinzlreiter, James Richard Perkins, Óscar Torreño Tirado, Tor Johan Mikael Karlsson, Juan Garcia Ranea, Andreas Mitterecker, Miguel Blanca, Oswaldo Trelles
CLOSER5
2013 Insights into polypharmacology from drug-domain associations
abstract
MOTIVATION: Polypharmacology (the ability of a single drug to affect multiple targets) is a key feature that may explain part of the decreasing success of conventional drug discovery strategies driven by the quest for drugs to act selectively on a single target. Most drug targets are proteins that are composed of domains (their structural and functional building blocks). RESULTS: In this work, we model drug-domain networks to explore the role of protein domains as drug targets and to explain drug polypharmacology in terms of the interactions between drugs and protein domains. We find that drugs are organized around a privileged set of druggable domains. CONCLUSIONS: Protein domains are a good proxy for drug targets, and drug polypharmacology emerges as a consequence of the multi-domain composition of proteins. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Aurelio A. Moya-García, Juan Garcia Ranea
Bioinform.2
2010 Finding the "Dark Matter" in Human and Yeast Protein Network Prediction and Modelling
abstract
Accurate modelling of biological systems requires a deeper and more complete knowledge about the molecular components and their functional associations than we currently have. Traditionally, new knowledge on protein associations generated by experiments has played a central role in systems modelling, in contrast to generally less trusted bio-computational predictions. However, we will not achieve realistic modelling of complex molecular systems if the current experimental designs lead to biased screenings of real protein networks and leave large, functionally important areas poorly characterised. To assess the likelihood of this, we have built comprehensive network models of the yeast and human proteomes by using a meta-statistical integration of diverse computationally predicted protein association datasets. We have compared these predicted networks against combined experimental datasets from seven biological resources at different level of statistical significance. These eukaryotic predicted networks resemble all the topological and noise features of the experimentally inferred networks in both species, and we also show that this observation is not due to random behaviour. In addition, the topology of the predicted networks contains information on true protein associations, beyond the constitutive first order binary predictions. We also observe that most of the reliable predicted protein associations are experimentally uncharacterised in our models, constituting the hidden or "dark matter" of networks by analogy to astronomical systems. Some of this dark matter shows enrichment of particular functions and contains key functional elements of protein networks, such as hubs associated with important functional areas like the regulation of Ras protein signal transduction in human cells. Thus, characterising this large and functionally important dark matter, elusive to established experimental designs, may be crucial for modelling biological systems. In any case, these predictions provide a valuable guide to these experimentally elusive regions.
Juan Garcia Ranea, Ian Morilla, Jonathan G. Lees, Adam James Reid, Corin Yeats, Andrew B. Clegg, Francisca Sánchez-Jiménez, Christine A. Orengo
PLoS Comput. Biol.1
2007 Predicting Protein Function with Hierarchical Phylogenetic Profiles: The Gene3D Phylo-Tuner Method Applied to Eukaryotic Genomes
abstract
"Phylogenetic profiling" is based on the hypothesis that during evolution functionally or physically interacting genes are likely to be inherited or eliminated in a codependent manner. Creating presence-absence profiles of orthologous genes is now a common and powerful way of identifying functionally associated genes. In this approach, correctly determining orthology, as a means of identifying functional equivalence between two genes, is a critical and nontrivial step and largely explains why previous work in this area has mainly focused on using presence-absence profiles in prokaryotic species. Here, we demonstrate that eukaryotic genomes have a high proportion of multigene families whose phylogenetic profile distributions are poor in presence-absence information content. This feature makes them prone to orthology mis-assignment and unsuited to standard profile-based prediction methods. Using CATH structural domain assignments from the Gene3D database for 13 complete eukaryotic genomes, we have developed a novel modification of the phylogenetic profiling method that uses genome copy number of each domain superfamily to predict functional relationships. In our approach, superfamilies are subclustered at ten levels of sequence identity-from 30% to 100%-and phylogenetic profiles built at each level. All the profiles are compared using normalised Euclidean distances to identify those with correlated changes in their domain copy number. We demonstrate that two protein families will "auto-tune" with strong co-evolutionary signals when their profiles are compared at the similarity levels that capture their functional relationship. Our method finds functional relationships that are not detectable by the conventional presence-absence profile comparisons, and it does not require a priori any fixed criteria to define orthologous genes.
Juan Garcia Ranea, Corin Yeats, Alastair Grant, Christine A. Orengo
PLoS Comput. Biol.1
2004 A structural perspective on genome evolution
abstract
At UCL we have developed several automated protocols for generating protein family resources (CATH; Gene3D). These resources can be used to perform comparative genome analyses in order to understand the evolution of protein families. Also to identify biologically and/or medically interesting families for which no structural data currently exists and which may therefore be important targets for structure genomics initiatives.The CATH domain structure database, established by Orengo and Thornton in 1993, now contains a significant proportion of protein structures from the PDB clustered into 1400 evolutionary families. Relationships have been identified using robust structure comparison methods (SSAP, CATHEDRAL). We have also benchmarked and optimised various 1D-profiles and HMM based protocols for assigning genome sequences to families within the resource (e.g. SAM-T99, SAMOSA, CATH-ISL).In this way we can assign structural data to a large proportion (up to 60%) of whole or partial sequences in completed genomes and >80% of genes coding for enzymes and other proteins in biochemical pathways. However, in order to include all families regardless of whether their structure is known or not, a new protein family resource has been developed (Gene3D). In Gene3D, complete genes have been clustered according to sequence similarity alone, using a robust clustering method (Pfscape). 120 completed genomes from all kingdoms have been clustered into 220,000 gene families, 70,000 of which contain 2 or more sequences. Subsequently, we have labelled those gene families for which CATH structural or Pfam functional domain annotations can be provided for all or part of the gene.Preliminary analysis of the genome annotations reveals that a significant proportion (up to 70%) of CATH annotated genes or gene regions in genomes are assigned to domain families that are common to all three kingdoms of life. However, only 20% of the genome sequences are assigned to gene families common to all kingdoms. Since a large proportion of these genes are multidomain proteins this supports the view that a great deal of functional diversity within the genomes has been achieved by combining domain modules in different ways.In collaboration with Professor Janet Thornton, we have analysed a subset of 56 bacterial genomes to determine the recurrence of specific domain structure families within the genomes. This revealed a small but essential group of universal, and in some cases, highly recurring domain families. For some size-dependent families, domain recurrence is highly correlated with increase in genome size, whilst in other size-independent families no correlation is observed. Statistical analysis allowed us to distinguish three groups. Within the size-dependent families we differentiated two groups: linearly-distributed and non-linearly-distributed. Functional annotation using the COGs revealed that these domains were predominantly involved in metabolism and regulation, respectively. Whilst a third group of Evenly-distributed size independent domains are primarily involved in protein translation and biosynthesis.By mapping CATH and Pfam domains families onto all the genome sequences in Gene3D we observe that a few hundred highly recurrent families are dominating at least 50% of whole or partial genome sequences. Many of these families are common to both prokaryotes and eukaryotes and are performing essential generic functions. In many of the largest families, significant divergence in sequence has been accompanied by modifications in structure and function. Targetting representatives in these families for structure determination will allow the structure genomics initiatives to map both fold and function space and reveal the mechanisms by which divergence in protein families promotes evolution of new functions.
David A. Lee, Alastair Grant, Ian Sillitoe, Mark Dibley, Juan Garcia Ranea, Christine A. Orengo
RECOMB5