Jordi Martorell-Marugan

dblp:210/9465 · DBLP profile ↗
← Back
13ranked-venue papers
5as first author
8since 2021 · last 2026
0000-0002-5186-0735ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 13 · 5 first-author · 8 since 2021
YearPublicationVenuePosition
2026 PathMED: an R toolkit for single-sample molecular scoring and machine learning with omics data
abstract
MOTIVATION: Molecular scoring is a popular approach for studying pathway-level functional alterations with omics data. Using molecular scores for tasks such as single-sample molecular characterisation, phenotype prediction or disease stratification has several advantages compared to using omics data directly. Molecular scores provide biological interpretability and are more generalisable across datasets, facilitating data integration and machine learning applications. However, numerous scoring methods are available through different software packages, and currently there is a lack of tools to easily use these scores for model training and prediction. RESULTS: We developed pathMED, an R/Bioconductor package that unifies various scoring methods in a simple framework. Furthermore, pathMED also contains a machine learning module to train and test models that use the calculated molecular scores to predict clinical outcomes. We demonstrate some of its potential applications in three use cases using public omics data. We showed the generalisability of machine learning models trained on transcriptomic scores in predicting clinical outcomes when deploying on proteomic scores. We also demonstrated the application of transcriptomics scores in predicting breast cancer treatment response and identifying pathways strongly associated to tumour biology and treatment response. Finally, we demonstrated the benefit of integrating a novel gene set dissection step into the analysis pipeline to resolve disease heterogeneity at the pathway level. AVAILABILITY: PathMED is freely available in the Bioconductor repository (https://bioconductor.org/packages/release/bioc/html/pathMED.html). Code to reproduce the analyses is publicly available at https://github.com/GENyO-BioInformatics/pathMED_article.
Jordi Martorell-Marugan, Iván Ellson-Lancho, Raúl López-Domínguez, Pablo Pedro Jurado-Bascón, Juan Antonio Villatoro-García, Frédéric Baribaud, Daniel Toro-Domínguez, Pedro Carmona-Saez
Bioinform.1
2025 Explainable deep neural networks for predicting sample phenotypes from single-cell transcriptomics
abstract
Recent advances in single-cell RNA-Sequencing (scRNA-Seq) technologies have revolutionized our ability to gather molecular insights into different phenotypes at the level of individual cells. The analysis of the resulting data poses significant challenges, and proper statistical methods are required to analyze and extract information from scRNA-Seq datasets. Sample classification based on gene expression data has proven effective and valuable for precision medicine applications. However, standard classification schemas are often not suitable for scRNA-Seq due to their unique characteristics, and new algorithms are required to effectively analyze and classify samples at the single-cell level. Furthermore, existing methods for this purpose have limitations in their usability. Those reasons motivated us to develop singleDeep, an end-to-end pipeline that streamlines the analysis of scRNA-Seq data training deep neural networks, enabling robust prediction and characterization of sample phenotypes. We used singleDeep to make predictions on scRNA-Seq datasets from different conditions, including systemic lupus erythematosus, Alzheimer's disease and coronavirus disease 2019. Our results demonstrate strong diagnostic performance, validated both internally and externally. Moreover, singleDeep outperformed traditional machine learning methods and alternative single-cell approaches. In addition to prediction accuracy, singleDeep provides valuable insights into cell types and gene importance estimation for phenotypic characterization. This functionality provided additional and valuable information in our use cases. For instance, we corroborated that some interferon signature genes are consistently relevant for autoimmunity across all immune cell types in lupus. On the other hand, we discovered that genes linked to dementia have relevant roles in specific brain cell populations, such as APOE in astrocytes.
Jordi Martorell-Marugan, Raúl López-Domínguez, Juan Antonio Villatoro-García, Daniel Toro-Domínguez, Marco Chierici, Giuseppe Jurman, Pedro Carmona-Saez
Briefings Bioinform.1
2025 Benchmarking single-sample gene set scoring methods for application in precision medicine
abstract
Gene set-based single-sample scoring methods are promising to elucidate patient level disease heterogeneity and enable functional interpretation of molecular data for precision medicine approaches. Despite the availability of numerous algorithms, their performance under different scenarios and for downstream applications for precision medicine approaches has not been systematically evaluated. In this study, we conducted a comprehensive survey of an exhaustive list of single-sample scoring methods to assess their stability and reproducibility performances under commo scenarios which include limitations of input data or data integration across studies. We also evaluated their performances for downstream patient stratification and clinical association analyses, as well as predictive modeling of disease states. The in-depth characterization of these scoring methods highlights the importance for a rational design of analysis strategies and provides fundamental insights into method selection under different scenarios or for different applications.
Daniel Toro-Domínguez, Iván Ellson-Lancho, Jordi Martorell-Marugan, Raúl López-Domínguez, Pedro Carmona-Saez, Marta E. Alarcón-Riquelme, Frédéric Baribaud
Briefings Bioinform.4
2024 Response to the letter 'testing the effectiveness of MyPROSLE in classifying patients with lupus nephritis'
abstract
Recently, a letter to the Editor entitled ‘Testing the Effectiveness of MyPROSLE in Classifying Patients with Lupus Nephritis’ has been submitted by Leventhal et al. to Briefings in Bioinformatics. In this letter, the authors test MyPROSLE, a web application we recently introduced [1], to characterizelupus patients from the molecular point of view. Leventhal et al. tested the application with independent datasets reporting that the software ‘did not perform sufficiently well to consider replacement of the standard kidney biopsy as a diagnostic procedure’. In this letter, we address in detail all the concerns described by Leventhal et al. First of all, we would like to thank the authors for their interest and evaluation of the web tool. Nevertheless, we want to remark that in this work we did not intend to provide software for the replacement of standard clinical diagnostic procedures, as they stated. In our manuscript, we present a scoring system to summarize the molecular portrait of each patient and machine learning models based on these features are one of the analyses used to demonstrate the utility of this scoring system. In this context, MyPROSLE was developed to apply this scoring system to gene expression datasets in order to predict clinical features based on transcriptomics data. The letter is focused on the performance of these models but, although transcriptomics profiles have emerged as a valuable resource for making new and significant discoveries in diagnosis, the integration of these profiles into clinical practice is still a distant goal. Consequently, our software was not designed to replace existing diagnostic approaches in the clinical setting but a system that can provide additional information for clinical decisions when sufficient quality RNA-Seq data is available. In the current scenario, it should be used for exploratory analysis and hypothesis generation. This concept is what we embodied in the final sentence of our original article: ‘Therefore, we set a precedent and an important advance in terms of personalized research’ (through the development of an analytical workflow) ‘oriented to a near future clinical practice within autoimmunity’.
Daniel Toro-Domínguez, Jordi Martorell-Marugan, Manuel Martínez-Bueno, Raúl López-Domínguez, Elena Carnero-Montoro, Guillermo Barturen, Daniel Goldman, Michelle Petri, Pedro Carmona-Saez, Marta E. Alarcón-Riquelme
Briefings Bioinform.2
2022 Scoring personalized molecular portraits identify Systemic Lupus Erythematosus subtypes and predict individualized drug responses, symptomatology and disease progression
abstract
OBJECTIVES: Systemic Lupus Erythematosus is a complex autoimmune disease that leads to significant worsening of quality of life and mortality. Flares appear unpredictably during the disease course and therapies used are often only partially effective. These challenges are mainly due to the molecular heterogeneity of the disease, and in this context, personalized medicine-based approaches offer major promise. With this work we intended to advance in that direction by developing MyPROSLE, an omic-based analytical workflow for measuring the molecular portrait of individual patients to support clinicians in their therapeutic decisions. METHODS: Immunological gene-modules were used to represent the transcriptome of the patients. A dysregulation score for each gene-module was calculated at the patient level based on averaged z-scores. Almost 6100 Lupus and 750 healthy samples were used to analyze the association among dysregulation scores, clinical manifestations, prognosis, flare and remission events and response to Tabalumab. Machine learning-based classification models were built to predict around 100 different clinical parameters based on personalized dysregulation scores. RESULTS: MyPROSLE allows to molecularly summarize patients in 206 gene-modules, clustered into nine main lupus signatures. The combination of these modules revealed highly differentiated pathological mechanisms. We found that the dysregulation of certain gene-modules is strongly associated with specific clinical manifestations, the occurrence of relapses or the presence of long-term remission and drug response. Therefore, MyPROSLE may be used to accurately predict these clinical outcomes. CONCLUSIONS: MyPROSLE (https://myprosle.genyo.es) allows molecular characterization of individual Lupus patients and it extracts key molecular information to support more precise therapeutic decisions.
Daniel Toro-Domínguez, Jordi Martorell-Marugan, Manuel Martínez-Bueno, Raúl López-Domínguez, Elena Carnero-Montoro, Guillermo Barturen, Daniel Goldman, Michelle Petri, Pedro Carmona-Saez, Marta E. Alarcón-Riquelme
Briefings Bioinform.2
2021 A survey of gene expression meta-analysis: methods and applications
abstract
The increasing use of high-throughput gene expression quantification technologies over the last two decades and the fact that most of the published studies are stored in public databases has triggered an explosion of studies available through public repositories. All this information offers an invaluable resource for reuse to generate new knowledge and scientific findings. In this context, great interest has been focused on meta-analysis methods to integrate and jointly analyze different gene expression datasets. In this work, we describe the main steps in the gene expression meta-analysis, from data preparation to the state-of-the art statistical methods. We also analyze the main types of applications and problems that can be approached in gene expression meta-analysis studies and provide a comparative overview of the available software and bioinformatics tools. Moreover, a practical guide for choosing the most appropriate method in each case is also provided.
Daniel Toro-Domínguez, Juan Antonio Villatoro-García, Jordi Martorell-Marugan, Yolanda Román-Montoya, Marta E. Alarcón-Riquelme, Pedro Carmona-Saez
Briefings Bioinform.3
2021 DREIMT: a drug repositioning database and prioritization tool for immunomodulation
abstract
MOTIVATION: Drug immunomodulation modifies the response of the immune system and can be therapeutically exploited in pathologies such as cancer and autoimmune diseases. RESULTS: DREIMT is a new hypothesis-generation web tool, which performs drug prioritization analysis for immunomodulation. DREIMT provides significant immunomodulatory drugs targeting up to 70 immune cells subtypes through a curated database that integrates 4960 drug profiles and ∼2600 immune gene expression signatures. The tool also suggests potential immunomodulatory drugs targeting user-supplied gene expression signatures. Final output includes drug-signature association scores, FDRs and downloadable plots and results tables. AVAILABILITYAND IMPLEMENTATION: http://www.dreimt.org. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Kevin Troulé, Hugo López-Fernández, Santiago García-Martín, Miguel Reboiro-Jato, Carlos Carretero-Puche, Jordi Martorell-Marugan, Guillermo Martín-Serrano, Pedro Carmona-Saez, Daniel Glez-Peña, Fátima Al-Shahrour, Gonzalo Gómez-López
Bioinform.6
2021 A comprehensive database for integrated analysis of omics data in autoimmune diseases
abstract
BACKGROUND: Autoimmune diseases are heterogeneous pathologies with difficult diagnosis and few therapeutic options. In the last decade, several omics studies have provided significant insights into the molecular mechanisms of these diseases. Nevertheless, data from different cohorts and pathologies are stored independently in public repositories and a unified resource is imperative to assist researchers in this field. RESULTS: Here, we present Autoimmune Diseases Explorer ( https://adex.genyo.es ), a database that integrates 82 curated transcriptomics and methylation studies covering 5609 samples for some of the most common autoimmune diseases. The database provides, in an easy-to-use environment, advanced data analysis and statistical methods for exploring omics datasets, including meta-analysis, differential expression or pathway analysis. CONCLUSIONS: This is the first omics database focused on autoimmune diseases. This resource incorporates homogeneously processed data to facilitate integrative analyses among studies.
Jordi Martorell-Marugan, Raúl López-Domínguez, Adrián García-Moreno, Daniel Toro-Domínguez, Juan Antonio Villatoro-García, Guillermo Barturen, Adoración Martín-Gómez, Kevin Troulé, Gonzalo Gómez-López, Fátima Al-Shahrour, Víctor González-Rumayor, María Peña-Chilet, Joaquín Dopazo, Julio Saez-Rodriguez, Marta E. Alarcón-Riquelme, Pedro Carmona-Saez
BMC Bioinform.1
2019 mCSEA: detecting subtle differentially methylated regions
abstract
MOTIVATION: The identification of differentially methylated regions (DMRs) among phenotypes is one of the main goals of epigenetic analysis. Although there are several methods developed to detect DMRs, most of them are focused on detecting relatively large differences in methylation levels and fail to detect moderate, but consistent, methylation changes that might be associated to complex disorders. RESULTS: We present mCSEA, an R package that implements a Gene Set Enrichment Analysis method to identify DMRs from Illumina450K and EPIC array data. It is especially useful for detecting subtle, but consistent, methylation differences in complex phenotypes. mCSEA also implements functions to integrate gene expression data and to detect genes with significant correlations among methylation and gene expression patterns. Using simulated datasets we show that mCSEA outperforms other tools in detecting DMRs. In addition, we applied mCSEA to a previously published dataset of sibling pairs discordant for intrauterine hyperglycemia exposure. We found several differentially methylated promoters in genes related to metabolic disorders like obesity and diabetes, demonstrating the potential of mCSEA to identify DMRs not detected by other methods. AVAILABILITY AND IMPLEMENTATION: mCSEA is freely available from the Bioconductor repository. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jordi Martorell-Marugan, Víctor González-Rumayor, Pedro Carmona-Saez
Bioinform.1
2019 ImaGEO: integrative gene expression meta-analysis from GEO database
abstract
SUMMARY: The Gene Expression Omnibus (GEO) database provides an invaluable resource of publicly available gene expression data that can be integrated and analyzed to derive new hypothesis and knowledge. In this context, gene expression meta-analysis (geMAs) is increasingly used in several fields to improve study reproducibility and discovering robust biomarkers. Nevertheless, integrating data is not straightforward without bioinformatics expertise. Here, we present ImaGEO, a web tool for geMAs that implements a complete and comprehensive meta-analysis workflow starting from GEO dataset identifiers. The application integrates GEO datasets, applies different meta-analysis techniques and provides functional analysis results in an easy-to-use environment. ImaGEO is a powerful and useful resource that allows researchers to integrate and perform meta-analysis of GEO datasets to lead robust findings for biomarker discovery studies. AVAILABILITY AND IMPLEMENTATION: ImaGEO is accessible at http://bioinfo.genyo.es/imageo/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Daniel Toro-Domínguez, Jordi Martorell-Marugan, Raúl López-Domínguez, Adrián García-Moreno, Víctor González-Rumayor, Marta E. Alarcón-Riquelme, Pedro Carmona-Saez
Bioinform.2
2018 Identification and visualization of differential isoform expression in RNA-seq time series
abstract
Motivation: As sequencing technologies improve their capacity to detect distinct transcripts of the same gene and to address complex experimental designs such as longitudinal studies, there is a need to develop statistical methods for the analysis of isoform expression changes in time series data. Results: Iso-maSigPro is a new functionality of the R package maSigPro for transcriptomics time series data analysis. Iso-maSigPro identifies genes with a differential isoform usage across time. The package also includes new clustering and visualization functions that allow grouping of genes with similar expression patterns at the isoform level, as well as those genes with a shift in major expressed isoform. Availability and implementation: The package is freely available under the LGPL license from the Bioconductor web site. Contact: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
María José Nueda, Jordi Martorell-Marugan, Cristina Martí, Sonia Tarazona, Ana Conesa
Bioinform.2
2017 Metagene projection characterizes GEN2.2 and CAL-1 as relevant human plasmacytoid dendritic cell models
abstract
MOTIVATION: Plasmacytoid dendritic cells (pDC) play a major role in the regulation of adaptive and innate immunity. Human pDC are difficult to isolate from peripheral blood and do not survive in culture making the study of their biology challenging. Recently, two leukemic counterparts of pDC, CAL-1 and GEN2.2, have been proposed as representative models of human pDC. Nevertheless, their relationship with pDC has been established only by means of particular functional and phenotypic similarities. With the aim of characterizing GEN2.2 and CAL-1 in the context of the main circulating immune cell populations we have performed microarray gene expression profiling of GEN2.2 and carried out an integrated analysis using publicly available gene expression datasets of CAL-1 and the main circulating primary leukocyte lineages. RESULTS: Our results show that GEN2.2 and CAL-1 share common gene expression programs with primary pDC, clustering apart from the rest of circulating hematopoietic lineages. We have also identified common differentially expressed genes that can be relevant in pDC biology. In addition, we have revealed the common and differential pathways activated in primary pDC and cell lines upon CpG stimulatio. AVAILABILITY AND IMPLEMENTATION: R code and data are available in the supplementary material. CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Pedro Carmona-Saez, Nieves Varela, María José Luque, Daniel Toro-Domínguez, Jordi Martorell-Marugan, Marta E. Alarcón-Riquelme, Concepción Marañón
Bioinform.5
2017 MetaGenyo: a web tool for meta-analysis of genetic association studies
abstract
BACKGROUND: Genetic association studies (GAS) aims to evaluate the association between genetic variants and phenotypes. In the last few years, the number of this type of study has increased exponentially, but the results are not always reproducible due to experimental designs, low sample sizes and other methodological errors. In this field, meta-analysis techniques are becoming very popular tools to combine results across studies to increase statistical power and to resolve discrepancies in genetic association studies. A meta-analysis summarizes research findings, increases statistical power and enables the identification of genuine associations between genotypes and phenotypes. Meta-analysis techniques are increasingly used in GAS, but it is also increasing the amount of published meta-analysis containing different errors. Although there are several software packages that implement meta-analysis, none of them are specifically designed for genetic association studies and in most cases their use requires advanced programming or scripting expertise. RESULTS: We have developed MetaGenyo, a web tool for meta-analysis in GAS. MetaGenyo implements a complete and comprehensive workflow that can be executed in an easy-to-use environment without programming knowledge. MetaGenyo has been developed to guide users through the main steps of a GAS meta-analysis, covering Hardy-Weinberg test, statistical association for different genetic models, analysis of heterogeneity, testing for publication bias, subgroup analysis and robustness testing of the results. CONCLUSIONS: MetaGenyo is a useful tool to conduct comprehensive genetic association meta-analysis. The application is freely available at http://bioinfo.genyo.es/metagenyo/ .
Jordi Martorell-Marugan, Daniel Toro-Domínguez, Marta E. Alarcón-Riquelme, Pedro Carmona-Saez
BMC Bioinform.1