Marta E. Alarcón-Riquelme

dblp:197/8393 · DBLP profile ↗
← Back
11ranked-venue papers
0as first author
7since 2021 · last 2025
0000-0002-7632-4154ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 11 · 7 since 2021
YearPublicationVenuePosition
2025 Benchmarking single-sample gene set scoring methods for application in precision medicine
abstract
Gene set-based single-sample scoring methods are promising to elucidate patient level disease heterogeneity and enable functional interpretation of molecular data for precision medicine approaches. Despite the availability of numerous algorithms, their performance under different scenarios and for downstream applications for precision medicine approaches has not been systematically evaluated. In this study, we conducted a comprehensive survey of an exhaustive list of single-sample scoring methods to assess their stability and reproducibility performances under commo scenarios which include limitations of input data or data integration across studies. We also evaluated their performances for downstream patient stratification and clinical association analyses, as well as predictive modeling of disease states. The in-depth characterization of these scoring methods highlights the importance for a rational design of analysis strategies and provides fundamental insights into method selection under different scenarios or for different applications.
Daniel Toro-Domínguez, Iván Ellson-Lancho, Jordi Martorell-Marugan, Raúl López-Domínguez, Pedro Carmona-Saez, Marta E. Alarcón-Riquelme, Frédéric Baribaud
Briefings Bioinform.7
2025 BiomiX, a user-friendly bioinformatic tool for democratized analysis and integration of multiomics data
abstract
BACKGROUND: Interpreting biological system changes requires interpreting vast amounts of multi-omics data. While user-friendly tools exist for single-omics analysis, integrating multiple omics still requires bioinformatics expertise, limiting accessibility for the broader scientific community. RESULTS: BiomiX tackles the bottleneck in high-throughput omics data analysis, enabling efficient and integrated analysis of multiomics data obtained from two cohorts. BiomiX incorporates diverse omics data, using DESeq2/Limma packages for transcriptomics, and quantifying metabolomics peak differences, evaluated via the Wilcoxon test with the False Discovery Rate correction. The metabolomics annotation for Liquid Chromatography-Mass Spectrometry untargeted metabolomics is additionally supported using the mass-to-charge ratio in the CEU Mass Mediator database and fragmentation spectra in the TidyMass package. Methylomics analysis is performed using the ChAMP R package. Finally, Multi-Omics Factor Analysis (MOFA) integration identifies shared sources of variation across omics data. BiomiX also generates statistics, report figures and integrates EnrichR and GSEA for biological process exploration and subgroup analysis based on user-defined gene panels enhancing condition subtyping. BiomiX fine-tunes MOFA models, to optimize factors number selection, distinguishing between cohorts and providing tools to interpret discriminative MOFA factors. The interpretation relies on innovative bibliography research on Pubmed, which provides the articles most related to the discriminant factor contributors. Furthermore, discriminant MOFA factors are correlated with clinical data, and the top contributing pathways are explored, all with the aim of guiding the user in factor interpretation. CONCLUSIONS: The analysis of single-omics and multi-omics integration in a standalone tool, along with MOFA implementation and its interpretability via literature, represents significant progress in the multi-omics field in line with the "Findable, Accessible, Interoperable, and Reusable" data principles. BiomiX offers a wide range of parameters and interactive data visualization, allowing for personalized analysis tailored to user needs. This R-based, user-friendly tool is compatible with multiple operating systems and aims to make multi-omics analysis accessible to non-experts in bioinformatics.
Cristian Iperi, Álvaro Fernández-Ochoa, Guillermo Barturen, Jacques-Olivier Pers, Nathan Foulquier, Eléonore Bettacchioli, Marta E. Alarcón-Riquelme, Divi Cornec, Anne Bordron, Christophe Jamin
BMC Bioinform.7
2024 Response to the letter 'testing the effectiveness of MyPROSLE in classifying patients with lupus nephritis'
abstract
Recently, a letter to the Editor entitled ‘Testing the Effectiveness of MyPROSLE in Classifying Patients with Lupus Nephritis’ has been submitted by Leventhal et al. to Briefings in Bioinformatics. In this letter, the authors test MyPROSLE, a web application we recently introduced [1], to characterizelupus patients from the molecular point of view. Leventhal et al. tested the application with independent datasets reporting that the software ‘did not perform sufficiently well to consider replacement of the standard kidney biopsy as a diagnostic procedure’. In this letter, we address in detail all the concerns described by Leventhal et al. First of all, we would like to thank the authors for their interest and evaluation of the web tool. Nevertheless, we want to remark that in this work we did not intend to provide software for the replacement of standard clinical diagnostic procedures, as they stated. In our manuscript, we present a scoring system to summarize the molecular portrait of each patient and machine learning models based on these features are one of the analyses used to demonstrate the utility of this scoring system. In this context, MyPROSLE was developed to apply this scoring system to gene expression datasets in order to predict clinical features based on transcriptomics data. The letter is focused on the performance of these models but, although transcriptomics profiles have emerged as a valuable resource for making new and significant discoveries in diagnosis, the integration of these profiles into clinical practice is still a distant goal. Consequently, our software was not designed to replace existing diagnostic approaches in the clinical setting but a system that can provide additional information for clinical decisions when sufficient quality RNA-Seq data is available. In the current scenario, it should be used for exploratory analysis and hypothesis generation. This concept is what we embodied in the final sentence of our original article: ‘Therefore, we set a precedent and an important advance in terms of personalized research’ (through the development of an analytical workflow) ‘oriented to a near future clinical practice within autoimmunity’.
Daniel Toro-Domínguez, Jordi Martorell-Marugan, Manuel Martínez-Bueno, Raúl López-Domínguez, Elena Carnero-Montoro, Guillermo Barturen, Daniel Goldman, Michelle Petri, Pedro Carmona-Saez, Marta E. Alarcón-Riquelme
Briefings Bioinform.10
2024 Pheno-Ranker: a toolkit for comparison of phenotypic data stored in GA4GH standards and beyond
abstract
BACKGROUND: Phenotypic data comparison is essential for disease association studies, patient stratification, and genotype-phenotype correlation analysis. To support these efforts, the Global Alliance for Genomics and Health (GA4GH) established Phenopackets v2 and Beacon v2 standards for storing, sharing, and discovering genomic and phenotypic data. These standards provide a consistent framework for organizing biological data, simplifying their transformation into computer-friendly formats. However, matching participants using GA4GH-based formats remains challenging, as current methods are not fully compatible, limiting their effectiveness. RESULTS: Here, we introduce Pheno-Ranker, an open-source software toolkit for individual-level comparison of phenotypic data. As input, it accepts JSON/YAML data exchange formats from Beacon v2 and Phenopackets v2 data models, as well as any data structure encoded in JSON, YAML, or CSV formats. Internally, the hierarchical data structure is flattened to one dimension and then transformed through one-hot encoding. This allows for efficient pairwise (all-to-all) comparisons within cohorts or for matching of a patient's profile in cohorts. Users have the flexibility to refine their comparisons by including or excluding terms, applying weights to variables, and obtaining statistical significance through Z-scores and p-values. The output consists of text files, which can be further analyzed using unsupervised learning techniques, such as clustering or multidimensional scaling (MDS), and with graph analytics. Pheno-Ranker's performance has been validated with simulated and synthetic data, showing its accuracy, robustness, and efficiency across various health data scenarios. A real data use case from the PRECISESADS study highlights its practical utility in clinical research. CONCLUSIONS: Pheno-Ranker is a user-friendly, lightweight software for semantic similarity analysis of phenotypic data in Beacon v2 and Phenopackets v2 formats, extendable to other data types. It enables the comparison of a wide range of variables beyond HPO or OMIM terms while preserving full context. The software is designed as a command-line tool with additional utilities for CSV import, data simulation, summary statistics plotting, and QR code generation. For interactive analysis, it also includes a web-based user interface built with R Shiny. Links to the online documentation, including a Google Colab tutorial, and the tool's source code are available on the project home page: https://github.com/CNAG-Biomedical-Informatics/pheno-ranker .
Ivo C. Leist, María Rivas-Torrubia, Marta E. Alarcón-Riquelme, Guillermo Barturen, Ivo Glynne Gut, Manuel Rueda
BMC Bioinform.3
2022 Scoring personalized molecular portraits identify Systemic Lupus Erythematosus subtypes and predict individualized drug responses, symptomatology and disease progression
abstract
OBJECTIVES: Systemic Lupus Erythematosus is a complex autoimmune disease that leads to significant worsening of quality of life and mortality. Flares appear unpredictably during the disease course and therapies used are often only partially effective. These challenges are mainly due to the molecular heterogeneity of the disease, and in this context, personalized medicine-based approaches offer major promise. With this work we intended to advance in that direction by developing MyPROSLE, an omic-based analytical workflow for measuring the molecular portrait of individual patients to support clinicians in their therapeutic decisions. METHODS: Immunological gene-modules were used to represent the transcriptome of the patients. A dysregulation score for each gene-module was calculated at the patient level based on averaged z-scores. Almost 6100 Lupus and 750 healthy samples were used to analyze the association among dysregulation scores, clinical manifestations, prognosis, flare and remission events and response to Tabalumab. Machine learning-based classification models were built to predict around 100 different clinical parameters based on personalized dysregulation scores. RESULTS: MyPROSLE allows to molecularly summarize patients in 206 gene-modules, clustered into nine main lupus signatures. The combination of these modules revealed highly differentiated pathological mechanisms. We found that the dysregulation of certain gene-modules is strongly associated with specific clinical manifestations, the occurrence of relapses or the presence of long-term remission and drug response. Therefore, MyPROSLE may be used to accurately predict these clinical outcomes. CONCLUSIONS: MyPROSLE (https://myprosle.genyo.es) allows molecular characterization of individual Lupus patients and it extracts key molecular information to support more precise therapeutic decisions.
Daniel Toro-Domínguez, Jordi Martorell-Marugan, Manuel Martínez-Bueno, Raúl López-Domínguez, Elena Carnero-Montoro, Guillermo Barturen, Daniel Goldman, Michelle Petri, Pedro Carmona-Saez, Marta E. Alarcón-Riquelme
Briefings Bioinform.10
2021 A survey of gene expression meta-analysis: methods and applications
abstract
The increasing use of high-throughput gene expression quantification technologies over the last two decades and the fact that most of the published studies are stored in public databases has triggered an explosion of studies available through public repositories. All this information offers an invaluable resource for reuse to generate new knowledge and scientific findings. In this context, great interest has been focused on meta-analysis methods to integrate and jointly analyze different gene expression datasets. In this work, we describe the main steps in the gene expression meta-analysis, from data preparation to the state-of-the art statistical methods. We also analyze the main types of applications and problems that can be approached in gene expression meta-analysis studies and provide a comparative overview of the available software and bioinformatics tools. Moreover, a practical guide for choosing the most appropriate method in each case is also provided.
Daniel Toro-Domínguez, Juan Antonio Villatoro-García, Jordi Martorell-Marugan, Yolanda Román-Montoya, Marta E. Alarcón-Riquelme, Pedro Carmona-Saez
Briefings Bioinform.5
2021 A comprehensive database for integrated analysis of omics data in autoimmune diseases
abstract
BACKGROUND: Autoimmune diseases are heterogeneous pathologies with difficult diagnosis and few therapeutic options. In the last decade, several omics studies have provided significant insights into the molecular mechanisms of these diseases. Nevertheless, data from different cohorts and pathologies are stored independently in public repositories and a unified resource is imperative to assist researchers in this field. RESULTS: Here, we present Autoimmune Diseases Explorer ( https://adex.genyo.es ), a database that integrates 82 curated transcriptomics and methylation studies covering 5609 samples for some of the most common autoimmune diseases. The database provides, in an easy-to-use environment, advanced data analysis and statistical methods for exploring omics datasets, including meta-analysis, differential expression or pathway analysis. CONCLUSIONS: This is the first omics database focused on autoimmune diseases. This resource incorporates homogeneously processed data to facilitate integrative analyses among studies.
Jordi Martorell-Marugan, Raúl López-Domínguez, Adrián García-Moreno, Daniel Toro-Domínguez, Juan Antonio Villatoro-García, Guillermo Barturen, Adoración Martín-Gómez, Kevin Troulé, Gonzalo Gómez-López, Fátima Al-Shahrour, Víctor González-Rumayor, María Peña-Chilet, Joaquín Dopazo, Julio Saez-Rodriguez, Marta E. Alarcón-Riquelme, Pedro Carmona-Saez
BMC Bioinform.15
2019 ImaGEO: integrative gene expression meta-analysis from GEO database
abstract
SUMMARY: The Gene Expression Omnibus (GEO) database provides an invaluable resource of publicly available gene expression data that can be integrated and analyzed to derive new hypothesis and knowledge. In this context, gene expression meta-analysis (geMAs) is increasingly used in several fields to improve study reproducibility and discovering robust biomarkers. Nevertheless, integrating data is not straightforward without bioinformatics expertise. Here, we present ImaGEO, a web tool for geMAs that implements a complete and comprehensive meta-analysis workflow starting from GEO dataset identifiers. The application integrates GEO datasets, applies different meta-analysis techniques and provides functional analysis results in an easy-to-use environment. ImaGEO is a powerful and useful resource that allows researchers to integrate and perform meta-analysis of GEO datasets to lead robust findings for biomarker discovery studies. AVAILABILITY AND IMPLEMENTATION: ImaGEO is accessible at http://bioinfo.genyo.es/imageo/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Daniel Toro-Domínguez, Jordi Martorell-Marugan, Raúl López-Domínguez, Adrián García-Moreno, Víctor González-Rumayor, Marta E. Alarcón-Riquelme, Pedro Carmona-Saez
Bioinform.6
2017 Metagene projection characterizes GEN2.2 and CAL-1 as relevant human plasmacytoid dendritic cell models
abstract
MOTIVATION: Plasmacytoid dendritic cells (pDC) play a major role in the regulation of adaptive and innate immunity. Human pDC are difficult to isolate from peripheral blood and do not survive in culture making the study of their biology challenging. Recently, two leukemic counterparts of pDC, CAL-1 and GEN2.2, have been proposed as representative models of human pDC. Nevertheless, their relationship with pDC has been established only by means of particular functional and phenotypic similarities. With the aim of characterizing GEN2.2 and CAL-1 in the context of the main circulating immune cell populations we have performed microarray gene expression profiling of GEN2.2 and carried out an integrated analysis using publicly available gene expression datasets of CAL-1 and the main circulating primary leukocyte lineages. RESULTS: Our results show that GEN2.2 and CAL-1 share common gene expression programs with primary pDC, clustering apart from the rest of circulating hematopoietic lineages. We have also identified common differentially expressed genes that can be relevant in pDC biology. In addition, we have revealed the common and differential pathways activated in primary pDC and cell lines upon CpG stimulatio. AVAILABILITY AND IMPLEMENTATION: R code and data are available in the supplementary material. CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Pedro Carmona-Saez, Nieves Varela, María José Luque, Daniel Toro-Domínguez, Jordi Martorell-Marugan, Marta E. Alarcón-Riquelme, Concepción Marañón
Bioinform.6
2017 CymeR: cytometry analysis using KNIME, docker and R
abstract
Summary: Here we present open-source software for the analysis of high-dimensional cytometry data using state of the art algorithms. Importantly, use of the software requires no programming ability, and output files can either be interrogated directly in CymeR or they can be used downstream with any other cytometric data analysis platform. Also, because we use Docker to integrate the multitude of components that form the basis of CymeR, we have additionally developed a proof-of-concept of how future open-source bioinformatic programs with graphical user interfaces could be developed. Availability and Implementation: CymeR is open-source software that ties several components into a single program that is perhaps best thought of as a self-contained data analysis operating system. Please see https://github.com/bmuchmore/CymeR/wiki for detailed installation instructions. Contact: [email protected] or [email protected].
B. Muchmore, Marta E. Alarcón-Riquelme
Bioinform.2
2017 MetaGenyo: a web tool for meta-analysis of genetic association studies
abstract
BACKGROUND: Genetic association studies (GAS) aims to evaluate the association between genetic variants and phenotypes. In the last few years, the number of this type of study has increased exponentially, but the results are not always reproducible due to experimental designs, low sample sizes and other methodological errors. In this field, meta-analysis techniques are becoming very popular tools to combine results across studies to increase statistical power and to resolve discrepancies in genetic association studies. A meta-analysis summarizes research findings, increases statistical power and enables the identification of genuine associations between genotypes and phenotypes. Meta-analysis techniques are increasingly used in GAS, but it is also increasing the amount of published meta-analysis containing different errors. Although there are several software packages that implement meta-analysis, none of them are specifically designed for genetic association studies and in most cases their use requires advanced programming or scripting expertise. RESULTS: We have developed MetaGenyo, a web tool for meta-analysis in GAS. MetaGenyo implements a complete and comprehensive workflow that can be executed in an easy-to-use environment without programming knowledge. MetaGenyo has been developed to guide users through the main steps of a GAS meta-analysis, covering Hardy-Weinberg test, statistical association for different genetic models, analysis of heterogeneity, testing for publication bias, subgroup analysis and robustness testing of the results. CONCLUSIONS: MetaGenyo is a useful tool to conduct comprehensive genetic association meta-analysis. The application is freely available at http://bioinfo.genyo.es/metagenyo/ .
Jordi Martorell-Marugan, Daniel Toro-Domínguez, Marta E. Alarcón-Riquelme, Pedro Carmona-Saez
BMC Bioinform.3