EDBT 2026 Demo / reviewers in the wild / expert
Pedro Carmona-Saez
dblp:13/3801
· DBLP profile ↗
22ranked-venue papers
3as first author
10since 2021 · last 2026
0000-0002-6173-7255ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 20 · 3 first-author · 10 since 2021Databases, data management, data science and information retrieval · 2Artificial intelligence and machine learning · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PathMED: an R toolkit for single-sample molecular scoring and machine learning with omics dataabstractMOTIVATION: Molecular scoring is a popular approach for studying pathway-level functional alterations with omics data. Using molecular scores for tasks such as single-sample molecular characterisation, phenotype prediction or disease stratification has several advantages compared to using omics data directly. Molecular scores provide biological interpretability and are more generalisable across datasets, facilitating data integration and machine learning applications. However, numerous scoring methods are available through different software packages, and currently there is a lack of tools to easily use these scores for model training and prediction. RESULTS: We developed pathMED, an R/Bioconductor package that unifies various scoring methods in a simple framework. Furthermore, pathMED also contains a machine learning module to train and test models that use the calculated molecular scores to predict clinical outcomes. We demonstrate some of its potential applications in three use cases using public omics data. We showed the generalisability of machine learning models trained on transcriptomic scores in predicting clinical outcomes when deploying on proteomic scores. We also demonstrated the application of transcriptomics scores in predicting breast cancer treatment response and identifying pathways strongly associated to tumour biology and treatment response. Finally, we demonstrated the benefit of integrating a novel gene set dissection step into the analysis pipeline to resolve disease heterogeneity at the pathway level. AVAILABILITY: PathMED is freely available in the Bioconductor repository (https://bioconductor.org/packages/release/bioc/html/pathMED.html). Code to reproduce the analyses is publicly available at https://github.com/GENyO-BioInformatics/pathMED_article. Jordi Martorell-Marugan, Iván Ellson-Lancho, Raúl López-Domínguez, Pablo Pedro Jurado-Bascón, Juan Antonio Villatoro-García, Frédéric Baribaud, Daniel Toro-Domínguez, Pedro Carmona-Saez |
Bioinform. | 9 |
| 2026 | Effect Size-Driven Pathway Meta-Analysis for Gene Expression DataabstractThe proliferation of omics datasets in public repositories has created unprecedented opportunities for biomedical research but has also posed significant challenges for their integration, particularly due to missing genes and platform-specific discrepancies. Traditional gene expression meta-analysis often focuses on individual genes, leading to data loss and limited biological insights when there are missing genes across different studies. To address these limitations, we propose GSEMA (Gene Set Enrichment Meta-Analysis), a novel methodology that leverages single-sample enrichment scoring to aggregate gene expression data into pathway-level matrices. By applying meta-analysis techniques to enrichment scores, GSEMA preserves the magnitude and directionality of effects, enabling the definition of pathway activity across datasets. Using simulated data and case studies on Systemic Lupus Erythematosus (SLE) and Parkinson's Disease (PD), we demonstrate that GSEMA outperforms other methods in controlling false positive rates while providing meaningful biological interpretations. GSEMA methodology is implemented as an R package available on CRAN repository. Juan Antonio Villatoro-García, Pablo Pedro Jurado-Bascón, Pedro Carmona-Saez |
IEEE Trans. Comput. Biol. Bioinform. | 3 |
| 2025 | Explainable deep neural networks for predicting sample phenotypes from single-cell transcriptomicsabstractRecent advances in single-cell RNA-Sequencing (scRNA-Seq) technologies have revolutionized our ability to gather molecular insights into different phenotypes at the level of individual cells. The analysis of the resulting data poses significant challenges, and proper statistical methods are required to analyze and extract information from scRNA-Seq datasets. Sample classification based on gene expression data has proven effective and valuable for precision medicine applications. However, standard classification schemas are often not suitable for scRNA-Seq due to their unique characteristics, and new algorithms are required to effectively analyze and classify samples at the single-cell level. Furthermore, existing methods for this purpose have limitations in their usability. Those reasons motivated us to develop singleDeep, an end-to-end pipeline that streamlines the analysis of scRNA-Seq data training deep neural networks, enabling robust prediction and characterization of sample phenotypes. We used singleDeep to make predictions on scRNA-Seq datasets from different conditions, including systemic lupus erythematosus, Alzheimer's disease and coronavirus disease 2019. Our results demonstrate strong diagnostic performance, validated both internally and externally. Moreover, singleDeep outperformed traditional machine learning methods and alternative single-cell approaches. In addition to prediction accuracy, singleDeep provides valuable insights into cell types and gene importance estimation for phenotypic characterization. This functionality provided additional and valuable information in our use cases. For instance, we corroborated that some interferon signature genes are consistently relevant for autoimmunity across all immune cell types in lupus. On the other hand, we discovered that genes linked to dementia have relevant roles in specific brain cell populations, such as APOE in astrocytes. Jordi Martorell-Marugan, Raúl López-Domínguez, Juan Antonio Villatoro-García, Daniel Toro-Domínguez, Marco Chierici, Giuseppe Jurman, Pedro Carmona-Saez |
Briefings Bioinform. | 7 |
| 2025 | Benchmarking single-sample gene set scoring methods for application in precision medicineabstractGene set-based single-sample scoring methods are promising to elucidate patient level disease heterogeneity and enable functional interpretation of molecular data for precision medicine approaches. Despite the availability of numerous algorithms, their performance under different scenarios and for downstream applications for precision medicine approaches has not been systematically evaluated. In this study, we conducted a comprehensive survey of an exhaustive list of single-sample scoring methods to assess their stability and reproducibility performances under commo scenarios which include limitations of input data or data integration across studies. We also evaluated their performances for downstream patient stratification and clinical association analyses, as well as predictive modeling of disease states. The in-depth characterization of these scoring methods highlights the importance for a rational design of analysis strategies and provides fundamental insights into method selection under different scenarios or for different applications. Daniel Toro-Domínguez, Iván Ellson-Lancho, Jordi Martorell-Marugan, Raúl López-Domínguez, Pedro Carmona-Saez, Marta E. Alarcón-Riquelme, Frédéric Baribaud |
Briefings Bioinform. | 6 |
| 2025 | A Survey of Preprocessing Techniques for Flow Cytometry Data in Classification TasksabstractFlow cytometry is an advanced technique for analyzing cellular heterogeneity in biomedical research and clinical diagnostics. Its ability to generate multiparametric data has facilitated advancements in disease classification problems, particularly addressing challenges in distinguishing cell populations and predicting disease outcomes. In these classification tasks, most works consist on either labeling individual cells based on their phenotypic markers or categorizing patient samples as healthy or diseased. However, the complexity of flow cytometry data, characterized by hurdles such as spectral overlap, wide dynamic ranges or batch effects require the usage of preprocessing strategies prior to modeling analysis. This research provides a comprehensive survey of current preprocessing techniques for flow cytometry data used in classification tasks and discusses their specific applications, focusing on four key aspects of data treatment: signal compensation and transformation, batch effect mitigation, imperfect data treatment, and feature selection and class balance. Emphasis is placed on standardizing preprocessing workflows and addressing computational and analytical difficulties posed by the size of modern flow cytometry datasets. The paper also includes a discussion on future opportunities for improving preprocessing pipelines to improve the reproducibility of flow cytometry-based classification models. In short, this work serves as a reference for experts in the field, consolidating best practices in preprocessing and providing guidelines for the development of methodologies that optimize the classification of flow cytometry data. David Núñez-Nepomuceno, José A. Sáez, Pedro Carmona-Saez |
IEEE Trans. Comput. Biol. Bioinform. | 3 |
| 2024 | Response to the letter 'testing the effectiveness of MyPROSLE in classifying patients with lupus nephritis'abstractRecently, a letter to the Editor entitled ‘Testing the Effectiveness of MyPROSLE in Classifying Patients with Lupus Nephritis’ has been submitted by Leventhal et al. to Briefings in Bioinformatics. In this letter, the authors test MyPROSLE, a web application we recently introduced [1], to characterizelupus patients from the molecular point of view. Leventhal et al. tested the application with independent datasets reporting that the software ‘did not perform sufficiently well to consider replacement of the standard kidney biopsy as a diagnostic procedure’. In this letter, we address in detail all the concerns described by Leventhal et al. First of all, we would like to thank the authors for their interest and evaluation of the web tool. Nevertheless, we want to remark that in this work we did not intend to provide software for the replacement of standard clinical diagnostic procedures, as they stated. In our manuscript, we present a scoring system to summarize the molecular portrait of each patient and machine learning models based on these features are one of the analyses used to demonstrate the utility of this scoring system. In this context, MyPROSLE was developed to apply this scoring system to gene expression datasets in order to predict clinical features based on transcriptomics data. The letter is focused on the performance of these models but, although transcriptomics profiles have emerged as a valuable resource for making new and significant discoveries in diagnosis, the integration of these profiles into clinical practice is still a distant goal. Consequently, our software was not designed to replace existing diagnostic approaches in the clinical setting but a system that can provide additional information for clinical decisions when sufficient quality RNA-Seq data is available. In the current scenario, it should be used for exploratory analysis and hypothesis generation. This concept is what we embodied in the final sentence of our original article: ‘Therefore, we set a precedent and an important advance in terms of personalized research’ (through the development of an analytical workflow) ‘oriented to a near future clinical practice within autoimmunity’. Daniel Toro-Domínguez, Jordi Martorell-Marugan, Manuel Martínez-Bueno, Raúl López-Domínguez, Elena Carnero-Montoro, Guillermo Barturen, Daniel Goldman, Michelle Petri, Pedro Carmona-Saez, Marta E. Alarcón-Riquelme |
Briefings Bioinform. | 9 |
| 2022 | Scoring personalized molecular portraits identify Systemic Lupus Erythematosus subtypes and predict individualized drug responses, symptomatology and disease progressionabstractOBJECTIVES: Systemic Lupus Erythematosus is a complex autoimmune disease that leads to significant worsening of quality of life and mortality. Flares appear unpredictably during the disease course and therapies used are often only partially effective. These challenges are mainly due to the molecular heterogeneity of the disease, and in this context, personalized medicine-based approaches offer major promise. With this work we intended to advance in that direction by developing MyPROSLE, an omic-based analytical workflow for measuring the molecular portrait of individual patients to support clinicians in their therapeutic decisions. METHODS: Immunological gene-modules were used to represent the transcriptome of the patients. A dysregulation score for each gene-module was calculated at the patient level based on averaged z-scores. Almost 6100 Lupus and 750 healthy samples were used to analyze the association among dysregulation scores, clinical manifestations, prognosis, flare and remission events and response to Tabalumab. Machine learning-based classification models were built to predict around 100 different clinical parameters based on personalized dysregulation scores. RESULTS: MyPROSLE allows to molecularly summarize patients in 206 gene-modules, clustered into nine main lupus signatures. The combination of these modules revealed highly differentiated pathological mechanisms. We found that the dysregulation of certain gene-modules is strongly associated with specific clinical manifestations, the occurrence of relapses or the presence of long-term remission and drug response. Therefore, MyPROSLE may be used to accurately predict these clinical outcomes. CONCLUSIONS: MyPROSLE (https://myprosle.genyo.es) allows molecular characterization of individual Lupus patients and it extracts key molecular information to support more precise therapeutic decisions. Daniel Toro-Domínguez, Jordi Martorell-Marugan, Manuel Martínez-Bueno, Raúl López-Domínguez, Elena Carnero-Montoro, Guillermo Barturen, Daniel Goldman, Michelle Petri, Pedro Carmona-Saez, Marta E. Alarcón-Riquelme |
Briefings Bioinform. | 9 |
| 2021 | A survey of gene expression meta-analysis: methods and applicationsabstractThe increasing use of high-throughput gene expression quantification technologies over the last two decades and the fact that most of the published studies are stored in public databases has triggered an explosion of studies available through public repositories. All this information offers an invaluable resource for reuse to generate new knowledge and scientific findings. In this context, great interest has been focused on meta-analysis methods to integrate and jointly analyze different gene expression datasets. In this work, we describe the main steps in the gene expression meta-analysis, from data preparation to the state-of-the art statistical methods. We also analyze the main types of applications and problems that can be approached in gene expression meta-analysis studies and provide a comparative overview of the available software and bioinformatics tools. Moreover, a practical guide for choosing the most appropriate method in each case is also provided. Daniel Toro-Domínguez, Juan Antonio Villatoro-García, Jordi Martorell-Marugan, Yolanda Román-Montoya, Marta E. Alarcón-Riquelme, Pedro Carmona-Saez |
Briefings Bioinform. | 6 |
| 2021 | DREIMT: a drug repositioning database and prioritization tool for immunomodulationabstractMOTIVATION: Drug immunomodulation modifies the response of the immune system and can be therapeutically exploited in pathologies such as cancer and autoimmune diseases. RESULTS: DREIMT is a new hypothesis-generation web tool, which performs drug prioritization analysis for immunomodulation. DREIMT provides significant immunomodulatory drugs targeting up to 70 immune cells subtypes through a curated database that integrates 4960 drug profiles and ∼2600 immune gene expression signatures. The tool also suggests potential immunomodulatory drugs targeting user-supplied gene expression signatures. Final output includes drug-signature association scores, FDRs and downloadable plots and results tables. AVAILABILITYAND IMPLEMENTATION: http://www.dreimt.org. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Kevin Troulé, Hugo López-Fernández, Santiago García-Martín, Miguel Reboiro-Jato, Carlos Carretero-Puche, Jordi Martorell-Marugan, Guillermo Martín-Serrano, Pedro Carmona-Saez, Daniel Glez-Peña, Fátima Al-Shahrour, Gonzalo Gómez-López |
Bioinform. | 8 |
| 2021 | A comprehensive database for integrated analysis of omics data in autoimmune diseasesabstractBACKGROUND: Autoimmune diseases are heterogeneous pathologies with difficult diagnosis and few therapeutic options. In the last decade, several omics studies have provided significant insights into the molecular mechanisms of these diseases. Nevertheless, data from different cohorts and pathologies are stored independently in public repositories and a unified resource is imperative to assist researchers in this field. RESULTS: Here, we present Autoimmune Diseases Explorer ( https://adex.genyo.es ), a database that integrates 82 curated transcriptomics and methylation studies covering 5609 samples for some of the most common autoimmune diseases. The database provides, in an easy-to-use environment, advanced data analysis and statistical methods for exploring omics datasets, including meta-analysis, differential expression or pathway analysis. CONCLUSIONS: This is the first omics database focused on autoimmune diseases. This resource incorporates homogeneously processed data to facilitate integrative analyses among studies. Jordi Martorell-Marugan, Raúl López-Domínguez, Adrián García-Moreno, Daniel Toro-Domínguez, Juan Antonio Villatoro-García, Guillermo Barturen, Adoración Martín-Gómez, Kevin Troulé, Gonzalo Gómez-López, Fátima Al-Shahrour, Víctor González-Rumayor, María Peña-Chilet, Joaquín Dopazo, Julio Saez-Rodriguez, Marta E. Alarcón-Riquelme, Pedro Carmona-Saez |
BMC Bioinform. | 16 |
| 2019 | mCSEA: detecting subtle differentially methylated regionsabstractMOTIVATION: The identification of differentially methylated regions (DMRs) among phenotypes is one of the main goals of epigenetic analysis. Although there are several methods developed to detect DMRs, most of them are focused on detecting relatively large differences in methylation levels and fail to detect moderate, but consistent, methylation changes that might be associated to complex disorders. RESULTS: We present mCSEA, an R package that implements a Gene Set Enrichment Analysis method to identify DMRs from Illumina450K and EPIC array data. It is especially useful for detecting subtle, but consistent, methylation differences in complex phenotypes. mCSEA also implements functions to integrate gene expression data and to detect genes with significant correlations among methylation and gene expression patterns. Using simulated datasets we show that mCSEA outperforms other tools in detecting DMRs. In addition, we applied mCSEA to a previously published dataset of sibling pairs discordant for intrauterine hyperglycemia exposure. We found several differentially methylated promoters in genes related to metabolic disorders like obesity and diabetes, demonstrating the potential of mCSEA to identify DMRs not detected by other methods. AVAILABILITY AND IMPLEMENTATION: mCSEA is freely available from the Bioconductor repository. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jordi Martorell-Marugan, Víctor González-Rumayor, Pedro Carmona-Saez |
Bioinform. | 3 |
| 2019 | ImaGEO: integrative gene expression meta-analysis from GEO databaseabstractSUMMARY: The Gene Expression Omnibus (GEO) database provides an invaluable resource of publicly available gene expression data that can be integrated and analyzed to derive new hypothesis and knowledge. In this context, gene expression meta-analysis (geMAs) is increasingly used in several fields to improve study reproducibility and discovering robust biomarkers. Nevertheless, integrating data is not straightforward without bioinformatics expertise. Here, we present ImaGEO, a web tool for geMAs that implements a complete and comprehensive meta-analysis workflow starting from GEO dataset identifiers. The application integrates GEO datasets, applies different meta-analysis techniques and provides functional analysis results in an easy-to-use environment. ImaGEO is a powerful and useful resource that allows researchers to integrate and perform meta-analysis of GEO datasets to lead robust findings for biomarker discovery studies. AVAILABILITY AND IMPLEMENTATION: ImaGEO is accessible at http://bioinfo.genyo.es/imageo/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Daniel Toro-Domínguez, Jordi Martorell-Marugan, Raúl López-Domínguez, Adrián García-Moreno, Víctor González-Rumayor, Marta E. Alarcón-Riquelme, Pedro Carmona-Saez |
Bioinform. | 7 |
| 2017 | Metagene projection characterizes GEN2.2 and CAL-1 as relevant human plasmacytoid dendritic cell modelsabstractMOTIVATION: Plasmacytoid dendritic cells (pDC) play a major role in the regulation of adaptive and innate immunity. Human pDC are difficult to isolate from peripheral blood and do not survive in culture making the study of their biology challenging. Recently, two leukemic counterparts of pDC, CAL-1 and GEN2.2, have been proposed as representative models of human pDC. Nevertheless, their relationship with pDC has been established only by means of particular functional and phenotypic similarities. With the aim of characterizing GEN2.2 and CAL-1 in the context of the main circulating immune cell populations we have performed microarray gene expression profiling of GEN2.2 and carried out an integrated analysis using publicly available gene expression datasets of CAL-1 and the main circulating primary leukocyte lineages. RESULTS: Our results show that GEN2.2 and CAL-1 share common gene expression programs with primary pDC, clustering apart from the rest of circulating hematopoietic lineages. We have also identified common differentially expressed genes that can be relevant in pDC biology. In addition, we have revealed the common and differential pathways activated in primary pDC and cell lines upon CpG stimulatio. AVAILABILITY AND IMPLEMENTATION: R code and data are available in the supplementary material. CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Pedro Carmona-Saez, Nieves Varela, María José Luque, Daniel Toro-Domínguez, Jordi Martorell-Marugan, Marta E. Alarcón-Riquelme, Concepción Marañón |
Bioinform. | 1 |
| 2017 | MetaGenyo: a web tool for meta-analysis of genetic association studiesabstractBACKGROUND: Genetic association studies (GAS) aims to evaluate the association between genetic variants and phenotypes. In the last few years, the number of this type of study has increased exponentially, but the results are not always reproducible due to experimental designs, low sample sizes and other methodological errors. In this field, meta-analysis techniques are becoming very popular tools to combine results across studies to increase statistical power and to resolve discrepancies in genetic association studies. A meta-analysis summarizes research findings, increases statistical power and enables the identification of genuine associations between genotypes and phenotypes. Meta-analysis techniques are increasingly used in GAS, but it is also increasing the amount of published meta-analysis containing different errors. Although there are several software packages that implement meta-analysis, none of them are specifically designed for genetic association studies and in most cases their use requires advanced programming or scripting expertise. RESULTS: We have developed MetaGenyo, a web tool for meta-analysis in GAS. MetaGenyo implements a complete and comprehensive workflow that can be executed in an easy-to-use environment without programming knowledge. MetaGenyo has been developed to guide users through the main steps of a GAS meta-analysis, covering Hardy-Weinberg test, statistical association for different genetic models, analysis of heterogeneity, testing for publication bias, subgroup analysis and robustness testing of the results. CONCLUSIONS: MetaGenyo is a useful tool to conduct comprehensive genetic association meta-analysis. The application is freely available at http://bioinfo.genyo.es/metagenyo/ . Jordi Martorell-Marugan, Daniel Toro-Domínguez, Marta E. Alarcón-Riquelme, Pedro Carmona-Saez |
BMC Bioinform. | 4 |
| 2011 | Finding closed frequent item sets by intersecting transactionsabstractMost known frequent item set mining algorithms work by enumerating candidate item sets and pruning infrequent candidates. An alternative method, which works by intersecting transactions, is much less researched. To the best of our knowledge, there are only two basic algorithms: a cumulative scheme, which is based on a repository with which new transactions are intersected, and the Carpenter algorithm, which enumerates and intersects candidate transaction sets. These approaches yield the set of so-called closed frequent item sets, since any such item set can be represented as the intersection of some subset of the given transactions. In this paper we describe a considerably improved implementation scheme of the cumulative approach, which relies on a prefix tree representation of the already found intersections. In addition, we present an improved way of implementing the Carpenter algorithm. We demonstrate that on specific data sets, which occur particularly often in the area of gene expression analysis, our implementations significantly outperform enumeration approaches to frequent item set mining. Christian Borgelt, Xiaoyuan Yang 0001, Rubén Nogales-Cadenas, Pedro Carmona-Saez, Alberto D. Pascual-Montano |
EDBT | 4 |
| 2008 | ChIPCodis: mining complex regulatory systems in yeast by concurrent enrichment analysis of chip-on-chip dataabstractMOTIVATION: Eukaryotic genes are often regulated by multiple transcription factors (TFs). Depending on the interactions among different TFs the expression of a gene can be tuned to respond to diverse environmental conditions. Chip-on-chip experiments provide a snapshot of which TF are in vivo bound to which genes in a particular condition, and have been applied to characterize the regulatory code of yeast under several experimental settings. ChIPCodis mines this data to provide new insights about how the expression of a particular group of genes is regulated. For a given list of yeast genes ChIPCodis determines which combinations of TFs are significantly over-represented in a series of environmental conditions. AVAILABILITY: http://chipcodis.dacya.ucm.es Federico Abascal, Pedro Carmona-Saez, José María Carazo, Alberto D. Pascual-Montano |
Bioinform. | 2 |
| 2006 | Integrated analysis of gene expression by association rules discoveryabstractBACKGROUND: Microarray technology is generating huge amounts of data about the expression level of thousands of genes, or even whole genomes, across different experimental conditions. To extract biological knowledge, and to fully understand such datasets, it is essential to include external biological information about genes and gene products to the analysis of expression data. However, most of the current approaches to analyze microarray datasets are mainly focused on the analysis of experimental data, and external biological information is incorporated as a posterior process. RESULTS: In this study we present a method for the integrative analysis of microarray data based on the Association Rules Discovery data mining technique. The approach integrates gene annotations and expression data to discover intrinsic associations among both data sources based on co-occurrence patterns. We applied the proposed methodology to the analysis of gene expression datasets in which genes were annotated with metabolic pathways, transcriptional regulators and Gene Ontology categories. Automatically extracted associations revealed significant relationships among these gene attributes and expression patterns, where many of them are clearly supported by recently reported work. CONCLUSION: The integration of external biological information and gene expression data can provide insights about the biological processes associated to gene expression programs. In this paper we show that the proposed methodology is able to integrate multiple gene annotations and expression data in the same analytic framework and extract meaningful associations among heterogeneous sources of data. An implementation of the method is included in the Engene software package. Pedro Carmona-Saez, Monica Chagoyen, Andrés Rodríguez Moreno, Oswaldo Trelles, José María Carazo, Alberto D. Pascual-Montano |
BMC Bioinform. | 1 |
| 2006 | Biclustering of gene expression data by non-smooth non-negative matrix factorizationabstractBACKGROUND: The extended use of microarray technologies has enabled the generation and accumulation of gene expression datasets that contain expression levels of thousands of genes across tens or hundreds of different experimental conditions. One of the major challenges in the analysis of such datasets is to discover local structures composed by sets of genes that show coherent expression patterns across subsets of experimental conditions. These patterns may provide clues about the main biological processes associated to different physiological states. RESULTS: In this work we present a methodology able to cluster genes and conditions highly related in sub-portions of the data. Our approach is based on a new data mining technique, Non-smooth Non-Negative Matrix Factorization (nsNMF), able to identify localized patterns in large datasets. We assessed the potential of this methodology analyzing several synthetic datasets as well as two large and heterogeneous sets of gene expression profiles. In all cases the method was able to identify localized features related to sets of genes that show consistent expression patterns across subsets of experimental conditions. The uncovered structures showed a clear biological meaning in terms of relationships among functional annotations of genes and the phenotypes or physiological states of the associated conditions. CONCLUSION: The proposed approach can be a useful tool to analyze large and heterogeneous gene expression datasets. The method is able to identify complex relationships among genes and conditions that are difficult to identify by standard clustering algorithms. Pedro Carmona-Saez, Roberto D. Pascual-Marqui, Francisco Tirado, José María Carazo, Alberto D. Pascual-Montano |
BMC Bioinform. | 1 |
| 2006 | A literature-based similarity metric for biological processesabstractBACKGROUND: Recent analyses in systems biology pursue the discovery of functional modules within the cell. Recognition of such modules requires the integrative analysis of genome-wide experimental data together with available functional schemes. In this line, methods to bridge the gap between the abstract definitions of cellular processes in current schemes and the interlinked nature of biological networks are required. RESULTS: This work explores the use of the scientific literature to establish potential relationships among cellular processes. To this end we have used a document based similarity method to compute pair-wise similarities of the biological processes described in the Gene Ontology (GO). The method has been applied to the biological processes annotated for the Saccharomyces cerevisiae genome. We compared our results with similarities obtained with two ontology-based metrics, as well as with gene product annotation relationships. We show that the literature-based metric conserves most direct ontological relationships, while reveals biologically sounded similarities that are not obtained using ontology-based metrics and/or genome annotation. CONCLUSION: The scientific literature is a valuable source of information from which to compute similarities among biological processes. The associations discovered by literature analysis are a valuable complement to those encoded in existing functional schemes, and those that arise by genome annotation. These similarities can be used to conveniently map the interlinked structure of cellular processes in a particular organism. Monica Chagoyen, Pedro Carmona-Saez, Concha Gil, José María Carazo, Alberto D. Pascual-Montano |
BMC Bioinform. | 2 |
| 2006 | Discovering semantic features in the literature: a foundation for building functional associationsabstractBACKGROUND: Experimental techniques such as DNA microarray, serial analysis of gene expression (SAGE) and mass spectrometry proteomics, among others, are generating large amounts of data related to genes and proteins at different levels. As in any other experimental approach, it is necessary to analyze these data in the context of previously known information about the biological entities under study. The literature is a particularly valuable source of information for experiment validation and interpretation. Therefore, the development of automated text mining tools to assist in such interpretation is one of the main challenges in current bioinformatics research. RESULTS: We present a method to create literature profiles for large sets of genes or proteins based on common semantic features extracted from a corpus of relevant documents. These profiles can be used to establish pair-wise similarities among genes, utilized in gene/protein classification or can be even combined with experimental measurements. Semantic features can be used by researchers to facilitate the understanding of the commonalities indicated by experimental results. Our approach is based on non-negative matrix factorization (NMF), a machine-learning algorithm for data analysis, capable of identifying local patterns that characterize a subset of the data. The literature is thus used to establish putative relationships among subsets of genes or proteins and to provide coherent justification for this clustering into subsets. We demonstrate the utility of the method by applying it to two independent and vastly different sets of genes. CONCLUSION: The presented method can create literature profiles from documents relevant to sets of genes. The representation of genes as additive linear combinations of semantic features allows for the exploration of functional associations as well as for clustering, suggesting a valuable methodology for the validation and interpretation of high-throughput experimental data. Monica Chagoyen, Pedro Carmona-Saez, Hagit Shatkay, José María Carazo, Alberto D. Pascual-Montano |
BMC Bioinform. | 2 |
| 2006 | bioNMF: a versatile tool for non-negative matrix factorization in biologyabstractBACKGROUND: In the Bioinformatics field, a great deal of interest has been given to Non-negative matrix factorization technique (NMF), due to its capability of providing new insights and relevant information about the complex latent relationships in experimental data sets. This method, and some of its variants, has been successfully applied to gene expression, sequence analysis, functional characterization of genes and text mining. Even if the interest on this technique by the bioinformatics community has been increased during the last few years, there are not many available simple standalone tools to specifically perform these types of data analysis in an integrated environment. RESULTS: In this work we propose a versatile and user-friendly tool that implements the NMF methodology in different analysis contexts to support some of the most important reported applications of this new methodology. This includes clustering and biclustering gene expression data, protein sequence analysis, text mining of biomedical literature and sample classification using gene expression. The tool, which is named bioNMF, also contains a user-friendly graphical interface to explore results in an interactive manner and facilitate in this way the exploratory data analysis process. CONCLUSION: bioNMF is a standalone versatile application which does not require any special installation or libraries. It can be used for most of the multiple applications proposed in the bioinformatics field or to support new research using this method. This tool is publicly available at http://www.dacya.ucm.es/apascual/bioNMF. Alberto D. Pascual-Montano, Pedro Carmona-Saez, Monica Chagoyen, Francisco Tirado, José María Carazo, Roberto D. Pascual-Marqui |
BMC Bioinform. | 2 |
| 2005 | Knowledge Discovery in the Identification of Differentially Expressed Genes in Tumoricidal Macrophage
Fazel Famili, Ziying Liu, Pedro Carmona-Saez, Alaka Mullick |
IDA | 3 |