Ana Conesa

dblp:60/4259 · DBLP profile ↗
← Back
23ranked-venue papers
2as first author
5since 2021 · last 2025
0000-0001-9597-311XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 23 · 2 first-author · 5 since 2021
YearPublicationVenuePosition
2025 MORE interpretable multi-omic regulatory networks to characterise phenotypes
abstract
Studying phenotype-specific regulatory mechanisms is crucial to understanding the molecular basis of diseases and other complex traits. However, existing approaches for constructing multi-omic regulatory networks MO-RN are scarce, and most cannot integrate diverse omics modalities, incorporate prior biological knowledge, or infer phenotype-specific networks. To address these challenges, we present MORE (Multi-Omics REgulation), a novel R package for inferring multi-modal regulatory networks. MORE is available at https://github.com/BiostatOmics/MORE and supports any number and type of omics layers while optionally incorporating prior regulatory knowledge. Leveraging advanced regression-based models and variable selection techniques, MORE identifies significant regulatory relationships. This tool also provides useful functionalities for the biological interpretation of MO-RN: network visualisations, differential regulatory networks, and functional enrichment analyses of key network features. We evaluated MORE on simulated multi-omic datasets and benchmarked it against state-of-the-art tools. Our tool consistently outperformed other methods regarding accuracy in identifying significant regulators, model goodness-of-fit, and computational efficiency. We further applied MORE to a multi-omic ovarian cancer dataset to uncover tumour subtype-specific regulatory mechanisms associated with distinct survival outcomes. This analysis revealed differential regulatory patterns to understand the molecular basis of each subtype. By addressing the limitations of methods for multi-omic network inference, MORE represents a valuable resource for studying regulatory systems. Its ability to construct phenotype-specific regulatory networks with high accuracy and interpretability positions it as a useful resource for researchers seeking to unravel the complexities of molecular interactions and regulatory mechanisms across diverse biological contexts.
Maider Aguerralde-Martin, Mónica Clemente-Císcar, Ana Conesa, Sonia Tarazona
Briefings Bioinform.3
2025 MOSim: bulk and single-cell multilayer regulatory network simulator
abstract
As multi-omics sequencing technologies advance, the need for simulation tools capable of generating realistic and diverse (bulk and single-cell) multi-omics datasets for method testing and benchmarking becomes increasingly important. We present MOSim, an R package that simulates both bulk (via mosim function) and single-cell (via sc_mosim function) multi-omics data. The mosim function generates bulk transcriptomics data (RNA-seq) and additional regulatory omics layers (ATAC-seq, miRNA-seq, ChIP-seq, Methyl-seq, and transcription factors), while sc_mosim simulates single-cell transcriptomics data (scRNA-seq) with scATAC-seq and transcription factors as regulatory layers. The tool supports various experimental designs, including simulation of gene co-expression patterns, biological replicates, and differential expression between conditions. MOSim enables users to generate quantification matrices for each simulated omics data type, capturing the heterogeneity and complexity of bulk and single-cell multi-omics datasets. Furthermore, MOSim provides differentially abundant features within each omics layer and elucidates the active regulatory relationships between regulatory omics and gene expression data at both bulk and single-cell levels. By leveraging MOSim, researchers will be able to generate realistic and customizable bulk and single-cell multi-omics datasets to benchmark and validate analytical methods specifically designed for the integrative analysis of diverse regulatory omics data.
Carolina Monzó, Maider Aguerralde-Martin, Carlos Martínez-Mira, Angeles Arzalluz-Luque, Ana Conesa, Sonia Tarazona
Briefings Bioinform.5
2024 scMaSigPro: differential expression analysis along single-cell trajectories
abstract
MOTIVATION: Understanding the dynamics of gene expression across different cellular states is crucial for discerning the mechanisms underneath cellular differentiation. Genes that exhibit variation in mean expression as a function of Pseudotime and between branching trajectories are expected to govern cell fate decisions. We introduce scMaSigPro, a method for the identification of differential gene expression patterns along Pseudotime and branching paths simultaneously. RESULTS: We assessed the performance of scMaSigPro using synthetic and public datasets. Our evaluation shows that scMaSigPro outperforms existing methods in controlling the False Positive Rate and is computationally efficient. AVAILABILITY AND IMPLEMENTATION: scMaSigPro is available as a free R package (version 4.0 or higher) under the GPL(≥2) license on GitHub at 'github.com/BioBam/scMaSigPro' and archived with version 0.03 on Zenodo at 'zenodo.org/records/12568922'.
Priyansh Srivastava, Marta Benegas Coll, Stefan Götz 0003, María José Nueda, Ana Conesa
Bioinform.5
2022 ECCB2022: the 21st European Conference on Computational Biology
abstract
This volume of Bioinformatics includes the proceedings papers of the 21st European Conference in Computational Biology (ECCB), an annual international conference for research in computational biology and bioinformatics. The conference is being held jointly with the Intelligent Systems for Molecular Biology (ISMB) Conference in the odd-numbered years and independently in the even-numbered years. This year, the 21st ECCB (ECCB2022) took place in a hybrid format in Sitges (Barcelona, Spain) from September 12 to September 21, 2022, under the motto of Planetary Health and Biodiversity. The positive experience of running ECCB2020 in a virtual format due to the COVID-19 pandemic and the still existing health concerns and travel restrictions have made ECCB2022 the first edition of the conference held in a hybrid format. Information on ECCB2022 and the conference’s satellite meetings can be found at http://eccb2022.org. With more than 850 in-person participants from academia and industry, ECCB is the leading European Conference on Computational Biology and Bioinformatics and the second largest internationally, next to ISMB. The work presented in ECCB is related to all domains of the field of computational biology and bioinformatics, ranging from molecular to systems biology level. The proceedings papers present new computational methodologies and tools for addressing challenging problems in the field, with a strong presence of predictive modelling approaches and multi-omics data integration (Fig. 1). Recovering the traditional in-person format allowed us to open the calls for Highlight Papers, Applications and Posters that were not included in the ECCB2020 programme. Word cloud representing the most frequent keywords accompanying the submissions received. The keywords ‘machine learning’, ‘deep learning’, ‘cancer’, ‘gene expression’ and ‘data integration’ were the most frequently used keywords ECCB2022 leveraged experiences from previous editions, including the last two ones, held online despite being initially programmed to be hosted in-person in Lyon, France (Dessimoz and Przytycka, 2021, ISMB/ECCB 2021) and Sitges, Barcelona, Spain (Capella-Gutierrez et al. 2020, ECCB2020). This was the second time that ECCB was hosted in Spain after the successful experience of the fourth edition back in 2005 in Madrid, Spain (Guigo et al., 2005, ECCB 2005). ECCB2022 was held under the auspices of the Spanish National Bioinformatics Institute (INB; http://inb-elixir.es), founded in 2003, which has been the node of ELIXIR in Spain since 2015. The Life Sciences Department of the Barcelona Supercomputing Center (BSC) was the main organiser of 2022’s conference. BSC is the national supercomputing facility in Spain with more than 800 R&D experts and professionals distributed across four domains: Engineering, Computer, Earth and Life Sciences. The Life Sciences Department includes more than 160 professionals, ranging from MSc and PhD students to postdoctoral researchers and research engineers. The department’s mission is to understand living organisms through theoretical and computational methods, from genomics, structural bioinformatics, molecular and cellular modelling and biomedical research infrastructures, all of them in the framework of European projects and infrastructures (ELIXIR in particular). ECCB2022 received manuscripts for Proceedings talks from institutions across 39 countries. The submissions were organised in five themes: (1) Data (organisation, management, categorization, integration, analysis of data, knowledge discovery), (2) Genes (expression, function, regulation, transcription, translation, geno-phenotype), (3) Genomes (sequence analysis, alignment, evolution, phylogeny, genetics, epigenetics, 3D conformation) (4) Proteins (structure, function, alterations, assemblies, interactions, design, proteomics) and (5) Systems (systems biology, pathways, molecular networks, dynamics, signalling, multi-scale modelling). All proceedings submissions were subjected to a rigorous peer-review process (two to four reviews per paper), organised by the members of the ECCB2022 Programme Committee of the corresponding theme. The Programme Committee, chaired by Ana Conesa and the respective area chair: Josep Lluís Gelpí, Artemis Hatzigeorgiou, Toni Gabaldón, Mark Wass and Patrick Aloy and one to three co-chairs per theme (Heidi Peterson, Rory Johnson, Irene Papatheodorou, Sonia Tarazona Campos, Stephane Rombauts, Shilpa Garg, Franca Fraternali, Jolanda van Leeuwen and Anaïs Baudot) and a Programme Committee with 185 reviewers (see the Supplementary File ‘ECCB2022 Committees’). The review process was handled by the EasyChair system (www.easychair.org), focusing on the impact and reproducibility of the submitted research, as well as its suitability for the ECCB2022 audience. Once the review was completed, the committee chairs and co-chairs selected 25 papers to be included in the ECCB2022 proceedings with an acceptance rate of 17.4% across the different areas. All Proceedings papers and supplementary files are freely available in the electronic form of Oxford University Press journal Bioinformatics as a special issue of September 2022 (https://academic.oup.com/bioinformatics/issue/38/Supplement_2). The ECCB2022 call for Highlight Papers invited in-person presentations of recently published works, i.e. up to 1 year since its publication. The track received a total of 79 submissions that were ranked according to relevance and impact on biology and/or medicine, and their suitability to be presented in front of a large and heterogeneous audience. The existence of complementary work additional to the published paper was also considered positively. Finally, 30% of the submissions were presented at the conference. The conference programme included an Applications track in which participants were invited to present their works on software, services and infrastructures developed in a broader context of academia and industry. The proposals were evaluated by a committee of six members chaired by Javier De Las Rivas and co-chaired by Sara Aibar. A total of 15 out of 29 submissions were selected to be presented at the conference after ranking them based on their impact and potential in bridging different researcher and professional profiles together. ECCB2022 hosted two poster sessions where researchers introduced and discussed their work in person. The Posters track received 567 contributions from authors affiliated with 54 different countries. The submissions were reviewed by a committee of 28 members, chaired by R. Gonzalo Parra and co-chaired by Handan Melike Dönertaş and Alexander Miguel Monzon. Overall the ECCB2022 tracks received the participation of authors from 66 countries. The countries with the higher number of accepted submissions in ECCB2022 tracks (Fig. 2) were Germany, Spain, the UK and France, jointly contributing to half of the overall accepted submissions. Authors from Belgium, Switzerland, Italy and Poland followed the ranking adding up a comparable number of accepted proposals (∼25–30 submissions) and the USA and India were the non-European countries with the higher number of authors with accepted submissions. Proportion of countries’ affiliation from the authors with accepted submissions at ECCB2022. The Programme Committee accepted a total number of 626 submissions, most of which were authored by researchers affiliated with institutions from Germany, Spain, the UK and France The programme incorporated a new track aligned with the motto of the conference: Climate Crisis and Health. The track consisted of two sessions with talks from invited speakers on air quality, heat stress, infectious diseases and biodiversity. Taking advantage of the hybrid format, ECCB2022 hosted a dedicated track to foster the interactions between different geographically distributed communities. During this edition, two sessions were dedicated to fostering interactions between ELIXIR, the pan-European infrastructure for data in Life Sciences and Latin America Bioinformatics societies. The main topic revolved around research data management and how to facilitate the knowledge exchange between communities. These joint sessions just represented the starting point for further collaborations. The programme also included four ELIXIR sessions in which the speakers showcased the latest outputs and services on data integration and data management strategies developed in this European research infrastructure context. Additionally, the Quest for Orthologs consortium organised a satellite meeting on the scope of the conference, which included a day and a half meeting and an open workshop to discuss new approaches for improving orthology predictions. Six distinguished keynote speakers presented their work: Prof. Cesar Hidalgo from the Universities of Toulouse, Manchester and Harvard, Prof. Raul Rabadan from Columbia University, Dr Ana Freitas from the Institute for Systems and Computer Engineering, Technology and Science—Technical University of Lisbon, Dr Maria Rodriguez-Martínez from IBM Research Europe, Dr Graciela Gonzalez-Hernandez from the University of Pennsylvania and Prof. Mar Albà from the Catalan Institution for Research and Advanced Studies and Hospital del Mar Medical Research Institute. The keynotes covered broad and diverse research areas, including evolutionary genomics, personalized medicine and new methods in artificial intelligence and natural language processing for computational modelling of biological data. Following the pilot experience from ECCB2020, this edition also had a programme of virtual workshops and tutorials under the umbrella of New Trends in Bioinformatics by ECCB. This format allowed the spread of 3-h virtual sessions during the week before ECCB2022, enabling the participation of attendees in multiple sessions and making it easier for the global community to join them. For the 2022 edition, the New Trends in Bioinformatics by ECCB combined a total of 9 virtual sessions and 10 in-person sessions. The virtual events were programmed in two blocks, the first from 13.30 to 16.30 and the second from 17.00 to 20.00 following the Central European Summer Time. All sessions were recorded so participants could visit them again through the ECCB2022 virtual platform and, starting on January 1, 2023, through the ISCB.tv channel. Among the 19 scheduled events were 11 workshops, 7 tutorials and 1 session led by ELIXIR. The workshops aimed to provide participants with the opportunity to discuss different perspectives on the cutting edge of a selected research area through presenting technical issues, exchanging research ideas and sharing practical experiences. The 11 workshops were selected out of a total of 17 applications. In contrast, the purpose of the tutorials program is to provide participants with specialized lectures and hands-on training on the most important and emerging topics in bioinformatics and computational biology research. The tutorials offered topics ranging from early and basic steps of computational analysis, e.g. machine learning, cellular processes modelling or data management, to advanced computational skills in important established topics, e.g. functional analyses of single-cell transcriptomics data. A total of 7 tutorials were selected as part of the New Trends in Bioinformatics by ECCB out of 10 applications. The programme of workshops and tutorials was round out by a workshop organised by ELIXIR on practical approaches for applying FAIR guiding principles in data reuse. The specific sessions for the New Trends in Bioinformatics by ECCB were as follows. Tutorials are denoted by a T, Workshops are denoted by a W and ELIXIR workshop contains an E, to form the events ID. NTB-T01 Computational challenges in phospho-proteomics and systems biology of cellular signalling, organised by Filipa Blasco Tavares Pereira-Lopes (Case Western Reserve University, USA), Marzieh Ayati (University of Texas Rio Grande Valley, USA), Serhan Yilmaz (Case Western Reserve University, USA), Daniela Schlatzer (Case Western Reserve University, USA), Mehmet Koyutürk (Case Western Reserve University, USA) and Mark Chance (Case Western Reserve University, USA). NTB-T02 To rarefy or not to rarefy microbiome data? What are the alpha diversity metrics?, organised by Violeta Larios-Serrato (Winter Genomics, México), Maira Nayeli Luis-Vargas (Winter Genomics, México), Karla Ruiz (Winter Genomics, México) and Kenya Contreras (Winter Genomics, México). NTB-T03 Deep learning for biological sequence data: from convolutional neural networks to transformers, organised by Panagiotis Alexiou (Masaryk University, Czech Republic), Petr Simecek (Masaryk University, Czech Republic), David Cechak (Masaryk University, Czech Republic) and Vlastimil Martinek (Masaryk University, Czech Republic). NTB-T04 Functional analysis of single-cell transcriptomics, organised by Pau Badia i Mompel (Heidelberg University, Germany), Robin Browaeys (VIB Center for Inflammation Research, Belgium) and Daniel Dimitrov (Heidelberg University, Germany). NTB-T05 Guidelines for the assessment and analysis of lrRNA-seq data for transcript identification and quantification (LRGASP challenge), organised by Ana Conesa (Institute for Integrative Systems Biology, Spain), Fairlie Reese (University of California at Irvine, USA), Dennis Mulligan (University of California at Santa Cruz, USA), Ying Chen (Genome Institute of Singapore, Singapore), Ralf Herwig (Max Planck Institute for Molecular Genetics, Germany) and Sílvia Carbonell-Sala (Center for Genomic Regulation, Spain). NTB-T06 Boost your data management planning, organised by Helena Schnitzer (ELIXIR Germany and Forschungszentrum Jülich, Germany) and Daniel Wibberg (ELIXIR Germany and Forschungszentrum Jülich, Germany). NTB-T07 Software containerization in bioinformatics: how to make reproducible, portable and reusable bioinformatics software and pipelines, organised by Giacomo Baruzzo (University of Padova, Italy), Barbara Di Camillo (University of Padova, Italy), Marco Cappellato (University of Padova, Italy), Giulia Cesaro (University of Padova, Italy), Mikele Milia (University of Padova, Italy). NTB-W01 Machine learning good practices—DOME recommendations for better machine learning in computational biology, organised by Jennifer Harrow (ELIXIR, UK), Fotis E. Psomopoulos (CERTH, Greece), Silvio Tosatto (University of Padova, Italy) and Leyla Jael García-Castro (ZB MED Information Centre for Life Sciences, Germany). NTB-W02 FAIRification of multi-omics metadata, organised by Gary Saunders (EATRIS-ERIC Data Director, The Netherlands), Emanuela Oldoni (EATRIS-ERIC Scientific Programme Manager, The Netherlands) and Anna Niehues (Bioinformatician at Radboud University Medical Center, The Netherlands). NTB-W03 Simulating cellular behaviours: advancing HPC-enabled computational biology, organised by Arnau Montagud (BSC, Spain), Marta Lloret-Llinares (EMBL-EBI, UK), Renata Giménez (BSC, Spain) and Mariola Tàrrega-Moltó (BSC, Spain). NTB-W04 Spatial transcriptomics and cell–cell communication modelling: new opportunities to study the cellular dynamics of biological systems, organised by Giacomo Baruzzo (University of Padova, Padova, Italy), Enrica Calura (University of Padova, Italy), Davide Risso (University of Padova, Italy), Chiara Romualdi (University of Padova, Italy), Gabriele Sales (University of Padova, Italy), Luz García-Alonso (Wellcome Sanger Institute, UK), Roser Vento-Tormo (Wellcome Sanger Institute, UK), Julio Saez-Rodriguez (Heidelberg University, Germany) and Yvan Saeys (VIB, Belgium). NTB-W05 Building high-quality reference genome assemblies of eukaryotes, organised by Nadège Guiglielmoni (University of Cologne, Germany), Joanna Malukiewicz (Deutsches Primatenzentrum, Germany; University of Sao Paulo, Brazil), Lino Ometto (University of Pavia, Italy) and Robert (University of Lausanne and Swiss Institute of Bioinformatics, Switzerland). NTB-W06 Tools and techniques to make sensitive data discoverable (use-cases, hands-on session of Beacon implementation), organised by Babita Singh (European Genome-phenome Archive, Spain) and Lauren Fromont (European Genome-phenome Archive, Spain). NTB-W07 Sex and gender dimension in biomedical research, organised by Àtia Cortés (BSC, Spain), Davide Cirillo (BSC, Spain) and Fatemeh Baghdadi (BSC, Spain). NTB-W08 Integration of large-scale data for reference genome development in biodiversity, organised by Shilpa Garg (University of Copenhagen, Denmark) and Josiah Kuja (University of Copenhagen, Denmark). NTB-W09 Annual European Bioinformatics Core Community (AEBC2) workshop 2022, organised by Camille Stephan-Otto Attolini (IRB Barcelona, Spain), Dieter Beule (Berlin Institute of Health, Germany), Sven Rahmann (Saarland University, Germany), Sven Nahnsen (University of Tübingen, Germany) and Daniel J. Stekhoven (ETH Zurich, Germany). NTB-W10 Computational modelling of immunological mechanisms: from statistical approaches to interpretable machine learning, organised by María Rodríguez-Martínez (IBM—Zurich Research Laboratory, Switzerland), Anna Niarakis (University of Évry Val d'Essonne and University of Paris-Saclay) and Matteo Barberis (University of Surrey, UK). NTB-W11 Novel challenges in the quest for orthologs, organised by Ingo Ebersberger (Goethe University Frankfurt, Germany), Michael Hiller (Senckenberg Society for Nature Research, Germany), Thomas Rattei (University of Vienna, Austria), Paul D. Thomas (University of Southern California, USA) and Sofia Kirke Forslund (Max-Delbrück-Centre for Molecular Medicine, Germany). NTB-EW01 ELIXIR | FAIR applied: a practical FAIRification guide for life science data from FAIRplus, organised by Tony Burdett (EMBL-EBI, UK), Ibrahim Emam (Imperial College London, UK), Nick Juty (University of Manchester, UK), (University of UK), (University of and (EMBL-EBI, UK). The New Trends in Bioinformatics by ECCB in format in with the European The was organised by the European of the for and by researchers to a to exchange research ideas and The was co-chaired by from Spain, and Maria from are The work and of the and participants of the have led to a and successful the last 10 making the an part of ECCB. Following previous of a number of by were to students and postdoctoral with to members affiliated in or countries to their at the conference and the New Trends in Bioinformatics by ECCB. A total of were received from from countries. a were by the ECCB2022 Committee and The ECCB2022 Committee is to gender for better and in this was considered in all the steps of the the of review out of 10 chairs and 9 out of 15 co-chairs were the and the made an important to the review by as as and a for the keynotes represented four out of the six distinguished The programme included jointly organised with the to further to in the field of bioinformatics. the of the were after in the field of bioinformatics and a was to to further to their contributions in the ECCB2020 established a of which was leveraged and by ECCB2022. The ECCB2022 of a and to and had to to in had been a of the of The ECCB2022 Committee to through their work to the conference’s The co-chairs and reviewers have been the of the conference and workshops, manuscripts and talks for and making this conference a are to them. ECCB2022 are to the ECCB2022 Committee and the ECCB2022 Committee for their and during the of the of ECCB2022. are also to the to for their and for contributing to the of ECCB2022 at the international and for This conference not be the of and ELIXIR the conference as a was the and the poster also further Sciences and The and Spanish Supercomputing were of ECCB2022. to all for their and make ECCB2022 an and conference. A number of to the of ECCB2022. their call of to make this conference and as part of Conference were for the members of the and the Life Sciences Department at BSC this conference: and the the Finally, the of a conference is its the conference for its ECCB2022 all of for us at the first hybrid edition of ECCB. the and the of previous of ECCB with the of ECCB2022 to the next conference in Supplementary Supplementary data are available at Bioinformatics This paper was published as part of a special issue by ECCB2022. of
Salvador Capella-Gutiérrez, Eva Alloza, Laura Rubinat-Ripoll, Ana Conesa, Alfonso Valencia
Bioinform.4
2022 MultiBaC: an R package to remove batch effects in multi-omic experiments
abstract
MOTIVATION: Batch effects in omics datasets are usually a source of technical noise that masks the biological signal and hampers data analysis. Batch effect removal has been widely addressed for individual omics technologies. However, multi-omic datasets may combine data obtained in different batches where omics type and batch are often confounded. Moreover, systematic biases may be introduced without notice during data acquisition, which creates a hidden batch effect. Current methods fail to address batch effect correction in these cases. RESULTS: In this article, we introduce the MultiBaC R package, a tool for batch effect removal in multi-omics and hidden batch effect scenarios. The package includes a diversity of graphical outputs for model validation and assessment of the batch effect correction. AVAILABILITY AND IMPLEMENTATION: MultiBaC package is available on Bioconductor (https://www.bioconductor.org/packages/release/bioc/html/MultiBaC.html) and GitHub (https://github.com/ConesaLab/MultiBaC.git). The data underlying this article are available in Gene Expression Omnibus repository (accession numbers GSE11521, GSE1002, GSE56622 and GSE43747). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Manuel Ugidos, María José Nueda, José Manuel Prats-Montalbán, Alberto Ferrer 0001, Ana Conesa, Sonia Tarazona
Bioinform.5
2020 Padhoc: a computational pipeline for pathway reconstruction on the fly
abstract
MOTIVATION: Molecular pathway databases represent cellular processes in a structured and standardized way. These databases support the community-wide utilization of pathway information in biological research and the computational analysis of high-throughput biochemical data. Although pathway databases are critical in genomics research, the fast progress of biomedical sciences prevents databases from staying up-to-date. Moreover, the compartmentalization of cellular reactions into defined pathways reflects arbitrary choices that might not always be aligned with the needs of the researcher. Today, no tool exists that allow the easy creation of user-defined pathway representations. RESULTS: Here we present Padhoc, a pipeline for pathway ad hoc reconstruction. Based on a set of user-provided keywords, Padhoc combines natural language processing, database knowledge extraction, orthology search and powerful graph algorithms to create navigable pathways tailored to the user's needs. We validate Padhoc with a set of well-established Escherichia coli pathways and demonstrate usability to create not-yet-available pathways in model (human) and non-model (sweet orange) organisms. AVAILABILITY AND IMPLEMENTATION: Padhoc is freely available at https://github.com/ConesaLab/padhoc. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Salvador Casaní-Galdón, Cécile Pereira, Ana Conesa
Bioinform.3
2020 MirCure: a tool for quality control, filter and curation of microRNAs of animals and plants
abstract
MOTIVATION: microRNAs (miRNAs) are essential components of gene expression regulation at the post-transcriptional level. miRNAs have a well-defined molecular structure and this has facilitated the development of computational and high-throughput approaches to predict miRNAs genes. However, due to their short size, miRNAs have often been incorrectly annotated in both plants and animals. Consequently, published miRNA annotations and miRNA databases are enriched for false miRNAs, jeopardizing their utility as molecular information resources. To address this problem, we developed MirCure, a new software for quality control, filtering and curation of miRNA candidates. MirCure is an easy-to-use tool with a graphical interface that allows both scoring of miRNA reliability and browsing of supporting evidence by manual curators. RESULTS: Given a list of miRNA candidates, MirCure evaluates a number of miRNA-specific features based on gene expression, biogenesis and conservation data, and generates a score that can be used to discard poorly supported miRNA annotations. MirCure can also curate and adjust the annotation of the 5p and 3p arms based on user-provided small RNA-seq data. We evaluated MirCure on a set of manually curated animal and plant miRNAs and demonstrated great accuracy. Moreover, we show that MirCure can be used to revisit previous bona fide miRNAs annotations to improve miRNA databases. AVAILABILITY AND IMPLEMENTATION: The MirCure software and all the additional scripts used in this project are publicly available at https://github.com/ConesaLab/MirCure. A Docker image of MirCure is available at https://hub.docker.com/r/conesalab/mircure. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Guillem Ylla, Ana Conesa
Bioinform.3
2019 A benchmarking of workflows for detecting differential splicing and differential expression at isoform level in human RNA-seq studies
abstract
Over the last few years, RNA-seq has been used to study alterations in alternative splicing related to several diseases. Bioinformatics workflows used to perform these studies can be divided into two groups, those finding changes in the absolute isoform expression and those studying differential splicing. Many computational methods for transcriptomics analysis have been developed, evaluated and compared; however, there are not enough reports of systematic and objective assessment of processing pipelines as a whole. Moreover, comparative studies have been performed considering separately the changes in absolute or relative isoform expression levels. Consequently, no consensus exists about the best practices and appropriate workflows to analyse alternative and differential splicing. To assist the adequate pipeline choice, we present here a benchmarking of nine commonly used workflows to detect differential isoform expression and splicing. We evaluated the workflows performance over different experimental scenarios where changes in absolute and relative isoform expression occurred simultaneously. In addition, the effect of the number of isoforms per gene, and the magnitude of the expression change over pipeline performances were also evaluated. Our results suggest that workflow performance is influenced by the number of replicates per condition and the conditions heterogeneity. In general, workflows based on DESeq2, DEXSeq, Limma and NOISeq performed well over a wide range of transcriptomics experiments. In particular, we suggest the use of workflows based on Limma when high precision is required, and DESeq2 and DEXseq pipelines to prioritize sensitivity. When several replicates per condition are available, NOISeq and Limma pipelines are indicated.
Gabriela Alejandra Merino, Ana Conesa, Elmer Andrés Fernández
Briefings Bioinform.2
2019 Building gene regulatory networks from scATAC-seq and scRNA-seq using Linked Self Organizing Maps
abstract
Rapid advances in single-cell assays have outpaced methods for analysis of those data types. Different single-cell assays show extensive variation in sensitivity and signal to noise levels. In particular, scATAC-seq generates extremely sparse and noisy datasets. Existing methods developed to analyze this data require cells amenable to pseudo-time analysis or require datasets with drastically different cell-types. We describe a novel approach using self-organizing maps (SOM) to link scATAC-seq regions with scRNA-seq genes that overcomes these challenges and can generate draft regulatory networks. Our SOMatic package generates chromatin and gene expression SOMs separately and combines them using a linking function. We applied SOMatic on a mouse pre-B cell differentiation time-course using controlled Ikaros over-expression to recover gene ontology enrichments, identify motifs in genomic regions showing similar single-cell profiles, and generate a gene regulatory network that both recovers known interactions and predicts new Ikaros targets during the differentiation process. The ability of linked SOMs to detect emergent properties from multiple types of highly-dimensional genomic data with very different signal properties opens new avenues for integrative analysis of heterogeneous data.
Camden Jansen, Ricardo N. Ramirez, Nicole C. El-Ali, David Gomez-Cabrero, Jesper Tegnér, Matthias Merkenschlager, Ana Conesa
PLoS Comput. Biol.7
2018 Identification and visualization of differential isoform expression in RNA-seq time series
abstract
Motivation: As sequencing technologies improve their capacity to detect distinct transcripts of the same gene and to address complex experimental designs such as longitudinal studies, there is a need to develop statistical methods for the analysis of isoform expression changes in time series data. Results: Iso-maSigPro is a new functionality of the R package maSigPro for transcriptomics time series data analysis. Iso-maSigPro identifies genes with a differential isoform usage across time. The package also includes new clustering and visualization functions that allow grouping of genes with similar expression patterns at the isoform level, as well as those genes with a shift in major expressed isoform. Availability and implementation: The package is freely available under the LGPL license from the Bioconductor web site. Contact: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
María José Nueda, Jordi Martorell-Marugan, Cristina Martí, Sonia Tarazona, Ana Conesa
Bioinform.5
2018 GRAM-CNN: a deep learning approach with local context for named entity recognition in biomedical text
abstract
Motivation: Best performing named entity recognition (NER) methods for biomedical literature are based on hand-crafted features or task-specific rules, which are costly to produce and difficult to generalize to other corpora. End-to-end neural networks achieve state-of-the-art performance without hand-crafted features and task-specific knowledge in non-biomedical NER tasks. However, in the biomedical domain, using the same architecture does not yield competitive performance compared with conventional machine learning models. Results: We propose a novel end-to-end deep learning approach for biomedical NER tasks that leverages the local contexts based on n-gram character and word embeddings via Convolutional Neural Network (CNN). We call this approach GRAM-CNN. To automatically label a word, this method uses the local information around a word. Therefore, the GRAM-CNN method does not require any specific knowledge or feature engineering and can be theoretically applied to a wide range of existing NER problems. The GRAM-CNN approach was evaluated on three well-known biomedical datasets containing different BioNER entities. It obtained an F1-score of 87.26% on the Biocreative II dataset, 87.26% on the NCBI dataset and 72.57% on the JNLPBA dataset. Those results put GRAM-CNN in the lead of the biological NER methods. To the best of our knowledge, we are the first to apply CNN based structures to BioNER problems. Availability and implementation: The GRAM-CNN source code, datasets and pre-trained model are available online at: https://github.com/valdersoul/GRAM-CNN. Contact: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Qile Zhu, Xiaolin Li 0001, Ana Conesa, Cécile Pereira
Bioinform.3
2017 The eBioKit, a stand-alone educational platform for bioinformatics
abstract
Bioinformatics skills have become essential for many research areas; however, the availability of qualified researchers is usually lower than the demand and training to increase the number of able bioinformaticians is an important task for the bioinformatics community. When conducting training or hands-on tutorials, the lack of control over the analysis tools and repositories often results in undesirable situations during training, as unavailable online tools or version conflicts may delay, complicate, or even prevent the successful completion of a training event. The eBioKit is a stand-alone educational platform that hosts numerous tools and databases for bioinformatics research and allows training to take place in a controlled environment. A key advantage of the eBioKit over other existing teaching solutions is that all the required software and databases are locally installed on the system, significantly reducing the dependence on the internet. Furthermore, the architecture of the eBioKit has demonstrated itself to be an excellent balance between portability and performance, not only making the eBioKit an exceptional educational tool but also providing small research groups with a platform to incorporate bioinformatics analysis in their research. As a result, the eBioKit has formed an integral part of training and research performed by a wide variety of universities and organizations such as the Pan African Bioinformatics Network (H3ABioNet) as part of the initiative Human Heredity and Health in Africa (H3Africa), the Southern Africa Network for Biosciences (SAnBio) initiative, the Biosciences eastern and central Africa (BecA) hub, and the International Glossina Genome Initiative.
Rafael Hernández-de-Diego, Etienne Pierre de Villiers, Tomas Klingström, Hadrien Gourlé, Ana Conesa, Erik Bongcam-Rudloff
PLoS Comput. Biol.5
2016 Qualimap 2: advanced multi-sample quality control for high-throughput sequencing data
abstract
MOTIVATION: Detection of random errors and systematic biases is a crucial step of a robust pipeline for processing high-throughput sequencing (HTS) data. Bioinformatics software tools capable of performing this task are available, either for general analysis of HTS data or targeted to a specific sequencing technology. However, most of the existing QC instruments only allow processing of one sample at a time. RESULTS: Qualimap 2 represents a next step in the QC analysis of HTS data. Along with comprehensive single-sample analysis of alignment data, it includes new modes that allow simultaneous processing and comparison of multiple samples. As with the first version, the new features are available via both graphical and command line interface. Additionally, it includes a large number of improvements proposed by the user community. AVAILABILITY AND IMPLEMENTATION: The implementation of the software along with documentation is freely available at http://www.qualimap.org. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Konstantin Okonechnikov, Ana Conesa, Fernando García-Alcalde
Bioinform.2
2016 RGmatch: matching genomic regions to proximal genes in omics data integration
abstract
BACKGROUND: The integrative analysis of multiple genomics data often requires that genome coordinates-based signals have to be associated with proximal genes. The relative location of a genomic region with respect to the gene (gene area) is important for functional data interpretation; hence algorithms that match regions to genes should be able to deliver insight into this information. RESULTS: In this work we review the tools that are publicly available for making region-to-gene associations. We also present a novel method, RGmatch, a flexible and easy-to-use Python tool that computes associations either at the gene, transcript, or exon level, applying a set of rules to annotate each region-gene association with the region location within the gene. RGmatch can be applied to any organism as long as genome annotation is available. Furthermore, we qualitatively and quantitatively compare RGmatch to other tools. CONCLUSIONS: RGmatch simplifies the association of a genomic region with its closest gene. At the same time, it is a powerful tool because the rules used to annotate these associations are very easy to modify according to the researcher's specific interests. Some important differences between RGmatch and other similar tools already in existence are RGmatch's flexibility, its wide range of user options, compatibility with any annotatable organism, and its comprehensive and user-friendly output.
Pedro Furió-Tarí, Ana Conesa, Sonia Tarazona
BMC Bioinform.2
2016 Separating common from distinctive variation
abstract
BACKGROUND: Joint and individual variation explained (JIVE), distinct and common simultaneous component analysis (DISCO) and O2-PLS, a two-block (X-Y) latent variable regression method with an integral OSC filter can all be used for the integrated analysis of multiple data sets and decompose them in three terms: a low(er)-rank approximation capturing common variation across data sets, low(er)-rank approximations for structured variation distinctive for each data set, and residual noise. In this paper these three methods are compared with respect to their mathematical properties and their respective ways of defining common and distinctive variation. RESULTS: The methods are all applied on simulated data and mRNA and miRNA data-sets from GlioBlastoma Multiform (GBM) brain tumors to examine their overlap and differences. When the common variation is abundant, all methods are able to find the correct solution. With real data however, complexities in the data are treated differently by the three methods. CONCLUSIONS: All three methods have their own approach to estimate common and distinctive variation with their specific strength and weaknesses. Due to their orthogonality properties and their used algorithms their view on the data is slightly different. By assuming orthogonality between common and distinctive, true natural or biological phenomena that may not be orthogonal at all might be misinterpreted.
Frans M. van der Kloet, Patricia Sebastián-León, Ana Conesa, Age K. Smilde, Johan A. Westerhuis
BMC Bioinform.3
2014 Next maSigPro: updating maSigPro bioconductor package for RNA-seq time series
abstract
MOTIVATION: The widespread adoption of RNA-seq to quantitatively measure gene expression has increased the scope of sequencing experimental designs to include time-course experiments. maSigPro is an R package specifically suited for the analysis of time-course gene expression data, which was developed originally for microarrays and hence was limited in its application to count data. RESULTS: We have updated maSigPro to support RNA-seq time series analysis by introducing generalized linear models in the algorithm to support the modeling of count data while maintaining the traditional functionalities of the package. We show a good performance of the maSigPro-GLM method in several simulated time-course scenarios and in a real experimental dataset. AVAILABILITY AND IMPLEMENTATION: The package is freely available under the LGPL license from the Bioconductor Web site (http://bioconductor.org).
María José Nueda, Sonia Tarazona, Ana Conesa
Bioinform.3
2012 Qualimap: evaluating next-generation sequencing alignment data
abstract
MOTIVATION: The sequence alignment/map (SAM) and the binary alignment/map (BAM) formats have become the standard method of representation of nucleotide sequence alignments for next-generation sequencing data. SAM/BAM files usually contain information from tens to hundreds of millions of reads. Often, the sequencing technology, protocol and/or the selected mapping algorithm introduce some unwanted biases in these data. The systematic detection of such biases is a non-trivial task that is crucial to drive appropriate downstream analyses. RESULTS: We have developed Qualimap, a Java application that supports user-friendly quality control of mapping data, by considering sequence features and their genomic properties. Qualimap takes sequence alignment data and provides graphical and statistical analyses for the evaluation of data. Such quality-control data are vital for highlighting problems in the sequencing and/or mapping processes, which must be addressed prior to further analyses. AVAILABILITY: Qualimap is freely available from http://www.qualimap.org.
Fernando García-Alcalde, Konstantin Okonechnikov, José Carbonell, Luis M. Cruz, Stefan Götz 0003, Sonia Tarazona, Joaquín Dopazo, Thomas F. Meyer, Ana Conesa
Bioinform.9
2011 Paintomics: a web based tool for the joint visualization of transcriptomics and metabolomics data
abstract
MOTIVATION: The development of the omics technologies such as transcriptomics, proteomics and metabolomics has made possible the realization of systems biology studies where biological systems are interrogated at different levels of biochemical activity (gene expression, protein activity and/or metabolite concentration). An effective approach to the analysis of these complex datasets is the joined visualization of the disparate biomolecular data on the framework of known biological pathways. RESULTS: We have developed the Paintomics web server as an easy-to-use bioinformatics resource that facilitates the integrated visual analysis of experiments where transcriptomics and metabolomics data have been measured on different conditions for the same samples. Basically, Paintomics takes complete transcriptomics and metabolomics datasets, together with lists of significant gene or metabolite changes, and paints this information on KEGG pathway maps. AVAILABILITY: Paintomics is freely available at http://www.paintomics.org.
Fernando García-Alcalde, Federico García-López, Joaquín Dopazo, Ana Conesa
Bioinform.4
2011 B2G-FAR, a species-centered GO annotation repository
abstract
MOTIVATION: Functional genomics research has expanded enormously in the last decade thanks to the cost reduction in high-throughput technologies and the development of computational tools that generate, standardize and share information on gene and protein function such as the Gene Ontology (GO). Nevertheless, many biologists, especially working with non-model organisms, still suffer from non-existing or low-coverage functional annotation, or simply struggle retrieving, summarizing and querying these data. RESULTS: The Blast2GO Functional Annotation Repository (B2G-FAR) is a bioinformatics resource envisaged to provide functional information for otherwise uncharacterized sequence data and offers data mining tools to analyze a larger repertoire of species than currently available. This new annotation resource has been created by applying the Blast2GO functional annotation engine in a strongly high-throughput manner to the entire space of public available sequences. The resulting repository contains GO term predictions for over 13.2 million non-redundant protein sequences based on BLAST search alignments from the SIMAP database. We generated GO annotation for approximately 150 000 different taxa making available 2000 species with the highest coverage through B2G-FAR. A second section within B2G-FAR holds functional annotations for 17 non-model organism Affymetrix GeneChips. CONCLUSIONS: B2G-FAR provides easy access to exhaustive functional annotation for 2000 species offering a good balance between quality and quantity, thereby supporting functional genomics research especially in the case of non-model organisms. AVAILABILITY: The annotation resource is available at http://www.b2gfar.org.
Stefan Götz 0003, Roland Arnold, Patricia Sebastián-León, Samuel Martín-Rodríguez, Patrick Tischler, Marc-André Jehl, Joaquín Dopazo, Thomas Rattei, Ana Conesa
Bioinform.9
2009 Functional assessment of time course microarray data
abstract
MOTIVATION: Time-course microarray experiments study the progress of gene expression along time across one or several experimental conditions. Most developed analysis methods focus on the clustering or the differential expression analysis of genes and do not integrate functional information. The assessment of the functional aspects of time-course transcriptomics data requires the use of approaches that exploit the activation dynamics of the functional categories to where genes are annotated. METHODS: We present three novel methodologies for the functional assessment of time-course microarray data. i) maSigFun derives from the maSigPro method, a regression-based strategy to model time-dependent expression patterns and identify genes with differences across series. maSigFun fits a regression model for groups of genes labeled by a functional class and selects those categories which have a significant model. ii) PCA-maSigFun fits a PCA model of each functional class-defined expression matrix to extract orthogonal patterns of expression change, which are then assessed for their fit to a time-dependent regression model. iii) ASCA-functional uses the ASCA model to rank genes according to their correlation to principal time expression patterns and assess functional enrichment on a GSA fashion. We used simulated and experimental datasets to study these novel approaches. Results were compared to alternative methodologies. RESULTS: Synthetic and experimental data showed that the different methods are able to capture different aspects of the relationship between genes, functions and co-expression that are biologically meaningful. The methods should not be considered as competitive but they provide different insights into the molecular and functional dynamic events taking place within the biological system under study.
María José Nueda, Patricia Sebastián-León, Sonia Tarazona, Francisco García-García 0002, Joaquín Dopazo, Alberto Ferrer 0001, Ana Conesa
BMC Bioinform.7
2007 Discovering gene expression patterns in time course microarray experiments by ANOVA-SCA
abstract
MOTIVATION: Designed microarray experiments are used to investigate the effects that controlled experimental factors have on gene expression and learn about the transcriptional responses associated with external variables. In these datasets, signals of interest coexist with varying sources of unwanted noise in a framework of (co)relation among the measured variables and with the different levels of the studied factors. Discovering experimentally relevant transcriptional changes require methodologies that take all these elements into account. RESULTS: In this work, we develop the application of the Analysis of variance-simultaneous component analysis (ANOVA-SCA) Smilde et al. Bioinformatics, (2005) to the analysis of multiple series time course microarray data as an example of multifactorial gene expression profiling experiments. We denoted this implementation as ASCA-genes. We show how the combination of ANOVA-modeling and a dimension reduction technique is effective in extracting targeted signals from data by-passing structural noise. The methodology is valuable for identifying main and secondary responses associated with the experimental factors and spotting relevant experimental conditions. We additionally propose a novel approach for gene selection in the context of the relation of individual transcriptional patterns to global gene expression signals. We demonstrate the methodology on both real and synthetic datasets. AVAILABILITY: ASCA-genes has been implemented in the statistical language R and is available at http://www.ivia.es/centrodegenomica/bioinformatics.htm. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
María José Nueda, Ana Conesa, Johan A. Westerhuis, Huub C. J. Hoefsloot, Age K. Smilde, Manuel Talón, Alberto Ferrer 0001
Bioinform.2
2006 maSigPro: a method to identify significantly differential expression profiles in time-course microarray experiments
abstract
MOTIVATION: Multi-series time-course microarray experiments are useful approaches for exploring biological processes. In this type of experiments, the researcher is frequently interested in studying gene expression changes along time and in evaluating trend differences between the various experimental groups. The large amount of data, multiplicity of experimental conditions and the dynamic nature of the experiments poses great challenges to data analysis. RESULTS: In this work, we propose a statistical procedure to identify genes that show different gene expression profiles across analytical groups in time-course experiments. The method is a two-regression step approach where the experimental groups are identified by dummy variables. The procedure first adjusts a global regression model with all the defined variables to identify differentially expressed genes, and in second a variable selection strategy is applied to study differences between groups and to find statistically significant different profiles. The methodology is illustrated on both a real and a simulated microarray dataset.
Ana Conesa, María José Nueda, Alberto Ferrer 0001, Manuel Talón
Bioinform.1
2005 Blast2GO: a universal tool for annotation, visualization and analysis in functional genomics research
abstract
SUMMARY: We present here Blast2GO (B2G), a research tool designed with the main purpose of enabling Gene Ontology (GO) based data mining on sequence data for which no GO annotation is yet available. B2G joints in one application GO annotation based on similarity searches with statistical analysis and highlighted visualization on directed acyclic graphs. This tool offers a suitable platform for functional genomics research in non-model species. B2G is an intuitive and interactive desktop application that allows monitoring and comprehension of the whole annotation and analysis process. AVAILABILITY: Blast2GO is freely available via Java Web Start at http://www.blast2go.de. SUPPLEMENTARY MATERIAL: http://www.blast2go.de -> Evaluation.
Ana Conesa, Stefan Götz 0003, Juan Miguel García-Gómez, Javier Terol, Manuel Talón, Montserrat Robles
Bioinform.1