Ivan G. Costa

dblp:00/4681 · also Ivan Gesteira Costa · DBLP profile ↗
← Back
40ranked-venue papers
5as first author
12since 2021 · last 2025
0000-0003-2890-8697ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 30 · 5 first-author · 10 since 2021Artificial intelligence and machine learning · 8Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Advances and challenges in cell-cell communication inference: a comprehensive review of tools, resources, and future directions
abstract
Recent advancements in high-resolution and high-throughput sequencing technologies have significantly enhanced the study of cell-cell communication inference using single-cell and spatial transcriptomics data. Over the past 6 years, this growing interest has led to the development of more than 100 bioinformatics tools and nearly 50 resources, primarily in the form of ligand-receptor databases. These tools vary widely in their requirements, scoring approaches, ability to infer inter- and/or intra-cellular communication, assumptions, and limitations. Similarly, cell-cell communication resources differ in many aspects, mainly in the number of annotated interactions, species coverage, and their focus on inter-cellular signaling or both inter- and intra-cellular communication. This abundance and diversity create challenges in identifying compatible and suitable tools and resources to meet specific user needs. In this collaborative effort, we aim to provide a comprehensive report on the current state of cell-cell communication analysis derived from single-cell or spatial transcriptomics data. The report reviews existing methods and resources, addressing all relevant aspects from the user's perspective. It also explores current limitations, pitfalls, and unresolved issues in cell-cell communication inference, offering an aggregated analysis of the existing literature on the topic. Furthermore, we highlight potential future directions in the field and consolidate the collected knowledge into CCC-Catalog (https://sysbiobig.gitlab.io/ccc-catalog), a centralized web platform designed to serve as a hub for bioinformaticians and researchers interested in cell-cell communication inference.
Giulia Cesaro, James Shiniti Nagai, Nicolò Gnoato, Alice Chiodi, Gaia Tussardi, Vanessa Klöker, Carmelo Vittorio Musumarra, Ettore Mosca, Ivan G. Costa, Barbara Di Camillo, Enrica Calura, Giacomo Baruzzo
Briefings Bioinform.9
2025 PILOT-GM-VAE: patient-level analysis of single-cell disease atlas with optimal transport of Gaussian mixture variational autoencoders
abstract
The analysis of single-cell disease atlases represents a challenge due to the presence of batch effects, low quality of disease samples, and the multiscale nature of the data, i.e. samples are described by different cell distributions. Because of these, few computational approaches are performing sample-level disease progression analysis so far. Here, we introduce Patient-Level Analysis with Optimal Transport based on Gaussian Mixture Variational Autoencoders (PILOT-GM-VAE). PILOT-GM-VAE explores the power of GM-VAE to estimate models describing complex single-cell distributions through efficient optimal transport algorithms for estimating the distance between Gaussian Mixtures. Extensive benchmarking on 12 single-cell disease atlases and competing approaches demonstrate the performance of PILOT-GM-VAE in sample-level clustering, sample-level trajectory inference, and batch correction tasks. Moreover, we performed a case study on a breast cancer disease atlas, where PILOT-GM-VAE highlighted cellular and molecular changes associated with breast cancer disease progression.
Mehdi Joodaki, Mina Shaigan, Samaneh Samiei, James Shiniti Nagai, Tiago Maié, Christoph Kuppe, Ivan G. Costa
Briefings Bioinform.7
2025 scACCorDiON: a clustering approach for explainable patient level cell-cell communication graph analysis
abstract
MOTIVATION: Combining single-cell sequencing with ligand-receptor (LR) analysis paves the way for the characterization of cell communication events in complex tissues. In particular, directed weighted graphs naturally represent cell-cell communication events. However, current computational methods cannot yet analyze sample-specific cell-cell communication events, as measured in single-cell data produced in large patient cohorts. Cohort-based cell-cell communication analysis presents many challenges, such as the nonlinear nature of cell-cell communication and the high variability given by the patient-specific single-cell RNAseq datasets. RESULTS: Here, we present scACCorDiON (single-cell Analysis of Cell-Cell Communication in Disease clusters using Optimal transport in Directed Networks), an optimal transport algorithm exploring node distances on the Markov Chain as the ground metric between directed weighted graphs. Benchmarking indicates that scACCorDiON performs a better clustering of samples according to their disease status than competing methods that use undirected graphs. We provide a case study of pancreas adenocarcinoma, where scACCorDion detects a sub-cluster of disease samples associated with changes in the tumor microenvironment. Our study case corroborates that clusters provide a robust and explainable representation of cell-cell communication events and that the expression of detected LR pairs is predictive of pancreatic cancer survival. AVAILABILITY AND IMPLEMENTATION: The code of scACCorDiON is available at https://scaccordion.readthedocs.io/en/latest/. and https://doi.org/10.5281/zenodo.15267648. The survival analysis package can be found at https://github.com/CostaLab/scACCorDiON.su.
James Shiniti Nagai, Tiago Maié, Michael T. Schaub, Ivan G. Costa
Bioinform.4
2024 Optimal Transport Distances for Directed, Weighted Graphs: A Case Study With Cell-Cell Communication Networks
abstract
Comparing graphs by means of optimal transport has recently gained significant attention, as the distances induced by optimal transport provide both a principled metric between graphs as well as an interpretable description of the associated changes between graphs in terms of a transport plan. As the lack of symmetry introduces challenges in the typically considered formulations, optimal transport distances for graphs have mostly been developed for undirected graphs. Here1, we propose two distance measures to compare directed graphs based on variants of optimal transport(OT): (i) an earth movers distance (Wasserstein) and (ii) a Gromov-Wasserstein (GW) distance. We evaluate these two distances and discuss their relative performance for both simulated graph data and real-world directed cell-cell communication graphs, inferred from single-cell RNA-seq data.
James Shiniti Nagai, Ivan G. Costa, Michael T. Schaub
ICASSP2
2024 SIngle cell level Genotyping Using scRna Data (SIGURD)
abstract
MOTIVATION: By accounting for variants within measured transcripts, it is possible to evaluate the status of somatic variants using single-cell RNA-sequencing (scRNA-seq) and to characterize their clonality. However, the sparsity (very few reads per transcript) or bias in protocols (favoring 3' ends of the transcripts) makes the chance of capturing somatic variants very unlikely. This can be overcome by targeted sequencing or the use of mitochondrial variants as natural barcodes for clone identification. Currently, available computational tools focus on genotyping, but do not provide functionality for combined analysis of somatic and mitochondrial variants and functional analysis such as characterization of gene expression changes in detected clones. RESULTS: Here, we propose SIGURD (SIngle cell level Genotyping Using scRna Data) (SIGURD), which is an R-based pipeline for the clonal analysis of scRNA-seq data. This allows the quantification of clones by leveraging both somatic and mitochondrial variants. SIGURD also allows for functional analysis after clonal detection: association of clones with cell populations, detection of differentially expressed genes across clones, and association of somatic and mitochondrial variants. Here, we demonstrate the power of SIGURD by analyzing single-cell data of colony-forming cells derived from patients with myeloproliferative neoplasms.
Martin Grasshoff, Milena Kalmer, Nicolas Chatain, Kim Kricheldorf, Angela Maurer, Ralf Weiskirchen, Steffen Koschmieder, Ivan G. Costa
Briefings Bioinform.8
2024 tRigon: an R package and Shiny App for integrative (path-)omics data analysis
abstract
BACKGROUND: Pathomics facilitates automated, reproducible and precise histopathology analysis and morphological phenotyping. Similar to molecular omics, pathomics datasets are high-dimensional, but also face large outlier variability and inherent data missingness, making quick and comprehensible data analysis challenging. To facilitate pathomics data analysis and interpretation as well as support a broad implementation we developed tRigon (Toolbox foR InteGrative (path-)Omics data aNalysis), a Shiny application for fast, comprehensive and reproducible pathomics analysis. RESULTS: tRigon is available via the CRAN repository ( https://cran.r-project.org/web/packages/tRigon ) with its source code available on GitLab ( https://git-ce.rwth-aachen.de/labooratory-ai/trigon ). The tRigon package can be installed locally and its application can be executed from the R console via the command 'tRigon::run_tRigon()'. Alternatively, the application is hosted online and can be accessed at https://labooratory.shinyapps.io/tRigon . We show fast computation of small, medium and large datasets in a low- and high-performance hardware setting, indicating broad applicability of tRigon. CONCLUSIONS: tRigon allows researchers without coding abilities to perform exploratory feature analyses of pathomics and non-pathomics datasets on their own using a variety of hardware.
David L. Hölscher, Michael Goedertier, Barbara Mara Klinkhammer, Patrick Droste, Ivan G. Costa, Peter Boor, Roman David Bülow
BMC Bioinform.5
2023 RGT: a toolbox for the integrative analysis of high throughput regulatory genomics data
abstract
BACKGROUND: Massive amounts of data are produced by combining next-generation sequencing with complex biochemistry techniques to characterize regulatory genomics profiles, such as protein-DNA interaction and chromatin accessibility. Interpretation of such high-throughput data typically requires different computation methods. However, existing tools are usually developed for a specific task, which makes it challenging to analyze the data in an integrative manner. RESULTS: We here describe the Regulatory Genomics Toolbox (RGT), a computational library for the integrative analysis of regulatory genomics data. RGT provides different functionalities to handle genomic signals and regions. Based on that, we developed several tools to perform distinct downstream analyses, including the prediction of transcription factor binding sites using ATAC-seq data, identification of differential peaks from ChIP-seq data, and detection of triple helix mediated RNA and DNA interactions, visualization, and finding an association between distinct regulatory factors. CONCLUSION: We present here RGT; a framework to facilitate the customization of computational methods to analyze genomic data for specific regulatory genomics problems. RGT is a comprehensive and flexible Python package for analyzing high throughput regulatory genomics data and is available at: https://github.com/CostaLab/reg-gen . The documentation is available at: https://reg-gen.readthedocs.io.
Chao-Chung Kuo, Fabio Ticconi, Mina Shaigan, Julia Gehrmann, Eduardo G. Gusmão, Manuel Allhoff, Martin Manolov, Martin Zenke, Ivan G. Costa
BMC Bioinform.10
2022 MOJITOO: a fast and universal method for integration of multimodal single-cell data
abstract
MOTIVATION: The advent of multi-modal single-cell sequencing techniques have shed new light on molecular mechanisms by simultaneously inspecting transcriptomes, epigenomes and proteomes of the same cell. However, to date, the existing computational approaches for integration of multimodal single-cell data are either computationally expensive, require the delineation of parameters or can only be applied to particular modalities. RESULTS: Here we present a single-cell multi-modal integration method, named Multi-mOdal Joint IntegraTion of cOmpOnents (MOJITOO). MOJITOO uses canonical correlation analysis for a fast and parameter free detection of a shared representation of cells from multimodal single-cell data. Moreover, estimated canonical components can be used for interpretation, i.e. association of modality-specific molecular features with the latent space. We evaluate MOJITOO using bi- and tri-modal single-cell datasets and show that MOJITOO outperforms existing methods regarding computational requirements, preservation of original latent spaces and clustering. AVAILABILITY AND IMPLEMENTATION: The software, code and data for benchmarking are available at https://github.com/CostaLab/MOJITOO and https://doi.org/10.5281/zenodo.6348128. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Mingbo Cheng, Ivan G. Costa
Bioinform.3
2022 Detection of cell markers from single cell RNA-seq with sc2marker
abstract
BACKGROUND: Single-cell RNA sequencing (scRNA-seq) allows the detection of rare cell types in complex tissues. The detection of markers for rare cell types is useful for further biological analysis of, for example, flow cytometry and imaging data sets for either physical isolation or spatial characterization of these cells. However, only a few computational approaches consider the problem of selecting specific marker genes from scRNA-seq data. RESULTS: Here, we propose sc2marker, which is based on the maximum margin index and a database of proteins with antibodies, to select markers for flow cytometry or imaging. We evaluated the performances of sc2marker and competing methods in ranking known markers in scRNA-seq data of immune and stromal cells. The results showed that sc2marker performed better than the competing methods in accuracy, while having a competitive running time.
Ronghui Li, Bella Banjanin, Rebekka K. Schneider, Ivan G. Costa
BMC Bioinform.4
2022 The area under the ROC curve as a measure of clustering quality
Pablo A. Jaskowiak, Ivan G. Costa, Ricardo J. G. B. Campello
Data Min. Knowl. Discov.2
2021 Deep learning-based clustering approaches for bioinformatics
abstract
Clustering is central to many data-driven bioinformatics research and serves a powerful computational method. In particular, clustering helps at analyzing unstructured and high-dimensional data in the form of sequences, expressions, texts and images. Further, clustering is used to gain insights into biological processes in the genomics level, e.g. clustering of gene expressions provides insights on the natural structure inherent in the data, understanding gene functions, cellular processes, subtypes of cells and understanding gene regulations. Subsequently, clustering approaches, including hierarchical, centroid-based, distribution-based, density-based and self-organizing maps, have long been studied and used in classical machine learning settings. In contrast, deep learning (DL)-based representation and feature learning for clustering have not been reviewed and employed extensively. Since the quality of clustering is not only dependent on the distribution of data points but also on the learned representation, deep neural networks can be effective means to transform mappings from a high-dimensional data space into a lower-dimensional feature space, leading to improved clustering results. In this paper, we review state-of-the-art DL-based approaches for cluster analysis that are based on representation learning, which we hope to be useful, particularly for bioinformatics research. Further, we explore in detail the training procedures of DL-based clustering algorithms, point out different clustering quality metrics and evaluate several DL-based approaches on three bioinformatics use cases, including bioimaging, cancer genomics and biomedical text mining. We believe this review and the evaluation results will provide valuable insights and serve a starting point for researchers wanting to apply DL-based unsupervised methods to solve emerging bioinformatics research problems.
Md. Rezaul Karim 0001, Oya Beyan, Achille Zappa, Ivan G. Costa, Dietrich Rebholz-Schuhmann, Michael Cochez, Stefan Decker
Briefings Bioinform.4
2021 CrossTalkeR: analysis and visualization of ligand-receptorne tworks
abstract
MOTIVATION: Ligand-receptor (LR) network analysis allows the characterization of cellular crosstalk based on single cell RNA-seq data. However, current methods typically provide a list of inferred LR interactions and do not allow the researcher to focus on specific cell types, ligands or receptors. In addition, most of these methods cannot quantify changes in crosstalk between two biological phenotypes. RESULTS: CrossTalkeR is a framework for network analysis and visualization of LR interactions. CrossTalkeR identifies relevant ligands, receptors and cell types contributing to changes in cell communication when contrasting two biological phenotypes, i.e. disease versus homeostasis. A case study on scRNA-seq of human myeloproliferative neoplasms reinforces the strengths of CrossTalkeR for characterization of changes in cellular crosstalk in disease. AVAILABILITY AND IMPLEMENTATION: CrosstalkeR is an R package available at: Github: https://github.com/CostaLab/CrossTalkeR. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
James Shiniti Nagai, Nils B. Leimkühler, Michael T. Schaub, Rebekka K. Schneider, Ivan G. Costa
Bioinform.5
2018 Data complexity meta-features for regression problems
Ana Carolina Lorena, Aron I. Maciel, Péricles B. C. Miranda, Ivan G. Costa, Ricardo B. C. Prudêncio
Mach. Learn.4
2016 Measuring the complexity of regression problems
abstract
Many works have attempted to characterize the complexity of classification problems by measures extracted from their learning datasets. These indexes provide indicatives of the inherent difficulty in solving a given classification problem. Although regression problems are equally frequent, there is a lack of studies in Machine Learning dedicated to understanding their complexity. This paper proposes some measures aimed to characterize the complexity of regression problems. They are experimentally evaluated on a set of synthetic datasets with different complexities. The results show that various measures and their combinations are able to distinguish simple linear problems from more complex variants.
Aron I. Maciel, Ivan G. Costa, Ana Carolina Lorena
IJCNN2
2016 A multiple kernel learning algorithm for drug-target interaction prediction
abstract
BACKGROUND: Drug-target networks are receiving a lot of attention in late years, given its relevance for pharmaceutical innovation and drug lead discovery. Different in silico approaches have been proposed for the identification of new drug-target interactions, many of which are based on kernel methods. Despite technical advances in the latest years, these methods are not able to cope with large drug-target interaction spaces and to integrate multiple sources of biological information. RESULTS: We propose KronRLS-MKL, which models the drug-target interaction problem as a link prediction task on bipartite networks. This method allows the integration of multiple heterogeneous information sources for the identification of new interactions, and can also work with networks of arbitrary size. Moreover, it automatically selects the more relevant kernels by returning weights indicating their importance in the drug-target prediction at hand. Empirical analysis on four data sets using twenty distinct kernels indicates that our method has higher or comparable predictive performance than 18 competing methods in all prediction tasks. Moreover, the predicted weights reflect the predictive quality of each kernel on exhaustive pairwise experiments, which indicates the success of the method to automatically reveal relevant biological sources. CONCLUSIONS: Our analysis show that the proposed data integration strategy is able to improve the quality of the predicted interactions, and can speed up the identification of new drug-target interactions as well as identify relevant information for the task. AVAILABILITY: The source code and data sets are available at www.cin.ufpe.br/~acan/kronrlsmkl/.
André C. A. Nascimento, Ricardo B. C. Prudêncio, Ivan G. Costa
BMC Bioinform.3
2015 Detecting differential peaks in ChIP-seq signals with ODIN
abstract
Bioinformatics (2014); 30(24), 3467–3475 doi: 10.1093/bioinformatics/btu722 In Figure 4 of the above article, there were several words missing from the figure legend due to a typesetting error. Please see below for the complete Figure 4.
Manuel Allhoff, Kristin Seré, Heike Chauvistré, Qiong Lin, Martin Zenke, Ivan G. Costa
Bioinform.6
2015 Impact of missing data imputation methods on gene expression clustering and classification
abstract
BACKGROUND: Several missing value imputation methods for gene expression data have been proposed in the literature. In the past few years, researchers have been putting a great deal of effort into presenting systematic evaluations of the different imputation algorithms. Initially, most algorithms were assessed with an emphasis on the accuracy of the imputation, using metrics such as the root mean squared error. However, it has become clear that the success of the estimation of the expression value should be evaluated in more practical terms as well. One can consider, for example, the ability of the method to preserve the significant genes in the dataset, or its discriminative/predictive power for classification/clustering purposes. RESULTS AND CONCLUSIONS: We performed a broad analysis of the impact of five well-known missing value imputation methods on three clustering and four classification methods, in the context of 12 cancer gene expression datasets. We employed a statistical framework, for the first time in this field, to assess whether different imputation methods improve the performance of the clustering/classification methods. Our results suggest that the imputation methods evaluated have a minor impact on the classification and downstream clustering analyses. Simple methods such as replacing the missing values by mean or the median values performed as well as more complex strategies. The datasets analyzed in this study are available at http://costalab.org/Imputation/ .
Marcílio Carlos Pereira de Souto, Pablo A. Jaskowiak, Ivan G. Costa
BMC Bioinform.3
2014 Detecting differential peaks in ChIP-seq signals with ODIN
abstract
MOTIVATION: Detection of changes in deoxyribonucleic acid (DNA)-protein interactions from ChIP-seq data is a crucial step in unraveling the regulatory networks behind biological processes. The simplest variation of this problem is the differential peak calling (DPC) problem. Here, one has to find genomic regions with ChIP-seq signal changes between two cellular conditions in the interaction of a protein with DNA. The great majority of peak calling methods can only analyze one ChIP-seq signal at a time and are unable to perform DPC. Recently, a few approaches based on the combination of these peak callers with statistical tests for detecting differential digital expression have been proposed. However, these methods fail to detect detailed changes of protein-DNA interactions. RESULTS: We propose an One-stage DIffereNtial peak caller (ODIN); an Hidden Markov Model-based approach to detect and analyze differential peaks (DPs) in pairs of ChIP-seq data. ODIN performs genomic signal processing, peak calling and p-value calculation in an integrated framework. We also propose an evaluation methodology to compare ODIN with competing methods. The evaluation method is based on the association of DPs with expression changes in the same cellular conditions. Our empirical study based on several ChIP-seq experiments from transcription factors, histone modifications and simulated data shows that ODIN outperforms considered competing methods in most scenarios.
Manuel Allhoff, Kristin Seré, Heike Chauvistré, Qiong Lin, Martin Zenke, Ivan G. Costa
Bioinform.6
2014 Detection of active transcription factor binding sites with the combination of DNase hypersensitivity and histone modifications
abstract
MOTIVATION: The identification of active transcriptional regulatory elements is crucial to understand regulatory networks driving cellular processes such as cell development and the onset of diseases. It has recently been shown that chromatin structure information, such as DNase I hypersensitivity (DHS) or histone modifications, significantly improves cell-specific predictions of transcription factor binding sites. However, no method has so far successfully combined both DHS and histone modification data to perform active binding site prediction. RESULTS: We propose here a method based on hidden Markov models to integrate DHS and histone modifications occupancy for the detection of open chromatin regions and active binding sites. We have created a framework that includes treatment of genomic signals, model training and genome-wide application. In a comparative analysis, our method obtained a good trade-off between sensitivity versus specificity and superior area under the curve statistics than competing methods. Moreover, our technique does not require further training or sequence information to generate binding location predictions. Therefore, the method can be easily applied on new cell types and allow flexible downstream analysis such as de novo motif finding. AVAILABILITY AND IMPLEMENTATION: Our framework is available as part of the Regulatory Genomics Toolbox. The software information and all benchmarking data are available at http://costalab.org/wp/dh-hmm. CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Eduardo G. Gusmão, Christoph Dieterich, Martin Zenke, Ivan G. Costa
Bioinform.4
2014 On the selection of appropriate distances for gene expression data clustering
abstract
BACKGROUND: Clustering is crucial for gene expression data analysis. As an unsupervised exploratory procedure its results can help researchers to gain insights and formulate new hypothesis about biological data from microarrays. Given different settings of microarray experiments, clustering proves itself as a versatile exploratory tool. It can help to unveil new cancer subtypes or to identify groups of genes that respond similarly to a specific experimental condition. In order to obtain useful clustering results, however, different parameters of the clustering procedure must be properly tuned. Besides the selection of the clustering method itself, determining which distance is going to be employed between data objects is probably one of the most difficult decisions. RESULTS AND CONCLUSIONS: We analyze how different distances and clustering methods interact regarding their ability to cluster gene expression, i.e., microarray data. We study 15 distances along with four common clustering methods from the literature on a total of 52 gene expression microarray datasets. Distances are evaluated on a number of different scenarios including clustering of cancer tissues and genes from short time-series expression data, the two main clustering applications in gene expression. Our results support that the selection of an appropriate distance depends on the scenario in hand. Moreover, in each scenario, given the very same clustering method, significant differences in quality may arise from the selection of distinct distance measures. In fact, the selection of an appropriate distance measure can make the difference between meaningful and poor clustering outcomes, even for a suitable clustering method.
Pablo A. Jaskowiak, Ricardo J. G. B. Campello, Ivan G. Costa
BMC Bioinform.3
2013 Discovering motifs that induce sequencing errors
abstract
BACKGROUND: Elevated sequencing error rates are the most predominant obstacle in single-nucleotide polymorphism (SNP) detection, which is a major goal in the bulk of current studies using next-generation sequencing (NGS). Beyond routinely handled generic sources of errors, certain base calling errors relate to specific sequence patterns. Statistically principled ways to associate sequence patterns with base calling errors have not been previously described. Extant approaches either incur decisive losses in power, due to relating errors with individual genomic positions rather than motifs, or do not properly distinguish between motif-induced and sequence-unspecific sources of errors. RESULTS: Here, for the first time, we describe a statistically rigorous framework for the discovery of motifs that induce sequencing errors. We apply our method to several datasets from Illumina GA IIx, HiSeq 2000, and MiSeq sequencers. We confirm previously known error-causing sequence contexts and report new more specific ones. CONCLUSIONS: Checking for error-inducing motifs should be included into SNP calling pipelines to avoid false positives. To facilitate filtering of sets of putative SNPs, we provide tracks of error-prone genomic positions (in BED format). AVAILABILITY: http://discovering-cse.googlecode.com.
Manuel Allhoff, Alexander Schönhuth, Marcel Martin, Ivan G. Costa, Sven Rahmann, Tobias Marschall
BMC Bioinform.4
2013 Proximity Measures for Clustering Gene Expression Microarray Data: A Validation Methodology and a Comparative Analysis
abstract
Cluster analysis is usually the first step adopted to unveil information from gene expression microarray data. Besides selecting a clustering algorithm, choosing an appropriate proximity measure (similarity or distance) is of great importance to achieve satisfactory clustering results. Nevertheless, up to date, there are no comprehensive guidelines concerning how to choose proximity measures for clustering microarray data. Pearson is the most used proximity measure, whereas characteristics of other ones remain unexplored. In this paper, we investigate the choice of proximity measures for the clustering of microarray data by evaluating the performance of 16 proximity measures in 52 data sets from time course and cancer experiments. Our results support that measures rarely employed in the gene expression literature can provide better results than commonly employed ones, such as Pearson, Spearman, and euclidean distance. Given that different measures stood out for time course and cancer data evaluations, their choice should be specific to each scenario. To evaluate measures on time-course data, we preprocessed and compiled 17 data sets from the microarray literature in a benchmark along with a new methodology, called Intrinsic Biological Separation Ability (IBSA). Both can be employed in future research to assess the effectiveness of new measures for gene time-course data.
Pablo A. Jaskowiak, Ricardo J. G. B. Campello, Ivan G. Costa
IEEE ACM Trans. Comput. Biol. Bioinform.3
2012 CLEVER: clique-enumerating variant finder
abstract
MOTIVATION: Next-generation sequencing techniques have facilitated a large-scale analysis of human genetic variation. Despite the advances in sequencing speed, the computational discovery of structural variants is not yet standard. It is likely that many variants have remained undiscovered in most sequenced individuals. RESULTS: Here, we present a novel internal segment size based approach, which organizes all, including concordant, reads into a read alignment graph, where max-cliques represent maximal contradiction-free groups of alignments. A novel algorithm then enumerates all max-cliques and statistically evaluates them for their potential to reflect insertions or deletions. For the first time in the literature, we compare a large range of state-of-the-art approaches using simulated Illumina reads from a fully annotated genome and present relevant performance statistics. We achieve superior performance, in particular, for deletions or insertions (indels) of length 20-100 nt. This has been previously identified as a remaining major challenge in structural variation discovery, in particular, for insert size based approaches. In this size range, we even outperform split-read aligners. We achieve competitive results also on biological data, where our method is the only one to make a substantial amount of correct predictions, which, additionally, are disjoint from those by split-read aligners. AVAILABILITY: CLEVER is open source (GPL) and available from http://clever-sv.googlecode.com. CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Tobias Marschall, Ivan G. Costa, Stefan Canzar, Markus Bauer 0001, Gunnar W. Klau, Alexander Schliep, Alexander Schönhuth
Bioinform.2
2012 Inferring epigenetic and transcriptional regulation during blood cell development with a mixture of sparse linear models
abstract
MOTIVATION: Blood cell development is thought to be controlled by a circuit of transcription factors (TFs) and chromatin modifications that determine the cell fate through activating cell type-specific expression programs. To shed light on the interplay between histone marks and TFs during blood cell development, we model gene expression from regulatory signals by means of combinations of sparse linear regression models. RESULTS: The mixture of sparse linear regression models was able to improve the gene expression prediction in relation to the use of a single linear model. Moreover, it performed an efficient selection of regulatory signals even when analyzing all TFs with known motifs (>600). The method identified interesting roles for histone modifications and a selection of TFs related to blood development and chromatin remodelling. AVAILABILITY: The method and datasets are available from http://www.cin.ufpe.br/~igcf/SparseMix. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Thaís Gaudencio do Rêgo, Helge G. Roider, Francisco de A. T. de Carvalho, Ivan G. Costa
Bioinform.4
2012 Analysis of complexity indices for classification problems: Cancer gene expression data
Ana Carolina Lorena, Ivan G. Costa, Newton Spolaôr, Marcílio Carlos Pereira de Souto
Neurocomputing2
2011 Classifying short gene expression time-courses with Bayesian estimation of piecewise constant functions
abstract
MOTIVATION: Analyzing short time-courses is a frequent and relevant problem in molecular biology, as, for example, 90% of gene expression time-course experiments span at most nine time-points. The biological or clinical questions addressed are elucidating gene regulation by identification of co-expressed genes, predicting response to treatment in clinical, trial-like settings or classifying novel toxic compounds based on similarity of gene expression time-courses to those of known toxic compounds. The latter problem is characterized by irregular and infrequent sample times and a total lack of prior assumptions about the incoming query, which comes in stark contrast to clinical settings and requires to implicitly perform a local, gapped alignment of time series. The current state-of-the-art method (SCOW) uses a variant of dynamic time warping and models time series as higher order polynomials (splines). RESULTS: We suggest to model time-courses monitoring response to toxins by piecewise constant functions, which are modeled as left-right Hidden Markov Models. A Bayesian approach to parameter estimation and inference helps to cope with the short, but highly multivariate time-courses. We improve prediction accuracy by 7% and 4%, respectively, when classifying toxicology and stress response data. We also reduce running times by at least a factor of 140; note that reasonable running times are crucial when classifying response to toxins. In conclusion, we have demonstrated that appropriate reduction of model complexity can result in substantial improvements both in classification performance and running time. AVAILABILITY: A Python package implementing the methods described is freely available under the GPL from http://bioinformatics.rutgers.edu/Software/MVQueries/.
Christoph Hafemeister, Ivan G. Costa, Alexander Schönhuth, Alexander Schliep
Bioinform.2
2011 Detection and interpretation of metabolite-transcript coresponses using combined profiling data
abstract
MOTIVATION: Studying the interplay between gene expression and metabolite levels can yield important information on the physiology of stress responses and adaptation strategies. Performing transcriptomics and metabolomics in parallel during time-series experiments represents a systematic way to gain such information. Several combined profiling datasets have been added to the public domain and they form a valuable resource for hypothesis generating studies. Unfortunately, detecting coresponses between transcript levels and metabolite abundances is non-trivial: they cannot be assumed to overlap directly with underlying biochemical pathways and they may be subject to time delays and obscured by considerable noise. RESULTS: Our aim was to predict pathway comemberships between metabolites and genes based on their coresponses to applied stress. We found that in the presence of strong noise and time-shifted responses, a hidden Markov model-based similarity outperforms the simpler Pearson correlation but performs comparably or worse in their absence. Therefore, we propose a supervised method that applies pathway information to summarize similarity statistics to a consensus statistic that is more informative than any of the single measures. Using four combined profiling datasets, we show that comembership between metabolites and genes can be predicted for numerous KEGG pathways; this opens opportunities for the detection of transcriptionally regulated pathways and novel metabolically related genes. AVAILABILITY: A command-line software tool is available at http://www.cin.ufpe.br/~igcf/Metabolites. CONTACT: [email protected]; [email protected]
Henning Redestig, Ivan G. Costa
Bioinform.2
2011 Predicting gene expression in T cell differentiation from histone modifications and transcription factor binding affinities by linear mixture models
abstract
BACKGROUND: The differentiation process from stem cells to fully differentiated cell types is controlled by the interplay of chromatin modifications and transcription factor activity. Histone modifications or transcription factors frequently act in a multi-functional manner, with a given DNA motif or histone modification conveying both transcriptional repression and activation depending on its location in the promoter and other regulatory signals surrounding it. RESULTS: To account for the possible multi functionality of regulatory signals, we model the observed gene expression patterns by a mixture of linear regression models. We apply the approach to identify the underlying histone modifications and transcription factors guiding gene expression of differentiated CD4+ T cells. The method improves the gene expression prediction in relation to the use of a single linear model, as often used by previous approaches. Moreover, it recovered the known role of the modifications H3K4me3 and H3K27me3 in activating cell specific genes and of some transcription factors related to CD4+ T differentiation.
Ivan G. Costa, Helge G. Roider, Thaís Gaudencio do Rêgo, Francisco de A. T. de Carvalho
BMC Bioinform.1
2010 Complexity measures of supervised classifications tasks: A case study for cancer gene expression data
abstract
Machine Learning algorithms have been widely used for gene expression data classification, despite the fact that these data have often intrinsic limitations, such as high dimensionality and a small number of examples. Few studies try to characterize to which extent these aspects can influence the performance of the classification models induced. In this paper we compute different measures characterizing the complexity of gene expression data sets for cancer diagnosis. We then investigate how these measures relate to the classification performances achieved by support vector machines, a popular Machine Learning technique usually employed in the analysis of gene expression data. The results obtained indicate that some of the complexity indices utilized are indeed successful in explaining the difficulty involved in the classification of cancer gene expression data.
Marcílio Carlos Pereira de Souto, Ana Carolina Lorena, Newton Spolaôr, Ivan G. Costa
IJCNN4
2010 PyMix - The Python mixture package - a tool for clustering of heterogeneous biological data
abstract
BACKGROUND: Cluster analysis is an important technique for the exploratory analysis of biological data. Such data is often high-dimensional, inherently noisy and contains outliers. This makes clustering challenging. Mixtures are versatile and powerful statistical models which perform robustly for clustering in the presence of noise and have been successfully applied in a wide range of applications. RESULTS: PyMix - the Python mixture package implements algorithms and data structures for clustering with basic and advanced mixture models. The advanced models include context-specific independence mixtures, mixtures of dependence trees and semi-supervised learning. PyMix is licenced under the GNU General Public licence (GPL). PyMix has been successfully used for the analysis of biological sequence, complex disease and gene expression data. CONCLUSIONS: PyMix is a useful tool for cluster analysis of biological data. Due to the general nature of the framework, PyMix can be applied to a wide range of applications and data sets.
Benjamin Georgi, Ivan G. Costa, Alexander Schliep
BMC Bioinform.2
2009 Mining Rules for the Automatic Selection Process of Clustering Methods Applied to Cancer Gene Expression Data
André C. A. Nascimento, Ricardo B. C. Prudêncio, Marcílio Carlos Pereira de Souto, Ivan G. Costa
ICANN (2)4
2009 Constrained mixture estimation for analysis and robust classification of clinical time series
abstract
MOTIVATION: Personalized medicine based on molecular aspects of diseases, such as gene expression profiling, has become increasingly popular. However, one faces multiple challenges when analyzing clinical gene expression data; most of the well-known theoretical issues such as high dimension of feature spaces versus few examples, noise and missing data apply. Special care is needed when designing classification procedures that support personalized diagnosis and choice of treatment. Here, we particularly focus on classification of interferon-beta (IFNbeta) treatment response in Multiple Sclerosis (MS) patients which has attracted substantial attention in the recent past. Half of the patients remain unaffected by IFNbeta treatment, which is still the standard. For them the treatment should be timely ceased to mitigate the side effects. RESULTS: We propose constrained estimation of mixtures of hidden Markov models as a methodology to classify patient response to IFNbeta treatment. The advantages of our approach are that it takes the temporal nature of the data into account and its robustness with respect to noise, missing data and mislabeled samples. Moreover, mixture estimation enables to explore the presence of response sub-groups of patients on the transcriptional level. We clearly outperformed all prior approaches in terms of prediction accuracy, raising it, for the first time, >90%. Additionally, we were able to identify potentially mislabeled samples and to sub-divide the good responders into two sub-groups that exhibited different transcriptional response programs. This is supported by recent findings on MS pathology and therefore may raise interesting clinical follow-up questions. AVAILABILITY: The method is implemented in the GQL framework and is available at http://www.ghmm.org/gql. Datasets are available at http://www.cin.ufpe.br/ approximately igcf/MSConst. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Ivan G. Costa, Alexander Schönhuth, Christoph Hafemeister, Alexander Schliep
Bioinform.1
2008 On the Complexity of Gene Expression Classification Data Sets
abstract
One of the main kinds of computational tasks regarding gene expression data is the construction of classifiers (models), often via some machine learning (ML) technique and given data sets, to automatically discriminate expression patterns from cancer (tumor) and normal tissues or from subtypes of cancers. A very distinctive characteristic of these data sets is its high dimensionality and the fewer number of data items. Such a characteristic makes the induction of accurate ML models difficult (e.g., it could lead to model overfitting). In this context, we present an empirical study on the complexity of the classification task of gene expression data sets, related to cancer, used for classification purposes. In order to do so, we measure the complexity of the ML models used to perform the tumors' classification. The results indicate that most of these data sets can be effectively discriminated by a simple linear function.
Ana Carolina Lorena, Ivan G. Costa, Marcílio Carlos Pereira de Souto
HIS2
2008 Comparative study on normalization procedures for cluster analysis of gene expression datasets
abstract
Normalization before clustering is often needed for proximity indices, such as Euclidian distance, which are sensitive to differences in the magnitude or scales of the attributes. The goal is to equalize the size or magnitude and the variability of these features. This can also be seen as a way to adjust the relative weighting of the attributes. In this context, we present a first large scale data driven comparative study of three normalization procedures applied to cancer gene expression data. The results are presented in terms of the recovering of the true cluster structure as found by five different clustering algorithms.
Marcílio Carlos Pereira de Souto, Daniel Araújo 0001, Ivan G. Costa, Rodrigo G. F. Soares, Teresa Bernarda Ludermir, Alexander Schliep
IJCNN3
2008 Ranking and selecting clustering algorithms using a meta-learning approach
abstract
We present a novel framework that applies a meta-learning approach to clustering algorithms. Given a dataset, our meta-learning approach provides a ranking for the candidate algorithms that could be used with that dataset. This ranking could, among other things, support non-expert users in the algorithm selection task. In order to evaluate the framework proposed, we implement a prototype that employs regression support vector machines as the meta-learner. Our case study is developed in the context of cancer gene expression micro-array datasets.
Marcílio Carlos Pereira de Souto, Ricardo B. C. Prudêncio, Rodrigo G. F. Soares, Daniel Araújo 0001, Ivan G. Costa, Teresa Bernarda Ludermir, Alexander Schliep
IJCNN5
2008 Inferring differentiation pathways from gene expression
abstract
MOTIVATION: The regulation of proliferation and differentiation of embryonic and adult stem cells into mature cells is central to developmental biology. Gene expression measured in distinguishable developmental stages helps to elucidate underlying molecular processes. In previous work we showed that functional gene modules, which act distinctly in the course of development, can be represented by a mixture of trees. In general, the similarities in the gene expression programs of cell populations reflect the similarities in the differentiation path. RESULTS: We propose a novel model for gene expression profiles and an unsupervised learning method to estimate developmental similarity and infer differentiation pathways. We assess the performance of our model on simulated data and compare it with favorable results to related methods. We also infer differentiation pathways and predict functional modules in gene expression data of lymphoid development. CONCLUSIONS: We demonstrate for the first time how, in principal, the incorporation of structural knowledge about the dependence structure helps to reveal differentiation pathways and potentially relevant functional gene modules from microarray datasets. Our method applies in any area of developmental biology where it is possible to obtain cells of distinguishable differentiation stages. AVAILABILITY: The implementation of our method (GPL license), data and additional results are available at http://algorithmics.molgen.mpg.de/Supplements/InfDif/. SUPPLEMENTARY INFORMATION: Supplementary data is available at Bioinformatics online.
Ivan G. Costa, Stefan Roepcke, Christoph Hafemeister, Alexander Schliep
ISMB1
2008 Clustering cancer gene expression data: a comparative study
abstract
BACKGROUND: The use of clustering methods for the discovery of cancer subtypes has drawn a great deal of attention in the scientific community. While bioinformaticians have proposed new clustering methods that take advantage of characteristics of the gene expression data, the medical community has a preference for using "classic" clustering methods. There have been no studies thus far performing a large-scale evaluation of different clustering methods in this context. RESULTS/CONCLUSION: We present the first large-scale analysis of seven different clustering methods and four proximity measures for the analysis of 35 cancer gene expression data sets. Our results reveal that the finite mixture of Gaussians, followed closely by k-means, exhibited the best performance in terms of recovering the true structure of the data sets. These methods also exhibited, on average, the smallest difference between the actual number of classes in the data sets and the best number of clusters as indicated by our validation criteria. Furthermore, hierarchical methods, which have been widely used by the medical community, exhibited a poorer recovery performance than that of the other methods evaluated. Moreover, as a stable basis for the assessment and comparison of different clustering methods for cancer gene expression data, this study provides a common group of data sets (benchmark data sets) to be shared among researchers and used for comparisons with new methods. The data sets analyzed in this study are available at http://algorithmics.molgen.mpg.de/Supplements/CompCancer/.
Marcílio Carlos Pereira de Souto, Ivan G. Costa, Daniel Araújo 0001, Teresa Bernarda Ludermir, Alexander Schliep
BMC Bioinform.2
2007 Semi-supervised learning for the identification of syn-expressed genes from fused microarray and in situ image data
abstract
BACKGROUND: Gene expression measurements during the development of the fly Drosophila melanogaster are routinely used to find functional modules of temporally co-expressed genes. Complimentary large data sets of in situ RNA hybridization images for different stages of the fly embryo elucidate the spatial expression patterns. RESULTS: Using a semi-supervised approach, constrained clustering with mixture models, we can find clusters of genes exhibiting spatio-temporal similarities in expression, or syn-expression. The temporal gene expression measurements are taken as primary data for which pairwise constraints are computed in an automated fashion from raw in situ images without the need for manual annotation. We investigate the influence of these pairwise constraints in the clustering and discuss the biological relevance of our results. CONCLUSION: Spatial information contributes to a detailed, biological meaningful analysis of temporal gene expression data. Semi-supervised learning provides a flexible, robust and efficient framework for integrating data sources of differing quality and abundance.
Ivan G. Costa, Roland Krause, Lennart Opitz, Alexander Schliep
BMC Bioinform.1
2005 The Graphical Query Language: a tool for analysis of gene expression time-courses
abstract
UNLABELLED: The Graphical Query Language (GQL) is a set of tools for the analysis of gene expression time-courses. They allow a user to pre-process the data, to query it for interesting patterns, to perform model-based clustering or mixture estimation, to include subsequent refinements of clusters and, finally, to use other biological resources to evaluate the results. Analyses are carried out in a graphical and interactive environment, allowing expert intervention in all stages of the data analysis. AVAILABILITY: The GQL package is freely available under the GNU general public license (GPL) at http://www.ghmm.org/gql
Ivan G. Costa, Alexander Schönhuth, Alexander Schliep
Bioinform.1
2005 Analyzing Gene Expression Time-Courses
abstract
Measuring gene expression over time can provide important insights into basic cellular processes. Identifying groups of genes with similar expression time-courses is a crucial first step in the analysis. As biologically relevant groups frequently overlap, due to genes having several distinct roles in those cellular processes, this is a difficult problem for classical clustering methods. We use a mixture model to circumvent this principal problem, with hidden Markov models (HMMs) as effective and flexible components. We show that the ensuing estimation problem can be addressed with additional labeled data-partially supervised learning of mixtures-through a modification of the Expectation-Maximization (EM) algorithm. Good starting points for the mixture estimation are obtained through a modification to Bayesian model merging, which allows us to learn a collection of initial HMMs. We infer groups from mixtures with a simple information-theoretic decoding heuristic, which quantifies the level of ambiguity in group assignment. The effectiveness is shown with high-quality annotation data. As the HMMs we propose capture asynchronous behavior by design, the groups we find are also asynchronous. Synchronous subgroups are obtained from a novel algorithm based on Viterbi paths. We show the suitability of our HMM mixture approach on biological and simulated data and through the favorable comparison with previous approaches. A software implementing the method is freely available under the GPL from http://ghmm.org/gql.
Alexander Schliep, Ivan G. Costa, Christine Steinhoff, Alexander Schönhuth
IEEE ACM Trans. Comput. Biol. Bioinform.2