Sampsa Hautaniemi

dblp:53/2962 · DBLP profile ↗
← Back
27ranked-venue papers
3as first author
8since 2021 · last 2026
0000-0002-7749-2694ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 25 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 FUSE: data-driven functional segmentation of DNA methylation data
abstract
SUMMARY: DNA methylation (DNAm) of neighbouring CpG sites is highly correlated, making DNAm function in terms of blocks. DNAm patterns and functionality are linked to both chromatin structure of DNA and gene regulation. Defining biologically meaningful DNA methylation blocks from whole-genome bisulfite sequencing (WGBS) data remains challenging, as most existing methods rely on fixed genomic windows rather than the observed methylation pattern. We present FUSE, a data-driven segmentation method that captures intrinsic methylation segments directly from WGBS data by jointly analyzing multiple samples. FUSE identifies spatially homogeneous methylation blocks shared across the input cohort while allowing different methylation states across samples. Applied to 61 WGBS samples from the ENCODE database, FUSE identified segments which overlap significantly with promoters, enhancers, and repetitive elements. FUSE was able to recover the true segment breakpoints in synthetic data with high sensitivity under increased levels of noise. As such, FUSE facilitates post hoc methylation analyses by aggregating coherent CpG sites into candidate segments for downstream differential methylation testing or other comparative studies. AVAILABILITY AND IMPLEMENTATION: FUSE is implemented as an R-package methFuse, available at https://github.com/holmsusa/methFuse and https://cran.r-project.org/package=methFuse. A GenomeSpy visualization of the data is available at https://csbi.ltdk.helsinki.fi/p/fuse_encode_gs/.
Susanna Holmström, Antti Häkkinen, Kari Lavikka, Giovanni Marchi, Sampsa Hautaniemi, Alexandra Lahtinen
Bioinform.5
2025 Jellyfish: integrative visualization of spatio-temporal tumor evolution and clonal dynamics
abstract
SUMMARY: Spatial and temporal intra-tumor heterogeneity drives tumor evolution and therapy resistance. Existing visualization tools often fail to capture both dimensions simultaneously. To address this, we developed Jellyfish, a tool that integrates phylogenetic and sample trees into a single plot, providing a holistic view of tumor evolution and capturing both spatial and temporal evolution. Available as a JavaScript library and R package, Jellyfish generates interactive visualizations from tumor phylogeny and clonal composition data. We demonstrate its ability to visualize complex subclonal dynamics using data from ovarian high-grade serous carcinoma. AVAILABILITY AND IMPLEMENTATION: Jellyfish is freely available with MIT license at https://github.com/HautaniemiLab/jellyfish (JavaScript library) and https://github.com/HautaniemiLab/jellyfisher (R package).
Kari Lavikka, Altti Ilari Maarala, Jaana Oikkonen, Sampsa Hautaniemi
Bioinform.4
2024 ECCB2024: The 23rd European Conference on Computational Biology
abstract
This volume of Bioinformatics includes the proceedings papers of the 23rd European Conference on Computational Biology (ECCB2024) to be held in Turku, Finland, from 16 September to 20 September 2024, under the theme Data and Algorithms for Health and Science. More information on the ECCB2024 conference is available at the conference website https://eccb2024.fi/. ECCB is one of the main international conferences in the field of computational biology and bioinformatics together with the Intelligent Systems for Molecular Biology (ISMB) and the Research in Computational Molecular Biology (RECOMB). It is held jointly with the ISMB conference in odd-numbered years and independently in even-numbered years. ECCB attracts scientists and industry professionals from diverse disciplines, including mathematics, statistics, computer science, biology, and medicine. Rapid technological advancements enable life scientists to gain increasingly detailed insights in complex biological systems. These improvements in measurement technologies, however, introduce new challenges for data interpretation and necessitate the development of advanced computational techniques to manage the complexity and volume of the data. Consequently, the field of computational biology is rapidly evolving with new algorithms, software, and databases. The ECCB2024 conference showcases cutting-edge developments in systems biology, artificial intelligence, single-cell and spatial technologies, data integration, and more, addressing the growing demand for sophisticated algorithms to enhance the analysis of large-scale biological and biomedical datasets. This is highlighted by the most frequent keywords accompanying all the accepted submissions in ECCB2024 (tutorials/workshops, proceedings, highlight talks, posters), with the most popular keywords including ‘machine learning’, ‘deep learning’, ‘single-cell rnaseq’, and ‘multiomics’ (Fig. 1). Most frequent keywords among all accepted submissions in ECCB2024. The barplot shows the frequencies of top keywords appearing in at least ten submissions, while the word cloud illustrates the relative frequencies of all keywords appearing in at least five submissions. The ECCB2024 edition features five keynote lectures by distinguished speakers: Sarah Teichmann (Cambridge Stem Cell Institute, University of Cambridge, UK), Peer Bork (EMBL—European Molecular Biology Laboratory, Germany), Ileana Cristea (Princeton University, US), Jussi Taipale (University of Cambridge, UK), and Fabian Theis (Helmholtz Munich Computational Health Center, Germany). In addition, ECCB2024 hosts a scientific debate on data sharing and privacy protection by Melissa Haendel (University of North Carolina, US) and Yves Moreau (KU Leuven, Belgium). The ECCB2024 conference covers a wide range of topics, focusing on methodological advancements in computational biology as well as innovative application of computational techniques to life sciences and medicine. To provide a cohesive overview of recent scientific progress, the conference presentations are organized under six broad themes: (i) Genomes, (ii) Proteins, (iii) Systems biology and multiomics, (iv) Single-cell omics, (v) Microbiomes and planetary health, and (vi) Digital health. The proceedings talks present new scientific contributions, while the highlight talks showcase already published cutting-edge science in computational biology and related further developments. Poster presentations provide an opportunity for the participants to discuss their recent work with other researchers in the field. Workshops and tutorials on specialized topics prior to the main conference program are platforms to share practical experiences and learn new skills. In addition to scientific presentation tracks, ECCB2024 has a separate ELIXIR track, overseen by ELIXIR, focusing on advancements in infrastructure and services within ELIXIR nodes in support of the expert groups known as ‘ELIXIR Communities’. ELIXIR is a distributed pan-European life science infrastructure for biological data that coordinates, integrates, and sustains bioinformatics resources across its member states. This enables academic and industry users to access data, tools, standards, computing, and training services for life science research. ELIXIR has selected ECCB as a primary dissemination platform, serving as a co-organizing sponsor. Apart from Community activities the track will also introduce three scientific areas as outlined in the 2024–2028 ELIXIR Scientific Programme (https://elixir-europe.org), enabling scientists to access and analyse life science data across ‘Cellular and Molecular Research’, ‘Biodiversity, food security & pathogens’, and ‘Human data & translational research’. ECCB2024 received an impressive number of 200 submissions of full manuscripts for the proceedings call, highlighting the growing interest and engagement within the community. These submissions were organized under the six conference themes. Each submission was subjected to a peer-review process, with at least two reviews per manuscript, managed by the members of the ECCB2024 Programme Committee. The Programme Committee was chaired by Laura Elo (University of Turku, Finland) and each conference theme had an area chair (Table 1). The Proceedings Review Committee had over 100 reviewers. Thematic areas of ECCB2024 proceedings talks.a The table lists the area chairs for each theme, the number of reviewed papers, the number of accepted papers, and the acceptance rate for each theme. Thematic areas of ECCB2024 proceedings talks.a The table lists the area chairs for each theme, the number of reviewed papers, the number of accepted papers, and the acceptance rate for each theme. The review process focused on the impact, reproducibility, and scientific quality of the submitted research, as well as its relevance, interest, and value for the ECCB2024 audience. Upon completion of the review, the Program Committee chairs selected 24 papers to be included in the ECCB2024 proceedings, with an acceptance rate of 12% across the different areas (Table 1). All proceedings papers are published open access in this Proceedings issue of September 2024 of Bioinformatics. The ECCB2024 call for Highlight Talks invited presentations of studies recently published in scientific journals since 1 March 2023, or accepted for publication. The existence of further developments related to the published paper was also positively considered. The call received a total of 212 proposals, which were ranked according to their relevance and impact on computational biology, as well as their suitability to be presented to the large and diverse audience. Ultimately, the Program Committee selected 23% of the submissions for presentation as a Highlight talk at the conference. ECCB2024 hosts two poster sessions where researchers can introduce and discuss their work under the conference themes. In total, nearly 700 posters were accepted to be presented. The poster submissions were reviewed by the Posters Committee: Sini Junttila, Asta Laiho, Tomi Suomi, and Sampsa Hautaniemi (late posters). In addition, the ELIXIR track selected posters to be presented at ECCB2024. Overall, the ECCB2024 tracks attracted participation of presenting authors from 48 countries. The countries with the highest number of presenting authors were Germany, Finland, the USA, the United Kingdom, and Spain, jointly contributing to half of the overall number (Fig. 2). Following these were Turkey, France, and China, each contributing 4%–5% of the presenting authors. Proportion of countries where the presenting author of each accepted submission had their primary affiliation. Countries with the proportion below 2% were grouped under ‘Other’. Exhibitor booths are open throughout the conference, including ELIXIR, ISCB, Oxford University Press, Royal Society Publishing and eLife Sciences Publications Ltd, as well as the two organizing institutions: University of Turku and CSC—IT Center for Science. The booths showcase the latest scientific literature in computational biology and bioinformatics, data stewardship, as well as new developments in hardware, software, and technology. In addition, the Visit Turku Archipelago (tourist office) is also represented. Before the main ECCB2024 conference, a satellite meeting, seven workshops, and nine tutorials are organized. The satellite meeting is organized by the Student Council of the International Society for Computational Biology (ISCB) by and for early-stage researchers. This 8th European Student Council Symposium (ESCS) continues the successful collaboration between ESCS and ECCB, building the future of research in computational biology. The 16 workshops/tutorials preceding the ECCB2024 main conference were selected out of a total of 24 proposals by the workshops and tutorials chair Bengt Persson (Uppsala University, Sweden), each running for half a day. The workshops foster discussions and exchange of ideas on a range of specialized or emerging topics in computational biology and provide opportunities to share practical experiences. The tutorials provide participants with lectures and hands-on training to enable learning about new areas of computational biology or important established topics. Following the tradition of previous ECCB conferences, the ISCB and ECCB sponsored a number of travel fellowships for students and postdoctoral fellows. These fellowships were primarily awarded to members presenting talks and those from low- or middle-income countries, facilitating their participation in the conference. A total of 68 applications were received from scientists across 26 countries. After a careful review, the ECCB2024 Organizing Committee in collaboration with the ISCB and ECCB awarded 10 and 18 fellowships, respectively. The ECCB2024 Code of Conduct is designed to ensure a safe and respectful environment for all attendees, outlining clear standards of behaviour to foster an inclusive conference atmosphere. In addition, the code provides specific procedures for attendees to follow if they feel these standards have been violated. This ensures that all participants have access to support and resources to address any concerns, reinforcing the commitment of ECCB2024 to uphold the highest ethical and professional standards during the conference. The ECCB2024 Organizing Committee is strongly committed to promote gender equity throughout the conference planning and execution. In particular, an important effort was made to ensure that the review panels included a balanced composition of female and male professionals across the world. In addition, three out of the seven distinguished keynote speakers are women. The conference program also features a collaborative workshop with the Bioinfo4Women initiative, discussing sex and gender bias in artificial intelligence. We would like to thank everyone who has contributed to the success of ECCB2024, ensuring it meets high standards of excellence. Special recognition is due to the Program Committee, including theme area chairs and reviewers, whose critical and dedicated efforts have been fundamental to the conference organization in selecting excellent tutorials, workshops, manuscripts, talks, and posters. We are equally thankful to the ECCB steering committee for their invaluable support and advice, especially ECCB2022 organizers for providing detailed information about the organization of the previous ECCB conference. The collaboration with the ISCB has been crucial in offering travel fellowships and promoting ECCB2024 globally, for which we are deeply grateful. Our gratitude also goes to all our financial sponsors, including our co-organizing sponsor ELIXIR. We also acknowledge the Oxford University Press production team for their work on the ECCB2024 Proceedings issue. Many individuals have played significant roles in the local organization, often exceeding their responsibilities, and we are immensely thankful for their dedication. Lastly, the conference would not be what it is without the diverse participants from around the world, whose scientific contributions, presentations, and discussions enrich ECCB2024. Thank you all for being there and for allowing us to enjoy science at ECCB2024 in Turku! None declared. This paper was published as part of a supplement financially supported by ECCB2024. All data are incorporated into the article. Additional information is available at the conference website https://eccb2024.fi.
Anu Kukkonen-Macchi, Sampsa Hautaniemi, Katharina F. Heil, Merja Heinäniemi, Lars Juhl Jensen, Sini Junttila, Lukas Käll, Asta Laiho, Peter Maccallum, Matti Nykter, Bengt Persson, Tomi Suomi, Tim Van Den Bossche, Tommi H. Nyrönen, Laura Elo
Bioinform.2
2022 POIBM: batch correction of heterogeneous RNA-seq datasets through latent sample matching
abstract
MOTIVATION: RNA sequencing and other high-throughput technologies are essential in understanding complex diseases, such as cancers, but are susceptible to technical factors manifesting as patterns in the measurements. These batch patterns hinder the discovery of biologically relevant patterns. Unbiased batch effect correction in heterogeneous populations currently requires special experimental designs or phenotypic labels, which are not readily available for patient samples in existing datasets. RESULTS: We present POIBM, an RNA-seq batch correction method, which learns virtual reference samples directly from the data. We use a breast cancer cell line dataset to show that POIBM exceeds or matches the performance of previous methods, while being blind to the phenotypes. Further, we analyze The Cancer Genome Atlas RNA-seq data to show that batch effects plague many cancer types; POIBM effectively discovers the true replicates in stomach adenocarcinoma; and integrating the corrected data in endometrial carcinoma improves cancer subtyping. AVAILABILITY AND IMPLEMENTATION: https://bitbucket.org/anthakki/poibm/ (archived at https://doi.org/10.5281/zenodo.6122436). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Susanna Holmström, Sampsa Hautaniemi, Antti Häkkinen
Bioinform.2
2021 Network-guided identification of cancer-selective combinatorial therapies in ovarian cancer
abstract
Each patient's cancer consists of multiple cell subpopulations that are inherently heterogeneous and may develop differing phenotypes such as drug sensitivity or resistance. A personalized treatment regimen should therefore target multiple oncoproteins in the cancer cell populations that are driving the treatment resistance or disease progression in a given patient to provide maximal therapeutic effect, while avoiding severe co-inhibition of non-malignant cells that would lead to toxic side effects. To address the intra- and inter-tumoral heterogeneity when designing combinatorial treatment regimens for cancer patients, we have implemented a machine learning-based platform to guide identification of safe and effective combinatorial treatments that selectively inhibit cancer-related dysfunctions or resistance mechanisms in individual patients. In this case study, we show how the platform enables prediction of cancer-selective drug combinations for patients with high-grade serous ovarian cancer using single-cell imaging cytometry drug response assay, combined with genome-wide transcriptomic and genetic profiles. The platform makes use of drug-target interaction networks to prioritize those combinations that warrant further preclinical testing in scarce patient-derived primary cells. During the case study in ovarian cancer patients, we investigated (i) the relative performance of various ensemble learning algorithms for drug response prediction, (ii) the use of matched single-cell RNA-sequencing data to deconvolute cell population-specific transcriptome profiles from bulk RNA-seq data, (iii) and whether multi-patient or patient-specific predictive models lead to better predictive accuracy. The general platform and the comparison results are expected to become useful for future studies that use similar predictive approaches also in other cancer types.
Liye He, Daria Bulanova, Jaana Oikkonen, Antti Häkkinen, Kaiyang Zhang, Erdogan Pekcan Erkan, Olli Carpén, Titta Joutsiniemi, Sakari Hietanen, Johanna Hynninen, Kaisa Huhtinen, Sampsa Hautaniemi, Anna Vähärautio, Jing Tang 0002, Krister Wennerberg, Tero Aittokallio
Briefings Bioinform.14
2021 Agile workflow for interactive analysis of mass cytometry data
abstract
MOTIVATION: Single-cell proteomics technologies, such as mass cytometry, have enabled characterization of cell-to-cell variation and cell populations at a single-cell resolution. These large amounts of data, require dedicated, interactive tools for translating the data into knowledge. RESULTS: We present a comprehensive, interactive method called Cyto to streamline analysis of large-scale cytometry data. Cyto is a workflow-based open-source solution that automates the use of state-of-the-art single-cell analysis methods with interactive visualization. We show the utility of Cyto by applying it to mass cytometry data from peripheral blood and high-grade serous ovarian cancer (HGSOC) samples. Our results show that Cyto is able to reliably capture the immune cell sub-populations from peripheral blood and cellular compositions of unique immune- and cancer cell subpopulations in HGSOC tumor and ascites samples. AVAILABILITYAND IMPLEMENTATION: The method is available as a Docker container at https://hub.docker.com/r/anduril/cyto and the user guide and source code are available at https://bitbucket.org/anduril-dev/cyto. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Julia Casado, Oskari Lehtonen, Ville Rantanen, Katja Kaipio, Luca Pasquini, Antti Häkkinen, Eleonora Petrucci, Johanna Hynninen, Sakari Hietanen, Olli Carpén, Mauro Biffoni, Anniina Färkkilä, Sampsa Hautaniemi
Bioinform.13
2021 FUNGI: FUsioN Gene Integration toolset
abstract
MOTIVATION: Fusion genes are both useful cancer biomarkers and important drug targets. Finding relevant fusion genes is challenging due to genomic instability resulting in a high number of passenger events. To reveal and prioritize relevant gene fusion events we have developed FUsionN Gene Identification toolset (FUNGI) that uses an ensemble of fusion detection algorithms with prioritization and visualization modules. RESULTS: We applied FUNGI to an ovarian cancer dataset of 107 tumor samples from 36 patients. Ten out of 11 detected and prioritized fusion genes were validated. Many of detected fusion genes affect the PI3K-AKT pathway with potential role in treatment resistance. AVAILABILITYAND IMPLEMENTATION: FUNGI and its documentation are available at https://bitbucket.org/alejandra_cervera/fungi as standalone or from Anduril at https://www.anduril.org. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Alejandra Cervera, Heidi Rausio, Tiia Kähkönen, Noora Andersson, Gabriele Partel, Ville Rantanen, Giulia Paciello, Elisa Ficarra, Johanna Hynninen, Sakari Hietanen, Olli Carpén, Rainer Lehtonen, Sampsa Hautaniemi, Kaisa Huhtinen
Bioinform.13
2021 PRISM: recovering cell-type-specific expression profiles from individual composite RNA-seq samples
abstract
MOTIVATION: A major challenge in analyzing cancer patient transcriptomes is that the tumors are inherently heterogeneous and evolving. We analyzed 214 bulk RNA samples of a longitudinal, prospective ovarian cancer cohort and found that the sample composition changes systematically due to chemotherapy and between the anatomical sites, preventing direct comparison of treatment-naive and treated samples. RESULTS: To overcome this, we developed PRISM, a latent statistical framework to simultaneously extract the sample composition and cell-type-specific whole-transcriptome profiles adapted to each individual sample. Our results indicate that the PRISM-derived composition-free transcriptomic profiles and signatures derived from them predict the patient response better than the composite raw bulk data. We validated our findings in independent ovarian cancer and melanoma cohorts, and verified that PRISM accurately estimates the composition and cell-type-specific expression through whole-genome sequencing and RNA in situ hybridization experiments. AVAILABILITYAND IMPLEMENTATION: https://bitbucket.org/anthakki/prism. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Antti Häkkinen, Kaiyang Zhang, Amjad Alkodsi, Noora Andersson, Erdogan Pekcan Erkan, Katja Kaipio, Tarja Lamminen, Naziha Mansuri, Kaisa Huhtinen, Anna Vähärautio, Olli Carpén, Johanna Hynninen, Sakari Hietanen, Rainer Lehtonen, Sampsa Hautaniemi
Bioinform.16
2020 qSNE: quadratic rate t-SNE optimizer with automatic parameter tuning for large datasets
abstract
MOTIVATION: Non-parametric dimensionality reduction techniques, such as t-distributed stochastic neighbor embedding (t-SNE), are the most frequently used methods in the exploratory analysis of single-cell datasets. Current implementations scale poorly to massive datasets and often require downsampling or interpolative approximations, which can leave less-frequent populations undiscovered and much information unexploited. RESULTS: We implemented a fast t-SNE package, qSNE, which uses a quasi-Newton optimizer, allowing quadratic convergence rate and automatic perplexity (level of detail) optimizer. Our results show that these improvements make qSNE significantly faster than regular t-SNE packages and enables full analysis of large datasets, such as mass cytometry data, without downsampling. AVAILABILITY AND IMPLEMENTATION: Source code and documentation are openly available at https://bitbucket.org/anthakki/qsne/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Antti Häkkinen, Juha Koiranen, Julia Casado, Katja Kaipio, Oskari Lehtonen, Eleonora Petrucci, Johanna Hynninen, Sakari Hietanen, Olli Carpén, Luca Pasquini, Mauro Biffoni, Rainer Lehtonen, Sampsa Hautaniemi
Bioinform.13
2019 Anduril 2: upgraded large-scale data integration framework
abstract
SUMMARY: Anduril is an analysis and integration framework that facilitates the design, use, parallelization and reproducibility of bioinformatics workflows. Anduril has been upgraded to use Scala for pipeline construction, which simplifies software maintenance, and facilitates design of complex pipelines. Additionally, Anduril's bioinformatics repository has been expanded with multiple components, and tutorial pipelines, for next-generation sequencing data analysis. AVAILABILITYAND IMPLEMENTATION: Freely available at http://anduril.org. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Alejandra Cervera, Ville Rantanen, Kristian Ovaska, Marko Laakso, Javier Nuñez-Fontarnau, Amjad Alkodsi, Julia Casado, Chiara Facciotto, Antti Häkkinen, Riku Louhimo, Sirkku Karinen, Kaiyang Zhang, Kari Lavikka, Lauri Lyly, Sampsa Hautaniemi
Bioinform.16
2018 Identifying differentially methylated sites in samples with varying tumor purity
abstract
Motivation: DNA methylation aberrations are common in many cancer types. A major challenge hindering comparison of patient-derived samples is that they comprise of heterogeneous collection of cancer and microenvironment cells. We present a computational method that allows comparing cancer methylomes in two or more heterogeneous tumor samples featuring differing, unknown fraction of cancer cells. The method is unique in that it allows comparison also in the absence of normal cell control samples and without prior tumor purity estimates, as these are often unavailable or unreliable in clinical samples. Results: We use simulations and next-generation methylome, RNA and whole-genome sequencing data from two cancer types to demonstrate that the method is accurate and outperforms alternatives. The results show that our method adapts well to various cancer types and to a wide range of tumor content, and works robustly without a control or with controls derived from various sources. Availability and implementation: The method is freely available at https://bitbucket.org/anthakki/dmml. Supplementary information: Supplementary data are available at Bioinformatics online.
Antti Häkkinen, Amjad Alkodsi, Chiara Facciotto, Kaiyang Zhang, Katja Kaipio, Sirpa Leppä, Olli Carpén, Seija Grénman, Johanna Hynninen, Sakari Hietanen, Rainer Lehtonen, Sampsa Hautaniemi
Bioinform.12
2018 PerPAS: Topology-Based Single Sample Pathway Analysis Method
abstract
Identification of intracellular pathways that play key roles in cancer progression and drug resistance is a prerequisite for developing targeted cancer treatments. The era of personalized medicine calls for computational methods that can function with one sample or a very small set of samples. Developing such methods is challenging because standard statistical approaches pose several limiting assumptions, such as number of samples, that prevent their application when approaches to one. We have developed a novel pathway analysis method called PerPAS to estimate pathway activity at a single sample level by integrating pathway topology and transcriptomics data. In addition, PerPAS is able to identify altered pathways between cancer and control samples as well as to identify key nodes that contribute to the pathway activity. In our case study using breast cancer data, we show that PerPAS can identify highly altered pathways that are associated with patient survival. PerPAS identified four pathways that were associated with patient survival and were successfully validated in three independent breast cancer cohorts. In comparison to two other pathway analysis methods that function at a single sample level, PerPAS had superior performance in both synthetic and breast cancer expression datasets. PerPAS is a free R package (http://csbi.ltdk.helsinki.fi/pub/czliu/perpas/).
Chengyu Liu 0002, Rainer Lehtonen, Sampsa Hautaniemi
IEEE ACM Trans. Comput. Biol. Bioinform.3
2015 Comparative analysis of methods for identifying somatic copy number alterations from deep sequencing data
abstract
Somatic copy-number alterations (SCNAs) are an important type of structural variation affecting tumor pathogenesis. Accurate detection of genomic regions with SCNAs is crucial for cancer genomics as these regions contain likely drivers of cancer development. Deep sequencing technology provides single-nucleotide resolution genomic data and is considered one of the best measurement technologies to detect SCNAs. Although several algorithms have been developed to detect SCNAs from whole-genome and whole-exome sequencing data, their relative performance has not been studied. Here, we have compared ten SCNA detection algorithms in both simulated and primary tumor deep sequencing data. In addition, we have evaluated the applicability of exome sequencing data for SCNA detection. Our results show that (i) clear differences exist in sensitivity and specificity between the algorithms, (ii) SCNA detection algorithms are able to identify most of the complex chromosomal alterations and (iii) exome sequencing data are suitable for SCNA detection.
Amjad Alkodsi, Riku Louhimo, Sampsa Hautaniemi
Briefings Bioinform.3
2013 Integrative Analysis of Deep Sequencing Data Identifies Estrogen Receptor Early Response Genes and Links ATAD3B to Poor Survival in Breast Cancer
abstract
Identification of responsive genes to an extra-cellular cue enables characterization of pathophysiologically crucial biological processes. Deep sequencing technologies provide a powerful means to identify responsive genes, which creates a need for computational methods able to analyze dynamic and multi-level deep sequencing data. To answer this need we introduce here a data-driven algorithm, SPINLONG, which is designed to search for genes that match the user-defined hypotheses or models. SPINLONG is applicable to various experimental setups measuring several molecular markers in parallel. To demonstrate the SPINLONG approach, we analyzed ChIP-seq data reporting PolII, estrogen receptor α (ERα), H3K4me3 and H2A.Z occupancy at five time points in the MCF-7 breast cancer cell line after estradiol stimulus. We obtained 777 ERa early responsive genes and compared the biological functions of the genes having ERα binding within 20 kb of the transcription start site (TSS) to genes without such binding site. Our results show that the non-genomic action of ERα via the MAPK pathway, instead of direct ERa binding, may be responsible for early cell responses to ERα activation. Our results also indicate that the ERα responsive genes triggered by the genomic pathway are transcribed faster than those without ERα binding sites. The survival analysis of the 777 ERα responsive genes with 150 primary breast cancer tumors and in two independent validation cohorts indicated the ATAD3B gene, which does not have ERα binding site within 20 kb of its TSS, to be significantly associated with poor patient survival.
Kristian Ovaska, Filomena Matarese, Korbinian Grote, Iryna Charapitsa, Alejandra Cervera, Chengyu Liu 0002, George Reid, Martin Seifert, Hendrik G. Stunnenberg, Sampsa Hautaniemi
PLoS Comput. Biol.10
2013 Genomic Region Operation Kit for Flexible Processing of Deep Sequencing Data
abstract
Computational analysis of data produced in deep sequencing (DS) experiments is challenging due to large data volumes and requirements for flexible analysis approaches. Here, we present a mathematical formalism based on set algebra for frequently performed operations in DS data analysis to facilitate translation of biomedical research questions to language amenable for computational analysis. With the help of this formalism, we implemented the Genomic Region Operation Kit (GROK), which supports various DS-related operations such as preprocessing, filtering, file conversion, and sample comparison. GROK provides high-level interfaces for R, Python, Lua, and command line, as well as an extension C++ API. It supports major genomic file formats and allows storing custom genomic regions in efficient data structures such as red-black trees and SQL databases. To demonstrate the utility of GROK, we have characterized the roles of two major transcription factors (TFs) in prostate cancer using data from 10 DS experiments. GROK is freely available with a user guide from >http://csbi.ltdk.helsinki.fi/grok/.
Kristian Ovaska, Lauri Lyly, Biswajyoti Sahu, Olli A. Jänne, Sampsa Hautaniemi
IEEE ACM Trans. Comput. Biol. Bioinform.5
2011 CNAmet: an R package for integrating copy number, methylation and expression data
abstract
SUMMARY: Gene copy number and DNA methylation alterations are key regulators of gene expression in cancer. Accordingly, genes that show simultaneous methylation, copy number and expression alterations are likely to have a key role in tumor progression. We have implemented a novel software package (CNAmet) for integrative analysis of high-throughput copy number, DNA methylation and gene expression data. To demonstrate the utility of CNAmet, we use copy number, DNA methylation and gene expression data from 50 glioblastoma multiforme and 188 ovarian cancer primary tumor samples. Our results reveal a synergistic effect of DNA methylation and copy number alterations on gene expression for several known oncogenes as well as novel candidate oncogenes. AVAILABILITY: CNAmet R-package and user guide are freely available under GNU General Public License at http://csbi.ltdk.helsinki.fi/CNAmet.
Riku Louhimo, Sampsa Hautaniemi
Bioinform.2
2010 Integrative platform to translate gene sets to networks
abstract
SUMMARY: We have implemented a computational platform (Moksiskaan) that integrates pathway, protein-protein interaction, genome and literature mining data to result in comprehensive networks for a list of genes or proteins. Moksiskaan is able to generate hypothetical pathways for these genes or proteins as well as estimate their activation statuses using regulation information in pathway repositories. An automatically generated result document provides a detailed description of the query genes, biological processes and drug targets. Moksiskaan networks can be downloaded to Cytoscape for further analysis. To demonstrate the utility of Moksiskaan, we use gene microarray and clinical data from >200 glioblastoma multiforme primary tumor samples and translate the resulting set of 124 survival-associated genes to a network. AVAILABILITY AND IMPLEMENTATION: Moksiskaan and user guide are freely available under GNU General Public License at http://csbi.ltdk.helsinki.fi/moksiskaan/
Marko Laakso, Sampsa Hautaniemi
Bioinform.2
2009 Comparison of Affymetrix data normalization methods using 6, 926 experiments across five array generations
abstract
BACKGROUND: Gene expression microarray technologies are widely used across most areas of biological and medical research. Comparing and integrating microarray data from different experiments would be very useful, but is currently very challenging due to the experimental and hybridization conditions, as well as data preprocessing and normalization methods. Furthermore, even in the case of the widely-used, industry-standard Affymetrix oligonucleotide microarrays, the various array generations have different probe sets representing different genes, hindering the data integration. RESULTS: In this study our objective is to find systematic approaches to normalize the data emerging from different Affymetrix array generations and from different laboratories. We compare and assess the accuracy of five normalization methods for Affymetrix gene expression data using 6,926 Affymetrix experiments from five array generations. The methods that we compare include 1) standardization, 2) housekeeping gene based normalization, 3) equalized quantile normalization, 4) Weibull distribution based normalization and 5) array generation based gene centering. Our results indicate that the best results are achieved when the data is normalized first within a sample and then between-samples with Array Generation based gene Centering (AGC) normalization. CONCLUSION: We conclude that with the AGC method integrating different Affymetrix datasets results in values that are significantly more comparable across the array generations than in the cases where no array generation based normalization is used. The AGC method was found to be the best method for normalizing the data from several different array generations, and achieve comparable gene values across thousands of samples.
Reija Autio, Sami Kilpinen, Matti Saarela, Olli-P. Kallioniemi, Sampsa Hautaniemi, Jaakko Astola
BMC Bioinform.5
2009 Advanced analysis and visualization of gene copy number and expression data
abstract
BACKGROUND: Gene copy number and gene expression values play important roles in cancer initiation and progression. Both can be measured with high-throughput microarrays and some methodologies to integrate and analyze these data exist. However, varying gene sets within different gene expression and copy number microarrays present significant challenges. RESULTS: We report an advanced version of earlier published CGH-Plotter that rapidly can identify amplified and deleted areas using gene copy number data. With CGH-Plotter v2, the copy number values can be filtered based on the genomic location in basepair units. After filtering, the values for the missing genes can be interpolated. Moreover, the effect of non-informative areas in the genome can be systematically removed by smoothing and interpolating. Further, we developed a tool (ECN) to illustrate the CGH-data values annotated based on the gene expression. The ECN-tool is a MATLAB toolbox enabling straightforward illustration of copy numbers annotated based on the gene expression levels. CONCLUSION: CGH-Plotter v2 provides two methods for analyzing copy number data; dynamic programming and genomic location based smoothing. With ECN-tool the data analyzed with CGH-Plotter v2 can easily be illustrated along the chromosomes individually or along the whole genome. ECN-tool plots the copy number data annotated based on the gene expression data, and it is easy to find the genes that are both over-expressed and amplified or under-expressed and deleted in the samples. From the resulting figures it is straightforward to select interesting genes.
Reija Autio, Matti Saarela, Anna-Kaarina Järvinen, Sampsa Hautaniemi, Jaakko Astola
BMC Bioinform.4
2007 Computational identification of candidate loci for recessively inherited mutation using high-throughput SNP arrays
abstract
MOTIVATION: Single nucleic polymorphisms (SNPs) are one of the most abundant genetic variations in the human genome. Recently, several platforms for high-throughput SNP analysis have become available, capable of measuring thousands of SNPs across the genome. Tools for analysing and visualizing these large genetic data sets in biologically relevant manner are rare. This hinders effective use of the SNP-array data in research on complex diseases, such as cancer. RESULTS: We describe a computational framework to analyse and visualize SNP-array data, and link the results in relevant databases. Our major objective is to develop methods for identifying DNA regions that likely harbour recessive mutations. Thus, the algorithms are designed to have high sensitivity and the identified regions are ranked using a scoring algorithm. We have also developed annotation tools that automatically query gene IDs, exon counts, microarray probe IDs, etc. In our case study, we apply the methods for identifying candidate regions for recessively inherited colorectal cancer predisposition and suggest directions for wet-lab experiments. AVAILABILITY: R-package implementation is available at http://www.ltdk.helsinki.fi/sysbio/csb/downloads/CohortComparator/
Marko Laakso, Sari Tuupanen, Auli Karhu, Rainer Lehtonen, Lauri A. Aaltonen, Sampsa Hautaniemi
Bioinform.6
2006 Relationships between probabilistic Boolean networks and dynamic Bayesian networks as models of gene regulatory networks
Harri Lähdesmäki, Sampsa Hautaniemi, Ilya Shmulevich, Olli Yli-Harja
Signal Process.2
2006 Jointly Analyzing Gene Expression and Copy Number Data in Breast Cancer Using Data Reduction Models
abstract
With the growing surge of biological measurements, the problem of integrating and analyzing different types of genomic measurements has become an immediate challenge for elucidating events at the molecular level. In order to address the problem of integrating different data types, we present a framework that locates variation patterns in two biological inputs based on the generalized singular value decomposition (GSVD). In this work, we jointly examine gene expression and copy number data and iteratively project the data on different decomposition directions defined by the projection angle theta in the GSVD. With the proper choice of theta, we locate similar and dissimilar patterns of variation between both data types. We discuss the properties of our algorithm using simulated data and conduct a case study with biologically verified results. Ultimately, we demonstrate the efficacy of our method on two genome-wide breast cancer studies to identify genes with large variation in expression and copy number across numerous cell line and tumor samples. Our method identifies genes that are statistically significant in both input measurements. The proposed method is useful for a wide variety of joint copy number and expression-based studies. Supplementary information is available online, including software implementations and experimental data.
John A. Berger, Sampsa Hautaniemi, Sanjit K. Mitra, Jaakko Astola
IEEE ACM Trans. Comput. Biol. Bioinform.2
2005 Modeling of signal-response cascades using decision tree analysis
abstract
MOTIVATION: Signal transduction cascades governing cell functional responses to stimulatory cues play crucial roles in cell regulatory systems and represent promising therapeutic targets for complex human diseases. however, mathematical analysis of how cell responses are governed by signaling activities is challenging due to their multivariate and non-linear nature. diverse computational methods are potentially available, but most are ineffective for protein-level data that is limited in extent and replication. RESULTS: We apply a decision tree approach to analyze the relationship of cell functional response to signaling activity across a spectrum of stimulatory cues. as a specific example, we studied five intracellular signals influencing fibroblast migration under eight conditions: four substratum fibronectin levels and presence versus absence of epidermal growth factor. we propose techniques for preprocessing and extending the experimental measurement set via interpolative modeling in order to gain statistical reliability. for this specific case study, our approach has 70% overall classification accuracy and the decision tree model reveals insights concerning the combined roles of the various signaling activities in governing cell migration speed. we conclude that decision tree methodology may facilitate elucidation of signal-response cascade relationships and generate experimentally testable predictions, which can be used as directions for future experiments.
Sampsa Hautaniemi, Sourabh Kharait, Akihiro Iwabu, Alan Wells, Douglas A. Lauffenburger
Bioinform.1
2004 Optimized LOWESS normalization parameter selection for DNA microarray data
abstract
BACKGROUND: Microarray data normalization is an important step for obtaining data that are reliable and usable for subsequent analysis. One of the most commonly utilized normalization techniques is the locally weighted scatterplot smoothing (LOWESS) algorithm. However, a much overlooked concern with the LOWESS normalization strategy deals with choosing the appropriate parameters. Parameters are usually chosen arbitrarily, which may reduce the efficiency of the normalization and result in non-optimally normalized data. Thus, there is a need to explore LOWESS parameter selection in greater detail. RESULTS AND DISCUSSION: In this work, we discuss how to choose parameters for the LOWESS method. Moreover, we present an optimization approach for obtaining the fraction of data points utilized in the local regression and analyze results for local print-tip normalization. The optimization procedure determines the bandwidth parameter for the local regression by minimizing a cost function that represents the mean-squared difference between the LOWESS estimates and the normalization reference level. We demonstrate the utility of the systematic parameter selection using two publicly available data sets. The first data set consists of three self versus self hybridizations, which allow for a quantitative study of the optimization method. The second data set contains a collection of DNA microarray data from a breast cancer study utilizing four breast cancer cell lines. Our results show that different parameter choices for the bandwidth window yield dramatically different calibration results in both studies. CONCLUSIONS: Results derived from the self versus self experiment indicate that the proposed optimization approach is a plausible solution for estimating the LOWESS parameters, while results from the breast cancer experiment show that the optimization procedure is readily applicable to real-life microarray data normalization. In summary, the systematic approach to obtain critical parameters in the LOWESS technique is likely to produce data that optimally meets assumptions made in the data preprocessing step and thereby makes studies utilizing the LOWESS method unambiguous and easier to repeat.
John A. Berger, Sampsa Hautaniemi, Anna-Kaarina Järvinen, Henrik Edgren, Sanjit K. Mitra, Jaakko Astola
BMC Bioinform.2
2003 CGH-Plotter: MATLAB toolbox for CGH-data analysis
abstract
CGH-Plotter is a MATLAB toolbox with a graphical user interface for the analysis of comparative genomic hybridization (CGH) microarray data. CGH-Plotter provides a tool for rapid visualization of CGH-data according to the locations of the genes along the genome. In addition, the CGH-Plotter identifies regions of amplifications and deletions, using k-means clustering and dynamic programming. The application offers a convenient way to analyze CGH-data and can also be applied for the analysis of cDNA microarray expression data. CGH-Plotter toolbox is platform independent and requires MATLAB 6.1 or higher to operate.
Reija Autio, Sampsa Hautaniemi, Päivikki Kauraniemi, Olli Yli-Harja, Jaakko Astola, Maija Wolf, Anne Kallioniemi
Bioinform.2
2003 A novel strategy for microarray quality control using Bayesian networks
abstract
MOTIVATION: High-throughput microarray technologies enable measurements of the expression levels of thousands of genes in parallel. However, microarray printing, hybridization and washing may create substantial variability in the quality of the data. As erroneous measurements may have a drastic impact on the results by disturbing the normalization schemes and by introducing expression patterns that lead to incorrect conclusions, it is crucial to discard low quality observations in the early phases of a microarray experiment. A typical microarray experiment consists of tens of thousands of spots on a microarray, making manual extraction of poor quality spots impossible. Thus, there is a need for a reliable and general microarray spot quality control strategy. RESULTS: We suggest a novel strategy for spot quality control by using Bayesian networks, which contain many appealing properties in the spot quality control context. We illustrate how a non-linear least squares based Gaussian fitting procedure can be used in order to extract features for a spot on a microarray. The features we used in this study are: spot intensity, size of the spot, roundness of the spot, alignment error, background intensity, background noise, and bleeding. We conclude that Bayesian networks are a reliable and useful model for microarray spot quality assessment. SUPPLEMENTARY INFORMATION: http://sigwww.cs.tut.fi/TICSP/SpotQuality/.
Sampsa Hautaniemi, Henrik Edgren, Petri Vesanen, Maija Wolf, Anna-Kaarina Järvinen, Olli Yli-Harja, Jaakko Astola, Olli-P. Kallioniemi, Outi Monni
Bioinform.1
2003 Analysis and Visualization of Gene Expression Microarray Data in Human Cancer Using Self-Organizing Maps
Sampsa Hautaniemi, Olli Yli-Harja, Jaakko Astola, Päivikki Kauraniemi, Anne Kallioniemi, Maija Wolf, Jimmy Ruiz, Spyro Mousses, Olli-P. Kallioniemi
Mach. Learn.1