VLDB 2026 Research / reviewers in the wild / expert
Antti Häkkinen
dblp:17/8426
· DBLP profile ↗
16ranked-venue papers
7as first author
5since 2021 · last 2026
0000-0002-8081-1588ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 16 · 7 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FUSE: data-driven functional segmentation of DNA methylation dataabstractSUMMARY: DNA methylation (DNAm) of neighbouring CpG sites is highly correlated, making DNAm function in terms of blocks. DNAm patterns and functionality are linked to both chromatin structure of DNA and gene regulation. Defining biologically meaningful DNA methylation blocks from whole-genome bisulfite sequencing (WGBS) data remains challenging, as most existing methods rely on fixed genomic windows rather than the observed methylation pattern. We present FUSE, a data-driven segmentation method that captures intrinsic methylation segments directly from WGBS data by jointly analyzing multiple samples. FUSE identifies spatially homogeneous methylation blocks shared across the input cohort while allowing different methylation states across samples. Applied to 61 WGBS samples from the ENCODE database, FUSE identified segments which overlap significantly with promoters, enhancers, and repetitive elements. FUSE was able to recover the true segment breakpoints in synthetic data with high sensitivity under increased levels of noise. As such, FUSE facilitates post hoc methylation analyses by aggregating coherent CpG sites into candidate segments for downstream differential methylation testing or other comparative studies. AVAILABILITY AND IMPLEMENTATION: FUSE is implemented as an R-package methFuse, available at https://github.com/holmsusa/methFuse and https://cran.r-project.org/package=methFuse. A GenomeSpy visualization of the data is available at https://csbi.ltdk.helsinki.fi/p/fuse_encode_gs/. Susanna Holmström, Antti Häkkinen, Kari Lavikka, Giovanni Marchi, Sampsa Hautaniemi, Alexandra Lahtinen |
Bioinform. | 2 |
| 2022 | POIBM: batch correction of heterogeneous RNA-seq datasets through latent sample matchingabstractMOTIVATION: RNA sequencing and other high-throughput technologies are essential in understanding complex diseases, such as cancers, but are susceptible to technical factors manifesting as patterns in the measurements. These batch patterns hinder the discovery of biologically relevant patterns. Unbiased batch effect correction in heterogeneous populations currently requires special experimental designs or phenotypic labels, which are not readily available for patient samples in existing datasets. RESULTS: We present POIBM, an RNA-seq batch correction method, which learns virtual reference samples directly from the data. We use a breast cancer cell line dataset to show that POIBM exceeds or matches the performance of previous methods, while being blind to the phenotypes. Further, we analyze The Cancer Genome Atlas RNA-seq data to show that batch effects plague many cancer types; POIBM effectively discovers the true replicates in stomach adenocarcinoma; and integrating the corrected data in endometrial carcinoma improves cancer subtyping. AVAILABILITY AND IMPLEMENTATION: https://bitbucket.org/anthakki/poibm/ (archived at https://doi.org/10.5281/zenodo.6122436). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Susanna Holmström, Sampsa Hautaniemi, Antti Häkkinen |
Bioinform. | 3 |
| 2021 | Network-guided identification of cancer-selective combinatorial therapies in ovarian cancerabstractEach patient's cancer consists of multiple cell subpopulations that are inherently heterogeneous and may develop differing phenotypes such as drug sensitivity or resistance. A personalized treatment regimen should therefore target multiple oncoproteins in the cancer cell populations that are driving the treatment resistance or disease progression in a given patient to provide maximal therapeutic effect, while avoiding severe co-inhibition of non-malignant cells that would lead to toxic side effects. To address the intra- and inter-tumoral heterogeneity when designing combinatorial treatment regimens for cancer patients, we have implemented a machine learning-based platform to guide identification of safe and effective combinatorial treatments that selectively inhibit cancer-related dysfunctions or resistance mechanisms in individual patients. In this case study, we show how the platform enables prediction of cancer-selective drug combinations for patients with high-grade serous ovarian cancer using single-cell imaging cytometry drug response assay, combined with genome-wide transcriptomic and genetic profiles. The platform makes use of drug-target interaction networks to prioritize those combinations that warrant further preclinical testing in scarce patient-derived primary cells. During the case study in ovarian cancer patients, we investigated (i) the relative performance of various ensemble learning algorithms for drug response prediction, (ii) the use of matched single-cell RNA-sequencing data to deconvolute cell population-specific transcriptome profiles from bulk RNA-seq data, (iii) and whether multi-patient or patient-specific predictive models lead to better predictive accuracy. The general platform and the comparison results are expected to become useful for future studies that use similar predictive approaches also in other cancer types. Liye He, Daria Bulanova, Jaana Oikkonen, Antti Häkkinen, Kaiyang Zhang, Erdogan Pekcan Erkan, Olli Carpén, Titta Joutsiniemi, Sakari Hietanen, Johanna Hynninen, Kaisa Huhtinen, Sampsa Hautaniemi, Anna Vähärautio, Jing Tang 0002, Krister Wennerberg, Tero Aittokallio |
Briefings Bioinform. | 4 |
| 2021 | Agile workflow for interactive analysis of mass cytometry dataabstractMOTIVATION: Single-cell proteomics technologies, such as mass cytometry, have enabled characterization of cell-to-cell variation and cell populations at a single-cell resolution. These large amounts of data, require dedicated, interactive tools for translating the data into knowledge. RESULTS: We present a comprehensive, interactive method called Cyto to streamline analysis of large-scale cytometry data. Cyto is a workflow-based open-source solution that automates the use of state-of-the-art single-cell analysis methods with interactive visualization. We show the utility of Cyto by applying it to mass cytometry data from peripheral blood and high-grade serous ovarian cancer (HGSOC) samples. Our results show that Cyto is able to reliably capture the immune cell sub-populations from peripheral blood and cellular compositions of unique immune- and cancer cell subpopulations in HGSOC tumor and ascites samples. AVAILABILITYAND IMPLEMENTATION: The method is available as a Docker container at https://hub.docker.com/r/anduril/cyto and the user guide and source code are available at https://bitbucket.org/anduril-dev/cyto. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Julia Casado, Oskari Lehtonen, Ville Rantanen, Katja Kaipio, Luca Pasquini, Antti Häkkinen, Eleonora Petrucci, Johanna Hynninen, Sakari Hietanen, Olli Carpén, Mauro Biffoni, Anniina Färkkilä, Sampsa Hautaniemi |
Bioinform. | 6 |
| 2021 | PRISM: recovering cell-type-specific expression profiles from individual composite RNA-seq samplesabstractMOTIVATION: A major challenge in analyzing cancer patient transcriptomes is that the tumors are inherently heterogeneous and evolving. We analyzed 214 bulk RNA samples of a longitudinal, prospective ovarian cancer cohort and found that the sample composition changes systematically due to chemotherapy and between the anatomical sites, preventing direct comparison of treatment-naive and treated samples. RESULTS: To overcome this, we developed PRISM, a latent statistical framework to simultaneously extract the sample composition and cell-type-specific whole-transcriptome profiles adapted to each individual sample. Our results indicate that the PRISM-derived composition-free transcriptomic profiles and signatures derived from them predict the patient response better than the composite raw bulk data. We validated our findings in independent ovarian cancer and melanoma cohorts, and verified that PRISM accurately estimates the composition and cell-type-specific expression through whole-genome sequencing and RNA in situ hybridization experiments. AVAILABILITYAND IMPLEMENTATION: https://bitbucket.org/anthakki/prism. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Antti Häkkinen, Kaiyang Zhang, Amjad Alkodsi, Noora Andersson, Erdogan Pekcan Erkan, Katja Kaipio, Tarja Lamminen, Naziha Mansuri, Kaisa Huhtinen, Anna Vähärautio, Olli Carpén, Johanna Hynninen, Sakari Hietanen, Rainer Lehtonen, Sampsa Hautaniemi |
Bioinform. | 1 |
| 2020 | qSNE: quadratic rate t-SNE optimizer with automatic parameter tuning for large datasetsabstractMOTIVATION: Non-parametric dimensionality reduction techniques, such as t-distributed stochastic neighbor embedding (t-SNE), are the most frequently used methods in the exploratory analysis of single-cell datasets. Current implementations scale poorly to massive datasets and often require downsampling or interpolative approximations, which can leave less-frequent populations undiscovered and much information unexploited. RESULTS: We implemented a fast t-SNE package, qSNE, which uses a quasi-Newton optimizer, allowing quadratic convergence rate and automatic perplexity (level of detail) optimizer. Our results show that these improvements make qSNE significantly faster than regular t-SNE packages and enables full analysis of large datasets, such as mass cytometry data, without downsampling. AVAILABILITY AND IMPLEMENTATION: Source code and documentation are openly available at https://bitbucket.org/anthakki/qsne/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Antti Häkkinen, Juha Koiranen, Julia Casado, Katja Kaipio, Oskari Lehtonen, Eleonora Petrucci, Johanna Hynninen, Sakari Hietanen, Olli Carpén, Luca Pasquini, Mauro Biffoni, Rainer Lehtonen, Sampsa Hautaniemi |
Bioinform. | 1 |
| 2019 | Anduril 2: upgraded large-scale data integration frameworkabstractSUMMARY: Anduril is an analysis and integration framework that facilitates the design, use, parallelization and reproducibility of bioinformatics workflows. Anduril has been upgraded to use Scala for pipeline construction, which simplifies software maintenance, and facilitates design of complex pipelines. Additionally, Anduril's bioinformatics repository has been expanded with multiple components, and tutorial pipelines, for next-generation sequencing data analysis. AVAILABILITYAND IMPLEMENTATION: Freely available at http://anduril.org. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Alejandra Cervera, Ville Rantanen, Kristian Ovaska, Marko Laakso, Javier Nuñez-Fontarnau, Amjad Alkodsi, Julia Casado, Chiara Facciotto, Antti Häkkinen, Riku Louhimo, Sirkku Karinen, Kaiyang Zhang, Kari Lavikka, Lauri Lyly, Sampsa Hautaniemi |
Bioinform. | 9 |
| 2018 | Identifying differentially methylated sites in samples with varying tumor purityabstractMotivation: DNA methylation aberrations are common in many cancer types. A major challenge hindering comparison of patient-derived samples is that they comprise of heterogeneous collection of cancer and microenvironment cells. We present a computational method that allows comparing cancer methylomes in two or more heterogeneous tumor samples featuring differing, unknown fraction of cancer cells. The method is unique in that it allows comparison also in the absence of normal cell control samples and without prior tumor purity estimates, as these are often unavailable or unreliable in clinical samples. Results: We use simulations and next-generation methylome, RNA and whole-genome sequencing data from two cancer types to demonstrate that the method is accurate and outperforms alternatives. The results show that our method adapts well to various cancer types and to a wide range of tumor content, and works robustly without a control or with controls derived from various sources. Availability and implementation: The method is freely available at https://bitbucket.org/anthakki/dmml. Supplementary information: Supplementary data are available at Bioinformatics online. Antti Häkkinen, Amjad Alkodsi, Chiara Facciotto, Kaiyang Zhang, Katja Kaipio, Sirpa Leppä, Olli Carpén, Seija Grénman, Johanna Hynninen, Sakari Hietanen, Rainer Lehtonen, Sampsa Hautaniemi |
Bioinform. | 1 |
| 2018 | SCIP: a single-cell image processor toolboxabstractSummary: Each cell is a phenotypically unique individual that is influenced by internal and external processes, operating in parallel. To characterize the dynamics of cellular processes one needs to observe many individual cells from multiple points of view and over time, so as to identify commonalities and variability. With this aim, we engineered a software, 'SCIP', to analyze multi-modal, multi-process, time-lapse microscopy morphological and functional images. SCIP is capable of automatic and/or manually corrected segmentation of cells and lineages, automatic alignment of different microscopy channels, as well as detect, count and characterize fluorescent spots (such as RNA tagged by MS2-GFP), nucleoids, Z rings, Min system, inclusion bodies, undefined structures, etc. The results can be exported into *mat files and all results can be jointly analyzed, to allow studying not only each feature and process individually, but also find potential relationships. While we exemplify its use on Escherichia coli, many of its functionalities are expected to be of use in analyzing other prokaryotes and eukaryotic cells as well. We expect SCIP to facilitate the finding of relationships between cellular processes, from small-scale (e.g. gene expression) to large-scale (e.g. cell division), in single cells and cell lineages. Availability and implementation: http://www.ca3-uninova.org/project_scip. Supplementary information: Supplementary data are available at Bioinformatics online. Leonardo Martins, Ramakanth Neeli-Venkata, Samuel M. D. Oliveira, Antti Häkkinen, Andre S. Ribeiro, José Manuel Fonseca |
Bioinform. | 4 |
| 2016 | Characterizing rate limiting steps in transcription from RNA production times in live cellsabstractMOTIVATION: Single-molecule measurements of live Escherichia coli transcription dynamics suggest that this process ranges from sub- to super-Poissonian, depending on the conditions and on the promoter. For its accurate quantification, we propose a model that accommodates all these settings, and statistical methods to estimate the model parameters and to select the relevant components. RESULTS: The new methodology has improved accuracy and avoids overestimating the transcription rate due to finite measurement time, by exploiting unobserved data and by accounting for the effects of discrete sampling. First, we use Monte Carlo simulations of models based on measurements to show that the methods are reliable and offer substantial improvements over previous methods. Next, we apply the methods on measurements of transcription intervals of different promoters in live E. coli, and show that they produce significantly different results, both in low- and high-noise settings, and that, in the latter case, they even lead to qualitatively different results. Finally, we demonstrate that the methods can be generalized for other similar purposes, such as for estimating gene activation kinetics. In this case, the new methods allow quantifying the inducer uptake dynamics as opposed to just comparing them between cases, which was not previously possible. We expect this new methodology to be a valuable tool for functional analysis of cellular processes using single-molecule or single-event microscopy measurements in live cells. AVAILABILITY AND IMPLEMENTATION: Source code is available under Mozilla Public License at http://www.cs.tut.fi/%7Ehakkin22/censored/ CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Antti Häkkinen, Andre S. Ribeiro |
Bioinform. | 1 |
| 2016 | Temperature-Dependent Model of Multi-step Transcription Initiation in Escherichia coli Based on Live Single-Cell MeasurementsabstractTranscription kinetics is limited by its initiation steps, which differ between promoters and with intra- and extracellular conditions. Regulation of these steps allows tuning both the rate and stochasticity of RNA production. We used time-lapse, single-RNA microscopy measurements in live Escherichia coli to study how the rate-limiting steps in initiation of the Plac/ara-1 promoter change with temperature and induction scheme. For this, we compared detailed stochastic models fit to the empirical data in maximum likelihood sense using statistical methods. Using this analysis, we found that temperature affects the rate limiting steps unequally, as nonlinear changes in the closed complex formation suffice to explain the differences in transcription dynamics between conditions. Meanwhile, a similar analysis of the PtetA promoter revealed that it has a different rate limiting step configuration, with temperature regulating different steps. Finally, we used the derived models to explore a possible cause for why the identified steps are preferred as the main cause for behavior modifications with temperature: we find that transcription dynamics is either insensitive or responds reciprocally to changes in the other steps. Our results suggests that different promoters employ different rate limiting step patterns that control not only their rate and variability, but also their sensitivity to environmental changes. Samuel M. D. Oliveira, Antti Häkkinen, Jason Lloyd-Price, Vinodh Kandavalli, Andre S. Ribeiro |
PLoS Comput. Biol. | 2 |
| 2015 | Estimation of GFP-tagged RNA numbers from temporal fluorescence intensity dataabstractMOTIVATION: MS2-GFP-tagging of RNA is currently the only method to measure intervals between consecutive transcription events in live cells. For this, new transcripts must be accurately detected from intensity time traces. RESULTS: We present a novel method for automatically estimating RNA numbers and production intervals from temporal data of cell fluorescence intensities that reduces uncertainty by exploiting temporal information. We also derive a robust variant, more resistant to outliers caused e.g. by RNAs moving out of focus. Using Monte Carlo simulations, we show that the quantification of RNA numbers and production intervals is generally improved compared with previous methods. Finally, we analyze data from live Escherichia coli and show statistically significant differences to previous methods. The new methods can be used to quantify numbers and production intervals of any fluorescent probes, which are present in low copy numbers, are brighter than the cell background and degrade slowly. AVAILABILITY: Source code is available under Mozilla Public License at http://www.cs.tut.fi/%7ehakkin22/jumpdet/. Antti Häkkinen, Andre S. Ribeiro |
Bioinform. | 1 |
| 2014 | Estimation of fluorescence-tagged RNA numbers from spot intensitiesabstractMOTIVATION: Present research on gene expression using live cell imaging and fluorescent proteins or tagged RNA requires accurate automated methods of quantification of these molecules from the images. Here, we propose a novel automated method for classifying pixel intensities of fluorescent spots to RNA numbers. RESULTS: The method relies on a new model of intensity distributions of tagged RNAs, for which we estimated parameter values in maximum likelihood sense from measurement data, and constructed a maximum a posteriori classifier to estimate RNA numbers in fluorescent RNA spots. We applied the method to estimate the number of tagged RNAs in individual live Escherichia coli cells containing a gene coding for an RNA with MS2-GFP binding sites. We tested the method using two constructs, coding for either 96 or 48 binding sites, and obtained similar distributions of RNA numbers, showing that the method is adaptive. We further show that the results agree with a method that uses time series data and with quantitative polymerase chain reaction measurements. Lastly, using simulated data, we show that the method is accurate in realistic parameter ranges. This method should, in general, be applicable to live single-cell measurements of low-copy number fluorescence-tagged molecules. AVAILABILITY AND IMPLEMENTATION: MATLAB extensions written in C for parameter estimation and finding decision boundaries are available under Mozilla public license at http://www.cs.tut.fi/%7ehakkin22/estrna/ CONTACT: [email protected]. Antti Häkkinen, Meenakshisundaram Kandhavelu, Stef Garasto, Andre S. Ribeiro |
Bioinform. | 1 |
| 2013 | CellAging: a tool to study segregation and partitioning in division in cell lineages of Escherichia coliabstractMOTIVATION: Cell division in Escherichia coli is morphologically symmetric. However, as unwanted protein aggregates are segregated to the cell poles and, after divisions, accumulate at older poles, generate asymmetries in sister cells' vitality. Novel single-molecule detection techniques allow observing aging-related processes in vivo, over multiple generations, informing on the underlying mechanisms. RESULTS: CellAging is a tool to automatically extract information on polar segregation and partitioning in division of aggregates in E.coli, and on cellular vitality. From time-lapse, parallel brightfield and fluorescence microscopy images, it performs cell segmentation, alignment of brightfield and fluorescence images, lineage construction and pole age determination, and it computes aging-related features. We exemplify its use by analyzing spatial distributions of fluorescent protein aggregates from images of cells across generations. AVAILABILITY: CellAging, instructions and an example are available at http://www.cs.tut.fi/%7esanchesr/cellaging/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Antti Häkkinen, Anantha Barathi Muthukrishnan, André Mora, José Manuel Fonseca, Andre S. Ribeiro |
Bioinform. | 1 |
| 2012 | Detecting sequence dependent transcriptional pauses from RNA and protein number time seriesabstractBACKGROUND: Evidence suggests that in prokaryotes sequence-dependent transcriptional pauses affect the dynamics of transcription and translation, as well as of small genetic circuits. So far, a few pause-prone sequences have been identified from in vitro measurements of transcription elongation kinetics. RESULTS: Using a stochastic model of gene expression at the nucleotide and codon levels with realistic parameter values, we investigate three different but related questions and present statistical methods for their analysis. First, we show that information from in vivo RNA and protein temporal numbers is sufficient to discriminate between models with and without a pause site in their coding sequence. Second, we demonstrate that it is possible to separate a large variety of models from each other with pauses of various durations and locations in the template by means of a hierarchical clustering and a random forest classifier. Third, we introduce an approximate likelihood function that allows to estimate the location of a pause site. CONCLUSIONS: This method can aid in detecting unknown pause-prone sequences from temporal measurements of RNA and protein numbers at a genome-wide scale and thus elucidate possible roles that these sequences play in the dynamics of genetic networks and phenotype. Frank Emmert-Streib, Antti Häkkinen, Andre S. Ribeiro |
BMC Bioinform. | 2 |
| 2010 | Effects of Transcriptional Pausing on Gene Expression DynamicsabstractStochasticity in gene expression affects many cellular processes and is a source of phenotypic diversity between genetically identical individuals. Events in elongation, particularly RNA polymerase pausing, are a source of this noise. Since the rate and duration of pausing are sequence-dependent, this regulatory mechanism of transcriptional dynamics is evolvable. The dependency of pause propensity on regulatory molecules makes pausing a response mechanism to external stress. Using a delayed stochastic model of bacterial transcription at the single nucleotide level that includes the promoter open complex formation, pausing, arrest, misincorporation and editing, pyrophosphorolysis, and premature termination, we investigate how RNA polymerase pausing affects a gene's transcriptional dynamics and gene networks. We show that pauses' duration and rate of occurrence affect the bursting in RNA production, transcriptional and translational noise, and the transient to reach mean RNA and protein levels. In a genetic repressilator, increasing the pausing rate and the duration of pausing events increases the period length but does not affect the robustness of the periodicity. We conclude that RNA polymerase pausing might be an important evolvable feature of genetic networks. Tiina Rajala, Antti Häkkinen, Shannon Healy, Olli Yli-Harja, Andre S. Ribeiro |
PLoS Comput. Biol. | 2 |