EDBT 2026 Demo / reviewers in the wild / expert
Leo Lahti
dblp:07/148
· DBLP profile ↗
16ranked-venue papers
4as first author
8since 2021 · last 2025
0000-0001-5537-637XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 12 · 3 first-author · 7 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-authorArtificial intelligence and machine learning · 2 · 1 first-authorTheory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Elementary methods provide more replicable results in microbial differential abundance analysisabstractDifferential abundance analysis (DAA) is a key component of microbiome studies. Although dozens of methods exist, there is currently no consensus on the preferred methods. While the correctness of results in DAA is an ambiguous concept and cannot be fully evaluated without setting the ground truth and employing simulated data, we argue that a well-performing method should be effective in producing highly reproducible results. We compared the performance of 14 DAA methods by employing datasets from 53 taxonomic profiling studies based on 16S rRNA gene or shotgun metagenomic sequencing. For each method, we examined how the results replicated between random partitions of each dataset and between datasets from separate studies. While certain methods showed good consistency, some widely used methods were observed to produce a substantial number of conflicting findings. Overall, when considering consistency together with sensitivity, the best performance was attained by analyzing relative abundances with a nonparametric method (Wilcoxon test or ordinal regression model) or linear regression/t-test. Moreover, a comparable performance was obtained by analyzing presence/absence of taxa with logistic regression. Juho Pelto, Kari Auranen, Janne V. Kujala, Leo Lahti |
Briefings Bioinform. | 4 |
| 2025 | Multi-omics time-series analysis in microbiome research: a systematic reviewabstractRecent developments in data generation have opened up unprecedented insights into living systems. It has been recognized that integrating and characterizing temporal variation simultaneously across multiple scales, from specific molecular interactions to entire ecosystems, is crucial for uncovering biological mechanisms and understanding the emergence of complex phenotypes. With the increasing number of studies incorporating multi-omics data sampled over time, it has become clear that integrated approaches are pivotal for these efforts. However, standard data analytical practices in longitudinal multi-omics are still shaping up and many of the available methods have not yet been widely evaluated and adopted. To address this gap, we performed the first systematic literature review that comprehensively categorizes, compares, and evaluates computational methods for longitudinal multi-omics integration, with a particular emphasis on four categories of the studies: (i) host and host-associated microbiome studies, (ii) microbiome-free host studies, (iii) host-free microbiome studies, and (iv) methodological framework studies. Our review highlights current methodological trends, identifies widely used and high-performing frameworks, and assesses each method across performance, interpretability, and ease of use. We further organize these methods into thematic groups-such as statistical modeling, machine learning, dimensionality reduction, and latent factor approaches-to provide a clear roadmap for future research and application. This work offers a critical foundation for advancing integrative longitudinal data science and supporting reproducible, scalable analysis in this rapidly evolving field. Moiz Khan Sherwani, Matti O. Ruuskanen, Dylan Feldner-Busztin, Panos Firbas Nisantzis, Gergely Boza, Ágnes Móréh, Tuomas Borman, Pande Putu Erawijantari, István Scheuring, Shyam Gopalakrishnan, Leo Lahti |
Briefings Bioinform. | 11 |
| 2025 | LimROTS: a hybrid method integrating empirical Bayes and reproducibility-optimized statistics for robust differential expression analysisabstractMOTIVATION: Differential expression analysis plays a vital role in omics research enabling precise identification of features that associate with different phenotypes. This process is critical for uncovering biological differences between conditions, such as disease versus healthy states. In proteomics, several statistical methods have been used, ranging from simple t-tests to more advanced methods like DEqMS, limma and ROTS. However, a flexible method for reproducibility-optimized statistics tailored for clinical omics data has been lacking. RESULTS: In this study, we developed LimROTS, a hybrid method that integrates a linear regression model and the empirical Bayes approach with reproducibility optimized statistics, to create a novel moderated ranking statistic, for robust and flexible analysis of proteomics data. We validated its performance using twenty-one proteomics gold standard spike-in datasets with different protein mixtures, MS instruments, and techniques for benchmarking. This hybrid approach improves accuracy and reproducibility of complex proteomics data, making LimROTS a powerful tool for high-dimensional omics data analysis. AVAILABILITY AND IMPLEMENTATION: LimROTS has been implemented as an R/Bioconductor package, available at https://doi.org/doi:10.18129/B9.bioc.LimROTS. Additionally, the code used in this study is available in GitHub repository https://github.com/AliYoussef96/LimROTSmanuscript. Ali Mostafa Anwar, Akewak Jeba, Leo Lahti, Eleanor Coffey |
Bioinform. | 3 |
| 2025 | HoloFoodR: a statistical programming framework for holo-omics data integration workflowsabstractSUMMARY: Holo-omics is an emerging research area that integrates multi-omic datasets from the host organism and its microbiome to study their interactions. Recently, curated and openly accessible holo-omic databases have been developed. The HoloFood database, for instance, provides nearly 10 000 holo-omic profiles for salmon and chicken under controlled treatments. However, bridging the gap between holo-omic data resources and algorithmic frameworks remains a challenge. Combining the latest advances in statistical programming with curated holo-omic data sets can facilitate the design of open and reproducible research workflows in the emerging field of holo-omics. AVAILABILITY AND IMPLEMENTATION: HoloFoodR R/Bioconductor package and the source code are available under the open-source Artistic License 2.0 at the package homepage https://doi.org/10.18129/B9.bioc.HoloFoodR. Tuomas Borman, Artur Sannikov, Robert D. Finn, Morten Tønsberg Limborg, Alexander B. Rogers, Varsha Kale, Kati Hanhineva, Leo Lahti |
Bioinform. | 8 |
| 2025 | Learning and teaching biological data science in the Bioconductor communityabstractModern biological research is increasingly data-intensive, leading to a growing demand for effective training in biological data science. In this article, we provide an overview of key resources and best practices available within the Bioconductor project-an open-source software community focused on omics data analysis. This guide serves as a valuable reference for both learners and educators in the field. Jenny Drnevich, Frederick Tan, Fabricio Almeida-Silva, Robert Castelo, Aedín C. Culhane, Sean R. Davis, Maria A. Doyle, Ludwig Geistlinger, Andrew R. Ghazi, Susan P. Holmes, Leo Lahti, Alexandru Mahmoud, Kozo Nishida, Marcel Ramos, Kevin C. Rue-Albrecht, David J. H. Shih, Laurent Gatto, Charlotte Soneson |
PLoS Comput. Biol. | 11 |
| 2024 | Algorithm 1047: FdeSolver, a Julia Package for Solving Fractional Differential EquationsabstractWe introduce FdeSolver, an open-source Julia package designed to solve fractional-order differential equations efficiently. The available solutions are based on product-integration rules, predictor–corrector algorithms, and the Newton-Raphson method. The package covers solutions for one-dimensional equations with orders of positive real numbers. For higher-dimensional systems, it supports orders up to one. Incommensurate derivatives are allowed and defined in the Caputo sense. Here, we summarize the implementation for a representative class of problems and compare it with available alternatives in Julia and MATLAB. Moreover, FdeSolver leverages the power and flexibility of the Julia environment to offer enhanced computational performance, and our development emphasizes adherence to the best practices of open research software. To highlight its practical utility, we demonstrate its capability in simulating microbial community dynamics and modeling the spread of COVID-19. This latter application involves fitting the order of derivatives grounded on real-world epidemiological data. Overall, these results highlight the efficiency, reliability, and practicality of the FdeSolver Julia package. Moein Khalighi, Giulio Benedetti, Leo Lahti |
ACM Trans. Math. Softw. | 3 |
| 2023 | Dealing with dimensionality: the application of machine learning to multi-omics dataabstractMOTIVATION: Machine learning (ML) methods are motivated by the need to automate information extraction from large datasets in order to support human users in data-driven tasks. This is an attractive approach for integrative joint analysis of vast amounts of omics data produced in next generation sequencing and other -omics assays. A systematic assessment of the current literature can help to identify key trends and potential gaps in methodology and applications. We surveyed the literature on ML multi-omic data integration and quantitatively explored the goals, techniques and data involved in this field. We were particularly interested in examining how researchers use ML to deal with the volume and complexity of these datasets. RESULTS: Our main finding is that the methods used are those that address the challenges of datasets with few samples and many features. Dimensionality reduction methods are used to reduce the feature count alongside models that can also appropriately handle relatively few samples. Popular techniques include autoencoders, random forests and support vector machines. We also found that the field is heavily influenced by the use of The Cancer Genome Atlas dataset, which is accessible and contains many diverse experiments. AVAILABILITY AND IMPLEMENTATION: All data and processing scripts are available at this GitLab repository: https://gitlab.com/polavieja_lab/ml_multi-omics_review/ or in Zenodo: https://doi.org/10.5281/zenodo.7361807. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Dylan Feldner-Busztin, Panos Firbas Nisantzis, Shelley Jane Edmunds, Gergely Boza, Fernando Racimo, Shyam Gopalakrishnan, Morten Tønsberg Limborg, Leo Lahti, Gonzalo G. de Polavieja |
Bioinform. | 8 |
| 2022 | Quantifying the impact of ecological memory on the dynamics of interacting communitiesabstractEcological memory refers to the influence of past events on the response of an ecosystem to exogenous or endogenous changes. Memory has been widely recognized as a key contributor to the dynamics of ecosystems and other complex systems, yet quantitative community models often ignore memory and its implications. Recent modeling studies have shown how interactions between community members can lead to the emergence of resilience and multistability under environmental perturbations. We demonstrate how memory can be introduced in such models using the framework of fractional calculus. We study how the dynamics of a well-characterized interaction model is affected by gradual increases in ecological memory under varying initial conditions, perturbations, and stochasticity. Our results highlight the implications of memory on several key aspects of community dynamics. In general, memory introduces inertia into the dynamics. This favors species coexistence under perturbation, enhances system resistance to state shifts, mitigates hysteresis, and can affect system resilience both ways depending on the time scale considered. Memory also promotes long transient dynamics, such as long-standing oscillations and delayed regime shifts, and contributes to the emergence and persistence of alternative stable states. Our study highlights the fundamental role of memory in communities, and provides quantitative tools to introduce it in ecological models and analyse its impact under varying conditions. Moein Khalighi, Guilhem Sommeria-Klein, Didier Gonze, Karoline Faust, Leo Lahti |
PLoS Comput. Biol. | 5 |
| 2018 | Open Data Science
Leo Lahti |
IDA | 1 |
| 2018 | A Hierarchical Ornstein-Uhlenbeck Model for Stochastic Time Series Analysis
Ville Laitinen, Leo Lahti |
IDA | 2 |
| 2017 | Linking Statistical and Ecological Theory: Hubbell's Unified Neutral Theory of Biodiversity as a Hierarchical Dirichlet ProcessabstractNeutral models which assume ecological equivalence between species provide null models for community assembly. In Hubbell's unified neutral theory of biodiversity (UNTB), many local communities are connected to a single metacommunity through differing immigration rates. Our ability to fit the full multisite UNTB has hitherto been limited by the lack of a computationally tractable and accurate algorithm. We show that a large class of neutral models with this mainland-island structure but differing local community dynamics converge in the large population limit to the hierarchical Dirichlet process. Using this approximation we developed an efficient Bayesian fitting strategy for the multisite UNTB. We can also use this approach to distinguish between neutral local community assembly given a nonneutral metacommunity distribution and the full UNTB where the metacommunity too assembles neutrally. We applied this fitting strategy to both tropical trees and a data set comprising 570$\,$851 sequences from 278 human gut microbiomes. The tropical tree data set was consistent with the UNTB but for the human gut neutrality was rejected at the whole community level. However, when we applied the algorithm to gut microbial species within the same taxon at different levels of taxonomic resolution, we found that species abundances within some genera were almost consistent with local community assembly. This was not true at higher taxonomic ranks. This suggests that the gut microbiota is more strongly niche constrained than macroscopic organisms, with different groups adopting different functional roles, but within those groups diversity may at least partially be maintained by neutrality. We also observed a negative correlation between body mass index and immigration rates within the family Ruminococcaceae. This provides a novel interpretation of the impact of obesity on the human microbiome as a relative increase in the importance of local growth versus external immigration within this key group of carbohydrate degrading organisms. Keith Harris, Todd L. Parsons, Umer Zeeshan Ijaz, Leo Lahti, Ian H. Holmes, Christopher Quince |
Proc. IEEE | 4 |
| 2013 | Cancer gene prioritization by integrative analysis of mRNA expression and DNA copy number data: a comparative reviewabstractA variety of genome-wide profiling techniques are available to investigate complementary aspects of genome structure and function. Integrative analysis of heterogeneous data sources can reveal higher level interactions that cannot be detected based on individual observations. A standard integration task in cancer studies is to identify altered genomic regions that induce changes in the expression of the associated genes based on joint analysis of genome-wide gene expression and copy number profiling measurements. In this review, we highlight common approaches to genomic data integration and provide a transparent benchmarking procedure to quantitatively compare method performances in cancer gene prioritization. Algorithms, data sets and benchmarking results are available at http://intcomp.r-forge.r-project.org. Leo Lahti, Martin Schäfer, Hans-Ulrich Klein, Silvio Bicciato, Martin Dugas |
Briefings Bioinform. | 1 |
| 2011 | Probabilistic Analysis of Probe Reliability in Differential Gene Expression Studies with Short Oligonucleotide ArraysabstractProbe defects are a major source of noise in gene expression studies. While existing approaches detect noisy probes based on external information such as genomic alignments, we introduce and validate a targeted probabilistic method for analyzing probe reliability directly from expression data and independently of the noise source. This provides insights into the various sources of probe-level noise and gives tools to guide probe design. Leo Lahti, Laura Elo, Tero Aittokallio, Samuel Kaski |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2010 | Global modeling of transcriptional responses in interaction networksabstractMOTIVATION: Cell-biological processes are regulated through a complex network of interactions between genes and their products. The processes, their activating conditions and the associated transcriptional responses are often unknown. Organism-wide modeling of network activation can reveal unique and shared mechanisms between tissues, and potentially as yet unknown processes. The same method can also be applied to cell-biological conditions in one or more tissues. RESULTS: We introduce a novel approach for organism-wide discovery and analysis of transcriptional responses in interaction networks. The method searches for local, connected regions in a network that exhibit coordinated transcriptional response in a subset of tissues. Known interactions between genes are used to limit the search space and to guide the analysis. Validation on a human pathway network reveals physiologically coherent responses, functional relatedness between tissues and coordinated, context-specific regulation of the genes. AVAILABILITY: Implementation is freely available in R and Matlab at http://www.cis.hut.fi/projects/mi/software/NetResponse Leo Lahti, Juha E. A. Knuuttila, Samuel Kaski |
Bioinform. | 1 |
| 2005 | Associative Clustering for Exploring Dependencies between Functional Genomics Data SetsabstractHigh-throughput genomic measurements, interpreted as cooccurring data samples from multiple sources, open up a fresh problem for machine learning: What is in common in the different data sets, that is, what kind of statistical dependencies are there between the paired samples from the different sets? We introduce a clustering algorithm for exploring the dependencies. Samples within each data set are grouped such that the dependencies between groups of different sets capture as much of pairwise dependencies between the samples as possible. We formalize this problem in a novel probabilistic way, as optimization of a Bayes factor. The method is applied to reveal commonalities and exceptions in gene expression between organisms and to suggest regulatory interactions in the form of dependencies between gene expression profiles and regulator binding patterns. Samuel Kaski, Janne Nikkilä, Janne Sinkkonen, Leo Lahti, Juha E. A. Knuuttila, Christophe Roos |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2004 | Associative Clustering
Janne Sinkkonen, Janne Nikkilä, Leo Lahti, Samuel Kaski |
ECML | 3 |