Tomi Suomi

dblp:214/2664 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
7since 2021 · last 2027
0000-0003-3639-979XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 10 · 2 first-author · 5 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2027 How contribution streaks form and their role in long-term retention in OSS projects
abstract
Abstract Contributor turnover is a major challenge in open-source software (OSS) projects, affecting project sustainability, quality, and many other dimensions of project performance. Although turnover has been widely studied in the OSS context, the importance of sustained contribution periods to long-term contributor survival remains poorly understood. In this context, contributor survival refers to the probability of sustained activity over time. We investigate the effect of contribution streaks, defined as consecutive months of contribution, on long-term survival, and analyze what kind of engagement, at different levels of ownership, predicts short-term contribution streaks. Using five years of data from 14 open-source projects, we applied landmark analysis to understand long-term survival based on peak streak lengths and logistic regression to identify the activities supporting short-term streaks. In both cases, we identified a three-month contribution streak to be an important threshold for sustained contributions. We found that development and communication related activities across ownership levels are important predictors of whether contributors continue their streaks, even though the relative importance shifts over time. In addition, we found that prior experience plays a stronger role early on, but its effect diminishes during later streak months. These findings suggest that contribution streaks are important predictors of long-term survival and provide practical insights for open-source practitioners on how to improve the retention of new and existing contributors through supporting short-term contribution streaks.
Tomi Suomi, Timo Aho, Petri Ihantola
Empir. Softw. Eng.1
2026 REACTOR: REgulon Activity analysis and Comparison Tool for single-cell transcriptOmics Research
abstract
SUMMARY: We introduce REACTOR, a computational tool designed to detect differential activity of transcriptional regulators and their target genes (regulons) in single-cell RNA-sequencing data. It expands the currently available framework for regulon analysis by introducing a robust statistical test to detect differential regulon activity between conditions, such as disease versus control, with multiple replicates. By contrasting different conditions, REACTOR enables identification of key condition- and cell type-specific regulons. To demonstrate the use of REACTOR, we illustrate its performance in a publicly available COVID-19 dataset. AVAILABILITY: REACTOR R-package together with an implementation vignette are available at https://www.github.com/elolab/REACTOR.
Markus Lindén, Sebastian I Zúñiga Norman, Tommi Välikangas, Sini Junttila, Tomi Suomi, Kalle T. Rytkönen, Laura Elo
Bioinform.5
2026 Clarifying the scope and capabilities of ROTS in differential expression analysis
abstract
SUMMARY: Recently, Anwar et al. introduced a method combining the ROTS reproducibility optimisation procedure with empirical Bayes variance estimation from limma. Here, we clarify several methodological aspects to support accurate interpretation of the results. We emphasise that ROTS is a general reproducibility optimisation framework rather than a single statistical test and demonstrate that benchmarking outcomes in the reported spike-in case studies are highly sensitive to analysis and evaluation choices. Furthermore, our reanalyses of the spike-in datasets do not support the reported conclusions, and we were unable to reproduce the results of the clinical Alzheimer's disease case study. These findings highlight the importance of transparent benchmarking practices and careful interpretation of comparative results. AVAILABILITY AND IMPLEMENTATION: The ROTS package is available through Bioconductor. The reanalyses were performed using the original code, with the minimal additions described in the manuscript.
Tomi Suomi, Jalmari Kettunen, Taneli Pusa, Laura Elo
Bioinform.1
2025 A Mosaic of Perspectives: Understanding Ownership in Software Engineering
abstract
Abstract Agile software development relies on self-organized teams, underlining the importance of individual responsibility. How developers take responsibility and build ownership are influenced by external factors such as architecture and development methods. This position paper examines the existing literature on ownership in software engineering and in psychology, and argues that a more comprehensive view of ownership in software engineering has a great potential in improving software team’s work. Initial positions on the issue are offered for discussion and to lay foundations for further research.
Tomi Suomi, Petri Ihantola, Tommi Mikkonen, Niko Mäkitalo
XP1
2024 ECCB2024: The 23rd European Conference on Computational Biology
abstract
This volume of Bioinformatics includes the proceedings papers of the 23rd European Conference on Computational Biology (ECCB2024) to be held in Turku, Finland, from 16 September to 20 September 2024, under the theme Data and Algorithms for Health and Science. More information on the ECCB2024 conference is available at the conference website https://eccb2024.fi/. ECCB is one of the main international conferences in the field of computational biology and bioinformatics together with the Intelligent Systems for Molecular Biology (ISMB) and the Research in Computational Molecular Biology (RECOMB). It is held jointly with the ISMB conference in odd-numbered years and independently in even-numbered years. ECCB attracts scientists and industry professionals from diverse disciplines, including mathematics, statistics, computer science, biology, and medicine. Rapid technological advancements enable life scientists to gain increasingly detailed insights in complex biological systems. These improvements in measurement technologies, however, introduce new challenges for data interpretation and necessitate the development of advanced computational techniques to manage the complexity and volume of the data. Consequently, the field of computational biology is rapidly evolving with new algorithms, software, and databases. The ECCB2024 conference showcases cutting-edge developments in systems biology, artificial intelligence, single-cell and spatial technologies, data integration, and more, addressing the growing demand for sophisticated algorithms to enhance the analysis of large-scale biological and biomedical datasets. This is highlighted by the most frequent keywords accompanying all the accepted submissions in ECCB2024 (tutorials/workshops, proceedings, highlight talks, posters), with the most popular keywords including ‘machine learning’, ‘deep learning’, ‘single-cell rnaseq’, and ‘multiomics’ (Fig. 1). Most frequent keywords among all accepted submissions in ECCB2024. The barplot shows the frequencies of top keywords appearing in at least ten submissions, while the word cloud illustrates the relative frequencies of all keywords appearing in at least five submissions. The ECCB2024 edition features five keynote lectures by distinguished speakers: Sarah Teichmann (Cambridge Stem Cell Institute, University of Cambridge, UK), Peer Bork (EMBL—European Molecular Biology Laboratory, Germany), Ileana Cristea (Princeton University, US), Jussi Taipale (University of Cambridge, UK), and Fabian Theis (Helmholtz Munich Computational Health Center, Germany). In addition, ECCB2024 hosts a scientific debate on data sharing and privacy protection by Melissa Haendel (University of North Carolina, US) and Yves Moreau (KU Leuven, Belgium). The ECCB2024 conference covers a wide range of topics, focusing on methodological advancements in computational biology as well as innovative application of computational techniques to life sciences and medicine. To provide a cohesive overview of recent scientific progress, the conference presentations are organized under six broad themes: (i) Genomes, (ii) Proteins, (iii) Systems biology and multiomics, (iv) Single-cell omics, (v) Microbiomes and planetary health, and (vi) Digital health. The proceedings talks present new scientific contributions, while the highlight talks showcase already published cutting-edge science in computational biology and related further developments. Poster presentations provide an opportunity for the participants to discuss their recent work with other researchers in the field. Workshops and tutorials on specialized topics prior to the main conference program are platforms to share practical experiences and learn new skills. In addition to scientific presentation tracks, ECCB2024 has a separate ELIXIR track, overseen by ELIXIR, focusing on advancements in infrastructure and services within ELIXIR nodes in support of the expert groups known as ‘ELIXIR Communities’. ELIXIR is a distributed pan-European life science infrastructure for biological data that coordinates, integrates, and sustains bioinformatics resources across its member states. This enables academic and industry users to access data, tools, standards, computing, and training services for life science research. ELIXIR has selected ECCB as a primary dissemination platform, serving as a co-organizing sponsor. Apart from Community activities the track will also introduce three scientific areas as outlined in the 2024–2028 ELIXIR Scientific Programme (https://elixir-europe.org), enabling scientists to access and analyse life science data across ‘Cellular and Molecular Research’, ‘Biodiversity, food security & pathogens’, and ‘Human data & translational research’. ECCB2024 received an impressive number of 200 submissions of full manuscripts for the proceedings call, highlighting the growing interest and engagement within the community. These submissions were organized under the six conference themes. Each submission was subjected to a peer-review process, with at least two reviews per manuscript, managed by the members of the ECCB2024 Programme Committee. The Programme Committee was chaired by Laura Elo (University of Turku, Finland) and each conference theme had an area chair (Table 1). The Proceedings Review Committee had over 100 reviewers. Thematic areas of ECCB2024 proceedings talks.a The table lists the area chairs for each theme, the number of reviewed papers, the number of accepted papers, and the acceptance rate for each theme. Thematic areas of ECCB2024 proceedings talks.a The table lists the area chairs for each theme, the number of reviewed papers, the number of accepted papers, and the acceptance rate for each theme. The review process focused on the impact, reproducibility, and scientific quality of the submitted research, as well as its relevance, interest, and value for the ECCB2024 audience. Upon completion of the review, the Program Committee chairs selected 24 papers to be included in the ECCB2024 proceedings, with an acceptance rate of 12% across the different areas (Table 1). All proceedings papers are published open access in this Proceedings issue of September 2024 of Bioinformatics. The ECCB2024 call for Highlight Talks invited presentations of studies recently published in scientific journals since 1 March 2023, or accepted for publication. The existence of further developments related to the published paper was also positively considered. The call received a total of 212 proposals, which were ranked according to their relevance and impact on computational biology, as well as their suitability to be presented to the large and diverse audience. Ultimately, the Program Committee selected 23% of the submissions for presentation as a Highlight talk at the conference. ECCB2024 hosts two poster sessions where researchers can introduce and discuss their work under the conference themes. In total, nearly 700 posters were accepted to be presented. The poster submissions were reviewed by the Posters Committee: Sini Junttila, Asta Laiho, Tomi Suomi, and Sampsa Hautaniemi (late posters). In addition, the ELIXIR track selected posters to be presented at ECCB2024. Overall, the ECCB2024 tracks attracted participation of presenting authors from 48 countries. The countries with the highest number of presenting authors were Germany, Finland, the USA, the United Kingdom, and Spain, jointly contributing to half of the overall number (Fig. 2). Following these were Turkey, France, and China, each contributing 4%–5% of the presenting authors. Proportion of countries where the presenting author of each accepted submission had their primary affiliation. Countries with the proportion below 2% were grouped under ‘Other’. Exhibitor booths are open throughout the conference, including ELIXIR, ISCB, Oxford University Press, Royal Society Publishing and eLife Sciences Publications Ltd, as well as the two organizing institutions: University of Turku and CSC—IT Center for Science. The booths showcase the latest scientific literature in computational biology and bioinformatics, data stewardship, as well as new developments in hardware, software, and technology. In addition, the Visit Turku Archipelago (tourist office) is also represented. Before the main ECCB2024 conference, a satellite meeting, seven workshops, and nine tutorials are organized. The satellite meeting is organized by the Student Council of the International Society for Computational Biology (ISCB) by and for early-stage researchers. This 8th European Student Council Symposium (ESCS) continues the successful collaboration between ESCS and ECCB, building the future of research in computational biology. The 16 workshops/tutorials preceding the ECCB2024 main conference were selected out of a total of 24 proposals by the workshops and tutorials chair Bengt Persson (Uppsala University, Sweden), each running for half a day. The workshops foster discussions and exchange of ideas on a range of specialized or emerging topics in computational biology and provide opportunities to share practical experiences. The tutorials provide participants with lectures and hands-on training to enable learning about new areas of computational biology or important established topics. Following the tradition of previous ECCB conferences, the ISCB and ECCB sponsored a number of travel fellowships for students and postdoctoral fellows. These fellowships were primarily awarded to members presenting talks and those from low- or middle-income countries, facilitating their participation in the conference. A total of 68 applications were received from scientists across 26 countries. After a careful review, the ECCB2024 Organizing Committee in collaboration with the ISCB and ECCB awarded 10 and 18 fellowships, respectively. The ECCB2024 Code of Conduct is designed to ensure a safe and respectful environment for all attendees, outlining clear standards of behaviour to foster an inclusive conference atmosphere. In addition, the code provides specific procedures for attendees to follow if they feel these standards have been violated. This ensures that all participants have access to support and resources to address any concerns, reinforcing the commitment of ECCB2024 to uphold the highest ethical and professional standards during the conference. The ECCB2024 Organizing Committee is strongly committed to promote gender equity throughout the conference planning and execution. In particular, an important effort was made to ensure that the review panels included a balanced composition of female and male professionals across the world. In addition, three out of the seven distinguished keynote speakers are women. The conference program also features a collaborative workshop with the Bioinfo4Women initiative, discussing sex and gender bias in artificial intelligence. We would like to thank everyone who has contributed to the success of ECCB2024, ensuring it meets high standards of excellence. Special recognition is due to the Program Committee, including theme area chairs and reviewers, whose critical and dedicated efforts have been fundamental to the conference organization in selecting excellent tutorials, workshops, manuscripts, talks, and posters. We are equally thankful to the ECCB steering committee for their invaluable support and advice, especially ECCB2022 organizers for providing detailed information about the organization of the previous ECCB conference. The collaboration with the ISCB has been crucial in offering travel fellowships and promoting ECCB2024 globally, for which we are deeply grateful. Our gratitude also goes to all our financial sponsors, including our co-organizing sponsor ELIXIR. We also acknowledge the Oxford University Press production team for their work on the ECCB2024 Proceedings issue. Many individuals have played significant roles in the local organization, often exceeding their responsibilities, and we are immensely thankful for their dedication. Lastly, the conference would not be what it is without the diverse participants from around the world, whose scientific contributions, presentations, and discussions enrich ECCB2024. Thank you all for being there and for allowing us to enjoy science at ECCB2024 in Turku! None declared. This paper was published as part of a supplement financially supported by ECCB2024. All data are incorporated into the article. Additional information is available at the conference website https://eccb2024.fi.
Anu Kukkonen-Macchi, Sampsa Hautaniemi, Katharina F. Heil, Merja Heinäniemi, Lars Juhl Jensen, Sini Junttila, Lukas Käll, Asta Laiho, Peter Maccallum, Matti Nykter, Bengt Persson, Tomi Suomi, Tim Van Den Bossche, Tommi H. Nyrönen, Laura Elo
Bioinform.12
2022 PhosPiR: an automated phosphoproteomic pipeline in R
abstract
Large-scale phosphoproteome profiling using mass spectrometry (MS) provides functional insight that is crucial for disease biology and drug discovery. However, extracting biological understanding from these data is an arduous task requiring multiple analysis platforms that are not adapted for automated high-dimensional data analysis. Here, we introduce an integrated pipeline that combines several R packages to extract high-level biological understanding from large-scale phosphoproteomic data by seamless integration with existing databases and knowledge resources. In a single run, PhosPiR provides data clean-up, fast data overview, multiple statistical testing, differential expression analysis, phosphosite annotation and translation across species, multilevel enrichment analyses, proteome-wide kinase activity and substrate mapping and network hub analysis. Data output includes graphical formats such as heatmap, box-, volcano- and circos-plots. This resource is designed to assist proteome-wide data mining of pathophysiological mechanism without a need for programming knowledge.
Ye Hong, Dani Flinkman, Tomi Suomi, Sami Pietilä, Peter James, Eleanor Coffey, Laura Elo
Briefings Bioinform.3
2022 Correction to: PhosPiR: an automated phosphoproteomic pipeline in R
abstract
This is a correction to: Ye Hong, Dani Flinkman, Tomi Suomi, Sami Pietilä, Peter James, Eleanor Coffey, Laura L Elo, PhosPiR: an automated phosphoproteomic pipeline in R, Briefings in Bioinformatics, Volume 23, Issue 1, January 2022, bbab510, https://doi.org/10.1093/bib/bbab510. In the originally published version of this manuscript, there was an error in the equal contribution author note. The equal contribution statement should apply to Dani Flinkman, Tomi Suomi, and Sami Pietilä only. This error has been corrected. The publisher apologizes for the error.
Ye Hong, Dani Flinkman, Tomi Suomi, Sami Pietilä, Peter James, Eleanor Coffey, Laura Elo
Briefings Bioinform.3
2018 A systematic evaluation of normalization methods in quantitative label-free proteomics
abstract
To date, mass spectrometry (MS) data remain inherently biased as a result of reasons ranging from sample handling to differences caused by the instrumentation. Normalization is the process that aims to account for the bias and make samples more comparable. The selection of a proper normalization method is a pivotal task for the reliability of the downstream analysis and results. Many normalization methods commonly used in proteomics have been adapted from the DNA microarray techniques. Previous studies comparing normalization methods in proteomics have focused mainly on intragroup variation. In this study, several popular and widely used normalization methods representing different strategies in normalization are evaluated using three spike-in and one experimental mouse label-free proteomic data sets. The normalization methods are evaluated in terms of their ability to reduce variation between technical replicates, their effect on differential expression analysis and their effect on the estimation of logarithmic fold changes. Additionally, we examined whether normalizing the whole data globally or in segments for the differential expression analysis has an effect on the performance of the normalization methods. We found that variance stabilization normalization (Vsn) reduced variation the most between technical replicates in all examined data sets. Vsn also performed consistently well in the differential expression analysis. Linear regression normalization and local regression normalization performed also systematically well. Finally, we discuss the choice of a normalization method and some qualities of a suitable normalization method in the light of the results of our evaluation.
Tommi Välikangas, Tomi Suomi, Laura Elo
Briefings Bioinform.2
2018 A comprehensive evaluation of popular proteomics software workflows for label-free proteome quantification and imputation
abstract
Label-free mass spectrometry (MS) has developed into an important tool applied in various fields of biological and life sciences. Several software exist to process the raw MS data into quantified protein abundances, including open source and commercial solutions. Each software includes a set of unique algorithms for different tasks of the MS data processing workflow. While many of these algorithms have been compared separately, a thorough and systematic evaluation of their overall performance is missing. Moreover, systematic information is lacking about the amount of missing values produced by the different proteomics software and the capabilities of different data imputation methods to account for them.In this study, we evaluated the performance of five popular quantitative label-free proteomics software workflows using four different spike-in data sets. Our extensive testing included the number of proteins quantified and the number of missing values produced by each workflow, the accuracy of detecting differential expression and logarithmic fold change and the effect of different imputation and filtering methods on the differential expression results. We found that the Progenesis software performed consistently well in the differential expression analysis and produced few missing values. The missing values produced by the other software decreased their performance, but this difference could be mitigated using proper data filtering or imputation methods. Among the imputation methods, we found that the local least squares (lls) regression imputation consistently increased the performance of the software in the differential expression analysis, and a combination of both data filtering and local least squares imputation increased performance the most in the tested data sets.
Tommi Välikangas, Tomi Suomi, Laura Elo
Briefings Bioinform.2
2018 Phosphonormalizer: an R package for normalization of MS-based label-free phosphoproteomics
abstract
Motivation: Global centering-based normalization is a commonly used normalization approach in mass spectrometry-based label-free proteomics. It scales the peptide abundances to have the same median intensities, based on an assumption that the majority of abundances remain the same across the samples. However, especially in phosphoproteomics, this assumption can introduce bias, as the samples are enriched during sample preparation which can mask the underlying biological changes. To address this possible bias, phosphopeptides quantified in both enriched and non-enriched samples can be used to calculate factors that mitigate the bias. Results: We present an R package phosphonormalizer for normalizing enriched samples in label-free mass spectrometry-based phosphoproteomics. Availability and implementation: The phosphonormalizer package is freely available under GPL ( > =2) license from Bioconductor (https://bioconductor.org/packages/phosphonormalizer). Contact: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Sohrab Saraei, Tomi Suomi, Otto Kauko, Laura Elo
Bioinform.2
2018 SimPhospho: a software tool enabling confident phosphosite assignment
abstract
Motivation: Mass spectrometry combined with enrichment strategies for phosphorylated peptides has been successfully employed for two decades to identify sites of phosphorylation. However, unambiguous phosphosite assignment is considered challenging. Given that site-specific phosphorylation events function as different molecular switches, validation of phosphorylation sites is of utmost importance. In our earlier study we developed a method based on simulated phosphopeptide spectral libraries, which enables highly sensitive and accurate phosphosite assignments. To promote more widespread use of this method, we here introduce a software implementation with improved usability and performance. Results: We present SimPhospho, a fast and user-friendly tool for accurate simulation of phosphopeptide tandem mass spectra. Simulated phosphopeptide spectral libraries are used to validate and supplement database search results, with a goal to improve reliable phosphoproteome identification and reporting. The presented program can be easily used together with the Trans-Proteomic Pipeline and integrated in a phosphoproteomics data analysis workflow. Availability and implementation: SimPhospho is open source and it is available for Windows, Linux and Mac operating systems. The software and its user's manual with detailed description of data analysis as well as test data can be found at https://sourceforge.net/projects/simphospho/. Supplementary information: Supplementary data are available at Bioinformatics online.
Veronika Suni, Tomi Suomi, Tomoya Tsubosaka, Susumu Y. Imanishi, Laura Elo, Garry L. Corthals
Bioinform.2
2017 ROTS: An R package for reproducibility-optimized statistical testing
abstract
Differential expression analysis is one of the most common types of analyses performed on various biological data (e.g. RNA-seq or mass spectrometry proteomics). It is the process that detects features, such as genes or proteins, showing statistically significant differences between the sample groups under comparison. A major challenge in the analysis is the choice of an appropriate test statistic, as different statistics have been shown to perform well in different datasets. To this end, the reproducibility-optimized test statistic (ROTS) adjusts a modified t-statistic according to the inherent properties of the data and provides a ranking of the features based on their statistical evidence for differential expression between two groups. ROTS has already been successfully applied in a range of different studies from transcriptomics to proteomics, showing competitive performance against other state-of-the-art methods. To promote its widespread use, we introduce here a Bioconductor R package for performing ROTS analysis conveniently on different types of omics data. To illustrate the benefits of ROTS in various applications, we present three case studies, involving proteomics and RNA-seq data from public repositories, including both bulk and single cell data. The package is freely available from Bioconductor (https://www.bioconductor.org/packages/ROTS).
Tomi Suomi, Fatemeh Seyednasrollah, Maria K. Jaakkola, Thomas Faux, Laura Elo
PLoS Comput. Biol.1