EDBT 2026 Demo / reviewers in the wild / expert
Laura Elo
dblp:62/2092 · also Laura L. Elo
· DBLP profile ↗
28ranked-venue papers
3as first author
13since 2021 · last 2026
0000-0001-5648-4532ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 28 · 3 first-author · 13 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | REACTOR: REgulon Activity analysis and Comparison Tool for single-cell transcriptOmics ResearchabstractSUMMARY: We introduce REACTOR, a computational tool designed to detect differential activity of transcriptional regulators and their target genes (regulons) in single-cell RNA-sequencing data. It expands the currently available framework for regulon analysis by introducing a robust statistical test to detect differential regulon activity between conditions, such as disease versus control, with multiple replicates. By contrasting different conditions, REACTOR enables identification of key condition- and cell type-specific regulons. To demonstrate the use of REACTOR, we illustrate its performance in a publicly available COVID-19 dataset. AVAILABILITY: REACTOR R-package together with an implementation vignette are available at https://www.github.com/elolab/REACTOR. Markus Lindén, Sebastian I Zúñiga Norman, Tommi Välikangas, Sini Junttila, Tomi Suomi, Kalle T. Rytkönen, Laura Elo |
Bioinform. | 7 |
| 2026 | Clarifying the scope and capabilities of ROTS in differential expression analysisabstractSUMMARY: Recently, Anwar et al. introduced a method combining the ROTS reproducibility optimisation procedure with empirical Bayes variance estimation from limma. Here, we clarify several methodological aspects to support accurate interpretation of the results. We emphasise that ROTS is a general reproducibility optimisation framework rather than a single statistical test and demonstrate that benchmarking outcomes in the reported spike-in case studies are highly sensitive to analysis and evaluation choices. Furthermore, our reanalyses of the spike-in datasets do not support the reported conclusions, and we were unable to reproduce the results of the clinical Alzheimer's disease case study. These findings highlight the importance of transparent benchmarking practices and careful interpretation of comparative results. AVAILABILITY AND IMPLEMENTATION: The ROTS package is available through Bioconductor. The reanalyses were performed using the original code, with the minimal additions described in the manuscript. Tomi Suomi, Jalmari Kettunen, Taneli Pusa, Laura Elo |
Bioinform. | 4 |
| 2024 | ECCB2024: The 23rd European Conference on Computational BiologyabstractThis volume of Bioinformatics includes the proceedings papers of the 23rd European Conference on Computational Biology (ECCB2024) to be held in Turku, Finland, from 16 September to 20 September 2024, under the theme Data and Algorithms for Health and Science. More information on the ECCB2024 conference is available at the conference website https://eccb2024.fi/. ECCB is one of the main international conferences in the field of computational biology and bioinformatics together with the Intelligent Systems for Molecular Biology (ISMB) and the Research in Computational Molecular Biology (RECOMB). It is held jointly with the ISMB conference in odd-numbered years and independently in even-numbered years. ECCB attracts scientists and industry professionals from diverse disciplines, including mathematics, statistics, computer science, biology, and medicine. Rapid technological advancements enable life scientists to gain increasingly detailed insights in complex biological systems. These improvements in measurement technologies, however, introduce new challenges for data interpretation and necessitate the development of advanced computational techniques to manage the complexity and volume of the data. Consequently, the field of computational biology is rapidly evolving with new algorithms, software, and databases. The ECCB2024 conference showcases cutting-edge developments in systems biology, artificial intelligence, single-cell and spatial technologies, data integration, and more, addressing the growing demand for sophisticated algorithms to enhance the analysis of large-scale biological and biomedical datasets. This is highlighted by the most frequent keywords accompanying all the accepted submissions in ECCB2024 (tutorials/workshops, proceedings, highlight talks, posters), with the most popular keywords including ‘machine learning’, ‘deep learning’, ‘single-cell rnaseq’, and ‘multiomics’ (Fig. 1). Most frequent keywords among all accepted submissions in ECCB2024. The barplot shows the frequencies of top keywords appearing in at least ten submissions, while the word cloud illustrates the relative frequencies of all keywords appearing in at least five submissions. The ECCB2024 edition features five keynote lectures by distinguished speakers: Sarah Teichmann (Cambridge Stem Cell Institute, University of Cambridge, UK), Peer Bork (EMBL—European Molecular Biology Laboratory, Germany), Ileana Cristea (Princeton University, US), Jussi Taipale (University of Cambridge, UK), and Fabian Theis (Helmholtz Munich Computational Health Center, Germany). In addition, ECCB2024 hosts a scientific debate on data sharing and privacy protection by Melissa Haendel (University of North Carolina, US) and Yves Moreau (KU Leuven, Belgium). The ECCB2024 conference covers a wide range of topics, focusing on methodological advancements in computational biology as well as innovative application of computational techniques to life sciences and medicine. To provide a cohesive overview of recent scientific progress, the conference presentations are organized under six broad themes: (i) Genomes, (ii) Proteins, (iii) Systems biology and multiomics, (iv) Single-cell omics, (v) Microbiomes and planetary health, and (vi) Digital health. The proceedings talks present new scientific contributions, while the highlight talks showcase already published cutting-edge science in computational biology and related further developments. Poster presentations provide an opportunity for the participants to discuss their recent work with other researchers in the field. Workshops and tutorials on specialized topics prior to the main conference program are platforms to share practical experiences and learn new skills. In addition to scientific presentation tracks, ECCB2024 has a separate ELIXIR track, overseen by ELIXIR, focusing on advancements in infrastructure and services within ELIXIR nodes in support of the expert groups known as ‘ELIXIR Communities’. ELIXIR is a distributed pan-European life science infrastructure for biological data that coordinates, integrates, and sustains bioinformatics resources across its member states. This enables academic and industry users to access data, tools, standards, computing, and training services for life science research. ELIXIR has selected ECCB as a primary dissemination platform, serving as a co-organizing sponsor. Apart from Community activities the track will also introduce three scientific areas as outlined in the 2024–2028 ELIXIR Scientific Programme (https://elixir-europe.org), enabling scientists to access and analyse life science data across ‘Cellular and Molecular Research’, ‘Biodiversity, food security & pathogens’, and ‘Human data & translational research’. ECCB2024 received an impressive number of 200 submissions of full manuscripts for the proceedings call, highlighting the growing interest and engagement within the community. These submissions were organized under the six conference themes. Each submission was subjected to a peer-review process, with at least two reviews per manuscript, managed by the members of the ECCB2024 Programme Committee. The Programme Committee was chaired by Laura Elo (University of Turku, Finland) and each conference theme had an area chair (Table 1). The Proceedings Review Committee had over 100 reviewers. Thematic areas of ECCB2024 proceedings talks.a The table lists the area chairs for each theme, the number of reviewed papers, the number of accepted papers, and the acceptance rate for each theme. Thematic areas of ECCB2024 proceedings talks.a The table lists the area chairs for each theme, the number of reviewed papers, the number of accepted papers, and the acceptance rate for each theme. The review process focused on the impact, reproducibility, and scientific quality of the submitted research, as well as its relevance, interest, and value for the ECCB2024 audience. Upon completion of the review, the Program Committee chairs selected 24 papers to be included in the ECCB2024 proceedings, with an acceptance rate of 12% across the different areas (Table 1). All proceedings papers are published open access in this Proceedings issue of September 2024 of Bioinformatics. The ECCB2024 call for Highlight Talks invited presentations of studies recently published in scientific journals since 1 March 2023, or accepted for publication. The existence of further developments related to the published paper was also positively considered. The call received a total of 212 proposals, which were ranked according to their relevance and impact on computational biology, as well as their suitability to be presented to the large and diverse audience. Ultimately, the Program Committee selected 23% of the submissions for presentation as a Highlight talk at the conference. ECCB2024 hosts two poster sessions where researchers can introduce and discuss their work under the conference themes. In total, nearly 700 posters were accepted to be presented. The poster submissions were reviewed by the Posters Committee: Sini Junttila, Asta Laiho, Tomi Suomi, and Sampsa Hautaniemi (late posters). In addition, the ELIXIR track selected posters to be presented at ECCB2024. Overall, the ECCB2024 tracks attracted participation of presenting authors from 48 countries. The countries with the highest number of presenting authors were Germany, Finland, the USA, the United Kingdom, and Spain, jointly contributing to half of the overall number (Fig. 2). Following these were Turkey, France, and China, each contributing 4%–5% of the presenting authors. Proportion of countries where the presenting author of each accepted submission had their primary affiliation. Countries with the proportion below 2% were grouped under ‘Other’. Exhibitor booths are open throughout the conference, including ELIXIR, ISCB, Oxford University Press, Royal Society Publishing and eLife Sciences Publications Ltd, as well as the two organizing institutions: University of Turku and CSC—IT Center for Science. The booths showcase the latest scientific literature in computational biology and bioinformatics, data stewardship, as well as new developments in hardware, software, and technology. In addition, the Visit Turku Archipelago (tourist office) is also represented. Before the main ECCB2024 conference, a satellite meeting, seven workshops, and nine tutorials are organized. The satellite meeting is organized by the Student Council of the International Society for Computational Biology (ISCB) by and for early-stage researchers. This 8th European Student Council Symposium (ESCS) continues the successful collaboration between ESCS and ECCB, building the future of research in computational biology. The 16 workshops/tutorials preceding the ECCB2024 main conference were selected out of a total of 24 proposals by the workshops and tutorials chair Bengt Persson (Uppsala University, Sweden), each running for half a day. The workshops foster discussions and exchange of ideas on a range of specialized or emerging topics in computational biology and provide opportunities to share practical experiences. The tutorials provide participants with lectures and hands-on training to enable learning about new areas of computational biology or important established topics. Following the tradition of previous ECCB conferences, the ISCB and ECCB sponsored a number of travel fellowships for students and postdoctoral fellows. These fellowships were primarily awarded to members presenting talks and those from low- or middle-income countries, facilitating their participation in the conference. A total of 68 applications were received from scientists across 26 countries. After a careful review, the ECCB2024 Organizing Committee in collaboration with the ISCB and ECCB awarded 10 and 18 fellowships, respectively. The ECCB2024 Code of Conduct is designed to ensure a safe and respectful environment for all attendees, outlining clear standards of behaviour to foster an inclusive conference atmosphere. In addition, the code provides specific procedures for attendees to follow if they feel these standards have been violated. This ensures that all participants have access to support and resources to address any concerns, reinforcing the commitment of ECCB2024 to uphold the highest ethical and professional standards during the conference. The ECCB2024 Organizing Committee is strongly committed to promote gender equity throughout the conference planning and execution. In particular, an important effort was made to ensure that the review panels included a balanced composition of female and male professionals across the world. In addition, three out of the seven distinguished keynote speakers are women. The conference program also features a collaborative workshop with the Bioinfo4Women initiative, discussing sex and gender bias in artificial intelligence. We would like to thank everyone who has contributed to the success of ECCB2024, ensuring it meets high standards of excellence. Special recognition is due to the Program Committee, including theme area chairs and reviewers, whose critical and dedicated efforts have been fundamental to the conference organization in selecting excellent tutorials, workshops, manuscripts, talks, and posters. We are equally thankful to the ECCB steering committee for their invaluable support and advice, especially ECCB2022 organizers for providing detailed information about the organization of the previous ECCB conference. The collaboration with the ISCB has been crucial in offering travel fellowships and promoting ECCB2024 globally, for which we are deeply grateful. Our gratitude also goes to all our financial sponsors, including our co-organizing sponsor ELIXIR. We also acknowledge the Oxford University Press production team for their work on the ECCB2024 Proceedings issue. Many individuals have played significant roles in the local organization, often exceeding their responsibilities, and we are immensely thankful for their dedication. Lastly, the conference would not be what it is without the diverse participants from around the world, whose scientific contributions, presentations, and discussions enrich ECCB2024. Thank you all for being there and for allowing us to enjoy science at ECCB2024 in Turku! None declared. This paper was published as part of a supplement financially supported by ECCB2024. All data are incorporated into the article. Additional information is available at the conference website https://eccb2024.fi. Anu Kukkonen-Macchi, Sampsa Hautaniemi, Katharina F. Heil, Merja Heinäniemi, Lars Juhl Jensen, Sini Junttila, Lukas Käll, Asta Laiho, Peter Maccallum, Matti Nykter, Bengt Persson, Tomi Suomi, Tim Van Den Bossche, Tommi H. Nyrönen, Laura Elo |
Bioinform. | 15 |
| 2023 | Cell-connectivity-guided trajectory inference from single-cell dataabstractMOTIVATION: Single-cell RNA-sequencing enables cell-level investigation of cell differentiation, which can be modelled using trajectory inference methods. While tremendous effort has been put into designing these methods, inferring accurate trajectories automatically remains difficult. Therefore, the standard approach involves testing different trajectory inference methods and picking the trajectory giving the most biologically sensible model. As the default parameters are often suboptimal, their tuning requires methodological expertise. RESULTS: We introduce Totem, an open-source, easy-to-use R package designed to facilitate inference of tree-shaped trajectories from single-cell data. Totem generates a large number of clustering results, estimates their topologies as minimum spanning trees, and uses them to measure the connectivity of the cells. Besides automatic selection of an appropriate trajectory, cell connectivity enables to visually pinpoint branching points and milestones relevant to the trajectory. Furthermore, testing different trajectories with Totem is fast, easy, and does not require in-depth methodological knowledge. AVAILABILITY AND IMPLEMENTATION: Totem is available as an R package at https://github.com/elolab/Totem. Johannes Smolander, Sini Junttila, Laura Elo |
Bioinform. | 3 |
| 2023 | VarSCAT: A computational tool for sequence context annotations of genomic variantsabstractThe sequence contexts of genomic variants play important roles in understanding biological significances of variants and potential sequencing related variant calling issues. However, methods for assessing the diverse sequence contexts of genomic variants such as tandem repeats and unambiguous annotations have been limited. Herein, we describe the Variant Sequence Context Annotation Tool (VarSCAT) for annotating the sequence contexts of genomic variants, including breakpoint ambiguities, flanking bases of variants, wildtype/mutated DNA sequences, variant nomenclatures, distances between adjacent variants, tandem repeat regions, and custom annotation with user customizable options. Our analyses demonstrate that VarSCAT is more versatile and customizable than the currently available methods or strategies for annotating variants in short tandem repeat (STR) regions or insertions and deletions (indels) with breakpoint ambiguity. Variant sequence context annotations of high-confidence human variant sets with VarSCAT revealed that more than 75% of all human individual germline and clinically relevant indels have breakpoint ambiguities. Moreover, we illustrate that more than 80% of human individual germline small variants in STR regions are indels and that the sizes of these indels correlated with STR motif sizes. VarSCAT is available from https://github.com/elolab/VarSCAT. Ning Wang 0032, Sofia Khan, Laura Elo |
PLoS Comput. Biol. | 3 |
| 2022 | PhosPiR: an automated phosphoproteomic pipeline in RabstractLarge-scale phosphoproteome profiling using mass spectrometry (MS) provides functional insight that is crucial for disease biology and drug discovery. However, extracting biological understanding from these data is an arduous task requiring multiple analysis platforms that are not adapted for automated high-dimensional data analysis. Here, we introduce an integrated pipeline that combines several R packages to extract high-level biological understanding from large-scale phosphoproteomic data by seamless integration with existing databases and knowledge resources. In a single run, PhosPiR provides data clean-up, fast data overview, multiple statistical testing, differential expression analysis, phosphosite annotation and translation across species, multilevel enrichment analyses, proteome-wide kinase activity and substrate mapping and network hub analysis. Data output includes graphical formats such as heatmap, box-, volcano- and circos-plots. This resource is designed to assist proteome-wide data mining of pathophysiological mechanism without a need for programming knowledge. Ye Hong, Dani Flinkman, Tomi Suomi, Sami Pietilä, Peter James, Eleanor Coffey, Laura Elo |
Briefings Bioinform. | 7 |
| 2022 | Correction to: PhosPiR: an automated phosphoproteomic pipeline in RabstractThis is a correction to: Ye Hong, Dani Flinkman, Tomi Suomi, Sami Pietilä, Peter James, Eleanor Coffey, Laura L Elo, PhosPiR: an automated phosphoproteomic pipeline in R, Briefings in Bioinformatics, Volume 23, Issue 1, January 2022, bbab510, https://doi.org/10.1093/bib/bbab510. In the originally published version of this manuscript, there was an error in the equal contribution author note. The equal contribution statement should apply to Dani Flinkman, Tomi Suomi, and Sami Pietilä only. This error has been corrected. The publisher apologizes for the error. Ye Hong, Dani Flinkman, Tomi Suomi, Sami Pietilä, Peter James, Eleanor Coffey, Laura Elo |
Briefings Bioinform. | 7 |
| 2022 | Estimating cell type-specific differential expression using deconvolutionabstractWhen a condition causes gene expression changes in a mixed tissue sample containing different types of cells, the changes can originate from either altered cell type composition or altered expression in some of the cell types. For example, in the context of type 1 diabetes it is an open debate if the pancreatic beta cells are dead (altered composition) or if they have, at least initially, just stopped insulin production (altered expression) [1–3]. Computational deconvolution is a free alternative to single cell and fluorescence-activated cell sorting (FACS) analyses to obtain cell type-specific information from readily analyzed bulk samples. Besides financial motivation, it also has the advantage of being applicable to old datasets, which are possibly very difficult to re-analyze in an experimental manner. There are many publicly available tools for computational deconvolution [4–6], and they have different input requirements and goals. In deconvolution, the bulk expression of a gene is typically considered as a linear combination of its expression levels from different cell types present in the sample, i.e. |$E = S \cdot C$|, where |$E$| is the observed bulk expression matrix with genes as rows and samples as columns, |$S$| is a cell type-specific expression matrix indicating how strongly each pure cell type (columns) expresses each gene (rows), and |$C$| is a cell type proportion matrix with cell types as rows and samples as columns. ‘Composition deconvolution’ (e.g. [7–12]) aims to estimate cell type composition (either proportions |$C$| or abundances) in different bulk samples, whereas ‘expression deconvolution’ (e.g. [13–18]) estimates cell type-specific gene expression profiles |$S$| (csGEPs). ‘Complete methods’ (e.g. [19–22]) do both tasks simultaneously. As mentioned above, the cell type-specific expression profiles might be altered by some factors (e.g. age, gender, disease) and, in such case, the assumption of the matrix |$S$| containing csGEPs as columns is overly simplistic. An important subtask related to expression deconvolution is to explore cell type-specific differentially expressed genes (csDEGs) between conditions. This can be done at three levels: associating differentially expressed genes (DEGs) detected from the bulk data with different cell types, identifying directly csDEGs, or defining csGEPs for each sample separately, i.e. personalized |$S$|. Personalizing cell type-specific expression is a very ambitious goal and there are only few tools available, with restrictions in their applicability. For example, ISOpure [23] is related to purifying tumor samples from the effect of immune cells, whereas CIBERSORTx [24] provides the personalization option only for datasets of limited size. Associating bulk findings with cell types is a more relaxed goal and, for instance, CellCODE [25] offers this option among other related functions. The downside of investigating bulk findings is that if cell type composition is also altered between sample groups, bulk findings are likely to contain plenty of genes that are strongly expressed by the cell type(s) with altered proportion rather than genes with altered expression in some cell types. Another issue is that genes altered in dominating cell types are likely to mask the DEGs of rare cell types at the bulk level, as discussed in [25]. Notably, csGEPs refer to columns of matrix |$S$| and the term is frequently used instead of |$S$| in this article. The issue with the term |$S$| is that it is technically no longer a matrix, but a three-dimensional tensor or a list of matrices if it is defined for samples or sample groups separately. Identifying csDEGs is a more relaxed goal than fully personalized |$S$|, but it provides more detailed insights than associating bulk DEGs with cell types. Few methods have been developed for the task [17, 25–27], but they have not been systematically investigated in the literature. Although composition deconvolution and expression deconvolution methods have been empirically compared [18, 28–30], there are no guidelines for selecting and using methods to identify csDEGs. Here, we address this issue by comparing nine different approaches, namely TOAST, csSAM, LRCDE, CARseq, Rodeo, qprog, CellDMC, TCA and DESeq2, from different practical perspectives, investigate which factors affect the accuracy and how much and offer insight when the end user can expect good results and when not. Here we evaluate nine methods for identifying csDEGs: tools for the analysis of heterogeneous tissues (TOAST) [26], cell type-specific significance analysis of microarrays (csSAM) [17], linear regression cell type-specific differential expression (LRCDE) [27], cell type aware analysis for RNA-seq data (CARseq) [31], robust deconvolution (Rodeo) [18], quadratic programming (qprog) [32], CellDMC [33], tensor composition analysis (TCA) [34] and DESeq2 [35]. Among these, TOAST, csSAM, LRCDE and CARseq are originally developed for detecting csDEGs from RNAseq data, whereas Rodeo and qprog are originally developed for expression deconvolution but we test here their utility to identify csDEGs. Methods CellDMC and TCA are originally designed for methylation data and DESeq2 represents a model not developed for deconvolution purposes of any kind. It can still be applied by defining cell type proportions and interaction terms of them and disease status as covariates. Further details about the methods and how any expression deconvolution method can be used for identifying csDEGs are available in section ‘Tested methods’. Besides accuracy, we tested the sensitivity of the methods to different factors (e.g. individual heterogeneity of csGEPs and outlier samples in the data), and how the end user can evaluate whether the obtained results are accurate. We utilized three semi-simulated datasets addressed as GSE60424, EMTAB9221 and GSE124742 in these tests. The datasets were constructed by first generating csGEPs with realistic individual variation and DEGs between two sample groups, 100 samples each. For the evaluation purposes, these csGEPs were then combined into bulk samples by calculating their sum weighted by the cell type proportions. As gold standard csDEGs, we considered DEGs identified using the csGEPs (false discovery rate (FDR) |$\leq $| 0.05). Datasets GSE60424 and EMTAB9221 are based on data from blood samples with different numbers of cell types present in them, whereas dataset GSE124742 involves measurements from pancreatic tissue samples. Further details regarding the datasets are available in section ‘Test data’. To evaluate the accuracy of the estimated csDEGs, we considered the overlap between the estimated csDEGs and the gold standard csDEGs. Among the tested methods, LRCDE detected over 5000 csDEGs with FDR |$\leq $| 0.05 from all tested semi-simulated datasets and all cell types. The contrast to the other tested methods was considerable; typically the methods identified fewer csDEGs than present in the corresponding gold standard. The opposite extreme was csSAM, which did not produce any detections with FDR |$\leq $| 0.05 from any cell type from any of the datasets. The exact numbers of significant findings from the different methods and cell types are available in Table S1. Because of the large variation in the number of significant detections, we compared the top most significant findings to the gold standard. The size of the evaluated top list was defined as the number of detections in the gold standard. Table 1 lists the overlaps between the known and estimated csDEGs in the different cell types and datasets. Notably, when the FDR cutoff 0.05 was used instead of the fixed top list size, the methods typically had high precision and low recall (Section 1 in Supplementary text), suggesting that the estimated csDEGs were correct, but many true csDEGs remained The proportion of genes between the known and estimated cell type differentially expressed The top list size, i.e. the number of gold standard detections, is in the the cell type DESeq2 is designed for it is not for dataset GSE60424 based on data and TCA we not for dataset GSE124742 results are The proportion of genes between the known and estimated cell type differentially expressed The top list size, i.e. the number of gold standard detections, is in the the cell type DESeq2 is designed for it is not for dataset GSE60424 based on data and TCA we not for dataset GSE124742 results are For dataset GSE60424, methods TOAST, CellDMC and TCA had the For dataset CARseq and TCA were the most For dataset TOAST, Rodeo, qprog, and CellDMC the LRCDE and DESeq2 estimates for csDEGs than the other methods, but between the methods’ were as in Table between the methods is not the only that can be from the in Table When the are compared the cell type proportion information (Section in Supplementary text), it can be that the cell type proportions the accuracy of the i.e. csDEGs from rare cell types were to than from As in the Supplementary cells, cells, and and cells had the proportion in datasets GSE60424, EMTAB9221 and cell types had also the in the corresponding datasets In the of and cells in the GSE124742 data with proportion of only cells were likely to than cells they are cells and beta cells, whereas cells are from i.e. they likely more from and beta In of Supplementary the effect of cell type is investigated in a more as we altered the proportion of cells in dataset EMTAB9221 and evaluated how the accuracy to csDEGs the cell type results the of cell type proportion and of the methods had accuracy when the proportion was to type proportion has also been to affect the accuracy of estimated csGEPs in expression deconvolution Although most of the methods were to few CARseq and expression deconvolution methods Rodeo and qprog were more For the with dataset EMTAB9221 for different methods were and Notably, as in section ‘Tested the of expression deconvolution methods Rodeo and qprog to be by the number of for in this to the for detecting csDEGs. the the can not be as in of Supplementary we investigated how the individual variation in csGEPs over samples the accuracy of detecting csDEGs. We the of i.e. of standard and to a fixed by the standard of the gene when the csGEPs for the bulk of variation of and were tested in dataset and they were the for all To investigate the effect of individual variation in only cell we also data where only the of variation were other cell types had their standard from gene to As this test was only methods TOAST, csSAM, CellDMC and TCA were results that individual heterogeneity had a on the accuracy with all the tested The was still when only cell type was altered and which that very heterogeneous cell type can also the results from other cell types. can be from dataset GSE124742 (Section of Supplementary The of variation over samples had of and in dataset EMTAB9221 of Supplementary for of variation from all cell types and The of variation in cells compared to cells the in the accuracy of csDEGs detected from them for all the tested of genes between the known and estimated top csDEGs when the of variation over was for each cell The was done by the standard of all cell types and or only of and whereas the other cell types had their The test was done with methods, and and CellDMC and and TCA and of genes between the known and estimated top csDEGs when the of variation over was for each cell The was done by the standard of all cell types and or only of and whereas the other cell types had their The test was done with methods, and and CellDMC and and TCA and this the cell type proportions and variation over samples as important factors the accuracy of the In we tested other important but to their was compared to cell type proportion and individual sample groups (Section in Supplementary text), of the present cell types (Section in Supplementary text), and from very rare cell types (Section in Supplementary Although the input matrix |$C$| the rare and difficult to cell types, individual variation in csGEPs is to Here we guidelines on how to that Although the end user not have about the sample heterogeneity of |$S$|, some information can be by investigating In this we compared the of a dataset section for to the accuracy of the identified csDEGs. Although these do not only of individual variation in but from cell types also into them, the still some to estimate the accuracy of the results as in Although it is to a cutoff for of the cell types had good accuracy when the and in the accuracy cutoff the in was from all cell types or from only did not have a on the between the accuracy and for most cell types. were an for of the csDEGs as a of for all cell types in dataset The data with very were from the datasets with a fixed of of the csDEGs as a of for all cell types in dataset The data with very were from the datasets with a fixed of As the bulk data to be analyzed can contain samples, csGEPs from of the other samples, we tested how the different methods are to such outlier samples. For this outlier sample and outlier samples were into each dataset and the accuracy of the detected csDEGs was Rodeo is designed to be robust few outlier samples and with such samples it had the in datasets GSE60424 and In dataset TCA and Rodeo were the top The of the other tested methods did not the for CARseq, which from the in dataset EMTAB9221 and csSAM, which them in dataset GSE124742 when the datasets contain outlier GSE60424, EMTAB9221 and when the datasets contain outlier GSE60424, EMTAB9221 and we tested how the methods different numbers of outlier samples. To evaluate we outlier samples to As in Rodeo had also in this test the most robust but with number of more methods still Notably, for most of the tested methods, accuracy of the detected csDEGs was the most to in cells, which is the cell type with the accuracy The other cell were more robust of different methods when different numbers of outlier samples are into dataset in is to the of the of different methods when different numbers of outlier samples are into dataset in is to the of the As findings can be more robust than we tested the of the results using the estimated csDEGs. and CellDMC were from these as they do not changes csGEPs for both sample groups for have been in Here, we on cell types and datasets with significant findings from the known csDEGs. This with dataset EMTAB9221 and cells from of the estimated lists many i.e. when there were only if gold standard the number of findings from different methods were also The precision and recall for cell types and datasets with more than gold standard findings are in Table in these the precision to be than the recall indicating that findings were more of an issue than DESeq2 was the only method with recall or at least with precision in most to the csSAM, CARseq, Rodeo, and TCA LRCDE and DESeq2, for cells in dataset EMTAB9221 DESeq2 had by the recall The significant findings are and discussed in in Supplementary on these it that if many findings can be detected from the estimated csDEGs, both and the csDEGs are likely accurate. the opposite is not few or no findings can either csDEGs (e.g. estimates for cells csDEGs but no findings as there is not much to (e.g. findings for beta cells or csDEGs but no findings if there be significant (e.g. findings for a single method can be as the but TOAST, CARseq, CellDMC and TCA had the In of over outlier samples with altered Rodeo has more In the CARseq, it was with which might its compared to the other tested method was in datasets EMTAB9221 and which are based on than in GSE60424, which is based on TCA is an and robust but it in some The that it in sample with of variation but with variation that the method not data with large heterogeneity between the samples. methods LRCDE to few rather than many csDEGs if FDR cutoff 0.05 is but the detections are high results that csDEGs can be if the individual variation is which can be evaluated using and cell type proportion is not tested factors such as sample groups, of cell types present in the samples, and from rare cell types had on the analysis from the estimated csDEGs can two if there are many (e.g. more than significant detections, the findings are likely and about can be as in any many findings that the estimated csDEGs used as an input for the analysis are accurate. in there are no no can be from Notably, when the accuracy of the results from the information the end user can we used to bulk expression to robust the of the bulk expression in to by different of the only genes with the expression in the bulk data were used to the these for the for other for instance, samples from tissues with very few expressed genes might not the the of individual variation cell types is not the only for the It is not whether the bulk data to be analyzed a sample with a altered cell type composition likely has a bulk expression from the of the samples. it is still not an outlier in a that it causes for the tested methods, the csGEPs are also The contain the variation from csGEPs instead of investigating them can the outlier if the of or few samples are of than of the of the samples, they can be identified as and from the bulk data and the input matrix The of an outlier is a of its the method to estimate csDEGs to be in the We are results from dataset GSE60424 were than from the other datasets the of variation (Section in Supplementary over the samples for the different cell types. The dominating cell type has variation than the other cell types present in the samples, but the is true for cells in dataset that is not likely the the dataset has more cell types than the other test datasets but them by cells or cells as did not the Another is that the csDEGs in GSE60424 are than based on observed cell types in the other two datasets ‘Test the number of detected gold standard csDEGs was typically to the other datasets, which this be the very dominating of which the proportions of the other cell types low (Section in Supplementary The of this is that all test data are cell type-specific datasets with sample were not publicly available, sample size has been to affect the accuracy of the estimated csDEGs and expression deconvolution [18, Although we did to realistic in test data, it is that some present in data are factors be related to in some of the samples, which are to be in In the context of deconvolution, the of the bulk data is the the effect of such be but on the other if a sample has much cells with low expression (e.g. its low expression not be to a to other samples. The semi-simulated test data from the effect of on the accuracy of the estimated csDEGs. It be to if is than the other or if different methods from different The of the effect on the accuracy also be of In the it has been that is a for composition deconvolution but for expression deconvolution and csDEGs the is still open to Besides the methods tested there are few tools to csDEGs that were from this for For example, CellCODE [25] and have input requirements that between them and the methods tested here for related to the number of present cell types (e.g. and restrictions about or tissues (e.g. first and recall of the no findings with The in the the cell type the number of detections from the known csDEGs. TCA with dataset GSE124742 and methods and CellDMC do not changes csGEPs to for results are first and recall of the no findings with The in the the cell type the number of detections from the known csDEGs. TCA with dataset GSE124742 and methods and CellDMC do not changes csGEPs to for results are GSE60424 is based on RNAseq data from with different available in with the In the data, cells, cells, cells and cells are analyzed from with different conditions. to generating cell type proportions to realistic individual variation in we them from with the and standard over the samples as in the EMTAB9221 was from single cell RNAseq data available in It cells from blood of with different of and The sample |$S$| was by first calculating the expression profiles for each cell type and sample and them by to the of the expression csGEPs of samples were from with and standard the For sensitivity we also a bulk data from cells of type by from cells to the dataset GSE124742 was constructed from single cell RNAseq data and the data is available in the It measurements of pancreatic samples from with 1 disease and this we utilized data from cells from and with We |$S$| for each individual by selecting and cells from each cell cells were not to only cells in this we also a bulk cells as where the number of cells their observed proportion in the To evaluate the accuracy of the we used csDEGs detected by from the known csGEPs in the samples with FDR cutoff 0.05 as a gold standard. Although the csDEGs were to the in the data not all of them are from the sample and only that can be identified with are considered as a gold standard. We then the proportion of these known csDEGs and an number of top detections from the methods as a of with bulk expression in both sample groups were from each dataset the The of accuracy was used in all tests. Notably, DESeq2 is designed for it is not for dataset GSE60424 based on and TCA for dataset results are as To test the effect of individual variation in |$S$|, we investigated how the accuracy when we the variation over the samples. For this we constructed bulk data as in the accuracy but the of variation of csGEPs over the samples were to fixed levels of or The was used for all which is not the with of variation is defined as standard to the and we it into the by the standard the expression at its were done for each of of variation and the accuracy obtained from them is For the of this test was done only with TOAST, csSAM, CellDMC and and test was used when the gold standard. Datasets EMTAB9221 and GSE124742 were utilized for this results from are in this and results from GSE124742 in the Supplementary are used to We did this test that the individual variation all cell types was and that only cell type in EMTAB9221 and beta cells in was the had their standard over the To evaluate the effect of outlier samples, we sample and samples with altered csGEPs into each The cell type proportions for these samples are from the as the of the samples. For each cell type and data the csGEPs for the are from using either or for an outlier as a The for an outlier were to i.e. the is \cdot and the is \cdot where and are the first and the and = To evaluate the effect of the number of outlier samples on the accuracy of detected csDEGs, we or into the dataset As detections can be more robust than we tested their as The analysis was done with analysis as its input requirements are and it has in It changes of DEGs and a list of analyzed genes as an The changes were for each cell type as csGEPs for samples by csGEPs for samples. In the gold the sample csGEPs were defined as over sample csGEPs for both sample groups, and estimated csGEPs were readily at instead of We tested nine methods, all into LRCDE CARseq Rodeo qprog CellDMC TCA and DESeq2 cell type proportion matrix sample information and a bulk expression matrix were as for all of the Among the nine tested methods, have been designed to csDEGs. is a deconvolution with not utilized in this For example, it can into different sample (e.g. age, gender, disease it can free composition deconvolution, and, it has been evaluated with methylation data as Methods and LRCDE are both based on linear In csSAM, the of between sample estimated csGEPs are estimated by the and in LRCDE, is utilized for the Although is an old and for this LRCDE also a to estimate cell type proportions using a CARseq is a designed to csDEGs from It instead of utilized linear Methods Rodeo and have been originally developed for expression deconvolution, but we tested here their utility to identify csDEGs. two were into this to their [18], but any expression deconvolution method be utilized for detecting csDEGs as |$S$| for and samples using any expression deconvolution between and This a matrix with the as rows and cell types as as in and For the sample groups and do 1 and how the is than the observed in be the number of the the observed estimates as In this This for all genes in all cell types. into FDR if Rodeo is based on robust linear regression and its is some outlier samples in the is based on quadratic programming and it was originally as a composition deconvolution method [32], but an expression deconvolution of which is used in this and TCA [34] are developed to estimate cell type-specific methylation but their input requirements can be also with RNAseq CellDMC and it has been to identify also significant findings with opposite of in different cell types. TCA is based on of matrix and it can (e.g. and DESeq2 is a and used to RNAseq data any on deconvolution or related detecting csDEGs. is an important of its and, we have not utilized it on dataset GSE60424, which is based on DESeq2 linear and it can for user defined covariates. Here, we have used the cell type proportions and the interaction terms of them and sample as covariates. the csDEGs are then defined based on the significance of the interaction TOAST, CARseq, CellDMC and TCA have among the tested methods, if the data not contain many Methods designed for methylation data also on RNAseq The most important factors the accuracy of the estimated cell type-specific differentially expressed genes are of the cell type and individual heterogeneity between the samples. can and whether the data is heterogeneous for this type of This publicly available gene expression data with either data and to or data from or The are GSE60424 and GSE124742 for data, and for The to the data are available as in the it and the in the the and the We to for and on about from the and of and and the of the is also by of and is of Computational and in at of computational and is a in the by in the of at the of are analysis and Maria K. Jaakkola, Laura Elo |
Briefings Bioinform. | 2 |
| 2022 | Benchmarking methods for detecting differential states between conditions from multi-subject single-cell RNA-seq dataabstractSingle-cell RNA-sequencing (scRNA-seq) enables researchers to quantify transcriptomes of thousands of cells simultaneously and study transcriptomic changes between cells. scRNA-seq datasets increasingly include multisubject, multicondition experiments to investigate cell-type-specific differential states (DS) between conditions. This can be performed by first identifying the cell types in all the subjects and then by performing a DS analysis between the conditions within each cell type. Naïve single-cell DS analysis methods that treat cells statistically independent are subject to false positives in the presence of variation between biological replicates, an issue known as the pseudoreplicate bias. While several methods have already been introduced to carry out the statistical testing in multisubject scRNA-seq analysis, comparisons that include all these methods are currently lacking. Here, we performed a comprehensive comparison of 18 methods for the identification of DS changes between conditions from multisubject scRNA-seq data. Our results suggest that the pseudobulk methods performed generally best. Both pseudobulks and mixed models that model the subjects as a random effect were superior compared with the naïve single-cell methods that do not model the subjects in any way. While the naïve models achieved higher sensitivity than the pseudobulk methods and the mixed models, they were subject to a high number of false positives. In addition, accounting for subjects through latent variable modeling did not improve the performance of the naïve methods. Sini Junttila, Johannes Smolander, Laura Elo |
Briefings Bioinform. | 3 |
| 2022 | scShaper: an ensemble method for fast and accurate linear trajectory inference from single-cell RNA-seq dataabstractMOTIVATION: Computational models are needed to infer a representation of the cells, i.e. a trajectory, from single-cell RNA-sequencing data that model cell differentiation during a dynamic process. Although many trajectory inference methods exist, their performance varies greatly depending on the dataset and hence there is a need to establish more accurate, better generalizable methods. RESULTS: We introduce scShaper, a new trajectory inference method that enables accurate linear trajectory inference. The ensemble approach of scShaper generates a continuous smooth pseudotime based on a set of discrete pseudotimes. We demonstrate that scShaper is able to infer accurate trajectories for a variety of trigonometric trajectories, including many for which the commonly used principal curves method fails. A comprehensive benchmarking with state-of-the-art methods revealed that scShaper achieved superior accuracy of the cell ordering and, in particular, the differentially expressed genes. Moreover, scShaper is a fast method with few hyperparameters, making it a promising alternative to the principal curves method for linear pseudotemporal ordering. AVAILABILITY AND IMPLEMENTATION: scShaper is available as an R package at https://github.com/elolab/scshaper. The test data are available at https://doi.org/10.5281/zenodo.5734488. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Johannes Smolander, Sini Junttila, Mikko S. Venäläinen, Laura Elo |
Bioinform. | 4 |
| 2022 | Tool evaluation for the detection of variably sized indels from next generation whole genome and targeted sequencing dataabstractInsertions and deletions (indels) in human genomes are associated with a wide range of phenotypes, including various clinical disorders. High-throughput, next generation sequencing (NGS) technologies enable the detection of short genetic variants, such as single nucleotide variants (SNVs) and indels. However, the variant calling accuracy for indels remains considerably lower than for SNVs. Here we present a comparative study of the performance of variant calling tools for indel calling, evaluated with a wide repertoire of NGS datasets. While there is no single optimal tool to suit all circumstances, our results demonstrate that the choice of variant calling tool greatly impacts the precision and recall of indel calling. Furthermore, to reliably detect indels, it is essential to choose NGS technologies that offer a long read length and high coverage coupled with specific variant calling tools. Ning Wang 0032, Vladislav Lysenkov, Katri Orte, Veli Kairisto, Juhani Aakko, Sofia Khan, Laura Elo |
PLoS Comput. Biol. | 7 |
| 2021 | Stable Iterative Variable SelectionabstractMOTIVATION: The emergence of datasets with tens of thousands of features, such as high-throughput omics biomedical data, highlights the importance of reducing the feature space into a distilled subset that can truly capture the signal for research and industry by aiding in finding more effective biomarkers for the question in hand. A good feature set also facilitates building robust predictive models with improved interpretability and convergence of the applied method due to the smaller feature space. RESULTS: Here, we present a robust feature selection method named Stable Iterative Variable Selection (SIVS) and assess its performance over both omics and clinical data types. As a performance assessment metric, we compared the number and goodness of the selected feature using SIVS to those selected by Least Absolute Shrinkage and Selection Operator regression. The results suggested that the feature space selected by SIVS was, on average, 41% smaller, without having a negative effect on the model performance. A similar result was observed for comparison with Boruta and caret RFE. AVAILABILITY AND IMPLEMENTATION: The method is implemented as an R package under GNU General Public License v3.0 and is accessible via Comprehensive R Archive Network (CRAN) via https://cran.r-project.org/package=sivs. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Mehrad Mahmoudian, Mikko S. Venäläinen, Riku Klén, Laura Elo |
Bioinform. | 4 |
| 2021 | ILoReg: a tool for high-resolution cell population identification from single-cell RNA-seq dataabstractMOTIVATION: Single-cell RNA-seq allows researchers to identify cell populations based on unsupervised clustering of the transcriptome. However, subpopulations can have only subtle transcriptomic differences and the high dimensionality of the data makes their identification challenging. RESULTS: We introduce ILoReg, an R package implementing a new cell population identification method that improves identification of cell populations with subtle differences through a probabilistic feature extraction step that is applied before clustering and visualization. The feature extraction is performed using a novel machine learning algorithm, called iterative clustering projection (ICP), that uses logistic regression and clustering similarity comparison to iteratively cluster data. Remarkably, ICP also manages to integrate feature selection with the clustering through L1-regularization, enabling the identification of genes that are differentially expressed between cell populations. By combining solutions of multiple ICP runs into a single consensus solution, ILoReg creates a representation that enables investigating cell populations with a high resolution. In particular, we show that the visualization of ILoReg allows segregation of immune and pancreatic cell populations in a more pronounced manner compared with current state-of-the-art methods. AVAILABILITY AND IMPLEMENTATION: ILoReg is available as an R package at https://bioconductor.org/packages/ILoReg. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Johannes Smolander, Sini Junttila, Mikko S. Venäläinen, Laura Elo |
Bioinform. | 4 |
| 2020 | Systematic evaluation of differential splicing tools for RNA-seq studiesabstractDifferential splicing (DS) is a post-transcriptional biological process with critical, wide-ranging effects on a plethora of cellular activities and disease processes. To date, a number of computational approaches have been developed to identify and quantify differentially spliced genes from RNA-seq data, but a comprehensive intercomparison and appraisal of these approaches is currently lacking. In this study, we systematically evaluated 10 DS analysis tools for consistency and reproducibility, precision, recall and false discovery rate, agreement upon reported differentially spliced genes and functional enrichment. The tools were selected to represent the three different methodological categories: exon-based (DEXSeq, edgeR, JunctionSeq, limma), isoform-based (cuffdiff2, DiffSplice) and event-based methods (dSpliceType, MAJIQ, rMATS, SUPPA). Overall, all the exon-based methods and two event-based methods (MAJIQ and rMATS) scored well on the selected measures. Of the 10 tools tested, the exon-based methods performed generally better than the isoform-based and event-based methods. However, overall, the different data analysis tools performed strikingly differently across different data sets or numbers of samples. Arfa Mehmood, Asta Laiho, Mikko S. Venäläinen, Aidan J. McGlinchey, Ning Wang 0032, Laura Elo |
Briefings Bioinform. | 6 |
| 2018 | A systematic evaluation of normalization methods in quantitative label-free proteomicsabstractTo date, mass spectrometry (MS) data remain inherently biased as a result of reasons ranging from sample handling to differences caused by the instrumentation. Normalization is the process that aims to account for the bias and make samples more comparable. The selection of a proper normalization method is a pivotal task for the reliability of the downstream analysis and results. Many normalization methods commonly used in proteomics have been adapted from the DNA microarray techniques. Previous studies comparing normalization methods in proteomics have focused mainly on intragroup variation. In this study, several popular and widely used normalization methods representing different strategies in normalization are evaluated using three spike-in and one experimental mouse label-free proteomic data sets. The normalization methods are evaluated in terms of their ability to reduce variation between technical replicates, their effect on differential expression analysis and their effect on the estimation of logarithmic fold changes. Additionally, we examined whether normalizing the whole data globally or in segments for the differential expression analysis has an effect on the performance of the normalization methods. We found that variance stabilization normalization (Vsn) reduced variation the most between technical replicates in all examined data sets. Vsn also performed consistently well in the differential expression analysis. Linear regression normalization and local regression normalization performed also systematically well. Finally, we discuss the choice of a normalization method and some qualities of a suitable normalization method in the light of the results of our evaluation. Tommi Välikangas, Tomi Suomi, Laura Elo |
Briefings Bioinform. | 3 |
| 2018 | A comprehensive evaluation of popular proteomics software workflows for label-free proteome quantification and imputationabstractLabel-free mass spectrometry (MS) has developed into an important tool applied in various fields of biological and life sciences. Several software exist to process the raw MS data into quantified protein abundances, including open source and commercial solutions. Each software includes a set of unique algorithms for different tasks of the MS data processing workflow. While many of these algorithms have been compared separately, a thorough and systematic evaluation of their overall performance is missing. Moreover, systematic information is lacking about the amount of missing values produced by the different proteomics software and the capabilities of different data imputation methods to account for them.In this study, we evaluated the performance of five popular quantitative label-free proteomics software workflows using four different spike-in data sets. Our extensive testing included the number of proteins quantified and the number of missing values produced by each workflow, the accuracy of detecting differential expression and logarithmic fold change and the effect of different imputation and filtering methods on the differential expression results. We found that the Progenesis software performed consistently well in the differential expression analysis and produced few missing values. The missing values produced by the other software decreased their performance, but this difference could be mitigated using proper data filtering or imputation methods. Among the imputation methods, we found that the local least squares (lls) regression imputation consistently increased the performance of the software in the differential expression analysis, and a combination of both data filtering and local least squares imputation increased performance the most in the tested data sets. Tommi Välikangas, Tomi Suomi, Laura Elo |
Briefings Bioinform. | 3 |
| 2018 | Phosphonormalizer: an R package for normalization of MS-based label-free phosphoproteomicsabstractMotivation: Global centering-based normalization is a commonly used normalization approach in mass spectrometry-based label-free proteomics. It scales the peptide abundances to have the same median intensities, based on an assumption that the majority of abundances remain the same across the samples. However, especially in phosphoproteomics, this assumption can introduce bias, as the samples are enriched during sample preparation which can mask the underlying biological changes. To address this possible bias, phosphopeptides quantified in both enriched and non-enriched samples can be used to calculate factors that mitigate the bias. Results: We present an R package phosphonormalizer for normalizing enriched samples in label-free mass spectrometry-based phosphoproteomics. Availability and implementation: The phosphonormalizer package is freely available under GPL ( > =2) license from Bioconductor (https://bioconductor.org/packages/phosphonormalizer). Contact: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Sohrab Saraei, Tomi Suomi, Otto Kauko, Laura Elo |
Bioinform. | 4 |
| 2018 | SimPhospho: a software tool enabling confident phosphosite assignmentabstractMotivation: Mass spectrometry combined with enrichment strategies for phosphorylated peptides has been successfully employed for two decades to identify sites of phosphorylation. However, unambiguous phosphosite assignment is considered challenging. Given that site-specific phosphorylation events function as different molecular switches, validation of phosphorylation sites is of utmost importance. In our earlier study we developed a method based on simulated phosphopeptide spectral libraries, which enables highly sensitive and accurate phosphosite assignments. To promote more widespread use of this method, we here introduce a software implementation with improved usability and performance. Results: We present SimPhospho, a fast and user-friendly tool for accurate simulation of phosphopeptide tandem mass spectra. Simulated phosphopeptide spectral libraries are used to validate and supplement database search results, with a goal to improve reliable phosphoproteome identification and reporting. The presented program can be easily used together with the Trans-Proteomic Pipeline and integrated in a phosphoproteomics data analysis workflow. Availability and implementation: SimPhospho is open source and it is available for Windows, Linux and Mac operating systems. The software and its user's manual with detailed description of data analysis as well as test data can be found at https://sourceforge.net/projects/simphospho/. Supplementary information: Supplementary data are available at Bioinformatics online. Veronika Suni, Tomi Suomi, Tomoya Tsubosaka, Susumu Y. Imanishi, Laura Elo, Garry L. Corthals |
Bioinform. | 5 |
| 2017 | Comparison of methods to detect differentially expressed genes between single-cell populationsabstractWe compared five statistical methods to detect differentially expressed genes between two distinct single-cell populations. Currently, it remains unclear whether differential expression methods developed originally for conventional bulk RNA-seq data can also be applied to single-cell RNA-seq data analysis. Our results in three diverse comparison settings showed marked differences between the different methods in terms of the number of detections as well as their sensitivity and specificity. They, however, did not reveal systematic benefits of the currently available single-cell-specific methods. Instead, our previously introduced reproducibility-optimization method showed good performance in all comparison settings without any single-cell-specific modifications. Maria K. Jaakkola, Fatemeh Seyednasrollah, Arfa Mehmood, Laura Elo |
Briefings Bioinform. | 4 |
| 2017 | ROTS: An R package for reproducibility-optimized statistical testingabstractDifferential expression analysis is one of the most common types of analyses performed on various biological data (e.g. RNA-seq or mass spectrometry proteomics). It is the process that detects features, such as genes or proteins, showing statistically significant differences between the sample groups under comparison. A major challenge in the analysis is the choice of an appropriate test statistic, as different statistics have been shown to perform well in different datasets. To this end, the reproducibility-optimized test statistic (ROTS) adjusts a modified t-statistic according to the inherent properties of the data and provides a ranking of the features based on their statistical evidence for differential expression between two groups. ROTS has already been successfully applied in a range of different studies from transcriptomics to proteomics, showing competitive performance against other state-of-the-art methods. To promote its widespread use, we introduce here a Bioconductor R package for performing ROTS analysis conveniently on different types of omics data. To illustrate the benefits of ROTS in various applications, we present three case studies, involving proteomics and RNA-seq data from public repositories, including both bulk and single cell data. The package is freely available from Bioconductor (https://www.bioconductor.org/packages/ROTS). Tomi Suomi, Fatemeh Seyednasrollah, Maria K. Jaakkola, Thomas Faux, Laura Elo |
PLoS Comput. Biol. | 5 |
| 2016 | Empirical comparison of structure-based pathway methodsabstractMultiple methods have been proposed to estimate pathway activities from expression profiles, and yet, there is not enough information available about the performance of those methods. This makes selection of a suitable tool for pathway analysis difficult. Although methods based on simple gene lists have remained the most common approach, various methods that also consider pathway structure have emerged. To provide practical insight about the performance of both list-based and structure-based methods, we tested six different approaches to estimate pathway activities in two different case study settings of different characteristics. The first case study setting involved six renal cell cancer data sets, and the differences between expression profiles of case and control samples were relatively big. The second case study setting involved four type 1 diabetes data sets, and the profiles of case and control samples were more similar to each other. In general, there were marked differences in the outcomes of the different pathway tools even with the same input data. In the cancer studies, the results of a tested method were typically consistent across the different data sets, yet different between the methods. In the more challenging diabetes studies, almost all the tested methods detected as significant only few pathways if any. Maria K. Jaakkola, Laura Elo |
Briefings Bioinform. | 2 |
| 2015 | Comparison of software packages for detecting differential expression in RNA-seq studiesabstractRNA-sequencing (RNA-seq) has rapidly become a popular tool to characterize transcriptomes. A fundamental research problem in many RNA-seq studies is the identification of reliable molecular markers that show differential expression between distinct sample groups. Together with the growing popularity of RNA-seq, a number of data analysis methods and pipelines have already been developed for this task. Currently, however, there is no clear consensus about the best practices yet, which makes the choice of an appropriate method a daunting task especially for a basic user without a strong statistical or computational background. To assist the choice, we perform here a systematic comparison of eight widely used software packages and pipelines for detecting differential expression between sample groups in a practical research setting and provide general guidelines for choosing a robust pipeline. In general, our results demonstrate how the data analysis tool utilized can markedly affect the outcome of the data analysis, highlighting the importance of this choice. Fatemeh Seyednasrollah, Asta Laiho, Laura Elo |
Briefings Bioinform. | 3 |
| 2011 | Probabilistic Analysis of Probe Reliability in Differential Gene Expression Studies with Short Oligonucleotide ArraysabstractProbe defects are a major source of noise in gene expression studies. While existing approaches detect noisy probes based on external information such as genomic alignments, we introduce and validate a targeted probabilistic method for analyzing probe reliability directly from expression data and independently of the noise source. This provides insights into the various sources of probe-level noise and gives tools to guide probe design. Leo Lahti, Laura Elo, Tero Aittokallio, Samuel Kaski |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2009 | Optimized detection of differential expression in global profiling experiments: case studies in clinical transcriptomic and quantitative proteomic datasetsabstractIdentification of reliable molecular markers that show differential expression between distinct groups of samples has remained a fundamental research problem in many large-scale profiling studies, such as those based on DNA microarray or mass-spectrometry technologies. Despite the availability of a wide spectrum of statistical procedures, the users of the high-throughput platforms are still facing the crucial challenge of deciding which test statistic is best adapted to the intrinsic properties of their own datasets. To meet this challenge, we recently introduced an adaptive procedure, named ROTS (Reproducibility-Optimized Test Statistic), which learns an optimal statistic directly from the given data, and whose relative benefits have previously been shown in comparison with state-of-the-art procedures for detecting differential expression. Using gene expression microarray and mass-spectrometry (MS)-based protein expression datasets as case studies, we illustrate here the practical usage and advantages of ROTS toward detecting reliable marker lists in clinical transcriptomic and proteomic studies. In a public leukemia microarray dataset, the procedure could improve the sensitivity of the gene marker lists detected with high specificity. When applied to a recent LC-MS dataset, involving plasma samples from severe burn patients, the procedure could identify several peptide markers that remained undetected in the conventional analysis, thus demonstrating the effectiveness of ROTS also for global quantitative proteomic studies. To promote its widespread usage, we have made freely available efficient implementations of ROTS, which are easily accessible either as a stand-alone R-package or as integrated in the open-source data analysis software Chipster. Laura Elo, Jukka Hiissa, Jarno Tuimala, Aleksi Kallio, Eija Korpelainen, Tero Aittokallio |
Briefings Bioinform. | 1 |
| 2008 | Missing value imputation improves clustering and interpretation of gene expression microarray dataabstractBACKGROUND: Missing values frequently pose problems in gene expression microarray experiments as they can hinder downstream analysis of the datasets. While several missing value imputation approaches are available to the microarray users and new ones are constantly being developed, there is no general consensus on how to choose between the different methods since their performance seems to vary drastically depending on the dataset being used. RESULTS: We show that this discrepancy can mostly be attributed to the way in which imputation methods have traditionally been developed and evaluated. By comparing a number of advanced imputation methods on recent microarray datasets, we show that even when there are marked differences in the measurement-level imputation accuracies across the datasets, these differences become negligible when the methods are evaluated in terms of how well they can reproduce the original gene clusters or their biological interpretations. Regardless of the evaluation approach, however, imputation always gave better results than ignoring missing data points or replacing them with zeros or average values, emphasizing the continued importance of using more advanced imputation methods. CONCLUSION: The results demonstrate that, while missing values are still severely complicating microarray data analysis, their impact on the discovery of biologically meaningful gene groups can - up to a certain degree - be reduced by using readily available and relatively fast imputation methods, such as the Bayesian Principal Components Algorithm (BPCA). Johannes Tuikkala, Laura Elo, Olli Nevalainen, Tero Aittokallio |
BMC Bioinform. | 2 |
| 2008 | Reproducibility-Optimized Test Statistic for Ranking Genes in Microarray StudiesabstractA principal goal of microarray studies is to identify the genes showing differential expression under distinct conditions. In such studies, the selection of an optimal test statistic is a crucial challenge, which depends on the type and amount of data under analysis. While previous studies on simulated or spike-in datasets do not provide practical guidance on how to choose the best method for a given real dataset, we introduce an enhanced reproducibility-optimization procedure, which enables the selection of a suitable gene- anking statistic directly from the data. In comparison with existing ranking methods, the reproducibilityoptimized statistic shows good performance consistently under various simulated conditions and on Affymetrix spike-in dataset. Further, the feasibility of the novel statistic is confirmed in a practical research setting using data from an in-house cDNA microarray study of asthma-related gene expression changes. These results suggest that the procedure facilitates the selection of an appropriate test statistic for a given dataset without relying on a priori assumptions, which may bias the findings and their interpretation. Moreover, the general reproducibilityoptimization procedure is not limited to detecting differential expression only but could be extended to a wide range of other applications as well. Laura Elo, Sanna Filen, Riitta Lahesmaa, Tero Aittokallio |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2007 | Systematic construction of gene coexpression networks with applications to human T helper cell differentiation processabstractMOTIVATION: Coexpression networks have recently emerged as a novel holistic approach to microarray data analysis and interpretation. Choosing an appropriate cutoff threshold, above which a gene-gene interaction is considered as relevant, is a critical task in most network-centric applications, especially when two or more networks are being compared. RESULTS: We demonstrate that the performance of traditional approaches, which are based on a pre-defined cutoff or significance level, can vary drastically depending on the type of data and application. Therefore, we introduce a systematic procedure for estimating a cutoff threshold of coexpression networks directly from their topological properties. Both synthetic and real datasets show clear benefits of our data-driven approach under various practical circumstances. In particular, the procedure provides a robust estimate of individual degree distributions, even from multiple microarray studies performed with different array platforms or experimental designs, which can be used to discriminate the corresponding phenotypes. Application to human T helper cell differentiation process provides useful insights into the components and interactions controlling this process, many of which would have remained unidentified on the basis of expression change alone. Moreover, several human-mouse orthologs showed conserved topological changes in both systems, suggesting their potential importance in the differentiation process. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Laura Elo, Henna Järvenpää, Matej Oresic, Riitta Lahesmaa, Tero Aittokallio |
Bioinform. | 1 |
| 2006 | Improving missing value estimation in microarray data with gene ontologyabstractMOTIVATION: Gene expression microarray experiments produce datasets with frequent missing expression values. Accurate estimation of missing values is an important prerequisite for efficient data analysis as many statistical and machine learning techniques either require a complete dataset or their results are significantly dependent on the quality of such estimates. A limitation of the existing estimation methods for microarray data is that they use no external information but the estimation is based solely on the expression data. We hypothesized that utilizing a priori information on functional similarities available from public databases facilitates the missing value estimation. RESULTS: We investigated whether semantic similarity originating from gene ontology (GO) annotations could improve the selection of relevant genes for missing value estimation. The relative contribution of each information source was automatically estimated from the data using an adaptive weight selection procedure. Our experimental results in yeast cDNA microarray datasets indicated that by considering GO information in the k-nearest neighbor algorithm we can enhance its performance considerably, especially when the number of experimental conditions is small and the percentage of missing values is high. The increase of performance was less evident with a more sophisticated estimation method. We conclude that even a small proportion of annotated genes can provide improvements in data quality significant for the eventual interpretation of the microarray experiments. AVAILABILITY: Java and Matlab codes are available on request from the authors. SUPPLEMENTARY MATERIAL: Available online at http://users.utu.fi/jotatu/GOImpute.html. Johannes Tuikkala, Laura Elo, Olli Nevalainen, Tero Aittokallio |
Bioinform. | 2 |