VLDB 2026 Research / reviewers in the wild / expert
Charlotte Soneson
dblp:81/8247
· DBLP profile ↗
12ranked-venue papers
6as first author
5since 2021 · last 2025
0000-0003-3833-2169ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 11 · 6 first-author · 5 since 2021Artificial intelligence and machine learning · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Learning and teaching biological data science in the Bioconductor communityabstractModern biological research is increasingly data-intensive, leading to a growing demand for effective training in biological data science. In this article, we provide an overview of key resources and best practices available within the Bioconductor project-an open-source software community focused on omics data analysis. This guide serves as a valuable reference for both learners and educators in the field. Jenny Drnevich, Frederick Tan, Fabricio Almeida-Silva, Robert Castelo, Aedín C. Culhane, Sean R. Davis, Maria A. Doyle, Ludwig Geistlinger, Andrew R. Ghazi, Susan P. Holmes, Leo Lahti, Alexandru Mahmoud, Kozo Nishida, Marcel Ramos, Kevin C. Rue-Albrecht, David J. H. Shih, Laurent Gatto, Charlotte Soneson |
PLoS Comput. Biol. | 18 |
| 2025 | Building a continuous benchmarking ecosystem in bioinformaticsabstractBenchmarking, which involves collecting reference datasets and demonstrating method performance, is a requirement for the development of new computational tools, but also becomes a domain of its own to achieve neutral comparisons of methods. Although a lot has been written about how to design and conduct benchmark studies, this Perspective sheds light on a wish list for a computational platform to orchestrate benchmark studies. We discuss various ideas for organizing reproducible software environments, formally defining benchmarks, orchestrating standardized workflows, and how they interface with computing infrastructure. Izaskun Mallona, Charlotte Soneson, Ben Carrillo, Almut Lütge, Daniel Incicau, Reto Gerber, Anthony Sonrel, Mark D. Robinson |
PLoS Comput. Biol. | 2 |
| 2025 | Eleven quick tips for writing a Bioconductor packageabstractinternational Charlotte Soneson, Lori A. Shepherd, Marcel Ramos, Kevin C. Rue-Albrecht, Johannes Rainer, Hervé Pagès, Vincent Carey |
PLoS Comput. Biol. | 1 |
| 2022 | monaLisa: an R/Bioconductor package for identifying regulatory motifsabstractSUMMARY: Proteins binding to specific nucleotide sequences, such as transcription factors, play key roles in the regulation of gene expression. Their binding can be indirectly observed via associated changes in transcription, chromatin accessibility, DNA methylation and histone modifications. Identifying candidate factors that are responsible for these observed experimental changes is critical to understand the underlying biological processes. Here, we present monaLisa, an R/Bioconductor package that implements approaches to identify relevant transcription factors from experimental data. The package can be easily integrated with other Bioconductor packages and enables seamless motif analyses without any software dependencies outside of R. AVAILABILITY AND IMPLEMENTATION: monaLisa is implemented in R and available on Bioconductor at https://bioconductor.org/packages/monaLisa with the development version hosted on GitHub at https://github.com/fmicompbio/monaLisa. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Dania Machlab, Lukas Burger, Charlotte Soneson, Filippo M. Rijli, Dirk Schübeler, Michael B. Stadler |
Bioinform. | 3 |
| 2021 | Preprocessing choices affect RNA velocity results for droplet scRNA-seq dataabstractExperimental single-cell approaches are becoming widely used for many purposes, including investigation of the dynamic behaviour of developing biological systems. Consequently, a large number of computational methods for extracting dynamic information from such data have been developed. One example is RNA velocity analysis, in which spliced and unspliced RNA abundances are jointly modeled in order to infer a 'direction of change' and thereby a future state for each cell in the gene expression space. Naturally, the accuracy and interpretability of the inferred RNA velocities depend crucially on the correctness of the estimated abundances. Here, we systematically compare five widely used quantification tools, in total yielding thirteen different quantification approaches, in terms of their estimates of spliced and unspliced RNA abundances in five experimental droplet scRNA-seq data sets. We show that there are substantial differences between the quantifications obtained from different tools, and identify typical genes for which such discrepancies are observed. We further show that these abundance differences propagate to the downstream analysis, and can have a large effect on estimated velocities as well as the biological interpretation. Our results highlight that abundance quantification is a crucial aspect of the RNA velocity analysis workflow, and that both the definition of the genomic features of interest and the quantification algorithm itself require careful consideration. Charlotte Soneson, Avi Srivastava, Rob Patro, Michael B. Stadler |
PLoS Comput. Biol. | 1 |
| 2020 | Tximeta: Reference sequence checksums for provenance identification in RNA-seqabstractCorrect annotation metadata is critical for reproducible and accurate RNA-seq analysis. When files are shared publicly or among collaborators with incorrect or missing annotation metadata, it becomes difficult or impossible to reproduce bioinformatic analyses from raw data. It also makes it more difficult to locate the transcriptomic features, such as transcripts or genes, in their proper genomic context, which is necessary for overlapping expression data with other datasets. We provide a solution in the form of an R/Bioconductor package tximeta that performs numerous annotation and metadata gathering tasks automatically on behalf of users during the import of transcript quantification files. The correct reference transcriptome is identified via a hashed checksum stored in the quantification output, and key transcript databases are downloaded and cached locally. The computational paradigm of automatically adding annotation metadata based on reference sequence checksums can greatly facilitate genomic workflows, by helping to reduce overhead during bioinformatic analyses, preventing costly bioinformatic mistakes, and promoting computational reproducibility. The tximeta package is available at https://bioconductor.org/packages/tximeta. Michael I. Love, Charlotte Soneson, Peter F. Hickey, Lisa K. Johnson, N. Tessa Pierce-Ward, Lori A. Shepherd, Martin Morgan, Rob Patro |
PLoS Comput. Biol. | 2 |
| 2018 | Towards unified quality verification of synthetic count data with countsimQCabstractSummary: Statistical tools for biological data analysis are often evaluated using synthetic data, designed to mimic the features of a specific type of experimental data. The generalizability of such evaluations depends on how well the synthetic data reproduce the main characteristics of the experimental data, and we argue that an assessment of this similarity should accompany any synthetic dataset used for method evaluation. We describe countsimQC, which provides a straightforward way to generate a stand-alone report that shows the main characteristics of (e.g. RNA-seq) count data and can be provided alongside a publication as verification of the appropriateness of any utilized synthetic data. Availability and implementation: countsimQC is implemented as an R package (for R versions ≥ 3.4) and is available from https://github.com/csoneson/countsimQC under a GPL (≥2) license. Contact: [email protected] or [email protected]. Charlotte Soneson, Mark D. Robinson |
Bioinform. | 1 |
| 2015 | Low Bias Local Intrinsic Dimension Estimation from Expected Simplex SkewnessabstractIn exploratory high-dimensional data analysis, local intrinsic dimension estimation can sometimes be used in order to discriminate between data sets sampled from different low-dimensional structures. Global intrinsic dimension estimators can in many cases be adapted to local estimation, but this leads to problems with high negative bias or high variance. We introduce a method that exploits the curse/blessing of dimensionality and produces local intrinsic dimension estimators that have very low bias, even in cases where the intrinsic dimension is higher than the number of data points, in combination with relatively low variance. We show that our estimators have a very good ability to classify local data sets by their dimension compared to other local intrinsic dimension estimators; furthermore we provide examples showing the usefulness of local intrinsic dimension estimation in general and our method in particular for stratification of real data sets. Kerstin Johnsson, Charlotte Soneson, Magnus Fontes |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2014 | compcodeR - an R package for benchmarking differential expression methods for RNA-seq dataabstractUNLABELLED: compcodeR is an R package for benchmarking of differential expression analysis methods, in particular, methods developed for analyzing RNA-seq data. The package provides functionality for simulating realistic RNA-seq count datasets, an interface to several of the most commonly used differential expression analysis methods and extensive functionality for evaluating and comparing different approaches on real and simulated data. AVAILABILITY AND IMPLEMENTATION: compcodeR is available from http://www.bioconductor.org/packages/release/bioc/html/compcodeR.html. Charlotte Soneson |
Bioinform. | 1 |
| 2013 | A comparison of methods for differential expression analysis of RNA-seq dataabstractBACKGROUND: Finding genes that are differentially expressed between conditions is an integral part of understanding the molecular basis of phenotypic variation. In the past decades, DNA microarrays have been used extensively to quantify the abundance of mRNA corresponding to different genes, and more recently high-throughput sequencing of cDNA (RNA-seq) has emerged as a powerful competitor. As the cost of sequencing decreases, it is conceivable that the use of RNA-seq for differential expression analysis will increase rapidly. To exploit the possibilities and address the challenges posed by this relatively new type of data, a number of software packages have been developed especially for differential expression analysis of RNA-seq data. RESULTS: We conducted an extensive comparison of eleven methods for differential expression analysis of RNA-seq data. All methods are freely available within the R framework and take as input a matrix of counts, i.e. the number of reads mapping to each genomic feature of interest in each of a number of samples. We evaluate the methods based on both simulated data and real RNA-seq data. CONCLUSIONS: Very small sample sizes, which are still common in RNA-seq experiments, impose problems for all evaluated methods and any results obtained under such conditions should be interpreted with caution. For larger sample sizes, the methods combining a variance-stabilizing transformation with the 'limma' method for differential expression analysis perform well under many different conditions, as does the nonparametric SAMseq method. Charlotte Soneson, Mauro Delorenzi |
BMC Bioinform. | 1 |
| 2011 | The projection score - an evaluation criterion for variable subset selection in PCA visualizationabstractBACKGROUND: In many scientific domains, it is becoming increasingly common to collect high-dimensional data sets, often with an exploratory aim, to generate new and relevant hypotheses. The exploratory perspective often makes statistically guided visualization methods, such as Principal Component Analysis (PCA), the methods of choice. However, the clarity of the obtained visualizations, and thereby the potential to use them to formulate relevant hypotheses, may be confounded by the presence of the many non-informative variables. For microarray data, more easily interpretable visualizations are often obtained by filtering the variable set, for example by removing the variables with the smallest variances or by only including the variables most highly related to a specific response. The resulting visualization may depend heavily on the inclusion criterion, that is, effectively the number of retained variables. To our knowledge, there exists no objective method for determining the optimal inclusion criterion in the context of visualization. RESULTS: We present the projection score, which is a straightforward, intuitively appealing measure of the informativeness of a variable subset with respect to PCA visualization. This measure can be universally applied to find suitable inclusion criteria for any type of variable filtering. We apply the presented measure to find optimal variable subsets for different filtering methods in both microarray data sets and synthetic data sets. We note also that the projection score can be applied in general contexts, to compare the informativeness of any variable subsets with respect to visualization by PCA. CONCLUSIONS: We conclude that the projection score provides an easily interpretable and universally applicable measure of the informativeness of a variable subset with respect to visualization by PCA, that can be used to systematically find the most interpretable PCA visualization in practical exploratory analysis. Magnus Fontes, Charlotte Soneson |
BMC Bioinform. | 2 |
| 2010 | Integrative analysis of gene expression and copy number alterations using canonical correlation analysisabstractBACKGROUND: With the rapid development of new genetic measurement methods, several types of genetic alterations can be quantified in a high-throughput manner. While the initial focus has been on investigating each data set separately, there is an increasing interest in studying the correlation structure between two or more data sets. Multivariate methods based on Canonical Correlation Analysis (CCA) have been proposed for integrating paired genetic data sets. The high dimensionality of microarray data imposes computational difficulties, which have been addressed for instance by studying the covariance structure of the data, or by reducing the number of variables prior to applying the CCA. In this work, we propose a new method for analyzing high-dimensional paired genetic data sets, which mainly emphasizes the correlation structure and still permits efficient application to very large data sets. The method is implemented by translating a regularized CCA to its dual form, where the computational complexity depends mainly on the number of samples instead of the number of variables. The optimal regularization parameters are chosen by cross-validation. We apply the regularized dual CCA, as well as a classical CCA preceded by a dimension-reducing Principal Components Analysis (PCA), to a paired data set of gene expression changes and copy number alterations in leukemia. RESULTS: Using the correlation-maximizing methods, regularized dual CCA and PCA+CCA, we show that without pre-selection of known disease-relevant genes, and without using information about clinical class membership, an exploratory analysis singles out two patient groups, corresponding to well-known leukemia subtypes. Furthermore, the variables showing the highest relevance to the extracted features agree with previous biological knowledge concerning copy number alterations and gene expression changes in these subtypes. Finally, the correlation-maximizing methods are shown to yield results which are more biologically interpretable than those resulting from a covariance-maximizing method, and provide different insight compared to when each variable set is studied separately using PCA. CONCLUSIONS: We conclude that regularized dual CCA as well as PCA+CCA are useful methods for exploratory analysis of paired genetic data sets, and can be efficiently implemented also when the number of variables is very large. Charlotte Soneson, Henrik Lilljebjörn, Thoas Fioretos, Magnus Fontes |
BMC Bioinform. | 1 |