EDBT 2026 Demo / reviewers in the wild / expert
Olga Vitek
dblp:09/4940
· DBLP profile ↗
32ranked-venue papers
2as first author
8since 2021 · last 2025
0000-0003-1728-1104ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 28 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 1Software engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Relative quantification of proteins and post-translational modifications in proteomic experiments with shared peptides: a weight-based approachabstractMOTIVATION: Bottom-up mass spectrometry-based proteomics studies changes in protein abundance and structure across conditions. Since the currency of these experiments are peptides, i.e. subsets of protein sequences that carry the quantitative information, conclusions at a different level must be computationally inferred. The inference is particularly challenging in situations where the peptides are shared by multiple proteins or post-translational modifications. While many approaches infer the underlying abundances from unique peptides, there is a need to distinguish the quantitative patterns when peptides are shared. RESULTS: We propose a statistical approach for estimating protein abundances, as well as site occupancies of post-translational modifications, based on quantitative information from shared peptides. The approach treats the quantitative patterns of shared peptides as convex combinations of abundances of individual proteins or modification sites, and estimates the abundance of each source in a sample together with the weights of the combination. In simulation-based evaluations, the proposed approach improved the precision of estimated fold changes between conditions. We further demonstrated the practical utility of the approach in experiments with diverse biological objectives, ranging from protein degradation and thermal proteome stability, to changes in protein post-translational modifications. AVAILABILITY AND IMPLEMENTATION: The approach is implemented in an open-source R package MSstatsWeightedSummary. The package is currently available at https://github.com/Vitek-Lab/MSstatsWeightedSummary (doi: 10.5281/zenodo.14662989). Code required to reproduce the results presented in this article can be found in a repository https://github.com/mstaniak/MWS_reproduction (doi: 10.5281/zenodo.14656053). Mateusz Staniak, Amanda M. Figueroa-Navedo, Devon Kohler, Meena Choi, Trent Hinkle, Tracy Kleinheinz, Robert Blake, Christopher M. Rose, Yingrong Xu, Pierre M. Jean Beltran, Malgorzata Bogdan, Olga Vitek |
Bioinform. | 14 |
| 2024 | <tt>MSIreg</tt>: an R package for unsupervised coregistration of mass spectrometry and H&E imagesabstractSUMMARY: Joint analysis of mass spectrometry images (MS images) and microscopy images of hematoxylin and eosin (H&E) stained tissues assists pathologists in characterizing the morphological structure of the tissues, and in performing diagnosis. Unfortunately, the analysis is undermined by substantial differences between these modalities in terms of aspect ratios, spatial resolution, number of channels in each image, as well as by large global or small local elastic spatial deformations of one image with respect to the other. Therefore, accurate coregistration of the images is a critical pre-requisite for their joint interpretation. We introduce MSIreg, an open-source R package for coregistration of MSI and H&E images. MSIreg is designed for high-dimensional MSI experiments where each spatial location is represented by thousands of mass features. Unlike most existing coregistration methods, MSIreg implements a landmark free workflow, and quantitative metrics for performance evaluation. We evaluate the performance of MSIreg on six case studies, including coregistration of contiguous tissues with large deformations, as well as simultaneous coregistration of 29 tissue microarray cores. AVAILABILITY AND IMPLEMENTATION: The R package, installation instructions, and fully reproducible vignettes describing methods and Case Studies are available open-source under the GPL-3.0 license at https://github.com/sslakkimsetty/msireg/. Sai Srikanth Lakkimsetty, Kylie A. Bemis, Verena Stehl, Peter Bronsert, Melanie Christine Föll, Olga Vitek |
Bioinform. | 7 |
| 2024 | <tt>Eliater</tt>: a Python package for estimating outcomes of perturbations in biomolecular networksabstractSUMMARY: We introduce Eliater, a Python package for estimating the effect of perturbation of an upstream molecule on a downstream molecule in a biomolecular network. The estimation takes as input a biomolecular network, observational biomolecular data, and a perturbation of interest, and outputs an estimated quantitative effect of the perturbation. We showcase the functionalities of Eliater in a case study of Escherichia coli transcriptional regulatory network. AVAILABILITY AND IMPLEMENTATION: The code, the documentation, and several case studies are available open source at https://github.com/y0-causal-inference/eliater. Sara Mohammad Taheri, Pruthvi Prakash Navada, Charles Tapley Hoyt, Jeremy Zucker, Karen Sachs, Benjamin M. Gyori, Olga Vitek |
Bioinform. | 7 |
| 2023 | A noise-robust deep clustering of biomolecular ions improves interpretability of mass spectrometric imagesabstractMOTIVATION: Mass Spectrometry Imaging (MSI) analyzes complex biological samples such as tissues. It simultaneously characterizes the ions present in the tissue in the form of mass spectra, and the spatial distribution of the ions across the tissue in the form of ion images. Unsupervised clustering of ion images facilitates the interpretation in the spectral domain, by identifying groups of ions with similar spatial distributions. Unfortunately, many current methods for clustering ion images ignore the spatial features of the images, and are therefore unable to learn these features for clustering purposes. Alternative methods extract spatial features using deep neural networks pre-trained on natural image tasks; however, this is often inadequate since ion images are substantially noisier than natural images. RESULTS: We contribute a deep clustering approach for ion images that accounts for both spatial contextual features and noise. In evaluations on a simulated dataset and on four experimental datasets of different tissue types, the proposed method grouped ions from the same source into a same cluster more frequently than existing methods. We further demonstrated that using ion image clustering as a pre-processing step facilitated the interpretation of a subsequent spatial segmentation as compared to using either all the ions or one ion at a time. As a result, the proposed approach facilitated the interpretability of MSI data in both the spectral domain and the spatial domain. AVAILABILITYAND IMPLEMENTATION: The data and code are available at https://github.com/DanGuo1223/mzClustering. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Melanie Christine Föll, Kylie A. Bemis, Olga Vitek |
Bioinform. | 4 |
| 2023 | Optimal adjustment sets for causal query estimation in partially observed biomolecular networksabstractCausal query estimation in biomolecular networks commonly selects a 'valid adjustment set', i.e. a subset of network variables that eliminates the bias of the estimator. A same query may have multiple valid adjustment sets, each with a different variance. When networks are partially observed, current methods use graph-based criteria to find an adjustment set that minimizes asymptotic variance. Unfortunately, many models that share the same graph topology, and therefore same functional dependencies, may differ in the processes that generate the observational data. In these cases, the topology-based criteria fail to distinguish the variances of the adjustment sets. This deficiency can lead to sub-optimal adjustment sets, and to miss-characterization of the effect of the intervention. We propose an approach for deriving 'optimal adjustment sets' that takes into account the nature of the data, bias and finite-sample variance of the estimator, and cost. It empirically learns the data generating processes from historical experimental data, and characterizes the properties of the estimators by simulation. We demonstrate the utility of the proposed approach in four biomolecular Case studies with different topologies and different data generation processes. The implementation and reproducible Case studies are at https://github.com/srtaheri/OptimalAdjustmentSet. Sara Mohammad Taheri, Vartika Tewari, Rohan Kapre, Ehsan Rahiminasab, Karen Sachs, Charles Tapley Hoyt, Jeremy Zucker, Olga Vitek |
Bioinform. | 8 |
| 2022 | Do-calculus enables estimation of causal effects in partially observed biomolecular pathwaysabstractMOTIVATION: Estimating causal queries, such as changes in protein abundance in response to a perturbation, is a fundamental task in the analysis of biomolecular pathways. The estimation requires experimental measurements on the pathway components. However, in practice many pathway components are left unobserved (latent) because they are either unknown, or difficult to measure. Latent variable models (LVMs) are well-suited for such estimation. Unfortunately, LVM-based estimation of causal queries can be inaccurate when parameters of the latent variables are not uniquely identified, or when the number of latent variables is misspecified. This has limited the use of LVMs for causal inference in biomolecular pathways. RESULTS: In this article, we propose a general and practical approach for LVM-based estimation of causal queries. We prove that, despite the challenges above, LVM-based estimators of causal queries are accurate if the queries are identifiable according to Pearl's do-calculus and describe an algorithm for its estimation. We illustrate the breadth and the practical utility of this approach for estimating causal queries in four synthetic and two experimental case studies, where structures of biomolecular pathways challenge the existing methods for causal query estimation. AVAILABILITY AND IMPLEMENTATION: The code and the data documenting all the case studies are available at https://github.com/srtaheri/LVMwithDoCalculus. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Sara Mohammad Taheri, Jeremy Zucker, Charles Tapley Hoyt, Karen Sachs, Vartika Tewari, Robert Osazuwa Ness, Olga Vitek |
Bioinform. | 7 |
| 2022 | Effective Use of Likert Scales in Visualization Evaluations: A Systematic ReviewabstractAbstract Likert scales are often used in visualization evaluations to produce quantitative estimates of subjective attributes, such as ease of use or aesthetic appeal. However, the methods used to collect, analyze, and visualize data collected with Likert scales are inconsistent among evaluations in visualization papers. In this paper, we examine the use of Likert scales as a tool for measuring subjective response in a systematic review of 134 visualization evaluations published between 2009 and 2019. We find that papers with both objective and subjective measures do not hold the same reporting and analysis standards for both aspects of their evaluation, producing less rigorous work for the subjective qualities measured by Likert scales. Additionally, we demonstrate that many papers are inconsistent in their interpretations of Likert data as discrete or continuous and may even sacrifice statistical power by applying nonparametric tests unnecessarily. Finally, we identify instances where key details about Likert item construction with the potential to bias participant responses are omitted from evaluation methodology reporting, inhibiting the feasibility and reliability of future replication studies. We summarize recommendations from other fields for best practices with Likert data in visualization evaluations, based on the results of our survey. A full copy of this paper and all supplementary material are available at https://osf.io/exbz8/ . Laura South, David Saffo, Olga Vitek, Cody Dunne, Michelle Borkin |
Comput. Graph. Forum | 3 |
| 2021 | Leveraging Structured Biological Knowledge for Counterfactual Inference: A Case Study of Viral PathogenesisabstractCounterfactual inference is a useful tool for comparing outcomes of interventions on complex systems. It requires us to represent the system in form of a structural causal model, complete with a causal diagram, probabilistic assumptions on exogenous variables, and functional assignments. Specifying such models can be extremely difficult in practice. The process requires substantial domain expertise, and does not scale easily to large systems, multiple systems, or novel system modifications. At the same time, many application domains, such as molecular biology, are rich in structured causal knowledge that is qualitative in nature. This article proposes a general approach for querying a causal biological knowledge graph, and converting the qualitative result into a quantitative structural causal model that can learn from data to answer the question. We demonstrate the feasibility, accuracy and versatility of this approach using two case studies in systems biology. The first demonstrates the appropriateness of the underlying assumptions and the accuracy of the results. The second demonstrates the versatility of the approach by querying a knowledge base for the molecular determinants of a severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2)-induced cytokine storm, and performing counterfactual inference to estimate the causal effect of medical countermeasures for severely ill patients. Jeremy Zucker, Kaushal Paneri, Sara Mohammad Taheri, Somya Bhargava, Pallavi Kolambkar, Craig Bakker, Jeremy Teuton, Charles Tapley Hoyt, Kristie L. Oxford, Robert Osazuwa Ness, Olga Vitek |
IEEE Trans. Big Data | 11 |
| 2020 | Deep multiple instance learning classifies subtissue locations in mass spectrometry images from tissue-level annotationsabstractMOTIVATION: Mass spectrometry imaging (MSI) characterizes the molecular composition of tissues at spatial resolution, and has a strong potential for distinguishing tissue types, or disease states. This can be achieved by supervised classification, which takes as input MSI spectra, and assigns class labels to subtissue locations. Unfortunately, developing such classifiers is hindered by the limited availability of training sets with subtissue labels as the ground truth. Subtissue labeling is prohibitively expensive, and only rough annotations of the entire tissues are typically available. Classifiers trained on data with approximate labels have sub-optimal performance. RESULTS: To alleviate this challenge, we contribute a semi-supervised approach mi-CNN. mi-CNN implements multiple instance learning with a convolutional neural network (CNN). The multiple instance aspect enables weak supervision from tissue-level annotations when classifying subtissue locations. The convolutional architecture of the CNN captures contextual dependencies between the spectral features. Evaluations on simulated and experimental datasets demonstrated that mi-CNN improved the subtissue classification as compared to traditional classifiers. We propose mi-CNN as an important step toward accurate subtissue classification in MSI, enabling rapid distinction between tissue types and disease states. AVAILABILITY AND IMPLEMENTATION: The data and code are available at https://github.com/Vitek-Lab/mi-CNN_MSI. Melanie Christine Föll, Veronika Volkmann, Kathrin Enderle-Ammour, Peter Bronsert, Oliver Schilling, Olga Vitek |
Bioinform. | 7 |
| 2020 | New mixture models for decoy-free false discovery rate estimation in mass spectrometry proteomicsabstractMOTIVATION: Accurate estimation of false discovery rate (FDR) of spectral identification is a central problem in mass spectrometry-based proteomics. Over the past two decades, target-decoy approaches (TDAs) and decoy-free approaches (DFAs) have been widely used to estimate FDR. TDAs use a database of decoy species to faithfully model score distributions of incorrect peptide-spectrum matches (PSMs). DFAs, on the other hand, fit two-component mixture models to learn the parameters of correct and incorrect PSM score distributions. While conceptually straightforward, both approaches lead to problems in practice, particularly in experiments that push instrumentation to the limit and generate low fragmentation-efficiency and low signal-to-noise-ratio spectra. RESULTS: We introduce a new decoy-free framework for FDR estimation that generalizes present DFAs while exploiting more search data in a manner similar to TDAs. Our approach relies on multi-component mixtures, in which score distributions corresponding to the correct PSMs, best incorrect PSMs and second-best incorrect PSMs are modeled by the skew normal family. We derive EM algorithms to estimate parameters of these distributions from the scores of best and second-best PSMs associated with each experimental spectrum. We evaluate our models on multiple proteomics datasets and a HeLa cell digest case study consisting of more than a million spectra in total. We provide evidence of improved performance over existing DFAs and improved stability and speed over TDAs without any performance degradation. We propose that the new strategy has the potential to extend beyond peptide identification and reduce the need for TDA on all analytical platforms. AVAILABILITYAND IMPLEMENTATION: https://github.com/shawn-peng/FDR-estimation. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yisu Peng, Shantanu Jain, Yong Fuga Li, Michal Gregus, Alexander R. Ivanov, Olga Vitek, Predrag Radivojac |
Bioinform. | 6 |
| 2019 | Evaluating Pan and Zoom Timelines and SlidersabstractPan and zoom timelines and sliders help us navigate large time series data. However, designing efficient interactions can be difficult. We study pan and zoom methods via crowd-sourced experiments on mobile and computer devices, asking which designs and interactions provide faster target acquisition. We find that visual context should be limited for low-distance navigation, but added for far-distance navigation; that timelines should be oriented along the longer axis, especially on mobile; and that, as compared to default techniques, double click, hold, and rub zoom appear to scale worse with task difficulty, whereas brush and especially ortho zoom seem to scale better. Software and data used in this research are available as open source. Michail Schwab, Olga Vitek, James Tompkin 0001, Jeff Huang 0002, Michelle Borkin |
CHI | 3 |
| 2019 | Integrating Markov processes with structural causal modeling enables counterfactual inference in complex systemsabstractThis manuscript contributes a general and practical framework for casting a Markov process model of a system at equilibrium as a structural causal model, and carrying out counterfactual inference. Markov processes mathematically describe the mechanisms in the system, and predict the system’s equilibrium behavior upon intervention, but do not support counterfactual inference. In contrast, structural causal models support counterfactual inference, but do not identify the mechanisms. This manuscript leverages the benefits of both approaches. We define the structural causal models in terms of the parameters and the equilibrium dynamics of the Markov process models, and counterfactual inference flows from these settings. The proposed approach alleviates the identifiability drawback of the structural causal models, in that the counterfactual inference is consistent with the counterfactual trajectories simulated from the Markov process model. We showcase the benefits of this framework in case studies of complex biomolecular systems with nonlinear dynamics. We illustrate that, in presence of Markov process model misspecification, counterfactual inference leverages prior data, and therefore estimates the outcome of an intervention more accurately than a direct simulation. Robert Osazuwa Ness, Kaushal Paneri, Olga Vitek |
NeurIPS | 3 |
| 2019 | Unsupervised segmentation of mass spectrometric ion images characterizes morphology of tissuesabstractMOTIVATION: Mass spectrometry imaging (MSI) characterizes the spatial distribution of ions in complex biological samples such as tissues. Since many tissues have complex morphology, treatments and conditions often affect the spatial distribution of the ions in morphology-specific ways. Evaluating the selectivity and the specificity of ion localization and regulation across morphology types is biologically important. However, MSI lacks algorithms for segmenting images at both single-ion and spatial resolution. RESULTS: This article contributes spatial-Dirichlet Gaussian mixture model (DGMM), an algorithm and a workflow for the analyses of MSI experiments, that detects components of single-ion images with homogeneous spatial composition. The approach extends DGMMs to account for the spatial structure of MSI. Evaluations on simulated and experimental datasets with diverse MSI workflows demonstrated that spatial-DGMM accurately segments ion images, and can distinguish ions with homogeneous and heterogeneous spatial distribution. We also demonstrated that the extracted spatial information is useful for downstream analyses, such as detecting morphology-specific ions, finding groups of ions with similar spatial patterns, and detecting changes in chemical composition of tissues between conditions. AVAILABILITY AND IMPLEMENTATION: The data and code are available at https://github.com/Vitek-Lab/IonSpattern. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Kylie A. Bemis, Catherine Rawlins, Jeffrey N. Agar, Olga Vitek |
Bioinform. | 5 |
| 2019 | On the Impact of Programming Languages on Code Quality: A Reproduction StudyabstractIn a 2014 article, Ray, Posnett, Devanbu, and Filkov claimed to have uncovered a statistically significant association between 11 programming languages and software defects in 729 projects hosted on GitHub. Specifically, their work answered four research questions relating to software defects and programming languages. With data and code provided by the authors, the present article first attempts to conduct an experimental repetition of the original study. The repetition is only partially successful, due to missing code and issues with the classification of languages. The second part of this work focuses on their main claim, the association between bugs and languages, and performs a complete, independent reanalysis of the data and of the statistical modeling steps undertaken by Ray et al. in 2014. This reanalysis uncovers a number of serious flaws that reduce the number of languages with an association with defects down from 11 to only 4. Moreover, the practical effect size is exceedingly small. These results thus undermine the conclusions of the original study. Correcting the record is important, as many subsequent works have cited the 2014 article and have asserted, without evidence, a causal link between the choice of programming language for a given task and the number of software defects. Causation is not supported by the data at hand; and, in our opinion, even after fixing the methodological flaws we uncovered, too many unaccounted sources of bias remain to hope for a meaningful comparison of bug rates across languages. Emery D. Berger, Celeste Hollenbeck, Petr Maj, Olga Vitek, Jan Vitek |
ACM Trans. Program. Lang. Syst. | 4 |
| 2018 | Statistical Inference of Peroxisome Dynamics
Cyril Galitzine, Pierre M. Jean Beltran, Ileana M. Cristea, Olga Vitek |
RECOMB | 4 |
| 2017 | A Bayesian Active Learning Experimental Design for Inferring Signaling Networks
Robert Osazuwa Ness, Karen Sachs, Parag Mallick, Olga Vitek |
RECOMB | 4 |
| 2017 | matter: an R package for rapid prototyping with larger-than-memory datasets on diskabstractSummary: We introduce matter , an R package for direct interactions with larger-than-memory datasets, stored in an arbitrary number of files of any size. matter is primarily designed for datasets in new and rapidly evolving file formats, which may lack extensive software support. matter enables a wide variety of data exploration and manipulation steps, and is extensible to many bioinformatics applications. It supports reproducible research by minimizing the need of converting and storing data in multiple formats. We illustrate the performance of matter in conjunction with the Bioconductor package Cardinal for analysis of high-resolution, high-throughput mass spectrometry imaging experiments. Availability: The package, vignettes, and examples of applications in several areas of bioinformatics are available open-source at www.bioconductor.org under the Artistic-2.0 license. Contact: [email protected]. Kylie A. Bemis, Olga Vitek |
Bioinform. | 2 |
| 2015 | Cardinal: an R package for statistical analysis of mass spectrometry-based imaging experimentsabstractAbstract Cardinal is an R package for statistical analysis of mass spectrometry-based imaging (MSI) experiments of biological samples such as tissues. Cardinal supports both Matrix-Assisted Laser Desorption/Ionization (MALDI) and Desorption Electrospray Ionization-based MSI workflows, and experiments with multiple tissues and complex designs. The main analytical functionalities include (1) image segmentation, which partitions a tissue into regions of homogeneous chemical composition, selects the number of segments and the subset of informative ions, and characterizes the associated uncertainty and (2) image classification, which assigns locations on the tissue to pre-defined classes, selects the subset of informative ions, and estimates the resulting classification error by (cross-) validation. The statistical methods are based on mixture modeling and regularization. Contact: [email protected] Availability and implementation: The code, the documentation, and examples are available open-source at www.cardinalmsi.org under the Artistic-2.0 license. The package is available at www.bioconductor.org. Kyle D. Bemis, April Harry, Livia Eberlin, Christina Ferreira, Stephanie M. van de Ven, Parag Mallick, Mark L. Stolowitz, Olga Vitek |
Bioinform. | 8 |
| 2015 | Using collective expert judgements to evaluate quality measures of mass spectrometry imagesabstractMOTIVATION: Imaging mass spectrometry (IMS) is a maturating technique of molecular imaging. Confidence in the reproducible quality of IMS data is essential for its integration into routine use. However, the predominant method for assessing quality is visual examination, a time consuming, unstandardized and non-scalable approach. So far, the problem of assessing the quality has only been marginally addressed and existing measures do not account for the spatial information of IMS data. Importantly, no approach exists for unbiased evaluation of potential quality measures. RESULTS: We propose a novel approach for evaluating potential measures by creating a gold-standard set using collective expert judgements upon which we evaluated image-based measures. To produce a gold standard, we engaged 80 IMS experts, each to rate the relative quality between 52 pairs of ion images from MALDI-TOF IMS datasets of rat brain coronal sections. Experts' optional feedback on their expertise, the task and the survey showed that (i) they had diverse backgrounds and sufficient expertise, (ii) the task was properly understood, and (iii) the survey was comprehensible. A moderate inter-rater agreement was achieved with Krippendorff's alpha of 0.5. A gold-standard set of 634 pairs of images with accompanying ratings was constructed and showed a high agreement of 0.85. Eight families of potential measures with a range of parameters and statistical descriptors, giving 143 in total, were evaluated. Both signal-to-noise and spatial chaos-based measures performed highly with a correlation of 0.7 to 0.9 with the gold standard ratings. Moreover, we showed that a composite measure with the linear coefficients (trained on the gold standard with regularized least squares optimization and lasso) showed a strong linear correlation of 0.94 and an accuracy of 0.98 in predicting which image in a pair was of higher quality. AVAILABILITY AND IMPLEMENTATION: The anonymized data collected from the survey and the Matlab source code for data processing can be found at: https://github.com/alexandrovteam/IMS_quality. Andrew Palmer, Ekaterina Ovchinnikova, Mikael Thuné, Régis Lavigne, Blandine Guével, Andrey Dyatlov, Olga Vitek, Charles Pineau, Mats Borén, Theodore Alexandrov |
Bioinform. | 7 |
| 2015 | Statistical elimination of spectral features with large between-run variation enhances quantitative protein-level conclusions in experiments with data-independent spectral acquisitionabstractMany proteomic investigations summarize the quantitative information across multiple spectral features into protein-level conclusions. Data-independent spectral acquisition (DIA) now generates a lot of interest, as it allows us to quantify many spectral features in a single run. However, the disadvantage of DIA experiments as compared, e.g., to Selected Reaction Monitoring (SRM) is that the features are subject to interferences and noise. We argue that between-run variation provides an additional insight for distinguishing good-quality and noisy DIA features. To appropriately use the quantitative between-run variation, it is important to account for the properties experimental design, and distinguish random artifacts from the biological changes. We have previously proposed a method (Chang et al., ASMS 2013) that accounts for the experimental design to eliminate features with low information content. In this project we furthermore emphasized that conducting regularization helps us avoid exploring every subset of features exhaustively, and allows us to conduct hypothesis tests later on so that we would be able to control the false discovery rate of the feature selection process. We evaluated our proposed approach by using three datasets that have some notion of ground truth: an extensive simulation study, a controlled mixture where proteins were spiked into a complex background in known concentrations, and a study of 232 plasma samples, where 18 proteins were quantified in both SWAH and SRM mode in presence of heavy labeled reference peptides. We worked on [ 1 ] protein-level estimates of fold changes between conditions, [ 2 ] sensitivity and specificity of detecting changes in protein abundance, and [ 3 ] accuracy of relative quantification of protein abundance in individual biological samples. A family of linear mixed models similar to that in MSstats http://www.msstats.org were fit to all the datasets. Then we conducted the regularization and hypothesis test to control the selection false discovery rate. The results demonstrated that our proposed feature selection approach enhanced sensitivity and specificity of the conclusions, was robust to the amount of noisy fragments, and increased the correlation of subject quantification between SRM and DIA workflows. Importantly, the performance exceeded that of the frequently used 'top 3' approach, which consists of using three spectral features with the highest average intensity between runs. Furthermore, we showed that our proposed approach outperforms using correlation to select the information features. Lin-Yang Cheng, Yansheng Liu, Ching-Yun Chang, Hannes L. Röst, Ruedi Aebersold, Olga Vitek |
BMC Bioinform. | 6 |
| 2014 | A framework for installable external tools in SkylineabstractUNLABELLED: Skyline is a Windows client application for targeted proteomics method creation and quantitative data analysis. The Skyline document model contains extensive mass spectrometry data from targeted proteomics experiments performed using selected reaction monitoring, parallel reaction monitoring and data-independent and data-dependent acquisition methods. Researchers have developed software tools that perform statistical analysis of the experimental data contained within Skyline documents. The new external tools framework allows researchers to integrate their tools into Skyline without modifying the Skyline codebase. Installed tools provide point-and-click access to downstream statistical analysis of data processed in Skyline. The framework also specifies a uniform interface to format tools for installation into Skyline. Tool developers can now easily share their tools with proteomics researchers using Skyline. AVAILABILITY AND IMPLEMENTATION: Skyline is available as a single-click self-updating web installation at http://skyline.maccosslab.org. This Web site also provides access to installable external tools and documentation. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Daniel Broudy, Trevor Killeen, Meena Choi, Nicholas Shulman, Deepak R. Mani, Susan E. Abbatiello, Deepak Mani, Rushdy Ahmad, Alexandria K. Sahu, Birgit Schilling, Kaipo Tamura, Yuval Boss, Vagisha Sharma, Bradford W. Gibson, Steven A. Carr, Olga Vitek, Michael J. MacCoss, Brendan MacLean |
Bioinform. | 16 |
| 2014 | MSstats: an R package for statistical analysis of quantitative mass spectrometry-based proteomic experimentsabstractUNLABELLED: MSstats is an R package for statistical relative quantification of proteins and peptides in mass spectrometry-based proteomics. Version 2.0 of MSstats supports label-free and label-based experimental workflows and data-dependent, targeted and data-independent spectral acquisition. It takes as input identified and quantified spectral peaks, and outputs a list of differentially abundant peptides or proteins, or summaries of peptide or protein relative abundance. MSstats relies on a flexible family of linear mixed models. AVAILABILITY AND IMPLEMENTATION: The code, the documentation and example datasets are available open-source at www.msstats.org under the Artistic-2.0 license. The package can be downloaded from www.msstats.org or from Bioconductor www.bioconductor.org and used in an R command line workflow. The package can also be accessed as an external tool in Skyline (Broudy et al., 2014) and used via graphical user interface. Meena Choi, Ching-Yun Chang, Timothy Clough, Daniel Broudy, Trevor Killeen, Brendan MacLean, Olga Vitek |
Bioinform. | 7 |
| 2013 | Shrinkage estimation of dispersion in Negative Binomial models for RNA-seq experiments with small sample sizeabstractMOTIVATION: RNA-seq experiments produce digital counts of reads that are affected by both biological and technical variation. To distinguish the systematic changes in expression between conditions from noise, the counts are frequently modeled by the Negative Binomial distribution. However, in experiments with small sample size, the per-gene estimates of the dispersion parameter are unreliable. METHOD: We propose a simple and effective approach for estimating the dispersions. First, we obtain the initial estimates for each gene using the method of moments. Second, the estimates are regularized, i.e. shrunk towards a common value that minimizes the average squared difference between the initial estimates and the shrinkage estimates. The approach does not require extra modeling assumptions, is easy to compute and is compatible with the exact test of differential expression. RESULTS: We evaluated the proposed approach using 10 simulated and experimental datasets and compared its performance with that of currently popular packages edgeR, DESeq, baySeq, BBSeq and SAMseq. For these datasets, sSeq performed favorably for experiments with small sample size in sensitivity, specificity and computational time. AVAILABILITY: http://www.stat.purdue.edu/∼ovitek/Software.html and Bioconductor. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Danni Yu, Wolfgang Huber, Olga Vitek |
Bioinform. | 3 |
| 2013 | Leading a Statistical Bioinformatics Lab: It's All About Finding BalanceabstractMy dissertation focused on a problem in structural biology, Olga Vitek |
PLoS Comput. Biol. | 1 |
| 2012 | Spatial segmentation and feature selection for desi imaging mass spectrometry data with spatially-aware sparse clusteringabstractRecent experimental advances in matrix-assisted laser desorption/ionization (MALDI) and desorption electrospray ionization (DESI) have demonstrated the usefulness of these technologies in the molecular imaging of biological samples. However, development of computational methods for the statistical interpretation and analysis of the chemical differences present in the distinct regions of these samples is still a major challenge. In this poster, we propose statistically-minded methods and computational tools for analyzing DESI imaging experiments. Specifically, we present techniques for signal processing and unsupervised multivariate image segmentation, which are also applicable to other imaging mass spectrometry (IMS) methods such as MALDI. Signal processing of DESI spectra typically involves binning to reduce dimensionality, but this inefficient for downstream analysis as it retains empty regions of the mass spectrum. In our proposed processing step, we apply a novel peak picking algorithm based on windowed smoothing splines that allows adaptive resolution based on spectral profile. With this approach, peaks are aligned using a recursive dynamic programming algorithm which accounts for the heterogenous nature of IMS data by making pairwise alignments between pixels based on their proximity. Peaks are then normalized using total ion count. In order to segment the sample into sub-regions of homogenous chemical composition in MALDI images, Alexandrov & Kobarg [ 1 ] proposed two efficient spatially-aware clustering techniques. We demonstrate that these approaches are also useful for DESI. Moreover, we extend one of these clustering methods using statistical regularization techniques that enable simultaneous feature selection of structurally-important peaks and facilitate interpretation. We evaluate the performance of the proposed methods in two applications. First, in a non-biological application of DESI-imaging, we recreate a painting from the clustering of its DESI mass spectra (Figure 1 ). Since the visual content of the painting is known, it can be used as a gold standard to evaluate the performance of these methods. In the second application, we present the spatial segmentation of a fetal pig section, and evaluate the performance of our methods by the quality of the mapping between the spatial segmentation and the morphological and functional structures (Figure 2 ). We show that statistical regularization improves accuracy and interpretation of the spatial segmentation over existing approaches. (A) Optical scan of a painting as an example of a non-biological sample and (B) the results of spatially-aware sparse clustering with C = 15 clusters. The discovered sub-regions often match up with the known colors used in the painting, but also shows regions that could not be accurately distinguished, such as the gray donkey in the lower right. (C) The mean mass spectrum of the cluster corresponding to the red houses with the 27 structurally-important ions selected via regularization for s = 0.2 marked in red and (D) the ion image for one such peak at m/z 419. (A) The fetal pig section and (B) the results of spatially-aware sparse clustering with C = 15 clusters. Note that the clusters are regions of homogenous chemical composition and often match up with visually distinguishable anatomical features, while also revealing possible hidden chemical structures in regions that are visually ambiguous. (C) The mean mass spectrum of the cluster corresponding to the liver and part of the brain with the 25 structurally-important ions selected via regularization for s = 0.6 marked in red and (D) the ion image for one such peak at m/z 888. Kyle D. Bemis, Livia Eberlin, Christina Ferreira, R. Cooks, Olga Vitek |
BMC Bioinform. | 5 |
| 2012 | Statistical protein quantification and significance analysis in label-free LC-MS experiments with complex designsabstractBACKGROUND: Liquid chromatography coupled with tandem mass spectrometry (LC-MS/MS) is widely used for quantitative proteomic investigations. The typical output of such studies is a list of identified and quantified peptides. The biological and clinical interest is, however, usually focused on quantitative conclusions at the protein level. Furthermore, many investigations ask complex biological questions by studying multiple interrelated experimental conditions. Therefore, there is a need in the field for generic statistical models to quantify protein levels even in complex study designs. RESULTS: We propose a general statistical modeling approach for protein quantification in arbitrary complex experimental designs, such as time course studies, or those involving multiple experimental factors. The approach summarizes the quantitative experimental information from all the features and all the conditions that pertain to a protein. It enables both protein significance analysis between conditions, and protein quantification in individual samples or conditions. We implement the approach in an open-source R-based software package MSstats suitable for researchers with a limited statistics and programming background. CONCLUSIONS: We demonstrate, using as examples two experimental investigations with complex designs, that a simultaneous statistical modeling of all the relevant features and conditions yields a higher sensitivity of protein significance analysis and a higher accuracy of protein quantification as compared to commonly employed alternatives. The software is available at http://www.stat.purdue.edu/~ovitek/Software.html. Timothy Clough, Safia Thaminy, Susanne Ragg, Ruedi Aebersold, Olga Vitek |
BMC Bioinform. | 5 |
| 2012 | A statistical model-building perspective to identification of MS/MS spectra with PeptideProphetabstractPeptideProphet is a post-processing algorithm designed to evaluate the confidence in identifications of MS/MS spectra returned by a database search. In this manuscript we describe the "what and how" of PeptideProphet in a manner aimed at statisticians and life scientists who would like to gain a more in-depth understanding of the underlying statistical modeling. The theory and rationale behind the mixture-modeling approach taken by PeptideProphet is discussed from a statistical model-building perspective followed by a description of how a model can be used to express confidence in the identification of individual peptides or sets of peptides. We also demonstrate how to evaluate the quality of model fit and select an appropriate model from several available alternatives. We illustrate the use of PeptideProphet in association with the Trans-Proteomic Pipeline, a free suite of software used for protein identification. Kelvin Ma, Olga Vitek, Alexey I. Nesvizhskii |
BMC Bioinform. | 2 |
| 2011 | Noise reduction in genome-wide perturbation screens using linear mixed-effect modelsabstractMOTIVATION: High-throughput perturbation screens measure the phenotypes of thousands of biological samples under various conditions. The phenotypes measured in the screens are subject to substantial biological and technical variation. At the same time, in order to enable high throughput, it is often impossible to include a large number of replicates, and to randomize their order throughout the screens. Distinguishing true changes in the phenotype from stochastic variation in such experimental designs is extremely challenging, and requires adequate statistical methodology. RESULTS: We propose a statistical modeling framework that is based on experimental designs with at least two controls profiled throughout the experiment, and a normalization and variance estimation procedure with linear mixed-effects models. We evaluate the framework using three comprehensive screens of Saccharomyces cerevisiae, which involve 4940 single-gene knock-out haploid mutants, 1127 single-gene knock-out diploid mutants and 5798 single-gene overexpression haploid strains. We show that the proposed approach (i) can be used in conjunction with practical experimental designs; (ii) allows extensions to alternative experimental workflows; (iii) enables a sensitive discovery of biologically meaningful changes; and (iv) strongly outperforms the existing noise reduction procedures. AVAILABILITY: All experimental datasets are publicly available at www.ionomicshub.org. The R package HTSmix is available at http://www.stat.purdue.edu/~ovitek/HTSmix.html. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Danni Yu, John Danku, Ivan Baxter, Olena K. Vatamaniuk, David E. Salt, Olga Vitek |
Bioinform. | 7 |
| 2011 | Identification and quantification of metabolites in 1H NMR spectra by Bayesian model selectionabstractMOTIVATION: Nuclear magnetic resonance (NMR) spectroscopy is widely used for high-throughput characterization of metabolites in complex biological mixtures. However, accurate interpretation of the spectra in terms of identities and abundances of metabolites can be challenging, in particular in crowded regions with heavy peak overlap. Although a number of computational approaches for this task have recently been proposed, they are not entirely satisfactory in either accuracy or extent of automation. RESULTS: We introduce a probabilistic approach Bayesian Quantification (BQuant), for fully automated database-based identification and quantification of metabolites in local regions of (1)H NMR spectra. The approach represents the spectra as mixtures of reference profiles from a database, and infers the identities and the abundances of metabolites by Bayesian model selection. We show using a simulated dataset, a spike-in experiment and a metabolomic investigation of plasma samples that BQuant outperforms the available automated alternatives in accuracy for both identification and quantification. AVAILABILITY: The R package BQuant is available at: http://www.stat.purdue.edu/~ovitek/BQuant-Web/. Shucha Zhang, Susanne Ragg, Daniel Raftery, Olga Vitek |
Bioinform. | 5 |
| 2011 | Computational Mass Spectrometry-Based ProteomicsabstractDOAJ is a unique and extensive index of diverse open access journals from around the world, driven by a growing community, committed to ensuring quality content is freely available online for everyone. Lukas Käll, Olga Vitek |
PLoS Comput. Biol. | 2 |
| 2009 | Getting Started in Computational Mass Spectrometry-Based ProteomicsabstractProteomics aims at a large-scale characterization of localization, abundance, post-translational modifications, and biomolecular interactions of the proteins in an organism, with the goal of understanding their function. An extensive insight can be obtained by identifying and quantifying the components of biological mixtures. For example, a) In studies of biomolecular networks, partners interacting with a protein can help determine its function. It is possible to experimentally isolate protein complexes, e.g., using tag affinity purification. Identification of the components of this mixture helps determine potential interactors [1]. b) Post-translational modifications such as phosphorylation play an important role in regulating biological processes, e.g., cellular growth and signaling. Identification and quantification of phosphorylated proteins and their substrates helps elucidate complex signaling pathway phosphorylation events [2]. c) Molecular biomarkers, i.e., proteins for which changes in abundance are indicative of an early onset of a disease or a therapy response, are of interest in clinical research. Identifying and quantifying components of a biofluid such as serum helps detect proteins with such discriminative ability [3]. d) A goal of genome annotation is the discovery and validation of protein-coding regions. Identifying peptides and proteins in a cell helps confirm and improve the annotations at the translational level, e.g., by confirming the presence of intron boundaries or alternative splicings [4].
Mass spectrometry is a method of choice for protein identification and quantification due to its sensitivity and to the versatility of the instrumentation [5],[6]. A typical “bottom-up” workflow experimentally digests the proteins into a mixture of peptides with an enzyme such as trypsin. This is necessary, in part, because the sensitivity of the mass spectrometer is much higher for peptides than for proteins. The peptides are then injected onto a liquid chromatography (LC) column from which they elute sequentially. The eluted peptides are ionized and separated by the mass spectrometer according to their ratio of mass to charge (m/z) in a mass spectrum (MS).
The collection of mass spectra obtained at different elution times forms an LC-MS run shown in Figure 1A. Peaks in the run correspond to peptide ions; however, the sequence of amino acids underlying each peak is unknown. For identification, the mass spectrometer isolates the biological material from a peak (called precursor ion in this context), and subjects it to a high-collision energy. The energy breaks the peptide at different amide bonds, and the resulting fragments are separated according to their m/z in a secondary spectrum (called MS2, MS/MS, or tandem MS), shown in Figure 1B. Distances between peaks in the MS/MS spectrum are used to infer the peptide sequence of the parent LC-MS peak.
Figure 1
Example of spectral data.
Peak intensity is related to the abundances of peptides, and can be used for relative quantification. With the label-free approach, a separate LC-MS run is obtained for each biological sample, and peaks are quantified and compared across runs. In stable isotopic labeling workflow, samples from different groups are labeled metabolically (e.g., in SILAC, where stable isotopes are included in the growth medium of an organism), or chemically (e.g., in ICAT or iTRAQ, where reacting chemical labels are applied after tryptic digestion). Several samples (e.g., one from each group) are then mixed, and their peaks are identified and quantified within the same run. Finally, a targeted workflow based, for example, on selected reaction monitoring (SRM) [7], increases sensitivity and specificity by monitoring signals from a list of predefined peptides.
The design of proteomic experiments, and subsequent analysis of the spectra, involves extensive computation and requires expertise at the intersection of computer science, engineering, and statistics. It presents exciting opportunities for both methodological and applied computational research. Olga Vitek |
PLoS Comput. Biol. | 1 |
| 2008 | Corra: Computational framework and tools for LC-MS discovery and targeted mass spectrometry-based proteomicsabstractBACKGROUND: Quantitative proteomics holds great promise for identifying proteins that are differentially abundant between populations representing different physiological or disease states. A range of computational tools is now available for both isotopically labeled and label-free liquid chromatography mass spectrometry (LC-MS) based quantitative proteomics. However, they are generally not comparable to each other in terms of functionality, user interfaces, information input/output, and do not readily facilitate appropriate statistical data analysis. These limitations, along with the array of choices, present a daunting prospect for biologists, and other researchers not trained in bioinformatics, who wish to use LC-MS-based quantitative proteomics. RESULTS: We have developed Corra, a computational framework and tools for discovery-based LC-MS proteomics. Corra extends and adapts existing algorithms used for LC-MS-based proteomics, and statistical algorithms, originally developed for microarray data analyses, appropriate for LC-MS data analysis. Corra also adapts software engineering technologies (e.g. Google Web Toolkit, distributed processing) so that computationally intense data processing and statistical analyses can run on a remote server, while the user controls and manages the process from their own computer via a simple web interface. Corra also allows the user to output significantly differentially abundant LC-MS-detected peptide features in a form compatible with subsequent sequence identification via tandem mass spectrometry (MS/MS). We present two case studies to illustrate the application of Corra to commonly performed LC-MS-based biological workflows: a pilot biomarker discovery study of glycoproteins isolated from human plasma samples relevant to type 2 diabetes, and a study in yeast to identify in vivo targets of the protein kinase Ark1 via phosphopeptide profiling. CONCLUSION: The Corra computational framework leverages computational innovation to enable biologists or other researchers to process, analyze and visualize LC-MS data with what would otherwise be a complex and not user-friendly suite of tools. Corra enables appropriate statistical analyses, with controlled false-discovery rates, ultimately to inform subsequent targeted identification of differentially abundant peptides by MS/MS. For the user not trained in bioinformatics, Corra represents a complete, customizable, free and open source computational platform enabling LC-MS-based proteomic workflows, and as such, addresses an unmet need in the LC-MS proteomics field. Mi-Youn K. Brusniak, Bernd Bodenmiller, David S. Campbell, Kelly Cooke, James Eddes, Andrew Garbutt, Hollis Lau, Simon Letarte, Lukas N. Mueller, Vagisha Sharma, Olga Vitek, Ruedi Aebersold, Julian D. Watts |
BMC Bioinform. | 11 |