Russell Schwartz

dblp:77/4437 · DBLP profile ↗
← Back
54ranked-venue papers
5as first author
9since 2021 · last 2025
0000-0002-4970-2252ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 49 · 4 first-author · 9 since 2021Theory of computation · 5 · 1 first-author
YearPublicationVenuePosition
2025 Marker selection strategies for circulating tumor DNA guided by phylogenetic inference
abstract
MOTIVATION: Blood-based profiling of tumor DNA ("liquid biopsy") offers great prospects for non-invasive early cancer diagnosis and clinical guidance, but requires further computational advances to become a robust quantitative assay of tumor clonal evolution. We propose new methods to better characterize tumor clonal dynamics from circulating tumor DNA (ctDNA), through application to two specific tasks: (i) applying longitudinal ctDNA data to refine phylogeny models of clonal evolution, and (ii) quantifying changes in clonal frequencies that may be indicative of treatment response or tumor progression. We pose these through a probabilistic framework for optimally identifying markers and using them to characterize clonal evolution. RESULTS: We first estimate a density over clonal tree models using bootstrap samples over pre-treatment tissue-based sequence data. We then refine these models over successive longitudinal samples. We use the resulting framework for modeling and refining tree densities to pose a set of optimization problems for selecting ctDNA markers to maximize measures of utility for reducing uncertainty in phylogeny models and quantifying clonal frequencies given the models. We tested our methods on synthetic data and showed them to be effective at refining tree densities and inferring clonal frequencies. Application to real tumor data further demonstrated the methods' effectiveness in refining a lineage model and assessing its clonal frequencies. The work shows the power of computational methods to improve marker selection, clonal lineage reconstruction, and clonal dynamics profiling for more precise and quantitative assays of somatic evolution and tumor progression. AVAILABILITY AND IMPLEMENTATION: https://github.com/CMUSchwartzLab/Mase-phi.git. (DOI: 10.5281/zenodo.14776163).
Xuecong Fu, Zhicheng Luo, Yueqian Deng, William Laframboise, David Bartlett, Russell Schwartz
Bioinform.6
2024 Determining Optimal Placement of Copy Number Aberration Impacted Single Nucleotide Variants in a Tumor Progression History
Chih Hao Wu, Suraj Joshi, Welles Robinson, Paul F. Robbins, Russell Schwartz, Süleyman Cenk Sahinalp, Salem Malikic
RECOMB5
2023 Ten simple rules for writing a PLOS Computational Biology quick tips article
abstract
At that time, the change in how we communicate sciences had already begun [2]: social media was blooming, the information deluge was ongoing, and the fear of missing information was anchored in each of us.The Ten Simple Rules collection filled a space in scientific publications where researchers eager to share their experiences, wisdom, and doubts could quickly do it using a colloquial narrative.The topics covered were broad themes in scientific practice, such as soft skills and career development, captured in a concise and quick-to-read format, something between a blog post and a scientific article.This type of article attracted much interest from readers.In 2018, the collection reached the milestone of 1,000 Rules [3], and today this figure is above 250 papers (2,500 Rules!) [4].With the increase of technical and scientific topics, in 2013, PLOS Computational Biology tried a new experience with a similar format-we introduced "Quick Tips" (QT) articles-with the attempt to make a clear distinction between the more specific and focused scientific activities and skills presented with resources, databases, and other tools in Quick Tips versus the broader themes presented in a Ten Simple Rules article."Ten Simple Rules for Writing a PLOS Ten Simple Rules Article" [5] explains the Ten Simple Rules concept, format, and reasoning very well and is still relevant today.Inspired by that article, we wrote this Ten Simple Rules paper intending to accomplish a similar task: explaining to our community what a Quick Tips article is about and how a Quick Tips article differs from a Ten Simple Rules article.The authors are the Section Editors of the Education Collection (PP and BFFO), which encompasses the Quick Tips, and the Section Editors of the Ten Simple Rules Collection (RS and SM).In our work at PLOS CB, we are routinely deliberating the merits of a submission being a Ten Simple Rules or a Quick Tips, and for this reason, we decided to put these Ten Simple Rules about Quick Tips together.Historically, several papers were submitted as Ten Simple Rules that should have been Quick Tips.Still, we will refrain from commiserating about things not done but rather present what we hope will be a clear distinction that will make it obvious in the future why articles are best considered as Quick Tips versus Ten Simple Rules (or vice-versa).We hope and plan for this present article to be helpful for future writers willing to contribute to this collection and to help clarify the differences between Ten Simple Rules and Quick Tips.Why is this not a Quick Tips article but a Ten Simple Rules article, and why are we not as endearing as Dashnow, Lonsdale, and Bourne were in doing a "Ten Quick Tips for Writing a PLOS Quick Tips Article"?The reason is simple.
Patricia M. Palagi, Russell Schwartz, Scott Markel, B. F. Francis Ouellette
PLoS Comput. Biol.2
2022 A Clonal Evolution Simulator for Planning Somatic Evolution Studies
Arjun Srivatsa, Haoyun Lei, Russell Schwartz
ISBRA3
2022 Reconstructing tumor clonal lineage trees incorporating single-nucleotide variants, copy number alterations and structural variations
abstract
MOTIVATION: Cancer develops through a process of clonal evolution in which an initially healthy cell gives rise to progeny gradually differentiating through the accumulation of genetic and epigenetic mutations. These mutations can take various forms, including single-nucleotide variants (SNVs), copy number alterations (CNAs) or structural variations (SVs), with each variant type providing complementary insights into tumor evolution as well as offering distinct challenges to phylogenetic inference. RESULTS: In this work, we develop a tumor phylogeny method, TUSV-ext, which incorporates SNVs, CNAs and SVs into a single inference framework. We demonstrate on simulated data that the method produces accurate tree inferences in the presence of all three variant types. We further demonstrate the method through application to real prostate tumor data, showing how our approach to coordinated phylogeny inference and clonal construction with all three variant types can reveal a more complicated clonal structure than is suggested by prior work, consistent with extensive polyclonal seeding or migration. AVAILABILITY AND IMPLEMENTATION: https://github.com/CMUSchwartzLab/TUSV-ext. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Xuecong Fu, Haoyun Lei, Yifeng Tao, Russell Schwartz
Bioinform.4
2022 Semi-deconvolution of bulk and single-cell RNA-seq data with application to metastatic progression in breast cancer
abstract
MOTIVATION: Identifying cell types and their abundances and how these evolve during tumor progression is critical to understanding the mechanisms of metastasis and identifying predictors of metastatic potential that can guide the development of new diagnostics or therapeutics. Single-cell RNA sequencing (scRNA-seq) has been especially promising in resolving heterogeneity of expression programs at the single-cell level, but is not always feasible, e.g. for large cohort studies or longitudinal analysis of archived samples. In such cases, clonal subpopulations may still be inferred via genomic deconvolution, but deconvolution methods have limited ability to resolve fine clonal structure and may require reference cell type profiles that are missing or imprecise. Prior methods can eliminate the need for reference profiles but show unstable performance when few bulk samples are available. RESULTS: In this work, we develop a new method using reference scRNA-seq to interpret sample collections for which only bulk RNA-seq is available for some samples, e.g. clonally resolving archived primary tissues using scRNA-seq from metastases. By integrating such information in a Quadratic Programming framework, our method can recover more accurate cell types and corresponding cell type abundances in bulk samples. Application to a breast tumor bone metastases dataset confirms the power of scRNA-seq data to improve cell type inference and quantification in same-patient bulk samples. AVAILABILITY AND IMPLEMENTATION: Source code is available on Github at https://github.com/CMUSchwartzLab/RADs.
Haoyun Lei, Xiaoyan A. Guo, Yifeng Tao, Xuecong Fu, Steffi Oesterreich, Adrian V. Lee, Russell Schwartz
Bioinform.8
2021 ConTreeDP: A consensus method of tumor trees based on maximum directed partition support problem
abstract
Phylogenetic inference has become a crucial tool for interpreting cancer genomic data, but continuing advances in our understanding of somatic mutability in cancer, genomic technologies for profiling it, and the scale of data available have created a persistent need for new algorithms able to deal with these challenges. One particular need has been for new forms of consensus tree algorithms, which present special challenges in the cancer space for dealing with heterogeneous data, short evolutionary time scales, and rapid mutation by a wide variety of somatic mutability mechanisms. We develop a new consensus tree method for clonal phylogenetics, ConTreeDP, based on a formulation of the Maximum Directed Partition Support Consensus Tree (MDPSCT) problem. We demonstrate theoretically and empirically that our approach can efficiently and accurately compute clonal consensus trees from cancer genomic data. Availability: https://github.conCMUSchwartzLab/ConTreeDP
Xuecong Fu, Russell Schwartz
BIBM2
2021 Tumor heterogeneity assessed by sequencing and fluorescence in situ hybridization (FISH) data
abstract
MOTIVATION: Computational reconstruction of clonal evolution in cancers has become a crucial tool for understanding how tumors initiate and progress and how this process varies across patients. The field still struggles, however, with special challenges of applying phylogenetic methods to cancers, such as the prevalence and importance of copy number alteration (CNA) and structural variation events in tumor evolution, which are difficult to profile accurately by prevailing sequencing methods in such a way that subsequent reconstruction by phylogenetic inference algorithms is accurate. RESULTS: In this work, we develop computational methods to combine sequencing with multiplex interphase fluorescence in situ hybridization to exploit the complementary advantages of each technology in inferring accurate models of clonal CNA evolution accounting for both focal changes and aneuploidy at whole-genome scales. By integrating such information in an integer linear programming framework, we demonstrate on simulated data that incorporation of FISH data substantially improves accurate inference of focal CNA and ploidy changes in clonal evolution from deconvolving bulk sequence data. Analysis of real glioblastoma data for which FISH, bulk sequence and single cell sequence are all available confirms the power of FISH to enhance accurate reconstruction of clonal copy number evolution in conjunction with bulk and optionally single-cell sequence data. AVAILABILITY AND IMPLEMENTATION: Source code is available on Github at https://github.com/CMUSchwartzLab/FISH_deconvolution. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Haoyun Lei, E. Michael Gertz, Alejandro A. Schäffer, Xuecong Fu, Yifeng Tao, Kerstin Heselmeyer-Haddad, Irianna Torres, Guibo Li, Liqin Xu, Yong Hou, Kui Wu 0005, Xulian Shi, Mike Dean, Thomas Ried, Russell Schwartz
Bioinform.15
2021 Assessing the contribution of tumor mutational phenotypes to cancer progression risk
abstract
Cancer occurs via an accumulation of somatic genomic alterations in a process of clonal evolution. There has been intensive study of potential causal mutations driving cancer development and progression. However, much recent evidence suggests that tumor evolution is normally driven by a variety of mechanisms of somatic hypermutability, which act in different combinations or degrees in different cancers. These variations in mutability phenotypes are predictive of progression outcomes independent of the specific mutations they have produced to date. Here we explore the question of how and to what degree these differences in mutational phenotypes act in a cancer to predict its future progression. We develop a computational paradigm using evolutionary tree inference (tumor phylogeny) algorithms to derive features quantifying single-tumor mutational phenotypes, followed by a machine learning framework to identify key features predictive of progression. Analyses of breast invasive carcinoma and lung carcinoma demonstrate that a large fraction of the risk of future clinical outcomes of cancer progression-overall survival and disease-free survival-can be explained solely from mutational phenotype features derived from the phylogenetic analysis. We further show that mutational phenotypes have additional predictive power even after accounting for traditional clinical and driver gene-centric genomic predictors of progression. These results confirm the importance of mutational phenotypes in contributing to cancer progression risk and suggest strategies for enhancing the predictive power of conventional clinical data or driver-centric biomarkers.
Yifeng Tao, Ashok Rajaraman, Xiaoyue Cui, Ziyi Cui, Haoran Chen 0006, Yuanqi Zhao, Jesse Eaton, Hannah Kim 0003, Jian Ma 0004, Russell Schwartz
PLoS Comput. Biol.10
2020 Metasubtract: an R-package to analytically produce leave-one-out meta-analysis GWAS summary statistics
abstract
SUMMARY: statistics from a meta-analysis of genome-wide association studies (meta-GWAS) can be used for many follow-up analyses. One valuable application is the creation of polygenic scores. However, if polygenic scores are calculated in a validation cohort that was part of the meta-GWAS consortium, this cohort is not independent and analyses will therefore yield inflated results. The R package 'MetaSubtract' was developed to subtract the results of the validation cohort from meta-GWAS summary statistics analytically. The statistical formulas for a meta-analysis were inverted to compute corrected summary statistics of a meta-GWAS leaving one (or more) cohort(s) out. These formulas have been implemented in MetaSubtract for different meta-analyses methods (fixed effects inverse variance or square root sample size weighted z-score) accounting for no, single or double genomic control correction. Results obtained by MetaSubtract correlate very well to those calculated using the traditional way, i.e. by performing a meta-analysis leaving out the validation cohort. In conclusion, MetaSubtract allows researchers to compute meta-GWAS summary statistics that are independent of the GWAS results of the validation cohort without requiring access to the cohort level GWAS results of the corresponding meta-GWAS consortium. AVAILABILITY AND IMPLEMENTATION: https://cran.r-project.org/web/packages/MetaSubtract. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Ilja M. Nolte, Russell Schwartz
Bioinform.2
2020 Efficient toolkit implementing best practices for principal component analysis of population genetic data
abstract
MOTIVATION: Principal component analysis (PCA) of genetic data is routinely used to infer ancestry and control for population structure in various genetic analyses. However, conducting PCA analyses can be complicated and has several potential pitfalls. These pitfalls include (i) capturing linkage disequilibrium (LD) structure instead of population structure, (ii) projected PCs that suffer from shrinkage bias, (iii) detecting sample outliers and (iv) uneven population sizes. In this work, we explore these potential issues when using PCA, and present efficient solutions to these. Following applications to the UK Biobank and the 1000 Genomes project datasets, we make recommendations for best practices and provide efficient and user-friendly implementations of the proposed solutions in R packages bigsnpr and bigutilsr. RESULTS: For example, we find that PC19-PC40 in the UK Biobank capture complex LD structure rather than population structure. Using our automatic algorithm for removing long-range LD regions, we recover 16 PCs that capture population structure only. Therefore, we recommend using only 16-18 PCs from the UK Biobank to account for population structure confounding. We also show how to use PCA to restrict analyses to individuals of homogeneous ancestry. Finally, when projecting individual genotypes onto the PCA computed from the 1000 Genomes project data, we find a shrinkage bias that becomes large for PC5 and beyond. We then demonstrate how to obtain unbiased projections efficiently using bigsnpr. Overall, we believe this work would be of interest for anyone using PCA in their analyses of genetic data, as well as for other omics data. AVAILABILITY AND IMPLEMENTATION: R packages bigsnpr and bigutilsr can be installed from either CRAN or GitHub (see https://github.com/privefl/bigsnpr). A tutorial on the steps to perform PCA on 1000G data is available at https://privefl.github.io/bigsnpr/articles/bedpca.html. All code used for this paper is available at https://github.com/privefl/paper4-bedpca/tree/master/code. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Florian Privé, Keurcien Luu, Michael G. B. Blum, John J. McGrath, Bjarni J. Vilhjálmsson, Russell Schwartz
Bioinform.6
2020 Robust and accurate deconvolution of tumor populations uncovers evolutionary mechanisms of breast cancer metastasis
abstract
MOTIVATION: Cancer develops and progresses through a clonal evolutionary process. Understanding progression to metastasis is of particular clinical importance, but is not easily analyzed by recent methods because it generally requires studying samples gathered years apart, for which modern single-cell sequencing is rarely an option. Revealing the clonal evolution mechanisms in the metastatic transition thus still depends on unmixing tumor subpopulations from bulk genomic data. METHODS: We develop a novel toolkit called robust and accurate deconvolution (RAD) to deconvolve biologically meaningful tumor populations from multiple transcriptomic samples spanning the two progression states. RAD uses gene module compression to mitigate considerable noise in RNA, and a hybrid optimizer to achieve a robust and accurate solution. Finally, we apply a phylogenetic algorithm to infer how associated cell populations adapt across the metastatic transition via changes in expression programs and cell-type composition. RESULTS: We validated the superior robustness and accuracy of RAD over alternative algorithms on a real dataset, and validated the effectiveness of gene module compression on both simulated and real bulk RNA data. We further applied the methods to a breast cancer metastasis dataset, and discovered common early events that promote tumor progression and migration to different metastatic sites, such as dysregulation of ECM-receptor, focal adhesion and PI3k-Akt pathways. AVAILABILITY AND IMPLEMENTATION: The source code of the RAD package, models, experiments and technical details such as parameters, is available at https://github.com/CMUSchwartzLab/RAD. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yifeng Tao, Haoyun Lei, Xuecong Fu, Adrian V. Lee, Jian Ma 0004, Russell Schwartz
Bioinform.6
2020 Predicting target genes of non-coding regulatory variants with IRT
abstract
SUMMARY: Interpreting genetic variants of unknown significance (VUS) is essential in clinical applications of genome sequencing for diagnosis and personalized care. Non-coding variants remain particularly difficult to interpret, despite making up a large majority of trait associations identified in genome-wide association studies (GWAS) analyses. Predicting the regulatory effects of non-coding variants on candidate genes is a key step in evaluating their clinical significance. Here, we develop a machine-learning algorithm, Inference of Connected expression quantitative trait loci (eQTLs) (IRT), to predict the regulatory targets of non-coding variants identified in studies of eQTLs. We assemble datasets using eQTL results from the Genotype-Tissue Expression (GTEx) project and learn to separate positive and negative pairs based on annotations characterizing the variant, gene and the intermediate sequence. IRT achieves an area under the receiver operating characteristic curve (ROC-AUC) of 0.799 using random cross-validation, and 0.700 for a more stringent position-based cross-validation. Further evaluation on rare variants and experimentally validated regulatory variants shows a significant enrichment in IRT identifying the true target genes versus negative controls. In gene-ranking experiments, IRT achieves a top-1 accuracy of 50% and top-3 accuracy of 90%. Salient features, including GC-content, histone modifications and Hi-C interactions are further analyzed and visualized to illustrate their influences on predictions. IRT can be applied to any VUS of interest and each candidate nearby gene to output a score reflecting the likelihood of regulatory effect on the expression level. These scores can be used to prioritize variants and genes to assist in patient diagnosis and GWAS follow-up studies. AVAILABILITY AND IMPLEMENTATION: Codes and data used in this work are available at https://github.com/miaecle/eQTL_Trees. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Zhenqin Wu, Nilah M. Ioannidis, James Zou 0001, Russell Schwartz
Bioinform.4
2020 IBDkin: fast estimation of kinship coefficients from identity by descent segments
abstract
MOTIVATION: Estimation of pairwise kinship coefficients in large datasets is computationally challenging because the number of related individuals increases quadratically with sample size. RESULTS: We present IBDkin, a software package written in C for estimating kinship coefficients from identity by descent (IBD) segments. We use IBDkin to estimate kinship coefficients for 7.95 billion pairs of individuals in the UK Biobank who share at least one detected IBD segment with length ≥ 4 cM. AVAILABILITY AND IMPLEMENTATION: https://github.com/YingZhou001/IBDkin. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Sharon R. Browning, Brian L. Browning, Russell Schwartz
Bioinform.4
2019 Tumor Copy Number Deconvolution Integrating Bulk and Single-Cell Sequencing Data
Haoyun Lei, Bochuan Lyu, E. Michael Gertz, Alejandro A. Schäffer, Xulian Shi, Kui Wu 0005, Guibo Li, Liqin Xu, Yong Hou, Mike Dean, Russell Schwartz
RECOMB11
2018 Reconstructing Tumor Evolution and Progression in Structurally Variant Cancer Cells
Russell Schwartz
COCOON1
2018 Deconvolution and phylogeny inference of structural variations in tumor genomic samples
abstract
Motivation: Phylogenetic reconstruction of tumor evolution has emerged as a crucial tool for making sense of the complexity of emerging cancer genomic datasets. Despite the growing use of phylogenetics in cancer studies, though, the field has only slowly adapted to many ways that tumor evolution differs from classic species evolution. One crucial question in that regard is how to handle inference of structural variations (SVs), which are a major mechanism of evolution in cancers but have been largely neglected in tumor phylogenetics to date, in part due to the challenges of reliably detecting and typing SVs and interpreting them phylogenetically. Results: We present a novel method for reconstructing evolutionary trajectories of SVs from bulk whole-genome sequence data via joint deconvolution and phylogenetics, to infer clonal sub-populations and reconstruct their ancestry. We establish a novel likelihood model for joint deconvolution and phylogenetic inference on bulk SV data and formulate an associated optimization algorithm. We demonstrate the approach to be efficient and accurate for realistic scenarios of SV mutation on simulated data. Application to breast cancer genomic data from The Cancer Genome Atlas shows it to be practical and effective at reconstructing features of SV-driven evolution in single tumors. Availability and implementation: Python source code and associated documentation are available at https://github.com/jaebird123/tusv.
Jesse Eaton, Russell Schwartz
Bioinform.3
2018 The development and application of bioinformatics core competencies to improve bioinformatics training and education
abstract
Bioinformatics is recognized as part of the essential knowledge base of numerous career paths in biomedical research and healthcare. However, there is little agreement in the field over what that knowledge entails or how best to provide it. These disagreements are compounded by the wide range of populations in need of bioinformatics training, with divergent prior backgrounds and intended application areas. The Curriculum Task Force of the International Society of Computational Biology (ISCB) Education Committee has sought to provide a framework for training needs and curricula in terms of a set of bioinformatics core competencies that cut across many user personas and training programs. The initial competencies developed based on surveys of employers and training programs have since been refined through a multiyear process of community engagement. This report describes the current status of the competencies and presents a series of use cases illustrating how they are being applied in diverse training contexts. These use cases are intended to demonstrate how others can make use of the competencies and engage in the process of their continuing refinement and application. The report concludes with a consideration of remaining challenges and future plans.
Nicola J. Mulder, Russell Schwartz, Michelle D. Brazas, Catherine Brooksbank, Bruno A. Gaëta, Sarah L. Morgan, Mark A. Pauley, Anne G. Rosenwald, Gabriella Rustici, Michael L. Sierk, Tandy J. Warnow, Lonnie R. Welch
PLoS Comput. Biol.2
2017 Automated deconvolution of structured mixtures from heterogeneous tumor genomic data
abstract
With increasing appreciation for the extent and importance of intratumor heterogeneity, much attention in cancer research has focused on profiling heterogeneity on a single patient level. Although true single-cell genomic technologies are rapidly improving, they remain too noisy and costly at present for population-level studies. Bulk sequencing remains the standard for population-scale tumor genomics, creating a need for computational tools to separate contributions of multiple tumor clones and assorted stromal and infiltrating cell populations to pooled genomic data. All such methods are limited to coarse approximations of only a few cell subpopulations, however. In prior work, we demonstrated the feasibility of improving cell type deconvolution by taking advantage of substructure in genomic mixtures via a strategy called simplicial complex unmixing. We improve on past work by introducing enhancements to automate learning of substructured genomic mixtures, with specific emphasis on genome-wide copy number variation (CNV) data, as well as the ability to process quantitative RNA expression data, and heterogeneous combinations of RNA and CNV data. We introduce methods for dimensionality estimation to better decompose mixture model substructure; fuzzy clustering to better identify substructure in sparse, noisy data; and automated model inference methods for other key model parameters. We further demonstrate their effectiveness in identifying mixture substructure in true breast cancer CNV data from the Cancer Genome Atlas (TCGA). Source code is available at https://github.com/tedroman/WSCUnmix.
Theodore Roman, Lu Xie, Russell Schwartz
PLoS Comput. Biol.3
2017 Derivative-Free Optimization of Rate Parameters of Capsid Assembly Models from Bulk in Vitro Data
abstract
The assembly of virus capsids proceeds by a complicated cascade of association and dissociation steps, the great majority of which cannot be directly experimentally observed. This has made capsid assembly a rich field for computational models, but there are substantial obstacles to model inference for such systems. Here, we describe progress on fitting kinetic rate constants defining capsid assembly models to experimental data, a difficult data-fitting problem because of the high computational cost of simulating assembly trajectories, the stochastic noise inherent to the models, and the limited and noisy data available for fitting. We evaluate the merits of data-fitting methods based on derivative-free optimization (DFO) relative to gradient-based methods used in prior work. We further explore the advantages of alternative data sources through simulation of a model of time-resolved mass spectrometry data, a technology for monitoring bulk capsid assembly that can be expected to provide much richer data than previously used static light scattering approaches. The results show that advances in both the data and the algorithms can improve model inference. More informative data sources lead to high-quality fits for all methods, but DFO methods show substantial advantages on less informative data sources that better represent current experimental practice.
Lu Xie, Gregory R. Smith, Russell Schwartz
IEEE ACM Trans. Comput. Biol. Bioinform.3
2016 Applying, Evaluating and Refining Bioinformatics Core Competencies (An Update from the Curriculum Task Force of ISCB's Education Committee)
abstract
The Curriculum Task Force (CTF) of ISCB’s Education Committee seeks to define curricular guidelines for those who educate or train bioinformatics professionals at all career stages. A recent report of the CTF [1] presented a draft set of bioinformatics core competencies, derived from the results of surveys of (1) core facility directors, (2) career opportunities, and (3) existing curricula. Since the publication of its 2014 report, the CTF has focused on the application of the guidelines in varied contexts to identify areas where refinement is needed. As a first step, the task force held an open meeting at the ISMB conference in July 2014. The ideas discussed at the meeting spawned four working groups (WGs), which focus on (i) defining core competencies for specific types and levels of bioinformatics training, (ii) mapping the curriculum guidelines and competencies to existing materials in order to identify the need for development of new materials, and (iii) identifying where revision of the guidelines may be valuable. The CTF is engaging the ISCB community through open WG meetings at ISCB’s official conferences. Thus far, the WGs have convened at the ISCB Great Lakes Bioinformatics Conference (Purdue University, May 2015) and at the ISMB/ECCB Conference (Dublin, Ireland, July 2015). Additionally, the CTF held a workshop at the Annual General Meeting of the Global Organization of Bioinformatics Learning, Education and Training (Cape Town, South Africa, November 2015). Specifically, the draft competencies have been employed in a wide range of activities and contexts (see Table 1 and [2–11]), including the development of new curricula, the analysis of existing curricula, and the creation of new roles involving bioinformatics. These activities have resulted in the identification of several areas where refinement would be useful: Table 1 Summary of the activities of the ISCB Curriculum Task Force. Identify different levels or phases of competency. It would be helpful to define different phases of competency development, or different levels of competency appropriate for distinct roles. Define competency profiles for disciplines that don’t fit into our current silos. Bioengineering provides an illustrative example of a discipline that requires core competency in bioinformatics but does not fit into our current categories. There are almost certainly others. It would be helpful if we could provide some guidance on how to produce ‘hybrid’ competency profiles, perhaps borrowing some competencies from the TF’s core set and others from different disciplines. The LifeTrain initiative (www.lifetrain.eu) [2, 3] is collecting competency profiles for a range of disciplines of relevance to the biomedical sciences and may provide a useful resource kit for this. Broaden the scope of the competency profiles in response to cutting-edge and emerging research. Current areas requiring improvement include incorporating competencies that capture a fundamental understanding of the biological principles central to analyzing biomolecular data, and broadening the user WG to include applications beyond medicine. Provide guidance on the evidence required to assess whether someone has acquired each competency. For undergraduate, Master’s and PhD programs, learning outcomes for each competency, perhaps with examples of appropriate means of assessment, would be valuable. For established professionals who need to assimilate competencies into their working lives, a different approach may be required (such as keeping a portfolio to capture evidence of competency); the CTF should seek guidance from relevant professional bodies, especially in regulated professions such as healthcare. Provide indicative course content or examples of programs that map to the competency requirements. We do not wish to prescribe what course providers should teach or how they should teach it; however, if a course provider is designing a course to meet a specific competency requirement, it may be helpful to find examples of other programs that do this successfully. One way of achieving this is by mapping existing training content to the TF’s competencies. Another way might be to provide an indication, perhaps based on several courses, of the course content that would meet the competency requirements. This would give course providers the freedom to build their own course syllabi without having to reinvent the wheel. Initiatives to collect examples of Creative Commons (or otherwise reusable) course materials will provide an extremely valuable bank of training materials that could be mapped to the core competencies.
Lonnie R. Welch, Catherine Brooksbank, Russell Schwartz, Sarah L. Morgan, Bruno A. Gaëta, Alastair M. Kilpatrick, Daniel Mietchen, Benjamin L. Moore, Nicola J. Mulder, Mark A. Pauley, William R. Pearson, Predrag Radivojac, Naomi Rosenberg, Anne G. Rosenwald, Gabriella Rustici, Tandy J. Warnow
PLoS Comput. Biol.3
2016 Classifying the Progression of Ductal Carcinoma from Single-Cell Sampled Data via Integer Linear Programming: A Case Study
abstract
Ductal Carcinoma In Situ (DCIS) is a precursor lesion of Invasive Ductal Carcinoma (IDC) of the breast. Investigating its temporal progression could provide fundamental new insights for the development of better diagnostic tools to predict which cases of DCIS will progress to IDC. We investigate the problem of reconstructing a plausible progression from single-cell sampled data of an individual with synchronous DCIS and IDC. Specifically, by using a number of assumptions derived from the observation of cellular atypia occurring in IDC, we design a possible predictive model using integer linear programming (ILP). Computational experiments carried out on a preexisting data set of 13 patients with simultaneous DCIS and IDC show that the corresponding predicted progression models are classifiable into categories having specific evolutionary characteristics. The approach provides new insights into mechanisms of clonal progression in breast cancers and helps illustrate the power of the ILP approach for similar problems in reconstructing tumor evolution scenarios under complex sets of constraints.
Daniele Catanzaro, Stanley Shackney, Alejandro A. Schäffer, Russell Schwartz
IEEE ACM Trans. Comput. Biol. Bioinform.4
2015 Inferring models of multiscale copy number evolution for single-tumor phylogenetics
abstract
MOTIVATION: Phylogenetic algorithms have begun to see widespread use in cancer research to reconstruct processes of evolution in tumor progression. Developing reliable phylogenies for tumor data requires quantitative models of cancer evolution that include the unusual genetic mechanisms by which tumors evolve, such as chromosome abnormalities, and allow for heterogeneity between tumor types and individual patients. Previous work on inferring phylogenies of single tumors by copy number evolution assumed models of uniform rates of genomic gain and loss across different genomic sites and scales, a substantial oversimplification necessitated by a lack of algorithms and quantitative parameters for fitting to more realistic tumor evolution models. RESULTS: We propose a framework for inferring models of tumor progression from single-cell gene copy number data, including variable rates for different gain and loss events. We propose a new algorithm for identification of most parsimonious combinations of single gene and single chromosome events. We extend it via dynamic programming to include genome duplications. We implement an expectation maximization (EM)-like method to estimate mutation-specific and tumor-specific event rates concurrently with tree reconstruction. Application of our algorithms to real cervical cancer data identifies key genomic events in disease progression consistent with prior literature. Classification experiments on cervical and tongue cancer datasets lead to improved prediction accuracy for the metastasis of primary cervical cancers and for tongue cancer survival. AVAILABILITY AND IMPLEMENTATION: Our software (FISHtrees) and two datasets are available at ftp://ftp.ncbi.nlm.nih.gov/pub/FISHtrees.
Salim Akhter Chowdhury, E. Michael Gertz, Darawalee Wangsa, Kerstin Heselmeyer-Haddad, Thomas Ried, Alejandro A. Schäffer, Russell Schwartz
Bioinform.7
2015 A simplicial complex-based approach to unmixing tumor progression data
abstract
BACKGROUND: Tumorigenesis is an evolutionary process by which tumor cells acquire mutations through successive diversification and differentiation. There is much interest in reconstructing this process of evolution due to its relevance to identifying drivers of mutation and predicting future prognosis and drug response. Efforts are challenged by high tumor heterogeneity, though, both within and among patients. In prior work, we showed that this heterogeneity could be turned into an advantage by computationally reconstructing models of cell populations mixed to different degrees in distinct tumors. Such mixed membership model approaches, however, are still limited in their ability to dissect more than a few well-conserved cell populations across a tumor data set. RESULTS: We present a method to improve on current mixed membership model approaches by better accounting for conserved progression pathways between subsets of cancers, which imply a structure to the data that has not previously been exploited. We extend our prior methods, which use an interpretation of the mixture problem as that of reconstructing simple geometric objects called simplices, to instead search for structured unions of simplices called simplicial complexes that one would expect to emerge from mixture processes describing branches along an evolutionary tree. We further improve on the prior work with a novel objective function to better identify mixtures corresponding to parsimonious evolutionary tree models. We demonstrate that this approach improves on our ability to accurately resolve mixtures on simulated data sets and demonstrate its practical applicability on a large RNASeq tumor data set. CONCLUSIONS: Better exploiting the expected geometric structure for mixed membership models produced from common evolutionary trees allows us to quickly and accurately reconstruct models of cell populations sampled from those trees. In the process, we hope to develop a better understanding of tumor evolution as well as other biological problems that involve interpreting genomic data gathered from heterogeneous populations of cells.
Theodore Roman, Amir Nayyeri, Brittany Terese Fasy, Russell Schwartz
BMC Bioinform.4
2014 Editorial
abstract
This special issue of Bioinformatics serves as the proceedings of the 22nd annual meeting of Intelligent Systems for Molecular Biology (ISMB), which took place in Boston, MA, July 11–15, 2014 (http://www.iscb.org/ismbeccb2014). The official conference of the International Society for Computational Biology (http://www.iscb.org/), ISMB, was accompanied by 12 Special Interest Group meetings of one or two days each, two satellite meetings, a High School Teachers Workshop and two half-day tutorials. Since its inception, ISMB has grown to be the largest international conference in computational biology and bioinformatics. It is expected to be the premiere forum in the field for presenting new research results, disseminating methods and techniques and facilitating discussions among leading researchers, practitioners and students in the field. The 37 papers in this volume were selected from 191 submissions divided into 13 research areas, collectively led by 24 Area Chairs. Each area’s chair or chairs selected an expert program committee for their subdiscipline and oversaw the reviewing process for that area. Program committee members themselves could optionally recruit subreviewers to assist in their reviews. By design, the Area Chairs included a mix of experienced individuals reappointed from previous years and experts newly recruited to ensure broad technical expertise and promote inclusivity of various elements of the research community. In total, the review process involved the 24 Area Chairs, 322 program committee members and an additional 131 external reviewers recruited as subreviewers by program committee members. Table 1 provides a summary of the areas, their chairs and the review process by area. Areas, cochairs and acceptance data Areas, cochairs and acceptance data The conference used a two-tier review system, a continuation and refinement of a process begun with ISMB 2013 in an effort to better ensure thorough and fair reviewing. Under the revised process, each of the 191 submissions was first reviewed by at least three expert referees, with a subset receiving between four and eight reviews, as needed. These formal reviews were frequently supplemented by online discussion among reviewers and Area Chairs to resolve points of dispute and reach a consensus on each paper. Among the 191 submissions, 29 were conditionally accepted for publication directly from the first round review based on an assessment of the reviewers that the paper was clearly above par for the conference. A subset of 16 papers were viewed as potentially in the top tier but raised significant questions the reviewers felt might be resolved by the authors in a response. The authors of these 16 papers were invited to submit revisions and responses to the round one criticisms for a second round review by the area chairs and other members of the program committee as needed. Nine of these 16 papers were judged to have addressed the concerns of the reviewers and were conditionally accepted for the conference proceedings, making a total of 38 conditional acceptances. In total, the two-tier review process involved 665 individual reviews. One conditionally accepted paper was subsequently withdrawn based on problems arising post-review, while the remaining 37 were approved for the final conference proceedings and presentation. We believe this two-tier system, more reflective of typical multi-round journal review procedures, provided a fairer process for ensuring only the highest quality original work was accepted within the tight timing constraints imposed by the conference scheduling. We recognize the process is not perfect and some outstanding work might have been rejected despite our best efforts. Nonetheless, we are hopeful that all authors received helpful feedback on their work and that most believed their submissions were judged fairly and expertly, if not always correctly. We are grateful to the Area Chairs, the members of the program committee and the external subreviewers for their outstanding efforts in conducting a thorough review process under tight time constraints. We also thank Steven Leard for his continuing support with the review process; the team at Oxford University Press for preparing this special proceedings volume; Conference Chairs Bonnie Berger and Janet Kelso; and the other members of the ISMB Steering Committee, including Burkhard Rost, Diane Kovats, Paul Horton and Reinhard Schneider, for their advice and supervision. We are also grateful to Nir Ben-Tal, the proceedings chair of ISMB 2013, for sharing his experience and various helpful documents on the review process.
Serafim Batzoglou, Russell Schwartz
Bioinform.2
2014 Algorithms to Model Single Gene, Single Chromosome, and Whole Genome Copy Number Changes Jointly in Tumor Phylogenetics
abstract
We present methods to construct phylogenetic models of tumor progression at the cellular level that include copy number changes at the scale of single genes, entire chromosomes, and the whole genome. The methods are designed for data collected by fluorescence in situ hybridization (FISH), an experimental technique especially well suited to characterizing intratumor heterogeneity using counts of probes to genetic regions frequently gained or lost in tumor development. Here, we develop new provably optimal methods for computing an edit distance between the copy number states of two cells given evolution by copy number changes of single probes, all probes on a chromosome, or all probes in the genome. We then apply this theory to develop a practical heuristic algorithm, implemented in publicly available software, for inferring tumor phylogenies on data from potentially hundreds of single cells by this evolutionary model. We demonstrate and validate the methods on simulated data and published FISH data from cervical cancers and breast cancers. Our computational experiments show that the new model and algorithm lead to more parsimonious trees than prior methods for single-tumor phylogenetics and to improved performance on various classification tasks, such as distinguishing primary tumors from metastases obtained from the same patient population.
Salim Akhter Chowdhury, Stanley Shackney, Kerstin Heselmeyer-Haddad, Thomas Ried, Alejandro A. Schäffer, Russell Schwartz
PLoS Comput. Biol.6
2014 Phenotypic Signatures Arising from Unbalanced Bacterial Growth
abstract
Fluctuations in the growth rate of a bacterial culture during unbalanced growth are generally considered undesirable in quantitative studies of bacterial physiology. Under well-controlled experimental conditions, however, these fluctuations are not random but instead reflect the interplay between intra-cellular networks underlying bacterial growth and the growth environment. Therefore, these fluctuations could be considered quantitative phenotypes of the bacteria under a specific growth condition. Here, we present a method to identify "phenotypic signatures" by time-frequency analysis of unbalanced growth curves measured with high temporal resolution. The signatures are then applied to differentiate amongst different bacterial strains or the same strain under different growth conditions, and to identify the essential architecture of the gene network underlying the observed growth dynamics. Our method has implications for both basic understanding of bacterial physiology and for the classification of bacterial strains.
Cheemeng Tan, Robert P. Smith 0004, Ming-Chi Tsai, Russell Schwartz, Lingchong You
PLoS Comput. Biol.4
2014 Bioinformatics Curriculum Guidelines: Toward a Definition of Core Competencies
abstract
Rapid advances in the life sciences and in related information technologies necessitate the ongoing refinement of bioinformatics educational programs in order to maintain their relevance. As the discipline of bioinformatics and computational biology expands and matures, it is important to characterize the elements that contribute to the success of professionals in this field. These individuals work in a wide variety of settings, including bioinformatics core facilities, biological and medical research laboratories, software development organizations, pharmaceutical and instrument development companies, and institutions that provide education, service, and training. In response to this need, the Curriculum Task Force of the International Society for Computational Biology (ISCB) Education Committee seeks to define curricular guidelines for those who train and educate bioinformaticians. The previous report of the task force summarized a survey that was conducted to gather input regarding the skill set needed by bioinformaticians [1]. The current article details a subsequent effort, wherein the task force broadened its perspectives by examining bioinformatics career opportunities, surveying directors of bioinformatics core facilities, and reviewing bioinformatics education programs.
Lonnie R. Welch, Fran Lewitter, Russell Schwartz, Catherine Brooksbank, Predrag Radivojac, Bruno A. Gaëta, Maria Victoria Schneider
PLoS Comput. Biol.3
2013 Phylogenetic analysis of multiprobe fluorescence in situ hybridization data from tumor cell populations
abstract
MOTIVATION: Development and progression of solid tumors can be attributed to a process of mutations, which typically includes changes in the number of copies of genes or genomic regions. Although comparisons of cells within single tumors show extensive heterogeneity, recurring features of their evolutionary process may be discerned by comparing multiple regions or cells of a tumor. A useful source of data for studying likely progression of individual tumors is fluorescence in situ hybridization (FISH), which allows one to count copy numbers of several genes in hundreds of single cells. Novel algorithms for interpreting such data phylogenetically are needed, however, to reconstruct likely evolutionary trajectories from states of single cells and facilitate analysis of tumor evolution. RESULTS: In this article, we develop phylogenetic methods to infer likely models of tumor progression using FISH copy number data and apply them to a study of FISH data from two cancer types. Statistical analyses of topological characteristics of the tree-based model provide insights into likely tumor progression pathways consistent with the prior literature. Furthermore, tree statistics from the resulting phylogenies can be used as features for prediction methods. This results in improved accuracy, relative to unstructured gene copy number data, at predicting tumor state and future metastasis. AVAILABILITY: Source code for software that does FISH tree building (FISHtrees) and the data on cervical and breast cancer examined here are available at ftp://ftp.ncbi.nlm.nih.gov/pub/FISHtrees. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Salim Akhter Chowdhury, Stanley Shackney, Kerstin Heselmeyer-Haddad, Thomas Ried, Alejandro A. Schäffer, Russell Schwartz
Bioinform.6
2013 Novel Multisample Scheme for Inferring Phylogenetic Markers from Whole Genome Tumor Profiles
abstract
Computational cancer phylogenetics seeks to enumerate the temporal sequences of aberrations in tumor evolution, thereby delineating the evolution of possible tumor progression pathways, molecular subtypes, and mechanisms of action. We previously developed a pipeline for constructing phylogenies describing evolution between major recurring cell types computationally inferred from whole-genome tumor profiles. The accuracy and detail of the phylogenies, however, depend on the identification of accurate, high-resolution molecular markers of progression, i.e., reproducible regions of aberration that robustly differentiate different subtypes and stages of progression. Here, we present a novel hidden Markov model (HMM) scheme for the problem of inferring such phylogenetically significant markers through joint segmentation and calling of multisample tumor data. Our method classifies sets of genome-wide DNA copy number measurements into a partitioning of samples into normal (diploid) or amplified at each probe. It differs from other similar HMM methods in its design specifically for the needs of tumor phylogenetics, by seeking to identify robust markers of progression conserved across a set of copy number profiles. We show an analysis of our method in comparison to other methods on both synthetic and real tumor data, which confirms its effectiveness for tumor phylogeny inference and suggests avenues for future advances.
Ayshwarya Subramanian, Stanley Shackney, Russell Schwartz
IEEE ACM Trans. Comput. Biol. Bioinform.3
2013 Coalescent-Based Method for Learning Parameters of Admixture Events from Large-Scale Genetic Variation Data
abstract
Detecting and quantifying the timing and the genetic contributions of parental populations to a hybrid population is an important but challenging problem in reconstructing evolutionary histories from genetic variation data. With the advent of high throughput genotyping technologies, new methods suitable for large-scale data are especially needed. Furthermore, existing methods typically assume the assignment of individuals into subpopulations is known, when that itself is a difficult problem often unresolved for real data. Here, we propose a novel method that combines prior work for inferring non reticulate population structures with an MCMC scheme for sampling over admixture scenarios to both identify population assignments and learn divergence times and admixture proportions for those populations using genome-scale admixed genetic variation data. We validated our method using coalescent simulations and a collection of real bovine and human variation data. On simulated sequences, our methods show better accuracy and faster run time than leading competitive methods in estimating admixture fractions and divergence times. Analysis on the real data further shows our methods to be effective at matching our best current knowledge about the relevant populations.
Ming-Chi Tsai, Guy E. Blelloch, R. Ravi 0001, Russell Schwartz
IEEE ACM Trans. Comput. Biol. Bioinform.4
2012 Novel Multi-sample Scheme for Inferring Phylogenetic Markers from Whole Genome Tumor Profiles
Ayshwarya Subramanian, Stanley Shackney, Russell Schwartz
ISBRA3
2012 A Report of the Curriculum Task Force of the ISCB Education Committee
abstract
The International Society for Computational Biology (ISCB) Education Committee (EduComm) promotes worldwide education and training in computational biology and bioinformatics and serves as a resource and advisor to organizations interested in developing educational programs. The topic of curricula for bioinformatics programs has long been of interest to ISCB and EduComm. Dr. Russ Altman, a founding board member and past president of ISCB, has been associated with one of the first bioinformatics degree programs (at Stanford University) and wrote an article on this topic [1]. Dr. Shoba Ranganathan, as chair of EduComm a decade ago, began organizing a yearly Workshop on Education in Bioinformatics (WEB) at Intelligent Systems for Molecular Biology (ISMB) meetings that generated exchange of information and many productive discussions. Curriculum development was one aspect of bioinformatics education covered in these sessions [2].
Lonnie R. Welch, Russell Schwartz, Fran Lewitter
PLoS Comput. Biol.2
2011 Phylogenetics of Heterogeneous Samples - (Keynote Talk)
Russell Schwartz
ISBRA1
2011 An Optimization-Based Sampling Scheme for Phylogenetic Trees
Navodit Misra, Guy E. Blelloch, R. Ravi 0001, Russell Schwartz
RECOMB4
2011 Approximate Dynamic Programming using Halfspace Queries and Multiscale Monge Decomposition
abstract
We consider the problem of approximating a signal P with another signal F consisting of a few piecewise constant segments. This problem arises naturally in applications including databases (e.g., histogram construction), speech recognition, computational biology (e.g., denoising aCGH data) and many more. Specifically, let P = (P1, P2, …, Pn), Pi ∊ ℝ for all i, be a signal and let C be a constant. Our goal is to find a function F : [n] → ℝ which optimizes the following objective function: The above optimization problem reduces to solving the following recurrence, which can be done using dynamic programming in O(n2) time: This recurrence arises naturally in several applications where one wants to approximate a given signal P with a signal F which ideally consists of few piecewise constant segments. Such applications include histogram construction in databases, determining DNA copy numbers in cancer cells from micro-array data, speech recognition, data mining and many others. In this work we present two new techniques for optimizing dynamic programming that can handle cost functions not treated by other standard methods. The basis of our first algorithm is the definition of a constant-shifted variant of the objective function that can be efficiently approximated using state of the art methods for range searching. Our technique approximates the optimal value of our objective function within additive ∊ error and runs in time, where δ is an arbitrarily small positive constant and . The second algorithm we provide solves a similar recurrence that's within a multiplicative factor of (1+∊) and runs in O(n log n/∊). The new technique introduced by our algorithm is the decomposition of the initial problem into a small (logarithmic) number of Monge optimization subproblems which we can speed up using existing techniques.
Gary L. Miller, Richard Peng, Russell Schwartz, Charalampos E. Tsourakakis
SODA3
2011 A Consensus Tree Approach for Reconstructing Human Evolutionary History and Detecting Population Substructure
abstract
The random accumulation of variations in the human genome over time implicitly encodes a history of how human populations have arisen, dispersed, and intermixed since we emerged as a species. Reconstructing that history is a challenging computational and statistical problem but has important applications both to basic research and to the discovery of genotype-phenotype correlations. We present a novel approach to inferring human evolutionary history from genetic variation data. We use the idea of consensus trees, a technique generally used to reconcile species trees from divergent gene trees, adapting it to the problem of finding robust relationships within a set of intraspecies phylogenies derived from local regions of the genome. Validation on both simulated and real data shows the method to be effective in recapitulating known true structure of the data closely matching our best current understanding of human evolutionary history. Additional comparison with results of leading methods for the problem of population substructure assignment verifies that our method provides comparable accuracy in identifying meaningful population subgroups in addition to inferring relationships among them. The consensus tree approach thus provides a promising new model for the robust inference of substructure and ancestry from large-scale genetic variation data.
Ming-Chi Tsai, Guy E. Blelloch, R. Ravi 0001, Russell Schwartz
IEEE ACM Trans. Comput. Biol. Bioinform.4
2010 A Consensus Tree Approach for Reconstructing Human Evolutionary History and Detecting Population Substructure
Ming-Chi Tsai, Guy E. Blelloch, R. Ravi 0001, Russell Schwartz
ISBRA4
2010 Generalized Buneman Pruning for Inferring the Most Parsimonious Multi-state Phylogeny
Navodit Misra, Guy E. Blelloch, R. Ravi 0001, Russell Schwartz
RECOMB4
2010 Robust unmixing of tumor states in array comparative genomic hybridization data
abstract
MOTIVATION: Tumorigenesis is an evolutionary process by which tumor cells acquire sequences of mutations leading to increased growth, invasiveness and eventually metastasis. It is hoped that by identifying the common patterns of mutations underlying major cancer sub-types, we can better understand the molecular basis of tumor development and identify new diagnostics and therapeutic targets. This goal has motivated several attempts to apply evolutionary tree reconstruction methods to assays of tumor state. Inference of tumor evolution is in principle aided by the fact that tumors are heterogeneous, retaining remnant populations of different stages along their development along with contaminating healthy cell populations. In practice, though, this heterogeneity complicates interpretation of tumor data because distinct cell types are conflated by common methods for assaying the tumor state. We previously proposed a method to computationally infer cell populations from measures of tumor-wide gene expression through a geometric interpretation of mixture type separation, but this approach deals poorly with noisy and outlier data. RESULTS: In the present work, we propose a new method to perform tumor mixture separation efficiently and robustly to an experimental error. The method builds on the prior geometric approach but uses a novel objective function allowing for robust fits that greatly reduces the sensitivity to noise and outliers. We further develop an efficient gradient optimization method to optimize this 'soft geometric unmixing' objective for measurements of tumor DNA copy numbers assessed by array comparative genomic hybridization (aCGH) data. We show, on a combination of semi-synthetic and real data, that the method yields fast and accurate separation of tumor states. CONCLUSIONS: We have shown a novel objective function and optimization method for the robust separation of tumor sub-types from aCGH data and have shown that the method provides fast, accurate reconstruction of tumor states from mixed samples. Better solutions to this problem can be expected to improve our ability to accurately identify genetic abnormalities in primary tumor samples and to infer patterns of tumor evolution. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
David Tolliver, Charalampos E. Tsourakakis, Ayshwarya Subramanian, Stanley Shackney, Russell Schwartz
Bioinform.5
2010 Applying unmixing to gene expression data for tumor phylogeny inference
abstract
BACKGROUND: While in principle a seemingly infinite variety of combinations of mutations could result in tumor development, in practice it appears that most human cancers fall into a relatively small number of "sub-types," each characterized a roughly equivalent sequence of mutations by which it progresses in different patients. There is currently great interest in identifying the common sub-types and applying them to the development of diagnostics or therapeutics. Phylogenetic methods have shown great promise for inferring common patterns of tumor progression, but suffer from limits of the technologies available for assaying differences between and within tumors. One approach to tumor phylogenetics uses differences between single cells within tumors, gaining valuable information about intra-tumor heterogeneity but allowing only a few markers per cell. An alternative approach uses tissue-wide measures of whole tumors to provide a detailed picture of averaged tumor state but at the cost of losing information about intra-tumor heterogeneity. RESULTS: The present work applies "unmixing" methods, which separate complex data sets into combinations of simpler components, to attempt to gain advantages of both tissue-wide and single-cell approaches to cancer phylogenetics. We develop an unmixing method to infer recurring cell states from microarray measurements of tumor populations and use the inferred mixtures of states in individual tumors to identify possible evolutionary relationships among tumor cells. Validation on simulated data shows the method can accurately separate small numbers of cell states and infer phylogenetic relationships among them. Application to a lung cancer dataset shows that the method can identify cell states corresponding to common lung tumor types and suggest possible evolutionary relationships among them that show good correspondence with our current understanding of lung tumor development. CONCLUSIONS: Unmixing methods provide a way to make use of both intra-tumor heterogeneity and large probe sets for tumor phylogeny inference, establishing a new avenue towards the construction of detailed, accurate portraits of common tumor sub-types and the mechanisms by which they develop. These reconstructions are likely to have future value in discovering and diagnosing novel cancer sub-types and in identifying targets for therapeutic development.
Russell Schwartz, Stanley Shackney
BMC Bioinform.1
2009 Network-Based Inference of Cancer Progression from Microarray Data
abstract
Cancer cells exhibit a common phenotype of uncontrolled cell growth, but this phenotype may arise from many different combinations of mutations. By inferring how cells evolve in individual tumors, a process called cancer progression, we may be able to identify important mutational events for different tumor types, potentially leading to new therapeutics and diagnostics. Prior work has shown that it is possible to infer frequent progression pathways by using gene expression profiles to estimate "distances" between tumors. Here, we apply gene network models to improve these estimates of evolutionary distance by controlling for correlations among coregulated genes. We test three variants of this approach: one using an optimized best-fit network, another using sampling to infer a high-confidence subnetwork, and one using a modular network inferred from clusters of similarly expressed genes. Application to lung cancer and breast cancer microarray data sets shows small improvements in phylogenies when correcting from the optimized network and more substantial improvements when correcting from the sampled or modular networks. Our results suggest that a network correction approach improves estimates of tumor similarity, but sophisticated network models are needed to control for the large hypothesis space and sparse data currently available.
Yongjin Park, Stanley Shackney, Russell Schwartz
IEEE ACM Trans. Comput. Biol. Bioinform.3
2008 Network-Based Inference of Cancer Progression from Microarray Data
Yongjin Park, Stanley Shackney, Russell Schwartz
ISBRA3
2008 Mixed Integer Linear Programming for Maximum-Parsimony Phylogeny Inference
abstract
Reconstruction of phylogenetic trees is a fundamental problem in computational biology. While excellent heuristic methods are available for many variants of this problem, new advances in phylogeny inference will be required if we are to be able to continue to make effective use of the rapidly growing stores of variation data now being gathered. In this paper, we present two integer linear programming (ILP) formulations to find the most parsimonious phylogenetic tree from a set of binary variation data. One method uses a flow-based formulation that can produce exponential numbers of variables and constraints in the worst case. The method has, however, proven extremely efficient in practice on datasets that are well beyond the reach of the available provably efficient methods, solving several large mtDNA and Y-chromosome instances within a few seconds and giving provably optimal results in times competitive with fast heuristics than cannot guarantee optimality. An alternative formulation establishes that the problem can be solved with a polynomial-sized ILP. We further present a web server developed based on the exponential-sized ILP that performs fast maximum parsimony inferences and serves as a front end to a database of precomputed phylogenies spanning the human genome.
Srinath Sridhar 0001, Fumei Lam, Guy E. Blelloch, R. Ravi 0001, Russell Schwartz
IEEE ACM Trans. Comput. Biol. Bioinform.5
2007 Efficiently Finding the Most Parsimonious Phylogenetic Tree Via Linear Programming
Srinath Sridhar 0001, Fumei Lam, Guy E. Blelloch, R. Ravi 0001, Russell Schwartz
ISBRA5
2007 Stochastic Modelling for Systems Biology: Darren J. Wilkinson
abstract
‘Stochastic Modelling for Systems Biology’ was designed to fill an important gap in the educational materials available for students learning about modelling methods for biological systems. Specifically, while stochastic models are emerging as perhaps the preferred method for modelling cellular and subcellular biochemistry in research practice, they remain unfamiliar to most of those who are not specialists in the field. The underlying mathematical and computational methods are well described in the literature of other fields, but the translation to biological practice is largely documented only in the current scientific literature. There are few teaching materials available for these models, particularly for beginning students in biological modelling who lack the background to follow the current scientific literature or the dense mathematical treatments available in texts from other fields. The material in this book arose out of a class the author teaches on stochastic systems biology to master's students in bioinformatics. The text therefore takes a practice-oriented approach to the material, assuming a limited background, focusing on practical considerations in model design and implementation, and making extensive use of example systems. The text provides a solid overview of the basics of stochastic kinetic modelling for the model developer. Chapter 1 introduces the topic by covering some basic concepts and applications of modelling for biology. Chapter 2 describes some representations of biochemical models that are used throughout the rest of the text. Chapters 3 through 5 then provide background material helpful in following the later sections. Chapter 3 covers some basic probability theory and a few important probability distributions, Chapter 4 techniques for sampling from probability distributions in general, and Chapter 5 concepts in Markov models including extensions to continuous time and space. Chapters 6 and 8 provide the bulk of the material specific to stochastic models in biology. Chapter 6 covers the basic theory, models, and methods behind standard Gillespie simulations of reaction chemistry. Chapter 8 provides a more thorough consideration of algorithmic issues in stochastic chemical modelling, including a survey of the leading methods for exact and approximate stochastic kinetic models. In between, Chapter 7 provides four extended examples: dimerization systems, Michaelis–Menten enzyme kinetics, a generic autoregulatory gene network, and the lac operon. Chapters 9 and 10 discuss techniques for model fitting, providing a useful if not exhaustive overview of some common methods. This is indeed a timely addition to the literature and nicely fills the gap Wilkinson identifies in the available teaching materials for biological modelling. The practical, hands-on approach Wilkinson takes will not satisfy all readers. Treatment of background theory is sparse and even non-specialists may find that they need to delve into the suggestions for further reading. But this approach makes the book more useful and accessible to the beginner than a denser but deeper text would be. The book is filled with useful practical advice on the gaps between theoretical models and realistic systems and data sets, as well as techniques for bridging these gaps in practice. There are many pointers to other texts and primary literature that should meet the needs of those requiring greater theoretical depth than this text provides. The text also has a companion website on which the author intends to keep an up-to-date directory of literature and tools for systems biology modelling. Wilkinson's practice-oriented approach is also reflected in the several extended examples presented in the text, which are likely to greatly help the beginner. Formal specifications for these examples are provided in the Systems Biology Markup Language (SBML), which should make it easy for readers to get hands-on practice with the models. There are some specific topics for which broader coverage would nonetheless have been desirable. Stochastic differential equations (SDEs) receive a limited, almost parenthetical treatment in the context of continuous state space Markov models. They are an important alternative to Gillespie-style models for stochastic simulation, though, and warrant a more thorough treatment of algorithms and examples in a text on stochastic models in systems biology. The discussion of parameter fitting is likewise briefer than might have been desirable. It covers several of the standard techniques used by statisticians—e.g. Gibbs sampling, Metropolis–Hastings and other MCMC methods—but omits consideration of the many continuous optimization methods that are also quite important to parameter tuning in practice. There are also some minor design decisions with which one could argue. For instance, the text presents code examples in R, a standard programming language for statisticians but an obscure choice for those from other backgrounds. The author reasonably defends this decision by noting that he must pick some language and that R is public domain and relatively readable. A more widely used language or even a generic pseudo-code might nonetheless have been more accessible. The discussion of many currently popular tools and standards, while a nice complement to the practice-oriented approach of the text, may also date it quickly. This book would be most useful for its originally intended purpose: an introductory graduate-level course on stochastic models for biology. The text is, by itself, a bit sparse for a full-semester graduate course. But it would lend itself well to a project-based class or to supplementation with current literature, much of which Wilkinson references in his suggestions for further reading. Some of the material, particularly the treatment of parameter inference methods, seems likely to demand a firmer background in probability and statistics than the text itself provides. An instructor might therefore do better to require a full introductory class on probability or statistics as a prerequisite and omit Wilkinson's background chapters. The text could also be useful as an independent study aid or reference for scientists with at least basic understanding of reaction systems to introduce themselves to the fundamentals of stochastic kinetic modelling. And although the text does not delve deeply enough for it to be sufficient for modelling specialists, the field is sufficiently young that even the expert is likely to learn something new from browsing it. Aside from a few minor criticisms, this is an excellent text that is likely to find an enthusiastic audience among instructors of systems biology and biological modelling.
Russell Schwartz
Briefings Bioinform.1
2007 Direct maximum parsimony phylogeny reconstruction from genotype data
abstract
BACKGROUND: Maximum parsimony phylogenetic tree reconstruction from genetic variation data is a fundamental problem in computational genetics with many practical applications in population genetics, whole genome analysis, and the search for genetic predictors of disease. Efficient methods are available for reconstruction of maximum parsimony trees from haplotype data, but such data are difficult to determine directly for autosomal DNA. Data more commonly is available in the form of genotypes, which consist of conflated combinations of pairs of haplotypes from homologous chromosomes. Currently, there are no general algorithms for the direct reconstruction of maximum parsimony phylogenies from genotype data. Hence phylogenetic applications for autosomal data must therefore rely on other methods for first computationally inferring haplotypes from genotypes. RESULTS: In this work, we develop the first practical method for computing maximum parsimony phylogenies directly from genotype data. We show that the standard practice of first inferring haplotypes from genotypes and then reconstructing a phylogeny on the haplotypes often substantially overestimates phylogeny size. As an immediate application, our method can be used to determine the minimum number of mutations required to explain a given set of observed genotypes. CONCLUSION: Phylogeny reconstruction directly from unphased data is computationally feasible for moderate-sized problem instances and can lead to substantially more accurate tree size inferences than the standard practice of treating phasing and phylogeny construction as two separate analysis stages. The difference between the approaches is particularly important for downstream applications that require a lower-bound on the number of mutations that the genetic region has undergone.
Srinath Sridhar 0001, Fumei Lam, Guy E. Blelloch, R. Ravi 0001, Russell Schwartz
BMC Bioinform.5
2007 Algorithms for Efficient Near-Perfect Phylogenetic Tree Reconstruction in Theory and Practice
abstract
We consider the problem of reconstructing near-perfect phylogenetic trees using binary character states (referred to as BNPP). A perfect phylogeny assumes that every character mutates at most once in the evolutionary tree, yielding an algorithm for binary character states that is computationally efficient but not robust to imperfections in real data. A near-perfect phylogeny relaxes the perfect phylogeny assumption by allowing at most a constant number of additional mutations. We develop two algorithms for constructing optimal near-perfect phylogenies and provide empirical evidence of their performance. The first simple algorithm is fixed parameter tractable when the number of additional mutations and the number of characters that share four gametes with some other character are constants. The second, more involved algorithm for the problem is fixed parameter tractable when only the number of additional mutations is fixed. We have implemented both algorithms and shown them to be extremely efficient in practice on biologically significant data sets. This work proves the BNPP problem fixed parameter tractable and provides the first practical phylogenetic tree reconstruction algorithms that find guaranteed optimal solutions while being easily implemented and computationally feasible for data sets of biologically meaningful size and complexity.
Srinath Sridhar 0001, Kedar Dhamdhere, Guy E. Blelloch, Eran Halperin, R. Ravi 0001, Russell Schwartz
IEEE ACM Trans. Comput. Biol. Bioinform.6
2006 Fixed Parameter Tractability of Binary Near-Perfect Phylogenetic Tree Reconstruction
Guy E. Blelloch, Kedar Dhamdhere, Eran Halperin, R. Ravi 0001, Russell Schwartz, Srinath Sridhar 0001
ICALP (1)5
2003 Haplotypes and informative SNP selection algorithms: don't block out information
abstract
It is widely hoped that variation in the human genome will provide a means of predicting risk of a variety of complex, chronic diseases. A major stumbling block to the successful identification of association between human DNA polymorphisms (SNPs) and variability in risk of complex diseases is the enormous number of SNPs in the human genome (4,9). The large number of SNPs results in unacceptably high costs for exhaustive genotyping, and so there is a broad effort to determine ways to select SNPs so as to maximize the informativeness of a subset.In this paper we contrast two methods for reducing the complexity of SNP variation: haplotype tagging, i.e. typing a subset of SNPs to identify segments of the genome that appear to be nearly unrecombined (haplotype blocks), and a new block-free model that we develop in this report. We present a statistic for comparing haplotype blocks and show that while the concept of haplotype blocks is reasonably robust there is substantial variability among block partitions. We develop a measure for selecting an informative subset of SNPs in a block free model. We show that the general version of this problem is NP-hard and give efficient algorithms for two important special cases of this problem.
Vineet Bafna, Bjarni V. Halldórsson, Russell Schwartz, Andrew G. Clark, Sorin Istrail
RECOMB3
2002 Methods for Inferring Block-Wise Ancestral History from Haploid Sequences
Russell Schwartz, Andrew G. Clark, Sorin Istrail
WABI1
2002 Algorithmic strategies for the single nucleotide polymorphism haplotype assembly problem
abstract
With the consensus human genome sequenced and many other sequencing projects at varying stages of completion, greater attention is being paid to the genetic differences among individuals and the abilities of those differences to predict phenotypes. A significant obstacle to such work is the difficulty and expense of determining haplotypes--sets of variants genetically linked because of their proximity on the genome--for large numbers of individuals for use in association studies. This paper presents some algorithmic considerations in a new approach for haplotype determination: inferring haplotypes from localised polymorphism data gathered from short genome 'fragments.' Formalised models of the biological system under consideration are examined, given a variety of assumptions about the goal of the problem and the character of optimal solutions. Some theoretical results and algorithms for handling haplotype assembly given the different models are then sketched. The primary conclusion is that some important simplified variants of the problem yield tractable problems while more general variants tend to be intractable in the worst case.
Ross Lippert, Russell Schwartz, Giuseppe Lancia, Sorin Istrail
Briefings Bioinform.2
2001 SNPs Problems, Complexity, and Algorithms
Giuseppe Lancia, Vineet Bafna, Sorin Istrail, Ross Lippert, Russell Schwartz
ESA5
2000 Local rule mechanism for selecting icosahedral shell geometry
Bonnie Berger, Jonathan A. King, Russell Schwartz, Peter W. Shor
Discret. Appl. Math.3