EDBT 2026 Demo / reviewers in the wild / expert
Adam B. Olshen
dblp:53/3321
· DBLP profile ↗
11ranked-venue papers
4as first author
3since 2021 · last 2025
0000-0002-8998-4514ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 11 · 4 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
8 papers |
Bioinformatics and computational biology · 99% Medical and health informatics · 1% |
Topics — the 20 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology › single-cell analysis › single-cell data preprocessing
demultiplexing |
0.9 | 1 | 2025 | SNACS: a tool for demultiplexing single-cell DNA sequencing data · Bioinform. 2025 |
Bioinformatics and computational biology
single-cell analysis |
0.9 | 1 | 2025 | SNACS: a tool for demultiplexing single-cell DNA sequencing data · Bioinform. 2025 |
Bioinformatics and computational biology › single-cell analysis › single-cell genomics
single-cell DNA sequencing |
0.9 | 1 | 2025 | SNACS: a tool for demultiplexing single-cell DNA sequencing data · Bioinform. 2025 |
Bioinformatics and computational biology
cancer genomics |
0.5 | 3 | 2025 | SNACS: a tool for demultiplexing single-cell DNA sequencing data · Bioinform. 2025 Clonality: an R package for testing clonal relatedness of two tumors from the same patient based on their genomic profiles · Bioinform. 2011 Integrative clustering of multiple genomic data types using a joint latent variable model with application to breast and lung cancer subtype analysis · Bioinform. 2009 |
Bioinformatics and computational biology › cancer genomics
tumor heterogeneity |
0.3 | 1 | 2025 | SNACS: a tool for demultiplexing single-cell DNA sequencing data · Bioinform. 2025 |
Bioinformatics and computational biology
gene regulation |
0.2 | 1 | 2013 | Assessing gene-level translational control from ribosome profiling · Bioinform. 2013 |
Bioinformatics and computational biology › transcriptomics
ribosome profiling |
0.2 | 1 | 2013 | Assessing gene-level translational control from ribosome profiling · Bioinform. 2013 |
Bioinformatics and computational biology › gene regulation › post-transcriptional regulation
translational control analysis |
0.2 | 1 | 2013 | Assessing gene-level translational control from ribosome profiling · Bioinform. 2013 |
Bioinformatics and computational biology › cancer genomics
multi-omics clustering |
0.1 | 2 | 2010 | Integrative clustering of multiple genomic data types using a joint latent variable model with application to breast and lung cancer subtype analysis · Bioinform. 2009 Integrative clustering of multiple genomic data types using a joint latent variable model with application to breast and lung cancer subtype analysis · Bioinform. 2010 |
Bioinformatics and computational biology › cancer genomics
copy number analysis |
0.1 | 1 | 2011 | Parent-specific copy number in paired tumor-normal studies using circular binary segmentation · Bioinform. 2011 |
Bioinformatics and computational biology › cancer genomics › chromosomal aberration detection
loss of heterozygosity detection |
0.1 | 1 | 2011 | Parent-specific copy number in paired tumor-normal studies using circular binary segmentation · Bioinform. 2011 |
Bioinformatics and computational biology › cancer genomics
cancer subtype analysis |
0.1 | 1 | 2010 | Integrative clustering of multiple genomic data types using a joint latent variable model with application to breast and lung cancer subtype analysis · Bioinform. 2010 |
Bioinformatics and computational biology › multi-omics data integration
integrative clustering |
0.1 | 1 | 2010 | Integrative clustering of multiple genomic data types using a joint latent variable model with application to breast and lung cancer subtype analysis · Bioinform. 2010 |
Bioinformatics and computational biology › gene expression analysis › microarray data analysis
array CGH analysis |
0.1 | 1 | 2007 | A faster circular binary segmentation algorithm for the analysis of array CGH data · Bioinform. 2007 |
Bioinformatics and computational biology › genomics › computational genomics
copy number segmentation |
0.1 | 1 | 2007 | A faster circular binary segmentation algorithm for the analysis of array CGH data · Bioinform. 2007 |
Medical and health informatics › oncology
cancer diagnosis |
0.0 | 1 | 2011 | Clonality: an R package for testing clonal relatedness of two tumors from the same patient based on their genomic profiles · Bioinform. 2011 |
Bioinformatics and computational biology › gene expression analysis › differential expression analysis
differentially expressed gene identification |
0.0 | 1 | 2002 | Deriving quantitative conclusions from microarray expression data · Bioinform. 2002 |
Bioinformatics and computational biology
gene expression analysis |
0.0 | 1 | 2002 | Deriving quantitative conclusions from microarray expression data · Bioinform. 2002 |
Bioinformatics and computational biology › gene expression analysis
sample classification |
0.0 | 1 | 2002 | Deriving quantitative conclusions from microarray expression data · Bioinform. 2002 |
Bioinformatics and computational biology › multi-omics data integration
genomic data integration |
0.0 | 1 | 2010 | Integrative clustering of multiple genomic data types using a joint latent variable model with application to breast and lung cancer subtype analysis · Bioinform. 2010 |
Methods — techniques the papers use, named apart from their topics
antibody-based cell sorting · 0.9SNP-based demultiplexing · 0.9expectation-maximization · 0.2parametric bootstrap · 0.2negative binomial · 0.2errors-in-variables regression · 0.2significance testing · 0.1loss of heterozygosity analysis · 0.1likelihood-ratio statistic · 0.1circular binary segmentation · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SNACS: a tool for demultiplexing single-cell DNA sequencing dataabstractMOTIVATION: Single-cell DNA sequencing (scDNA-seq) and multi-modal profiling with the addition of cell-surface antibodies (scDAb-seq) have recently provided key insights into cancer heterogeneity. Scaling these technologies across large patient cohorts, however, is cost and time prohibitive. Multiplexing, in which cells from unique patients are pooled into a single experiment, offers a possible solution. While multiplexing methods exist for scRNAseq, accurate demultiplexing in scDNAseq remains an unmet need. RESULTS: Here, we introduce SNACS: single-nucleotide polymorphism and antibody-based cell sorting. SNACS relies on a combination of patient-level cell-surface identifiers and natural variation in genetic polymorphisms to demultiplex scDNAseq data. We demonstrated the performance of SNACS on a dataset consisting of multi-sample experiments from patients with leukemia where we knew truth from single-sample experiments from the same patients. Using SNACS, accuracy ranged from 0.948 to 0.991 versus 0.552 to 0.934 using demultiplexing methods from the single-cell literature. AVAILABILITY AND IMPLEMENTATION: SNACS is available at https://github.com/olshena/SNACS. Vanessa E. Kennedy, Ritu Roy, Cheryl A. C. Peretz, Andrew Koh, Elaine Tran, Catherine C. Smith, Adam B. Olshen |
Bioinform. | 7 |
| 2024 | Torch-eCpG: a fast and scalable eQTM mapper for thousands of molecular phenotypes with graphical processing unitsabstractBACKGROUND: Gene expression may be regulated by the DNA methylation of regulatory elements in cis, distal, and trans regions. One method to evaluate the relationship between DNA methylation and gene expression is the mapping of expression quantitative trait methylation (eQTM) loci (also called expression associated CpG loci, eCpG). However, no open-source tools are available to provide eQTM mapping. In addition, eQTM mapping can involve a large number of comparisons which may prevent the analyses due to limitations of computational resources. Here, we describe Torch-eCpG, an open-source tool to perform eQTM mapping that includes an optimized implementation that can use the graphical processing unit (GPU) to reduce runtime. RESULTS: We demonstrate the analyses using the tool are reproducible, up to 18 × faster using the GPU, and scale linearly with increasing methylation loci. CONCLUSIONS: Torch-eCpG is a fast, reliable, and scalable tool to perform eQTM mapping. Source code for Torch-eCpG is available at https://github.com/kordk/torch-ecpg . Kord M. Kober, Liam Berger, Ritu Roy, Adam B. Olshen |
BMC Bioinform. | 4 |
| 2023 | Does multi-way, long-range chromatin contact data advance 3D genome reconstruction?abstractBACKGROUND: Methods for inferring the three-dimensional (3D) configuration of chromatin from conformation capture assays that provide strictly pairwise interactions, notably Hi-C, utilize the attendant contact matrix as input. More recent assays, in particular split-pool recognition of interactions by tag extension (SPRITE), capture multi-way interactions instead of solely pairwise contacts. These assays yield contacts that straddle appreciably greater genomic distances than Hi-C, in addition to instances of exceptionally high-order chromatin interaction. Such attributes are anticipated to be consequential with respect to 3D genome reconstruction, a task yet to be undertaken with multi-way contact data. However, performing such 3D reconstruction using distance-based reconstruction techniques requires framing multi-way contacts as (pairwise) distances. Comparing approaches for so doing, and assessing the resultant impact of long-range and multi-way contacts, are the objectives of this study. RESULTS: We obtained 3D reconstructions via multi-dimensional scaling under a variety of weighting schemes for mapping SPRITE multi-way contacts to pairwise distances. Resultant configurations were compared following Procrustes alignment and relationships were assessed between associated Procrustes root mean square errors and key features such as the extent of multi-way and/or long-range contacts. We found that these features had surprisingly limited influence on 3D reconstruction, a finding we attribute to their influence being diminished by the preponderance of pairwise contacts. CONCLUSION: Distance-based 3D genome reconstruction using SPRITE multi-way contact data is not appreciably affected by the weighting scheme used to convert multi-way interactions to pairwise distances. Adam B. Olshen, Mark R. Segal |
BMC Bioinform. | 1 |
| 2013 | Assessing gene-level translational control from ribosome profilingabstractMOTIVATION: The translational landscape of diverse cellular systems remains largely uncharacterized. A detailed understanding of the control of gene expression at the level of messenger RNA translation is vital to elucidating a systems-level view of complex molecular programs in the cell. Establishing the degree to which such post-transcriptional regulation can mediate specific phenotypes is similarly critical to elucidating the molecular pathogenesis of diseases such as cancer. Recently, methods for massively parallel sequencing of ribosome-bound fragments of messenger RNA have begun to uncover genome-wide translational control at codon resolution. Despite its promise for deeply characterizing mammalian proteomes, few analytical methods exist for the comprehensive analysis of this paired RNA and ribosome data. RESULTS: We describe the Babel framework, an analytical methodology for assessing the significance of changes in translational regulation within cells and between conditions. This approach facilitates the analysis of translation genome-wide while allowing statistically principled gene-level inference. Babel is based on an errors-in-variables regression model that uses the negative binomial distribution and draws inference using a parametric bootstrap approach. We demonstrate the operating characteristics of Babel on simulated data and use its gene-level inference to extend prior analyses significantly, discovering new translationally regulated modules under mammalian target of rapamycin (mTOR) pathway signaling control. Adam B. Olshen, Andrew C. Hsieh, Craig R. Stumpf, Richard A. Olshen, Davide Ruggero, Barry S. Taylor |
Bioinform. | 1 |
| 2011 | Parent-specific copy number in paired tumor-normal studies using circular binary segmentationabstractMOTIVATION: High-throughput techniques facilitate the simultaneous measurement of DNA copy number at hundreds of thousands of sites on a genome. Older techniques allow measurement only of total copy number, the sum of the copy number contributions from the two parental chromosomes. Newer single nucleotide polymorphism (SNP) techniques can in addition enable quantifying parent-specific copy number (PSCN). The raw data from such experiments are two-dimensional, but are unphased. Consequently, inference based on them necessitates development of new analytic methods. METHODS: We have adapted and enhanced the circular binary segmentation (CBS) algorithm for this purpose with focus on paired test and reference samples. The essence of paired parent-specific CBS (Paired PSCBS) is to utilize the original CBS algorithm to identify regions of equal total copy number and then to further segment these regions where there have been changes in PSCN. For the final set of regions, calls are made of equal parental copy number and loss of heterozygosity (LOH). PSCN estimates are computed both before and after calling. RESULTS: The methodology is evaluated by simulation and on glioblastoma data. In the simulation, PSCBS compares favorably to established methods. On the glioblastoma data, PSCBS identifies interesting genomic regions, such as copy-neutral LOH. AVAILABILITY: The Paired PSCBS method is implemented in an open-source R package named PSCBS, available on CRAN (http://cran.r-project.org/). Adam B. Olshen, Henrik Bengtsson, Pierre Neuvial, Paul T. Spellman, Richard A. Olshen, Venkatraman E. Seshan |
Bioinform. | 1 |
| 2011 | Clonality: an R package for testing clonal relatedness of two tumors from the same patient based on their genomic profilesabstractSUMMARY: If a cancer patient develops multiple tumors, it is sometimes impossible to determine whether these tumors are independent or clonal based solely on pathological characteristics. Investigators have studied how to improve this diagnostic challenge by comparing the presence of loss of heterozygosity (LOH) at selected genetic locations of tumor samples, or by comparing genomewide copy number array profiles. We have previously developed statistical methodology to compare such genomic profiles for an evidence of clonality. We assembled the software for these tests in a new R package called 'Clonality'. For LOH profiles, the package contains significance tests. The analysis of copy number profiles includes a likelihood ratio statistic and reference distribution, as well as an option to produce various plots that summarize the results. AVAILABILITY: Bioconductor (http://bioconductor.org/packages/release/bioc/html/Clonality.html) and http://www.mskcc.org/mskcc/html/13287.cfm. Irina Ostrovnaya, Venkatraman E. Seshan, Adam B. Olshen, Colin B. Begg |
Bioinform. | 3 |
| 2010 | Integrative clustering of multiple genomic data types using a joint latent variable model with application to breast and lung cancer subtype analysisabstractMOTIVATION: The molecular complexity of a tumor manifests itself at the genomic, epigenomic, transcriptomic and proteomic levels. Genomic profiling at these multiple levels should allow an integrated characterization of tumor etiology. However, there is a shortage of effective statistical and bioinformatic tools for truly integrative data analysis. The standard approach to integrative clustering is separate clustering followed by manual integration. A more statistically powerful approach would incorporate all data types simultaneously and generate a single integrated cluster assignment. METHODS: We developed a joint latent variable model for integrative clustering. We call the resulting methodology iCluster. iCluster incorporates flexible modeling of the associations between different data types and the variance-covariance structure within data types in a single framework, while simultaneously reducing the dimensionality of the datasets. Likelihood-based inference is obtained through the Expectation-Maximization algorithm. RESULTS: We demonstrate the iCluster algorithm using two examples of joint analysis of copy number and gene expression data, one from breast cancer and one from lung cancer. In both cases, we identified subtypes characterized by concordant DNA copy number changes and gene expression as well as unique profiles specific to one or the other in a completely automated fashion. In addition, the algorithm discovers potentially novel subtypes by combining weak yet consistent alteration patterns across data types. AVAILABILITY: R code to implement iCluster can be downloaded at http://www.mskcc.org/mskcc/html/85130.cfm Ronglai Shen, Adam B. Olshen, Marc Ladanyi |
Bioinform. | 2 |
| 2010 | A classification model for distinguishing copy number variants from cancer-related alterationsabstractBACKGROUND: Both somatic copy number alterations (CNAs) and germline copy number variants (CNVs) that are prevalent in healthy individuals can appear as recurrent changes in comparative genomic hybridization (CGH) analyses of tumors. In order to identify important cancer genes CNAs and CNVs must be distinguished. Although the Database of Genomic Variants (DGV) contains a list of all known CNVs, there is no standard methodology to use the database effectively. RESULTS: We develop a prediction model that distinguishes CNVs from CNAs based on the information contained in the DGV and several other variables, including segment's length, height, closeness to a telomere or centromere and occurrence in other patients. The models are fitted on data from glioblastoma and their corresponding normal samples that were collected as part of The Cancer Genome Atlas project and hybridized to Agilent 244 K arrays. CONCLUSIONS: Using the DGV alone CNVs in the test set can be correctly identified with about 85% accuracy if the outliers are removed before segmentation and with 72% accuracy if the outliers are included, and additional variables improve the prediction by about 2-3% and 12%, respectively. Final models applied to data from ovarian tumors have about 90% accuracy with all the variables and 86% accuracy with the DGV alone. Irina Ostrovnaya, Gouri Nanjangud, Adam B. Olshen |
BMC Bioinform. | 3 |
| 2009 | Integrative clustering of multiple genomic data types using a joint latent variable model with application to breast and lung cancer subtype analysisabstractAbstract Motivation: The molecular complexity of a tumor manifests itself at the genomic, epigenomic, transcriptomic and proteomic levels. Genomic profiling at these multiple levels should allow an integrated characterization of tumor etiology. However, there is a shortage of effective statistical and bioinformatic tools for truly integrative data analysis. The standard approach to integrative clustering is separate clustering followed by manual integration. A more statistically powerful approach would incorporate all data types simultaneously and generate a single integrated cluster assignment. Methods: We developed a joint latent variable model for integrative clustering. We call the resulting methodology iCluster. iCluster incorporates flexible modeling of the associations between different data types and the variance–covariance structure within data types in a single framework, while simultaneously reducing the dimensionality of the datasets. Likelihood-based inference is obtained through the Expectation–Maximization algorithm. Results: We demonstrate the iCluster algorithm using two examples of joint analysis of copy number and gene expression data, one from breast cancer and one from lung cancer. In both cases, we identified subtypes characterized by concordant DNA copy number changes and gene expression as well as unique profiles specific to one or the other in a completely automated fashion. In addition, the algorithm discovers potentially novel subtypes by combining weak yet consistent alteration patterns across data types. Availability: R code to implement iCluster can be downloaded at http://www.mskcc.org/mskcc/html/85130.cfm Contact: [email protected] Supplementary information: Supplementary data are available at Bioinformatics online. Ronglai Shen, Adam B. Olshen, Marc Ladanyi |
Bioinform. | 2 |
| 2007 | A faster circular binary segmentation algorithm for the analysis of array CGH dataabstractMOTIVATION: Array CGH technologies enable the simultaneous measurement of DNA copy number for thousands of sites on a genome. We developed the circular binary segmentation (CBS) algorithm to divide the genome into regions of equal copy number. The algorithm tests for change-points using a maximal t-statistic with a permutation reference distribution to obtain the corresponding P-value. The number of computations required for the maximal test statistic is O(N2), where N is the number of markers. This makes the full permutation approach computationally prohibitive for the newer arrays that contain tens of thousands markers and highlights the need for a faster algorithm. RESULTS: We present a hybrid approach to obtain the P-value of the test statistic in linear time. We also introduce a rule for stopping early when there is strong evidence for the presence of a change. We show through simulations that the hybrid approach provides a substantial gain in speed with only a negligible loss in accuracy and that the stopping rule further increases speed. We also present the analyses of array CGH data from breast cancer cell lines to show the impact of the new approaches on the analysis of real data. AVAILABILITY: An R version of the CBS algorithm has been implemented in the "DNAcopy" package of the Bioconductor project. The proposed hybrid method for the P-value is available in version 1.2.1 or higher and the stopping rule for declaring a change early is available in version 1.5.1 or higher. E. S. Venkatraman, Adam B. Olshen |
Bioinform. | 2 |
| 2002 | Deriving quantitative conclusions from microarray expression dataabstractMOTIVATION: The last few years have seen the development of DNA microarray technology that allows simultaneous measurement of the expression levels of thousands of genes. While many methods have been developed to analyze such data, most have been visualization-based. Methods that yield quantitative conclusions have been diverse and complex. RESULTS: We present two straightforward methods for identifying specific genes whose expression is linked with a phenotype or outcome variable as well as for systematically predicting sample class membership: (1) a conservative, permutation-based approach to identifying differentially expressed genes; (2) an augmentation of K-nearest-neighbor pattern classification. Our analyses replicate the quantitative conclusions of Golub et al. (1999; Science, 286, 531-537) on leukemia data, with better classification results, using far simpler methods. With the breast tumor data of Perou et al. (2000; Nature, 406, 747-752), the methods lend rigorous quantitative support to the conclusions of the original paper. In the case of the lymphoma data in Alizadeh et al. (2000; Nature, 403, 503-511), our analyses only partially support the conclusions of the original authors. AVAILABILITY: The software and supplementary information are available freely to researchers at academic and non-profit institutions at http://cc.ucsf.edu/jain/public Adam B. Olshen, Ajay N. Jain |
Bioinform. | 1 |