EDBT 2026 Demo / reviewers in the wild / expert
Mark D. M. Leiserson
dblp:23/6320
· DBLP profile ↗
15ranked-venue papers
5as first author
2since 2021 · last 2023
0000-0002-1034-4363ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 14 · 5 first-author · 2 since 2021Artificial intelligence and machine learning · 1Human-computer interaction and ubiquitous computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
10 papers |
Bioinformatics and computational biology · 79% Computational science and engineering · 21% | |
| Databases, data mining, and information retrieval
1 paper |
Data mining · 100% |
Topics — the 17 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology
cancer genomics |
1.6 | 5 | 2019 | Modeling clinical and molecular covariates of mutational process activity in cancer · Bioinform. 2019 A Sticky Multinomial Mixture Model of Strand-Coordinated Mutational Processes in Cancer · RECOMB 2019 Hierarchical HotNet: identifying hierarchies of altered subnetworks · Bioinform. 2018 |
Computational science and engineering › latent variable model
mixture model |
0.8 | 2 | 2020 | A Mixture Model for Signature Discovery from Sparse Mutation Data · RECOMB 2020 A Sticky Multinomial Mixture Model of Strand-Coordinated Mutational Processes in Cancer · RECOMB 2019 |
Computational science and engineering › numerical linear algebra
matrix factorization |
0.4 | 1 | 2020 | Matrix (factorization) reloaded: flexible methods for imputing genetic interactions with cross-species and side information · Bioinform. 2020 |
Bioinformatics and computational biology › cancer genomics › mutational signature analysis
mutation signature discovery |
0.4 | 1 | 2020 | A Mixture Model for Signature Discovery from Sparse Mutation Data · RECOMB 2020 |
Bioinformatics and computational biology › cancer genomics
mutational signature analysis |
0.4 | 1 | 2019 | Modeling clinical and molecular covariates of mutational process activity in cancer · Bioinform. 2019 |
Bioinformatics and computational biology › biostatistics › statistical bioinformatics
statistical genomics |
0.4 | 1 | 2019 | A Sticky Multinomial Mixture Model of Strand-Coordinated Mutational Processes in Cancer · RECOMB 2019 |
Bioinformatics and computational biology › cancer genomics
cancer gene prediction |
0.3 | 1 | 2018 | Hierarchical HotNet: identifying hierarchies of altered subnetworks · Bioinform. 2018 |
Bioinformatics and computational biology
functional genomics |
0.3 | 1 | 2018 | A Multi-species Functional Embedding Integrating Sequence and Network Structure · RECOMB 2018 |
Bioinformatics and computational biology › network bioinformatics › biological network analysis
network analysis |
0.3 | 1 | 2018 | Hierarchical HotNet: identifying hierarchies of altered subnetworks · Bioinform. 2018 |
Bioinformatics and computational biology
statistical genetics |
0.2 | 1 | 2016 | A weighted exact test for mutually exclusive mutations in cancer · Bioinform. 2016 |
Bioinformatics and computational biology › comparative genomics
cross-species data integration |
0.1 | 1 | 2020 | Matrix (factorization) reloaded: flexible methods for imputing genetic interactions with cross-species and side information · Bioinform. 2020 |
Bioinformatics and computational biology › statistical genetics › gene-gene interaction
genetic interaction analysis |
0.1 | 1 | 2011 | Inferring Mechanisms of Compensation from E-MAP and SGA Data Using Local Search Algorithms for Max Cut · RECOMB 2011 |
Bioinformatics and computational biology
systems biology |
0.1 | 1 | 2011 | Inferring Mechanisms of Compensation from E-MAP and SGA Data Using Local Search Algorithms for Max Cut · RECOMB 2011 |
Bioinformatics and computational biology › biological network
network biology |
0.1 | 1 | 2018 | A Multi-species Functional Embedding Integrating Sequence and Network Structure · RECOMB 2018 |
Bioinformatics and computational biology › systems bioinformatics
pathway analysis |
0.1 | 1 | 2009 | Evaluating Between-Pathway Models with Expression Data · RECOMB 2009 |
Data mining
pattern mining |
0.1 | 1 | 2015 | CoMEt: A Statistical Approach to Identify Combinations of Mutually Exclusive Alterations in Cancer · RECOMB 2015 |
Bioinformatics and computational biology › statistical genetics › gene-gene interaction
epistasis |
0.0 | 1 | 2011 | Inferring Mechanisms of Compensation from E-MAP and SGA Data Using Local Search Algorithms for Max Cut · RECOMB 2011 |
Methods — techniques the papers use, named apart from their topics
mixture model · 0.4kernel side information · 0.4extensible matrix factorization · 0.4non-negative matrix factorization · 0.4multinomial mixture model · 0.4markov model · 0.4bayesian topic modeling · 0.4sequence embedding · 0.3permutation testing · 0.3network embedding · 0.3statistical methods · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | A mutation-level covariate model for mutational signaturesabstractMutational processes and their exposures in particular genomes are key to our understanding of how these genomes are shaped. However, current analyses assume that these processes are uniformly active across the genome without accounting for potential covariates such as strand or genomic region that could impact such activities. Here we suggest the first mutation-covariate models that explicitly model the effect of different covariates on the exposures of mutational processes. We apply these models to test the impact of replication strand on these processes and compare them to strand-oblivious models across a range of data sets. Our models capture replication strand specificity, point to signatures affected by it, and score better on held-out data compared to standard models that do not account for mutation-level covariate information. Itay Kahane, Mark D. M. Leiserson, Roded Sharan |
PLoS Comput. Biol. | 2 |
| 2021 | A data-driven approach for constructing mutation categories for mutational signature analysisabstractMutational processes shape the genomes of cancer patients and their understanding has important applications in diagnosis and treatment. Current modeling of mutational processes by identifying their characteristic signatures views each base substitution in a limited context of a single flanking base on each side. This context definition gives rise to 96 categories of mutations that have become the standard in the field, even though wider contexts have been shown to be informative in specific cases. Here we propose a data-driven approach for constructing a mutation categorization for mutational signature analysis. Our approach is based on the assumption that tumor cells that are exposed to similar mutational processes, show similar expression levels of DNA damage repair genes that are involved in these processes. We attempt to find a categorization that maximizes the agreement between mutation and gene expression data, and show that it outperforms the standard categorization over multiple quality measures. Moreover, we show that the categorization we identify generalizes to unseen data from different cancer types, suggesting that mutation context patterns extend beyond the immediate flanking bases. Gal Gilad, Mark D. M. Leiserson, Roded Sharan |
PLoS Comput. Biol. | 2 |
| 2020 | A Mixture Model for Signature Discovery from Sparse Mutation Data
Itay Sason, Yuexi Chen, Mark D. M. Leiserson, Roded Sharan |
RECOMB | 3 |
| 2020 | Matrix (factorization) reloaded: flexible methods for imputing genetic interactions with cross-species and side informationabstractMOTIVATION: Mapping genetic interactions (GIs) can reveal important insights into cellular function and has potential translational applications. There has been great progress in developing high-throughput experimental systems for measuring GIs (e.g. with double knockouts) as well as in defining computational methods for inferring (imputing) unknown interactions. However, existing computational methods for imputation have largely been developed for and applied in baker's yeast, even as experimental systems have begun to allow measurements in other contexts. Importantly, existing methods face a number of limitations in requiring specific side information and with respect to computational cost. Further, few have addressed how GIs can be imputed when data are scarce. RESULTS: In this article, we address these limitations by presenting a new imputation framework, called Extensible Matrix Factorization (EMF). EMF is a framework of composable models that flexibly exploit cross-species information in the form of GI data across multiple species, and arbitrary side information in the form of kernels (e.g. from protein-protein interaction networks). We perform a rigorous set of experiments on these models in matched GI datasets from baker's and fission yeast. These include the first such experiments on genome-scale GI datasets in multiple species in the same study. We find that EMF models that exploit side and cross-species information improve imputation, especially in data-scarce settings. Further, we show that EMF outperforms the state-of-the-art deep learning method, even when using strictly less data, and incurs orders of magnitude less computational cost. AVAILABILITY: Implementations of models and experiments are available at: https://github.com/lrgr/EMF. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jason Fan, Xuan Cindy Li, Mark Crovella, Mark D. M. Leiserson |
Bioinform. | 4 |
| 2019 | What are the Biases in My Word Embedding?abstractThis paper presents an algorithm for enumerating biases in word embeddings. The algorithm exposes a large number of offensive associations related to sensitive features such as race and gender on publicly available embeddings, including a supposedly "debiased" embedding. These biases are concerning in light of the widespread use of word embeddings. The associations are identified by geometric patterns in word embeddings that run parallel between people's names and common lower-case tokens. The algorithm is highly unsupervised: it does not even require the sensitive features to be pre-specified. This is desirable because: (a) many forms of discrimination?such as racial discrimination-are linked to social constructs that may vary depending on the context, rather than to categories with fixed definitions; and (b) it makes it easier to identify biases against intersectional groups, which depend on combinations of sensitive features. The inputs to our algorithm are a list of target tokens, e.g. names, and a word embedding. It outputs a number of Word Embedding Association Tests (WEATs) that capture various biases present in the data. We illustrate the utility of our approach on publicly available word embeddings and lists of names, and evaluate its output using crowdsourcing. We also show how removing names may not remove potential proxy bias. Nathaniel Swinger, Maria De-Arteaga, Neil Thomas Heffernan IV, Mark D. M. Leiserson, Adam Tauman Kalai |
AIES | 4 |
| 2019 | A Sticky Multinomial Mixture Model of Strand-Coordinated Mutational Processes in Cancer
Itay Sason, Damian Wójtowicz, Welles Robinson, Mark D. M. Leiserson, Teresa M. Przytycka, Roded Sharan |
RECOMB | 4 |
| 2019 | Modeling clinical and molecular covariates of mutational process activity in cancerabstractMOTIVATION: Somatic mutations result from processes related to DNA replication or environmental/lifestyle exposures. Knowing the activity of mutational processes in a tumor can inform personalized therapies, early detection, and understanding of tumorigenesis. Computational methods have revealed 30 validated signatures of mutational processes active in human cancers, where each signature is a pattern of single base substitutions. However, half of these signatures have no known etiology, and some similar signatures have distinct etiologies, making patterns of mutation signature activity hard to interpret. Existing mutation signature detection methods do not consider tumor-level clinical/demographic (e.g. smoking history) or molecular features (e.g. inactivations to DNA damage repair genes). RESULTS: To begin to address these challenges, we present the Tumor Covariate Signature Model (TCSM), the first method to directly model the effect of observed tumor-level covariates on mutation signatures. To this end, our model uses methods from Bayesian topic modeling to change the prior distribution on signature exposure conditioned on a tumor's observed covariates. We also introduce methods for imputing covariates in held-out data and for evaluating the statistical significance of signature-covariate associations. On simulated and real data, we find that TCSM outperforms both non-negative matrix factorization and topic modeling-based approaches, particularly in recovering the ground truth exposure to similar signatures. We then use TCSM to discover five mutation signatures in breast cancer and predict homologous recombination repair deficiency in held-out tumors. We also discover four signatures in a combined melanoma and lung cancer cohort-using cancer type as a covariate-and provide statistical evidence to support earlier claims that three lung cancers from The Cancer Genome Atlas are misdiagnosed metastatic melanomas. AVAILABILITY AND IMPLEMENTATION: TCSM is implemented in Python 3 and available at https://github.com/lrgr/tcsm, along with a data workflow for reproducing the experiments in the paper. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Welles Robinson, Roded Sharan, Mark D. M. Leiserson |
Bioinform. | 3 |
| 2018 | A Multi-species Functional Embedding Integrating Sequence and Network Structure
Mark D. M. Leiserson, Jason Fan, Anthony Cannistra, Inbar Fried, Tim Lim, Thomas Schaffner, Mark Crovella, Benjamin Hescott |
RECOMB | 1 |
| 2018 | Hierarchical HotNet: identifying hierarchies of altered subnetworksabstractMotivation: The analysis of high-dimensional 'omics data is often informed by the use of biological interaction networks. For example, protein-protein interaction networks have been used to analyze gene expression data, to prioritize germline variants, and to identify somatic driver mutations in cancer. In these and other applications, the underlying computational problem is to identify altered subnetworks containing genes that are both highly altered in an 'omics dataset and are topologically close (e.g. connected) on an interaction network. Results: We introduce Hierarchical HotNet, an algorithm that finds a hierarchy of altered subnetworks. Hierarchical HotNet assesses the statistical significance of the resulting subnetworks over a range of biological scales and explicitly controls for ascertainment bias in the network. We evaluate the performance of Hierarchical HotNet and several other algorithms that identify altered subnetworks on the problem of predicting cancer genes and significantly mutated subnetworks. On somatic mutation data from The Cancer Genome Atlas, Hierarchical HotNet outperforms other methods and identifies significantly mutated subnetworks containing both well-known cancer genes and candidate cancer genes that are rarely mutated in the cohort. Hierarchical HotNet is a robust algorithm for identifying altered subnetworks across different 'omics datasets. Availability and implementation: http://github.com/raphael-group/hierarchical-hotnet. Supplementary information: Supplementary material are available at Bioinformatics online. Matthew A. Reyna, Mark D. M. Leiserson, Benjamin J. Raphael |
Bioinform. | 2 |
| 2016 | A weighted exact test for mutually exclusive mutations in cancerabstractMOTIVATION: The somatic mutations in the pathways that drive cancer development tend to be mutually exclusive across tumors, providing a signal for distinguishing driver mutations from a larger number of random passenger mutations. This mutual exclusivity signal can be confounded by high and highly variable mutation rates across a cohort of samples. Current statistical tests for exclusivity that incorporate both per-gene and per-sample mutational frequencies are computationally expensive and have limited precision. RESULTS: We formulate a weighted exact test for assessing the significance of mutual exclusivity in an arbitrary number of mutational events. Our test conditions on the number of samples with a mutation as well as per-event, per-sample mutation probabilities. We provide a recursive formula to compute P-values for the weighted test exactly as well as a highly accurate and efficient saddlepoint approximation of the test. We use our test to approximate a commonly used permutation test for exclusivity that conditions on per-event, per-sample mutation frequencies. However, our test is more efficient and it recovers more significant results than the permutation test. We use our Weighted Exclusivity Test (WExT) software to analyze hundreds of colorectal and endometrial samples from The Cancer Genome Atlas, which are two cancer types that often have extremely high mutation rates. On both cancer types, the weighted test identifies sets of mutually exclusive mutations in cancer genes with fewer false positives than earlier approaches. AVAILABILITY AND IMPLEMENTATION: See http://compbio.cs.brown.edu/projects/wext for software. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Mark D. M. Leiserson, Matthew A. Reyna, Benjamin J. Raphael |
Bioinform. | 1 |
| 2015 | CoMEt: A Statistical Approach to Identify Combinations of Mutually Exclusive Alterations in Cancer
Mark D. M. Leiserson, Hsin-Ta Wu, Fabio Vandin, Benjamin J. Raphael |
RECOMB | 1 |
| 2013 | Genecentric: a package to uncover graph-theoretic structure in high-throughput epistasis dataabstractBACKGROUND: New technology has resulted in high-throughput screens for pairwise genetic interactions in yeast and other model organisms. For each pair in a collection of non-essential genes, an epistasis score is obtained, representing how much sicker (or healthier) the double-knockout organism will be compared to what would be expected from the sickness of the component single knockouts. Recent algorithmic work has identified graph-theoretic patterns in this data that can indicate functional modules, and even sets of genes that may occur in compensatory pathways, such as a BPM-type schema first introduced by Kelley and Ideker. However, to date, any algorithms for finding such patterns in the data were implemented internally, with no software being made publically available. RESULTS: Genecentric is a new package that implements a parallelized version of the Leiserson et al. algorithm (J Comput Biol 18:1399-1409, 2011) for generating generalized BPMs from high-throughput genetic interaction data. Given a matrix of weighted epistasis values for a set of double knock-outs, Genecentric returns a list of generalized BPMs that may represent compensatory pathways. Genecentric also has an extension, GenecentricGO, to query FuncAssociate (Bioinformatics 25:3043-3044, 2009) to retrieve GO enrichment statistics on generated BPMs. Python is the only dependency, and our web site provides working examples and documentation. CONCLUSION: We find that Genecentric can be used to find coherent functional and perhaps compensatory gene sets from high throughput genetic interaction data. Genecentric is made freely available for download under the GPLv2 from http://bcb.cs.tufts.edu/genecentric. Andrew Gallant, Mark D. M. Leiserson, Maxim Kachalov, Lenore Cowen, Benjamin Hescott |
BMC Bioinform. | 2 |
| 2013 | Simultaneous Identification of Multiple Driver Pathways in CancerabstractDistinguishing the somatic mutations responsible for cancer (driver mutations) from random, passenger mutations is a key challenge in cancer genomics. Driver mutations generally target cellular signaling and regulatory pathways consisting of multiple genes. This heterogeneity complicates the identification of driver mutations by their recurrence across samples, as different combinations of mutations in driver pathways are observed in different samples. We introduce the Multi-Dendrix algorithm for the simultaneous identification of multiple driver pathways de novo in somatic mutation data from a cohort of cancer samples. The algorithm relies on two combinatorial properties of mutations in a driver pathway: high coverage and mutual exclusivity. We derive an integer linear program that finds set of mutations exhibiting these properties. We apply Multi-Dendrix to somatic mutations from glioblastoma, breast cancer, and lung cancer samples. Multi-Dendrix identifies sets of mutations in genes that overlap with known pathways - including Rb, p53, PI(3)K, and cell cycle pathways - and also novel sets of mutually exclusive mutations, including mutations in several transcription factors or other genes involved in transcriptional regulation. These sets are discovered directly from mutation data with no prior knowledge of pathways or gene interactions. We show that Multi-Dendrix outperforms other algorithms for identifying combinations of mutations and is also orders of magnitude faster on genome-scale data. Software available at: http://compbio.cs.brown.edu/software. Mark D. M. Leiserson, Dima Blokh, Roded Sharan, Benjamin J. Raphael |
PLoS Comput. Biol. | 1 |
| 2011 | Inferring Mechanisms of Compensation from E-MAP and SGA Data Using Local Search Algorithms for Max Cut
Mark D. M. Leiserson, Diana Tatar, Lenore Cowen, Benjamin Hescott |
RECOMB | 1 |
| 2009 | Evaluating Between-Pathway Models with Expression Data
Benjamin Hescott, Mark D. M. Leiserson, Lenore Cowen, Donna K. Slonim |
RECOMB | 2 |