Dana Pe'er

dblp:37/6966 · DBLP profile ↗
← Back
15ranked-venue papers
2as first author
5since 2021 · last 2025
0000-0002-9259-8817ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Deep learning architectures and training · 36% Generative modeling · 24% Optimization for machine learning · 24%
Interdisciplinary, comprehensive, and emerging computing
9 papers
Bioinformatics and computational biology · 100%
Databases, data mining, and information retrieval
2 papers
Machine learning and data management · 100%
Theoretical computer science
1 paper
Algorithms and data structures · 50% Computational geometry · 50%

Topics — the 30 heaviest of 34, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning and data management
optimal transport
1.622025
Wasserstein Flow Matching: Generative Modeling Over Families of Distributions · ICML 2025
Wasserstein Wormhole: Scalable Optimal Transport Distance with Transformer · ICML 2024
Bioinformatics and computational biology › biological network › network biology › network inference
gene regulatory network inference
1.552024
REUNION: transcription factor binding prediction and regulatory association inference from single-cell multi-omics data · Bioinform. 2024
scKINETICS: inference of regulatory velocity with single-cell transcriptomics data · Bioinform. 2023
MinReg: A Scalable Algorithm for Learning Parsimonious Regulatory Networks in Yeast and Mammals · J. Mach. Learn. Res. 2006
Machine learning › Generative modeling
flow matching
0.912025
Wasserstein Flow Matching: Generative Modeling Over Families of Distributions · ICML 2025
Machine learning › Deep learning architectures and training
autoencoder
0.812024
Wasserstein Wormhole: Scalable Optimal Transport Distance with Transformer · ICML 2024
Machine learning › Deep learning architectures and training › autoencoder
transformer autoencoder
0.812024
Wasserstein Wormhole: Scalable Optimal Transport Distance with Transformer · ICML 2024
Bioinformatics and computational biology › gene regulation
transcription factor binding site prediction
0.812024
REUNION: transcription factor binding prediction and regulatory association inference from single-cell multi-omics data · Bioinform. 2024
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
discrete latent variable model
0.712023
Gradient Estimation for Binary Latent Variables via Gradient Variance Clipping · AAAI 2023
Machine learning › Deep learning architectures and training › neural network training
discrete latent variable training
0.712023
Gradient Estimation for Binary Latent Variables via Gradient Variance Clipping · AAAI 2023
Machine learning › Optimization for machine learning
gradient estimation
0.712023
Gradient Estimation for Binary Latent Variables via Gradient Variance Clipping · AAAI 2023
Machine learning › Optimization for machine learning
variance reduction
0.712023
Gradient Estimation for Binary Latent Variables via Gradient Variance Clipping · AAAI 2023
Machine learning › Generative modeling
variational autoencoder
0.712023
Gradient Estimation for Binary Latent Variables via Gradient Variance Clipping · AAAI 2023
Bioinformatics and computational biology › single-cell analysis
single-cell transcriptomics
0.712023
scKINETICS: inference of regulatory velocity with single-cell transcriptomics data · Bioinform. 2023
Bioinformatics and computational biology › transcriptomics
spatial transcriptomics
0.312025
Wasserstein Flow Matching: Generative Modeling Over Families of Distributions · ICML 2025
Visual content generation and editing
3d shape generation
0.312025
Wasserstein Flow Matching: Generative Modeling Over Families of Distributions · ICML 2025
Bioinformatics and computational biology › genomics
computational genomics
0.212016
Dirichlet Process Mixture Model for Correcting Technical Variation in Single-Cell Gene Expression Data · ICML 2016
Bioinformatics and computational biology › single-cell analysis › cell clustering
single-cell clustering
0.212016
Dirichlet Process Mixture Model for Correcting Technical Variation in Single-Cell Gene Expression Data · ICML 2016
Bioinformatics and computational biology › single-cell analysis › single-cell RNA sequencing
single-cell RNA-seq analysis
0.212016
Dirichlet Process Mixture Model for Correcting Technical Variation in Single-Cell Gene Expression Data · ICML 2016
Bioinformatics and computational biology
single-cell analysis
0.212024
Wasserstein Wormhole: Scalable Optimal Transport Distance with Transformer · ICML 2024
Bioinformatics and computational biology
single-cell biology
0.212024
Wasserstein Wormhole: Scalable Optimal Transport Distance with Transformer · ICML 2024
Computational geometry › graph drawing › geometric embedding
distance-preserving embedding
0.212024
Wasserstein Wormhole: Scalable Optimal Transport Distance with Transformer · ICML 2024
Algorithms and data structures › numerical linear algebra › dimensionality reduction
multidimensional scaling
0.212024
Wasserstein Wormhole: Scalable Optimal Transport Distance with Transformer · ICML 2024
Machine learning › Optimization for machine learning › gradient estimation
stochastic gradient estimation
0.212023
Gradient Estimation for Binary Latent Variables via Gradient Variance Clipping · AAAI 2023
Bioinformatics and computational biology › single-cell analysis
RNA velocity
0.212023
scKINETICS: inference of regulatory velocity with single-cell transcriptomics data · Bioinform. 2023
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model › mixture model
bayesian mixture model
0.112016
Dirichlet Process Mixture Model for Correcting Technical Variation in Single-Cell Gene Expression Data · ICML 2016
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › bayesian nonparametric model
dirichlet process mixture model
0.112016
Dirichlet Process Mixture Model for Correcting Technical Variation in Single-Cell Gene Expression Data · ICML 2016
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models › structure learning › graphical model learning
bayesian network learning
0.112006
MinReg: A Scalable Algorithm for Learning Parsimonious Regulatory Networks in Yeast and Mammals · J. Mach. Learn. Res. 2006
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models
bayesian network
0.112005
Learning Module Networks · J. Mach. Learn. Res. 2005
Machine learning › Probabilistic and Bayesian machine learning › structured models
graphical models
0.112005
Learning Module Networks · J. Mach. Learn. Res. 2005
Machine learning › Deep learning architectures and training
modular network
0.112005
Learning Module Networks · J. Mach. Learn. Res. 2005
Bioinformatics and computational biology
signal transduction
0.012013
Can CAD cure cancer? · DAC 2013

Methods — techniques the papers use, named apart from their topics

optimal transport · 6.5flow matching · 3.5attention mechanism · 3.5transformer · 3.0multidimensional scaling · 3.0sequence feature learning · 0.8semi-supervised learning · 0.8information theory · 0.8gradient variance clipping · 0.7expectation-maximization · 0.7dynamical model · 0.7bitflip-1 · 0.7DisARM · 0.7hierarchical bayesian model · 0.2gibbs sampling · 0.2single-cell resolution experiments · 0.2siRNA silencing · 0.2
YearPublicationVenuePosition
2025 Wasserstein Flow Matching: Generative Modeling Over Families of Distributions
abstract
Generative modeling typically concerns transporting a single source distribution to a target distribution via simple probability flows. However, in fields like computer graphics and single-cell genomics, samples themselves can be viewed as distributions, where standard flow matching ignores their inherent geometry. We propose Wasserstein flow matching (WFM), which lifts flow matching onto families of distributions using the Wasserstein geometry. Notably, WFM is the first algorithm capable of generating distributions in high dimensions, whether represented analytically (as Gaussians) or empirically (as point-clouds). Our theoretical analysis establishes that Wasserstein geodesics constitute proper conditional flows over the space of distributions, making for a valid FM objective. Our algorithm leverages optimal transport theory and the attention mechanism, demonstrating versatility across computational regimes: exploiting closed-form optimal transport paths for Gaussian families, while using entropic estimates on point-clouds for general distributions. WFM successfully generates both 2D \& 3D shapes and high-dimensional cellular microenvironments from spatial transcriptomics data. Code is available at [WassersteinFlowMatching](https://github.com/WassersteinFlowMatching/WassersteinFlowMatching/).
Doron Haviv, Aram-Alexandre Pooladian, Dana Pe'er, Brandon Amos
ICML3
2024 Wasserstein Wormhole: Scalable Optimal Transport Distance with Transformer
abstract
Optimal transport (OT) and the related Wasserstein metric ($W$) are powerful and ubiquitous tools for comparing distributions. However, computing pairwise Wasserstein distances rapidly becomes intractable as cohort size grows. An attractive alternative would be to find an embedding space in which pairwise Euclidean distances map to OT distances, akin to standard multidimensional scaling (MDS). We present Wasserstein Wormhole, a transformer-based autoencoder that embeds empirical distributions into a latent space wherein Euclidean distances approximate OT distances. Extending MDS theory, we show that our objective function implies a bound on the error incurred when embedding non-Euclidean distances. Empirically, distances between Wormhole embeddings closely match Wasserstein distances, enabling linear time computation of OT distances. Along with an encoder that maps distributions to embeddings, Wasserstein Wormhole includes a decoder that maps embeddings back to distributions, allowing for operations in the embedding space to generalize to OT spaces, such as Wasserstein barycenter estimation and OT interpolation. By lending scalability and interpretability to OT approaches, Wasserstein Wormhole unlocks new avenues for data analysis in the fields of computational geometry and single-cell biology.
Doron Haviv, Russell Z. Kunes, Thomas Dougherty, Cassandra Burdziak, Tal Nawy, Anna Gilbert 0001, Dana Pe'er
ICML7
2024 REUNION: transcription factor binding prediction and regulatory association inference from single-cell multi-omics data
abstract
MOTIVATION: Profiling of gene expression and chromatin accessibility by single-cell multi-omics approaches can help to systematically decipher how transcription factors (TFs) regulate target gene expression via cis-region interactions. However, integrating information from different modalities to discover regulatory associations is challenging, in part because motif scanning approaches miss many likely TF binding sites. RESULTS: We develop REUNION, a framework for predicting genome-wide TF binding and cis-region-TF-gene "triplet" regulatory associations using single-cell multi-omics data. The first component of REUNION, Unify, utilizes information theory-inspired complementary score functions that incorporate TF expression, chromatin accessibility, and target gene expression to identify regulatory associations. The second component, Rediscover, takes Unify estimates as input for pseudo semi-supervised learning to predict TF binding in accessible genomic regions that may or may not include detected TF motifs. Rediscover leverages latent chromatin accessibility and sequence feature spaces of the genomic regions, without requiring chromatin immunoprecipitation data for model training. Applied to peripheral blood mononuclear cell data, REUNION outperforms alternative methods in TF binding prediction on average performance. In particular, it recovers missing region-TF associations from regions lacking detected motifs, which circumvents the reliance on motif scanning and facilitates discovery of novel associations involving potential co-binding transcriptional regulators. Newly identified region-TF associations, even in regions lacking a detected motif, improve the prediction of target gene expression in regulatory triplets, and are thus likely to genuinely participate in the regulation. AVAILABILITY AND IMPLEMENTATION: All source code is available at https://github.com/yangymargaret/REUNION.
Dana Pe'er
Bioinform.2
2023 Gradient Estimation for Binary Latent Variables via Gradient Variance Clipping
abstract
Gradient estimation is often necessary for fitting generative models with discrete latent variables, in contexts such as reinforcement learning and variational autoencoder (VAE) training. The DisARM estimator achieves state of the art gradient variance for Bernoulli latent variable models in many contexts. However, DisARM and other estimators have potentially exploding variance near the boundary of the parameter space, where solutions tend to lie. To ameliorate this issue, we propose a new gradient estimator bitflip-1 that is lower variance at the boundaries of the parameter space. As bitflip-1 has complementary properties to existing estimators, we introduce an aggregated estimator, unbiased gradient variance clipping (UGC) that uses either a bitflip-1 or a DisARM gradient update for each coordinate. We theoretically prove that UGC has uniformly lower variance than DisARM. Empirically, we observe that UGC achieves the optimal value of the optimization objectives in toy experiments, discrete VAE training, and in a best subset selection problem.
Russell Z. Kunes, Mingzhang Yin, Max Land, Doron Haviv, Dana Pe'er, Simon Tavaré
AAAI5
2023 scKINETICS: inference of regulatory velocity with single-cell transcriptomics data
abstract
MOTIVATION: Transcriptional dynamics are governed by the action of regulatory proteins and are fundamental to systems ranging from normal development to disease. RNA velocity methods for tracking phenotypic dynamics ignore information on the regulatory drivers of gene expression variability through time. RESULTS: We introduce scKINETICS (Key regulatory Interaction NETwork for Inferring Cell Speed), a dynamical model of gene expression change which is fit with the simultaneous learning of per-cell transcriptional velocities and a governing gene regulatory network. Fitting is accomplished through an expectation-maximization approach designed to learn the impact of each regulator on its target genes, leveraging biologically motivated priors from epigenetic data, gene-gene coexpression, and constraints on cells' future states imposed by the phenotypic manifold. Applying this approach to an acute pancreatitis dataset recapitulates a well-studied axis of acinar-to-ductal transdifferentiation whilst proposing novel regulators of this process, including factors with previously appreciated roles in driving pancreatic tumorigenesis. In benchmarking experiments, we show that scKINETICS successfully extends and improves existing velocity approaches to generate interpretable, mechanistic models of gene regulatory dynamics. AVAILABILITY AND IMPLEMENTATION: All python code and an accompanying Jupyter notebook with demonstrations are available at http://github.com/dpeerlab/scKINETICS.
Cassandra Burdziak, Chujun Julia Zhao, Doron Haviv, Direna Alonso-Curbelo, Scott W. Lowe, Dana Pe'er
Bioinform.6
2016 Dirichlet Process Mixture Model for Correcting Technical Variation in Single-Cell Gene Expression Data
abstract
We introduce an iterative normalization and clustering method for single-cell gene expression data. The emerging technology of single-cell RNA-seq gives access to gene expression measurements for thousands of cells, allowing discovery and characterization of cell types. However, the data is confounded by technical variation emanating from experimental errors and cell type-specific biases. Current approaches perform a global normalization prior to analyzing biological signals, which does not resolve missing data or variation dependent on latent cell types. Our model is formulated as a hierarchical Bayesian mixture model with cell-specific scalings that aid the iterative normalization and clustering of cells, teasing apart technical variation from biological signals. We demonstrate that this approach is superior to global normalization followed by clustering. We show identifiability and weak convergence guarantees of our method and present a scalable Gibbs inference algorithm. This method improves cluster inference in both synthetic and real single-cell data compared with previous methods, and allows easy interpretation and recovery of the underlying structure and cell types.
Sandhya Prabhakaran, Elham Azizi, Ambrose J. Carr, Dana Pe'er
ICML4
2013 Can CAD cure cancer?
abstract
Eukaryotic cells have complex regulatory systems that sense adversity (e.g. DNA damage, heat shock, external death-induction signals) and respond by invoking programmed cell-death or apoptosis. Cancer cells have evolved the ability to thwart such sensory information and associated regulation. As biologists begin to understand cells in circuit-like terms, we can also begin to derive and simulate druggable targets of cellular networks that cause a cancerous cell to kill its self. Towards this goal, we have performed preliminary experiments that test the impact of siRNA (RNA silencing) on diverse cellular signaling pathways, at unprecedented single-cell resolution. We propose ways in which a CAD system can process such data to automatically derive drug targets.
Smita Krishnaswamy, Bernd Bodenmiller, Dana Pe'er
DAC3
2012 Using systems and structure biology tools to dissect cellular phenotypes
abstract
The Center for the Multiscale Analysis of Genetic Networks (MAGNet, http://magnet.c2b2.columbia.edu) was established in 2005, with the mission of providing the biomedical research community with Structural and Systems Biology algorithms and software tools for the dissection of molecular interactions and for the interaction-based elucidation of cellular phenotypes. Over the last 7 years, MAGNet investigators have developed many novel analysis methodologies, which have led to important biological discoveries, including understanding the role of the DNA shape in protein-DNA binding specificity and the discovery of genes causally related to the presentation of malignant phenotypes, including lymphoma, glioma, and melanoma. Software tools implementing these methodologies have been broadly adopted by the research community and are made freely available through geWorkbench, the Center's integrated analysis platform. Additionally, MAGNet has been instrumental in organizing and developing key conferences and meetings focused on the emerging field of systems biology and regulatory genomics, with special focus on cancer-related research.
Aris Floratos, Barry Honig, Dana Pe'er, Andrea Califano
J. Am. Medical Informatics Assoc.3
2010 JISTIC: Identification of Significant Targets in Cancer
abstract
BACKGROUND: Cancer is caused through a multistep process, in which a succession of genetic changes, each conferring a competitive advantage for growth and proliferation, leads to the progressive conversion of normal human cells into malignant cancer cells. Interrogation of cancer genomes holds the promise of understanding this process, thus revolutionizing cancer research and treatment. As datasets measuring copy number aberrations in tumors accumulate, a major challenge has become to distinguish between those mutations that drive the cancer versus those passenger mutations that have no effect. RESULTS: We present JISTIC, a tool for analyzing datasets of genome-wide copy number variation to identify driver aberrations in cancer. JISTIC is an improvement over the widely used GISTIC algorithm. We compared the performance of JISTIC versus GISTIC on a dataset of glioblastoma copy number variation, JISTIC finds 173 significant regions, whereas GISTIC only finds 103 significant regions. Importantly, the additional regions detected by JISTIC are enriched for oncogenes and genes involved in cell-cycle and proliferation. CONCLUSIONS: JISTIC is an easy-to-install platform independent implementation of GISTIC that outperforms the original algorithm detecting more relevant candidate genes and regions. The software and documentation are freely available and can be found at: http://www.c2b2.columbia.edu/danapeerlab/html/software.html.
Felix Sanchez-Garcia, Uri David Akavia, Eyal Mozes, Dana Pe'er
BMC Bioinform.4
2006 MinReg: A Scalable Algorithm for Learning Parsimonious Regulatory Networks in Yeast and Mammals
abstract
In recent years, there has been a growing interest in applying Bayesian networks and their extensions to reconstruct regulatory networks from gene expression data. Since the gene expression domain involves a large number of variables and a limited number of samples, it poses both computational and statistical challenges to Bayesian network learning algorithms. Here we define a constrained family of Bayesian network structures suitable for this domain and devise an efficient search algorithm that utilizes these structural constraints to find high scoring networks from data. Interestingly, under reasonable assumptions on the underlying probability distribution, we can provide performance guarantees on our algorithm. Evaluation on real data from yeast and mouse, demonstrates that our method cannot only reconstruct a high quality model of the yeast regulatory network, but is also the first method to scale to the complexity of mammalian networks and successfully reconstructs a reasonable model over thousands of variables.
Dana Pe'er, Amos Tanay, Aviv Regev
J. Mach. Learn. Res.1
2005 Learning Module Networks
abstract
Methods for learning Bayesian networks can discover dependency structure between observed variables. Although these methods are useful in many applications, they run into computational and statistical problems in domains that involve a large number of variables. In this paper, we consider a solution that is applicable when many variables have similar behavior. We introduce a new class of models, module networks, that explicitly partition the variables into modules, so that the variables in each module share the same parents in the network and the same conditional probability distribution. We define the semantics of module networks, and describe an algorithm that learns the modules' composition and their dependency structure from data. Evaluation on real data in the domains of gene expression and the stock market shows that module networks generalize better than Bayesian networks, and that the learned module network structure reveals regularities that are obscured in learned Bayesian networks.
Eran Segal, Dana Pe'er, Aviv Regev, Daphne Koller, Nir Friedman
J. Mach. Learn. Res.2
2003 Learning Module Networks
Eran Segal, Dana Pe'er, Aviv Regev, Daphne Koller, Nir Friedman
UAI2
2002 Minreg: Inferring an active regulator set
abstract
Regulatory relations between genes are an important component of molecular pathways. Here, we devise a novel global method that uses a set of gene expression profiles to find a small set of relevant active regulators, identify the genes that they regulate, and automatically annotate them. We show that our algorithm is capable of handling a large number of genes in a short time and is robust to a wide range of parameters. We apply our method to a combined dataset of S. cerevisiae expression profiles, and validate the resulting model of regulation by cross-validation and extensive biological analysis of the selected regulators and their derived annotations.
Dana Pe'er, Aviv Regev, Amos Tanay
ISMB1
2000 Using Bayesian networks to analyze expression data
abstract
DNA hybridization arrays simultaneously measure the expression level for thousands of genes. These measurements provide a “snapshot” of transcription levels within the cell. A major challenge in computational biology is to uncover, from such measurements, gene/protein interactions and key biological features of cellular systems.
Nir Friedman, Michal Linial, Iftach Nachman, Dana Pe'er
RECOMB4
1999 Learning Bayesian Network Structure from Massive Datasets: The "Sparse Candidate" Algorithm
Nir Friedman, Iftach Nachman, Dana Pe'er
UAI3